Morning, {{first name | folks}}! A lot happened in AI this week. We’ve got an AI beating 676 human forecasters, a Gemini security test going very wrong, $1 billion going into independent AI evaluation, China cooling its humanoid robot IPO rush, and NVIDIA tackling AI inference benchmarking.
Today’s Top 5
AI Beats 676 Human Forecasters: Mantic beat every human and AI competitor in this summer’s Metaculus Cup and raised $25 million to expand its forecasting platform.
Gemini Accidentally Hacked Real Companies: Gemini broke into three real companies during a security test after confusing them with a fictional target.
Anthropic Puts $1 Billion Into AI Safety: Anthropic and Accenture each plan to invest at least $1 billion over five years in independent evaluation of Anthropic’s AI systems.
China Slows the Humanoid Robot IPO Rush: Chinese regulators are scrutinizing humanoid robot valuations and whether growing demand reflects real commercial deployments.
NVIDIA Launches AIPerf for AI Inference: AIPerf lets teams test LLM inference under high-concurrency workloads and replay production traffic to measure real-world performance.
Mantic beat every human participant in this summer's Metaculus Cup, a serious forecasting tournament, including professional forecasters, something experts predicted was still a decade away as recently as last year. It called Colombia's presidential frontrunner a week before humans did, and it beat every other AI system in the competition too, not just humans, meaning the edge came from Mantic's own tuning, not just a strong base model.
It's already being used by hedge funds and Fortune 500 companies for things like M&A timing and macro predictions. Today it raised $25M, led by Radical Ventures, with Microsoft's own venture arm and Hugging Face's co-founder both backing it.

Google confirmed Gemini broke into three real companies' systems during a May security test, and the cause is almost absurd, the fictional company it was supposed to be testing against happened to share a name with a real one, and Gemini went after the real target instead. It guessed its way into one system through password attempts and found exposed credentials in a public repo for the other two. Internet access wasn't even supposed to be available during the test at all.
Google's model did stop itself the moment it realized it had hit real companies, not the fictional target. The bigger picture matters more than this one incident though, the same evaluator says the exact same testing-procedure gaps already caused similar incidents at OpenAI, Anthropic, and Meta. This isn't one lab's mistake anymore, it's a pattern across the entire industry's safety testing process.

Anthropic is partnering with Accenture to embed independent evaluators inside the company, real follow-through on Dario Amodei's "We Must Pace the Frontier" essay from earlier this month. Both companies expect to invest at least $1 billion each over five years. These evaluators get employee-level access, watching models take shape during actual training, sitting in on internal decisions, talking to Anthropic staff directly, not just testing a finished product from the outside.
Anthropic's clear that this doesn't let them off the hook, they still own the safety of their own models. This just makes their claims independently checkable. There's no industry standard yet for how any of this should actually work, Anthropic says they're building it as they go, and more evaluator partnerships are coming in the next few weeks.

China is putting the brakes on a wave of humanoid robot IPOs as regulators question whether soaring valuations and government-backed projects reflect real commercial demand. The scrutiny intensified after Unitree Robotics surged more than fivefold on its Shanghai debut, then fell 55% from its peak.
Several Chinese robotics companies are preparing to go public, but investors are now being asked to look harder at factory deployments, order volumes, and recurring demand. Some private-market valuations have already been cut by 30% to 50%, showing how quickly the market is moving from excitement around humanoid robots to questions about which companies can turn the technology into a real business.

NVIDIA has introduced AIPerf, an open-source benchmarking tool for testing LLM inference under heavy production workloads. It is designed to measure latency, throughput, GPU usage and other performance metrics across different models, serving systems, and hardware setups. It can also generate high-concurrency traffic without the benchmark itself becoming the limiting factor.
AIPerf can replay production traces and simulate different traffic patterns, including bursty and concurrent workloads. That gives AI teams a way to test whether their inference stack can actually handle the traffic they expect, rather than relying on performance numbers from simple benchmark runs.
Other AI Signals:
Qwen released Qwen-Image-2.1, a compact 7B open image model that can natively generate transparent-background images and edit using up to 10 reference photos while keeping a person's or product's identity consistent. Small, efficient, and free to self-host, the same efficiency theme is running through AI releases lately.
Meta Recovers From a US Outage. Facebook and Instagram briefly went down for thousands of users before service largely returned the same evening.
xAI Releases Grok Voice Transcribe 2.0. The new speech-to-text model supports real-time and batch transcription, speaker identification, timestamps, and dozens of languages.
Nscale Files for a US IPO. The AI cloud company has filed to list on the NYSE under the ticker NSCL, although the size and price of the offering have not been set.
TSMC Expands Advanced Chip Packaging in Taiwan. A new 88.7-hectare park in Kaohsiung will include a TSMC technology lab and training center for advanced packaging.
AI Tools to Try:
Naoma: AI agent that engages website visitors and turns interested prospects into meetings.
Voiskey: Voice typing that lets you dictate naturally across the apps you already use.
Kilo Code: Run coding agents from your phone and keep working on projects away from your computer.
Figr: Generates product screens using your existing design system, user flows, and product context.



