Morning, folks! AI agents are starting to touch the real world, a benchmark finally got a proper blind test, Z.ai pulled a sneaky model launch, and humanoid robots are finding out that demos are easier than real work. Let’s get into it.
Today's Top 5
AI Agents Just Started Running Physics Labs Better Than Humans Do: Anthropic is letting agents control real lab equipment. At QuEra, one tuned lasers overnight and took qubit stability from 58% to 99.3%.
Google Just Made It Impossible to Cheat on an AI Benchmark: Google tested Gemini without letting either side see the benchmark data. No shared questions, no shared model weights, no trusting everyone to play fair.
Z.ai Hid Its New Model in Plain Sight: Z.ai quietly released its new model as “Ox Alpha,” let everyone test it, and then revealed it was theirs. It’s now open-source, and Z.ai says it gets close to Opus-level coding performance at a fraction of the price.
Marvell’s Becoming More Dependent on the AI Boom: Marvell’s data-center revenue jumped 46% to $2.17B. That’s now 79% of its business, so the AI buildout is becoming a very big bet for the company.
China’s Humanoid Robot Boom Just Hit a Reality Check: China could ship 100,000 humanoids this year, but some struggled with basic tasks at a robot competition. Making the robot is one problem. Getting it to do useful work is another.
Anthropic opened a research preview letting AI agents directly run lab equipment, microscopes, liquid handlers, robotic arms, and stuff that normally takes weeks to wire together and doesn't talk to itself at all. At QuEra, a quantum computing company, an AI agent spent one night tuning the lasers that keep their qubits stable. Success rate went from 58% to 99.3%. Recovery time dropped from 150 seconds to 6.
This isn't staying in the lab either. AWS, Danaher, Tecan, Universal Robots, Hugging Face, and Raspberry Pi are already building support for it. Anthropic's still holding back the full open-source release, though. Claude's physical reasoning still has real gaps, researchers actually had to explain to it that foaming samples were a physics problem, not a software bug.

Google DeepMind ran the world's first double-blind evaluation of a frontier AI model, testing Gemini Flash Lite against confidential benchmarks where neither side could see the other's data. DeepMind never saw the test questions. The evaluators, including Singapore's AI Safety Institute and MLCommons, never saw the model weights. It's built on Google Cloud's confidential computing, cryptographically enforced, not just a pinky promise.
The problem it's solving is bigger than it sounds. Every benchmark score you've ever read, including every one this newsletter has covered, relies on trusting a model hasn't somehow seen its own exam in advance. This is the first real attempt to make that trust unnecessary instead of just assumed.

Z.ai quietly tested GLM-5.3-Flash as “Ox Alpha,” a free mystery model that shot to the top of OpenRouter and OpenCode before developers figured out it was theirs. Z.ai confirmed it on August 26 and released the 320B-parameter model under an MIT license, with 18B active parameters and weights on Hugging Face.
The pricing is the real story, though. Z.ai says it's getting close to Claude Opus 4.8 on coding benchmarks at roughly a tenth of the price, while Ox Alpha was served entirely on Chinese AI chips. The benchmark numbers are still mostly Z.ai's own, so independent testing will matter, but cheap, open, frontier-level coding on Chinese hardware is a combination worth watching.

Marvell just posted another record quarter, with revenue up 37% from a year ago. But the number that caught my attention was $2.17 billion from data centers, up 46%.
That's now 79% of the company's revenue. Marvell expects another record quarter next and raised its outlook again, with AI-related bookings still strong. For a company that makes the chips and infrastructure behind data centers, that's a great place to be right now. It's also becoming a pretty concentrated bet on one very expensive AI buildout.

China could ship around 100,000 humanoid robots this year, but the bigger question is getting harder to ignore: can they actually do useful work? At Beijing's World Humanoid Robot Games, some robots struggled with fairly basic tasks, while Unitree's valuation has already fallen nearly 30% from its post-IPO peak.
The interesting part is that building a robot that can run or fight is one thing. Getting it to reliably handle messy factory work is another. That's where the real commercial opportunity is, and even Unitree says useful applications are still some way off. The robot boom is moving from impressive demos to a much tougher test: whether anyone actually needs them.
Other AI Signals:
OpenAI opened commercial operations in Brazil, now one of ChatGPT's three largest markets. Brazilians send about 215 million messages a day. Weekly Codex users grew 11x since January and daily interactions nearly 30x. An OpenAI-funded study estimates AI could add nearly R$1 trillion to Brazil's economy by 2030.
Anthropic's opening 10,000 free or discounted Claude seats for scientists worldwide, plus up to $50K in project credits through its AI for Science program. The more interesting detail: it's enrolled its first participants in a controlled program letting life science researchers access Mythos-class models, the ones normally locked down over dual-use risk, work directly with the US government to manage it.
Google released Gemini Omni 1.1 Flash, adding scene extension (up to 40 seconds), precise first/last-frame control, and 4K upscaling for AI-generated video. It's already live inside Adobe Firefly, Figma Weave, and Runway, real production use, not just a demo.
Jensen Huang pushed back on the “circular financing” criticism around Nvidia's growing investments in AI companies, arguing the risk is low because the GPUs can simply be redeployed if a customer fails. The debate is getting harder to ignore as Nvidia increasingly funds the same ecosystem buying its chips.
Nvidia started shipping Vera, its first CPU built for AI agents. AWS has received the first servers, with OpenAI, Anthropic, Oracle, and SpaceX AI also testing them. Oracle plans to deploy hundreds of thousands, giving Nvidia a bigger role in AI infrastructure beyond GPUs.
Today’s AI Tools to Try:
Putty: Build websites and simple tools with AI, with a collaborative workflow for experimenting and iterating on ideas.
Akta: Get real-time private-company data for deal sourcing, lead research, and outbound workflows.
Verse: Create autonomous AI workers by describing the role or task you want them to handle.
Bloxks App: Turn an app idea into a working mobile product without writing code.



