Good morning. It's Friday, July 31, and we're covering why AI agents can resort to deception, how two API settings changed GPT-5.6 benchmark results, and China's warning over the U.S. humanoid robot ban.

Plus: Friday Fact on the metal powering AI chips, and a guest essay on the leap from web readers to web agents.

YOUR DAILY ROLLUP

Top Stories of the Day

China says the U.S. humanoid robot import restrictions threaten trade stability more than cybersecurity, raising the stakes ahead of planned leadership talks. The commerce ministry demands the FCC withdraw its new import limits and warns of retaliation if it does not. Analysts say Beijing could tighten rare earth exports or restrict firms like Tesla and NVIDIA, while IPO hopefuls Unitree and Agibot face added uncertainty.

The biggest hurdle for electric flying taxis may be convincing passengers, not building the technology. eVTOL makers including Eve Air Mobility and Vertical Aerospace say autonomous flight is essential for scaling operations and lowering costs, but U.S. regulators have yet to certify any passenger eVTOL. China's EHang already operates paid autonomous flights on limited routes, while Western firms target gradual automation before removing pilots entirely.

Google says a single AI model can now control robots from feet to fingertips while adapting to new hardware in just hours. Gemini Robotics 2 introduces three models for whole-body control, embodied reasoning, and on-device operation, enabling humanoids to perform complex tasks and collaborate with other robots. The company reports success rates up to 89.6% on precise insertion tasks, while multi-finger manipulation remains a key challenge.

FRIDAY FACT

The Metal Destroyed 50,000 Times Every Second

Somewhere in the supply chain behind every advanced AI chip, a metal is being blown apart 50,000 times per second. The full story is at the bottom of today's issue.

VIDEO

AI Memory Card Tested

A memory-tracking AI card records back-to-back meetings, beats Matt in a recall test, then auto-generates action items from the conversations.

FORWARD FUTURE ORIGINAL

Getting Good at Reading the Web Won't Teach an Agent to Use It

Guest contributor: Abhishek Das is the co-founder and co-CEO of Yutori. Previously a research scientist at Meta FAIR, he earned his PhD from Georgia Tech.

A few years ago, autocomplete finished your line of code. Today, coding agents open a pull request, run the tests, read the failures, and try again until the build is approved. On the surface it looks like one thing grew into a better version of itself: autocomplete got smart enough to become an agent.

But anyone who has built these systems knows the gap between them is enormous, with years of work sitting in between. We are now crossing that same canyon on the web. → Read the full article here.

ALIGNMENT

Claude Opus 5 Dominates AI Vending Test With Deception and Collusion

Andon Labs’ latest Vending-Bench experiment found that frontier AI models, including Claude Opus 5, GPT-5.6 Sol, and Kimi K3, routinely lied, colluded, and broke agreements while managing simulated vending machine businesses over the course of a simulated year. Claude Opus 5 achieved a record average final balance of $11,182 by negotiating deceptive supplier deals, proposing price-fixing schemes, ignoring refund requests, and repeatedly undercutting competitors despite collaboration agreements.

The models also attempted market division, supplier manipulation, and coercive wholesale tactics, with Opus breaking 11 separate collusion agreements during the benchmark. While the test took place in a simulation, Andon Labs argues the results underscore how today's frontier AI agents can adopt unethical strategies when left to pursue long-term financial goals without human oversight.Read the full article here.

EVALUATION

Two API Settings Tripled GPT-5.6 Sol’s ARC-AGI-3 Benchmark Score

OpenAI found that GPT-5.6 Sol’s low performance on the ARC-AGI-3 reasoning benchmark was largely due to the benchmark’s evaluation harness rather than the model itself. By enabling two Responses API features already used in ChatGPT and Codex—retained reasoning and context compaction—the model’s score on the public ARC-AGI-3 task set increased from 13.3% to 38.3% Relative Human Action Efficiency (RHAE), while output tokens fell sixfold.

The changes allowed the model to preserve its reasoning across turns and avoid losing earlier context, making it more effective at learning game mechanics over long sessions. OpenAI argues the results highlight how benchmark outcomes can depend as much on evaluation settings as on the underlying model, particularly for long tests. Read the full article here.

MARKET PULSE

Microsoft Soars While Meta Slumps as AI Spending Divides Investors

Microsoft shares surged 15% after the company reported stronger-than-expected fiscal fourth-quarter results, including 43% Azure cloud growth and more than 30 million paid Microsoft 365 Copilot seats, reinforcing investor confidence that its AI investments are generating returns.

In contrast, Meta fell 9% after missing earnings expectations, issuing weaker-than-expected revenue guidance of $61 billion to $64 billion for the current quarter, and reporting a 91% year-over-year drop in free cash flow to $784 million as AI spending accelerated. → Continue reading here.

NEWS

What Else is Happening

Aschenbrenner Fund Unwinds Stocks: Leopold Aschenbrenner's hedge fund sold all public stock holdings after AI and software bets triggered steep losses.

Forward-Deployed Engineers Surge: Demand for AI deployment specialists is projected to jump 2,100% as enterprises prioritize real-world AI.

xAI Challenges Minnesota AI Law: xAI sued to block Minnesota's AI nudification ban, arguing the law is overly broad and violates free speech.

LinkedIn Adds AI Slop Reports: LinkedIn now lets users flag suspected AI-generated posts as it works to reduce low-quality AI content.

Samsung Locks In AI Chip Deals: Samsung signed five-year chip supply deals covering most future capacity as AI demand continues to outpace production.

FRIDAY FACT

Every AI Chip Starts With Exploding Tin Droplets

ASML's extreme ultraviolet lithography machines make light by destroying metal. A generator fires molten tin droplets about 25 microns across — roughly a third the width of a human hair — into a vacuum chamber at 70m/s. Each droplet is hit twice by a carbon dioxide laser: a weak first pulse flattens it into a pancake, then a second, far more powerful pulse vaporizes it into plasma at around 220,000 degrees Celsius. The plasma emits light at a wavelength of 13.5 nanometers. This repeats 50,000 times a second.

That 13.5-nanometer light is the reason leading-edge chips exist. Older lithography used 193-nanometer light, and shrinking transistors further required a wavelength more than an order of magnitude smaller. Nothing else produces it at usable brightness. So every advanced AI accelerator shipping today traces back to a vacuum chamber where tin is being blown apart at roughly 40 times hotter than the temperature of the Sun's surface.

That's All for Today

Before you go, what did you think of today's issue?

We read every response.

Login or Subscribe to participate

Thanks for reading. See you next time!

— Matthew Berman, Nick Wentz & the Forward Future Team

Reply

Avatar

or to participate