This week made verification the scarce asset. A machine-generated geometry proof survived specialist review while BenchJack broke benchmark graders. Google bundled agents, Search interfaces, and media generation into one platform just as Sora's closure disrupted a film. Anthropic's multibillion-dollar compute supply accelerated cyber discovery into a much smaller human patch pipeline.
1. Proof review and benchmark hardening solve opposite verification problems
An internal OpenAI model produced a counterexample in the planar unit-distance problem, and external mathematicians checked the proof and published explanatory remarks. BenchJack attacked the inverse case: agents received benchmark environments and found 219 ways to earn rewards without completing the intended tasks.
Mathematical review asks whether a surprising artifact is correct; benchmark hardening asks whether a score implies the claimed behavior. Both place an independent verifier after machine search. The difference is adversarial: a proof checker tests one artifact, while an agent benchmark must assume the system may actively exploit its grader.
Sources: OpenAI's geometry result · External mathematicians' remarks · BenchJack preprint, version 1
2. Google's integrated agent stack trades implementation speed for continuity risk
Google spread the Antigravity harness across hosted Managed Agents, generated Search mini-apps, and conversational video in Gemini Omni. In the same week, an OpenAI-backed animated film replaced Sora after the consumer service closed during production.
The contrast separates convenience from durability. A provider bundle can supply models, sandboxes, state, interfaces, and distribution with little custom orchestration; it can also change several dependencies at once. Exportable assets, versioned skills, retained prompts, trace access, API deprecation terms, and tested substitutes define the actual exit path.
Sources: Google's Managed Agents announcement · Google I/O keynote · Google's Gemini Omni launch · The Next Web on Critterz replacing Sora
3. Compute scales discovery faster than the patch pipeline
SpaceX's filing disclosed an Anthropic commitment reaching $1.25 billion per month for GPU capacity through May 2029. Project Glasswing partners meanwhile reported more than 10,000 high- or critical-severity findings, while Anthropic's open-source dashboard showed 1,596 disclosures but only 97 upstream patches by May 22.
The financial input and security output are both large; the completed remediation count is not. More compute can multiply candidate generation, but reproduction, prioritization, maintainer communication, patch creation, and downstream deployment remain serial human and institutional steps. Patch latency and installed-fix coverage therefore measure defensive value better than candidates per GPU.
Sources: WIRED on Anthropic's SpaceX contract · Anthropic's Glasswing update · Anthropic's disclosure dashboard