Scale split into configuration, capital, and topology. Opus 4.8 exposed effort and speed as priced controls; Anthropic closed a $65 billion round; OpenAI gated biological work by participant and project; Groq and Apple reportedly rearranged cloud and device infrastructure; and Paris 2.0 distributed video training across independent GPUs.
1. Claude Opus 4.8 adds effort controls while Mythos remains gated
Anthropic released Claude Opus 4.8 at the same regular API price as Opus 4.7: $5 per million input tokens and $25 per million output tokens. It added user-selectable effort, a Claude Code dynamic-workflows preview, and a fast mode that runs 2.5 times faster at $10 and $50 per million tokens.
Anthropic's benchmark and honesty claims remain its own evaluations, and the promised Mythos expansion is conditional on safeguards still under development. Opus 4.8 consequently has no single operating point: effort and fast mode alter both output and price, so a migration result is meaningful only when those settings are fixed.
Sources: Anthropic's Opus 4.8 announcement · Reuters on Opus 4.8 and Mythos
2. Anthropic raises $65 billion at a $965 billion valuation
Anthropic announced a $65 billion Series H at a $965 billion post-money valuation, including $15 billion in previously committed hyperscaler investments. The company says its annualized revenue run rate crossed $47 billion in May and that proceeds will fund compute, safety research, products, and partnerships.
The $47 billion figure annualizes a recent company-reported period; $965 billion is the post-money equity price set by this round. Anthropic's gigawatt-scale agreements across Amazon, Google, and SpaceX turn the financing into an infrastructure story: capital secures the capacity on which model growth depends.
Sources: Anthropic's Series H announcement · Reuters on the financing
3. OpenAI launches restricted access through Rosalind Biodefense
OpenAI launched Rosalind Biodefense to sponsor qualified developers building pandemic-preparedness and biosecurity applications with GPT-Rosalind. It is also extending trusted access to selected U.S. government and allied public-health partners, with initial work spanning DNA screening, epidemiological modeling, early detection, diagnostics, and medical countermeasures.
Rosalind's immediate product is a governed channel into GPT-Rosalind; public-health impact has yet to be measured. Credibility rests on whether approved projects produce expert-reviewed defensive results without allowing the same access to drift into higher-risk biological work.
Sources: OpenAI's Rosalind Biodefense announcement
4. Groq reportedly seeks $650 million for an inference-cloud pivot
Groq is reportedly raising up to $650 million from existing investors as it shifts toward an AI inference neocloud built on its own chips and systems. Axios says Disruptive and Infinitum will backstop the round after a reported $20 billion Nvidia licensing transaction that moved much of Groq's senior team.
Groq has announced neither a completed raise nor the capacity of the proposed cloud. The Nvidia transaction makes continuity the decisive question: investors are financing a service whose senior team and core hardware economics may already be changing.
Sources: Axios on Groq's reported raise · TechCrunch's follow-up
5. Apple reportedly splits Gemini-powered Siri across device and cloud
Apple is reportedly distilling Google's large Gemini models for some on-device Siri tasks while preparing cloud execution for more complex requests. Ars Technica, citing The Information, says Apple has also explored Nvidia Confidential Computing after encountering difficulty running the full models on its own Private Cloud Compute infrastructure.
The three companies have not confirmed this architecture, so it remains a report about Apple's design work. If shipped, the routing decision becomes the privacy promise: users can judge hybrid Siri only if the interface reveals when a request leaves the phone and which cloud receives it.
Sources: Ars Technica on Apple's reported Gemini architecture
6. Paris 2.0 tests video diffusion across decentralized GPUs
The Paris 2.0 preprint describes a text-to-video diffusion model trained through decentralized computation rather than one monolithic cluster. Under matched data and total compute at low resolution, the authors report a Frechet Video Distance of 279.01 versus 561.04 for their monolithic baseline, alongside higher text-video similarity and aesthetic scores.
The result applies to one matched, low-resolution experiment in a six-page preprint. Its significance is narrower and useful: decentralized training has cleared a video-coherence test under equal compute, shifting the next question to whether that advantage survives production resolution and unreliable workers.