daily

AI Adjacent Daily Briefing – March 3, 2026

March 3, 2026

OpenAI narrows its Pentagon surveillance terms, Anthropic approaches a $20B revenue run rate, and ExpGuard tests specialist moderation.

Overview

OpenAI narrowed the surveillance language in its Pentagon agreement after public criticism, while Anthropic's reported revenue approached a scale that makes the government dispute commercially consequential. ExpGuard research addressed a separate control failure: general moderation models can miss harmful requests expressed in specialist language.

Developments

1. OpenAI expressly narrows domestic surveillance in its defense pact

OpenAI amended its Pentagon agreement to prohibit intentional domestic surveillance of U.S. persons and said intelligence-agency use would be covered by a separate agreement. The revision followed criticism that the original language left too much ambiguity around domestic surveillance.

The amendment narrows one contractual boundary while creating another negotiation point for intelligence work. OpenAI now has to distinguish defense users, intelligence users, and prohibited domestic targeting inside a shared technical stack, with any enforcement gap threatening both the contract and its claim to retain safety control.

Sources: OpenAI agreement · Reuters

2. Anthropic approaches a reported $20B revenue run rate

Anthropic was approaching a $20 billion annualized revenue run rate during its Pentagon confrontation, Bloomberg reported on March 3. The figure places the company among the fastest-growing software suppliers even as a federal designation threatens defense and contractor distribution.

That scale changes the leverage on both sides. The government can close a valuable procurement channel without threatening Anthropic's immediate survival, while Anthropic can absorb a defense setback more readily than a smaller supplier; contractor spillover remains the larger risk because it can reach commercial revenue outside direct federal sales.

Sources: Bloomberg

3. ExpGuard tests moderation against specialist harmful prompts

ExpGuard introduces a moderation model and 58,928-example dataset spanning financial, medical, and legal prompts with domain-expert annotations. In the March 2 preprint, the authors report gains over WildGuard of up to 8.9 percentage points for prompt classification and 15.3 points for response classification on their specialist test set.

Specialist vocabulary creates an evasion surface because a dangerous request can resemble legitimate professional discussion to a general guardrail. ExpGuard's advantage is therefore concentrated in the three selected sectors; extension to another regulated domain brings a fresh labeling burden and a new boundary between expert assistance and harmful instruction.

Sources: ExpGuard preprint