GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Pentagon Tests OpenAI, Google, and Grok to Replace Claude — $200M Contract in Dispute

The Pentagon has been running competitive AI model evaluations since March 1, 2026 — three days after Defense Secretary Pete Hegseth designated Anthropic a supply-chain risk and initiated the process of removing Claude from classified military networks.

According to Bloomberg, 25 designated military personnel classified as “power users” are testing models from OpenAI, Google, and xAI’s Grok on GenAI.mil, a platform that runs separately from the Pentagon’s Maven Smart System. Emil Michael, the undersecretary of defense for research and engineering, told Bloomberg Television that talks with Anthropic are suspended while the company’s legal challenge proceeds.

How the contract collapsed

In July 2025, Anthropic signed a two-year, $200 million contract to integrate Claude into classified defense networks. Claude was deployed on Maven Smart System, the digital mission control platform used for classified operations. By early 2026, Claude had reportedly been used in intelligence analysis and operational planning related to Iran.

The falling out was structural: Anthropic refused to remove restrictions on use of its models for mass surveillance and lethal autonomous weapons systems. When Hegseth issued an ultimatum in late February, Anthropic held its position. On February 27, he designated the company a supply-chain risk, triggering the phased replacement.

Contractors who had integrated Claude into defence systems were given a six-month window to find alternatives.

What the alternatives offer

OpenAI has renegotiated its agreement with the DoD explicitly to allow unrestricted lawful use — a direct contrast to Anthropic’s refusal. The models under evaluation include GPT-5.5 and Codex for coding and reasoning tasks, Google Gemini for sensor integration and robotics-adjacent workflows, and xAI’s Grok for real-time data processing.

The Pentagon’s CTO framed the Anthropic situation as a catalyst for deliberate diversification. Michael: “One of the problems, without getting into the specifics about Anthropic, is that we had one primary provider on classified networks. That doesn’t work for the Department of War.”

Early test results suggest the models respond differently to identical queries. The DoD is exploring whether making final evaluation results public would be possible — which would be an unusual window into how defence agencies actually benchmark AI tools.

The NSA contradiction

The evaluation is occurring alongside an active internal contradiction: the NSA was testing Anthropic’s Claude Mythos Preview on its own classified networks as recently as April 2026, even as the broader Pentagon was phasing out Claude.

The agencies are separately structured, but the schism is visible. The same Mythos model the Pentagon is trying to replace is being considered by the NSA for offensive and defensive cyber operations — which sit closer to the capabilities Mythos demonstrably has.

The Trump administration is simultaneously defending the Anthropic supply-chain risk designation in federal court and exploring whether it can access Mythos for cyber defence through a separate channel. Anthropic is contesting the designation, which it says could cost it billions in revenue.

The precedent for AI defence contracts

The Pentagon’s publicly stated conclusion from this episode is that single-provider dependency is a procurement failure, not a vendor problem. The six-month transition timeline it set is now a baseline expectation: any AI vendor operating in classified environments should assume its contract can be replaced within two quarters if the relationship breaks down.

API compatibility and standardised interfaces have gone from nice-to-have to procurement requirements. The DoD is treating model substitutability the way it treats spare parts.

Key numbers

  • $200M: Anthropic’s two-year DoD contract value, now in dispute
  • March 1, 2026: day competitive testing began, 3 days after Hegseth’s designation
  • 25: military “power users” running current evaluations
  • 6 months: contractor transition window to find Claude alternatives
  • Models in evaluation: GPT-5.5/Codex, Google Gemini, xAI Grok