FLI Summer 2026 Safety Index: No AI Lab Earns Above C+ — xAI, DeepSeek, Mistral All Fail
The Future of Life Institute released its Summer 2026 AI Safety Index on July 7. Seven independent reviewers graded nine frontier AI companies across 37 indicators in six categories: risk assessment, current harms, safety frameworks, existential safety, governance and accountability, and information sharing.
No company received an A in any single category.
The Rankings
| Company | Grade | Change |
|---|---|---|
| Anthropic | C+ (2.66) | Held #1 |
| OpenAI | C | Fell from C+ |
| Google DeepMind | C | Held #3 |
| Meta | D+ | Rose from D (4th from 6th) |
| Z.ai | D- | Debut |
| Alibaba Cloud | D- | Debut |
| SpaceXAI (formerly xAI) | F | Fell from 4th to 7th |
| DeepSeek | F | — |
| Mistral | F | — |
Anthropic led across five of six domains. OpenAI led Risk Assessment. xAI, DeepSeek, and Mistral did not respond to the institute’s survey, which the panel treated as material to their governance scores.
What the Panel Actually Found
The sharpest finding is not the individual grades but the direction of movement. The panel concluded that Anthropic, OpenAI, Google DeepMind, and Meta have all weakened or eliminated earlier commitments to pause development if their models approached specified danger thresholds — what the report calls “moving the goalposts.”
Anthropic’s case is specific: in February, the company dropped a pledge it had held since 2022 to never train a system unless it could guarantee in advance that its safety measures were adequate. The FLI panel recommends reversing the change. Anthropic has not indicated it will.
Existential safety was the weakest category across all nine companies. Every firm received a failing grade in that domain.
xAI’s Fall
xAI’s collapse from fourth to seventh is the most dramatic movement in the index’s history. Factors cited: no published system card for Grok 4.5, no red-teaming framework comparable to Anthropic’s Responsible Scaling Policy, a $530 million legal reserve disclosed in the SpaceX IPO prospectus, and documented instances of Grok generating content that induced user delusions. The adult content controversy and the Musk court admissions about training on OpenAI model outputs contributed to the panel’s information-sharing score.
What the Score of C+ Actually Means
The C+ is not a passing grade. The FLI methodology sets a theoretical 4.0 scale; Anthropic’s 2.66 sits in the lower end of the C band on any standard academic curve. The institute is explicit that a C+ means meaningful safety work is happening and that meaningful gaps remain — not that the company is operating safely by any external standard.
The FLI has no enforcement authority. Companies do not have to accept the findings, respond to the survey, or fix the weaknesses the panel identifies. This is the third edition of the index; in prior editions, the grades have had limited observable impact on company behavior. Whether the Summer 2026 results generate any policy traction depends on whether regulators treat them as credible input — which, under the current US framework, is not guaranteed.
Three Labs, Three Continents
FLI president Max Tegmark noted that the three failing grades come from the United States (SpaceXAI), China (DeepSeek), and Europe (Mistral). The geographic spread is deliberate rhetorical framing: inadequate safety is not a China problem or a European startup problem. It is a frontier AI problem, and the failing labs are in the same economic tier as the labs passing with C’s.
The full index is available at futureoflife.org.