GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

Special Report: The AI Index 2026 — Signals From a Frontier That Refuses to Slow Down

The frontier kept moving in 2025. The Stanford AI Index 2026, the ninth edition of the field’s most exhaustive scorecard, puts numbers on how far. Generative AI reached roughly 53% population adoption within three years, faster than either the PC or the internet. Performance on SWE-bench Verified climbed from 60% to near 100% of the human baseline in a single year. The US–China gap between leading models collapsed to 2.7% as of March 2026. US private AI investment hit $285.9 billion, more than 23 times China’s. And the hardware stack that makes all of it possible still runs through a single foundry in Taiwan.

The Index is, in a sense, the sector’s annual 10-K. Nearly 400 pages across nine chapters, assembled by the Human-Centered AI Institute at Stanford, built from benchmark data, regulatory filings, investment records, labour statistics, public opinion surveys and the kind of patient accounting that independent market data tends to get wrong. It is the document against which most other “state of AI” narratives should be stress-tested this year. Here are the signals that matter.

Capability: Acceleration, Not Plateau

The headline signal is that capability is not flattening.

The Index’s coverage of technical performance is blunt: SWE-bench Verified, the coding benchmark most correlated with real-world developer productivity, went from 60% to near 100% of the human baseline in twelve months. On OSWorld, a computer-use benchmark for agents, accuracy jumped from 12% to 66.3%. MLE-bench, which measures machine-learning engineering ability, went from 17% to 64.4%. Cybench, which tests unguided cybersecurity tasks, rose from 15% to 93%.

“On a key coding benchmark—SWE-bench Verified—performance rose from 60% to near 100% of meeting the human baseline in a single year.”

The math story is the same shape. Gemini Deep Think earned a gold medal at the International Mathematical Olympiad. On FrontierMath Tier 4, a benchmark of research-grade problems, accuracy reached 31.3%. On MathArena it hit 96.97%. Several frontier systems now meet or exceed human baselines on PhD-level science questions, competition mathematics, and multimodal reasoning.

Adoption ran alongside capability. Organisational adoption rose to 88%. Generative AI is regularly used by 79% of surveyed organisations. Among university students globally, 80% now use generative AI for schoolwork, up from 40% in 2023. In the United States, over 80% of high school and college students use AI for school-related tasks. Roughly half of US adults were familiar with AI-related news by 2025, against 26% in 2022.

Speed of uptake is the market-defining number. Per the Index, generative AI reached 53% of the population within three years, with Singapore leading at 61% and the United Arab Emirates at 54%. The United States ranks 24th at 28.3%.

The important analytical caveat is that capability is uneven, not monotonic. The same systems that solve Olympiad problems read analogue clocks correctly only 50.1% of the time. Robots succeed in only 12% of real household tasks, even as simulated robotic manipulation on RLBench reaches 89.4%. AI agents have crossed 66% success on OSWorld, but AI agent deployment remains in single digits across nearly all business functions. Treat the frontier as jagged, not smooth.

Convergence: The US–China Gap Effectively Closed

The second signal is that the national-leaderboard trade has largely played out.

As of March 2026 the top US model led its Chinese counterpart by just 2.7%, or 39 Elo points on the Arena Leaderboard. The top closed-weight model led the top open-weight model by 3.4%, 49 points. US and Chinese systems have traded the lead multiple times since early 2025. In February 2025, DeepSeek-R1 briefly matched the top US model outright.

“As of March 2026 Anthropic’s top model leads by just 2.7%.”

Two things are true at once. The United States still produces more top-tier models: 1,618 cumulative notable models between 2018 and 2025, against 849 from China. The Index also shows that the United States released 50 notable models in 2025, China released 30, and South Korea released 5. On the other hand, China leads in publication volume, citations and patent grants; South Korea leads the world in AI patents per capita; and China’s share of the top 100 most-cited AI papers rose from 33 in 2021 to 41 in 2024.

The moat argument has to be rewritten accordingly. “Best model” is now a variable that changes quarter to quarter, not a durable franchise. That reframes the strategic question from who can build the best model to who can sustain compute, distribution, talent, and post-training data pipelines at frontier scale.

Concentration Risk: One Foundry, One Grid, One Chip Vendor

The third signal is that the physical substrate of the frontier is more concentrated, not less.

Per the Index, a single company, TSMC, fabricates almost every leading AI chip. A TSMC-US expansion began operations in 2025, but the global AI hardware supply chain still routes through one foundry in Taiwan. This is the Taiwan exposure trade every serious AI portfolio has to price.

Above the chip layer, concentration compounds. Nvidia accounts for over 60% of total compute, with Google and Amazon supplying much of the remainder and Huawei holding a small but growing share. Global AI compute capacity reached 17.1 million H100-equivalents in 2025, having grown roughly 3.3 times per year since 2022.

The power layer tells the same story. The United States hosts 5,427 data centres, more than ten times any other country, and consumes more energy than any other region. AI data centre power capacity reached 29.6 GW by the end of 2025, comparable to New York state at peak demand. AI chip power accounts for 11.8 GW of that. Grok 4’s estimated training emissions came in at 72,816 tons of CO₂ equivalent. Annual GPT-4o inference water use alone may exceed the drinking water needs of 12 million people.

Industry now produces over 90% of notable frontier models. In 2025 alone, Epoch AI identified one notable AI model originating from academia, compared to 87 from industry. Across 2003 to 2025, industry accounts for 91.58% of notable models, industry–academia collaboration 5.26%, academia 1.05%.

Call it what it is. The frontier is industrial, private, geographically thin, and dependent on a small number of chokepoints: one foundry, one dominant compute vendor, one country’s grid, and a handful of labs.

Responsible AI Is Not Keeping Up

The fourth signal is that governance and safety are lagging capability.

Documented AI incidents rose to 362 in 2025, up from 233 in 2024. The OECD AI Incidents and Hazards Monitor tracked a six-month moving average of 326 incidents, peaking at 435 in January 2026. Almost all leading developers disclose capability benchmark results, the Index finds, but reporting on responsible AI benchmarks remains spotty.

Transparency is running in the wrong direction. The Foundation Model Transparency Index average score dropped to 40 in 2025, against 58 in 2024. Parameter counts, training data details, and compute budgets are increasingly withheld by the most capable systems. OpenAI, Anthropic and Google each withhold multiple key disclosure fields.

Capability itself is fragile in ways the headline benchmarks obscure. GPT-4o’s accuracy dropped from 98.2% to 64.4% when handling first-person false beliefs. DeepSeek R1 fell from above 90% to 14.4% under the same pressure. On AA-Omniscience, hallucination rates across 26 models ranged from 22% to 94%. Error rates on widely used benchmarks reached 42% on GSM8K and 31% on MMLU 5Sub. On the NOHARM clinical benchmark, leading LLMs produced between 11.8 and 14.6 severely harmful recommendations per 100 cases, 76.6% of them errors of omission.

Corporate governance has adapted unevenly. 11% of surveyed businesses reported no responsible AI policies in place, down from 24% in 2024; AI-specific governance roles grew 17% in 2025. The top barriers to scaling agentic AI systems were security and risk concerns (62%), followed by knowledge and training gaps (59%), budget constraints (48%) and regulatory uncertainty (41%).

Safety is not an abstract concern in this dataset. It is a concrete reporting gap.

The Economic Signal: Capital Tilts American, Entry Level Gets Squeezed

The fifth signal is the sheer capital gradient between jurisdictions and the first concrete labour-market footprint.

US private AI investment reached $285.9 billion in 2025, 23.1 times the $12.4 billion invested in China and 48.5 times the UK’s $5.9 billion. Global private AI investment totalled $344.66 billion, up 127.5% year over year. Generative AI private investment alone was $170.87 billion, up 200%. Global corporate AI investment reached $581.69 billion. There were 3,499 newly funded AI companies globally, of which 1,953 were in the United States, 172 in the UK and 161 in China. Twenty-eight funding events exceeded $1 billion.

Consumer value has kept pace. US consumer surplus from generative AI reached $172 billion annually by early 2026, up 54% from $112 billion, with the median value per user roughly tripling from $3.40 to $11.40. Productivity gains in the Index’s survey of studies run from 14% to 15% in customer support, 26% in software development, up to 50% in marketing output. There is one important counter-datapoint: in one specific study of experienced open-source developers, AI tooling was associated with a 19% productivity drop, a reminder that gains are not uniformly distributed.

The labour signal is the sharpest one. US developers aged 22 to 25 saw employment fall nearly 20% from 2024, even as headcount for older developers continued to grow. This is the first quantified evidence in a major independent dataset that the entry-level displacement thesis has a real footprint at the start of the software career ladder. AI agent deployment still sitting in single digits across business functions is why the headline number is not worse.

US vs China: The Ledger at a Glance

MetricUnited StatesChina
Notable AI models released in 20255030
Cumulative notable models, 2018–20251,618849
Private AI investment, 2025$285.9bn$12.4bn
Newly funded AI companies, 20251,953161
Top-model Arena gap vs. rival, Mar 2026Lead−2.7% / −39 Elo
Publications, citations, patent grantsTrailingLeads across all three
Industrial robot installations, 202434,200295,000 (54.4% of global)
Data centres5,427 (more than 10× next country)n/a
Public investment benchmark$20.4bn federal AI-related contracts/grants/OTAs, 2013–24Estimated $184bn in guidance funds, 2000–23

Read the two columns together and the thesis writes itself. The United States is the capital and hardware country. China is the deployment and volume country. Neither is close to the other on the other’s strength.

Talent: The Flow Into America Has Reversed

The sixth signal is a labour-market inversion that the Index treats as structural.

The number of AI researchers and developers moving to the United States has dropped 89% since 2017, with 80% of that decline in the last year alone. Switzerland and Singapore lead the world in AI researchers and developers per capita. Israel has the highest concentration of AI talent at 2.10%, followed by Singapore at 1.82% and Luxembourg at 1.60%. The number of new AI PhDs graduating in the United States and Canada increased 22% from 2022 to 2024, but the growth went disproportionately into academia, not industry; 65% of new AI PhDs entered industry, 31.59% academia, and 1.96% government.

Gender representation in the talent pool has barely moved. Women represent 34.3% of AI talent in the United States. In India, men list AI skills at 1.5 times the rate of women. The Index flags this plainly: no meaningful progress in any country since 2010.

The scenario to watch is whether the United States’ capital gradient is strong enough to pull through the talent lever even as raw migration compresses. Nothing in this year’s data says it is.

Policy Divergence and the Sovereignty Frame

The seventh signal is that the world’s policy regimes are actively diverging.

In the EU, the AI Act’s first prohibitions took effect on 2 February 2025, and general-purpose AI obligations began to apply on 2 August 2025. Italy passed Law No. 132/2025 on 17 September 2025. South Korea enacted the Framework Act on the Development of Artificial Intelligence and the Creation of a Foundation for Trust. Japan passed the Act on the Promotion of Research and Development and Utilization of Artificial Intelligence–Related Technology.

The United States moved the other way. The Trump administration’s “Removing Barriers to American Leadership in AI” executive order was signed on 23 January 2025. Executive Order 14110 was revoked. On 12 December 2025, “Ensuring a National Policy Framework for Artificial Intelligence” replaced the previous federal posture. At the state level, 150 AI-related bills passed into law in 2025, up from fewer than 10 in 2020; California alone enacted 20 in 2025. Fifty-eight federal AI-related regulations were issued in 2025, 28 of them by the Executive Office of the President. AI-related congressional hearings went from 5 witnesses in 2017 to 102 in 2025, 37% from industry.

Below the G7 level, AI sovereignty has become the organising frame. More than half of newly adopted national AI strategies came from developing countries entering the policy landscape for the first time. Europe and Central Asia expanded from 3 to 44 AI supercomputing clusters between 2018 and 2025. North America reached 41, East Asia and Pacific 27. Data localisation measures from 2000 to 2024 reached 77 in East Asia and Pacific and 71 in Sub-Saharan Africa, against just 3 in North America.

Between 2013 and 2024, US public investment in AI-related contracts, grants and OTAs totalled $20.4 billion, of which $15.9 billion was grants, $3.9 billion contracts, and $650 million OTAs. European nations committed $3.7 billion in the same period (UK $1.6 billion, Germany $505 million, France $320 million). Estimates put Chinese state guidance funds allocated to AI at $184 billion between 2000 and 2023.

Three regulatory postures, three different levels of state capital, one global model market. That is a policy arbitrage map, not a consensus.

The eighth signal is an expert–public divide that will shape political outcomes before it shapes models.

Per the Index, 73% of experts expect AI to have a positive impact on how people do their jobs, against just 23% of the US public, a 50-percentage-point gap. On economic impact, the split is 69% experts against 21% public. On medical care it narrows to 84% experts against 44% public. On K–12 education it is 61% against 24%. 64% of US adults expect AI to lead to fewer jobs over the next 20 years, against 39% of AI experts.

Regulatory trust is stratified sharply. Globally, 54% trust governments to regulate AI responsibly, led by Singapore at 81% and Indonesia at 76%. The United States is at 31%, the lowest in its peer set. Trust in specific regulators is 53% for the EU, 37% for the United States and 27% for China. Inside the United States, 41% of respondents say federal AI regulation will not go far enough, 27% say it will go too far, and over 33% are unsure. Globally, 59% say AI benefits outweigh drawbacks, up from 55% in 2024, even as 52% say AI makes them nervous, up from 50%.

The durable consequence is that the political window is wide open for regulation in jurisdictions other than the United States. The EU is already spending that window. Developing countries are drafting strategies into it. The United States is betting, currently, that faster deployment beats consensus-building.

What to Watch Over the Next Twelve Months

The ninth signal is really a set of arrows pointing into 2026.

  1. Data-wall dynamics. Synthetic data is still not replacing real data in pre-training. Post-training techniques and data-quality work are the ones carrying most of the gains. The question is whether post-training can compound fast enough to outpace a flattening pre-training curve.
  2. Open-weight convergence. OLMo 3.1 Think 32B achieves comparable results on several benchmarks to Grok 4 despite nearly 90 times fewer parameters. The top closed-weight model leads the top open-weight model by only 3.4%. Watch whether this gap narrows or snaps back as closed labs push capability.
  3. Agent reliability. OSWorld at 66.3% is impressive as an academic benchmark and unusable as an enterprise substrate. The gap between 66.3% task success and the single-digit deployment number is the reliability question the sector has to solve next.
  4. Clinical evidence. 258 AI medical devices were authorised by the FDA in 2025 through September, taking the cumulative total to 1,357. Only 2.4% of FDA-authorised devices with clinical studies had randomised trial data behind them. Diagnostic multi-agent systems are hitting 85.5% accuracy on complex published cases against 20% for unaided physicians, and ambient documentation systems are saving up to 83% of note-writing time in Epic-integrated hospitals. Regulatory evidence has to catch up with deployment.
  5. Taiwan exposure. One foundry, one grid, one dominant chip vendor. The AI market’s structural risk is not a training breakthrough. It is a geopolitical interruption in the hardware supply chain.

Bottom Line

The Index’s contribution this year is not a single headline number. It is the quiet accumulation of signals pointing in the same direction. Capability is accelerating. The US–China gap is effectively closed. Industry has almost completely absorbed model production. Capital is tilting harder toward the United States. The hardware stack is more concentrated, not less. The safety reporting baseline is weakening. And the first durable labour-market footprint has appeared at the bottom of the software career ladder.

Strip away the narrative layer and what the Index actually offers is an instrument panel. Anyone tracking the model race should read it the way a portfolio manager reads a 10-K: for what the numbers constrain, not for what the prose celebrates. The full report is available in full at the Stanford AI Index 2026.