GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

4mo ago benchmark

Microsoft Research's Webwright Hits 60.1% on Odysseys — 35 Points Above Previous Web Agent SOTA

Microsoft Research's AI Frontiers lab released Webwright, a terminal-native web agent that scores 60.1% on the Odysseys long-horizon benchmark, a 35-point jump above Opus 4.6's previous state-of-the-art 44.5%.

4mo ago research

George Hotz: AI Coding Agents Front-Load Progress, Then Stall — Large Orgs Are Most Exposed

George Hotz published a pointed critique of AI coding agents: they produce an initial burst of progress then fail to deliver polish, and large organizations with weak feedback loops are the most at risk from the resulting slop.

4mo ago research

ZEDA Cuts 50% of MoE Expert Compute on Already-Deployed Models — 20% Faster Inference, Near-Zero Accuracy Loss

A new technique called ZEDA retrofits static Mixture-of-Experts models to dynamically skip unnecessary experts at inference time — without retraining from scratch. Tested on Qwen3-30B-A3B and GLM-4.7-Flash, it eliminates over half of expert FLOPs with a 1.20x end-to-end speedup.

4mo ago research

Coding Agents Lose 30 Assertion Points When Backend Constraints Stack — New Paper Names the Failure Mode

A systematic study of LLM agents across 100 backend code generation tasks and 8 web frameworks finds that agents capable of simple greenfield generation degrade by 30 points or more as structural constraints accumulate. The problem has a name: constraint decay.

4mo ago research

HBM Is Now 63% of AI Chip Component Costs — Up From 52% Just 18 Months Ago

Epoch AI's chip cost analysis finds high-bandwidth memory has gone from just over half to nearly two-thirds of total AI chip component spending between Q1 2024 and Q4 2025, with absolute HBM spend nearly tripling to $32 billion annually.

4mo ago release

Anthropic Embeds Mythos-1 in Claude Code Source — Late-June Commercial Window Exposed

Source code strings confirm 'claude-mythos-1-preview' is wired into Claude Code and Claude Security. Users briefly saw the model in the UI before it was pulled. Anthropic's timeline points to a restricted commercial launch alongside Opus 4.8, expected late June to early July.

4mo ago release

NVIDIA Open-Sources Nemotron-Labs Diffusion: One Checkpoint, Three Modes, 6.4x the Throughput

NVIDIA released Nemotron-Labs Diffusion on May 23 — 3B, 8B, and 14B open-weight models that run autoregressive, diffusion, and self-speculative generation from a single checkpoint. The 8B base pulled 228,000 downloads in under two days. Throughput claims are vendor-reported and await independent verification.

4mo ago policy

Trump Canceled the AI Oversight EO Hours Before Signing — Sacks Called in the Morning, Musk and Zuckerberg Followed

President Trump abruptly postponed a White House AI executive order ceremony on May 21 after David Sacks, Elon Musk, and Mark Zuckerberg lobbied against it. The draft would have established 90-day voluntary pre-release reviews for frontier AI. The administration has not said when or whether it will return.

4mo ago benchmark

WebArena Hit 74.3% While One in Three Enterprises Report 25%+ Agent Failure Rates: Stanford AI Index 2026

Stanford HAI's ninth annual AI Index shows agents within four percentage points of human performance on structured benchmarks — and the same models reading analog clocks correctly only 50.1% of the time. Three quarters of enterprises now report double-digit agent failure rates in production.

4mo ago research

Microsoft Fara1.5 Hits 63% on Mind2Web and 86.6% on WebVoyager With a 9B Open-Weight Browser Agent

Microsoft Research ships Fara1.5 — three open-weight computer-use models at 4B, 9B, and 27B that outperform every comparable-size model on web task benchmarks. Paired with MagenticBrain and MagenticLite, it is a full local agent stack designed to run on modest hardware.

4mo ago benchmark

Cerebras Runs Kimi K2.6 at 981 Tokens/Second — 6.7x Faster Than Any GPU Cloud on a 1T-Parameter Model

Cerebras benchmarked its wafer-scale hardware on Kimi K2.6, reporting 981 tokens/second — 6.7x faster than the next GPU cloud provider on a 1-trillion-parameter model, validated by Artificial Analysis.

4mo ago funding

Anthropic's $30B Round Is Closing Next Week — Sequoia, Dragoneer, Altimeter, Greenoaks Co-Lead at $900B

Bloomberg reports Anthropic will close a round tracking above $30B at a $900B valuation as soon as next week, vaulting it ahead of OpenAI. Sequoia, Dragoneer, Altimeter, and Greenoaks each commit roughly $2B. Revenue run rate is now $30B, up from $14B in February.

4mo ago funding

Wall Street Files Lab ETFs for All Five Frontier AI Ecosystems — Anthropic, OpenAI, DeepMind, Meta, xAI

Harbor Capital filed for five actively managed Lab ETFs on May 22, each targeting the public-market ecosystem of one AI lab. Confirmed by Bloomberg ETF analyst James Seyffart. The move financializes AI lab dominance without requiring direct stakes in private companies.

4mo ago benchmark

Gemini 3.5 Flash's 1M-Token Context Window: 77% Retrieval Accuracy at 128K, 27% at Full Scale — and Google Skipped General Reasoning Benchmarks

Independent analysis of Gemini 3.5 Flash finds its headline 1M-token context window degrades sharply in practice: MRCR v2 retrieval accuracy falls from 77.3% at 128K tokens to 26.6% at 1M. Google published only agentic benchmarks at I/O — no MMLU, GPQA, or competition math results.

4mo ago research

Glasswing Month One: UK AISI Confirms Mythos Solves Both Cyber Ranges End-to-End, 2,100 Vulnerabilities Patched — Public Release Still Blocked

Anthropic's first Glasswing progress report: UK AI Security Institute confirms Mythos Preview is the first model to complete both multistep cyberattack simulations end-to-end. Claude Security patched 2,100 vulnerabilities in three weeks. Anthropic will not release Mythos-class models publicly until substantially stronger safeguards exist.