Live Feed
Microsoft Research's Webwright Hits 60.1% on Odysseys — 35 Points Above Previous Web Agent SOTA
Microsoft Research's AI Frontiers lab released Webwright, a terminal-native web agent that scores 60.1% on the Odysseys long-horizon benchmark, a 35-point jump above Opus 4.6's previous state-of-the-art 44.5%.
George Hotz: AI Coding Agents Front-Load Progress, Then Stall — Large Orgs Are Most Exposed
George Hotz published a pointed critique of AI coding agents: they produce an initial burst of progress then fail to deliver polish, and large organizations with weak feedback loops are the most at risk from the resulting slop.
ZEDA Cuts 50% of MoE Expert Compute on Already-Deployed Models — 20% Faster Inference, Near-Zero Accuracy Loss
A new technique called ZEDA retrofits static Mixture-of-Experts models to dynamically skip unnecessary experts at inference time — without retraining from scratch. Tested on Qwen3-30B-A3B and GLM-4.7-Flash, it eliminates over half of expert FLOPs with a 1.20x end-to-end speedup.
Coding Agents Lose 30 Assertion Points When Backend Constraints Stack — New Paper Names the Failure Mode
A systematic study of LLM agents across 100 backend code generation tasks and 8 web frameworks finds that agents capable of simple greenfield generation degrade by 30 points or more as structural constraints accumulate. The problem has a name: constraint decay.
HBM Is Now 63% of AI Chip Component Costs — Up From 52% Just 18 Months Ago
Epoch AI's chip cost analysis finds high-bandwidth memory has gone from just over half to nearly two-thirds of total AI chip component spending between Q1 2024 and Q4 2025, with absolute HBM spend nearly tripling to $32 billion annually.
Anthropic Embeds Mythos-1 in Claude Code Source — Late-June Commercial Window Exposed
Source code strings confirm 'claude-mythos-1-preview' is wired into Claude Code and Claude Security. Users briefly saw the model in the UI before it was pulled. Anthropic's timeline points to a restricted commercial launch alongside Opus 4.8, expected late June to early July.
NVIDIA Open-Sources Nemotron-Labs Diffusion: One Checkpoint, Three Modes, 6.4x the Throughput
NVIDIA released Nemotron-Labs Diffusion on May 23 — 3B, 8B, and 14B open-weight models that run autoregressive, diffusion, and self-speculative generation from a single checkpoint. The 8B base pulled 228,000 downloads in under two days. Throughput claims are vendor-reported and await independent verification.
Trump Canceled the AI Oversight EO Hours Before Signing — Sacks Called in the Morning, Musk and Zuckerberg Followed
President Trump abruptly postponed a White House AI executive order ceremony on May 21 after David Sacks, Elon Musk, and Mark Zuckerberg lobbied against it. The draft would have established 90-day voluntary pre-release reviews for frontier AI. The administration has not said when or whether it will return.
WebArena Hit 74.3% While One in Three Enterprises Report 25%+ Agent Failure Rates: Stanford AI Index 2026
Stanford HAI's ninth annual AI Index shows agents within four percentage points of human performance on structured benchmarks — and the same models reading analog clocks correctly only 50.1% of the time. Three quarters of enterprises now report double-digit agent failure rates in production.
Microsoft Fara1.5 Hits 63% on Mind2Web and 86.6% on WebVoyager With a 9B Open-Weight Browser Agent
Microsoft Research ships Fara1.5 — three open-weight computer-use models at 4B, 9B, and 27B that outperform every comparable-size model on web task benchmarks. Paired with MagenticBrain and MagenticLite, it is a full local agent stack designed to run on modest hardware.
Cerebras Runs Kimi K2.6 at 981 Tokens/Second — 6.7x Faster Than Any GPU Cloud on a 1T-Parameter Model
Cerebras benchmarked its wafer-scale hardware on Kimi K2.6, reporting 981 tokens/second — 6.7x faster than the next GPU cloud provider on a 1-trillion-parameter model, validated by Artificial Analysis.
Anthropic's $30B Round Is Closing Next Week — Sequoia, Dragoneer, Altimeter, Greenoaks Co-Lead at $900B
Bloomberg reports Anthropic will close a round tracking above $30B at a $900B valuation as soon as next week, vaulting it ahead of OpenAI. Sequoia, Dragoneer, Altimeter, and Greenoaks each commit roughly $2B. Revenue run rate is now $30B, up from $14B in February.
Wall Street Files Lab ETFs for All Five Frontier AI Ecosystems — Anthropic, OpenAI, DeepMind, Meta, xAI
Harbor Capital filed for five actively managed Lab ETFs on May 22, each targeting the public-market ecosystem of one AI lab. Confirmed by Bloomberg ETF analyst James Seyffart. The move financializes AI lab dominance without requiring direct stakes in private companies.
Gemini 3.5 Flash's 1M-Token Context Window: 77% Retrieval Accuracy at 128K, 27% at Full Scale — and Google Skipped General Reasoning Benchmarks
Independent analysis of Gemini 3.5 Flash finds its headline 1M-token context window degrades sharply in practice: MRCR v2 retrieval accuracy falls from 77.3% at 128K tokens to 26.6% at 1M. Google published only agentic benchmarks at I/O — no MMLU, GPQA, or competition math results.
Glasswing Month One: UK AISI Confirms Mythos Solves Both Cyber Ranges End-to-End, 2,100 Vulnerabilities Patched — Public Release Still Blocked
Anthropic's first Glasswing progress report: UK AI Security Institute confirms Mythos Preview is the first model to complete both multistep cyberattack simulations end-to-end. Claude Security patched 2,100 vulnerabilities in three weeks. Anthropic will not release Mythos-class models publicly until substantially stronger safeguards exist.