GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

3mo ago benchmark

MLPerf Training v6.0: NVIDIA Sweeps All 7 Benchmarks, DeepSeek-V3 671B Trained in 2 Minutes

MLCommons released MLPerf Training v6.0 results on June 16. NVIDIA won every benchmark and was the only platform to submit across all seven. CoreWeave trained DeepSeek-V3's 671-billion-parameter MoE model to target quality in 2.02 minutes on 8,192 GB300 NVL72 GPUs — the fastest Available-cloud result in the benchmark's history.

3mo ago funding

SpaceX Signs the Cursor Deal: $60B Stock Acquisition Closes the Option, Grok V9-Medium Ships Alongside

SpaceX converted its April option into a definitive $60B all-stock merger agreement for Anysphere, the company behind Cursor. The deal is set to close Q3 2026, and Grok V9-Medium — trained on Cursor developer workflows — is simultaneously rolling out to users.

3mo ago ai

OpenAI's Leaked Financials Put Hard Numbers on Altman's Compute Bet

Audited documents show OpenAI generated $13.07B in 2025 revenue against a $20.92B operating loss and $34B cost base — with $17.2B of that going to Microsoft. The leak lands as OpenAI's confidential S-1 sits at the SEC.

3mo ago benchmark

GLM 5.2 Debuts at #10 on Agent Arena — Z.ai's Open-Weight Model Cracks the Proprietary Tier

Z.ai's GLM 5.2 (Max) entered the Agent Arena leaderboard on June 16 at rank 10 with a 4.37% composite score, landing above Claude Opus 4.8 and Claude Sonnet 4.6. It is the only MIT-licensed open-weight model in the top 10.

3mo ago research

Claude Code's Core Is a While-Loop. The Infrastructure Around It Has Five Layers.

A new paper reverse-engineers Claude Code's TypeScript source and finds the agent loop itself is trivial. Everything that makes it work lives in the surrounding harness: a 7-mode permission classifier, a 5-layer context compaction pipeline, and four extensibility mechanisms.

3mo ago benchmark

Artificial Analysis Overhauls Its Intelligence Index: τ³-Bench Banking Replaces Telecom, DeepSeek V4 Pro Costs 45x Less Than Opus 4.8 Per Task

Artificial Analysis has released Intelligence Index v4.1 — a methodological overhaul that replaces τ²-Bench Telecom with τ³-Bench Banking, upgrades GDPval to v2 with human-baseline ELO and 250-turn horizons, and adds per-task cost and time metrics. Claude Fable 5 leads at 60 but is unavailable; Opus 4.8 is the smartest available model at 56.

3mo ago funding

AlphaGo Creator Raises Europe's Largest-Ever Seed Round — $1.1B for a London Lab Chasing Superintelligence via Reinforcement Learning

Ineffable Intelligence, founded by AlphaGo architect David Silver, has closed a $1.1 billion seed round and signed a cloud partnership with Google to deploy NVIDIA Vera Rubin NVL72 GPUs at scale. The lab's stated target is a 'superlearner' — an AI that generates and learns from its own experience without relying on human-generated training data.

3mo ago release

Alibaba Releases Qwen-Robot Suite: Three Foundation Models Unifying Navigation, Manipulation, and World Simulation Across 20+ Robot Types

Alibaba's Qwen team has shipped three physical AI foundation models — RobotNav, RobotManip, and RobotWorld — trained on a 38,100-hour cross-embodiment corpus under a single world model. The suite bridges the gap between Qwen's vision-language understanding and physical robot control, enabling one model to operate across 20+ embodiment types through a natural-language action interface.

3mo ago benchmark

Fable 5 Enters Arena Search at #3 — Claude Opus 4.6 Holds the Summit at 1252 Elo

Arena added Claude Fable 5 and Opus 4.8 to its Search leaderboard on June 15. Fable 5 debuts at rank 3 with 1237 Elo — 15 points behind Claude Opus 4.6 and 3 behind GPT-5.5. With only 5,007 battles accumulated, the score is still volatile.

3mo ago benchmark

CoreWeave First to Deploy NVIDIA Vera Rubin NVL72 at Rack Scale — 10x Inference Efficiency, MLPerf v6.0 Record

CoreWeave validated NVIDIA's Vera Rubin NVL72 in June, becoming the first AI cloud to run the architecture at full rack scale. The 72-GPU, 36-CPU platform delivers 10x inference per watt and one-tenth the cost per million tokens vs Blackwell. Stock jumped 14% on the news; the company leads MLPerf v6.0 inference with a 2x performance record.

3mo ago release

AWS Kiro Adds $100/Month Pro Max Tier and Self-Correcting Subagents — Amazon Q Developer Sunsets April 2027

Kiro Pro Max lands at $100/month with 5,000 credits and all frontier models, filling the gap between $40 Pro+ and $200 Power. A self-correcting subagent pipeline also shipped this week. Amazon Q Developer new signups closed May 15; full platform EOL is April 2027.

3mo ago model

OpenAI Pulled GPT-5.2 From ChatGPT 61 Days Before Its Documented End-of-Life

A June 10 model picker rollout silently removed GPT-5.2 from ChatGPT despite published deprecation docs listing August 10 as the shutdown date. Pro subscribers at $100/month found out through a missing model option, not an announcement.

3mo ago research

3.2 Million Math Records: AI Lifts Completion Speed 23%, Drops Proctored Retention 25%

A study tracking 3.2 million ALEKS math problems across 10 years found AI assistance sped up completion on word problems but reduced proctored retention scores by 25%. Graph problems, which resist AI handoff, showed no effect. The friction that slows problem-solving is also what makes knowledge stick.

3mo ago release

Anthropic Launches Claude Corps: $150M to Place 1,000 AI Fellows at 400 Nonprofits

Anthropic is committing $150 million to embed 1,000 early-career fellows inside 400+ nonprofit organizations for a year. No degree required. CodePath and Social Finance operate the program.

3mo ago policy

Semafor: China-Linked Group Had Already Accessed Mythos Before the Export Order Landed

New reporting ties the White House export control on Fable 5 and Mythos to a specific intelligence finding: a China-linked group had accessed Mythos and could distill its capabilities into a separate model. The jailbreak dispute was a secondary concern.