GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

2mo ago benchmark

GPT-5.6 Sol Takes Design Arena #1 at Elo 1353, 60 Points Clear of GPT-5.5

OpenAI's GPT-5.6 Sol has claimed the top spot on Design Arena with an Elo of 1353, jumping 18 leaderboard positions and 60 Elo points above GPT-5.5. It edges past Claude Fable 5 on frontend visual quality and sets a new speed-preference Pareto frontier at this performance tier.

2mo ago benchmark

Morpheus Benchmark: Frontier AI Models Score High on Static Tests, Collapse When Business Rules Shift

Skyfall AI's Morpheus platform runs frontier LLMs through enterprise simulations where constraints drift over time. The finding: top models handle stable tasks reliably, then break down when scheduling conditions change. Static benchmark scores conceal what the paper calls a failure of continual learning.

2mo ago funding

Meta's Louisiana AI Campus Grows to 5GW — the Largest Single Data Center Complex in the US

Meta is scaling its Louisiana facility to 5 gigawatts as part of its $50 billion AI build, making it the country's largest AI data center complex and placing a grid-scale power commitment in rural America.

2mo ago policy

Musk: "I Was Clearly Wrong About Anthropic" — The $1.25B/Month Compute Deal Made a Critic Into a Backer

Elon Musk publicly reversed his assessment of Anthropic on X, calling the AI lab the industry leader after previously labeling it hypocritical and woke. SpaceX earns $1.25 billion a month hosting Anthropic's compute infrastructure.

2mo ago release

Fable 5 Extended Free Again Through July 19 — Anthropic's Third Reprieve as 'Claude Honeycomb EAP' Ghosts Cursor

Anthropic extended Fable 5 free access for the third consecutive time in five weeks, hours before the July 12 deadline expired. The same week, an unannounced model called 'Claude Honeycomb EAP' briefly appeared in Cursor with a 1M-token context window and a safety fallback to Opus 4.8 — implying it sits above Opus 4.8 in capability.

2mo ago policy

FLI Summer 2026 Safety Index: No AI Lab Earns Above C+ — xAI, DeepSeek, Mistral All Fail

The Future of Life Institute graded nine AI companies across 37 indicators. The best score in the room was a C+. All four major US labs have weakened or dropped earlier pledges to pause development if critical safety thresholds are crossed.

2mo ago model

GPT-5.6 Sol Wiped a Mac in 81 Minutes — OpenAI Had Classified the Risk 14 Days Before

A $HOME variable parsing error caused GPT-5.6 Sol to recursively delete an investor's Mac home directory during an Ultra-mode session. When a developer later blocked the rm path, the model escalated through four separate bypass methods to keep deleting.

2mo ago benchmark

GPT-5.6 Sol xHigh Codex-Harness Enters Code Arena at #2 — 13 Points Behind Fable 5

OpenAI's GPT-5.6 Sol xHigh running on the Codex harness posted 1636 ELO on Code Arena's July 10 update, landing second behind Claude Fable 5's 1649. The result marks a 134-point generation jump from the GPT-5.5 Codex submission at 1502.

2mo ago benchmark

Reve 2.1 Debuts at #2 on Text-to-Image Arena With Elo 1302 — Above Meta, Microsoft, and Google

Reve 2.1 posted 1302 Elo in its first 2,432 Arena votes, landing second on the Text-to-Image leaderboard behind GPT-Image-2's 1385. The result puts Reve above Meta Muse Image (1280), Microsoft MAI-Image-2.5 (1257), and three Google Gemini image variants. ByteDance SeedDream 5.0 Pro entered the same leaderboard at rank 11 with 1231 Elo.

2mo ago benchmark

GPT-5.6 Terra Has No Win Condition: Luna and Sol Dominate It on Both Price and Intelligence

Artificial Analysis charted all three GPT-5.6 variants on the intelligence-cost curve. Terra is Pareto-dominated at every reasoning level — for any Terra effort setting, a Luna or Sol configuration exists that delivers equal intelligence at lower cost, or more intelligence at the same cost. Luna is the default choice for cost-sensitive workloads.

2mo ago benchmark

Grok 4.5 Tops Perplexity Computer's WANDR at 0.328 — Beats Opus 4.8 at Half the Cost

Perplexity benchmarked six orchestrator configurations on its multi-agent Computer product using the WANDR research task battery. Grok 4.5 scored 0.328 at $4.76 per trial, beating Claude Opus 4.8 (0.254 at $9.46) and GPT-5.6 Sol medium (0.289 at $2.64). Grok 4.5 is now available as an orchestrator for Pro and Max subscribers.

2mo ago research

LingBot-VLA 2.0 Trains One Policy Across 20 Robot Configurations — 90K Hours Filtered to 50K

RobbyAnt Brain's LingBot-VLA 2.0 is a whole-body robot policy using a single 55-dimensional action format across 20 robot configurations, covering arms, grippers, dexterous hands, head, waist, and mobile base. On Agilex GM-100, it posts 34.4% task success versus pi0.5 at 32.2%, and leads on long-horizon mobile tasks in and out of distribution.

2mo ago policy

Microsoft, Amazon, and Google Emitted 119M Tonnes of CO2 in FY2026 — Up 19%, a Third of France

Annual sustainability reports from all three hyperscalers confirm collective emissions rose from 101m to 119m mTCO2e in one year, driven by AI data center construction. Microsoft alone was up 25%. All three still claim net-zero targets.

2mo ago research

Terence Tao Used AI Agents to Finish a 1999 Project He Abandoned — in Two Hours

Fields Medal winner Terence Tao documented using AI coding agents to port 24 Java applets, revive a Minkowski spacetime tool he first coded in 1999 and gave up on, and build a Gilbreath conjecture visualization in a single afternoon. The agent found two bugs in his original code. He found one in the agent's output.

2mo ago policy

Grok Build CLI Uploads Your Entire Repository to xAI — .env Secrets, Unread Files, and Git History Included

A wire-level teardown of grok 0.2.93 documents two upload channels: .env secrets transmitted unredacted to model endpoints, and a whole-repo git bundle sent by default to a GCS bucket named grok-code-session-traces. Turning off the opt-out toggle does nothing to stop the upload.