GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

2mo ago release

Poolside's Laguna S 2.1: 118B MoE Hits 70.2% Terminal-Bench at $0.10 Per Million

Poolside released Laguna S 2.1 on July 21, an open-weight 118B MoE model with only 8B active parameters. It scores 70.2% on Terminal-Bench 2.1 and 40.4% on DeepSWE at $0.10/$0.20 per million tokens under the OpenMDW-1.1 license.

2mo ago release

Google Ships Three Gemini Models: 3.6 Flash at $1.50/M, 49% DeepSWE, and a Restricted Cyber Tier

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on July 21. The workhorse model cuts output tokens 17%, beats its predecessor on DeepSWE 49% vs 37%, and arrives cheaper than 3.5 Flash. The cyber model is government-only.

2mo ago research

Xiaomi-Robotics-1 Scales a Robot Policy Model on 100,000 Hours of Real-World Data

Xiaomi's new VLA foundation model applies LLM-style two-stage training to robotics: 100,000 hours of embodiment-free pre-training across 1,700 scenarios, then 7,200 hours of real-robot post-training. It is the largest data-scale effort yet to test whether scaling laws hold for robot policy models.

2mo ago benchmark

AISI: Open-Weight Models Are Now 4-7 Months Behind Closed Frontier on Cyber

The UK AI Security Institute's first public comparative cyber benchmark finds GLM-5.2 matching closed models from four to seven months ago. The gap was 6-10 months in early 2025. GPT-5.6 Sol completed a 32-step autonomous attack range in 7 of 10 attempts when given 100M inference tokens per run.

2mo ago benchmark

Kimi K3 Leads Harvey LAB-AA at 26.7%, Nearly Double Fable 5 — Then Fixed 15 Bugs Fable Refused

On Artificial Analysis's 120-task legal agent benchmark, Kimi K3 scores 26.7% against Claude Fable 5's 14.2%. A separate $250 run patched 15 critical security bugs that Codex and Fable both declined on guardrail grounds.

2mo ago policy

Anthropic's $1.5B Author Settlement Approved — Training Was Fair Use, Storing 7M Books Was Not

A federal judge approved Anthropic's $1.5 billion copyright settlement on July 20, closing the first major US AI training case. The ruling preserves fair use for AI training while drawing a hard line at pirated book libraries.

2mo ago research

GPT-5.6 Sol Formally Proves Erdős Unit Distance Conjecture in Lean — 1.2 Million Lines From Axioms in 3 Weeks

Kevin Buzzard, Imperial College mathematician and Lean library maintainer, documents three AI-driven breakthroughs across 7 weeks: ChatGPT disproving the Erdős Unit Distance conjecture in May, Yann LeCun's Logical Intelligence formalizing the proof within days, and OpenAI's Sol generating 1.2 million lines of Lean code to complete the full formalization from mathematical axioms. Half of mathlib's nine-year output, in three weeks.

2mo ago release

Thinking Machines Lab Ships Inkling: 975B Open-Weight MoE, 77.6% SWE-bench, Enters Agent Arena

Mira Murati's Thinking Machines Lab releases Inkling — a 975B Mixture-of-Experts model trained on 45 trillion tokens across text, image, audio, and video, available Apache 2.0 with full weights. Architecture mirrors DeepSeek-V3. Added to Text, Code, and Agent Arena leaderboards within five days of launch.

2mo ago funding

Hut 8 Seals Second $9.8B Lease to Fully Commercialize 1GW Beacon Point at $19.6B

The same investment-grade tenant that signed Phase 1 has doubled its contracted capacity to 704 MW with a second 15-year lease, filling the entire 1GW Beacon Point campus. Hut 8's total portfolio aggregate base-term contract value reaches $26.6 billion.

2mo ago research

GPT-5.6 Found a $500K WordPress RCE for $25 — 500 Million Sites Were Exposed

A security researcher used GPT-5.6 and $25 in API credits to find wp2shell, twin pre-authentication RCE flaws in WordPress Core. Every 6.9 and 7.0 site was vulnerable until emergency patches shipped. Public exploits are now circulating.

2mo ago policy

US Power Companies Are Using Eminent Domain to Seize Land for AI Data Center Infrastructure

Power utilities in Georgia and other states are invoking eminent domain — the government's power to compel land sales — to build transmission lines serving AI data centers. Homeowners facing forced sales call it theft. Georgia Power says it's a last resort used in under 1% of cases.

2mo ago research

Context Is the Hidden Layer: 300-Run Study Shows Agent Setup Quality Predicts Failure Before the Agent Runs

A new paper from ProofAgent validates context-engineering quality as a leading indicator of AI agent reliability across 300 multi-turn evaluations and 7,500 turns. The weakest context is often the cheapest — and the most dangerous.

2mo ago research

Claude Fable Helps Disprove the 85-Year-Old Jacobian Conjecture — Verified Arithmetic, One Tweet

Levent Alpoge, a number theorist at Anthropic, used Claude Fable 5 to produce an explicit counterexample to the Jacobian Conjecture — an open problem since 1939 and one of Smale's 23 hardest questions for the 21st century. The counterexample is three lines of arithmetic anyone can verify by hand.

2mo ago research

Google's TPU Push Into Neoclouds Is Hitting Nvidia's Distribution Moat — CoreWeave, Nebius, and Lambda Won't Switch

Google is trying to sell TPUs through neocloud providers to compete with Nvidia in the market that sits between hyperscalers and AI labs. The Information reports CoreWeave, Nebius, and Lambda have said no. Smaller providers and a Blackstone vehicle are the fallback.

2mo ago research

AI Access Collapses 'I Don't Know' Rate From 44% to 3% — Accuracy Drops by Two-Thirds, Confidence Doubles

A three-university controlled study gave participants access to an AI that was deliberately wrong on the test questions. Accuracy fell from 27% to 9%. Confidence climbed from 30% to 76%. Monetary incentives barely moved the numbers.