GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%

Live Feed

1mo ago policy

Amodei: Anthropic Does Not Want to Ban Open Weights — It Wants Chip Controls, Anti-Distillation, and Global Safety Tests

Dario Amodei published Anthropic's formal position on open-weights models July 28, explicitly rejecting calls for bans while breaking with the Microsoft-NVIDIA coalition on distillation enforcement and mandatory safety testing.

1mo ago benchmark

Claude Opus 5 Sets ARC-AGI-3 Record at 30.2% — 4x the Previous Best

Anthropic's Claude Opus 5 scored 30.2% on ARC-AGI-3, nearly four times GPT-5.6 Sol's prior record of 7.8%. It also tops the Artificial Analysis Intelligence Index and undercuts Fable 5's per-task cost by 70%.

1mo ago policy

Google Loses SerpApi DMCA Bid: Public Search Results Are Not a Copyright Lock

A California federal court dismissed Google's core anti-circumvention claims against SerpApi, limiting how far platforms can stretch the DMCA to block scraping of public search results. Google gets 21 days to rework a narrower image-related claim.

1mo ago release

Microsoft Puts MAI-Cyber-1-Flash Inside MDASH: 96% CyberGym at Half the Cost

Microsoft's first cyber-specific model routes 90% of MDASH vulnerability work away from expensive frontier models and reserves GPT-5.4 for the hardest 10%. The combined system scores 95.95% on CyberGym, roughly 12 points above Mythos, while cutting MDASH model cost by 50%.

1mo ago policy

OpenAI and Anthropic Hit $3.17M Q2 Lobbying Record as AI Policy Moves to Washington

OpenAI and Anthropic spent a combined $3.17 million on federal lobbying in Q2 2026, up 23% from Q1 and roughly double year over year. Anthropic alone spent $1.97 million, more than Nvidia, as cybersecurity, copyright, cloud, defense procurement, and model oversight moved into the policy stack.

1mo ago research

You Can Predict RL Gains Before Running Them: Pretraining Loss Is the Oracle

A new paper finds that pretraining validation loss predicts post-RL pass@1 with correlation up to 0.99, and RL's compute-optimal share maxes out at 28% — before that ceiling, labs are burning GPU time they could have priced out in advance.

1mo ago benchmark

Claude Opus 5 Enters LiveBench at #3 Overall — Beats Fable 5 on Agentic Coding 61.3% to 46.9%

Claude Opus 5 Thinking xHigh debuted at 80.3 overall on LiveBench, landing third behind GPT-5.6 Sol and Fable 5. On the metric that matters most for production agents, it beats Fable 5 by 14.4 points — and costs $0.487 per successful task versus Fable 5's $1.573.

1mo ago research

Refused Isn't Cleared: Agent Memory Files Carry Injections Into the Next Session

A University of Washington study finds that AI agents often refuse prompt injection attacks in memory files but leave the malicious instruction intact — where it waits for the next session, a cheaper model, or a subagent that won't say no.

1mo ago benchmark

LiveBench Agentic Coding: Terra Leads at 68.0%, Fable 5 Posts the Worst Score in the Frontier Tier at 46.9%

The July 2026 LiveBench data shows GPT-5.6 Terra leading agentic coding at 68.0%, with the cheapest cost per successful task in the frontier tier at $0.497. Fable 5 scores last among top-tier models at 46.9% agentic coding while costing $1.573 per task, 3.2x Terra's price.

1mo ago research

LLMs Just Made Formal Software Verification Practical: A Google Engineer Proved zstd in Lean 4

Adam Langley, the Google security engineer behind BoringSSL, formally verified a Zstandard decompressor in Lean 4 using LLM-assisted proof automation. His conclusion: the 10x proof overhead that made formal verification economically unviable for production software no longer applies.

1mo ago release

Gemini 3.5 Flash Lite Targets the Sub-Agent Layer at $0.30/$2.50 Per Million

Google released Gemini 3.5 Flash Lite on July 21, a model designed specifically for sub-agent execution in multi-agent workflows. At 174 tokens per second and a 90% cache discount, it undercuts Flash on price while keeping the 1M context window.

1mo ago benchmark

Anthropic's Project Pilot Puts Claude Behind the Controls of a Surveillance Drone

Anthropic has published a new autonomous drone benchmark, Drone-Bench, testing whether frontier AI models can pilot a flying drone to locate and follow a target without human input. The work surfaces real capability thresholds and safety risks Anthropic says its Frontier Red Team now tracks systematically.

1mo ago research

Terence Tao Tells ICM 2026 That AI Reasoning for Science Is Becoming Measurable and Cheap

At the International Congress of Mathematicians 2026, Fields Medal winner Terence Tao delivered a public lecture on mathematics and AI, arguing that frontier models' reasoning capabilities are now trackable across scientific domains and cheap enough that labs should be doing it systematically. His talk arrives weeks after Claude Fable 5 produced a verifiable counterexample to the 85-year-old Jacobian Conjecture.

2mo ago release

Microsoft's MAI-Voice-2-Flash Undercuts OpenAI by Up to 89% in Enterprise Voice

Microsoft released MAI-Voice-2-Flash into public preview alongside MAI-Image-2.5-Pro, claiming up to 89% lower cost than OpenAI equivalents. The company holds a $13B+ stake in OpenAI and now offers cheaper competing models on the same Azure platform.

2mo ago release

OpenAI's Jalapeño Chip Goes Official: Broadcom-Designed, 50% Cheaper Than Nvidia, Ships End-2026

OpenAI and Broadcom unveiled Jalapeño on June 24, OpenAI's first custom inference chip, codenamed Titan. Broadcom CEO Hock Tan claims 50% better cost efficiency over standard AI GPUs. Deployment starts end-2026, full commercial scale by H1 2028.