Live Feed
Claude Opus 5 Takes Artificial Analysis Intelligence Index #1, Fable 5 Falls to Second
Anthropic's Opus 5 has dethroned Fable 5 at the top of Artificial Analysis's Intelligence Index. Fable 5 held the benchmark at 64.9 since its June launch. Opus 5 now leads across SWE-bench Verified, AA Intelligence Index, AA-Briefcase, and LiveBench Agentic Coding.
Anthropic Expands Claude Voice to All Three Model Families With Cross-App Actions
Claude voice mode previously ran only Haiku. The update lets users switch between Opus, Sonnet, and Haiku mid-conversation without losing context, and connects speech to actions in Gmail, Calendar, Slack, Canva, and Notion. Free tier stays on Haiku with one connected tool.
Situational Awareness Lost 67% in July. The AI Thesis Held. The Leverage Didn't.
Leopold Aschenbrenner's AI infrastructure hedge fund went up 439% in the first half of 2026, then lost 67% of that in a single month. The same stocks still represent a real bet on AI compute — but he was running four-to-one leverage when they dropped 35-47%.
Axe Compute Hits $3B in Contracted Blackwell Clusters as B300 Enterprise Demand Proves Out
Pittsburgh-based Axe Compute locked a new $1.5B five-year deal to deploy 9,200 NVIDIA B300 GPUs, pushing its 2026 signed contracts above $3 billion. $534M in customer prepayments arrives in 30 days, covering the bulk of GPU and infrastructure capex before a single cluster goes live.
Kimi K3 Ties Fable 5 on LiveBench Agentic Coding at 62.2% and Costs $0.35 vs $1.44 Per Task
The latest LiveBench scores place Moonshot AI's open-weight Kimi K3 and Anthropic's Fable 5 at exactly 62.2% on agentic coding, with Kimi costing $0.35 per successful task versus Fable 5's $1.44. Claude Opus 5 leads both at 65.2% at $0.70 per task.
Celeris-1 Hits 1,664 Tokens/Second: A Diffusion LLM Enters the Production API Tier
The second commercially available diffusion language model posts 158ms p50 response times at 75.9% MMLU-Pro accuracy — 5x Mercury 2 on speed, 12 points ahead on quality.
Fubon: Google Plans 12–15 Million TPU v9 Chips by 2028, Intel Foundry Required
The volume would approach Nvidia's projected 12.4 million GPU shipments in the same year. TPU v9 shifts to four compute dies per package, more than doubling capacity consumption from 2027 and forcing Google to split production between TSMC and Intel.
GPT-5.6 Sol Triples on ARC-AGI-3 Without a Weight Change — Two API Settings Did It
Retained reasoning and compaction moved Sol from 13.3% to 38.3% on the public task set while cutting output tokens sixfold. On the official ARC Prize leaderboard, Sol is at 7.8% and Opus 5 at 30.2%.
DeepSeek Plans 1GW Data Center in Inner Mongolia — Efficiency Lab Bets on Infrastructure
DeepSeek is building at least 1GW of AI compute capacity in Ulanqab, Inner Mongolia, combining its own facility with leased capacity from third-party operators. Partial operations by late 2027. The company has raised $7.4B from Tencent, CATL, and others to fund the shift.
Meta: Naive RL Fails for Code Optimization. Rebuilding the Timing Pipeline Lifted Qwen 2.5 7B by 13 Points.
A new Meta paper diagnoses why standard RL for code speed optimization almost always fails: runtime is a noisy, sparse signal, and small defects anywhere in the evaluation chain corrupt the training. A full pipeline rebuild took Qwen 2.5 7B from 18.0% to 31.3% on the top-50% speed threshold.
Nvidia Forms Open Secure AI Alliance: GLM 5.2 Cleaned Up What Fable 5 Refused to Touch
Nvidia led dozens of AI companies in forming OSAA to build open-source defensive cybersecurity AI. The immediate trigger: Fable 5 refused to assist with HuggingFace incident response, and Nvidia's quantized build of GLM 5.2 completed the task instead.
AMD Trains Its Own MoE on Instinct GPUs and Opens Every Checkpoint — But the Weights Aren't Commercially Free
AMD released Instella-MoE-16B-A3B, a 16B MoE with 2.8B active parameters trained end-to-end on MI300X and MI325X with ROCm. Every training stage, data mixture, and config is published. ResearchRAIL limits commercial weight use; the MIT-licensed training code is fully open.
South Korea's Two 700B+ Models Release 48 Hours Apart — Elimination Round Starts August 8
SK Telecom's A.X K2 (688B) and LG AI Research's K-EXAONE 2.0 (750B) both shipped under Apache 2.0 within 48 hours. A.X K2 leads on Korean and math; K-EXAONE 2.0 leads on long-context retrieval. A.X K2 scores 9.3 on BrowseComp. Government evaluation August 8-11 cuts the field from four to three.
Disney Drops GitHub Copilot for OpenAI Codex as Enterprise Coding Market Consolidates
Disney is pulling GitHub Copilot and several other AI coding tools from US engineering teams in August, standardizing on OpenAI Codex. The same week, Cursor removed cost transparency from its usage page ahead of an expected SpaceX acquisition close.
Explorative Models Train on the Best of K Guesses — A New Pretraining Axis Targets End-to-End Generation
A new pretraining paradigm generates K candidate outputs at each training step and learns only from the best match, eliminating mode blurring without requiring VAE or codec pipelines. Researchers call exploration a third pretraining axis alongside scale and architecture.