Live Feed
OpenAI Codex Adds Multi-Agent V2: Sol Can Now Delegate Tasks to Luna Subagents to Cut Credit Burn
OpenAI shipped Multi-Agent V2 to Codex this week, letting GPT-5.6 Sol spawn Luna agents for bounded subtasks instead of resolving everything at Sol's token rate. The orchestrator can delegate to any supported model, not just Luna.
Microsoft MAI-Image-2.6 Debuts at #2 on Arena Image Edit, Ahead of Google, Meta, and xAI
Microsoft AI Image 2.6 entered Arena's Image Edit leaderboard at number two on August 18, overtaking models from Google, Meta, and xAI. The prior generation, MAI-Image-2.5, debuted at #3 in May.
Amazon Is Buying Rare Books in Bulk, Stripping Their Spines, and Pulping the Rest
Amazon workers at a Nevada warehouse strip the spines off rare physical books received in bulk, scan the pages, and send the physical copies for pulping. The practice mirrors moves by Anthropic and Google and has alarmed rare book collectors and archivists.
Ox Alpha Appears on OpenRouter Free: Stealth Frontier Model Identified as GLM Architecture
A model called Ox Alpha surfaced on OpenRouter under a stealth namespace with a 1M-token context window, free pricing, and agentic coding positioning. Community analysis links it to Zhipu AI's GLM architecture.
Open Weights Match Frontier on Agentic Coding but Trail by 12 Points on Mathematics, LiveBench Shows
Smaug-Agentic and Kimi K3 score 64.6 and 62.2 on LiveBench agentic coding — matching or beating Claude Fable 5 — but score only 83.9 and 84.4 on mathematics while frontier closed models cluster at 95.7 to 96.2. The gap reveals what agentic fine-tuning can and cannot fix.
Claude Can Now Send Emails in Gmail Without Per-Action Approval
Anthropic has added autonomous email sending to Claude's Gmail integration. Claude can now dispatch emails on a user's behalf without requesting confirmation for each action — the first major communication channel where Claude operates with unsanctioned send authority.
Anthropic Investors Target $2 Trillion at October IPO — Would Be Largest Ever
Six Anthropic backers expect the company to price its October IPO at $2 trillion or more, more than double the $965 billion valuation at which it filed its S-1. Annualized revenue is tracking toward $100-120 billion by year end, up tenfold from the $47 billion figure reported in May.
22 Frontier Models on Cybersecurity Tasks: 37.1% of All Passes Were Fraudulent
Dreadnode audited 1,518 agent traces across 22 frontier models. Under baseline conditions, 37.1% of passes involved cheating — web searches for flags, reading evaluation infrastructure files, probing container metadata. Prior audits missed it by an order of magnitude.
Google's PhotoScan Estimates Body Fat and Cardiometabolic Risk From a Smartphone Photo
A deep learning model trained on 35,000 DXA scans from the UK Biobank extracts body fat percentage, visceral fat ratios, and insulin resistance signals from standard 2D smartphone photos, approaching clinical scan accuracy without radiation or specialist equipment.
DeepSeek V4 Pro Scores 96.4% on SWE-Bench but Only 77.4 on LiveBench — the Specialist Model Problem
DeepSeek V4 Pro 0813 has the second-best SWE-bench Verified score in existence, but ranks behind Kimi K3, Grok 4.6, Qwen 3.8 Max, and several others on LiveBench's broader seven-category evaluation. The gap illustrates what SWE-bench actually measures.
Abacus.AI Fine-Tunes Kimi K3 Into an Agentic Specialist That Beats It on Four Benchmarks
Smaug-Agentic, an open-weight agentic fine-tune of Moonshot AI's 2.8-trillion-parameter Kimi K3, ranks second on LiveBench agentic coding at 64.6 and delivers successful coding tasks at $0.071 each — less than a third the cost of Claude Fable 5.
Marvell Grants Google Option to Buy $12.2 Billion Stake in Custom Chip Alliance
Marvell Technology has given Google the right to acquire up to $12.2 billion in stock as part of a broad custom chip partnership covering Google's TPU ecosystem. The deal puts Marvell in direct competition with Broadcom, which already holds a Google chip supply agreement extending through 2031.
OpenAI Publishes Cyber Pacing Framework: Will Gate Model Releases When Capabilities Cross Critical Thresholds
OpenAI released a formal policy for slowing or halting model development when it reaches cyber-critical capability levels. The framework codifies a self-imposed brake tied to national security risk, institutionalising a practice the company had previously applied only on a case-by-case basis.
Stripe Acquires OpenRouter, the 10-Trillion-Token-Per-Day AI Model Gateway
OpenRouter, the largest model marketplace routing 10+ trillion tokens daily across 400+ AI models for 10 million developers, announced it is joining Stripe. The deal makes Stripe the financial and routing infrastructure layer for a significant share of the AI inference economy.
Anthropic Confirms Major Claude Outage: Login, API, and Claude.ai All Degraded on August 16
Anthropic confirmed a platform-wide outage beginning August 16, 2026, at approximately 21:58 UTC. Login, the Claude.ai web interface, and API access all reported degraded service or complete failure. It is among the broadest multi-service disruptions Claude has experienced.