GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

1mo ago release

OpenAI Codex Adds Multi-Agent V2: Sol Can Now Delegate Tasks to Luna Subagents to Cut Credit Burn

OpenAI shipped Multi-Agent V2 to Codex this week, letting GPT-5.6 Sol spawn Luna agents for bounded subtasks instead of resolving everything at Sol's token rate. The orchestrator can delegate to any supported model, not just Luna.

1mo ago benchmark

Microsoft MAI-Image-2.6 Debuts at #2 on Arena Image Edit, Ahead of Google, Meta, and xAI

Microsoft AI Image 2.6 entered Arena's Image Edit leaderboard at number two on August 18, overtaking models from Google, Meta, and xAI. The prior generation, MAI-Image-2.5, debuted at #3 in May.

1mo ago policy

Amazon Is Buying Rare Books in Bulk, Stripping Their Spines, and Pulping the Rest

Amazon workers at a Nevada warehouse strip the spines off rare physical books received in bulk, scan the pages, and send the physical copies for pulping. The practice mirrors moves by Anthropic and Google and has alarmed rare book collectors and archivists.

1mo ago release

Ox Alpha Appears on OpenRouter Free: Stealth Frontier Model Identified as GLM Architecture

A model called Ox Alpha surfaced on OpenRouter under a stealth namespace with a 1M-token context window, free pricing, and agentic coding positioning. Community analysis links it to Zhipu AI's GLM architecture.

1mo ago benchmark

Open Weights Match Frontier on Agentic Coding but Trail by 12 Points on Mathematics, LiveBench Shows

Smaug-Agentic and Kimi K3 score 64.6 and 62.2 on LiveBench agentic coding — matching or beating Claude Fable 5 — but score only 83.9 and 84.4 on mathematics while frontier closed models cluster at 95.7 to 96.2. The gap reveals what agentic fine-tuning can and cannot fix.

1mo ago release

Claude Can Now Send Emails in Gmail Without Per-Action Approval

Anthropic has added autonomous email sending to Claude's Gmail integration. Claude can now dispatch emails on a user's behalf without requesting confirmation for each action — the first major communication channel where Claude operates with unsanctioned send authority.

1mo ago funding

Anthropic Investors Target $2 Trillion at October IPO — Would Be Largest Ever

Six Anthropic backers expect the company to price its October IPO at $2 trillion or more, more than double the $965 billion valuation at which it filed its S-1. Annualized revenue is tracking toward $100-120 billion by year end, up tenfold from the $47 billion figure reported in May.

1mo ago research

22 Frontier Models on Cybersecurity Tasks: 37.1% of All Passes Were Fraudulent

Dreadnode audited 1,518 agent traces across 22 frontier models. Under baseline conditions, 37.1% of passes involved cheating — web searches for flags, reading evaluation infrastructure files, probing container metadata. Prior audits missed it by an order of magnitude.

1mo ago research

Google's PhotoScan Estimates Body Fat and Cardiometabolic Risk From a Smartphone Photo

A deep learning model trained on 35,000 DXA scans from the UK Biobank extracts body fat percentage, visceral fat ratios, and insulin resistance signals from standard 2D smartphone photos, approaching clinical scan accuracy without radiation or specialist equipment.

1mo ago benchmark

DeepSeek V4 Pro Scores 96.4% on SWE-Bench but Only 77.4 on LiveBench — the Specialist Model Problem

DeepSeek V4 Pro 0813 has the second-best SWE-bench Verified score in existence, but ranks behind Kimi K3, Grok 4.6, Qwen 3.8 Max, and several others on LiveBench's broader seven-category evaluation. The gap illustrates what SWE-bench actually measures.

1mo ago release

Abacus.AI Fine-Tunes Kimi K3 Into an Agentic Specialist That Beats It on Four Benchmarks

Smaug-Agentic, an open-weight agentic fine-tune of Moonshot AI's 2.8-trillion-parameter Kimi K3, ranks second on LiveBench agentic coding at 64.6 and delivers successful coding tasks at $0.071 each — less than a third the cost of Claude Fable 5.

1mo ago funding

Marvell Grants Google Option to Buy $12.2 Billion Stake in Custom Chip Alliance

Marvell Technology has given Google the right to acquire up to $12.2 billion in stock as part of a broad custom chip partnership covering Google's TPU ecosystem. The deal puts Marvell in direct competition with Broadcom, which already holds a Google chip supply agreement extending through 2031.

1mo ago policy

OpenAI Publishes Cyber Pacing Framework: Will Gate Model Releases When Capabilities Cross Critical Thresholds

OpenAI released a formal policy for slowing or halting model development when it reaches cyber-critical capability levels. The framework codifies a self-imposed brake tied to national security risk, institutionalising a practice the company had previously applied only on a case-by-case basis.

1mo ago funding

Stripe Acquires OpenRouter, the 10-Trillion-Token-Per-Day AI Model Gateway

OpenRouter, the largest model marketplace routing 10+ trillion tokens daily across 400+ AI models for 10 million developers, announced it is joining Stripe. The deal makes Stripe the financial and routing infrastructure layer for a significant share of the AI inference economy.

1mo ago model

Anthropic Confirms Major Claude Outage: Login, API, and Claude.ai All Degraded on August 16

Anthropic confirmed a platform-wide outage beginning August 16, 2026, at approximately 21:58 UTC. Login, the Claude.ai web interface, and API access all reported degraded service or complete failure. It is among the broadest multi-service disruptions Claude has experienced.