GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%

Live Feed

1mo ago benchmark

Claude Opus 5 Takes Artificial Analysis Intelligence Index #1, Fable 5 Falls to Second

Anthropic's Opus 5 has dethroned Fable 5 at the top of Artificial Analysis's Intelligence Index. Fable 5 held the benchmark at 64.9 since its June launch. Opus 5 now leads across SWE-bench Verified, AA Intelligence Index, AA-Briefcase, and LiveBench Agentic Coding.

1mo ago release

Anthropic Expands Claude Voice to All Three Model Families With Cross-App Actions

Claude voice mode previously ran only Haiku. The update lets users switch between Opus, Sonnet, and Haiku mid-conversation without losing context, and connects speech to actions in Gmail, Calendar, Slack, Canva, and Notion. Free tier stays on Haiku with one connected tool.

1mo ago funding

Situational Awareness Lost 67% in July. The AI Thesis Held. The Leverage Didn't.

Leopold Aschenbrenner's AI infrastructure hedge fund went up 439% in the first half of 2026, then lost 67% of that in a single month. The same stocks still represent a real bet on AI compute — but he was running four-to-one leverage when they dropped 35-47%.

1mo ago funding

Axe Compute Hits $3B in Contracted Blackwell Clusters as B300 Enterprise Demand Proves Out

Pittsburgh-based Axe Compute locked a new $1.5B five-year deal to deploy 9,200 NVIDIA B300 GPUs, pushing its 2026 signed contracts above $3 billion. $534M in customer prepayments arrives in 30 days, covering the bulk of GPU and infrastructure capex before a single cluster goes live.

1mo ago benchmark

Kimi K3 Ties Fable 5 on LiveBench Agentic Coding at 62.2% and Costs $0.35 vs $1.44 Per Task

The latest LiveBench scores place Moonshot AI's open-weight Kimi K3 and Anthropic's Fable 5 at exactly 62.2% on agentic coding, with Kimi costing $0.35 per successful task versus Fable 5's $1.44. Claude Opus 5 leads both at 65.2% at $0.70 per task.

1mo ago model

Celeris-1 Hits 1,664 Tokens/Second: A Diffusion LLM Enters the Production API Tier

The second commercially available diffusion language model posts 158ms p50 response times at 75.9% MMLU-Pro accuracy — 5x Mercury 2 on speed, 12 points ahead on quality.

1mo ago research

Fubon: Google Plans 12–15 Million TPU v9 Chips by 2028, Intel Foundry Required

The volume would approach Nvidia's projected 12.4 million GPU shipments in the same year. TPU v9 shifts to four compute dies per package, more than doubling capacity consumption from 2027 and forcing Google to split production between TSMC and Intel.

1mo ago benchmark

GPT-5.6 Sol Triples on ARC-AGI-3 Without a Weight Change — Two API Settings Did It

Retained reasoning and compaction moved Sol from 13.3% to 38.3% on the public task set while cutting output tokens sixfold. On the official ARC Prize leaderboard, Sol is at 7.8% and Opus 5 at 30.2%.

1mo ago funding

DeepSeek Plans 1GW Data Center in Inner Mongolia — Efficiency Lab Bets on Infrastructure

DeepSeek is building at least 1GW of AI compute capacity in Ulanqab, Inner Mongolia, combining its own facility with leased capacity from third-party operators. Partial operations by late 2027. The company has raised $7.4B from Tencent, CATL, and others to fund the shift.

1mo ago research

Meta: Naive RL Fails for Code Optimization. Rebuilding the Timing Pipeline Lifted Qwen 2.5 7B by 13 Points.

A new Meta paper diagnoses why standard RL for code speed optimization almost always fails: runtime is a noisy, sparse signal, and small defects anywhere in the evaluation chain corrupt the training. A full pipeline rebuild took Qwen 2.5 7B from 18.0% to 31.3% on the top-50% speed threshold.

1mo ago policy

Nvidia Forms Open Secure AI Alliance: GLM 5.2 Cleaned Up What Fable 5 Refused to Touch

Nvidia led dozens of AI companies in forming OSAA to build open-source defensive cybersecurity AI. The immediate trigger: Fable 5 refused to assist with HuggingFace incident response, and Nvidia's quantized build of GLM 5.2 completed the task instead.

1mo ago release

AMD Trains Its Own MoE on Instinct GPUs and Opens Every Checkpoint — But the Weights Aren't Commercially Free

AMD released Instella-MoE-16B-A3B, a 16B MoE with 2.8B active parameters trained end-to-end on MI300X and MI325X with ROCm. Every training stage, data mixture, and config is published. ResearchRAIL limits commercial weight use; the MIT-licensed training code is fully open.

1mo ago release

South Korea's Two 700B+ Models Release 48 Hours Apart — Elimination Round Starts August 8

SK Telecom's A.X K2 (688B) and LG AI Research's K-EXAONE 2.0 (750B) both shipped under Apache 2.0 within 48 hours. A.X K2 leads on Korean and math; K-EXAONE 2.0 leads on long-context retrieval. A.X K2 scores 9.3 on BrowseComp. Government evaluation August 8-11 cuts the field from four to three.

1mo ago release

Disney Drops GitHub Copilot for OpenAI Codex as Enterprise Coding Market Consolidates

Disney is pulling GitHub Copilot and several other AI coding tools from US engineering teams in August, standardizing on OpenAI Codex. The same week, Cursor removed cost transparency from its usage page ahead of an expected SpaceX acquisition close.

1mo ago research

Explorative Models Train on the Best of K Guesses — A New Pretraining Axis Targets End-to-End Generation

A new pretraining paradigm generates K candidate outputs at each training step and learns only from the best match, eliminating mode blurring without requiring VAE or codec pipelines. Researchers call exploration a third pretraining axis alongside scale and architecture.