GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

19d ago release

IBM Releases Granite 4.2 8B: Dense Reasoning at $0.10 Input per Million Tokens, 12-Language Support

IBM's Granite 4.2 8B brings multi-step reasoning, code generation, and multilingual dialogue to 12 languages at $0.10/$0.15 per million tokens. The 131K-context dense model targets agentic workflows on a single inference host.

19d ago benchmark

OpenAI Tests GPT-6 Astra on 20 Recent Chrome V8 Vulnerabilities

ExploitBench uses 20 high-severity V8 flaws from 13 stable Chrome releases between June and August 2026. Astra beat GPT-5.6 Sol on arbitrary code execution while using fewer output tokens, sharpening the case for phased cybersecurity access.

19d ago funding

SpaceXAI Brings Colossus 2 Online at 946 MW With $35.8B in Memphis Compute

Colossus 2 became operational on September 3 with roughly 946 MW of IT power and 1.1 million H100-equivalent compute units. The build uses temporary gas turbines as a bridge to a 1.2 GW plant due in mid-2027, turning SpaceXAI into a merchant compute provider as much as a model lab.

20d ago research

Spotify's Portal Cut Claude Code Token Use by 90% — by Solving the I/O Problem

Spotify Engineering published findings from Portal, an internal proxy that intercepts Claude Code's filesystem reads and caches repeated calls. Token consumption dropped 90% per session. The finding reframes where AI coding agent cost actually comes from.

20d ago funding

Gimlet Labs Raises $300M at $3B to Route AI Inference Across Any Chip

Andreessen Horowitz leads a $300M Series B in Gimlet Labs, which claims the first multi-silicon inference cloud. The round values the company at $3B — six months after its $80M Series A. Total funding is now $392M.

20d ago model

OpenAI, Anthropic, and xAI All Went Down Simultaneously on September 3 — No Shared Cause Confirmed

Three frontier AI providers suffered overlapping outages on the same morning. xAI traced it to a Memphis compute center. OpenAI cited a routing error. Anthropic said nothing publicly. AWS, Azure, and Cloudflare all reported no issues.

20d ago release

GitHub HydraFusion Matches Opus 5 Quality on TerminalBench at 67% Lower Estimated Cost

GitHub Copilot's new research preview selects from models across multiple providers at runtime, using three execution patterns to hit frontier quality without always paying frontier prices. On TerminalBench 2.1, it beat Opus 5 by 4.9 points at 67% lower estimated cost.

20d ago research

Claude Formalized Fermat's Last Theorem in 11 Days — 13 Million Lines of Lean, 29,500 Proofs

Working largely autonomously, Claude produced the first end-to-end computer-checked proof of Fermat's Last Theorem. It wrote 13 million lines of Lean code and proved 29,500 intermediate theorems, five times the size of the entire Mathlib library.

20d ago research

16,893 Sessions: Claude Code Skips Web Search, Codex Always Does, Agents Disagree More Than They Agree

A controlled study of 16,893 coding agent sessions reveals sharply different tool-selection behaviors between Claude Code, Codex, and Cursor. Repo context overrides brand preference. Some tools get mentioned in every session and chosen in almost none.

20d ago benchmark

FrontierCode Diamond: Claude Opus 4.8 Merges 13.4% of PRs Under Real Maintainer Standards

Cognition's new benchmark measures whether AI code would survive actual code review, not just pass tests. The best model clears the bar on 1 in 7 tasks. The rest are much worse.

20d ago benchmark

LiveBench September 2026: Fable 5.1 Leads at 83.4, Muse Spark 1.3 Scores 81.6 at $0.22 per Task

September LiveBench rankings confirm Claude Fable 5.1 at the top with 83.4 overall, 0.4 points above Fable 5. Meta's Muse Spark 1.3 enters at 81.6 — matching GPT-5.6 Sol on agentic coding (64.1 vs 56.2) while costing $0.22 per task against Sol's $0.52.

20d ago benchmark

GPT-6 Astra Matches Fable 5 on Coding Agents but Costs 75% More Per Task Than Its Own Predecessor

Artificial Analysis benchmarks GPT-6 Astra across two indices: it ties Fable 5 on the Coding Agent Index at less than half the price, while landing at the same Intelligence Index score as GPT-5.6 Sol (61) at 75% higher per-task cost. The split verdict turns on which workload you are running.

21d ago research

Google Doubles Its AI Chip Rollout Cadence, Targeting Two Generations Per Year

Google's AI infrastructure chief says the company is shifting from a two-year design cycle to releasing two custom chips per year, with ambitions to move even faster. The change puts Google's internal silicon on a cadence that rivals NVIDIA's accelerated Blackwell roadmap.

21d ago benchmark

GPT-6 Astra Scores 99.9% on ARC-AGI-3, Leaving Every Competitor Below 31%

ARC Prize results show GPT-6 Astra effectively solved the newest iteration of the benchmark, used fewer actions than the median human on 96% of levels, and widened the gap between OpenAI and the rest of the frontier to a degree not seen on any prior AGI evaluation.

21d ago release

MBZUAI Releases K2 Horizon: Six Fully Open Models From 0.9B to 375B, First Complete Agentic Fleet

Abu Dhabi's Institute of Foundation Models at MBZUAI launches K2 Horizon, six models spanning wearable edge to enterprise deployment. Every model ships with weights, training code, intermediate checkpoints, and data recipes — the most comprehensive open release to date and the first fleet to expose the full agentic training lifecycle.