GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

1mo ago release

Grok 4.6 Lands in GitHub Copilot — SpaceXAI's 95.6% SWE-Bench Model Reaches Every VS Code Developer

SpaceXAI pushes Grok 4.6 into Copilot's model picker at the same $2/$6 per million pricing as Grok 4.5, but with a fundamentally different model underneath: 95.6% SWE-bench Verified, Intelligence Index 61, and a new entry on Arena's Code and Text leaderboards.

1mo ago funding

Databricks Closes $5B at $190B — Up 42% in Six Months, Coatue Leads

Databricks has closed a $5 billion strategic round at a $190 billion valuation, up from $134 billion in February 2026 and $62 billion in January 2025. Capital targets Lakebase, the Genie AI coworker, and Unity AI Gateway.

1mo ago benchmark

DeepSeek V4-Pro 0813 Hits 96.4% on Neutral SWE-Bench — Open Weights at #2, 59x Cheaper Than Claude Opus 5

Vals.ai's neutral bash harness puts DeepSeek V4-Pro-0813 at 96.4% on SWE-bench Verified, second overall and ahead of GPT-5.6 Sol. The MIT-licensed open model costs $0.022 per task against Claude Opus 5's $1.29.

1mo ago funding

Jeff Dean's Discovery Loop in Talks for $1B at $10B to Automate Science

Discovery Loop, the public benefit corporation Jeff Dean founded after 27 years at Google, is raising $1 billion at a $10 billion valuation. The startup's stated aim is to automate machine learning, science, and technology.

1mo ago benchmark

Terminal-Bench 3.0 Ships: Frontier Agents Drop From 84% to 34% as the Bar Resets

Scale Labs and the Harbor community ship Terminal-Bench 3.0 on July 23, replacing tasks frontier models were approaching 80% on with a harder problem set. GPT-5.6 Sol Max and Claude Fable 5 Max now lead at 34% — the new ceiling — while Grok 4.6 reaches 26% and Grok 4.5 falls to 15.7%.

1mo ago model

DeepSeek V4 Pro and Flash Pricing Jumps Up to 12x on August 16 as Peak/Off-Peak Rates Replace Flat Billing

DeepSeek restructures API pricing around peak and off-peak windows effective August 16, with V4-Pro cache-hit input rates rising more than 12x at peak. The change coincides with V4-Pro's official GA and lands while the model scores just one point above V4-Flash on the Artificial Analysis Intelligence Index.

1mo ago research

Google Audits 75 AI Research Papers and Finds Evidence Failures in Every System Tested

Google Cloud AI Research ran five autonomous research agents across 75 papers on five tasks and found every baseline had at least one systematic evidence failure — fabricated citations, scores that did not reproduce, or method claims that were not in the submitted code. ScientistOne proposes a Chain-of-Evidence framework to close the gap.

1mo ago release

MiniMax H3 Ships Open Weights: 33B Omni Video Model Challenges Closed API Leaders at $0.08/s

MiniMax releases H3 as open weights — a 33B diffusion transformer that accepts text, image, video, and audio inputs and outputs 4-15 second videos with native stereo audio. On Artificial Analysis Arena, it scores 1193 on I2V and 1237 on T2V, entering the leaderboard's top competitive tier.

1mo ago benchmark

Ramp July: Anthropic Reaches 43.5% of US Businesses as Fable 5 Stalls at 6% of Internal Tokens

Anthropic widened its business AI lead over OpenAI to 43.5% vs 39.7% in Ramp's July index — growing five times faster. But inside Anthropic's own token mix, Fable 5 accounts for just 6% of consumption one month after launch, and that share is not rising. Price is the constraint.

1mo ago release

L&T Deploys India's First 10,000-GPU Single NVLink Cluster for Together AI in Chennai

Larsen & Toubro has secured a $1.05B-$1.57B contract to deploy 10,000 NVIDIA B300 Blackwell Ultra GPUs as one interconnected NVLink fabric at Vyoma.AI's Chennai campus — India's first unified AI supercluster, not a distributed collection of servers.

1mo ago release

Google Takes Control of Crusoe's 1.8 GW Wyoming Campus — Project Jade Is Now Project Tembo

County records confirm Google has replaced Crusoe Energy as the development partner on a 716-acre Cheyenne campus with 42 natural gas turbines, fuel cells, and a $7B+ power investment targeting early 2028 service.

1mo ago release

Cerebras Runs GPT-5.6 Sol at 750 Tokens Per Second — 7x Faster Than Fable 5, No Quality Loss

OpenAI and Cerebras launch Ultrafast Mode, a new API tier powered by Cerebras wafer-scale silicon that delivers GPT-5.6 Sol at 750 output tokens per second, completing Humanity's Last Exam in 11 hours versus Fable 5's 78.

1mo ago release

DeepSeek Harness v0.1: MIT-Licensed Coding Agent With Swappable Plugins Enters the Claude Code Fight

DeepSeek open-sources Harness, a Codex and Claude Code rival built on a swappable plugin system called Cordis, as the lab simultaneously raises API cache prices 6x for agent-heavy workflows.

1mo ago release

Gemini 3.7 Flash: DeepSWE Jumps 16 Points to 65.3%, Priced at Half of What 3.6 Flash Cost at Launch

Google ships Gemini 3.7 Flash three weeks after 3.6, with FrontierCode up to 43.6%, DeepSWE at 65.3%, and an introductory price of $0.75/M input — 50% cheaper than 3.6 Flash was at launch.

1mo ago benchmark

OpenAI Enterprise: 64% of Output Tokens Now Run Agent Loops — Chat Has Become the Minority Workload

New data from OpenAI shows that as of June 2026, agentic AI — defined as Codex tokens — accounts for 64% of combined enterprise output. The same frontier models power both chat and agents; the difference is how the compute gets spent.