Live Feed
Grok 4.6 Lands in GitHub Copilot — SpaceXAI's 95.6% SWE-Bench Model Reaches Every VS Code Developer
SpaceXAI pushes Grok 4.6 into Copilot's model picker at the same $2/$6 per million pricing as Grok 4.5, but with a fundamentally different model underneath: 95.6% SWE-bench Verified, Intelligence Index 61, and a new entry on Arena's Code and Text leaderboards.
Databricks Closes $5B at $190B — Up 42% in Six Months, Coatue Leads
Databricks has closed a $5 billion strategic round at a $190 billion valuation, up from $134 billion in February 2026 and $62 billion in January 2025. Capital targets Lakebase, the Genie AI coworker, and Unity AI Gateway.
DeepSeek V4-Pro 0813 Hits 96.4% on Neutral SWE-Bench — Open Weights at #2, 59x Cheaper Than Claude Opus 5
Vals.ai's neutral bash harness puts DeepSeek V4-Pro-0813 at 96.4% on SWE-bench Verified, second overall and ahead of GPT-5.6 Sol. The MIT-licensed open model costs $0.022 per task against Claude Opus 5's $1.29.
Jeff Dean's Discovery Loop in Talks for $1B at $10B to Automate Science
Discovery Loop, the public benefit corporation Jeff Dean founded after 27 years at Google, is raising $1 billion at a $10 billion valuation. The startup's stated aim is to automate machine learning, science, and technology.
Terminal-Bench 3.0 Ships: Frontier Agents Drop From 84% to 34% as the Bar Resets
Scale Labs and the Harbor community ship Terminal-Bench 3.0 on July 23, replacing tasks frontier models were approaching 80% on with a harder problem set. GPT-5.6 Sol Max and Claude Fable 5 Max now lead at 34% — the new ceiling — while Grok 4.6 reaches 26% and Grok 4.5 falls to 15.7%.
DeepSeek V4 Pro and Flash Pricing Jumps Up to 12x on August 16 as Peak/Off-Peak Rates Replace Flat Billing
DeepSeek restructures API pricing around peak and off-peak windows effective August 16, with V4-Pro cache-hit input rates rising more than 12x at peak. The change coincides with V4-Pro's official GA and lands while the model scores just one point above V4-Flash on the Artificial Analysis Intelligence Index.
Google Audits 75 AI Research Papers and Finds Evidence Failures in Every System Tested
Google Cloud AI Research ran five autonomous research agents across 75 papers on five tasks and found every baseline had at least one systematic evidence failure — fabricated citations, scores that did not reproduce, or method claims that were not in the submitted code. ScientistOne proposes a Chain-of-Evidence framework to close the gap.
MiniMax H3 Ships Open Weights: 33B Omni Video Model Challenges Closed API Leaders at $0.08/s
MiniMax releases H3 as open weights — a 33B diffusion transformer that accepts text, image, video, and audio inputs and outputs 4-15 second videos with native stereo audio. On Artificial Analysis Arena, it scores 1193 on I2V and 1237 on T2V, entering the leaderboard's top competitive tier.
Ramp July: Anthropic Reaches 43.5% of US Businesses as Fable 5 Stalls at 6% of Internal Tokens
Anthropic widened its business AI lead over OpenAI to 43.5% vs 39.7% in Ramp's July index — growing five times faster. But inside Anthropic's own token mix, Fable 5 accounts for just 6% of consumption one month after launch, and that share is not rising. Price is the constraint.
L&T Deploys India's First 10,000-GPU Single NVLink Cluster for Together AI in Chennai
Larsen & Toubro has secured a $1.05B-$1.57B contract to deploy 10,000 NVIDIA B300 Blackwell Ultra GPUs as one interconnected NVLink fabric at Vyoma.AI's Chennai campus — India's first unified AI supercluster, not a distributed collection of servers.
Google Takes Control of Crusoe's 1.8 GW Wyoming Campus — Project Jade Is Now Project Tembo
County records confirm Google has replaced Crusoe Energy as the development partner on a 716-acre Cheyenne campus with 42 natural gas turbines, fuel cells, and a $7B+ power investment targeting early 2028 service.
Cerebras Runs GPT-5.6 Sol at 750 Tokens Per Second — 7x Faster Than Fable 5, No Quality Loss
OpenAI and Cerebras launch Ultrafast Mode, a new API tier powered by Cerebras wafer-scale silicon that delivers GPT-5.6 Sol at 750 output tokens per second, completing Humanity's Last Exam in 11 hours versus Fable 5's 78.
DeepSeek Harness v0.1: MIT-Licensed Coding Agent With Swappable Plugins Enters the Claude Code Fight
DeepSeek open-sources Harness, a Codex and Claude Code rival built on a swappable plugin system called Cordis, as the lab simultaneously raises API cache prices 6x for agent-heavy workflows.
Gemini 3.7 Flash: DeepSWE Jumps 16 Points to 65.3%, Priced at Half of What 3.6 Flash Cost at Launch
Google ships Gemini 3.7 Flash three weeks after 3.6, with FrontierCode up to 43.6%, DeepSWE at 65.3%, and an introductory price of $0.75/M input — 50% cheaper than 3.6 Flash was at launch.
OpenAI Enterprise: 64% of Output Tokens Now Run Agent Loops — Chat Has Become the Minority Workload
New data from OpenAI shows that as of June 2026, agentic AI — defined as Codex tokens — accounts for 64% of combined enterprise output. The same frontier models power both chat and agents; the difference is how the compute gets spent.