GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —

Live Feed

5mo ago benchmark

Claude Opus 4.7 Leads Tool Use by 9 Points, Trails GPT-5.4 on Web Search — The Benchmark Split That Matters for Agent Selection

Full benchmark data for Claude Opus 4.7 shows the largest jump was MCP-Atlas (+14.6 to 77.3%), which measures multi-step tool orchestration — where it now leads GPT-5.4 by 9.2 points. The regression is BrowseComp, down 4.7 points to 79.3% while GPT-5.4 leads at 89.3%. Where Anthropic and OpenAI differ is now clearly visible in the numbers.

5mo ago release

Alibaba Ships Qwen3.6-35B-A3B Free: 3B Active Params Score 73.4% SWE-Bench, Beat Gemma 4 by 21 Points

Alibaba releases Qwen3.6-35B-A3B under Apache 2.0 — a sparse MoE with only 3 billion active parameters per forward pass that reaches 73.4% on SWE-bench Verified, 21 points ahead of Google Gemma 4 26B A4B despite activating fewer parameters per token.

5mo ago funding

Sequoia Raises $7B AI Fund — Double Its Last, First Under New Leadership

Sequoia Capital has raised approximately $7 billion for a new expansion fund, nearly double its 2022 comparable vehicle. It is the first major raise under co-stewards Alfred Lin and Pat Grady, and a direct signal that late-stage AI investing has fundamentally repriced.

5mo ago policy

UK Sovereign AI Unit Opens: £500M Government VC Fund, 1M GPU Hours Per Startup, Visas in 24 Hours

The UK government launched its Sovereign AI Unit on April 16 — a £500M state-backed VC fund structured to operate at the pace of industry. First equity investment goes to Callosum. Six more startups get supercomputer access. Chaired by Balderton partner James Wise, the unit is already in discussions with 30 additional companies.

5mo ago funding

Accel Raises $5B on Anthropic 4x Returns — AI's Growth Stage Has Officially Arrived

Accel closes a $5B vehicle targeting 20–25 late-stage AI investments at an average cheque of $200M, riding paper gains on Anthropic (invested at $183B, now ~$800B) and Cursor (backed at $9.9B, now ~$50B). It lands in a Q1 2026 that deployed a record $297B in global venture capital.

5mo ago funding

Allbirds Dumps Shoes, Becomes NewBird AI — Stock Up 700% on $50M Compute Pivot

The $4B sustainable footwear brand, reduced to a $21M shell, sold its IP for $39M and raised $50M to lease AI compute hardware. The most extreme AI pivot yet is also the most revealing data point on where compute demand sits right now.

5mo ago funding

Anthropic Draws $800B Investor Offers as OpenAI Disputes the Revenue Behind the Number

Investor offers have valued Anthropic at up to $800 billion — more than double its February raise — while OpenAI's chief revenue officer circulated an internal memo accusing Anthropic of inflating its $30 billion run rate by $8 billion through gross accounting practices.

5mo ago benchmark

100 Teams, 300 Humanoids: Beijing's Robotic Half-Marathon Goes Autonomous on April 19

The second annual Beijing E-Town Humanoid Robot Half-Marathon runs April 19 with 100+ competing teams — five times last year's field — and nearly 40% of entrants operating under full autonomous navigation. Remote-controlled robots face a 1.2× time penalty designed to make human supervision economically uncompetitive.

5mo ago benchmark

Arena Opens Image-to-WebDev Leaderboard; Kimi K2.5 and Grok 4.20 Beta Enter Document Rankings

Chatbot Arena launched its Image-to-WebDev evaluation on April 15, testing models on visual-to-code conversion. The Document Arena simultaneously added Moonshot's multimodal Kimi K2.5 and xAI's Grok 4.20 Beta — expanding the most credible comparative view of how frontier models handle long-form document tasks.

5mo ago benchmark

AI Coding Agents Ship Insecure Code 87% of the Time — Even When All Tests Pass

Endor Labs measured 200 real tasks across 108 open-source projects and found that 87% of AI-generated code contains at least one exploitable vulnerability, regardless of whether functional tests pass. The best security performer, OpenAI Codex with GPT-5.4, scored just 17.3% on security correctness.

5mo ago funding

Fluidstack Raises $1B at $18B Valuation — AI Data Center Demand Triples Its Value in Four Months

The specialized AI data center startup jumped from a $7.5B valuation in December 2025 to $18B, making it one of the fastest-appreciating infrastructure bets in tech. The $1B round reflects mounting hyperscaler capacity constraints with no relief in sight.

5mo ago release

Google Ships Gemini 3.1 Flash TTS: Audio Tags, 70+ Languages, 1,211 Elo on AA Leaderboard

Google DeepMind has launched Gemini 3.1 Flash TTS in preview, bringing director-level voice control through inline audio tags, native multi-speaker dialogue, and 70+ language coverage. It ranks second on the Artificial Analysis TTS leaderboard at Elo 1,211 — beating ElevenLabs v3 — at $1.00/M text tokens and $20.00/M audio tokens.

5mo ago release

Google Launches Native Gemini App for macOS — The Last Major AI Lab on Desktop

Google has shipped a 100% native Swift Gemini app for macOS 15+, with an Option+Space hotkey, live screen sharing, and access to Gemini 3 Fast, Thinking, and Pro models. ChatGPT and Claude Mac apps arrived roughly a year ago.

5mo ago funding

Jane Street Commits $7B to CoreWeave — Wall Street's Biggest Quant Shop Is Now a Frontier Lab Client

Jane Street signs a $6B AI cloud agreement with CoreWeave and takes a $1B equity stake at $109 per share, making the quantitative trading firm the fifth-largest shareholder in the AI cloud provider. The deal includes access to next-gen NVIDIA Vera Rubin compute and marks CoreWeave's third major contract announcement in seven days.

5mo ago release

OpenAI Launches GPT-5.4-Cyber With Thousands of Vetted Defenders — Directly Contrasting Anthropic's 40-Partner Glasswing

GPT-5.4-Cyber, a fine-tuned variant of OpenAI's flagship model with reduced refusal boundaries for security work, launches one week after Anthropic's Mythos announcement. BNY, Citi, CrowdStrike, NVIDIA, Oracle, Cisco, and Zscaler are among day-one partners. $10M cybersecurity grant program attached.