GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —

Live Feed

5mo ago benchmark

DeepSeek Puts Three V4 Variants on Arena Simultaneously — 80.6% SWE-Bench Puts Pro at the Frontier Tier

DeepSeek V4-Pro, V4-Pro-Thinking, and V4-Flash-Thinking all entered Chatbot Arena's Text and Code leaderboards on April 23. Independent SWE-Bench Verified results for V4 Pro sit at 80.6% — within range of Gemini 3.1 Pro Preview and Kimi K2.6.

5mo ago benchmark

OpenAI Codex + GPT-5.5 Takes Terminal-Bench 2.0 at 82.0% — First Time OpenAI's Own Agent Leads

GPT-5.5 paired with OpenAI's Codex agent posts 82.0% on Terminal-Bench 2.0, edging out ForgeCode's 81.8% with GPT-5.4. The 0.2-point margin is notable for a different reason: OpenAI's own scaffold finally beats the third-party wrapper that held the top spot.

5mo ago funding

NVIDIA Joins $30B Vast Data Round — The GPU Maker Moves Into AI Storage

NVIDIA has taken a stake in Vast Data's latest funding round, valuing the AI storage company at $30 billion. With $4 billion in bookings and positive free cash flow, Vast is one of the few AI infrastructure players that isn't burning toward profitability — and NVIDIA's investment signals where the next layer of the stack is being contested.

5mo ago release

DeepSeek V4 Flash Ships at $0.14/M: 79% SWE-Bench Verified on 13B Active Parameters

DeepSeek's efficient V4 variant goes live today with 284B total / 13B active MoE parameters, scoring 79.0% on SWE-Bench Verified at 18x less per token than GPT-5.4. The full V4 family — Pro and Flash — is now on API with 1M standard context and MIT-licensed weights.

5mo ago release

DeepSeek-V4-Pro: Open-Source Coding SOTA, 1.6T Params, 1M Context, MIT License

DeepSeek released V4-Pro — 1.6T parameters, 49B activated, 1M context — claiming best open-source model on coding and reasoning. LiveCodeBench 93.5% beats all frontier models. Terminal-Bench 67.9%. Weights are MIT licensed and on HuggingFace now.

5mo ago research

Perplexity Research: Post-Training, Not Base Model Choice, Determines Search Quality

Perplexity's published SFT + RL pipeline turns open Qwen models into search tools that match or beat GPT-5.4 on factual benchmarks at lower cost. The core claim is that search quality is a function of how you tune, not what you start with.

5mo ago benchmark

GPT-5.5 Leads Intelligence Index at 60 — and Hallucinates at 86%, the Highest Rate at the Frontier

Artificial Analysis' full benchmark sweep of GPT-5.5 reveals a split performance profile: the model sets the highest-ever AA-Omniscience accuracy at 57% while posting an 86% hallucination rate — more than double Claude Opus 4.7 max at 36%. The cost story is more favourable.

5mo ago research

80% of Claude's Weekly Users Earn $100K+. Only 37% of Meta AI's Do.

An Epoch AI/Ipsos survey of 5,000 US adults reveals a sharp income stratification in AI adoption. Claude is the most income-concentrated major AI platform by a significant margin. Meta AI is running the opposite strategy entirely.

5mo ago model

Anthropic Postmortem: Three Overlapping Claude Code Changes Caused Six Weeks of Quality Degradation

Anthropic published an engineering postmortem tracing Claude Code quality complaints to three separate changes made between March 4 and April 16. All were reversed by April 20. The API was never affected. Usage limits reset for all subscribers on April 23.

5mo ago release

OpenAI Launches GPT-5.5: 82.7% Terminal-Bench, 58.6% SWE-Bench Pro, Cheaper Per Codex Task

OpenAI released GPT-5.5, codenamed Spud, to Plus, Pro, Business, and Enterprise users in ChatGPT and Codex. It sets a new Terminal-Bench 2.0 record at 82.7%, matches GPT-5.4 per-token latency, and uses fewer tokens per Codex task. API access pending safety review.

5mo ago research

Sony's Ace Robot Beats Elite Table Tennis Players Under Official ITTF Rules — Published in Nature

Sony AI published a Nature cover paper on April 23 describing Ace, the first robot to defeat elite and professional human table tennis players under official tournament rules. Nine active-pixel cameras plus three event-based gaze systems give it 20.2ms end-to-end latency. Trained entirely in simulation via model-free reinforcement learning, it transferred directly to physical competition.

5mo ago funding

SpaceX's $60B Cursor Option Is Not a Purchase. It Is a Wager on Who Stays Alive.

SpaceX locked in a one-year option to acquire Cursor for $60 billion, or pay $10 billion for the joint work. The number is the headline. The logic underneath it is about something colder: every lab Cursor depends on to run its product is now shipping a direct competitor.

5mo ago benchmark

SeedDance 2.0 Holds Both Arena Video #1 Spots at Elo 1450 — 79 Points Ahead of Google Veo 3.1

ByteDance's Dreamina SeedDance 2.0 720p ranks #1 on both Arena's Text-to-Video and Image-to-Video leaderboards, with Elo 1450 and 1449 respectively. It leads Google Veo 3.1 audio 1080p by 79 points on T2V — and does it at 720p.

5mo ago funding

DeepSeek First External Raise Doubles to $20B in Five Days as Tencent and Alibaba Move In

DeepSeek is in active talks with Tencent and Alibaba at a valuation exceeding $20 billion — double the $10 billion floor it was seeking less than a week ago. The round would be the company's first external funding since its 2023 founding, targeting at least $300 million.

5mo ago funding

Google Cloud Commits $750M to Accelerate Agentic AI Across Its 120,000-Partner Ecosystem

At Cloud Next in Las Vegas, Google Cloud announced a $750 million fund for its global partner network — covering prototyping, deployment, upskilling, and enterprise agent development. Named enterprise-ready agents from Adobe, Atlassian, Salesforce, Workday and others ship inside Gemini Enterprise on day one.