Live Feed
DeepSeek Puts Three V4 Variants on Arena Simultaneously — 80.6% SWE-Bench Puts Pro at the Frontier Tier
DeepSeek V4-Pro, V4-Pro-Thinking, and V4-Flash-Thinking all entered Chatbot Arena's Text and Code leaderboards on April 23. Independent SWE-Bench Verified results for V4 Pro sit at 80.6% — within range of Gemini 3.1 Pro Preview and Kimi K2.6.
OpenAI Codex + GPT-5.5 Takes Terminal-Bench 2.0 at 82.0% — First Time OpenAI's Own Agent Leads
GPT-5.5 paired with OpenAI's Codex agent posts 82.0% on Terminal-Bench 2.0, edging out ForgeCode's 81.8% with GPT-5.4. The 0.2-point margin is notable for a different reason: OpenAI's own scaffold finally beats the third-party wrapper that held the top spot.
NVIDIA Joins $30B Vast Data Round — The GPU Maker Moves Into AI Storage
NVIDIA has taken a stake in Vast Data's latest funding round, valuing the AI storage company at $30 billion. With $4 billion in bookings and positive free cash flow, Vast is one of the few AI infrastructure players that isn't burning toward profitability — and NVIDIA's investment signals where the next layer of the stack is being contested.
DeepSeek V4 Flash Ships at $0.14/M: 79% SWE-Bench Verified on 13B Active Parameters
DeepSeek's efficient V4 variant goes live today with 284B total / 13B active MoE parameters, scoring 79.0% on SWE-Bench Verified at 18x less per token than GPT-5.4. The full V4 family — Pro and Flash — is now on API with 1M standard context and MIT-licensed weights.
DeepSeek-V4-Pro: Open-Source Coding SOTA, 1.6T Params, 1M Context, MIT License
DeepSeek released V4-Pro — 1.6T parameters, 49B activated, 1M context — claiming best open-source model on coding and reasoning. LiveCodeBench 93.5% beats all frontier models. Terminal-Bench 67.9%. Weights are MIT licensed and on HuggingFace now.
Perplexity Research: Post-Training, Not Base Model Choice, Determines Search Quality
Perplexity's published SFT + RL pipeline turns open Qwen models into search tools that match or beat GPT-5.4 on factual benchmarks at lower cost. The core claim is that search quality is a function of how you tune, not what you start with.
GPT-5.5 Leads Intelligence Index at 60 — and Hallucinates at 86%, the Highest Rate at the Frontier
Artificial Analysis' full benchmark sweep of GPT-5.5 reveals a split performance profile: the model sets the highest-ever AA-Omniscience accuracy at 57% while posting an 86% hallucination rate — more than double Claude Opus 4.7 max at 36%. The cost story is more favourable.
80% of Claude's Weekly Users Earn $100K+. Only 37% of Meta AI's Do.
An Epoch AI/Ipsos survey of 5,000 US adults reveals a sharp income stratification in AI adoption. Claude is the most income-concentrated major AI platform by a significant margin. Meta AI is running the opposite strategy entirely.
Anthropic Postmortem: Three Overlapping Claude Code Changes Caused Six Weeks of Quality Degradation
Anthropic published an engineering postmortem tracing Claude Code quality complaints to three separate changes made between March 4 and April 16. All were reversed by April 20. The API was never affected. Usage limits reset for all subscribers on April 23.
OpenAI Launches GPT-5.5: 82.7% Terminal-Bench, 58.6% SWE-Bench Pro, Cheaper Per Codex Task
OpenAI released GPT-5.5, codenamed Spud, to Plus, Pro, Business, and Enterprise users in ChatGPT and Codex. It sets a new Terminal-Bench 2.0 record at 82.7%, matches GPT-5.4 per-token latency, and uses fewer tokens per Codex task. API access pending safety review.
Sony's Ace Robot Beats Elite Table Tennis Players Under Official ITTF Rules — Published in Nature
Sony AI published a Nature cover paper on April 23 describing Ace, the first robot to defeat elite and professional human table tennis players under official tournament rules. Nine active-pixel cameras plus three event-based gaze systems give it 20.2ms end-to-end latency. Trained entirely in simulation via model-free reinforcement learning, it transferred directly to physical competition.
SpaceX's $60B Cursor Option Is Not a Purchase. It Is a Wager on Who Stays Alive.
SpaceX locked in a one-year option to acquire Cursor for $60 billion, or pay $10 billion for the joint work. The number is the headline. The logic underneath it is about something colder: every lab Cursor depends on to run its product is now shipping a direct competitor.
SeedDance 2.0 Holds Both Arena Video #1 Spots at Elo 1450 — 79 Points Ahead of Google Veo 3.1
ByteDance's Dreamina SeedDance 2.0 720p ranks #1 on both Arena's Text-to-Video and Image-to-Video leaderboards, with Elo 1450 and 1449 respectively. It leads Google Veo 3.1 audio 1080p by 79 points on T2V — and does it at 720p.
DeepSeek First External Raise Doubles to $20B in Five Days as Tencent and Alibaba Move In
DeepSeek is in active talks with Tencent and Alibaba at a valuation exceeding $20 billion — double the $10 billion floor it was seeking less than a week ago. The round would be the company's first external funding since its 2023 founding, targeting at least $300 million.
Google Cloud Commits $750M to Accelerate Agentic AI Across Its 120,000-Partner Ecosystem
At Cloud Next in Las Vegas, Google Cloud announced a $750 million fund for its global partner network — covering prototyping, deployment, upskilling, and enterprise agent development. Named enterprise-ready agents from Adobe, Atlassian, Salesforce, Workday and others ship inside Gemini Enterprise on day one.