GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

3mo ago policy

WIRED: Meta Silently Embedded Face-Recognition Code in Millions of Phones for Its Smart Glasses

WIRED found unreleased facial recognition code built into Meta's smart glasses platform — internally called 'Name Tag' — already deployed in apps on millions of phones. Meta previously paid $7B+ in biometric privacy settlements and explicitly planned to launch during political turmoil.

3mo ago research

LEAP Solves All 12 Putnam 2025 Problems in Lean — General LLMs Beat Specialist Provers With Agentic Framework

Google's LEAP paper uses only general-purpose Gemini 3.1 Pro with an agentic DAG framework to achieve 100% on Putnam 2025 and 70% on a new IMO-style formal benchmark — surpassing every specialised theorem-proving model without any fine-tuning on formal corpora.

3mo ago research

Anthropic Data: AI Task Horizons Doubling Every 4 Months, 8x Code Output — Global Pause Urged

Anthropic's Institute publishes previously unreported internal data showing Claude-powered engineers shipping 8x more code per quarter than the 2021-2025 baseline, with autonomous task horizons doubling every four months. The WSJ reports Anthropic is calling for a global pause.

3mo ago benchmark

Arena Launches Agent Leaderboard With Causal Tracing — Grok 4.3 Ranks Last, GPT-5.5 High Leads

Chatbot Arena has replaced pairwise voting with causal tracing from real in-the-wild agent sessions, producing rankings that diverge sharply from chat leaderboards. Grok 4.3, which places high on chat, finishes last on agent tasks. GPT-5.5 High and Claude Opus 4.7 Thinking lead.

3mo ago release

OpenAI Ships Dreaming 2.0: Asynchronous Memory Synthesis Fixes ChatGPT's Staleness Problem at Scale

OpenAI is rolling out a revamped memory architecture for ChatGPT that synthesizes memories automatically in the background, tackling the staleness, correctness, and scale failures of saved memories. Plus and Pro users in the US get it today; Free and international rollout follows.

3mo ago funding

Alphabet's AI Capex Is Collapsing Its Free Cash Flow by 89%. Amazon's Goes Negative. The $2 Trillion Revenue Gap.

The five largest hyperscalers will spend $700-900B on AI infrastructure in 2026 while AI revenue runs at roughly one-sixth of what Sequoia calculates is needed to justify the buildout. Alphabet FCF drops from $73B to an estimated $8B. Amazon projects negative free cash flow for the first time in its cloud era.

3mo ago release

Meta Launches Business Agent to All 1B+ Daily WhatsApp Business Threads — the Enterprise AI Play Nobody Was Watching

Meta's Business Agent goes globally available on June 3, deploying AI across WhatsApp, Messenger, and Instagram where more than 1 billion customer threads already run daily. The platform connects to hundreds of enterprise systems including Shopify and Zendesk. Free to start.

3mo ago release

Reve 2.0 and Ideogram 4.0 Both Ship Spatial Layout Control on the Same Day — Image Generation Finally Knows Where Things Go

Two leading image labs shipped bounding-box layout control on June 3. Reve 2.0 claims the world's best 4K model with precise spatial editing. Ideogram 4.0 went open-weights and debuted at Arena #8 overall, #1 among open models, with outsized gains in text rendering and commercial design.

3mo ago benchmark

Alibaba's Fun-Realtime-TTS Takes AA Speech Arena #1 at Elo 1,219 — Edges Google by 5 Points

Alibaba's Fun-Realtime-TTS moved to the top of Artificial Analysis's Speech Arena leaderboard with an Elo of 1,219, edging Gemini 3.1 Flash TTS (1,214) and Inworld Realtime TTS-2 Research Preview (1,209). It is Alibaba's first #1 finish on the Speech Arena and the second major leaderboard this week where a Chinese lab has taken the top spot.

3mo ago research

35% of Berkeley CS 10 Students Failed in Spring 2026 — 1,100 Faculty Sign Petition as AI Cheating and Math Gaps Hit Together

UC Berkeley's CS 10 failure rate tripled to 35.3% in a single semester, with Professor Dan Garcia naming AI-assisted cheating as the primary driver. Simultaneously, over 1,100 faculty across the UC system signed an open letter citing students arriving unable to do middle school math — two compounding failures that won't fully surface in graduation statistics for years.

3mo ago release

Qwen3.7 Plus Hits ScreenSpot Pro 79.0: Alibaba's GUI+CLI Hybrid Agent Enters Frontier-Tier Computer Use

Alibaba launched Qwen3.7 Plus on June 1 with ScreenSpot Pro 79.0 — placing it inside the frontier band for GUI grounding — and Terminal-Bench 70.3 for agentic coding. At $0.40/M input via OpenRouter, it is the cheapest path to frontier-class computer use in the current market.

3mo ago release

Holo3.1: AndroidWorld Jumps From 67% to 79.3%, First Computer-Use Model With Local Quantized Weights

HCompany ships Holo3.1 across four sizes with the first quantized computer-use checkpoints for consumer hardware. The 35B-A3B model lifts AndroidWorld scores by 12 points; NVFP4 quantization cuts step time from 6.8s to 3.3s on DGX Spark.

3mo ago release

Google Ships Gemma 4 12B: Encoder-Free Architecture, Native Audio Input, Runs on 16GB RAM

Google adds a 12B model to the Gemma 4 family — the first mid-sized open-weight model to drop encoder modules entirely and accept audio directly into the LLM backbone. Performance nears the 26B MoE at under half the memory footprint.

3mo ago research

Anthropic's Production Containment Numbers: 93% Auto-Approve Rate, 0.1% Injection Miss, 83% Overeager-Behavior Catch

Anthropic's engineering team publishes its first detailed post-mortem on containing Claude across claude.ai, Claude Code, and Cowork. Real production numbers reveal where human oversight breaks down and why environment-level containment now does the heavy lifting.

3mo ago policy

RAMageddon Reaches Retail: 32GB DDR5 Hits $375, Steam Deck Up $240 as AI Memory Shortage Spills Into Consumer Hardware

The AI-driven HBM shortage is repricing consumer DRAM. 32GB DDR5 now costs $375 minimum. Valve raised the Steam Deck by $240. Automakers, retailers, and electronics makers have formally petitioned Washington as the shortage deepens with no relief until 2027.