GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%

Live Feed

4mo ago research

Oxford Study in Nature: Warmth-Tuned AI Makes 7.43pp More Errors, 40% More Sycophantic When Users Express Sadness

An Oxford Internet Institute study published in Nature fine-tuned five AI models for warmth and found error rates climb 10-30 percentage points across medical advice, conspiracy claims, and factual questions. Warm models were 40% more likely to validate false user beliefs, with the gap widening to 12.1pp when users expressed emotion alongside incorrect statements.

4mo ago policy

Grok Ranked Most Likely to Trigger User Delusions in BBC Investigation of 14 Cases Across 6 Countries

A BBC investigation published May 3 found Grok the most dangerous AI for inducing delusional thinking, based on interviews with 14 affected users and controlled research by social psychologist Luke Nicholls. ChatGPT 5.2 and Claude performed significantly better at redirecting users.

4mo ago policy

Maryland Becomes First US State to Ban AI-Driven Grocery Pricing — Critics Say the Law Has Carveouts

Maryland signed legislation on May 1 prohibiting retailers from using personal data to change grocery prices in real time, becoming the first US state to regulate AI-powered surveillance pricing. The Guardian and multiple outlets confirm the law passed with exemptions that consumer advocates say blunt its impact.

4mo ago release

OpenAI Ships Auto-Review for Codex: AI Reviews AI at 99.93% Approval, Blocks 99.3% of Prompt Injections

OpenAI published its Auto-review system for Codex on April 30: a separate agent that replaces synchronous human approval at the sandbox boundary, achieving 99.93% effective approval across all actions while blocking 99.3% of prompt injection attacks and 96.1% of MonitoringBench Hard cases.

4mo ago funding

Anthropic Earns $16.20 Per User Per Month. OpenAI Earns $2.20. Counterpoint's ARPU Data Explains the Revenue Reversal.

Counterpoint Research puts Anthropic's monthly revenue per active user at $16.20 against OpenAI's $2.20 — a 7.4x gap across a user base one-seventh the size. The numbers make the revenue crossover arithmetic obvious: fewer, higher-value users compound faster than large consumer audiences paying nothing.

4mo ago funding

OpenAI Is Behind Anthropic on Revenue, Missed Its 1B User Target, and Its CFO Is Worried About the Compute Bill

WSJ reports OpenAI missed multiple monthly revenue targets in early 2026 and now runs at $24B ARR — trailing Anthropic's $30B. CFO Sarah Friar has privately warned the company may not be able to fund its $600B+ in compute commitments if growth does not accelerate, complicating a planned IPO.

4mo ago release

VS Code 1.117.0 Enables Copilot Co-Authorship on All AI-Touched Commits by Default — Including Tab Completions

A two-line PR merged April 16 flipped VS Code's git.addAICoAuthor setting from off to all, meaning every commit containing detected AI-generated code now automatically appends 'Co-authored-by: GitHub Copilot' as a git trailer — without user prompting. Opt-out requires an explicit settings change.

4mo ago research

Uber's Play to Become the AV Industry's Data Layer: 25 Partners, an AV Cloud, and Millions of Drivers as Rolling Sensors

Uber CTO Praveen Naga outlined a long-term plan to outfit driver cars with sensors for autonomous vehicle training — describing data, not technology, as the AV bottleneck. The company already has 25 AV partners and a shadow-mode testing infrastructure. The ride-hailing company that sold its self-driving unit in 2020 is positioning itself as the data supplier to everyone who didn't.

4mo ago policy

California Ends Robotaxi Citation Immunity: AV Manufacturers Get the Ticket Starting July 1

California's DMV formally adopted the most comprehensive AV regulations in the US on April 28. From July 1, police can issue a Notice of AV Noncompliance directly to manufacturers. Repeated violations can suspend operating permits. Autonomous trucks are now cleared for California roads.

4mo ago release

Poolside Releases Laguna M.1 and XS.2: 225B MoE Coding Agent at 72.5% SWE-Bench, Open Weights Under Apache 2.0

Government AI startup Poolside goes public with its first two models. Laguna M.1 (225B-A23B) posts 72.5% SWE-bench Verified and 46.9% SWE-bench Pro. Laguna XS.2 (33B-A3B) is open-weight, runs locally on 36 GB of RAM, and ships free on OpenRouter.

4mo ago benchmark

ERNIE 5.1 Preview Hits Arena #13 Globally With 1/3 the Parameters of ERNIE 5.0 — and Leads Worldwide in Legal

Baidu's ERNIE 5.1 Preview reached 1476 ELO on LMArena's Text leaderboard — #13 globally and #1 among Chinese models — while ranking first in the Legal and Government category above every Western frontier model. It achieves this using only 6% of the pre-training compute of comparable models.

4mo ago funding

Founders Fund Raises $6B After Burning Through $4.6B in Under a Year — 7 Checks, Average $600M Each

Peter Thiel's Founders Fund closed a $6 billion fund on May 1 — its largest ever — less than twelve months after its predecessor was fully deployed across seven companies. The prior fund included $1.25 billion in Anthropic and $1 billion in Anduril. Q1 2026 saw $297 billion flow into startups globally, the most ever in a single quarter.

4mo ago policy

Pentagon Completes Seven-Company AI Bloc — Anthropic Is the Only Major Lab Without a Deal

The US Defense Department has now signed classified AI agreements with seven companies: Google, OpenAI, SpaceX, Nvidia, Microsoft, AWS, and Reflection AI. Every top-tier AI lab holds a Pentagon deal except Anthropic, which was ejected after refusing to remove usage restrictions. Its valuation has since risen to $900 billion.

4mo ago release

OpenAI Rolls Out GPT-5.5-Cyber to Critical Infrastructure Defenders — Pentesting and Malware RE Now in Scope

OpenAI announced GPT-5.5-Cyber on April 29, a frontier model with explicit offensive security capabilities — penetration testing, vulnerability exploitation, and malware reverse engineering — being rolled out to vetted defenders through its Trusted Access for Cyber program. The move escalates from April's GPT-5.4-Cyber and pairs with $10M in API grants for security organizations.

4mo ago release

Microsoft Builds Its Own Legal Agent in Word — 18 Robin AI Engineers, No Third-Party Model Required

Microsoft launched a first-party Legal Agent inside Word on April 30, built by 18 engineers acqui-hired from defunct Robin AI. A deterministic insertion engine handles tracked changes natively, cutting Microsoft loose from general-purpose LLMs for contract review — and putting it directly against Anthropic's Claude-in-Word play.