GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

4mo ago research

RPCS3, Godot, SDL and llama.cpp Have All Banned AI-Generated PRs. GitHub Is Considering a Kill Switch.

Five major open-source projects formally banned undisclosed AI-generated pull requests in six months. GitHub opened a community discussion in February calling the volume a 'critical issue' and is evaluating tools to let maintainers block submissions outright.

4mo ago policy

Maryland Ratepayers Face $1.6B Decade Bill for Virginia's Data Centers — FERC Complaint Filed

Maryland's consumer advocate filed a FERC complaint on May 7, arguing PJM's $2B charge on the state is driven almost entirely by Northern Virginia data center demand. At stake: $1.6 billion in residential and commercial bills over ten years.

4mo ago research

AI2's OLMo Hybrid Reaches Same Accuracy With 49% Fewer Training Tokens

Allen Institute for AI releases OLMo Hybrid, a 7B model combining transformer layers with Gated DeltaNet linear RNNs. It matches OLMo 3's MMLU accuracy using 49% fewer training tokens, with scaling law projections of 1.3-1.9x efficiency gains across model sizes from 1B to 70B parameters.

4mo ago funding

Anthropic Locks $1.8B Cloud Deal With Akamai — CDN Giant's Stock Jumps 27%

Anthropic has signed a $1.8 billion, seven-year infrastructure deal with Akamai, the largest contract in the CDN company's 28-year history. Akamai stock surged 27% on the news. The deal is Anthropic's third major compute commitment in weeks, as CEO Dario Amodei confirms 80x year-over-year growth in Claude usage.

4mo ago research

NVIDIA Star Elastic: One Checkpoint, Three Deployable Model Sizes — 30B, 23B, and 12B

NVIDIA ships Star Elastic, a training method that embeds multiple nested submodels inside a single reasoning checkpoint. One training run on Nemotron Nano v3 produces 30B, 23B, and 12B variants that slice out at deployment time with no additional training. Benchmarks show 1.9x lower latency versus standard single-model control.

4mo ago research

LLMorphism: New Paper Names the Cognitive Bias That Makes Humans Think Like LLMs

A preprint by Valerio Capraro introduces LLMorphism — the growing tendency for people to model their own cognition on LLM architecture. The paper argues the bias spreads through analogical transfer and linguistic metaphor, with downstream effects on education, healthcare, and legal responsibility.

4mo ago release

DeepSeek Ships Image Recognition to All Users at 10x Fewer Tokens Than Rivals, V4.1 Due in June

DeepSeek has rolled out its image recognition mode to all users, processing an 800x800 image in roughly 90 tokens versus 870-1,100 for mainstream models. The company has briefed investors that V4.1 arrives in June.

4mo ago release

Gemini API File Search Gets Multimodal RAG, Custom Metadata, and Page-Level Citations

Google has shipped three new capabilities to the Gemini API File Search tool: multimodal support for images and documents, custom metadata filters, and per-page citations. For enterprise RAG builders, page-level citations are the operative upgrade.

4mo ago funding

Isomorphic Labs in Talks for $2B+ to Scale AI Drug Discovery Engine Beyond AlphaFold

Google DeepMind's drug discovery spinoff is in advanced talks to raise more than $2 billion. Thrive Capital leads; Alphabet is participating. The money goes into IsoDDE, a drug design engine that doubles AlphaFold 3's score on the hardest binding-pocket tasks.

4mo ago benchmark

Arena's Seven Web Dev Categories Reveal Which Model Actually Wins Your Use Case

Aggregate Code Arena rankings hide significant model specialization. Claude Opus 4.7 Thinking leads broadly across all seven web dev categories; GPT-5.5 High dominates gaming and interactive simulations; Meta's Muse Spark tops brand marketing, informational websites, and consumer product apps.

4mo ago funding

IREN Locks $3.1B ARR, Moves Into Spain, and Acquires Mirantis — the Bitcoin Miner's AI Pivot Is Complete

IREN's Q1 2026 update reveals $3.1B in contracted annual recurring revenue against a $3.7B year-end target. Two acquisitions — Nostrum (490MW in Spain, 1GW+ pipeline) and Mirantis (software orchestration) — extend the company's reach beyond its US GPU farms. 1,210MW of 2027 capacity is now under construction.

4mo ago release

ERNIE 5.1 Launches at 6% of Industry Pretraining Cost — Arena Search #4 Globally, Beats DeepSeek-V4-Pro on Agents

Baidu's GA release compresses ERNIE 5.0 to one-third the total parameters, slashes pretraining compute to 6% of comparable models, and lands at 1,223 Elo on the Arena Search leaderboard — the only Chinese model in the global top four.

4mo ago funding

NVIDIA Has Bet $40B on AI Equity Deals in 2026 — Most of It Flows Back to Its Own Customers

The chipmaker has committed more than $40 billion to equity investments in AI companies in the first five months of 2026 — $30 billion of it in OpenAI — plus seven multi-billion deals in public companies and roughly two dozen private rounds. Wall Street calls it circular. Jensen Huang calls it a moat.

4mo ago release

OpenAI Ships GPT-Realtime-2 With GPT-5-Class Reasoning and Moves the Realtime API to GA

GPT-Realtime-2 scores 96.6% on Big Bench Audio — 15.2 points above its predecessor — and lifted Zillow's adversarial call success rate from 69% to 95%. Realtime API exits beta alongside two companion models for translation and transcription.

4mo ago research

Microsoft Research: Frontier LLMs Corrupt 25% of Document Content in Delegated Workflows

A large-scale study from Microsoft Research tests 19 LLMs across 52 professional domains in simulated long-horizon delegation. Even frontier models — Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4 — lose or corrupt an average of 25% of document content after 20 delegated interactions. Agentic tool use makes it worse. Python is the only domain where most models are ready.