GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

4mo ago research

Pwn2Own Berlin 2026: OpenAI Codex Hacked Three Ways, Claude Code Flagged for Known Bugs, AI Tools Now a First-Class Attack Category

AI coding agents debuted as an official Pwn2Own category at OffensiveCon Berlin. Codex was exploited by three independent teams using three separate techniques. Claude Code attempts returned bug collisions — meaning known vulnerabilities were already in the codebase, unpatched.

4mo ago research

Glasswing: 10,000 Vulnerabilities in One Month, Verizon Joins as First Telecom

Anthropic's Project Glasswing has found more than 10,000 high- or critical-severity vulnerabilities across critical internet infrastructure in its first month. Cloudflare found 2,000 bugs at 10x its normal rate. Verizon just became the first telecom to join. And the bottleneck is no longer finding vulnerabilities — it's patching them.

4mo ago policy

Newsom Signs First-in-Nation AI Labor Executive Order as Meta Cuts 8,000 Jobs

California Governor Gavin Newsom signed an executive order on May 21 directing state agencies to develop early-warning systems, revise WARN Act rules, and explore worker ownership models in response to AI-driven job displacement. It is the first state-level policy framework in the US to treat AI labor disruption as a government problem rather than a company problem.

4mo ago policy

Pentagon Tests OpenAI, Google, and Grok to Replace Claude — $200M Contract in Dispute

The US military began testing AI models from OpenAI, Google, and xAI on March 1, three days after designating Anthropic a supply-chain risk. Twenty-five power users are now running competitive evaluations on GenAI.mil. Anthropic is fighting the designation in court. The NSA is still using Mythos internally.

4mo ago research

BitCPM-CANN: First Open-Source 1.58-Bit LLM Built End-to-End on Huawei Ascend 910B

ModelBest, Tsinghua University, and OpenBMB have released BitCPM-CANN, the first ternary large language model trained entirely on Huawei Ascend 910B NPUs — not ported, but natively built on Chinese hardware. At 1.58-bit weights, an 8B model retains 95-97% of full-precision accuracy with a 6x lower memory footprint.

4mo ago research

Google Genie Connects 280 Billion Street View Images to Its World Model — Waymo Trains on It Already

Google DeepMind's Genie 3 world model now draws on 280 billion Street View images across 110 countries, turning two decades of real-world photography into an interactive, promptable simulator. Waymo has been using Genie to train its autonomous vehicles on rare-event scenarios.

4mo ago model

DeepSeek Locks In V4 Pro's 75% Price Cut as Permanent — $0.87/M Output From June 1

DeepSeek has confirmed that its 75% promotional discount on V4 Pro becomes the permanent price on June 1, 2026. Output drops from $3.48 to $0.87 per million tokens — one-tenth the cost of GPT-5.5 at $30/M and a fraction of Claude Opus 4.7.

4mo ago research

AI's Memory Appetite Has Doubled DRAM Prices — Samsung's 18-Day Strike Is Now in Progress

PC DRAM up 100% in a single quarter. LPDDR5 heading from $10/GB to nearly $20/GB. Samsung's 18-day strike, which controls 40% of world supply, started May 21. Gartner says the sub-$500 laptop segment disappears by 2028.

4mo ago policy

ArXiv Bans Authors 1 Year for Unchecked AI Output — Fake Citations Up 10x Since 2023

The preprint server powering ML research now has teeth: a 1-year submission ban for papers with incontrovertible LLM artifacts. Hallucinated citations hit 1 in 277 papers in early 2026, up from 1 in 2,828 three years ago.

4mo ago research

Tri Dao's CODA Fuses Every Non-Attention Transformer Op Into the GEMM — Training's Last Memory Bottleneck Has a Fix

A paper from FlashAttention creator Tri Dao and collaborators introduces CODA, a GPU kernel abstraction that reparameterizes normalization, activations, and residual updates as GEMM epilogue programs. The operations run while the output tile is still on-chip, eliminating the memory round trips that have become a primary bottleneck in otherwise optimized training stacks.

4mo ago research

Google Sold So Much TPU Capacity to Anthropic and Meta That DeepMind Researchers Are Queuing for Scraps

Bloomberg reports that Google DeepMind researchers compete with paying customers for access to the chips they built. Andrew Dai left the company after failing to secure compute for a Gemini improvement he discovered while playing a board game. Ioannis Antonoglou, a long-tenured contributor, also departed.

4mo ago policy

France's AION Consortium Places $10B Bid for EU AI Gigafactory — 288,000 H100s, 1GW Target

Eight French heavyweights — Iliad, EDF, Capgemini, Orange, and Ardian — have lodged the largest single-country bid inside the EU's 20B InvestAI programme. The 200MW initial phase targets 1GW of total capacity and would double France's current compute footprint.

4mo ago research

Max Planck Paper Proposes Multi-Stream LLMs to Break the Single-Thread Agent Bottleneck

Researchers at MPI Tubingen show that replacing sequential message formats with parallel computation streams lets language models think while reading, act while thinking, and monitor themselves while generating — fixing a structural limitation present in every deployed agent today.

4mo ago release

Waymo Pauses Atlanta Service After Robotaxis Drive Into Floods — Two Cities Down, Same Failure Mode

Waymo has halted operations in Atlanta after multiple robotaxis drove into flooded streets. The pause follows a recall issued last week after a vehicle was swept into a creek in a separate market. Two commercial deployments suspended in a week over the same adverse-weather routing failure.

4mo ago benchmark

Cursor Composer 2.5 Takes #3 on AA Coding Agent Index at $0.07 Per Task — 60x Cheaper Than Codex

Artificial Analysis places Composer 2.5 third on its Coding Agent Index with a score of 62, behind only Claude Opus 4.7 (max) at 66 and GPT-5.5 (xhigh) in Codex at 65. Standard variant costs $0.07 per task versus $4.82 for Codex — making it the only agent above 60 on the cost-quality Pareto frontier.