GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

3mo ago policy

Meta Is the Only Major US Lab Without a Pre-Release AI Safety Deal

The Trump administration is pressing Meta to join a voluntary government review program for frontier AI models. OpenAI, Anthropic, Google, xAI, and Microsoft have all agreed. Meta has not.

3mo ago policy

NSA Loses Mythos Access After White House Export Order — Early Commercial Testers Keep Theirs

Parts of the NSA have lost access to Anthropic's Mythos 5 after Trump administration export restrictions took effect. Early commercial testers retained access through a preview carve-out. The split exposes a structural problem: the US intelligence community lost a tool mid-deployment with no clear timeline for return.

3mo ago release

Google Bakes Computer Use Into Gemini 3.5 Flash — Browser, Mobile, and Desktop in One Model

Computer use is now a native capability in Gemini 3.5 Flash, not a separate model. Google says it delivers its best agentic computer use performance yet, adds enterprise injection-stop safeguards, and opens it through both the Gemini API and Enterprise Agent Platform.

3mo ago release

OpenAI Ships Jalapeño: First Custom Inference Chip, 9-Month Tape-Out, Gigawatt Deployment Planned

OpenAI and Broadcom unveiled Jalapeño, OpenAI's first purpose-built AI accelerator. Designed in nine months with help from OpenAI's own models, it targets inference workloads for ChatGPT, Codex, and the API — with early testing showing substantially better performance per watt than current SOTA.

3mo ago policy

42 State AGs Open a Sweeping Investigation Into OpenAI — Five Days After Its IPO Filing

A coalition of 42 state attorneys general has launched an investigation into OpenAI, with New York serving a subpoena covering advertising, user safety, minor and senior data, health data, and internal AI policies. The probe must be disclosed in OpenAI's S-1 and lands as the company prepares to list at an $852 billion valuation.

3mo ago funding

SpaceX Acquires Cursor for $60B in Stock — Anysphere Becomes the Coding Layer of the SpaceX AI Stack

SpaceX has agreed to buy Anysphere, parent of AI coding tool Cursor, for $60 billion in class A stock. The deal closes the loop on a strategy spanning xAI foundation models, Colossus training compute, and now the most-used AI-first IDE in professional software development.

3mo ago funding

Google Backs $3.2B Lake Mariner Data Center to Push TPUs as Nvidia Alternative — Mirroring Nvidia's Own Lock-In Playbook

Google is financing a $3.2 billion data center near Niagara Falls, New York, that will run exclusively on its custom TPU processors. TeraWulf supplies capacity, Fluidstack brokers it, and Anthropic is the end customer. The structure copies how Nvidia built its dominance by underwriting the infrastructure its chips run on.

3mo ago benchmark

Artificial Analysis Ships Speech-to-Speech Index: GPT-Realtime-2 Leads at 77.2%, Grok Voice Tops Two of Three Dimensions

Artificial Analysis launched a composite Speech-to-Speech benchmark covering Speech Reasoning, Conversational Dynamics, and Agentic Performance. GPT-Realtime-2 (High) leads the overall index at 77.2%, but Grok Voice Think Fast 1.0 tops both the reasoning and agentic dimensions and costs 27% less per hour.

3mo ago benchmark

OpenAI Retires SWE-bench Verified: 59% of Audited Tasks Have Broken Tests, Frontier Models All Saw the Answers

OpenAI found two critical flaws that make SWE-bench Verified unsuitable for frontier model evaluation: 59.4% of audited problems have test cases that reject functionally correct solutions, and every frontier model tested can reproduce the gold-patch fix verbatim, indicating contamination.

3mo ago funding

NVIDIA's $2B Coherent Investment Secures the Photonic Layer Wiring AI Factories Together

Coherent broke ground on an expanded Sherman, Texas facility to scale InP wafer production for AI data centers. NVIDIA invested $2B in the company in March — locking in a supply of the compound semiconductors that move data between GPU clusters at light speed.

3mo ago policy

Nadella Warns AI Power Is Too Concentrated — From Inside One of Three Companies Concentrating It

Microsoft CEO Satya Nadella gave a WSJ exclusive calling out AI infrastructure concentration as a systemic risk. The warning lands as Microsoft commits $190B in 2026 AI capex and launches its own frontier models to reduce its OpenAI dependency.

3mo ago benchmark

Sakana's Fugu Ultra Posts 73.7% SWE-Bench Pro — Collective AI Beats Frontier Monoliths, 10 Days After Fable Ban

Tokyo-based Sakana AI launched Fugu and Fugu Ultra on June 22, a multi-model orchestration system scoring 73.7% SWE-Bench Pro and 93.2% LiveCodeBench — ahead of Opus 4.8 and GPT-5.5 — without training a single frontier model itself.

3mo ago release

ByteDance Launches Seed 2.1: Elo 1539 on Arena Frontend, Claims Top GDPVal, Seedance 2.5 Ships Alongside

Seed 2.1 Pro goes live on Volcano Engine today with the highest GDPVal score and Arena Code Frontend rank of 1539 (#8 globally). ByteDance ships Pro, Turbo, and weekly-updating Evolving variants, plus Seedance 2.5 for 30-second video generation.

3mo ago funding

SpaceX Signs $6.3B Reflection AI Deal — Colossus 2 Becomes the Frontier's Shared Compute Rack

Reflection AI locks $150M/month through 2029 for GB300 access at Colossus 2, bringing SpaceX's committed AI compute revenue to $2.32B per month. The Nvidia-backed open-source lab is the third frontier tenant at infrastructure originally built for xAI.

3mo ago model

GLM-5.2 Runs on a Mac: 744B Frontier Open Model Fits 239GB at 2-Bit

Unsloth's Dynamic GGUF quantization compresses Z.ai's 744B open-weight GLM-5.2 from 1.5TB to 239GB, retaining 82% accuracy and fitting inside a 256GB Mac Studio or a single H100 node with CPU offloading.