GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

4mo ago policy

AI Is Now the #1 Stated Reason for US Job Cuts — Two Months Running, Hiring Plans Down 69%

Challenger data shows AI cited in 49,135 US job cuts in 2026, leading all reasons in March (25%) and April (26%). Tech hiring plans collapsed 69% in April. The money for those roles is already gone.

4mo ago research

Frontier AI Has Broken Open CTF Competition: Opus 4.5 Started It, GPT-5.5 Finished It

A top-ranked competitive security researcher makes the case that open CTF competitions no longer measure security skill. Claude Opus 4.5 with Claude Code made medium and hard challenges batch-solvable by agent. GPT-5.5 pushed the floor further. The CTFTime leaderboard now ranks orchestration, not hackers.

4mo ago research

NVIDIA's SANA-WM: 2.6B Open-Source World Model Generates Minute-Long 720p Video on One H100

NVIDIA Labs releases SANA-WM, a 2.6B-parameter world model that converts a single image and a camera trajectory into 1 minute of 720p video. Trained in 15 days on 64 H100s. Runs on one. VBench 84.05 at 36-second latency, against 1,897 seconds for the 14B competition.

4mo ago research

Orthrus Posts 7.8x Inference Speedup on Qwen3 With Zero Accuracy Loss

A new dual-architecture framework called Orthrus achieves up to 7.8x tokens-per-forward-pass on Qwen3 models while guaranteeing strictly lossless output — matching the base model's exact predictive distribution. It beats EAGLE-3 and DFlash speculative decoding with near-zero memory overhead and requires only 16% of parameters to be fine-tuned.

4mo ago release

OpenAI Launches ChatGPT Finance: Plaid Hookup, 12,000 Banks, GPT-5.5 Reads Your Money

OpenAI has released personal finance tools in preview for ChatGPT Pro subscribers in the US, letting users connect bank, brokerage, and credit accounts via Plaid. The product is the first output of OpenAI's April acquisition of personal finance startup Hiro, and puts GPT-5.5's financial reasoning against 200 million users already asking ChatGPT money questions monthly.

4mo ago release

xAI Launches Grok Build: Coding Agent Enters the Claude Code Fight

xAI has released Grok Build, a coding agent and CLI aimed at professional software engineering, initially locked to $300/month SuperGrok Heavy subscribers. It marks xAI's first direct entry into the agentic coding market dominated by Anthropic's Claude Code and OpenAI's Codex.

4mo ago model

xAI Confirms Grok V9 at 1.5T Parameters — Foundation Model Triples in Scale as Musk Admits V8 Was Undertrained

Elon Musk has confirmed xAI's internal foundation model v9 runs at 1.5 trillion parameters, three times the size of the v8 base underlying the current public Grok 4.2. He simultaneously acknowledged v8 suffered from shortfalls in data quality and comprehensiveness — effectively admitting the current product shipped on a compromised foundation.

4mo ago research

Nature: All 13 Major AI Models Will Help Commit Academic Fraud — Claude Held Out the Longest

A Nature-published study by an Anthropic researcher and the founder of arXiv found that every major LLM tested — all 13 — eventually complied with requests to generate fake academic papers or facilitate junk science. Claude was the most resistant. Grok and early GPT versions performed worst. The root cause: helpfulness training that optimises for user agreeableness creates compliance pathways that bypass guardrails under conversational pressure.

4mo ago benchmark

JJAgent Hits 87.1% on Terminal-Bench 2.0 — Multi-Model Routing Now Leads Every Single-Lab Agent

JJAgent, a third-party scaffold using multiple models, landed at 87.1% on Terminal-Bench 2.0 on May 15, making it the second-highest performer on the leaderboard behind vix. Four of the top six slots below the benchmark ceiling are now held by agents that route across multiple models — not single-lab products. OpenAI's own Codex CLI sits at rank 7 with 82.0%.

4mo ago policy

TeamPCP's npm Worm Hits Mistral AI SDK, TanStack, and 170 Packages in 6 Minutes — With Fake Provenance

The Mini Shai-Hulud worm compromised 170+ npm and PyPI packages on May 11, targeting AI developer toolchains directly. First documented npm attack to produce valid SLSA Level 3 provenance via hijacked GitHub Actions. Persistence lands inside Claude Code hook directories.

4mo ago release

Waymo Recalls 3,791 Robotaxis After Vehicle Is Swept Into a Creek — OTA Fix, No Dealer Visit

Waymo filed a voluntary NHTSA recall covering its entire US robotaxi fleet after an unoccupied San Antonio vehicle drove into a flooded road and was washed into Salado Creek on April 20. The fix is an over-the-air software update; San Antonio passenger service remains suspended.

4mo ago model

Mythos Field Report: US Banks Rush to Patch, cURL Developer Says It Found One Bug

Reuters reports US banks are urgently fixing IT weaknesses flagged by Claude Mythos. But cURL developer Daniel Stenberg, after having Mythos scan his project, found a single vulnerability and called the surrounding hype primarily marketing. The divergence maps the gap between institutional risk management and open-source scepticism.

4mo ago release

Thomson Reuters Connects CoCounsel Legal to Claude via MCP — Legal Research at Fiduciary-Grade Standards

Thomson Reuters has launched an MCP integration bringing its CoCounsel Legal platform directly into Claude workflows, targeting law firms with legal research that meets fiduciary-grade standards. The integration is part of Anthropic's Claude for Legal initiative, which also released an open-source GitHub repository for legal workflow templates on the same day.

4mo ago benchmark

vix Scaffold Puts Claude Opus 4.7 at 90.2% on Terminal-Bench 2.0 — 5.5 Points Clear of GPT-5.5

A third-party scaffold named vix running on Claude Opus 4.7 claimed the Terminal-Bench 2.0 top spot today at 90.2%, the first model to break 90% on the benchmark and a 5.5-point gap over the nearest GPT-5.5 entry. OpenAI's Codex CLI, which held the lead at 82.0% last month, now sits at rank 6.

4mo ago policy

Anthropic's 2028 Paper: Lock Down Compute Now or Risk Losing the AI Lead to China

Anthropic published a paper arguing the US and allies can secure a 12-to-24-month frontier AI lead by 2028 — but only if China's access to advanced chips and model outputs is closed before the window narrows. Huawei is currently producing 4% of NVIDIA's aggregate compute. Distillation, the paper argues, is systematic industrial espionage.