GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

4mo ago policy

Vatican Forms First AI Commission as 'Magnifica Humanitas' Launches May 25 With Anthropic Co-Founder

Pope Leo XIV created the Catholic Church's first formal AI coordination body on May 16. His debut encyclical launches May 25 alongside Christopher Olah, the Anthropic co-founder who built Claude.

4mo ago policy

Nine-Juror Unanimous Verdict: Musk Loses OpenAI Lawsuit on Statute of Limitations

A jury in Oakland cleared Sam Altman, Greg Brockman, and Microsoft of all claims in Musk's three-week trial. The jury found Musk was aware of OpenAI's commercial direction as early as 2021 — too late to sue. Altman is now free to push toward an IPO without a $150 billion disgorgement cloud overhead.

4mo ago release

Anthropic Acquires Stainless — the SDK and MCP Server Layer Comes In-House

Stainless has generated every official Anthropic SDK since the Claude API launched. Anthropic is buying them. The acquisition plants Anthropic's developer tooling stack — TypeScript, Python, Go, Java, Kotlin, CLIs, and MCP server generation — inside the company building the models.

4mo ago research

Two Papers Cut Agent Architecture Orthodoxy: One Strong Model Beats Pipelines, and Grep Beats Embeddings on Search

A Stanford paper shows single-agent LLMs match or beat multi-agent systems under equal compute budgets on multi-hop reasoning. A separate paper finds direct terminal search raises BrowseComp-Plus accuracy from 69% to 80% while reducing cost. Together, they challenge two default assumptions in production agent engineering.

4mo ago research

VulnSage: Alibaba's Multi-Agent Framework Found 146 Zero-Days — AI Exploit Generation Works on Messy Real Code

Alibaba's VulnSage uses a five-stage multi-agent pipeline to generate working exploits from vulnerability reports, outperforming prior tools by 34.64% on SecBench.js and finding 146 zero-days in production packages. The paper shows AI exploit generation now works on the messy, real-world code where symbolic and fuzzing tools have historically failed.

4mo ago release

OpenBMB's MiniCPM-o 4.5 Ships Full-Duplex Omnimodal at 9B — Beats Qwen3-Omni-30B-A3B on 12GB RAM

OpenBMB released MiniCPM-o 4.5, a 9B open-weight model that processes video, audio, and text on a single shared timeline. It runs under 12GB RAM, outperforms Qwen3-Omni-30B-A3B on omni-modal benchmarks, and ships with code, weights, and a technical report under open licence.

4mo ago release

Anthropic Splits Claude Into Two Budgets: Agent SDK Gets Non-Rollover API Credits From June 15

Claude subscribers get a dedicated monthly credit for programmatic usage from June 15, billed at full API rates with no rollover. The subsidy that made third-party agents viable on flat subscriptions is gone.

4mo ago release

OpenAI Merges ChatGPT and Codex Under Brockman — 900M Users, One Agentic Platform

Greg Brockman formally takes OpenAI's product lead, merging ChatGPT, Codex, and the developer API into a single organisation. Fidji Simo's absence is now formalised into a structural shift ahead of the IPO.

4mo ago model

Google Enters I/O Trailing Mythos by 13 Points and GPT-5.5 by 8 on SWE-Bench — Analysts Expect a Point Release, Not Gemini 4.0

Google I/O 2026 opens May 19 with Google ranked third on every major coding benchmark. Analysts expect Gemini 3.5 rather than 4.0. The company's answer is not a new model number — it's Gemini Intelligence, a proactive agentic layer across phones, watches, cars, and glasses.

4mo ago policy

EU's May 27 Tech Sovereignty Package Would Bar AWS, Azure, and Google Cloud From Government Health and Finance Data

The European Commission's Cloud and AI Development Act, due for adoption May 27, would require sovereign EU cloud infrastructure for sensitive government data in health, finance, and judiciary — restricting AWS, Azure, and Google Cloud from workloads that represent the core of Europe's government AI market.

4mo ago benchmark

Full-Duplex Bench v3: Gemini Live 3.1 Is Fastest and Goes Silent 22% of the Time

A new benchmark on real human speech — disfluencies, interruptions, mid-sentence corrections — shows Gemini Live 3.1's 4.25-second latency comes with a 22-point response reliability penalty. GPT-Realtime leads on accuracy but runs 2.3x slower.

4mo ago release

Google I/O Is 48 Hours Out: Gemini 4.0, Project Astra Production API, and Android XR on the Line

Google's May 19 keynote arrives with its benchmark standing under pressure from Claude Opus 4.7 and GPT-5.5. What's confirmed, what's expected, and what the numbers need to say.

4mo ago policy

OpenAI and Malta Deploy ChatGPT Plus to All 525,000 Citizens — the World's First National AI Rollout

Malta becomes the first sovereign government to give all citizens paid AI access, gated behind a government AI literacy course from the University of Malta. Free for one year after course completion.

4mo ago policy

Trump-Xi Summit Clears Nvidia H200 Sales to 10 Chinese Firms — Beijing Blocks Every Delivery

The US approved H200 chip sales to Alibaba, Tencent, ByteDance, and 7 others at 75,000 chips per buyer. Jensen Huang flew to Alaska to board Air Force One. Not one chip has been delivered.

4mo ago research

LLMs Are Converging to the Same Output — and the Spam Economy Is Built on It

Multiple independent LLM runs produce identical prose. Two Gemini 2.5 Flash-Lite sessions, same prompt, both open with 'The old lighthouse keeper, Elias...' Cheap agents bake that convergence into industrial-scale spam with no one checking the mistakes.