GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —

Live Feed

5mo ago release

Cloudflare Ships Five Agent Products in Three Days — Project Think, Mesh, and the Bid to Own Agent Infrastructure

Cloudflare ran "Agents Week" April 15–17, launching Project Think (a batteries-included Agents SDK with durable execution and sub-agents), Cloudflare Mesh (private networking for agent fleets), Agent Lee (dashboard AI assistant), Redirects for AI Training, and an Agent Readiness Score. It is the most concentrated infrastructure land-grab since the cloud wars.

5mo ago funding

Meta to Cut 8,000 Jobs Starting May 20 — The Largest AI-Driven Headcount-to-Compute Swap in Tech History

Meta will begin laying off roughly 10% of its global workforce on May 20, 2026, with additional rounds expected later in the year. The capital is being redirected into chips, data centres, and model training — making this the single largest AI-driven reallocation of labour to compute in tech history.

5mo ago model

MiniMax Open-Sources M2.7 at 1495 GDPval-AA ELO and 57% Terminal Bench — Self-Evolving Agent Hits Top-5

MiniMax has released M2.7 under an open-source license, posting 1495 ELO on GDPval-AA (highest at release), 56.22% on SWE-Pro, and 57.0% on Terminal Bench 2. A Chinese lab has put a freely self-hostable model inside the global top tier on agentic capability benchmarks.

5mo ago research

MIT, Oxford, CMU Study: 10 Minutes of AI Assistance Measurably Reduces Independent Problem-Solving Persistence

A multi-institution study across 1,200 participants finds that direct AI assistance improves immediate performance but reduces problem-solving persistence within minutes of the AI being removed. The damage is locatable: it comes from answer-outsourcing, not AI use itself.

5mo ago policy

Bank of England, Fed, and Treasury Convene Emergency Sessions Over Claude Mythos — White House Extends Access to Federal Agencies

Finance ministers and central bank governors worldwide are holding crisis meetings after Claude Mythos demonstrated the ability to chain previously unknown vulnerabilities across banking and critical infrastructure software. The White House is now planning Mythos access for major federal agencies.

5mo ago release

OpenAI Codex Becomes a macOS Desktop Agent — Computer Use, Memory, and 90+ Plugins Ship Together

OpenAI has transformed Codex from a coding assistant into a persistent macOS agent that can see, click, browse, generate images, and remember user habits across sessions. The update is a direct escalation against Anthropic's Claude Code and positions Codex for OpenAI's planned super-app.

5mo ago release

Perplexity Launches Personal Computer: A 24/7 Mac Agent That Owns Your Files and Apps

Perplexity Personal Computer turns any Mac into a persistent AI agent with access to local files, native apps, and the browser simultaneously — running on 20+ frontier models via Perplexity servers, available to Max subscribers now.

5mo ago research

Tencent Open-Sources HY-World 2.0: 3D Scene Construction From Text, Image, or Video

Tencent Hunyuan releases HY-World 2.0 under open weights, shifting the world-model paradigm from pixel-stream prediction to persistent 3D asset generation — meshes and Gaussian splats importable into Blender, Unity, and Isaac Sim.

5mo ago funding

AMD Locks In France's First Exascale Supercomputer — Alice Recoque Breaks NVIDIA's Hold on Sovereign AI

AMD and the French government signed a multi-year Letter of Intent in Paris to accelerate France's national AI strategy. Alice Recoque, France's planned first exascale supercomputer, will run on AMD silicon — making it one of the largest sovereign AI deployments outside NVIDIA's ecosystem.

5mo ago policy

Anthropic Deliberately Trained Opus 4.7 to Be Less Capable — and Published That Fact

In the Opus 4.7 system card, Anthropic explicitly states the model does not advance its capability frontier, that Mythos Preview outperforms it on every relevant evaluation, and that Anthropic actively worked to reduce Opus 4.7's cybersecurity capabilities during training. The release is less a product launch and more a staged-release governance experiment.

5mo ago release

Claude Opus 4.7 Launches at $5/$25 Per Million, CursorBench Climbs to 70%

Anthropic has shipped Claude Opus 4.7 as a same-price upgrade over Opus 4.6, with stronger long-running coding performance, higher-resolution vision and tighter instruction-following. Launch metrics show 70% on CursorBench, 90.9% on BigLaw Bench and a 13% lift on Anthropic’s 93-task coding set.

5mo ago funding

Factory Raises $150M at $1.5B to Build Enterprise AI Coding Agents

Factory closed a $150M round led by Khosla Ventures with Sequoia, Insight Partners, and Blackstone. The three-year-old startup builds model-agnostic AI agents for enterprise engineering teams at Morgan Stanley, Ernst & Young, and Palo Alto Networks. Its Droid scaffold already sits at rank 6 on Terminal-Bench 2.0.

5mo ago benchmark

GLM-5.1 Tops Its Open-Weight Class on AA Index at 44 — Frontier Leaders Hold a 23-Point Lead

Z.AI's GLM-5.1 ranks #1 among open-weight non-reasoning models on Artificial Analysis Intelligence Index v4.0 with a score of 44, and #28 overall out of 474 models. The gap to joint frontier leaders GPT-5.4 and Gemini 3.1 Pro Preview (both 57) is 23 points. At $1.40/$4.40 per million tokens, its API pricing sits 2.5× above the class average.

5mo ago policy

Jensen Huang: ASICs Are 'Not Sensible', and DeepSeek V4 on Huawei Ascend Would Be a 'Horrible Outcome' for America

In a wide-ranging interview with Dwarkesh Patel, Nvidia CEO Jensen Huang dismissed the ASIC challenge to GPU dominance as a graveyard of cancelled projects, then warned that DeepSeek V4 optimising for Huawei's Ascend 950 — rather than Nvidia hardware — would be a geopolitical loss for the United States.

5mo ago release

OpenAI's GPT-Rosalind Posts 0.751 on BixBench — First Domain-Specific Model for Drug Discovery, Restricted to US Enterprise

OpenAI releases GPT-Rosalind, a frontier reasoning model fine-tuned for biochemistry, genomics, and drug discovery. It leads all evaluated models on BixBench at 0.751 Pass@1, outperforms GPT-5.4 on 6 of 11 LABBench2 tasks, and is initially restricted to qualified US enterprise customers through a Trusted Access programme.