GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —

Live Feed

5mo ago model

Anthropic Launches Claude Managed Agents: $0.08/Hr to Run Your AI Agents in the Cloud

Anthropic opened Claude Managed Agents to public beta on April 8, 2026, offering a fully managed runtime for long-running AI agents at $0.08 per session-hour plus standard token costs. Notion, Asana, Sentry, Rakuten, and Atlassian are already in production.

5mo ago model

Same Model, 17 More Fixes: Augment Code Leads SWE-Bench Pro by Outscaffolding Claude Code

Augment Code's Auggie CLI scored 51.80% on SWE-bench Pro — 2 points above Cursor and 17 problems above Claude Code — despite all three agents running the same Claude Opus 4.5 model. The entire performance gap came from Augment's semantic codebase indexing engine, not from the underlying model.

5mo ago model

DeepSeek R2 Scores 92.7% AIME on 32B Dense Model — Runs on a Single RTX 4090

DeepSeek has abandoned its 671B MoE architecture for a 32-billion-parameter dense transformer that scores 92.7% on AIME 2025, runs on consumer hardware with a single 24 GB GPU, and undercuts Western frontier reasoning APIs by roughly 70% on token cost. MIT license.

5mo ago model

Meta Enters the Frontier: Muse Spark Debuts at #4 on the Intelligence Index

Meta Superintelligence Labs launched Muse Spark on April 8 — the first frontier-class model from Meta since Llama 4 and their first proprietary, closed-weights release. It scores 52 on the Artificial Analysis Intelligence Index, placing it 4th behind Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6.

5mo ago policy

OpenAI Publishes Superintelligence Transition Blueprint — Wealth Redistribution, Workforce Protection, Federal Preemption

OpenAI released a landmark policy document in April 2026 outlining how it believes governments should manage the transition to superintelligent AI: universal income mechanisms, workforce safeguards, energy investment, and federal preemption of state AI laws.

5mo ago model

AI API Pricing in 2026: A 97% Drop in Three Years and Still Falling

GPT-4 cost $30 per million input tokens when it launched in 2023. Today, comparable-quality models are under $1/M — a 97% reduction in three years. Anthropic averages $6.27/M across its lineup; Meta averages $0.17/M. The price war is far from over.

5mo ago model

The AI Infrastructure Arms Race: Google, CoreWeave, and Why Compute Beats Models

Google committed $5B to finance a Texas data center for Anthropic. CoreWeave closed an $8.5B debt facility anchored by Meta. OpenAI's $122B raise is primarily an infrastructure bet. The race is no longer about who has the best model — it's about who controls the electricity.

5mo ago release

Anthropic Launches Claude Opus 4.6 Fast Mode at $30/$150 Per Million — 6x the Standard Price

Anthropic has added a speed-optimised variant of Claude Opus 4.6 that doubles throughput at a 6x cost premium. At $30/M input and $150/M output, Fast Mode targets latency-sensitive production pipelines willing to pay for it.

5mo ago model

Anthropic's Claude Mythos Preview Posts 93.9% on SWE-Bench Verified — a Generation Ahead of Opus 4.6

Anthropic's unreleased Mythos Preview scores 77.8% on SWE-Bench Pro and 93.9% on SWE-Bench Verified — 24 points and 13 points ahead of Opus 4.6 respectively. On SWE-Bench Multimodal it scores 59.0% versus 27.1% for Opus 4.6.

5mo ago model

China's Z.AI Releases GLM-5.1, Tops SWE-Bench Pro Ahead of GPT-5.4 and Claude Opus 4.6

Z.AI has released GLM-5.1, scoring 58.4 on SWE-Bench Pro — the highest of any model, clearing GPT-5.4 (57.7) and Claude Opus 4.6 (57.3). It runs on zero Nvidia hardware.

5mo ago policy

OpenAI, Anthropic and Google Form First Shared Intelligence Pact Against Chinese AI Theft

The three largest US AI labs are sharing real-time threat data on Chinese adversarial distillation through the Frontier Model Forum — the first time direct competitors have operationalised a joint defence against a named adversary. Microsoft, the fourth FMF founder, is not a signatory.

5mo ago model

Gemini 3.1 Pro Preview Matches GPT-5.4 on Intelligence Index at 20% Lower Cost

Gemini 3.1 Pro Preview prices at $2.00/$12.00 per million tokens and ties GPT-5.4 at score 57 on Artificial Analysis's Intelligence Index — making it the most cost-efficient top-tier option for developers who need frontier reasoning without flagship pricing. The gap to Claude Opus 4.6 is 60% on input tokens.

5mo ago benchmark

DeepSeek V4 Wins Code, Gemini Flash-Lite Wins Content: The 2026 LLM Cost-Per-Task Rankings

Per-token pricing is the wrong lens for AI spend in 2026. A 155-model analysis shows that cache hit rates, token efficiency, and batch discounts restructure which model is cheapest — completely differently by task type. A model at $0.30/M input can cost more per task than one at $2.00/M.

5mo ago benchmark

SWE-Rebench Exposes Contamination: DeepSeek V3 Scores Drop 46% on Fresh Tests

Decontaminated benchmark SWE-rebench shows Chinese models losing up to 18 percentage points versus SWE-bench Verified scores — evidence that a large share of their reported coding ability comes from training data memorisation, not genuine problem-solving.

5mo ago model

Anthropic Signs Gigawatt-Scale TPU Deal With Google and Broadcom, Compute Comes Online in 2027

Anthropic has secured multiple gigawatts of next-generation TPU capacity through a new agreement with Google Cloud and Broadcom, expected to come online starting in 2027. The deal is the largest compute commitment Anthropic has announced and extends a November 2025 pledge to invest $50 billion in US computing infrastructure.