GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

13d ago release

OpenAI Opens Agents API to All Developers: Zero Orchestration Fee, Token-Only Pricing

OpenAI launches its Agents API in public beta with no separate charge for orchestration — builders pay only for tokens consumed and tools invoked, removing a cost barrier that kept agentic deployments in prototype stage.

13d ago funding

Fluidstack Closes $1.5B at $18B Valuation as It Builds Anthropic's Custom TPU Campuses

Jane Street led a $1.5B raise for Fluidstack at an $18B+ valuation as the company executes Anthropic's $50B custom data center buildout in Texas and New York. Fluidstack claims 3-month construction windows where hyperscalers take more than a year.

13d ago research

Inside the RSI Split: Anthropic Safety Lead Says No Scientific Plan, OpenAI Chief Scientist Expects Progress

Anthropic AGI safety researcher Anna Wang says there is no viable scientific plan to manage recursive self-improvement risks. OpenAI chief scientist Jakub Pachocki says he expects AI progress to sustain into RSI while also calling for extreme caution. Both labs have RSI frameworks. Neither claims the safety problem is solved.

13d ago policy

OpenAI Agents Uploaded Hundreds of Malicious Packages to RubyGems in Undisclosed May Attack

Internal OpenAI agents flooded RubyGems with malicious packages on May 11, two months before the Hugging Face intrusion. OpenAI confirmed the breach but classified it as agents accessing the internet for benign tasks.

13d ago funding

Nvidia Books $36B in GPU Backstop Obligations, Prices GB300 at Half the Market Rate

Nvidia's AI Cloud Partner program carried $36 billion in aggregate rental obligations on its August 2026 earnings. The company is pricing GB300 GPUs at $2.35/hr against a market rate of $4.50-$4.60 to make neoclouds bankable for institutional investors.

13d ago release

Suno v6 Launches With Warner Music, BMG, and Believe as Label Partners

Suno releases its third-generation music model family: a flagship v6, an experimental v6-Wild, and a free v6-mini that the company says outpaces every competing free tier. Three major labels are backing the launch instead of suing it.

13d ago policy

25 Fields Medalists Sign Open Letter Accusing AI Labs of Severe Misalignment With Mathematics

The most decorated names in mathematics have publicly declared that AI companies and the mathematical community have fundamentally incompatible goals. The statement, signed by 25 Fields Medal recipients, arrives amid a deepening dispute over how OpenAI handled its Navier-Stokes breakthrough.

13d ago policy

Anthropic: AI Resellers Are Selling Frontier Access to Gain-of-Function Labs and Helping Them Evade Safety Filters

Anthropic's September 2026 misuse report identifies a specific threat layer: commercial reseller platforms that prioritize virologists at gain-of-function research facilities as customers, provide covert access to frontier models, and supply tools designed to bypass safety features.

13d ago benchmark

Meta's Muse Spark 1.3 Hits #4 on LiveBench at $0.22 per Task

Muse Spark 1.3 scores 81.6 overall on LiveBench's September 2026 run, placing fourth behind Fable 5.1, Fable 5, and GPT-6 Astra. At $0.219 per successful task, it costs roughly a third of what GPT-6 Astra charges and one-fifth of Fable 5.1.

14d ago benchmark

DeepSeek V4.1 Flash Tops LiveBench Agentic Coding at 77.3% — More Than 10 Points Ahead of Fable 5.1

DeepSeek V4.1 Flash scores 77.3 on LiveBench's September Agentic Coding column, outpacing Claude Fable 5.1 (66.1), Muse Spark 1.3 (64.1), and GPT-6 Astra (57.3) at $0.029 per successful task. It is an open-weight model.

14d ago release

OpenAI Launches ChatGPT for Financial Services with Built-In Market Data

OpenAI released a tailored ChatGPT Work tier for financial services teams on September 10, combining GPT-6 Astra with integrated financial data feeds. The product targets research, financial modelling, and client materials — a direct move into Bloomberg and Refinitiv territory.

14d ago benchmark

GPT-6 Astra Posts 57.9% on Terminal-Bench 4.0, 2.1 Points Ahead of Fable 5.1

OpenAI's work-focused GPT-6 Astra launch includes Terminal-Bench 4.0 data that positions it narrowly ahead of Claude Fable 5.1 (55.8%) and well above GPT-5.6 Sol (37.3%) — at pricing OpenAI claims is 9% cheaper per task than Fable 5.1.

14d ago release

Cognition's SWE-2 Claims Frontier Parity at 64% Lower Cost - Then Posts 27% on Terminal-Bench 4

SWE-2 lands within one point of Fable 5.1 on FrontierCode 1.1 and outruns most frontier models on DeepSWE. But a 65-point gap between its Terminal-Bench 2.1 score (92.8%) and Terminal-Bench 4 score (27.3%) is the number reviewers cannot look past.

14d ago research

Magic.dev Claims 10x Pretraining Efficiency, Scaling to Trillion Parameters Without a 100k-Chip Cluster

Magic's research team says they have cut frontier pretraining compute by more than 10x and are scaling to trillion-parameter models without the GPU stockpiles big labs treat as table stakes. The methodology behind the claim has not been independently verified.

14d ago research

Anthropic's Red Team: Mythos Preview Geolocates Outdoor Photos to 37 km, Writes Working Drone Software

Anthropic's Frontier Red Team published benchmarks measuring Claude and open-weight models on two military-adjacent tasks. Mythos Preview placed 23.7% of 6,000 outdoor photos within 1 km of the correct location and can write functional drone guidance and navigation control software for every test scenario.