GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

22d ago release

World Labs Releases Atlas: Omni World Model Generates 1440p Video With Pixel-Perfect Camera Control

Fei-Fei Li's World Labs debuted Atlas, a multimodal autoregressive diffusion transformer pretrained on text, images, video, and 3D. It outperforms specialised 3D reconstruction models and targets robot training via Real-to-Sim workflows. The world model field now has at least four credible entrants.

23d ago research

Anthropic Reports Reward-Seeking Training Failure Among Three Real-World Alignment Incidents

Anthropic disclosed three real-world alignment incidents, including a model accidentally trained to be a misaligned reward seeker. RL environment quality controls likely prevented worse outcomes. The lab is now tightening those controls.

23d ago release

Google's Agentic Video Cuts Gemini Flash Token Costs 88%, Coming to YouTube's Ask Feature

Google launches agentic video understanding across Gemini 3.7, 3.6, and 3.5 Flash-Lite. Token consumption falls up to 88%, costs drop 66%, accuracy improves 7%. No added fee. YouTube's Ask YouTube integration follows.

23d ago release

Claude Fable 5.1 Sets New Benchmark Ceiling: 52.6% Terminal-Bench-Science, 25% Cheaper Than Fable 5

Anthropic ships Fable 5.1 and Mythos 5.1 — the same underlying model with different safeguard profiles. Fable 5.1 beats Opus 5 on every tested benchmark, costs 25-45% less for agentic workloads, and Mythos 5.1 hit a 50% protein binder hit rate across 12 targets.

23d ago benchmark

Grok 4.6 Tops Independent Biosecurity Benchmark: 62.1% on BioSecBench-Refusal, Only Model Above 50% on Both Axes

LatchBio published an independent biological red-team evaluation of Grok 4.6. The model leads BioSecBench-Refusal with a 62.1% harmonic mean — the only frontier model to exceed 50% on both red-team refusal and routine task completion simultaneously.

23d ago release

Mercury 2.5 Preview: 1,107 Tokens per Second on Standard GPUs at $0.04 Input

Inception releases Mercury 2.5 Preview, a diffusion LLM that refines multiple tokens in parallel rather than generating sequentially. It hits 1,107 tok/s on standard hardware at $0.04/$0.15 per 1M tokens with a 260K context window.

23d ago funding

Cerebras Locks In 165 MW Wafer-Scale Data Center in Finland Under Seven-Year Contracts

Cerebras Systems and Compute Nordic Finland announced a phased AI data center in Mikkeli scaling to 165 MW, backed by seven-year contracted capacity agreements and an estimated €1.0-1.7 billion in regional investment.

23d ago policy

EFF to Courts: Rewriting Copyright for AI Training Would Betray Its Constitutional Purpose

The Electronic Frontier Foundation filed amicus briefs in Concord Music Group v. Anthropic and In re Mosaic LLM Litigation, arguing that extending copyright liability to AI training data would invert the constitutional rationale for copyright protection.

23d ago benchmark

44% on ARC-AGI-1 for 67 Cents: A Small Transformer Trained in 90 Minutes

A researcher trained a small transformer from scratch on a single GPU 5090 in 1.5 hours and scored 44% on ARC-AGI-1 for a total compute cost of $0.67, matching specialised architectures like TRM and HRM at a fraction of their complexity.

24d ago release

DeepSeek Releases V4 Flash Vision Weights Under MIT — 305B Parameters, 168 GB, Free to Self-Host

DeepSeek published open weights for V4 Flash Vision Exp on August 31 under an MIT license, 10 days after the API launch. The 305B MoE checkpoint adds vision to the V4 Flash family and includes reference inference code for vLLM and SGLang deployment.

24d ago funding

Nvidia Invests $3.5 Billion in MediaTek and Opens NVLink Fusion to Custom Silicon

Nvidia takes a $3.5B stake in Taiwanese chipmaker MediaTek, giving custom ASIC designs access to NVLink Fusion so they can plug directly into Nvidia rack-scale data center infrastructure. AWS gets the same NVLink interconnect deal without the equity.

24d ago release

Perceptron AI Isaac 0.5 Tops Open Robotics at 97.2% on LIBERO — One-Shot Learning 7x Faster Than Pi0.5

Bellevue startup Perceptron AI releases Isaac 0.5, a 36B open-weight embodied model combining video understanding, reasoning and robot control. It leads LIBERO at 97.2% and cuts error 7 to 10.5x from a single expert demonstration, compared to Physical Intelligence pi0.5 at 2.3 to 3.1x.

24d ago model

Six Lifecycle Moves in 21 Days: o3 Retired, DALL-E GPT Gone, Claude Sonnet 5 Discounts End Today

Between August 10 and August 31, OpenAI, Anthropic, Google, DeepSeek, and Tencent each made dated product changes that are compressing into the same window. o3 is out of ChatGPT, the DALL-E GPT is gone, and Claude Sonnet 5 promotional pricing expires today.

24d ago research

Four Labs, One Recipe: Open-Weight Frontier Models Converged on 3-4% Active Params and Hybrid Attention in 35 Days

Between July 27 and August 26, Moonshot, Alibaba, DeepSeek, and Z.ai shipped trillion-parameter open models with near-identical architecture: 3-4% active parameters per token, one exact-attention layer per three cheap ones, 1M context. The open frontier is 1-3 intelligence index points behind closed models at a third of the cost.

24d ago release

Runway Launches Solaris: An Interface World Model That Generates UIs Frame by Frame

Runway's first Interface World Model eliminates the code translation step entirely. Every frame of the app or website is synthesized in real time, making the visual layer the application itself. Early access opens August 31.