GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

26d ago benchmark

Harbor-Index 1.0: 82 Tasks, 54 Benchmarks, No Frontier Agent Clears 30%

tbench.ai distills 6,627 agent evaluation tasks down to 82 rigorously curated problems. GPT-5.5 on Codex CLI leads at 25.6%. The ceiling exists not because the tasks are unsolvable but because broken tasks were removed first.

27d ago research

Samsung's LPDDR5X PIM Delivers 614 GB/s Inside One Memory Package — 8x What Exits the Chip

At Hot Chips 2026, Samsung showed a 16 GB LPDDR5X package with compute units embedded beside DRAM banks, delivering 614 GB/s internal bandwidth versus 76.8 GB/s at the external pins. Tokens-per-second for AI inference scales with memory bandwidth, not FLOPs — and Samsung is moving the bandwidth inside.

27d ago release

Qwen3.8 Flash Ships with 1M Context at $0.16 Input — Cheaper on Output Than GPT-4.1 Mini

Alibaba's Qwen3.8 Flash lands with a 1M-token context window, multimodal input (text, image, video), and pricing at $0.16/$0.47 per million tokens — undercutting GPT-4.1 mini on output and sitting well below most frontier-adjacent models.

27d ago research

Anthropic: AI Agents Outperform 28 Safety Researchers at Fixing Deception, Sycophancy, and Jailbreaks

Automated alignment researchers built on Claude Opus 4.8 beat experienced human safety engineers across all 10 alignment failure modes, generalise to 4.7x larger models, and do it in 2,400 training examples. Human guidance doesn't help.

27d ago research

OpenRouter Data: GPT-5.6 Discounts Drove 13.8x Luna Token Surge — 32% of Users Stuck Around

OpenRouter's post-mortem on the GPT-5.6 Terra and Luna discount window shows a textbook Jevons paradox: daily Luna token usage jumped 13.8x, Terra 5.6x, and three-quarters of the share gained came from rival labs rather than from within the OpenAI model family.

27d ago policy

Federal Judge Strikes Down Anthropic Blacklisting as Unlawful First Amendment Retaliation

U.S. District Judge Rita Lin ruled Thursday that the Trump administration's February designation of Anthropic as a supply-chain security risk was unlawful retaliation under the First Amendment and violated due process under the Fifth. The Pentagon designation is vacated.

27d ago benchmark

TASTE Benchmark: Frontier Models Score Near Chance on AI Safety Research Judgment

Anthropic's new TASTE benchmark finds that even top frontier models cannot reliably judge AI safety research proposals. Fable 5 reaches 60% agreement with expert researchers — other frontrunners, including Opus 5 and GPT-5.6 Sol, score near the 50% chance baseline.

27d ago release

Agnes 2.5 Pro Beta: Singapore's $0.10/M Reasoning Model Scores 49 on AA Intelligence Index

Sapiens AI released Agnes 2.5 Pro Beta on August 26, jumping from Intelligence Index 39 to 49 while cutting input cost 78% to $0.10 per million tokens. At 91% GPQA Diamond and 70% Terminal-Bench v2.1, it undercuts DeepSeek V4 Flash on price while trailing by 3 index points.

27d ago benchmark

Google DeepMind Pilots World's First Double-Blind AI Evaluation With Cryptographic Confidentiality

Google DeepMind is running the first double-blind evaluation of a proprietary frontier AI model, using privacy-preserving compute to prevent test data from reaching training pipelines. Partners include the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons.

27d ago benchmark

AgentBeats Opens 13-Category Agent Benchmark Registry Beyond Coding and Customer Service

AgentBeats has published an open agent evaluation registry spanning 13 task domains including DeFi, cybersecurity, healthcare, finance, legal, and agent safety. Its AgentX-AgentBeats competition is the platform's first major structured evaluation event.

28d ago benchmark

GPT-5.6-sol-xhigh and Grok-4.5 Enter Arena.ai Search Arena as Frontier Reasoning Gets Formal Search Benchmarking

Arena.ai's Search Arena added OpenAI's maximum-effort reasoning model and xAI's general-purpose Grok-4.5 on August 24. Notable: it is grok-4.5, not the two-week-old coding-focused grok-4.6, entering search evaluation — a deliberate split in how xAI positions its model lineup.

28d ago research

Mercury 2 for Search: The Case for Running 100 LLM Calls Per Query

Inception published the technical argument on August 26 for why diffusion decoding changes search pipeline economics. At 1,000 tokens per second, Mercury 2 makes it viable to run query rewriting, reranking, and snippet summarization 50 to 100 times per query within standard interactive latency budgets.

28d ago model

Celeris-1 at 1,664 Tokens Per Second Tops Artificial Analysis Speed Rankings on Commodity GPUs

A new diffusion language model from Lightspeed-backed Celeris reaches 1,664 tokens per second on standard AWS hardware, 8% ahead of Cerebras and 3.5x faster than Groq. The two fastest models on Artificial Analysis are now both diffusion architectures.

28d ago benchmark

GLM-5.3 Max Enters Agent Arena — Z.AI Completes Four-Venue Push in Six Days

Z.AI's GLM-5.3 Max was added to Arena.ai's Agent Arena on August 24, completing a six-day expansion across Text Arena, Code Arena, GDPval-AA v2 (ELO 1769, #3 globally), and Agent Arena. GLM-5.3 Flash entered Code Arena:WebDev two days later.

28d ago release

GLM-5.3-Flash: 320B MoE, AA Index 57, DeepSWE 63.4 — All Running on Chinese Chips at $0.045/Task

Z.ai's GLM-5.3-Flash posts an Artificial Analysis Intelligence Index score of 57 at $0.045 per task — one-tenth the cost of models at comparable intelligence. On DeepSWE v1.1 it scores 63.4 versus GLM-5.2's 46.2, and nearly matches Claude Opus 4.8 on Z.ai's internal coding benchmark. The entire inference stack runs on Chinese domestic chips.