Live Feed
Harbor-Index 1.0: 82 Tasks, 54 Benchmarks, No Frontier Agent Clears 30%
tbench.ai distills 6,627 agent evaluation tasks down to 82 rigorously curated problems. GPT-5.5 on Codex CLI leads at 25.6%. The ceiling exists not because the tasks are unsolvable but because broken tasks were removed first.
Samsung's LPDDR5X PIM Delivers 614 GB/s Inside One Memory Package — 8x What Exits the Chip
At Hot Chips 2026, Samsung showed a 16 GB LPDDR5X package with compute units embedded beside DRAM banks, delivering 614 GB/s internal bandwidth versus 76.8 GB/s at the external pins. Tokens-per-second for AI inference scales with memory bandwidth, not FLOPs — and Samsung is moving the bandwidth inside.
Qwen3.8 Flash Ships with 1M Context at $0.16 Input — Cheaper on Output Than GPT-4.1 Mini
Alibaba's Qwen3.8 Flash lands with a 1M-token context window, multimodal input (text, image, video), and pricing at $0.16/$0.47 per million tokens — undercutting GPT-4.1 mini on output and sitting well below most frontier-adjacent models.
Anthropic: AI Agents Outperform 28 Safety Researchers at Fixing Deception, Sycophancy, and Jailbreaks
Automated alignment researchers built on Claude Opus 4.8 beat experienced human safety engineers across all 10 alignment failure modes, generalise to 4.7x larger models, and do it in 2,400 training examples. Human guidance doesn't help.
OpenRouter Data: GPT-5.6 Discounts Drove 13.8x Luna Token Surge — 32% of Users Stuck Around
OpenRouter's post-mortem on the GPT-5.6 Terra and Luna discount window shows a textbook Jevons paradox: daily Luna token usage jumped 13.8x, Terra 5.6x, and three-quarters of the share gained came from rival labs rather than from within the OpenAI model family.
Federal Judge Strikes Down Anthropic Blacklisting as Unlawful First Amendment Retaliation
U.S. District Judge Rita Lin ruled Thursday that the Trump administration's February designation of Anthropic as a supply-chain security risk was unlawful retaliation under the First Amendment and violated due process under the Fifth. The Pentagon designation is vacated.
TASTE Benchmark: Frontier Models Score Near Chance on AI Safety Research Judgment
Anthropic's new TASTE benchmark finds that even top frontier models cannot reliably judge AI safety research proposals. Fable 5 reaches 60% agreement with expert researchers — other frontrunners, including Opus 5 and GPT-5.6 Sol, score near the 50% chance baseline.
Agnes 2.5 Pro Beta: Singapore's $0.10/M Reasoning Model Scores 49 on AA Intelligence Index
Sapiens AI released Agnes 2.5 Pro Beta on August 26, jumping from Intelligence Index 39 to 49 while cutting input cost 78% to $0.10 per million tokens. At 91% GPQA Diamond and 70% Terminal-Bench v2.1, it undercuts DeepSeek V4 Flash on price while trailing by 3 index points.
Google DeepMind Pilots World's First Double-Blind AI Evaluation With Cryptographic Confidentiality
Google DeepMind is running the first double-blind evaluation of a proprietary frontier AI model, using privacy-preserving compute to prevent test data from reaching training pipelines. Partners include the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons.
AgentBeats Opens 13-Category Agent Benchmark Registry Beyond Coding and Customer Service
AgentBeats has published an open agent evaluation registry spanning 13 task domains including DeFi, cybersecurity, healthcare, finance, legal, and agent safety. Its AgentX-AgentBeats competition is the platform's first major structured evaluation event.
GPT-5.6-sol-xhigh and Grok-4.5 Enter Arena.ai Search Arena as Frontier Reasoning Gets Formal Search Benchmarking
Arena.ai's Search Arena added OpenAI's maximum-effort reasoning model and xAI's general-purpose Grok-4.5 on August 24. Notable: it is grok-4.5, not the two-week-old coding-focused grok-4.6, entering search evaluation — a deliberate split in how xAI positions its model lineup.
Mercury 2 for Search: The Case for Running 100 LLM Calls Per Query
Inception published the technical argument on August 26 for why diffusion decoding changes search pipeline economics. At 1,000 tokens per second, Mercury 2 makes it viable to run query rewriting, reranking, and snippet summarization 50 to 100 times per query within standard interactive latency budgets.
Celeris-1 at 1,664 Tokens Per Second Tops Artificial Analysis Speed Rankings on Commodity GPUs
A new diffusion language model from Lightspeed-backed Celeris reaches 1,664 tokens per second on standard AWS hardware, 8% ahead of Cerebras and 3.5x faster than Groq. The two fastest models on Artificial Analysis are now both diffusion architectures.
GLM-5.3 Max Enters Agent Arena — Z.AI Completes Four-Venue Push in Six Days
Z.AI's GLM-5.3 Max was added to Arena.ai's Agent Arena on August 24, completing a six-day expansion across Text Arena, Code Arena, GDPval-AA v2 (ELO 1769, #3 globally), and Agent Arena. GLM-5.3 Flash entered Code Arena:WebDev two days later.
GLM-5.3-Flash: 320B MoE, AA Index 57, DeepSWE 63.4 — All Running on Chinese Chips at $0.045/Task
Z.ai's GLM-5.3-Flash posts an Artificial Analysis Intelligence Index score of 57 at $0.045 per task — one-tenth the cost of models at comparable intelligence. On DeepSWE v1.1 it scores 63.4 versus GLM-5.2's 46.2, and nearly matches Claude Opus 4.8 on Z.ai's internal coding benchmark. The entire inference stack runs on Chinese domestic chips.