GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

30d ago release

Thomson Reuters Builds Its Own Frontier Legal LLM for $40M — Trained on Westlaw, Built to Cut Anthropic Costs

Thomson Reuters launched Thomson 1.0, a proprietary LLM built on Qwen 3.5 and trained on Westlaw, Practical Law, Checkpoint, and Reuters content. Total investment: $40M across talent and compute. Final training run cost: $450,000. First deployment: tabular analysis in CoCounsel Legal.

30d ago research

Inference Engine Bugs Let LLMs Execute Code on GPU Host Machines, Security Analysis Finds

A detailed analysis published on LessWrong identifies inference engines like vLLM and SGLang as the attack surface between a malicious LLM and the GPU host running its weights. vLLM's CVE-2025-9141 — an eval() call on tool parameters that shipped despite Gemini flagging it as critical — is the concrete proof of concept.

30d ago release

Microsoft Ships Agent Lightning v1.0: A Coding Agent That Improves Other Agents

Agent Lightning is a installable skill for Claude Code, Codex, and GitHub Copilot that takes an editable agent plus a benchmark and systematically improves prompts, tools, and reasoning settings. It uses harnessed agentic RL, where the production harness runs the training loop directly.

1mo ago research

High Bandwidth Flash Emerges as a Third AI Memory Tier at Hot Chips 2026

A Hot Chips 2026 tutorial mapped out how HBF — NAND flash packaged like HBM — could extend GPU memory capacity for MoE inference without replacing HBM. vLLM expert offloading and sparse KV cache are the first target workloads. No products ship yet.

1mo ago research

NVIDIA Sets Server-Grade Hardware Requirements for RISC-V CUDA at Hot Chips 2026

NVIDIA's Hot Chips 2026 presentation outlines the hardware checklist RISC-V CPUs must clear before CUDA can run on them. Few existing RISC-V systems qualify; SiFive is the first named partner with a working demo.

1mo ago policy

Anthropic Screens Candidates on Mission vs. Equity as Dario Flags Culture Drift

Anthropic's hiring process includes a culture interview asking candidates to weigh the company's mission against future share price. CEO Dario Amodei has questioned whether newer employees are joining for the right reasons, according to Axios.

1mo ago research

NVIDIA AVO Scores 100% on ARC-AGI-3 as the Harness Emerges as the Real AI Moat

NVIDIA's AVO architecture achieved a perfect score on ARC-AGI-3 by building around persistent state, tool use, and long-horizon recovery. Its VP says the industry has been watching the wrong thing: it's the harness, not the model.

1mo ago funding

Starcloud Raises $450M for Orbital AI Compute While Racing Falcon 9's 2028 Retirement

The orbital data centre startup closed a $250M extension to reach $450M raised at a $2.3B valuation, with NVIDIA among the investors. It is the only company running an H100 GPU in orbit. Its biggest near-term risk is the launch calendar.

1mo ago benchmark

LiveBench Cost-Per-Task Metric Shows 9x Gap: Gemini 3.7 Flash High Beats Claude Fable 5 on Value

Claude Fable 5 leads LiveBench at 83.0 but costs $1.439 per successful task. Gemini 3.7 Flash High scores 78.8 for $0.157 — nine times cheaper for four fewer quality points. The August 2026 data reframes model selection as an economics problem.

1mo ago research

Microsoft ThinkingBox: Best Agent Hits 65% Pass@1, Collapses to 25% Across 20 Consecutive Trials

Microsoft's ThinkingBox-Bench tests 12 models across 507 business tasks run 20 times each. The top model's first-attempt success rate drops 40 points when required to succeed consistently, what Microsoft calls the discovery-reliability gap.

1mo ago benchmark

Ox Alpha Posts 80% on DeepSWE: 15 Points Ahead of Fable 5, Zhipu AI Suspected

Developer testing puts the anonymous OpenRouter model stealth/ox-alpha at 80% on a 10-task DeepSWE run, ahead of Claude Fable 5 Max at 65%, GLM-5.3 Max and Grok 4.6 at 62%, and GPT-5.6 Sol at 52%. A second subset run returned 63%.

1mo ago model

Anthropic's Best Model Is Struggling to Win Users as Cheaper Rivals Close the Gap

The Financial Times reports that Claude Opus 5, Anthropic's top-performing model, is seeing weak user uptake while cheaper tools capture the market. The finding arrives three months before Anthropic's expected IPO.

1mo ago research

ArXiv Paper Claims Non-Invasive EEG Can Decode What You Are Silently Reading

A 45-page ML paper (arXiv:2608.20186) reports successful decoding of silently read text from surface EEG alone, no surgery required. The claim is extraordinary. Independent replication will determine whether it holds.

1mo ago release

OpenAI Cuts GPT-5.6 Sol 20%: First Model ID Reprice Since 2022, Three-Month Window Creates API Uncertainty

OpenAI dropped GPT-5.6 Sol pricing by 20% across the API, Codex credits, and ChatGPT Work. It is the first time OpenAI has changed pricing on a live model ID since 2022. The discount runs three months, leaving developers uncertain about what happens in November.

1mo ago benchmark

LiveBench: Every Frontier Model Drops 15-26 Points in Agentic Coding vs Overall Score

The latest LiveBench leaderboard shows Claude Fable 5, GPT-5.6 Sol, GPT-5.5 Thinking, and Claude Opus 5 all score 15 to 26 points lower in Agentic Coding than in their headline number. The gap is largest for reasoning-heavy models. Kimi K3 open-weight ties Claude Fable 5 in the category at half the cost.