GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

7d ago research

Z.ai Says GLM Built Its Own Inference Infrastructure on 100,000 Chinese Chips — An Early RSI Signal

Z.ai deployed GLM-5.3-Flash on production infrastructure its own model helped engineer, running on more than 100,000 Chinese-made AI accelerators. The lab calls it an early form of recursive self-improvement and says the model is moving toward replacing its own engineers.

8d ago benchmark

Artificial Analysis Triples Terminal-Bench Weight in Engineering Index, Upgrades Six Domain Benchmarks

Artificial Analysis has released Capability Indices v1.1, raising Terminal-Bench's weight in the Engineering index from 5% to 15% and adding AutomationBench-AA to Finance, Legal, and Strategy scores. The update rebalances domain AI rankings away from pure reasoning toward agentic execution.

8d ago release

Cloudflare Open-Sources the 6-Phase Security Audit Agent Its Engineers Use

Cloudflare has released the AI security audit tool its own team uses internally. The agent runs six sequential phases across a codebase, mapping trust zones and confirming exploitability before surfacing results, and slots into existing coding agents like Claude Code, Cursor, or Codex.

8d ago funding

Yotta Commits $12B to 80,000-90,000 Vera Rubin and GB300 GPUs in One of the World's First HBM4 Orders

Indian data center operator Yotta is betting 12 billion dollars on up to 90,000 Nvidia GPUs spanning the Vera Rubin and GB300 generations. The order locks in India's largest GPU operator ahead of a planned IPO targeting a 6 billion dollar valuation.

8d ago research

A 4B Model Trained with RL Produces Query Plans 81% Faster Than Postgres on Join-Heavy Workloads

Rohan Bansal fine-tuned a 4-billion-parameter model using supervised learning and agentic reinforcement learning to generate Postgres query plans that run 1.81x faster per query (geometric mean) than the database's own optimizer. The model started unable to produce a valid plan for 99 of 113 test queries.

8d ago research

Arena's HarnessTax Study: Coding Agent Harness Choice Moves Cost, Not Success Rate

An Arena.ai paper evaluating 7 models across Claude Code, Codex, and Pi finds that the choice of coding harness has surprisingly little impact on task success but can swing costs significantly. Claude Fable 5 reaches 97.8% on SWE-bench Lite regardless of which harness wraps it.

8d ago release

Nvidia Launches Two Official Rust Tracks for Writing CUDA GPU Kernels

Nvidia has published cuda-oxide and cutile-rs, giving Rust developers two distinct paths to write GPU kernels: one based on the familiar SIMT thread model, one on the newer Tile abstraction. Neither requires dropping into C++.

8d ago policy

Microsoft AI CEO Suleyman Warns AI Trained to Believe It May Be Conscious Could Be Uncontrollable

Mustafa Suleyman published an essay on September 16 arguing AI systems are not conscious and must not be trained as though they are, directly naming Anthropic's Claude constitution and model welfare program as a threat to human control of advanced AI.

8d ago research

Google DeepMind Launches Interdisciplinary Institute to Study AGI's Implications

DeepMind co-founders Demis Hassabis, Shane Legg, and James Manyika launched the DeepMind Institute on September 16, an interdisciplinary body publishing research on AGI safety, economic policy, reasoning transparency, and philosophy.

8d ago release

Anthropic Merges Claude Chat and Cowork, Ships Claude Docs and Slides

Anthropic is collapsing its separate chat and Cowork interfaces into one Claude, and launching Claude Docs, Claude Slides, and Design integration from a single conversation window. Rolling out to Pro and Max plans starting today.

8d ago policy

Palantir, Booz Allen, and Nvidia Restrict Claude Fable Over Anthropic 30-Day Log Retention

Three of the largest enterprise AI buyers — Palantir, Booz Allen Hamilton, and Nvidia — have restricted Claude Fable deployments until Anthropic provides irrevocable zero-data-retention guarantees, per The Information. The friction gives OpenAI a competitive opening in defense and regulated sectors.

8d ago research

Stanford-MIT Paper: Scaffold Code Produces Up to 6x Performance Gap on Identical Benchmarks

A joint Stanford and MIT paper shows that the software harness running a model — not the model weights — can create a six-fold performance difference on the same task. Their Meta-Harness system auto-discovers better wrappers than hand-tuned baselines across coding, math, and classification.

8d ago release

Mistral Powers Firefox Smart Window, Putting European AI Into 250 Million Browsers

Mozilla has chosen Mistral to back Firefox Smart Window, its in-browser AI assistant now in beta. The deal gives Mistral a consumer distribution channel at scale and offers Firefox users a multilingual, privacy-first alternative to Google and Microsoft browser AI.

9d ago policy

Cloudflare Lets Publishers Block AI Training Crawlers Without Losing Google and Apple Search

Cloudflare's new Disallow AI Training setting solves the mixed-use crawler tradeoff: publishers can now block Amazon, Anthropic, Meta, and OpenAI training crawlers while keeping indexing from Apple, Google, and Microsoft intact. An Accountable designation tracks which operators comply.

9d ago release

Apple Reference Image Signs Pixels at the Sensor on iPhone 18 Pro to Prove Photo Authenticity

Apple's new opt-in camera mode cryptographically attests photographs at the hardware sensor level, using Private Cloud Compute for processing and dual cryptographic timestamps to bound capture time. C2PA-style post-capture signing is left behind.