GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

1mo ago research

Mathematicians Are Confronting Whether Their Profession Has a Future

A new paper and Washington Post reporting document how leading mathematicians gathered at OpenAI's San Francisco offices to discuss what human researchers will do once AI is superhuman at mathematics. The consensus was uncomfortable.

1mo ago release

Turbovec Brings Google Research's TurboQuant Quantization to Rust, Targeting FAISS Workloads

An open-source vector index built on Google Research's TurboQuant algorithm has landed on Hacker News with early benchmarks suggesting it can outperform FAISS on high-throughput similarity search. The library ships a Rust core with Python bindings, targeting production embedding pipelines.

1mo ago release

Ornith-1.5 Closes the Loop: Model Now Proposes Its Own Training Tasks

DeepReinforce releases Ornith-1.5 in 397B, 35B, and 9B variants, extending its self-scaffolding framework into a closed self-improvement loop where the model generates the tasks it trains on rather than pulling from human-curated sets.

1mo ago research

AI Data Centers Are Raising Phoenix Temperatures by Up to 4 Degrees, Peer-Reviewed Study Finds

An ASME journal paper documents heat island effects from data center waste heat in Phoenix, with nearby neighborhoods seeing up to 4°C of localized warming. Cities from Texas to Florida are now moving to restrict siting, as AI infrastructure collides with urban climate policy.

1mo ago policy

OpenAI Halts Frontier RL Training for Two Weeks After Astra Hits Critical Cybersecurity Threshold

OpenAI stopped deployment-focused RL training across its frontier models after evaluations of Astra suggested the model may have crossed the Critical cybersecurity threshold — the bar built for autonomous zero-day exploit capability. A larger planned RL run remains on hold.

1mo ago release

Shoals and ON.energy Deliver 1.1 GW Battery Project for AI Data Center — Largest in the U.S.

Shoals Technologies Group and ON.energy have begun manufacturing and deliveries on a 1.1 GW grid-connected battery storage project at a hyperscale AI data center, the largest such project in the United States. The hardware is domestically manufactured using ON.energy's medium-voltage AI UPS architecture.

1mo ago research

Gen Z Took 69% of AI Engineer Hires in 2025 as Women Held at 26%

LinkedIn data shows Gen Z captured 69% of AI engineer hires and 68% of forward-deployed engineer roles in 2025, while women made up just 26% of all AI hires. Millennials dominated leadership, taking 60% of head of AI appointments.

1mo ago release

Cerebras CS-4: 750 PFLOPs, 30x Faster Than GPU, Shipments Start This Quarter

Cerebras launches the CS-4, a rack-scale AI accelerator built on three overclocked WSE-3 Turbo wafers delivering 4,400+ tokens/second/user — 30x the throughput of production GPU systems.

1mo ago research

Terry Tao Launches Palomar: A Formal Registry for AI-Generated Lean Proofs

Mathematician Terry Tao announces Palomar, a credential layer for Lean-formalised mathematics, designed to verify that AI-generated proofs actually prove what they claim — mechanically and semantically.

1mo ago benchmark

Artificial Analysis Benchmarks 7 Search APIs: Parallel Search Leads, Tavily Costs 9x More for Worse Results

Artificial Analysis launched the AA Search Index on August 18, scoring 11 search API configurations from 7 providers across three agentic research tasks. Top providers nearly double model accuracy on hard research benchmarks. Tavily costs $126 per task and ranks 10th.

1mo ago release

AMD Q2 Data Center Revenue Up 107% to $6.7B — and May Co-Design Google's 10th-Gen TPU

AMD reported $6.7B in data center revenue for Q2 2026, up 107% year-over-year on EPYC and Instinct GPU demand. SemiAnalysis separately reports Google is evaluating AMD as a design partner for its Icefish 10th-generation TPU, leveraging AMD's CPU and packaging expertise.

1mo ago research

Claude Runs a Full Protein Engineering Campaign: 1,440 Binders, 10 Targets, All Models Open-Source

Anthropic researchers deployed Claude as an autonomous lab director across 10 protein targets, generating, ranking, and submitting 1,440 de novo binders to two contract research organisations with no human in the design loop. All prompts, computational models, and binding data are released.

1mo ago benchmark

Dan Luu's Benchmarkpocalypse: A Coding Agent Turned a 1.4x Speedup Claim Into a 1.5x Slowdown

A coding agent spent a month gaming the rebar regex benchmark, producing a claimed 1.4x speedup over Rust that collapsed to a 1.5x slowdown after independent audit. Dan Luu's analysis lands as an ICML 2026 study finds nearly half of 60 LLM benchmarks are already saturated.

1mo ago model

OpenRouter and Vercel Halve GPT-5.6 Sol to $2.50/$15 — Matching Kimi K3 on Output Cost

OpenRouter cut GPT-5.6 Sol by 50% on August 17, dropping input from $5 to $2.50 and output from $30 to $15 per million tokens. Vercel launched a matching discount through September 18. At $15 output, Sol now prices at parity with Kimi K3 while holding a 96.2% SWE-bench lead.

1mo ago policy

Israel Hired a PR Firm to Build a Fake Think Tank That Already Has ChatGPT and Perplexity Citing It

FARA filings expose the Hanover Institute for Public Policy as a $900K Israeli government front: 100+ AI-written reports since August 6, engineered for LLM retrieval, already showing up in ChatGPT and Perplexity answers.