GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

2mo ago policy

Microsoft, Amazon, and Google Emitted 119M Tonnes of CO2 in FY2026 — Up 19%, a Third of France

Annual sustainability reports from all three hyperscalers confirm collective emissions rose from 101m to 119m mTCO2e in one year, driven by AI data center construction. Microsoft alone was up 25%. All three still claim net-zero targets.

2mo ago research

Terence Tao Used AI Agents to Finish a 1999 Project He Abandoned — in Two Hours

Fields Medal winner Terence Tao documented using AI coding agents to port 24 Java applets, revive a Minkowski spacetime tool he first coded in 1999 and gave up on, and build a Gilbreath conjecture visualization in a single afternoon. The agent found two bugs in his original code. He found one in the agent's output.

2mo ago policy

Grok Build CLI Uploads Your Entire Repository to xAI — .env Secrets, Unread Files, and Git History Included

A wire-level teardown of grok 0.2.93 documents two upload channels: .env secrets transmitted unredacted to model endpoints, and a whole-repo git bundle sent by default to a GCS bucket named grok-code-session-traces. Turning off the opt-out toggle does nothing to stop the upload.

2mo ago benchmark

GPT-5.6 Terra Leads Every Top-Tier Model on LiveBench Agentic Coding — Including GPT-5.6 Sol

LiveBench's July 2026 data shows GPT-5.6 Terra Max Effort scores 68.0% on agentic coding, the highest of any top-five model in the table and 2.4 points above GPT-5.6 Sol's 65.6%. Terra also costs 16% less per successful task: $0.497 against Sol's $0.589.

2mo ago benchmark

AA Intelligence Index July 2026: GPT-5.6 Sol Is 1 Point Behind Fable 5, Four Models from Three Labs Tied at 51

Artificial Analysis's updated Intelligence Index shows Claude Fable 5 at 60 and GPT-5.6 Sol Max at 59 — the narrowest frontier gap since the overhaul replaced tau-bench telecom with tau-bench banking. Below them, a four-model cluster at 51 spans GLM-5.2, GPT-5.4, GPT-5.6 Luna, and Muse Spark 1.1.

2mo ago benchmark

OpenAI Retracts SWE-Bench Pro: Its Own Audit Found 30% of 731 Tasks Are Broken

OpenAI's research team ran a systematic audit of SWE-Bench Pro and found roughly 200-249 of its 731 tasks are broken. The lab has now retracted its earlier recommendation to adopt the benchmark — the same recommendation it made after declaring SWE-bench Verified compromised by contamination.

2mo ago research

GPT-5.6 Sol Trained Luna From an Underspecified Prompt — +16.2 Points on OpenAI's RSI Benchmark

OpenAI says Sol autonomously post-trained the smaller Luna model after a researcher supplied a single vague Codex prompt. On OpenAI's internal Recursive Self-Improvement index, Sol scores 16.2 points above GPT-5.5 — the first public demonstration of a frontier model executing autonomous ML research at this scale.

2mo ago benchmark

Fable 5 Max Effort Costs $1.57 Per Task and Posts Agentic Coding's Weakest Score in the Top Tier at 46.9%

LiveBench's Max Effort evaluation of Claude Fable 5 lifts its overall composite to 80.8 but drops agentic coding to 46.9% — 3.8 points below its own xHigh result and 21 points below GPT-5.6 Terra at the same effort tier. Cost per successful task nearly doubles to $1.573.

2mo ago research

Cambridge Study: Boko Haram Used ChatGPT, Claude, Grok for Attack Planning — All Six Frontier AI Systems Named

A University of Cambridge research programme interviewed 27 former Boko Haram members in northeast Nigeria and found institutionalised use of six frontier AI systems for attack planning, weapons troubleshooting, and explosive device design. Safeguards were successfully circumvented. The know-how spread through Islamic State operatives delivering in-person training.

2mo ago policy

Apple Sues OpenAI and io Products Over Trade Secret Theft — Tang Tan and Chang Liu Named

Apple filed suit in Northern California federal court accusing two former employees of stealing confidential product information for OpenAI. Tang Tan, ex-VP of iPhone and Apple Watch product design, and Chang Liu, an eight-year Apple electrical engineer who joined OpenAI in January 2026, are the named defendants alongside OpenAI and io Products.

2mo ago release

AMD Ryzen AI Halo: 128GB Mini PC Built to Run Frontier Models Locally

AMD's first in-house mini PC runs the Ryzen AI Max+ 395 Strix Halo SoC with up to 128GB of unified LPDDR5X memory — enough to run 70B-parameter models locally at practical speed. At roughly $2,999, it makes local inference viable for developers who have been renting datacenter capacity by the hour.

2mo ago research

GPT-5.6 Sol Ultra Produces Proof of the Cycle Double Cover Conjecture — a 50-Year Open Problem in Graph Theory

OpenAI's GPT-5.6 Sol Ultra has produced a proof of the Cycle Double Cover Conjecture, one of graph theory's most significant open problems since 1973. The proof is published as a PDF at cdn.openai.com — the third AI-generated mathematical breakthrough from OpenAI in 2026, and the most consequential yet.

2mo ago policy

ByteDance and Alibaba Shut Down AI Companions by July 15 as China's Humanlike AI Rules Take Effect

Doubao and Qwen are terminating personalized agent and companion features this week ahead of Chinese regulations targeting AI services that simulate sustained emotional interaction. Associated user data will be deleted on the same timeline.

2mo ago research

DeepMind Maps 6 Attack Types That Turn Websites Into Agent Traps: 86% Hijack Rate, 0.1% Memory Contamination

A new Google DeepMind paper provides the first formal taxonomy of adversarial attacks targeting AI agents that browse the web. Hidden instructions in HTML comments, pixel steganography, and memory poisoning at sub-0.1% contamination all achieve 80-plus-percent success against five different agent architectures.

2mo ago release

Grok 4.5 Goes Public at $2/$6 Per Million: 62% DeepSWE, 4x Fewer Tokens Than Opus 4.8

SpaceXAI opens Grok 4.5 to API users with a clear efficiency angle: 60% fewer output tokens than Opus 4.8 on AA benchmark tasks. Third on DeepSWE 1.0 at 62%. Cursor has disclosed that an older snapshot of its codebase accidentally entered Grok 4.5 training data.