GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —

Live Feed

5mo ago benchmark

Claude Opus 4.7 Hits 1505 Arena ELO Tied With Gemini 3.1 Pro — Three Days After Joining the Leaderboard

Claude Opus 4.7 Thinking reached 1505 Arena ELO as of April 18, tied with Gemini 3.1 Pro Preview at the top of the Chatbot Arena leaderboard. Claude Opus 4.7 standard sits at 1503, matching Opus 4.6 Thinking. In the Coding Arena, Claude 4.7 Thinking leads all models at 1565.

5mo ago release

GitHub Copilot Freezes New Subscriptions and Locks Opus 4.7 to Paid Pro+ Tier

GitHub has paused new signups for all individual Copilot plans and stripped Opus model access from the Pro tier. Opus 4.7 is now exclusive to Pro+. Anthropic Opus 4.5 and 4.6 are being retired from Pro+ entirely. Existing Pro subscribers can still upgrade.

5mo ago model

Moonshot AI Open-Sources Kimi K2.6: Leads HLE-Full and SWE-Bench Pro Over Every Frontier Proprietary Model

Moonshot AI releases Kimi K2.6, a 1-trillion-parameter open-weight MoE that posts 54.0 on HLE-Full with tools — the hardest frontier knowledge benchmark — ahead of GPT-5.4 (52.1), Claude Opus 4.6 (53.0), and Gemini 3.1 Pro (51.4). SWE-Bench Pro score of 58.6 also leads all comers.

5mo ago release

Microsoft Fairwater Goes Live Ahead of Schedule — $3.3B Wisconsin Campus with Hundreds of Thousands of NVIDIA GB200s

CEO Satya Nadella says Fairwater, Microsoft's flagship AI data center in Mount Pleasant, Wisconsin, is going live ahead of schedule. The $3.3B campus houses one of the densest concentrations of NVIDIA GB200 chips in the world, and Microsoft has already received approval to add 15 more buildings to the site.

5mo ago release

OpenAI Codex Chronicle Captures Your Mac Screen to Build AI Context — Stores It Unencrypted, Skips EU

Chronicle, a research preview inside Codex Desktop for macOS, takes periodic screenshots, sends them to OpenAI servers for processing, and writes summaries to local unencrypted Markdown files. The system helps Codex remember active tasks without user prompting — and is explicitly excluded from the EU, UK, and Switzerland.

5mo ago release

Alibaba Releases Qwen3.6 Max Preview Free — AA Score 52 Places It Atop Every Chinese Model

Alibaba released Qwen3.6 Max Preview today, scoring 52 on the Artificial Analysis Intelligence Index and ranking #2 among all tracked models. Free via Qwen Studio during preview, with paid API access coming via Alibaba Cloud. Beats Plus across six agentic coding and reasoning benchmarks.

5mo ago policy

Context.ai OAuth Breach Pivots Into Vercel — A Textbook AI Tool Supply Chain Attack

Attackers compromised Context.ai using Lumma infostealer malware, harvested OAuth tokens, replayed them to access a Vercel employee's Google Workspace, then pivoted into Vercel's internal systems. ShinyHunters has claimed credit. Mandiant and CrowdStrike are both on the case.

5mo ago benchmark

Berkeley RDI Backs AgentBeats With $1M+ Competition to Standardise Agent Evaluation Across 13 Domains

AgentBeats, a centralised agent benchmark registry built on the Agentified Agent Assessment (AAA) paradigm, launches with backing from Berkeley's Responsible Decentralized Intelligence lab and a $1M+ competition. The platform covers 13 agent domains including coding, healthcare, legal, DeFi, and agent safety, and integrates tau2-bench natively.

5mo ago research

AI Medical Chatbots Hit 95% in Lab Tests — Then Drop to 35% When Real Patients Talk to Them

New research and a BBC investigation expose a systematic accuracy collapse in AI health chatbots: near-perfect performance on clean, structured medical cases gives way to 35% accuracy when real users describe symptoms the way real people actually do — incomplete, distracted, non-linear. A single phrasing change can flip advice from 'rest at home' to 'go to hospital now.'

5mo ago benchmark

Honor's Lightning Runs 50:26 — Humanoid Robot Breaks Human Half-Marathon World Record in Beijing

At the 2026 Beijing E-Town Humanoid Robot Half-Marathon, smartphone maker Honor's 'Lightning' robot completed 21.1 km in 50 minutes 26 seconds, breaking the human men's world record of 57:20. Last year's winning robot took over 2 hours 40 minutes. Twelve months of progress, measured in asphalt.

5mo ago release

DeepSeek V4 Is Weeks Away: 1 Trillion Parameters, Huawei Chips, Apache 2.0

Multiple credible sources including Reuters and The Information point to a late-April launch for DeepSeek V4 — built entirely on Huawei Ascend hardware, estimated at 1 trillion parameters with 37 billion active per token, and expected under Apache 2.0. The model has already missed two earlier windows.

5mo ago model

Gemma 4's 2B Edge Model Runs Fully Offline on iPhone — 1.5 GB, Apple Neural Engine, No Cloud Required

Google's Gemma 4 E2B and E4B edge variants are running on-device on iPhones via the Apple Neural Engine, distributed through Locally AI and Google's own AI Edge Gallery. The 1.5 GB quantized footprint fits comfortably in a modern iPhone's memory budget, with all inference staying on-device.

5mo ago model

xAI Drops Grok 4.3 Beta at 500B Parameters With Zero Fanfare — 1T Version Due April 22

Grok 4.3 appeared in the model selector on April 17 with no blog post, no press release, and no announcement. Elon Musk confirmed the beta runs 500 billion parameters. The full 1 trillion parameter model completes training around April 22-23.

5mo ago research

LLMs De-Anonymize Online Users at 68% Match Rate — 9 in 10 Guesses Correct

A new paper demonstrates that LLMs can link anonymous accounts to real identities across platforms with 68% recall and 90% precision, shattering the assumption that pseudonyms protect privacy when public writing is available.

5mo ago policy

NSA Deploys Anthropic Mythos for Cyber Work While Its Own Parent Department Calls Anthropic a Supply Chain Risk

The National Security Agency is using Claude Mythos Preview for offensive and defensive cyber operations despite the Department of Defense — which oversees the NSA — formally flagging Anthropic as a supply chain risk. The split creates a public contradiction at the centre of the US government AI security posture.