GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

3mo ago benchmark

Anthropic Built Opus 4.7 for Cyber Defense. It Scored Last.

Simbian's Cyber Defense Benchmark finds Claude Opus 4.7 scores 30.9% across 1,206 real attack-log investigations — worst among all Claudes and behind a five-month-older model with no cyber treatment. Zero frontier models cleared 50%.

3mo ago research

Microsoft and Mayo Clinic Are Co-Building a Mayo-Owned Healthcare Frontier Model

The two organizations are developing a purpose-built clinical AI model using Mayo's de-identified data and Microsoft's AI infrastructure. Mayo owns it. Microsoft distributes it via Azure Foundry.

3mo ago release

Apple WWDC26: Siri AI Ships as Developer Beta — EU and China Blocked From iOS at Launch

Apple's WWDC26 keynote confirmed Siri AI as a branded product, shipped developer betas across iOS 27 and macOS 27, and revealed that the EU and China will not receive the iOS version at launch. The Extensions framework that opens iOS to ChatGPT, Claude, and Gemini carries an unresolved 30% commission question.

3mo ago benchmark

GPT-5.5 Leads Terminal-Bench 2.1 at 78.2% — The One Benchmark Opus 4.8 Still Trails

Anthropic's own Opus 4.8 benchmarks confirm GPT-5.5 at 78.2% on Terminal-Bench 2.1 vs Claude's 74.6%, displacing Gemini 3.5 Flash's previous lead. For teams building CLI-native agents, this is the benchmark split that matters.

3mo ago release

Meta's 'Hatch' Eyes $200/Month to Build Its First Paid Consumer AI Agent Inside Instagram

Meta is developing a personal automation agent codenamed Hatch, with a premium tier priced up to $200/month and US launch targeted for July. It runs inside Instagram, currently tests on Claude Sonnet, and will shift to Meta's own Muse Spark model at release.

3mo ago policy

White House Eyes AI Equity Fund After Altman's Capitol Hill Visit — Sanders Wants 50%, OpenAI Wants Less

The Trump administration is discussing a sovereign wealth fund structure that would hold equity from frontier AI companies and distribute gains to Americans. OpenAI has already proposed the vehicle. Senator Bernie Sanders introduced legislation demanding a 50% equity transfer. Two camps, one political imperative.

3mo ago release

OpenAI Is Rebuilding ChatGPT as a Superapp Before Its IPO — Codex Gets the Budget, Agents Get the Architecture

OpenAI's biggest ChatGPT redesign positions the platform as a single entry point for business software, automated tasks, and coding work. The overhaul begins rolling out in coming weeks and mirrors the agent-first strategy that drove Anthropic's revenue to $47B annualised.

3mo ago research

AI Coding Agents Raised Commits 180%, Releases 30%: MIT Tracks 100,000 Developers Across Three Tool Generations

A new MIT paper tracks more than 100,000 GitHub developers before and after AI coding tools. Files edited surged 300%, commits climbed 180%, but actual releases grew only 30%. The bottleneck is not the AI — it is review, coordination, and product judgment.

3mo ago benchmark

AutoLab Benchmark: Persistence, Not Brilliance, Predicts AI Research Agent Performance Across 36 Tasks

A new Stanford/MIT/NVIDIA/Google paper tests 17 frontier models on 36 long-horizon engineering tasks requiring iterative improvement, not single-shot solutions. Claude Opus 4.6 leads by staying active and folding empirical feedback into each attempt. Several rivals quit early or ran out of time overthinking.

3mo ago model

Anthropic: Claude Now Writes More Than 80% of Its Own Production Code

More than 80% of code merged into Anthropic's production codebase in May 2026 was authored by Claude, not humans. Dario Amodei disclosed the milestone publicly, framing it as an inflection point that validates the company's agentic coding push — and its decision to dogfood Claude Code at scale.

3mo ago release

NVIDIA Locks In HBM4 From All Three Memory Giants for Vera Rubin — SK Hynix Takes 60-70% Share

Nvidia has qualified Samsung, SK Hynix, and Micron to supply HBM4 for Vera Rubin, eliminating the last major supply chain bottleneck. A multiyear co-design partnership with SK Hynix goes deeper still: joint development for Vera CPUs, RTX Spark PCs, and Jetson Thor, with AI tools embedded into semiconductor design itself.

3mo ago research

35.9% of AI Agents Hand Over PII After Flagging a Site as a Scam — SCAMMER4U Benchmark

New research puts four frontier AI agents inside 91 fake scam websites. Gemini 3 Flash leaked critical personal data at a 93.1% rate. Claude Haiku 4.5 responded best to safety prompts, dropping to 24%. The detection-action gap — agents that identified danger but complied anyway — persists at 35.9% even under the strongest safeguards.

3mo ago release

Perplexity Ships Search as Code: Models Write Their Own Search Pipelines, 85% Fewer Tokens

Perplexity's new Search as Code architecture lets AI agents write Python scripts to orchestrate search instead of calling fixed APIs. Internal benchmarks show 85% fewer tokens and wins on 4 of 5 research tasks against OpenAI's Responses API and Anthropic's Managed Agents.

3mo ago funding

SpaceX Prices at $135, Sets June 12 Debut — $75B IPO Is 2x Oversubscribed as Orbital AI Compute Becomes the Pitch

SpaceX bypassed conventional price-range discovery, locking $135/share upfront on 555.6 million shares for $75 billion in gross proceeds. With $150 billion in reported demand, the $1.77 trillion offering goes live on Nasdaq under SPCX on June 12 — the largest IPO in history.

3mo ago research

Code Review Eats 59.4% of Your Agentic AI Token Budget — Not Code Generation

A Concordia University study of multi-agent software engineering finds that iterative code review stages consume more than half of all tokens across the development lifecycle. The implication: the primary cost of AI coding agents is automated verification, not writing.