GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

4mo ago model

Microsoft Cancels Thousands of Claude Code Licenses by June 30 — Copilot CLI Takes Over the Agentic Stack

Microsoft's Experiences and Devices group — responsible for Windows, Microsoft 365, and Teams — has been ordered off Claude Code by June 30. EVP Rajesh Jha cited control over Microsoft's own repos and security requirements. Developer preference still runs toward Anthropic.

4mo ago research

Mythos Chained 2 Unknown macOS Kernel Bugs Into a Working Exploit in 5 Days — Apple Is Investigating

Security researchers at Palo Alto firm Calif used early Mythos access in April to discover two unknown macOS kernel vulnerabilities and build a privilege escalation exploit that bypasses Apple's memory integrity protections. Apple has confirmed it is investigating.

4mo ago research

146,932 Hallucinated Citations Entered the Scientific Record in 2025. Now Venues Are Fighting Back.

A new paper auditing 111 million references across arXiv, bioRxiv, SSRN, and PubMed Central finds a sharp post-2024 rise in AI-fabricated citations with no sign of plateauing. ICLR 2026 desk-rejected 647 submissions. ICML and ACM CCS have announced similar policies. arXiv's CS chair is publicly affirming author liability.

4mo ago policy

Ontario Audit: 12 of 20 Government-Approved AI Medical Scribes Got the Drug Wrong

Ontario's Auditor General found that most AI scribes approved for use by the province's doctors hallucinated treatment plans, misidentified drugs, and missed mental health details in standardised tests. The province weighted 'accuracy of medical notes' at 4% of the procurement score. Domestic presence in Ontario was weighted at 30%.

4mo ago release

OpenAI Brings Codex to Mobile: 4 Million Weekly Users Can Now Direct AI Coding Agents From Their Phones

OpenAI launched Codex in the ChatGPT mobile app on May 14, giving developers remote access to running agents, live diffs, approval queues, and terminal output from any device. A secure relay layer keeps the host machine unreachable from the public internet.

4mo ago policy

Six State AGs Ask SEC to Probe Sam Altman's Investments Before OpenAI's IPO

Republican attorneys general from six states and the House Oversight Committee have opened parallel investigations into whether Sam Altman's personal stakes in Helion, Stoke Space, and other startups constitute conflicts of interest with OpenAI's commercial relationships, putting direct pressure on the company's IPO path.

4mo ago policy

Apple-OpenAI Partnership Fractures as OpenAI Explores Legal Action

Bloomberg reports the Apple-OpenAI relationship has deteriorated badly enough that OpenAI is now exploring legal options against Apple, as Gemini replaces ChatGPT as Siri's AI backbone and OpenAI builds competing consumer hardware with Jony Ive.

4mo ago policy

Anthropic Commits $200M With the Gates Foundation to Vaccines, Global Health, and K-12 Education

Anthropic and the Gates Foundation are deploying $200M over four years to accelerate vaccine discovery for polio and HPV, improve disease forecasting in low-income countries, and deliver AI tutoring to students in sub-Saharan Africa and India.

4mo ago release

OpenRouter Ships Human-in-the-Loop Approval Gates and Session-ID Stickiness for Production Agents

OpenRouter's May 11 update adds two primitives production agent teams have been building in-house: HITL tool calls that pause an agent loop for human review, and session-id stickiness that routes multi-turn conversations to the same provider backend for better cache performance.

4mo ago release

Apple Rebuilds Siri as an Always-On Agent for iOS 27, on a $1B Gemini Foundation

Apple is scrapping the legacy Siri architecture for iOS 27 and replacing it with a system-wide AI agent built on Google Gemini at $1B/year. A dedicated Siri app, Dynamic Island integration, and third-party agent extensions debut at WWDC 2026 on June 8.

4mo ago funding

Cerebras Prices IPO at $185, Raises $5.55B — Largest US Tech Debut Since Uber in 2019

AI chipmaker Cerebras priced its Nasdaq IPO at $185 per share on May 13, above the $150-160 marketed range, on 20x oversubscription. Fully diluted valuation: $56.4B. Ticker CBRS begins trading May 14.

4mo ago policy

Princeton Ends 133-Year Honor Code Precedent: AI Cheating Forces First Exam Proctoring Since 1893

Princeton faculty voted nearly unanimously on May 11 to require proctors for all in-person exams, ending a practice students won in 1893. Generative AI made cheating easier and harder to catch simultaneously. Effective July 1.

4mo ago release

Anthropic Launches Claude for Small Business: 15 Agentic Workflows, 7 Connectors, No IT Required

Anthropic's fifth market-specific Claude bundle integrates with QuickBooks, PayPal, HubSpot, Canva, DocuSign, and Google and Microsoft Workspace in a single toggle install. Small businesses account for 44% of US GDP; most AI adoption had stopped at the chat window.

4mo ago benchmark

Seven Days After OpenAI's Clinical Benchmark Dropped, a Healthcare Specialist Topped It: Corti 60.5, ChatGPT for Clinicians 59.0

Corti Symphony scored 60.5 on HealthBench Professional, OpenAI's own physician-authored benchmark — surpassing ChatGPT for Clinicians (59.0) and GPT-5.4 with extended reasoning (48.1). On the adversarial red-teaming slice, Corti scored 87.7. GPT-5.4 scored 30.3. Physicians scored 30.

4mo ago research

Microsoft's 100-Agent AI Harness Found 16 Windows CVEs — Including an Unauthenticated IKEv2 RCE

MDASH, Microsoft's multi-model agentic scanning harness, found 16 vulnerabilities patched in May Patch Tuesday — including 4 critical RCE flaws. It topped the CyberGym public benchmark at 88.45%, five points above the next entrant, and achieved 100% recall on confirmed Windows kernel bugs.