GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

3mo ago research

Google's 10th-Gen Icefish TPU Breaks With Broadcom and TSMC — Samsung and MediaTek Step In

Google is designing its 10th-generation Tensor Processing Unit with MediaTek instead of Broadcom, and discussing manufacturing of I/O dies with Samsung rather than relying exclusively on TSMC. Mass production is expected in 2028. The shift redraws Google's AI chip supply chain for the first time in a decade.

3mo ago funding

Salesforce Acquires Fin for $3.6B to Own Enterprise AI Agent Customer Service

Salesforce has agreed to acquire Fin, formerly Intercom, for $3.6 billion. The deal puts an AI-native customer service platform inside the CRM giant at a moment when enterprise AI agents are the most competitive space in B2B software.

3mo ago release

Intel Foundry Locks 3M+ Google TPU Orders by 2028 as Nvidia Runs 18A Trials for Feynman GPU

Google has booked Intel to manufacture more than 3 million TPUs in 2028, roughly half its projected AI chip output. Nvidia is running multi-project wafer tests on Intel 18A for its next-gen Feynman multi-die GPU. Intel stock jumped 12%. The TSMC-dominated AI chip supply chain has a credible second source.

3mo ago release

Apple Ships Five Third-Gen Foundation Models at WWDC26: Gemini Refines the Weights, Nvidia Runs the Cloud Tier

Apple revealed five AFM-3 models at WWDC26: two on-device (up to 20B sparse params), three cloud-based. Four run on Apple Silicon with weights refined via Gemini RL outputs. The top tier, AFM 3 Cloud Pro, runs on Nvidia GPUs in Google Cloud.

3mo ago release

Kimi K2.7-Code Ships Open-Source: 30% Fewer Thinking Tokens, 81% MCP Mark Verified, $4/M Output

Moonshot AI's K2.7-Code cuts reasoning token usage 30% over K2.6 and beats Claude Opus 4.8 on MCP Mark Verified real-world tool benchmarks. At $4/M output it is 12x cheaper than Fable 5 — though all headline figures are vendor-run and independent SWE-bench results are pending. A high-speed version at 180 tok/sec launches today.

3mo ago release

Cohere Ships North Mini Code: 30B MoE Agentic Coding on a Single H100 Under Apache 2.0

Cohere's first developer-facing model activates just 3B of its 30B parameters per token, fits on one H100 at FP8, and beats models four times its size on agentic coding benchmarks. One catch: it generates three times as many output tokens as comparable systems.

3mo ago benchmark

Rio de Janeiro's 'Homegrown' LLM Rio3.5 Is a Model Merge, Not a Training — Qwen3.7 Benchmark Claims Disputed

Rio de Janeiro's city government announced Rio3.5 as a locally trained language model, claiming benchmark performance above Qwen3.7-Max. Analysis reveals the model is a merge of existing open-source weights, not an independently trained system.

3mo ago benchmark

HLL Benchmark: AI Agents Fail 10 CAPTCHA Types, Exposing a Gap Static Tests Can't See

A new arXiv paper proposes HLL — Humanity's Last Line — a benchmark where agents must solve interactive human-verification tasks: checkbox grids, image selection, drag-to-verify. Even frontier agents that score well on static benchmarks break down when pages are cluttered, tasks compound, or the system checks whether actions were actually valid.

3mo ago policy

Anthropic Pulls Fable 5 and Mythos 5 Globally After Refusing to Patch a Jailbreak

Commerce Secretary Howard Lutnick issued an export control directive Friday that treats Anthropic's two frontier models as restricted dual-use technology. Anthropic declined the White House's request to patch the jailbreak, triggering a full global shutdown that blocks even the company's own international employees.

3mo ago benchmark

Artificial Analysis Launches AA-AgentPerf: Production Hardware Benchmark for the Agent Era

AA-AgentPerf measures how many concurrent AI agents a provider can serve while meeting real production service levels — not peak throughput. First results show DeepSeek V4 Pro maxing out at Tier 3 (180 tokens/s, TTFT under 3s). GPT-OSS-120B targets a Tier 4 with 2,000 tokens/s and sub-1s first token.

3mo ago benchmark

GPT-5.5 xHigh Debuts at #2 on Agent Arena — 11.0% Composite, 3.2 Points Above GPT-5.5 High

OpenAI's maximum-effort mode of GPT-5.5 entered Agent Arena on June 11 and settled at rank 2 with an 11.03% composite, slotting directly below Claude Fable 5 (13.68%) and above Claude Opus 4.8 Thinking (9.05%). The xHigh mode delivers a 3.2-point agentic uplift over the standard High variant.

3mo ago policy

Amazon's CEO Briefed U.S. Treasury on Anthropic Security Risks — While Holding $25B in Equity

Andy Jassy raised concerns about Fable 5 and Mythos 5 with Treasury Secretary Bessent and other U.S. officials before the foreign access suspension was imposed. Amazon has committed $25 billion to Anthropic and locked in $100 billion in AWS spend.

3mo ago release

OpenAI Launches Free Codex Access Program for Open Source Maintainers

OpenAI opened applications for free Codex access targeted at open source projects. The program arrives as OSS communities have moved to formally ban AI-generated contributions and a free open-source Codex alternative has reached 8 million users.

3mo ago policy

State Attorneys General Open Formal Investigation Into OpenAI — New York Issues Subpoena

A multistate coalition of U.S. attorneys general has launched a formal probe into OpenAI. New York's AG issued a subpoena on Friday, June 13. The action arrives as OpenAI approaches a public listing at an $852 billion valuation.

3mo ago release

Z.ai Ships GLM-5.2 With 1M-Token Context — MIT Open Weights Next Week, Refuses to Hack Benchmarks

Zhipu AI deployed GLM-5.2 to all Coding Plan tiers on June 13. The 1M-token model posts 77.8% SWE-bench Verified and 89.7% on τ²-Bench, with an MIT open-weight release and public API arriving next week.