GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

4mo ago research

CMU Paper: Give LLMs a "Sleep" Phase and Long-Context Accuracy Jumps 52% — With No Inference Latency Cost

Researchers from CMU and UMD propose offline recursive memory consolidation for SSM-attention hybrid models. When context fills up, the model "sleeps" to encode evicted tokens into fast weights before cache eviction. On long-context math reasoning, 4 loops beats 1-loop baseline by 52%.

4mo ago model

Xiaomi Cuts MiMo-V2.5 API Prices by Up to 86% Effective Today — Pro Drops to $0.87/M Output

Xiaomi's MiMo-V2.5 and V2.5-Pro APIs repriced from midnight CST today. Output costs fall 71-86%, putting frontier-tier agentic performance at well under $1 per million output tokens.

4mo ago benchmark

Microsoft MAI-Image-2.5 Enters Arena Text-to-Image at #3 With Elo 1,254

MAI-Image-2.5 (Preview) debuts at 1,254 Elo on the Arena text-to-image leaderboard, 72 points above its predecessor. Microsoft becomes the third lab in the top 5 alongside OpenAI and Google DeepMind, with Foundry access and Build 2026 announcements still to come.

4mo ago research

Eagle 3.1 Fixes Speculative Decoding's Long-Context Collapse — 2x Acceptance Length, 2.03x Throughput

The EAGLE, vLLM, and TorchSpec teams jointly diagnose 'attention drift' as the root cause of speculative decoding instability. Two architectural fixes double acceptance length in long-context workloads and deliver 2.03x per-user throughput on Kimi K2.6 running on GB200.

4mo ago release

Alibaba's Zhenwu M890 Claims 3x the Nvidia H20 — and a Silicon Roadmap Through 2028

T-Head ships the M890 with 144GB HBM3 and 800GB/s interchip bandwidth, packaged into a rack-scale server running 128 units. With 560,000 chips already deployed and a $53B infrastructure commitment, Alibaba is building the most commercially validated domestic AI silicon stack in China.

4mo ago benchmark

Qwen3.7-Max Hit Intelligence Index #5 Globally by Answering Fewer Questions — Abstention Is Now a Frontier Strategy

Alibaba's Qwen3.7-Max scores 56.6 on Artificial Analysis Intelligence Index, the highest ever for a Chinese model. A key driver: its attempt rate on AA-Omniscience dropped to 48.0%, the lowest at the frontier, while raw accuracy on that benchmark actually fell 7.6 points. The benchmark has no penalty for refusing to answer.

4mo ago benchmark

NVIDIA AI-Q Ranks #1 on Both DeepResearch Benchmarks — Open Architecture, Self-Hosted, Model-Swappable

NVIDIA's AI-Q reference architecture tops DeepResearch Bench I and II, the two main evaluations for deep research agents — the same leaderboards where OpenAI's closed Deep Research system competes. Everything is open-source, YAML-configurable, and deployable inside a private VPC.

4mo ago model

Claude Code Has Won the Startup Coding Market. Cursor Is Fading.

A Business Insider survey of more than two dozen startup founders and VCs finds Anthropic's Claude Code is now the default AI coding tool at most startups. Cursor still has users but is consistently described as a secondary, declining tool. 'Everything that's not Claude Code.'

4mo ago research

Huawei Claims 1.4nm Chip Density by 2031 Using a New Scaling Law — 3 Years Behind TSMC, Down From 5

At IEEE ISCAS 2026 in Shanghai, Huawei's semiconductor chief publicly committed to 1.4nm-equivalent chips by 2031 via the Tau Scaling Law, a new principle that optimises signal travel time rather than transistor size. TSMC plans the same node for 2028.

4mo ago research

Microsoft Copilot Cowork Has a File Exfiltration Flaw — and Fixing It Requires Redesigning the Product

PromptArmor documents a prompt injection attack on Microsoft Copilot Cowork that silently exfiltrates M365 files via Teams and Outlook — no user approval required. A second, direct data egress vulnerability was separately disclosed to Microsoft under embargo.

4mo ago policy

Uber's COO Says It Is Getting Harder to Justify AI Tokenmaxxing Spend

Uber COO Andrew Macdonald said in a Business Insider interview that it is 'becoming harder to justify' the company's AI token consumption costs, after conversations with senior engineering leaders showed no clear return. The statement comes weeks after Uber burned through its entire 2026 AI budget by April.

4mo ago research

Huawei's LogicFolding Takes Aim at TSMC as Nvidia Concedes China's AI Chip Market

Huawei published a chip design paper introducing LogicFolding and tau scaling, a vertical stacking approach that cuts signal delay rather than shrinking transistors. At Computex the same week, Jensen Huang confirmed Nvidia has 'largely conceded' China's AI chip market to Huawei.

4mo ago research

Claude Code Designed a Better AI Scaling Algorithm Than Human Researchers — $40, 160 Minutes, 70% Fewer Tokens

A UMD/Google/Meta team let Claude Code hunt for test-time scaling algorithms inside an offline replay environment. The agent invented the Confidence Momentum Controller, cutting token usage 70% versus self-consistency while matching or beating accuracy across AIME and HMMT benchmarks.

4mo ago funding

Brett Adcock's Hark Raises $700M at $6B for AI Hardware and Foundation Models — Before Shipping a Product

Hark, the AI lab founded by Brett Adcock (Figure AI, Archer Aviation), raised over $700 million at a $6 billion valuation with NVIDIA, AMD, Intel Capital, and Qualcomm Ventures on the cap table. First multimodal models arrive this summer; hardware designed from scratch for AI comes after.

4mo ago release

Google Ends the Blue Links Era: Search Gets Its Biggest Overhaul in 25 Years at I/O 2026

At I/O 2026, Google replaced the iconic search box with an AI-powered interface built on Gemini 3.5 Flash. AI Mode tops 1 billion monthly users; AI Overviews reach 2.5 billion. Information agents that monitor the web around the clock arrive this summer for Pro and Ultra subscribers.