Live Feed
CMU Paper: Give LLMs a "Sleep" Phase and Long-Context Accuracy Jumps 52% — With No Inference Latency Cost
Researchers from CMU and UMD propose offline recursive memory consolidation for SSM-attention hybrid models. When context fills up, the model "sleeps" to encode evicted tokens into fast weights before cache eviction. On long-context math reasoning, 4 loops beats 1-loop baseline by 52%.
Xiaomi Cuts MiMo-V2.5 API Prices by Up to 86% Effective Today — Pro Drops to $0.87/M Output
Xiaomi's MiMo-V2.5 and V2.5-Pro APIs repriced from midnight CST today. Output costs fall 71-86%, putting frontier-tier agentic performance at well under $1 per million output tokens.
Microsoft MAI-Image-2.5 Enters Arena Text-to-Image at #3 With Elo 1,254
MAI-Image-2.5 (Preview) debuts at 1,254 Elo on the Arena text-to-image leaderboard, 72 points above its predecessor. Microsoft becomes the third lab in the top 5 alongside OpenAI and Google DeepMind, with Foundry access and Build 2026 announcements still to come.
Eagle 3.1 Fixes Speculative Decoding's Long-Context Collapse — 2x Acceptance Length, 2.03x Throughput
The EAGLE, vLLM, and TorchSpec teams jointly diagnose 'attention drift' as the root cause of speculative decoding instability. Two architectural fixes double acceptance length in long-context workloads and deliver 2.03x per-user throughput on Kimi K2.6 running on GB200.
Alibaba's Zhenwu M890 Claims 3x the Nvidia H20 — and a Silicon Roadmap Through 2028
T-Head ships the M890 with 144GB HBM3 and 800GB/s interchip bandwidth, packaged into a rack-scale server running 128 units. With 560,000 chips already deployed and a $53B infrastructure commitment, Alibaba is building the most commercially validated domestic AI silicon stack in China.
Qwen3.7-Max Hit Intelligence Index #5 Globally by Answering Fewer Questions — Abstention Is Now a Frontier Strategy
Alibaba's Qwen3.7-Max scores 56.6 on Artificial Analysis Intelligence Index, the highest ever for a Chinese model. A key driver: its attempt rate on AA-Omniscience dropped to 48.0%, the lowest at the frontier, while raw accuracy on that benchmark actually fell 7.6 points. The benchmark has no penalty for refusing to answer.
NVIDIA AI-Q Ranks #1 on Both DeepResearch Benchmarks — Open Architecture, Self-Hosted, Model-Swappable
NVIDIA's AI-Q reference architecture tops DeepResearch Bench I and II, the two main evaluations for deep research agents — the same leaderboards where OpenAI's closed Deep Research system competes. Everything is open-source, YAML-configurable, and deployable inside a private VPC.
Claude Code Has Won the Startup Coding Market. Cursor Is Fading.
A Business Insider survey of more than two dozen startup founders and VCs finds Anthropic's Claude Code is now the default AI coding tool at most startups. Cursor still has users but is consistently described as a secondary, declining tool. 'Everything that's not Claude Code.'
Huawei Claims 1.4nm Chip Density by 2031 Using a New Scaling Law — 3 Years Behind TSMC, Down From 5
At IEEE ISCAS 2026 in Shanghai, Huawei's semiconductor chief publicly committed to 1.4nm-equivalent chips by 2031 via the Tau Scaling Law, a new principle that optimises signal travel time rather than transistor size. TSMC plans the same node for 2028.
Microsoft Copilot Cowork Has a File Exfiltration Flaw — and Fixing It Requires Redesigning the Product
PromptArmor documents a prompt injection attack on Microsoft Copilot Cowork that silently exfiltrates M365 files via Teams and Outlook — no user approval required. A second, direct data egress vulnerability was separately disclosed to Microsoft under embargo.
Uber's COO Says It Is Getting Harder to Justify AI Tokenmaxxing Spend
Uber COO Andrew Macdonald said in a Business Insider interview that it is 'becoming harder to justify' the company's AI token consumption costs, after conversations with senior engineering leaders showed no clear return. The statement comes weeks after Uber burned through its entire 2026 AI budget by April.
Huawei's LogicFolding Takes Aim at TSMC as Nvidia Concedes China's AI Chip Market
Huawei published a chip design paper introducing LogicFolding and tau scaling, a vertical stacking approach that cuts signal delay rather than shrinking transistors. At Computex the same week, Jensen Huang confirmed Nvidia has 'largely conceded' China's AI chip market to Huawei.
Claude Code Designed a Better AI Scaling Algorithm Than Human Researchers — $40, 160 Minutes, 70% Fewer Tokens
A UMD/Google/Meta team let Claude Code hunt for test-time scaling algorithms inside an offline replay environment. The agent invented the Confidence Momentum Controller, cutting token usage 70% versus self-consistency while matching or beating accuracy across AIME and HMMT benchmarks.
Brett Adcock's Hark Raises $700M at $6B for AI Hardware and Foundation Models — Before Shipping a Product
Hark, the AI lab founded by Brett Adcock (Figure AI, Archer Aviation), raised over $700 million at a $6 billion valuation with NVIDIA, AMD, Intel Capital, and Qualcomm Ventures on the cap table. First multimodal models arrive this summer; hardware designed from scratch for AI comes after.
Google Ends the Blue Links Era: Search Gets Its Biggest Overhaul in 25 Years at I/O 2026
At I/O 2026, Google replaced the iconic search box with an AI-powered interface built on Gemini 3.5 Flash. AI Mode tops 1 billion monthly users; AI Overviews reach 2.5 billion. Information agents that monitor the web around the clock arrive this summer for Pro and Ultra subscribers.