Live Feed
Setting Image Inputs to 'detail: low' Costs More on Reasoning Models, OpenRouter Finds
An OpenRouter benchmark across 5 models shows that low-detail image inputs on GPT-5.5 produce 13.8 points less accuracy and cost 13% more than auto-detail. Reasoning models compensate for downsampled images by thinking 1.6x harder, burning tokens that outweigh any image savings.
Apple Commits $30B to Broadcom for 15 Billion US-Made Chips
Apple's expanded Broadcom partnership is its largest American manufacturing commitment to date. More than 15 billion chips will be produced on US soil, extending a two-decade supply relationship that now runs straight through AI infrastructure.
FuriosaAI Deploys Its 5nm Inference Chip at Equinix Lisbon as Europe Bets on Alternative AI Silicon
The South Korean AI chip startup is installing its RNGD accelerator at Equinix's Lisbon LS2 data center — its first major European deployment. The 5nm Tensor Contraction Processor runs at 512 TFLOPS FP8 within a strict 180W envelope, targeting Nvidia's inference dominance.
The Central Bankers Are Scared: BIS Warns AI's $100B Debt Stack Could Seed the Next Financial Crisis
The Bank for International Settlements issued one of its sharpest warnings yet about AI infrastructure debt. Hyperscaler bond issuance topped $100B in 2025. Private credit funds have quadrupled their AI exposure. Circular financing between chipmakers, labs, and compute providers makes real demand impossible to verify.
Beijing Is Moving to Lock Advanced Chinese AI Models Inside Its Borders
China is in talks with Alibaba, ByteDance, and Z.ai about restricting foreign access to advanced AI models, Reuters reports. The controls would cover both closed APIs and open-weight downloads — and arrive just as Chinese models have become competitive with US frontier labs.
Brookfield and Bloom Energy Expand AI Power Partnership to $25B — Fuel Cells Replace Grid Waiting
Brookfield and Bloom Energy have scaled their AI infrastructure collaboration to $25 billion, betting that on-site solid oxide fuel cells can deliver power faster and with less community friction than new grid connections. The model integrates power, compute, and capital from day one.
Artificial Analysis Launches Harvey LAB-AA: 120 Private Legal Tasks Test AI Agents on Real Deliverables
Artificial Analysis and Harvey have published the Harvey LAB-AA benchmark, running AI agents against 120 private legal tasks across 24 practice areas. The scoring system reports both an all-pass rate and a criterion pass rate — a methodology gap the legal AI market has been missing.
AI Audits Cloudflare's Go Crypto Library and Finds Something — ZKSecurity Opens a Series
ZKSecurity published AI-assisted audit findings from Cloudflare's Circl, a production Go cryptographic library. Cryptographic code has historically resisted automated analysis. A numbered series suggests the workflow is repeatable.
SK Telecom Targets 15GW to Become Asia's AI Infrastructure Hub — Korea Names It a National Revolution
SK Telecom announced a 15GW AI data center buildout anchored at Ulsan, with 5GW coming online from 2029. Seoul's AI G3 strategy — a bid to rank alongside the US and China in global AI — has found its infrastructure arm.
DeepSeek Is Building Its Own Inference Chip to Break From Nvidia
The Chinese lab behind DeepSeek V4 is developing a proprietary inference chip, sources say. Nvidia shares fell pre-market on the news. The move mirrors the silicon independence plays of every major US hyperscaler — and arrives as US export controls cut off DeepSeek from the latest Nvidia hardware.
OpenAI Ships GPT-Realtime-2.1: Reasoning Comes to Voice Agents, p95 Latency Drops 25%
OpenAI released two new Realtime API models on July 6 — gpt-realtime-2.1 and gpt-realtime-2.1-mini. The mini model is the first reasoning model available for real-time voice, at the same price as its predecessor. Both cut p95 voice latency 25% through improved caching.
Fable 5 Goes Credit-Only Today: $10/$50 Per Million Ends Subscription-Included Frontier Access
Starting July 7, every Claude subscriber pays metered rates to use Fable 5 — $10 per million input tokens, $50 per million output. That is double Opus 4.8's rate and the most expensive price Anthropic has ever published for a generally available model.
Chinese AI Models Now Hold 30% of US Enterprise Token Use — and Peaked at 46%
CNBC data shows Chinese models' share of US company tokens on OpenRouter has stayed above 30% every week since February 8, up from 4.5% in the first half of 2025. GLM-5.2 grew 27x on Vercel in its first week. The driver is cost: Chinese open models run 60-90% cheaper than US frontier alternatives.
Tencent Ships Full Hy3 Under Apache 2.0: Search Agent Champion That Cedes Code to GLM-5.2
Tencent's Hy3 goes production with a commercial Apache 2.0 license, a 20x token consumption jump since preview, and benchmark leadership on search and tool orchestration. It trails GLM-5.2 by 6.2 points on SWE-bench Verified but wins the infrastructure argument: 295B total, 21B active, half the memory of its nearest rival.
ATOM Report: China Passed the US in Open-Weight AI Downloads -- 1.15B to 723M, Qwen in the Lead
A new paper measuring the global open-model ecosystem finds China crossed the US in total downloads in summer 2025 and now leads 1.15B to 723M. Qwen accounts for most of the gap; DeepSeek controls the large-model segment above 250B parameters. After size and age adjustments, some US models still show strong momentum.