Live Feed
Qwen3.8-2.4T and 27B Both Ship Apache 2.0: First Time Alibaba Releases Max-Class Model as Open Weights
Alibaba dropped weights for both Qwen3.8-27B and Qwen3.8-2.4T-A95B on Hugging Face on August 12-14 under Apache 2.0, marking the first time it has released a Max-class MoE flagship alongside a smaller dense variant simultaneously.
Artificial Analysis Launches Optima: Build Your Own LLM Benchmark, Compare Models by Cost Per Task
Optima lets teams benchmark AI models against their own data, agent traces, or use-case descriptions — then compare results by quality, cost per task, and time per task. The platform uses the same grading methods behind GDPval-AA and AA-Briefcase.
China's Instagram Ships Its First AI Model: dots3-note Scores 75.1 on Terminal-Bench 2.1 Under Apache 2.0
Xiaohongshu's AI lab open-sources dots3-note Preview, a 280B MoE with 16B active parameters, multimodal input, and a novel long-horizon RL training method — the first release in a planned three-model family.
Qwen Hits 3 Billion Downloads in Six Months — More Than Meta and Google Combined
Alibaba's Qwen family has accumulated 3 billion global downloads since February 2026, across 460-plus open models and 300,000-plus community derivatives, eclipsing Meta at 227 million and Google at 418 million according to a Hugging Face state of open models report.
35 Trajectories Beat Reasoning Mode: Microsoft Paper Replaces Test-Time Compute With Distilled Skills
A Microsoft Research paper shows that collecting 35-50 agent trajectories and distilling failure patterns into a markdown skill file recovers 55-100% of the reasoning premium at 2.9-4.5x lower token cost. On two benchmarks, the skilled non-reasoning model outperformed reasoning mode entirely.
Apple Completes China LLM Training With Alibaba — First Foreign AI Model Cleared by Beijing
Apple has finished training a China-specific large language model built with Alibaba. China's AI regulator registered Apple Intelligence in July. If deployed as expected, Apple would be the first foreign company with a proprietary AI model approved in China.
45-Agent Swarm Found 12x More Vulnerabilities Than Parallel Agents — Anthropic Names 3 Systemic Failure Modes in Multi-Agent Systems
Anthropic's new research on multi-agent systems documents coordination gaps, conformity failures, and epistemic breakdowns. A vulnerability-finding swarm of 45 agents found 266 bugs vs 21 for independent agents — but the two methods shared only 12 findings.
PerceptionBench: No Frontier Model Breaks 60% on Visual Perception — GPT-5.6 Sol Leads at 59.7%
Moonshot AI's new benchmark isolates visual perception from reasoning across 10 atomic sub-skills. The best frontier model scores 59.7%. The weakest category is object hallucination, where GPT-5.6 Sol manages 26.9% and a weaker Gemini 3.5 Flash beats it at 50.6%.
Google's $15B India AI Hub Takes Shape in Andhra Pradesh — Largest Overseas AI Build on Earth
Google's $15B data centre hub in Andhra Pradesh, breaking ground in April 2026, is the company's largest AI investment anywhere outside the United States. Three sites, 601 acres, TPU and GPU compute, and subsea cable infrastructure.
Riot Platforms Exits Bitcoin Mining With $9.8B in AI Contracts — Anthropic Deal Closes the Pivot
Anthropic signed a $9.1B, 20-year lease with Riot Platforms for 191 megawatts at Rockdale, Texas. Combined with a January AMD deal, Riot now holds $9.8B in contracted AI revenue and has effectively converted from Bitcoin miner to AI data center landlord.
Anthropic's August 2026 Risk Report Names Two Unreleased Frontier Models — All Claude Lines Provisionally Hit CB-1
Anthropic's biannual safety assessment discloses two internal frontier-class models that were not publicly released as of July 15, including one broadly comparable to Mythos 5. Every Claude line has now provisionally crossed the CB-1 catastrophic-harm threshold since Opus 4.
Debian Puts AI Policy to an 8-Way Vote — From Social Contract Ban to Climate Refusal
The Debian project opened voting today on 8 competing AI/LLM policy proposals, ranging from a total Social Contract ban to a climate-based rejection. Voting runs August 15 to 28. This is the project's first binding decision on AI contributions after two earlier attempts failed.
GLM-5.3 Tops CyberGym at 84.5%, Finds a Cursor Vulnerability, Ships Open Weights in Two Weeks
Z.ai's GLM-5.3 posts 66.9% on DeepSWE (+20 points over GLM-5.2), leads CyberGym at 84.5% ahead of GPT-5.6 Sol and Mythos 5, and reportedly found a serious vulnerability in Cursor before public release. Open weights follow once safety hardening is done.
Arena Adds Gemini 3.7 Flash High and Grok 4.6 High Across Three Leaderboards in 48 Hours
Six models entered Arena across multiple leaderboards between August 11 and 13. Gemini 3.7 Flash High now has head-to-head preference data in Agent, Text, and Code arenas. Grok 4.6 High — carrying 95.6% SWE-bench — is live in Code and Text.
SK Hynix Locks Korea's HBM Capacity Through 2033 — and the US Has No Domestic Wafer Fab
SK Hynix's board vote commits its AI memory roadmap to Korea through 2033. The US has zero HBM wafer fabrication capacity, with its earliest realistic domestic timeline arriving in 2028 — a packaging facility only, not wafer production. Every NVIDIA GPU and Google TPU runs on Korean-made HBM.