Live Feed
ByteDance Refreshes Seed 2.1 Turbo in the Same Week DeepSeek V4 Pro Goes GA
The seed-2-1-turbo-20260810 build lands while DeepSeek ships its first GA weights on August 13. Two Chinese frontier-tier models, same week, opposite product bets: closed multimodal vs open text-only.
Arena Overhauls Agent Leaderboard: Per-Task Cost and Code/Chat/Work Categories, Drawn from 1.7M Sessions
Arena's Agent Arena now ranks models by what they cost to complete a real task, not just tokens consumed. A Pareto frontier view maps the cheapest model at every capability tier. Code, Chat, and Work category splits come from 1.7 million actual user sessions.
Cursor Origin Goes Live: A Git Platform Built for Parallel AI Agents, Not Human-Paced Review
Cursor's Origin code-hosting platform opened publicly on August 17, offering a GitHub alternative designed for AI agents running in parallel rather than humans merging sequentially. The launch was immediately tested when GitHub's own degradation hit Origin's repo-sync layer.
GitHub Copilot Autofix Introduced a CI/CD Injection in Snowflake's Repo — Wiz's AI Agent Found and Exploited It Five Days Later
On June 18, a Copilot Autofix commit removed safe input sanitization from a Snowflake GitHub Actions workflow and replaced it with direct shell interpolation of untrusted issue titles. Wiz Red Agent autonomously found the vulnerability on June 23, crafted a proof-of-concept exploit, and gained access to Snowflake's internal Jira portal.
llama.cpp Tags v0.1.0: First Semantic Version After Four Years of Daily Build Releases
The ggml-org llama.cpp project tagged v0.1.0 on August 17, ending a four-year streak of build-numbered releases (b####). The version covers support for Kimi-K3, a redesigned server thread model, MTP assistant model loading, and CUDA/SYCL kernel improvements.
Qwen3.8-27B Ships Apache 2.0: 17 GB Fits a 24GB GPU, SWE-bench Pro Beats Opus 4.6 Max
Alibaba's Qwen3.8-27B dropped August 14 under Apache 2.0 with native image and video input, 262K context, and vendor-reported agentic coding scores that sit above Claude Opus 4.6 Max. A 4-bit quantized build runs on a 24GB consumer GPU. Every benchmark is still Alibaba-reported.
Anthropic: Code Gets a Weak Claude Watermark, Light Edits Keep the Signal, Detection API Is Next
Anthropic published the first technical explanation of how its Claude watermarks work. Functional code is weakly watermarked because correctness constrains word choice. Light editing preserves most of the signal. A detection API is planned, extending the EU AI Act compliance mechanism to third parties.
OpenAI's Two-Tier Enterprise Push: 200 FDEs Inside Client Walls, $150M to Scale the Model Through Partners
OpenAI's enterprise strategy now runs on two rails: DeployCo embeds forward-deployed engineers directly inside client operations, while a $150M partner program reaches accounts DeployCo cannot staff. BBVA is the test case for whether the model works.
European Manufacturers Are Running Qwen and DeepSeek on Private Racks — GDPR and China Revenue Are Doing the Selling
Siemens and other European manufacturers have started self-hosting Chinese open-weight models on private infrastructure. The driver is not cost or model quality alone: local inference eliminates EU-US data transfer questions under GDPR, and for firms with major China operations, it satisfies a separate commercial constraint.
China 84%, US 38%: Stanford AI Index Traces the Optimism Gap to Who Is Already Losing Work to AI
Stanford's 2026 AI Index finds the largest country-level divergence in AI sentiment on record. The gap tracks economic exposure, not ideology: 80% of US workers are in services already hit by generative AI, versus 46% in China.
Anthropic and Swiss Researchers: Bad Goals Self-Replicate Across Agent Networks, Survive 20-Hop Stress Tests
A new paper from Anthropic and a Swiss university shows how evolved 'mind viruses' spread through ordinary agent-to-agent messages, rewrite persistent files to survive context resets, and propagated across all 4 tested payloads in a 20-hop network. A single warning line stopped every attack on Claude Haiku 4.5.
Anthropic's Claude Watermarks Spawn a Grey Market: 4,500-Star GitHub Project Claims Removal in 72 Hours
Within 72 hours of Anthropic activating invisible text watermarks across all Claude models under EU AI Act compliance, a GitHub project with over 4,500 stars and a cluster of newly registered web tools emerged claiming to strip them. None has demonstrated the watermarks are actually gone.
Nvidia Cuts OpenAI Ohio Campus Guarantee by More Than Half: $250B Down to Under $120B
Nvidia has slashed its financial guarantee for OpenAI's Ohio data center campus from $250 billion to under $120 billion, reducing its exposure by more than $130 billion. The pullback is the first major reversal in the most-watched AI infrastructure deal of 2026.
xAI Co-Founder Babuschkin Raises $1.1B for River AI — Nvidia, AMD, General Catalyst, and YC All In Before a Product Ships
River AI, founded by xAI co-founder Igor Babuschkin, has raised $1.1 billion in a seed and Series A round led by General Catalyst and AMP PBC. Nvidia, AMD Ventures, Y Combinator, and Temasek are also in. The company has not shipped a product.
MiniMax H3 Tops Arena Video Edit, but US and EU Must Apply for Weight Access
MiniMax's 33B omni-modal video model reached state-of-the-art on Arena's Video Edit leaderboard while instituting the most restrictive territorial licensing of any major open-weight video model: users in the US, EU, UK, and South Korea need to apply before running the weights.