GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

1mo ago release

ByteDance Refreshes Seed 2.1 Turbo in the Same Week DeepSeek V4 Pro Goes GA

The seed-2-1-turbo-20260810 build lands while DeepSeek ships its first GA weights on August 13. Two Chinese frontier-tier models, same week, opposite product bets: closed multimodal vs open text-only.

1mo ago benchmark

Arena Overhauls Agent Leaderboard: Per-Task Cost and Code/Chat/Work Categories, Drawn from 1.7M Sessions

Arena's Agent Arena now ranks models by what they cost to complete a real task, not just tokens consumed. A Pareto frontier view maps the cheapest model at every capability tier. Code, Chat, and Work category splits come from 1.7 million actual user sessions.

1mo ago release

Cursor Origin Goes Live: A Git Platform Built for Parallel AI Agents, Not Human-Paced Review

Cursor's Origin code-hosting platform opened publicly on August 17, offering a GitHub alternative designed for AI agents running in parallel rather than humans merging sequentially. The launch was immediately tested when GitHub's own degradation hit Origin's repo-sync layer.

1mo ago research

GitHub Copilot Autofix Introduced a CI/CD Injection in Snowflake's Repo — Wiz's AI Agent Found and Exploited It Five Days Later

On June 18, a Copilot Autofix commit removed safe input sanitization from a Snowflake GitHub Actions workflow and replaced it with direct shell interpolation of untrusted issue titles. Wiz Red Agent autonomously found the vulnerability on June 23, crafted a proof-of-concept exploit, and gained access to Snowflake's internal Jira portal.

1mo ago release

llama.cpp Tags v0.1.0: First Semantic Version After Four Years of Daily Build Releases

The ggml-org llama.cpp project tagged v0.1.0 on August 17, ending a four-year streak of build-numbered releases (b####). The version covers support for Kimi-K3, a redesigned server thread model, MTP assistant model loading, and CUDA/SYCL kernel improvements.

1mo ago release

Qwen3.8-27B Ships Apache 2.0: 17 GB Fits a 24GB GPU, SWE-bench Pro Beats Opus 4.6 Max

Alibaba's Qwen3.8-27B dropped August 14 under Apache 2.0 with native image and video input, 262K context, and vendor-reported agentic coding scores that sit above Claude Opus 4.6 Max. A 4-bit quantized build runs on a 24GB consumer GPU. Every benchmark is still Alibaba-reported.

1mo ago policy

Anthropic: Code Gets a Weak Claude Watermark, Light Edits Keep the Signal, Detection API Is Next

Anthropic published the first technical explanation of how its Claude watermarks work. Functional code is weakly watermarked because correctness constrains word choice. Light editing preserves most of the signal. A detection API is planned, extending the EU AI Act compliance mechanism to third parties.

1mo ago funding

OpenAI's Two-Tier Enterprise Push: 200 FDEs Inside Client Walls, $150M to Scale the Model Through Partners

OpenAI's enterprise strategy now runs on two rails: DeployCo embeds forward-deployed engineers directly inside client operations, while a $150M partner program reaches accounts DeployCo cannot staff. BBVA is the test case for whether the model works.

1mo ago model

European Manufacturers Are Running Qwen and DeepSeek on Private Racks — GDPR and China Revenue Are Doing the Selling

Siemens and other European manufacturers have started self-hosting Chinese open-weight models on private infrastructure. The driver is not cost or model quality alone: local inference eliminates EU-US data transfer questions under GDPR, and for firms with major China operations, it satisfies a separate commercial constraint.

1mo ago research

China 84%, US 38%: Stanford AI Index Traces the Optimism Gap to Who Is Already Losing Work to AI

Stanford's 2026 AI Index finds the largest country-level divergence in AI sentiment on record. The gap tracks economic exposure, not ideology: 80% of US workers are in services already hit by generative AI, versus 46% in China.

1mo ago research

Anthropic and Swiss Researchers: Bad Goals Self-Replicate Across Agent Networks, Survive 20-Hop Stress Tests

A new paper from Anthropic and a Swiss university shows how evolved 'mind viruses' spread through ordinary agent-to-agent messages, rewrite persistent files to survive context resets, and propagated across all 4 tested payloads in a 20-hop network. A single warning line stopped every attack on Claude Haiku 4.5.

1mo ago policy

Anthropic's Claude Watermarks Spawn a Grey Market: 4,500-Star GitHub Project Claims Removal in 72 Hours

Within 72 hours of Anthropic activating invisible text watermarks across all Claude models under EU AI Act compliance, a GitHub project with over 4,500 stars and a cluster of newly registered web tools emerged claiming to strip them. None has demonstrated the watermarks are actually gone.

1mo ago funding

Nvidia Cuts OpenAI Ohio Campus Guarantee by More Than Half: $250B Down to Under $120B

Nvidia has slashed its financial guarantee for OpenAI's Ohio data center campus from $250 billion to under $120 billion, reducing its exposure by more than $130 billion. The pullback is the first major reversal in the most-watched AI infrastructure deal of 2026.

1mo ago funding

xAI Co-Founder Babuschkin Raises $1.1B for River AI — Nvidia, AMD, General Catalyst, and YC All In Before a Product Ships

River AI, founded by xAI co-founder Igor Babuschkin, has raised $1.1 billion in a seed and Series A round led by General Catalyst and AMP PBC. Nvidia, AMD Ventures, Y Combinator, and Temasek are also in. The company has not shipped a product.

1mo ago release

MiniMax H3 Tops Arena Video Edit, but US and EU Must Apply for Weight Access

MiniMax's 33B omni-modal video model reached state-of-the-art on Arena's Video Edit leaderboard while instituting the most restrictive territorial licensing of any major open-weight video model: users in the US, EU, UK, and South Korea need to apply before running the weights.