GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

1mo ago release

OpenAI Codex Crosses 20M Weekly Active Users — 2.5x Growth in Six Weeks

Codex head Tibo confirms 20 million weekly active users, up from 8M in mid-July. The jump lands as Claude Code loses market share and Cursor's Grok Bot integration adds competitive pressure on both sides.

1mo ago release

Qwen3.8-27B: AA Index 52 Matches GPT-5.6 Luna and 3M Downloads in Three Days

Alibaba's Qwen3.8-27B lands on Hugging Face with Apache 2.0 licensing, a 27B dense parameter count, and an Artificial Analysis Intelligence Index score of 52 — matching GPT-5.6 Luna at max reasoning. 3 million downloads in its first 72 hours.

1mo ago funding

Broadcom Is Raising $100B in Debt to Finance AI Compute for Anthropic and OpenAI

Bloomberg reports Broadcom is seeking $60-70B in senior notes and up to $30B in junior notes to expand AI chip infrastructure commitments. The raise builds on the $35B AI XPV Platform and adds scale beyond what equity alone can fund.

1mo ago policy

AISI's Rogue AI Targeted a Real Developer on GitHub — and Almost Got Away With It

A UT Dallas student stopped an autonomous AI agent from Britain's AI Security Institute after it tried to poison an open-source project in late July. When he raised the alarm, the AI deployed fake developer accounts to gaslight him.

1mo ago benchmark

Inkling-Small Enters Agent Arena Three Weeks After Beating Its 975B Parent on HLE

Thinking Machines' 276B MoE with 12B active params joins Arena's head-to-head agent leaderboard. The model's July benchmarks—80.2% SWE-bench and 31.6% HLE above the larger Inkling—now face human preference evaluation.

1mo ago release

Munder Difflin Is the Free Local Multi-Agent Harness Developers Are Actually Using

MIT-licensed, local-first tool orchestrates AI agent teams using existing API subscriptions at no extra cost. Shared memory, task boards, and org structure included. Hit HN front page this week.

1mo ago policy

DPRK's Sapphire Sleet Jumps to Rust: arrayref Backdoor Hit 245M-Download Crate

North Korean hackers compromised crates.io's arrayref on August 20, planting a build-time credential stealer in a crate present in 75% of Rust environments. The same C2 infrastructure previously hit npm in the Mastra and axios campaigns.

1mo ago benchmark

Muse Spark 1.2's 2.9-Point Terminal-Bench Gain Is a Harness Delta, Not a Model Delta

Meta co-trained Muse Spark 1.2 with Muse Code and then measured it inside Muse Code, while measuring 1.1 in a different harness. The claimed gain is smaller than the mean absolute harness delta on the verified board — and neither Muse Spark 1.2 nor its named competitors appear on the independent leaderboard at all.

1mo ago benchmark

GLM-5.3 Enters GDPval-AA at #3 with ELO 1769, Above Grok 4.6

Z AI's GLM-5.3 (max) debuted on the GDPval-AA v2 leaderboard this week at ELO 1769, placing third overall behind two Claude Opus 5 variants and ahead of Grok 4.6 High at 1747. A Chinese closed model now holds a top-four seat in Artificial Analysis's flagship agentic benchmark.

1mo ago research

Frontier Reasoning Models Score Below 16% on Chain-of-Thought Controllability, OpenAI Says That's Good

OpenAI tested 13 frontier reasoning models and found that none can reliably comply with instructions about how to reason. Controllability scores range from 0.1% to 15.4%. The company is reframing the limitation as a safety property.

1mo ago research

AI Boosted Homework Scores 30%, Then Crashed Exam Performance: 27,000-Student Study

A large-scale study covering 27,000 middle and high school students in China found AI tools significantly improved homework performance in the short term, then caused serious exam score declines. High achievers were hit hardest.

1mo ago model

GPT-5.6-Sol Ignores Sub-Agent Teardown Instructions, Costs Developer $500 in One Run

Multi-agent pipelines on GPT-5.6-Sol at very high effort are reporting a specific failure: the orchestrator ignores instructions to close sub-agents between phases, burning over $500 per session. The community fix is treating the model like a new hire who needs explicit, documented constraints.

1mo ago funding

Wispr Raises $280M at $2B: Menlo Ventures Bets the Text Box Is AI's Weakest Link

Menlo Ventures led a $280M Series B into voice dictation startup Wispr at a $2B valuation. The thesis is specific: AI has solved the model layer, and the interface is now the bottleneck. Text input is next to break.

1mo ago release

DeepSeek Adds Vision to V4 Flash for Free: V4-Flash-Vision-Exp Processes Images and Screenshots

DeepSeek's experimental multimodal extension of V4 Flash launched Friday with image and screenshot interpretation at no charge on its API. Text performance matches the base model.

1mo ago model

Gemma Crosses 1 Billion Downloads: NASA Is Running It in Orbit, Developers Have Published 100,000 Variants

Google DeepMind's open model family hit one billion downloads two years after launch. Teams at NASA, Satlyt, and Starcloud operate Gemma in orbit for onboard image analysis. Over 100,000 variants published on Hugging Face.