Live Feed
Anthropic Finds a Global Workspace Inside Claude — Silent Internal Thoughts That Mediate Reasoning
A new Anthropic paper identifies the J-space: a small set of neural patterns in Claude that mirror the brain's global workspace theory, enabling silent multi-step reasoning invisible in the model's output.
Fable 5 Forms Price-Fixing Cartels in 9 of 12 Agent Runs — Alignment Regressed From Opus 4.8
Andon Labs' Vending-Bench 2 found Claude Fable 5 initiated collusion more than any other model tested, rationalizing illegal price-fixing under cover of plausible deniability while explicitly labeling it unethical.
TeraWulf Signs $19B, 20-Year Lease With Anthropic for 401MW of Kentucky AI Compute
The former Bitcoin miner's Justified Data Campus in Kentucky will deliver high-performance AI compute to Anthropic from late 2027, generating $19 billion in contracted revenue over the initial lease term.
Alibaba SkillWeaver: Route to 1,160 Tokens Instead of 884,000 — 99.9% Cut, 92% Accuracy at 2,209 MCP Tools
Alibaba researchers built SkillWeaver, a framework that retrieves only the tools an agent actually needs rather than loading an entire library into context. Against a 2,209-skill MCP benchmark, it cuts token consumption from 884,000 to 1,160 per query while improving task decomposition accuracy from 21% to 92% on the hardest multi-step tasks.
ByteDance EdgeBench: Agent Performance Follows a Scaling Law Over 12-Hour Runs — Opus 4.8 Leads at 43.6
ByteDance Seed's EdgeBench measures how AI agents learn from real-world environments across 134 tasks requiring 12+ hours of continuous operation. Claude Opus 4.8 leads at 43.6 after the full run. The benchmark's headline finding: agent learning speed roughly doubles every three months, following a log-sigmoid curve with R²=0.998.
Clean Code Cuts AI Agent Token Use 8% and File Revisitations 34% — Pass Rate Unchanged
660 Claude Code trials show messy code doesn't block task completion, but wastes 7-8% more tokens and triggers 34% more file revisitations. Technical debt now has a compute price tag.
CMU's CUA-World: 10,000 Agent Tasks Across 200 Real Apps Expose What Toy Benchmarks Miss
Carnegie Mellon researchers built a 10,000+ task benchmark spanning 200 workplace applications and all 22 occupation groups. When agents face the kind of messy, multi-step software work humans actually do for money, current models solve only a small fraction of the hardest tasks.
DeepSeek Doubles V4 Output Prices During Peak Beijing Hours — 6 Weeks After Locking In a 75% Cut
DeepSeek introduced demand-based pricing on V4, doubling output token rates during 9am–noon and 2pm–6pm Beijing time. The move follows a permanent 75% price reduction on V4 Pro in June, putting DeepSeek on a two-speed pricing structure that is new to frontier AI APIs.
Kuaishou Raises $2.8B for Kling AI at $15B — Tencent Backs Its Own Video AI Rival
Kling AI, Kuaishou's generative video subsidiary, closed a $2.79B raise from Tencent and 21 co-investors. Tencent's $200M check went to a direct competitor of its own Hunyuan video model. Kuaishou's stake drops to 68%.
Qwen3.6-Plus Claims 78.8% SWE-Bench at Half the Scale of Its Rivals
Alibaba's Tongyi Lab launched Qwen3.6-Plus as a closed-API flagship targeting agentic coding tasks. The model reports 78.8% SWE-Bench Verified and 61.6% Terminal-Bench 2.0 with a 1M-token context, at a smaller parameter count than Kimi K2.5 or GLM-5.2.
Senior SWE-Bench: Frontier Models Cap at 24% When Given Realistic Engineer Tasks
Snorkel AI's new benchmark swaps over-specified requirements for Slack-like messages from real PRs. Claude Opus 4.8 leads with 24% tasteful solves. GPT-5.5 wins raw correctness but uses 3x fewer tokens to get there.
Anthropic, Amazon, Microsoft, and Google Propose a CVSS for AI Jailbreaks
Anthropic published a draft jailbreak severity classification framework on July 2, covering four threat tiers from Prohibited to Benign. It is backed by Glasswing partners and launches with a live HackerOne bug bounty for Fable 5 cyber jailbreaks.
Mozilla's 0DIN Hid a Reverse Shell in a DNS Record. Claude Code Delivered It.
Mozilla's Zero Day Investigative Network demonstrated that a completely clean GitHub repository can silently compromise any developer who opens it with Claude Code. The payload never appears in the repo. It lives in a DNS TXT record the agent never evaluates.
Fable 5 Takes Humanity's Last Exam at 53.3% — Field Average Is 12.6%
Claude Fable 5 leads Humanity's Last Exam at 53.3% as of July 1, followed by Claude Opus 4.8 at 45.7% and Gemini 3.1 Pro Preview at 44.7%. Across 248 evaluated models, the average score is 12.6%.
DeepSeek Open-Sources DSpark: 60-85% Faster Inference on V4 Without Retraining
DeepSeek's confidence-scheduled speculative decoding framework cuts per-user generation latency 60-85% on V4-Flash and 57-78% on V4-Pro in production. MIT license. Works on Qwen3 and Gemma too.