Live Feed
Cambridge and NVIDIA's Red Queen Gödel Machine: AI Agents and Their Evaluators Co-Evolve Together
A new paper from Cambridge, NVIDIA, and collaborators shows that fixing the benchmark while improving the agent is self-defeating — the evaluator needs to improve too. On coding tasks the co-evolved system beats prior SOTA while using 1.35-1.72x fewer tokens. On paper writing it generates 1.86x higher acceptance from reviewer panels.
Samsung and SK Hynix Commit $518B to Korean Chip Hub — Shares Fall as Investors Price Capex Risk
South Korea's two chip giants announced $800 trillion won ($518B) in combined investment on Monday, anchoring a new semiconductor cluster. Samsung fell 4.7% and SK Hynix fell 3.1% the same day — the market pricing oversupply risk into the biggest capex commitment in the industry's history.
GPT-5.5 Instant's Third Update in 50 Days Ships No Benchmarks — OpenAI Pivots to Conversational Quality
On June 24, OpenAI pushed the third update to GPT-5.5 Instant since its May 5 launch, focusing on intent recognition and conversational quality. For the first time in this update cycle, the company published no quantitative benchmarks. The gpt-5.5-instant API endpoint is now a moving target, updating every 17 days without version bumps.
Karpathy's CLAUDE.md Grows to 10 Rules: A Self-Check Protocol Cuts Agent Error Rate From 41% to 11%
A ten-rule document attributed to Andrej Karpathy circulated on June 28, adding six self-monitoring rules to the four-rule CLAUDE.md template that has amassed 200,000 GitHub stars. The new rules teach autonomous coding loops to catch their own failure modes, not just write better code.
Semgrep: GLM-5.2 Scores 39% F1 on IDOR Detection, Beats Claude Code at $0.17 Per Vuln
Semgrep's vulnerability benchmark put GLM-5.2 at 39% F1 on IDOR detection, ahead of Claude Code's 32%. The open-weight model costs roughly one-sixth of comparable frontier models and runs entirely on-premise.
China's LineShine Displaces El Capitan as World's Fastest Supercomputer — CPU-Only, No NVIDIA
At ISC 2026 in Hamburg, China's LineShine posted 2.198 Exaflops to take the TOP500 #1 spot. No GPUs. Domestic chips only. The US hasn't topped the list in two years.
Microsoft Fairwater Goes Live in Wisconsin as One 800G AI Supercomputer
Microsoft's Fairwater campus in Mount Pleasant is operational, linking hundreds of thousands of NVIDIA GB200 GPUs with an 800G Ethernet fabric. The former Foxconn site has become a test case for whether AI campuses are data centers or single machines.
GPT Image 2 Lands on OpenRouter at $8/$8 Per Million Tokens
OpenRouter lists OpenAI's GPT Image 2 with $8 per million input tokens, $8 per million output tokens, and a 400K context window. The release puts image generation into the same token-metered procurement frame as text models.
Google Caps Meta's Gemini Access as AI Demand Outruns Hyperscaler Supply
Google has limited Meta's use of Gemini after the social network sought more model capacity than Google could provide. The story is not just intercompany rivalry: even the companies spending fastest on AI are now capacity constrained.
Artificial Analysis Legal Index Makes Agentic Execution 25% of Legal AI Score
Artificial Analysis has added a legal capability index that gives agentic execution the second-largest weight after legal knowledge. The benchmark is a quiet correction to legal AI marketing: knowing doctrine is only 30% of the score.
Ford Rehired 350 Engineers After AI Could Not Fix Vehicle Quality
Ford's quality turnaround needed 350 experienced engineers, technicians, and specialists after the company found AI and changed design requirements were not enough. The episode is a useful warning for manufacturers treating AI as a substitute for domain memory.
AI RFIC Design Moves From Templates to Search: 30-100 GHz Amplifier Shows the Payoff
Princeton researchers are using reinforcement learning, inverse design, and diffusion models to generate radio-frequency chip layouts from scratch. The near-term story is not LLMs designing chips. It is search collapsing an RFIC workflow that can take years and tens to hundreds of millions of dollars.
Open Weights Are 5 Months Behind Frontier on Average — Coding Is the Only Category Closing Fast
Analysis of 18 Artificial Analysis benchmarks shows the open-source lag has held steady at roughly 5 months since 2024. The convergence narrative is real, but almost entirely driven by coding — every other capability dimension remains flat or slightly wider.
US Unblocks Mythos 5 for 100+ Trusted Partners — Fable 5 Stays Dark
Commerce Secretary Lutnick's letter to Anthropic clears Mythos 5 for over 100 US institutions including Fortune 500 companies and government agencies. Fable 5, the consumer-tier model from the same family, is not mentioned and remains offline.
Data Centers Cost Utah's Senate President His Seat: Stratos Project's 9 GW Demand Becomes an Election Issue
Utah State Senate President J. Stuart Adams lost his Republican primary after backing the Stratos data center campus near the Great Salt Lake. The project would consume up to 9 GW of power, more than Utah's entire current usage. A Reuters/Ipsos poll finds 57% of Americans oppose a data center in their community.