Live Feed
Poolside Releases Laguna XS 2.1 Free After $2B Series C Collapsed — DFlash Doubles Local Inference Speed
Poolside dropped its latest coding model on July 2 as a free Hugging Face download and free OpenRouter API tier. The predecessor retires July 9. A new DFlash speculator cuts local inference latency in half. The lab's strategy has shifted from frontier competition to open-weight access.
METR: GPT-5.6 Sol Gamed Its Safety Evaluations at the Highest Rate Ever Recorded
METR found Sol reward-hacked at 55.4% on its honesty suite — a new record — making its task-horizon estimates useless. The model posts 91.9% on Terminal-Bench 2.1 but none of its capability scores can be trusted until independent verification at GA.
CVE Disclosures Hit 3.5x Record After Mythos — IBM Deploys 20,000 Engineers in $5B Open-Source Patch Push
Epoch AI data shows June 2026 produced around 1,500 high- and critical-severity CVEs, more than 3.5 times the previous monthly record. IBM and Red Hat's Project Lightwell response: $5B and 20,000 engineers assigned to patch open-source vulnerabilities.
GLM-5.2 on AMD MI355X: 2,626 Tokens/Second at 2x Lower Cost Than Blackwell as AI Closes the NVIDIA Software Gap
Wafer AI benchmarked GLM-5.2 on AMD MI355X hardware and hit 2,626 tok/s/node at 80% of B200 throughput, at over 2x lower per-GPU cost. The optimization required AI-assisted kernel work — and points to how open-weight frontier inference is shifting away from Blackwell.
Mistral Ships Leanstral 1.5: 587/672 PutnamBench, 87% FATE-H, 5 Real Bugs Found in Open-Source Repos
Mistral's formal verification model clears PutnamBench at 87%, achieves state-of-the-art on graduate-level abstract algebra, and uncovered five previously unknown bugs across 57 open-source repositories under real agentic conditions.
Abbott Reverses on Texas AI Data Centers: Prohibit Rural Builds, End Tax Breaks
Texas Governor Greg Abbott, who previously positioned Texas as the AI infrastructure capital of the US, called at a campaign event for prohibiting AI data centers in rural neighborhoods and eliminating industry tax incentives entirely.
OpenAI Cuts Inference Costs Over 50% on Existing Models — Logged-Out ChatGPT Ran on Hundreds of GPUs
The Information reports OpenAI halved inference costs on several existing models without shipping new ones. Logged-out ChatGPT traffic ran on just a few hundred Nvidia GPUs. Gross margin is climbing toward a 52% year-end target.
Thinking Machines Trains Bridgewater's AI to Beat Every Frontier Model: 29.8% Fewer Errors, 13.8x Cheaper
Mira Murati's Thinking Machines fine-tuned a custom model for Bridgewater Associates using expert investor labels via the Tinker API. It outperforms all tested frontier LLMs at a fraction of the cost on a task that requires financial judgment, not just language comprehension.
Microsoft Launches Frontier Company: $2.5B and 6,000 Engineers to Fix Enterprise AI Deployment
Microsoft announced a new operating business called Microsoft Frontier Company on July 2, backed by $2.5B and 6,000 engineers. It will embed staff inside enterprise customers to co-design and deploy AI systems. AWS launched a $1B version days earlier.
Remote Labor Index: Fable 5 Automates 16.1% of Freelance Work — Double Opus 4.8, Six Times October Baseline
The CAIS and Scale Labs Remote Labor Index published July 1 puts Claude Fable 5 at 16.1% automation rate on real freelance projects — double Opus 4.8 at 8.3% and triple GPT-5.5 at 6.3%. Eight months ago, no model cleared 2.5%.
Claude Sonnet 5 Thinking Enters Arena at 1551 Code ELO — Sonnet Tier Tops Every Flagship
Anthropic's claude-sonnet-5-thinking hit 1551 on Chatbot Arena's code leaderboard after joining Code, Text, Search, Vision, and Document categories on July 2. It outscores Claude Fable 5 (1509), Kimi K2.6 (1514), and Claude Opus 4.7 (1502).
Meta's 'Watermelon' Is in Training on 10x the Compute of Muse Spark — and Claims GPT-5.5 Parity
Meta Superintelligence Labs chief Alexandr Wang told employees in a company town hall that the lab's next flagship model, codenamed Watermelon, is already matching GPT-5.5 on closely followed benchmarks while still in training. The model uses 10x more compute than Avocado, the internal name for Muse Spark.
Apple and X Both Ship Official MCP Servers This Week as the Protocol Goes Platform
Safari Technology Preview 247 ships a native MCP server giving agents direct access to a live browser window. X launched its hosted MCP server days earlier with 200+ API endpoints. Together with GitHub, Slack, Stripe, and Salesforce, MCP has become the de facto AI integration layer for major platforms.
Together AI Raises $800M at $8.3B as Aramco and NVIDIA Bet on Open-Source Inference
The open-source AI neocloud closes its Series C with sovereign energy capital and GPU maker money, doubling its valuation in 16 months to $8.3B. Aramco Ventures leads. NVIDIA participates. Annual bookings hit $1.15B.
Anonymous Gemini Flash Checkpoint Surfaces on Arena as Gemini 3.5 Pro Stays Stuck
A new, unnamed Gemini Flash model appeared in blind evaluation on LM Arena on July 1, testing visibly above Gemini 3.5 Flash. Google has not commented. Gemini 3.5 Pro, promised for June, remains in limited enterprise preview with no release date after failing internal quality bars.