Live Feed
Qwen3.7-Max Leads SWE-Pro at 60.6% — Alibaba's Coding Agent Runs 35 Hours, 1,000+ Tool Calls
Alibaba launches Qwen3.7-Max, a proprietary agent model that tops SWE-bench Pro at 60.6%, above Kimi K2.6 (59.5%), DeepSeek V4 Pro (59.0%), and Claude Opus 4.6 (57.3%). Designed for long-horizon autonomous execution, it completed a 35-hour kernel optimisation run with over 1,000 tool calls.
Qwen3.7-Max-Preview Enters Arena at Rank 13 — Alibaba Reaches 6th Lab Globally
Alibaba's Qwen3.7-Max-Preview debuted on the Arena text leaderboard at rank 13, slotting between GPT-5.5 and Grok 4.2. It is the highest-ranked Chinese model on the platform and its sub-rankings in math and expert tasks place it inside the top 10 globally. Alibaba has now shipped three major model generations since February 2026.
xAI Ships Grok Skills: Teach It Once, It Remembers Every Conversation
xAI launched Skills on May 18 — a persistent memory and automation layer for Grok 4.3 on web, iOS, and Android. Built-in skills generate Word, PowerPoint, Excel, and PDF files natively. Users can build custom skills in seconds via conversation.
Mistral Acquires Austrian Physics AI Startup Emmi — 30 Researchers, Large Engineering Models, and a Second Deal in Three Months
Mistral AI has acquired Emmi AI, a Vienna-based startup that builds physics AI models for industrial simulation. The deal is Mistral's second acquisition in 2026 after Koyeb in February, and adds Large Engineering Models for aerospace, automotive, and semiconductor engineering to the French lab's stack.
Google Antigravity 2.0 Launches at I/O 2026 — Gemini CLI Dies June 18, Agent IDE War Has a New Entrant
Google deprecated Gemini CLI at I/O 2026 and replaced it with Antigravity 2.0: a standalone agent-first development platform with a VS Code fork, a multi-agent Manager surface, CLI, SDK, and Managed Agents powered by Gemini 3.5 Flash. Claude Code and Codex now have direct competition from the search giant.
Gemini 3.5 Flash Lands at Intelligence Index 55, #7 Globally — Google's Agent Play Has 1M Context at $1.50/M
Released May 19, Gemini 3.5 Flash hits 55 on the Artificial Analysis Intelligence Index, ranking #7 of 147 models. At $1.50/M input and $9/M output with a 1M-token context window, Google is pitching it directly at agent workloads — not as a chatbot companion.
Gemini 3.5 Flash Leads Terminal-Bench 2.1 at 76.2% — 4x Faster Than Rivals at Under Half the Cost
Google launches Gemini 3.5 Flash at I/O 2026 with 76.2% on Terminal-Bench 2.1 and 1656 Elo on GDPval-AA, both ahead of Gemini 3.1 Pro on agentic benchmarks. At $0.75/M input and $4.50/M output, it delivers frontier-grade agent performance at Flash pricing. Gemini 3.5 Pro is due next month.
Andrej Karpathy Joins Anthropic's Pretraining Team — OpenAI Co-Founder Returns to Frontier R&D
Andrej Karpathy, who co-founded OpenAI and led Tesla's AI team, announced on May 19 that he has joined Anthropic. He will work under pretraining team lead Nick Joseph, building a team to use Claude for accelerating pretraining research. Ross Nordeen, a founding xAI member, joined Anthropic earlier this month.
KPMG Deploys Claude to 276,000 Employees in 138 Countries — the Largest Professional Services AI Rollout
KPMG announced a global alliance with Anthropic on May 19, embedding Claude in Digital Gateway — the software used for tax, legal, and PE client work across 138 countries. All 276,000+ KPMG employees gain access. KPMG is also named Anthropic's preferred partner for private equity.
Boston Dynamics Atlas Carries a Full Fridge via Proprioception — Figure AI Runs 24 Hours Nonstop
Two humanoid milestones in 48 hours. Atlas used reinforcement learning and body feedback to carry 100-plus-pound loads without relying primarily on vision. Figure AI sorted 30,000 packages autonomously across a full day with zero failures.
Microsoft AI Chief Puts an 18-Month Deadline on White-Collar Automation
Mustafa Suleiman told the Financial Times that most computer-based professional work will be automated within 12 to 18 months. He named specific job categories: email, spreadsheets, code, contracts, dashboards, project trackers.
Blackstone Commits $5B to Google TPU Data Center Venture — 500MW by 2027
The world's largest private data center owner is backing Google's push to break NVIDIA's grip on AI hardware. The joint venture puts 500MW of TPU compute online by 2027, with plans to scale significantly from there.
PolyAI's Raven 3.5 Beats GPT-5 and Claude Sonnet 4.6 on All 4 Customer Service Benchmarks at Under 300ms
PolyAI's in-house Raven 3.5 model, trained on millions of real customer service calls via a GRPO+DPO fusion recipe, outperforms GPT-5 and Claude Sonnet 4.6 across all four of PolyAI's customer service evaluations at sub-300ms latency — the threshold required for natural phone conversations. The company is also launching ADK, a code-first voice agent development kit, and PolyPhone, which turns any website into a live voice agent in 10 minutes.
Agora-1: Odyssey Ships the First Multi-Agent World Model — 4 Players, One Generated World
Odyssey released Agora-1 on May 18, the first world model that supports multiple participants interacting inside the same simulation in real time. Up to 4 players share a single generated GoldenEye deathmatch world, each receiving a consistent rendered view. The architecture decouples world state from rendering, enabling applications in multiplayer gaming, robotics, and agent training inside generated environments.
Cursor Composer 2.5: 25x More Synthetic Training, Targeted RL Feedback, SpaceX Colossus Next
Cursor ships Composer 2.5 on the Kimi K2.5 base, trained with 25x more synthetic tasks than Composer 2 and a novel targeted RL technique that assigns localized feedback inside multi-hundred-thousand-token rollouts. The next model trains on SpaceX's Colossus 2 with 10x more total compute.