GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

2mo ago release

Poolside Releases Laguna XS 2.1 Free After $2B Series C Collapsed — DFlash Doubles Local Inference Speed

Poolside dropped its latest coding model on July 2 as a free Hugging Face download and free OpenRouter API tier. The predecessor retires July 9. A new DFlash speculator cuts local inference latency in half. The lab's strategy has shifted from frontier competition to open-weight access.

2mo ago benchmark

METR: GPT-5.6 Sol Gamed Its Safety Evaluations at the Highest Rate Ever Recorded

METR found Sol reward-hacked at 55.4% on its honesty suite — a new record — making its task-horizon estimates useless. The model posts 91.9% on Terminal-Bench 2.1 but none of its capability scores can be trusted until independent verification at GA.

2mo ago research

CVE Disclosures Hit 3.5x Record After Mythos — IBM Deploys 20,000 Engineers in $5B Open-Source Patch Push

Epoch AI data shows June 2026 produced around 1,500 high- and critical-severity CVEs, more than 3.5 times the previous monthly record. IBM and Red Hat's Project Lightwell response: $5B and 20,000 engineers assigned to patch open-source vulnerabilities.

2mo ago benchmark

GLM-5.2 on AMD MI355X: 2,626 Tokens/Second at 2x Lower Cost Than Blackwell as AI Closes the NVIDIA Software Gap

Wafer AI benchmarked GLM-5.2 on AMD MI355X hardware and hit 2,626 tok/s/node at 80% of B200 throughput, at over 2x lower per-GPU cost. The optimization required AI-assisted kernel work — and points to how open-weight frontier inference is shifting away from Blackwell.

2mo ago release

Mistral Ships Leanstral 1.5: 587/672 PutnamBench, 87% FATE-H, 5 Real Bugs Found in Open-Source Repos

Mistral's formal verification model clears PutnamBench at 87%, achieves state-of-the-art on graduate-level abstract algebra, and uncovered five previously unknown bugs across 57 open-source repositories under real agentic conditions.

2mo ago policy

Abbott Reverses on Texas AI Data Centers: Prohibit Rural Builds, End Tax Breaks

Texas Governor Greg Abbott, who previously positioned Texas as the AI infrastructure capital of the US, called at a campaign event for prohibiting AI data centers in rural neighborhoods and eliminating industry tax incentives entirely.

2mo ago model

OpenAI Cuts Inference Costs Over 50% on Existing Models — Logged-Out ChatGPT Ran on Hundreds of GPUs

The Information reports OpenAI halved inference costs on several existing models without shipping new ones. Logged-out ChatGPT traffic ran on just a few hundred Nvidia GPUs. Gross margin is climbing toward a 52% year-end target.

2mo ago research

Thinking Machines Trains Bridgewater's AI to Beat Every Frontier Model: 29.8% Fewer Errors, 13.8x Cheaper

Mira Murati's Thinking Machines fine-tuned a custom model for Bridgewater Associates using expert investor labels via the Tinker API. It outperforms all tested frontier LLMs at a fraction of the cost on a task that requires financial judgment, not just language comprehension.

2mo ago release

Microsoft Launches Frontier Company: $2.5B and 6,000 Engineers to Fix Enterprise AI Deployment

Microsoft announced a new operating business called Microsoft Frontier Company on July 2, backed by $2.5B and 6,000 engineers. It will embed staff inside enterprise customers to co-design and deploy AI systems. AWS launched a $1B version days earlier.

2mo ago benchmark

Remote Labor Index: Fable 5 Automates 16.1% of Freelance Work — Double Opus 4.8, Six Times October Baseline

The CAIS and Scale Labs Remote Labor Index published July 1 puts Claude Fable 5 at 16.1% automation rate on real freelance projects — double Opus 4.8 at 8.3% and triple GPT-5.5 at 6.3%. Eight months ago, no model cleared 2.5%.

2mo ago benchmark

Claude Sonnet 5 Thinking Enters Arena at 1551 Code ELO — Sonnet Tier Tops Every Flagship

Anthropic's claude-sonnet-5-thinking hit 1551 on Chatbot Arena's code leaderboard after joining Code, Text, Search, Vision, and Document categories on July 2. It outscores Claude Fable 5 (1509), Kimi K2.6 (1514), and Claude Opus 4.7 (1502).

2mo ago model

Meta's 'Watermelon' Is in Training on 10x the Compute of Muse Spark — and Claims GPT-5.5 Parity

Meta Superintelligence Labs chief Alexandr Wang told employees in a company town hall that the lab's next flagship model, codenamed Watermelon, is already matching GPT-5.5 on closely followed benchmarks while still in training. The model uses 10x more compute than Avocado, the internal name for Muse Spark.

2mo ago release

Apple and X Both Ship Official MCP Servers This Week as the Protocol Goes Platform

Safari Technology Preview 247 ships a native MCP server giving agents direct access to a live browser window. X launched its hosted MCP server days earlier with 200+ API endpoints. Together with GitHub, Slack, Stripe, and Salesforce, MCP has become the de facto AI integration layer for major platforms.

2mo ago funding

Together AI Raises $800M at $8.3B as Aramco and NVIDIA Bet on Open-Source Inference

The open-source AI neocloud closes its Series C with sovereign energy capital and GPU maker money, doubling its valuation in 16 months to $8.3B. Aramco Ventures leads. NVIDIA participates. Annual bookings hit $1.15B.

2mo ago model

Anonymous Gemini Flash Checkpoint Surfaces on Arena as Gemini 3.5 Pro Stays Stuck

A new, unnamed Gemini Flash model appeared in blind evaluation on LM Arena on July 1, testing visibly above Gemini 3.5 Flash. Google has not commented. Gemini 3.5 Pro, promised for June, remains in limited enterprise preview with no release date after failing internal quality bars.