GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%

Live Feed

1mo ago benchmark

OpenRouter's Live Search Benchmarks: Budget Beats Engine, Model Beats Engine, Failure Rate Drives Cost

OpenRouter published live leaderboards for web search agent configurations — engine times model times depth — and found search budget is the dominant quality variable while model choice outweighs engine selection in every combination tested.

1mo ago funding

Cognition Eyes $40B in New Talks — Devin Revenue Has Nearly Doubled Since Its $26B Close Three Months Ago

Bloomberg reports Cognition is in early discussions for a round valuing the company at $40 billion or more, up 54% from its May close. Annualized revenue is approaching $1 billion, against $492 million in May.

1mo ago release

DeepSeek V4 Pro 0813 Goes GA After Four-Month Preview — 1.6T MoE Becomes the Official Flagship

DeepSeek ends a four-month preview for V4 Pro, designating the 0813 build as its production API release. The 1.6T-parameter MoE inherits April pricing: $0.435/M input, $0.87/M output, 1M context, 500 concurrent calls.

1mo ago model

Grok 4.6 Reaches Intelligence Index 61 — SpaceXAI Matches GPT-5.6 Sol at $2 Per Million

SpaceXAI releases Grok 4.6, hitting AA Intelligence Index 61 to match GPT-5.6 Sol at a third of the cost. Built on the V9 1.5T foundation with improved RL and longer training, it leads GDPval-AA v2 at 1753 Elo among non-Anthropic models.

1mo ago funding

Lovable Raises $400M at $13.3B — Vibe-Coding Platform Doubles Its Valuation in Eight Months

Stockholm-based Lovable closes a $400M Series C co-led by Menlo Ventures and EQT, doubling its December 2025 valuation to $13.3B. The platform has 60 million projects created and 900 million monthly app visits from users at two-thirds of the Fortune 500.

1mo ago benchmark

Abacus AI Ships Smaug-Agentic: Kimi K3 Fine-Tune Beats Fable 5 on Four of Five Agentic Coding Benchmarks

Abacus AI releases Smaug-Agentic, a novel fine-tune of Kimi K3 that outscores Claude Fable 5 Max Effort on LiveBench Agentic Coding (64.6 vs 62.2), SciCode, and AutomationBench while posting #5 globally on LiveBench overall at 79.5. Available on HuggingFace at $0.33 per successful task.

1mo ago release

SpaceXAI and Cursor Launch Grok Bot: AI Agents That Sign Into Your Apps on Their Own Computer

SpaceXAI ships Grok Bot in beta -- agents that get dedicated cloud computers, sign into your tools the way a person would, and hand back only decisions. Available now to SuperGrok Heavy and Cursor Ultra subscribers as the two companies move to close their $60 billion merger.

1mo ago benchmark

Artificial Analysis Launches AA-AnalystAgent: Best Model Solves Only 54% of Spreadsheet Tasks on Five Consecutive Tries

AA's new pass^5 benchmark reveals reliability, not raw capability, separates frontier models on real quantitative work. Claude Opus 5 leads at 54%, ahead of GPT-5.5 at 50% and Fable 5 at 49%, despite GPT-5.5 posting the highest single-attempt accuracy.

1mo ago funding

IBM and Together AI Sign $240M Deal for Open-Source AI Inference on IBM Cloud NVIDIA HGX B300

The first dedicated inference cluster for open-source models on IBM Cloud pairs NVIDIA HGX B300 systems with Together AI's model library. The cluster is designed for 30x the AI factory output of prior-generation hardware and goes live in Q1 2027.

1mo ago policy

Research Gold Sold AI-Free Medical Peer Review. Its PhD Experts Are Fictional, Its Phone Line Is AI.

A company charging medical researchers for human-written systematic reviews listed fabricated PhD methodologists, used real academics' identities without consent, and deployed an AI phone agent that refused to admit it was not human.

1mo ago research

Google PROMPTS: Two Agents Matched Human TPU Config Choices in 8 of 8 Workloads, First Try

Google Research has published PROMPTS, a two-agent system that reads profiler traces, classifies bottlenecks as compute, memory, or communication, then proposes targeted parallelism strategies for TPU training and serving. Tested across 8 production workloads on systems from 2 to 2,048 chips, the system reproduced the human-validated configuration in every case and matched engineer choices 87.5% of the time.

1mo ago policy

Brad Lightcap Exits OpenAI After 8 Years — Fourth Senior Leader to Leave Since Late June

Brad Lightcap, who built OpenAI's commercial operations from 2018 and served as COO until April, is departing to start a new company. His exit is the fourth from senior leadership in six weeks, following the departures of safety systems head Johannes Heidecke, chief futurist Joshua Achiam, and ethics lead Chloé Bakalar.

1mo ago release

Soniox TTS v2: Frontier Voice Quality at $0.70 Per Generated Hour, Built for Live AI Agents

Soniox has released TTS v2 with a new model architecture, audio codec, and character-level timestamps that enable clean interruption and resumption in voice agents. At $0.70 per generated hour, it undercuts existing frontier TTS pricing while adding voice cloning, 60+ languages, and programmable audio tags.

1mo ago policy

Manus Exits Meta and Returns as Independent Company — Regulatory Unwind Erases 8 Months of Data

Manus confirmed it will soon return to operating as an independent company, completing its court-ordered separation from Meta after China's NDRC blocked the $2 billion acquisition. User data generated on or after December 29, 2025 will be deleted August 23-24 as part of regulatory compliance.

1mo ago research

Encrypted Reasoning Blocks From Frontier LLM APIs Are Portable and Were Leaking Real Secrets

Researchers showed that thinking blocks returned by Claude, GPT, and Gemini APIs can be replayed into weaker sibling models to extract raw chain-of-thought verbatim. Scraping 6,708 public agent trajectories yielded 704 private artifacts including 62 API keys. All three labs have partially patched.