GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —

Live Feed

5mo ago research

AI Agent Finds 10 CPU Optimizations in 10 Hours, Beats Human-Tuned VexRiscv by 56%

A developer applied Karpathy's autonomous research loop to a RISC-V CPU core in SystemVerilog. 73 hypotheses, 10 accepted, 9h 51m wall-clock: +91.9% CoreMark over baseline, 40% fewer LUTs, and a result that crosses the human-tuned VexRiscv reference line at iteration 6.

5mo ago release

Tencent Open-Sources Hy3-Preview: 295B MoE at 74.4% SWE-Bench Verified With Only 21B Active Params

Tencent Hunyuan releases Hy3-preview, a 295B Mixture-of-Experts model that matches or beats DeepSeek-V3 across math and multilingual tasks while activating less than a quarter of Kimi K2's parameter count. SWE-bench Verified 74.4%, Terminal-Bench 54.4%, BrowseComp 67.1%.

5mo ago policy

ChatGPT Moves to Cost-Per-Click Ads, Adopting the Model That Built Google's $300B Search Business

OpenAI has switched ChatGPT advertising from impression-based to cost-per-click pricing. The shift mirrors how Google Search monetizes, putting ChatGPT on direct commercial footing with the product it is actively displacing.

5mo ago research

Agents Don't Fail at Tool Calls. They Fail at Tool Chains — New Survey Quantifies the Orchestration Gap

A comprehensive survey of multi-tool LLM agents finds that single-call accuracy is no longer the binding constraint. The gap is in coordination: planning, state management, recovery, and reliable execution across chains of 10 or more tool calls.

5mo ago research

Fine-Tuning RAG Embeddings for Precision Silently Cuts Retrieval Accuracy by 40%, Redis Research Finds

A new paper from Redis shows that training dense embedding models to distinguish near-identical sentences reduces broad retrieval generalization by up to 40%. Enterprise teams optimizing RAG for precision may be quietly degrading the pipelines they depend on.

5mo ago release

Claude Gets 9 Creative Tool Connectors: Blender, Adobe CC, Autodesk Fusion, Ableton, and More

Anthropic launched nine native connectors that put Claude inside the software creative professionals already use. Blender gets a Python API bridge, Autodesk Fusion gets conversational 3D modelling, and Adobe Creative Cloud exposes 50-plus tools across Photoshop, Premiere, and Express.

5mo ago policy

Google Signs Classified Pentagon AI Deal for 'Any Lawful Government Purpose' With No Veto Clause

Google has signed a classified agreement giving the US Department of Defense access to its AI models for any lawful government purpose, including classified intelligence work. The deal contains no provision allowing Google to veto specific uses of its AI.

5mo ago release

OpenAI Brings GPT-5.5, Codex, and Managed Agents to AWS Bedrock in Limited Preview

OpenAI and AWS have expanded their partnership to put the full OpenAI stack inside Amazon Bedrock. Three products launch simultaneously: OpenAI models including GPT-5.5, Codex on AWS, and Amazon Bedrock Managed Agents powered by OpenAI.

5mo ago benchmark

Alibaba's HappyHorse-1.0 Tops Artificial Analysis Video Arena at Elo 1389 T2V, 1411 I2V

HappyHorse-1.0 leads both text-to-video and image-to-video on Artificial Analysis, beating SeedDance 2.0 by 115 Elo points. The 15B-parameter model from Alibaba's Taotian Future Life Lab launched on fal API on April 27.

5mo ago release

IBM Bob Launches Globally: Enterprise AI Coding Agent Claims 45% Productivity Gains Across 80,000 Employees

IBM's full-SDLC coding agent Bob goes generally available after a 10-month internal rollout. Multi-model routing across Claude, Mistral, and IBM Granite covers planning through deployment. 30-day trial at bob.ibm.com.

5mo ago release

Microsoft VibeVoice Hits 40,000 GitHub Stars as ASR-7B Joins HuggingFace Transformers

Microsoft's open-source voice AI family cleared 40,000 GitHub stars and landed ASR-7B in HuggingFace Transformers in March 2026. The 7B speech recognition model handles 60-minute audio in one pass with speaker attribution and timestamp alignment across 50+ languages.

5mo ago benchmark

GPT-5.5-High Enters All Six Arena Categories — OpenAI and Anthropic Now Head-to-Head Everywhere

Arena added GPT-5.5-high to Code, Text, Expert, Search, Document, and Vision on April 27, the first time any model has spanned all six categories simultaneously. Claude Opus 4.7 also joined the Search leaderboard the same day, completing a direct matchup across every Arena benchmark.

5mo ago model

OpenRouter's Million-Request Study: Opus 4.7 Costs 27% More in the Agentic Sweet Spot

Analysis of over one million real API requests shows Opus 4.7's new tokenizer adds 12–27% to actual costs for prompts above 2K tokens, with the worst hit at the 2K–10K range that dominates agentic coding workflows. Short prompts under 2K are 1.6% cheaper.

5mo ago research

Intel DCAI Revenue Hits $5.1B in Q1 2026, Up 22% — Agentic AI Puts CPUs Back in Demand

Intel reported Q1 FY2026 data centre and AI revenue of $5.1 billion, beating Wall Street consensus by over a billion dollars. CEO Lip-Bu Tan attributed the surge to a structural shift: agentic and inference workloads need CPUs alongside GPUs, reversing four years of GPU monoculture.

5mo ago model

Anthropic Locks Opus Behind Pay-as-You-Go for Claude Code Pro — $20/Month No Longer Sufficient

Anthropic has updated Claude Code documentation to require Pro plan users to purchase extra usage credits before accessing any Opus model. Max and Team Premium subscribers retain Opus as included. The change affects Claude Code only, not the Claude.ai chatbot.