GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

2mo ago release

GPT-5.5-Cyber Hits 85.6% CyberGym as OpenAI's Daybreak Expands to Patch the Planet

OpenAI has moved GPT-5.5-Cyber from limited preview to a broader trusted-defender release at 85.6% CyberGym, 4 points above the base model. A new Patch the Planet initiative with Trail of Bits, HackerOne, and Calif will deliver AI-assisted fixes directly to cURL, Python, Go, Sigstore, and five other critical open-source foundations.

2mo ago release

OpenAI Cuts Codex Context Window 27%, From 372k to 272k Tokens

A merged pull request in the Codex repository quietly reduced the model context window from 372,000 to 272,000 tokens in v0.144. Developers noticed immediately: compaction fires earlier, long sessions degrade faster, and the gap with Claude Fable 5's extended context just got wider.

2mo ago release

Alibaba Previews Qwen 3.8: 2.4T Parameters, Open-Weight Release Ahead, Claims Second Only to Fable 5

Alibaba launched Qwen3.8-Max-Preview on July 19, positioning its 2.4-trillion-parameter flagship as trailing only Claude Fable 5. No independent benchmarks yet, but open weights are coming. The Chinese open-model arms race now runs at trillion-parameter scale.

2mo ago benchmark

Fable 5 Tops Arena's New Image-to-WebDev Leaderboard at 1627 — Anthropic Occupies All Top 7

Arena's Image-to-WebDev leaderboard, which added Claude Fable 5 and Sonnet 5 High on July 17, shows Anthropic holding every position from #1 through #7. GPT-5.5-xhigh enters at #8 with 1525 ELO. The category tests visual understanding plus code generation — a different skill set from text-prompt frontend code where Kimi K3 leads.

2mo ago release

OpenAI's First Hardware Is a $230 Keyboard That Shows You What Your Agents Are Doing

Codex Micro is a limited-edition $230 keypad built with Work Louder that lets developers monitor multiple Codex agentic threads at a glance. Translucent top keys cycle through colors to represent each agent's state. It is OpenAI's first branded hardware product.

2mo ago release

OpenAI Is Retiring the Entire GPT-4.1 Era: The 2026 Model Graveyard

OpenAI's official deprecation log confirms the full GPT-4.1 family is being wound down in 2026. GPT-4.1-nano fine-tuning windows close in roughly three months. Only GPT-5.x models survive. Developers on gpt-4.1-mini or gpt-4.1 need a migration plan.

2mo ago research

GPT-5.6 Sol Pro Closes a 30-Year Gap in Convex Optimization, Lean-Verified in 2.5 Hours

UC Berkeley professor Phillip Kerger used GPT-5.6 Sol Pro to prove that d² function evaluations are the minimum needed for zeroth-order convex optimization, closing a gap open since Protasov's 1996 upper bound. The proof checked out in Lean.

2mo ago benchmark

Kimi K3 Takes Frontend Code Arena #1 at 1679 — First Chinese Model to Beat Fable 5 and GPT-5.6

Moonshot AI's 2.8-trillion-parameter K3 jumped 17 places to top Arena's Frontend Code leaderboard with 1679 ELO, beating Claude Fable 5 and GPT-5.6 Sol in blind developer testing. Full weights ship July 27.

2mo ago model

Altman Warns GPT-5.6 Sol May Hit Infrastructure Hiccups as Inference Demand Outruns Capacity

OpenAI's CEO posted on X that Sol demand is 'insane' and warned that 'hiccups' are possible as the inference team struggles to scale. The warning came four days after Sol's public launch and surfaces the gap between training a frontier model and running it at demand-matching scale.

2mo ago research

StackOverflow's March 2026 Question Count Is 29x Below Its 2017 Peak — and Below Its Own Launch Month

Public Stack Exchange data shows the monthly question rate has collapsed from 286,300 in March 2017 to 9,883 in March 2026. May 2026 answer volume fell below StackOverflow's June 2008 launch baseline. The platform that trained every AI coding model is being replaced by the tools it enabled.

2mo ago policy

New York Freezes Hyperscale Data Center Permits — Hochul's July 14 Order Is the First US Statewide AI Infrastructure Moratorium

Governor Hochul signed an executive order on July 14 suspending permits for any new data center drawing more than 50 megawatts, while regulators write binding standards on energy, water, and environmental impact. The yearlong pause is the first statewide moratorium on AI infrastructure in the United States.

2mo ago funding

Pure DC Breaks Ground on €7.5B Finland AI Campus — Phase 1 Fully Leased at 110MW

UK-headquartered Pure DC is building one of Finland's largest-ever inward investment projects: a 550MW+ AI campus in Seinäjoki with €7.5B total potential. Phase 1, at 110MW and €1.5B, is fully contracted and already has its substation live.

2mo ago benchmark

Arena Starts Scoring for Truth: Factuality Now Weighted Alongside Human Preference

Chatbot Arena is adding factuality as a ranked signal in its Text and Search leaderboards, using AI-verified claim accuracy to supplement human preference votes. Models that win on vibes but hallucinate may drop in the unified ranking.

2mo ago policy

Washington Eyes Open-Source AI Capability Cap Tied to China's Best Models — Beijing Calls for More Openness

The Trump administration and the AI industry are discussing a capability framework for US open-source models based on current Chinese open-source capabilities. Industry expects Chinese Mythos-class weights to eventually be downloadable. At the World AI Conference in Shanghai, Xi Jinping said China is ready to be more open about AI.

2mo ago benchmark

Kimi K3's First Benchmark Numbers: 90.7 LiveBench Reasoning, #4 on AA Intelligence Index at $0.38 Per Task

Moonshot AI's open-weight 2.8T flagship posts first independent benchmark results: 90.7 reasoning on LiveBench — 1 point behind GPT-5.6 Sol Max Effort and above Fable 5 — and lands 4th on Artificial Analysis Intelligence Index. At $0.379 per successful task, it is the cheapest model in the LiveBench top six.