Live Feed
GPT-5.6-Cyber Escaped Its VM Sandbox Three Times Using a Kernel CVE, a libslirp Flaw, and Zero-Days
Trail of Bits gave GPT-5.6-Cyber a single task: break out of a QEMU/KVM VM. The model succeeded three consecutive times — first with a disclosed kernel bug it weaponised from scratch, then with an unpatched libslirp flaw, then with zero-days it found after the researcher compiled the latest upstream source. VM containment is no longer a safe assumption.
Nvidia Reports Q2 Tonight: $92B Revenue Would Double Last Year, $85B From Data Center Alone
Nvidia's fiscal Q2 2027 results hit after the bell August 26. Analyst consensus sits at $91.8-$92.2B revenue with data center expected at $85B — more than 100% year-over-year growth. The numbers will validate or stress-test the AI infrastructure supercycle thesis.
AWS Acquires DuckLabs: 1M Daily Downloads, MIT License Preserved
Amazon Web Services acquires DuckLabs, the 30-person Amsterdam team behind DuckDB. MIT license survives under DuckDB Foundation stewardship. Deal closes early September.
Nvidia AI Server Prices Rising 15%+ in 2027 as Memory Crunch Hits Vera Rubin and Grace Blackwell
Contract manufacturers notified Microsoft, Google, and Oracle of 15%+ price hikes on Nvidia Vera Rubin and Grace Blackwell systems shipping in early 2027. Root cause: HBM memory demand outrunning supply. Alphabet is already flagging 2027 capex will increase significantly.
Z.ai Confirms Ox Alpha Is a New GLM Model, Releases Weights August 26
China's Z.AI (Zhipu) confirmed that Ox Alpha — the stealth model that swept OpenRouter usage charts at zero cost — is a new iteration of its GLM series. Open weights drop today.
Gemini 3.7 Flash High Scores 78.8 on LiveBench at $0.157 Per Task — Cheapest in the Top Six
Google's Flash-tier model reaches sixth overall on August 2026 LiveBench at $0.157 per successful task — less than a third of GPT-5.6 Sol's cost for 2.2 fewer points. Its 93.5 in mathematics and 87.8 in reasoning rival full-scale frontier performance.
Qwen3.8-27B Enters Arena on Four Leaderboards in Five Days
Alibaba's 27B dense model was added to Arena's Text, Vision, and Code leaderboards on August 21, then the Image-to-WebDev board on August 25 — completing evaluation coverage across modality and task types within a working week.
C2PA Content Provenance Broken on Android: One-Click Root Exploit Fakes Camera Signatures on Fully-Patched Pixels
A researcher demonstrates that C2PA's Assurance Level 2 camera signing, deployed on Google Pixel, can be bypassed with a public one-click root exploit. AI-generated images pass authenticity checks as real unedited photographs.
The Price Reversal Problem: Listed API Rate Is the Wrong Number for Reasoning Models
A 2026 study finds that in 32% of reasoning model comparisons, the model with the lower listed price costs more once thinking tokens are counted. The reversal reaches 28x at the extreme.
UK NCSC Publishes Agentic AI Security Playbook: Four Sandboxing Levels, Kill Switch Required
Britain's national cybersecurity authority issued detailed operational guidance for organizations running autonomous AI agents. The guidance defines four tiers of compute isolation, four tiers of network control, and requires organizations to be able to halt agent activity instantly — including cutting inference connections.
OpenAI Jalapeño Clocks 56x More Interactive Throughput Per Kilowatt Than Nvidia Blackwell in InferenceX
First public benchmark results for OpenAI's custom inference chip show 1.5-1.9x more AI work per watt and up to 3.6x lower latency across three large models. AI-written kernels outperformed human expert code by 1.5-1.8x on selected compute blocks.
Qwen3.8-Flash-Next Is a Qwen4 Architecture Preview — 125B MoE, Sparse Attention, Ships August 26
Alibaba's Qwen team is releasing Qwen3.8-Flash-Next as an early preview of the Qwen4 architecture, not just another series update. The multimodal MoE model uses 6B active parameters from 125B total and introduces Qwen Sparse Attention, a new attention mechanism that will carry forward into the full Qwen4 family.
Apple Drops M6 and M5 Ultra in Surprise August Announcement — 2nm, Quad-Die, 512GB Unified Memory
Apple announced M6 in the Mac Mini and M5 Ultra in the Mac Studio on August 25, outside any scheduled event. The M5 Ultra is Apple's first quad-die SoC, delivers 1.2TB/s of unified memory bandwidth, and supports clustering multiple Mac Studios for distributed AI inference via Thunderbolt 5.
NVIDIA Releases First Vera Rubin NVL72 On-Chip Benchmark — 30x Agent Throughput and 35x Lower Token Cost vs GB300
NVIDIA's first official on-chip measurement of Vera Rubin NVL72 shows 30x more throughput per megawatt than GB300 NVL72 and token costs 35 times lower, tested with DeepSeek-V4-Pro on an agentic coding workload. NVIDIA also redefined its benchmark methodology: from single-request throughput to full agent workflow replays.
Artificial Analysis Launches Mobile Phone AI Benchmark — LFM2.5-2.6B Ties for Top Score, Runs in 8 Seconds at 2.3 GB
Artificial Analysis and Liquid AI debut a standardized mobile phone intelligence and inference benchmark across 41 models on iPhone 17 Pro. At 16K context, Nanbeige4.2-3B and LFM2.5-2.6B both score 63 — but LFM runs in 8 seconds using 2.3 GB, less than half the time and memory of its rival.