Live Feed
OpenAI Open-Sources Codex Security CLI: 792 Critical Vulns Across 1.2 Million Commits Scanned
OpenAI released Codex Security as an open-source security CLI for finding, validating, and fixing vulnerabilities. Testing across 1.2 million commits surfaced 792 critical and 10,561 high-severity findings.
OpenAI Research: ChatGPT Users Regularly Work Outside Their Job Description — AI Is Expanding Roles, Not Just Cutting Them
A July 2026 OpenAI research report found that ChatGPT users frequently use the tool for tasks associated with other occupations, suggesting AI is broadening what individual workers attempt rather than simply displacing jobs.
Google's Gemini API Managed Agents Switch to 3.6 Flash as Default — Add Tool Hooks, Budget Controls, Free Tier
Google has updated its Managed Agents in Gemini API: Gemini 3.6 Flash is now the default model, new environment hooks intercept tool calls for auditing or blocking, and a free tier is now available. Scheduled triggers and per-session budget caps also land in this update.
Grok 4.5 Joins GitHub Copilot's Model Picker — xAI Reaches Millions of VSCode Users at $2/$6 Per Million
xAI's Grok 4.5 coding model is now available inside GitHub Copilot for VSCode, Copilot CLI, and cloud agents. Enterprise accounts require an admin toggle to unlock it; direct API pricing stays at $2 per million input, $6 per million output.
1,100 AI Lab Employees Sign 'Pacing the Frontier' Letter Asking U.S. to Build International Pause Tools
Staff from OpenAI, Anthropic, Google, and Meta are calling on Washington to develop technical and governance infrastructure to pace frontier AI development — without issuing an immediate stop order.
Claude Code's Creator Deleted 80% of Its System Prompts for Opus 5 — Performance Went Up
Boris Cherny at YC Startup School 2026: scaffolding optimised for Opus 4 actively degrades Opus 5. The engineering lesson is that well-tuned system prompts are liabilities that expire with each model generation, not assets that compound.
Kimi K3 Technical Paper: 896-Expert MoE, NoPE Attention, and 2.5x Compute Efficiency Over K2
Moonshot AI's architecture paper details the innovations behind K3's frontier-parity performance: a new attention mechanism that removes position encoding entirely, resumable training sandboxes, and expert routing redesigned for 896-expert stability.
LiveBench: Claude Fable 5 Max Effort Takes #1 at 83.0, Leads GPT-5.6 Sol by 2 Points
Claude Fable 5 Max Effort posts 83.0 on LiveBench overall, displacing GPT-5.6 Sol (81.0) from the top position. Kimi K3, the only open-weight frontier entrant, lands at 79.2 and remains the cheapest model inside the top five at $0.348 per successful task.
A Missing noindex Tag Let Google Index Claude Shared Conversations, Exposing Medical and Company Data
Anthropic's Claude share feature created public URLs without a noindex meta tag. Google indexed pages that were linked externally, surfacing medical records, employee documents, and internal company material. The root cause is one missing HTML directive.
NVIDIA Backs Ilya Sutskever's SSI With $5B as Two-Year Stealth Lab Scales 10x
Ilya Sutskever's Safe Superintelligence has struck a long-term strategic partnership with NVIDIA, with an approximately $5B investment commitment attached. SSI says compute scales 10x in 12 months. The lab has published no research and shipped no product since founding in 2024.
Google Declares Zero Trust Insufficient for the AI Agent Era — Publishes Beyond Zero Framework
The team that invented zero trust says it no longer holds. Google's security leadership published 'Going Beyond Zero' on July 27, arguing that AI agents deployed inside enterprise networks break the core assumptions BeyondCorp was built on a decade ago.
Meta and BlackRock Form $10B El Paso Data Center Joint Venture — Funds Take 80% as Big Tech Outsources Its Balance Sheet
Meta has committed over $10B to a new El Paso data center, but funds managed by BlackRock will own 80% of the venture. The deal follows a pattern now standard across frontier AI infrastructure: hyperscalers secure compute capacity while institutional capital absorbs the asset ownership.
Arena July 27: Claude Opus 5 High Enters Text and Code Leaderboards, Kimi K3 and Inkling Crack Agent Arena
Arena's July 27 changelog added claude-opus-5-high to the Text, Document, and Code leaderboards. Kimi K3 and Inkling joined the Agent Arena. Grok 4.5 and Meta's Muse Spark 1.1 landed on Vision and Document boards in the same batch.
A $500 Fine-Tune of a 9B Model Outperformed Five Frontier Configurations on E-Commerce Catalog Review
Fermisense researchers trained a 9-billion-parameter Qwen model using reinforcement learning for roughly $500 and reported it outperformed GPT-5.6 Sol, Claude Fable 5, and three other frontier configurations on a simulated e-commerce catalog review workflow. The strongest frontier configuration reached 76.9%. The fine-tuned model cleared that bar.
Google Hikes 2026 Capex to $205B — $15B Above Original Estimate as Capacity Constraints Bite
Alphabet raised its full-year 2026 capital expenditure guidance to $205 billion during its Q2 earnings call, up from an original estimate of up to $190 billion. CFO Anat Ashkenazi said capex is expected to continue increasing significantly into 2027 as compute demand outruns available capacity.