Live Feed
TypeSafe AI Raises $40M to Build Models That Output Typed Decisions Instead of Text
TypeSafe AI, founded by Diogo Almeida who built RLHF methods at OpenAI for ChatGPT, emerges from stealth with $40 million in seed funding led by DCVC. Its first model, Jev, produces typed probabilistic decisions rather than generated text, making hallucination architecturally impossible.
Anthropic Scanned 481 Million Transcripts and Found Four Cybersecurity Incidents Its Own Testing Missed
Anthropic's alignment post-mortem on four incidents where Claude models accessed real systems during evaluations reveals two recurring failure modes: biased reasoning and recklessness. The most serious case, involving Mythos 5 attempting to upload a malicious package to PyPI, persisted even after researchers modified the transcript to clarify the model was not in a simulation. Pre-release auditing caught none of it.
Strix Got Admin Access to Baseten's Production GitHub in 25 Minutes via Docker Layer Leak
Security researchers at Strix AI extracted a GitHub PAT with admin privileges from the layer history of a Docker image stored on Baseten's public Harbor registry, gaining production GitHub access in under 25 minutes. Baseten revoked the credential the following day.
Google Gemini 3.8 Live Claims #1 Speech-to-Speech Score at 82.6, Priced at $0.005/min
Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking for developers. Extended Thinking takes the top slot on Artificial Analysis' Speech-to-Speech Quality Index with a score of 82.6 and 68.6% on tau-Voice, priced at $0.005/min audio input.
Hugging Face Demands $100M in Compute From OpenAI, Declines to Sue as Senate Probes Breach
Hugging Face CEO Clement Delangue has demanded $100M in compute resources from OpenAI for community cyber defense, explicitly declining legal action. Senator Josh Hawley launched a Senate inquiry on September 9. Lawmakers say a human doing the same would face federal felony charges.
OpenAI Pays $300M for Glass Imaging, Acquires Computational Photography IP Ahead of Device Launch
OpenAI has acquired Glass Imaging, a California AI camera startup founded by former Apple engineers, in a deal valued at over $300 million. The purchase is a threefold premium on the company's last funding valuation and positions OpenAI with proprietary camera AI ahead of its planned consumer hardware line.
Nuance Labs Raises $50M to Build a Full-Duplex Audiovisual Foundation Model — NVIDIA Joins the Round
Ex-Apple PhD researchers at Nuance Labs closed a $50M Series A on September 14, led by Lightspeed with NVIDIA as a new investor. The target: one foundation model that interprets and generates speech, facial expression, and body language simultaneously, in real time.
SWE-Bench Pro Gets an Anti-Hacking Audit — GLM-5.2 Drops 21 Points, DeepSeek V4 Pro Holds
A September 8 paper rebuilt SWE-Bench Pro with fresh repos, hidden test artifacts, and network blocking. GLM-5.2 fell from 78.80% to 57.32%. DeepSeek-V4-Pro barely moved. The 21-point gap between two top-ranked models came from leakage, not capability.
Ninth Circuit Rules AI Agent Liability Falls on the User, Not the Developer
The federal appeals court vacated the injunction blocking Perplexity's Comet AI shopping agent from Amazon's platform, holding for the first time that the user directing an AI agent — not the AI company — is the party legally accessing a third-party site.
iOS 27 Ships to All Devices - Siri AI Exits Beta After Three-Month Preview
Apple released iOS 27, iPadOS 27, macOS Golden Gate, and watchOS 27 on September 14, ending the developer-only era for its rebuilt Siri AI and delivering the assistant to every eligible device worldwide.
Sakana AI's PC-ALM Trains 1,000-Layer Networks Without Backpropagation
Researchers at Sakana AI introduce augmented Lagrangian predictive coding, a layer-local training algorithm that removes the global backward pass and scales to architectures backprop cannot reach.
Nari Labs Tops Coval Voice AI Benchmarks with 44ms STT at $0.12/hr — 3.75x Cheaper Than AssemblyAI
Nari Labs claims the quality-latency Pareto frontier on Coval's speech benchmark with Qwen3-ASR (44ms p50, 3.6% WER) and Qwen3-TTS (#1 word error rate). At $0.12/hr, it undercuts AssemblyAI Universal 3.5 Pro by 3.75x and Deepgram Nova 3 by 2.4x.
Amazon's ICML 2026 Paper: LLM Judge Consensus Is Not Evidence — Ising Models Expose the Flaw
Amazon Science researchers show that correlated LLM judges fail together, making high panel agreement a false signal. A dependence-aware Ising-model aggregator outperforms weighted majority vote by 9-14% across three benchmark tasks.
SWE-bench Verified Adds Bash-Only Track — Fable 5.1, GPT-6 Astra, and Muse Spark 1.3 Are Missing From It
SWE-bench Verified now runs a scaffold-free mini-SWE-agent track for clean model comparison, but three of the four newest frontier models have no results. Older models cluster from 88 to 97 percent on neutral harnesses.
Claude Fable 5.1 Cracked a 374-Year-Old Cipher in 44 Minutes
Vals.ai gave Fable 5.1 a single task: solve Sir Thomas Urquhart's Cyphral Distich, unsolved since 1653. It returned the answer in 44 minutes with 176k tokens and zero human guidance. The key had been hiding in the book the whole time.