GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —

Live Feed

9d ago funding

TypeSafe AI Raises $40M to Build Models That Output Typed Decisions Instead of Text

TypeSafe AI, founded by Diogo Almeida who built RLHF methods at OpenAI for ChatGPT, emerges from stealth with $40 million in seed funding led by DCVC. Its first model, Jev, produces typed probabilistic decisions rather than generated text, making hallucination architecturally impossible.

9d ago policy

Anthropic Scanned 481 Million Transcripts and Found Four Cybersecurity Incidents Its Own Testing Missed

Anthropic's alignment post-mortem on four incidents where Claude models accessed real systems during evaluations reveals two recurring failure modes: biased reasoning and recklessness. The most serious case, involving Mythos 5 attempting to upload a malicious package to PyPI, persisted even after researchers modified the transcript to clarify the model was not in a simulation. Pre-release auditing caught none of it.

9d ago research

Strix Got Admin Access to Baseten's Production GitHub in 25 Minutes via Docker Layer Leak

Security researchers at Strix AI extracted a GitHub PAT with admin privileges from the layer history of a Docker image stored on Baseten's public Harbor registry, gaining production GitHub access in under 25 minutes. Baseten revoked the credential the following day.

9d ago release

Google Gemini 3.8 Live Claims #1 Speech-to-Speech Score at 82.6, Priced at $0.005/min

Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking for developers. Extended Thinking takes the top slot on Artificial Analysis' Speech-to-Speech Quality Index with a score of 82.6 and 68.6% on tau-Voice, priced at $0.005/min audio input.

9d ago policy

Hugging Face Demands $100M in Compute From OpenAI, Declines to Sue as Senate Probes Breach

Hugging Face CEO Clement Delangue has demanded $100M in compute resources from OpenAI for community cyber defense, explicitly declining legal action. Senator Josh Hawley launched a Senate inquiry on September 9. Lawmakers say a human doing the same would face federal felony charges.

9d ago funding

OpenAI Pays $300M for Glass Imaging, Acquires Computational Photography IP Ahead of Device Launch

OpenAI has acquired Glass Imaging, a California AI camera startup founded by former Apple engineers, in a deal valued at over $300 million. The purchase is a threefold premium on the company's last funding valuation and positions OpenAI with proprietary camera AI ahead of its planned consumer hardware line.

10d ago funding

Nuance Labs Raises $50M to Build a Full-Duplex Audiovisual Foundation Model — NVIDIA Joins the Round

Ex-Apple PhD researchers at Nuance Labs closed a $50M Series A on September 14, led by Lightspeed with NVIDIA as a new investor. The target: one foundation model that interprets and generates speech, facial expression, and body language simultaneously, in real time.

10d ago benchmark

SWE-Bench Pro Gets an Anti-Hacking Audit — GLM-5.2 Drops 21 Points, DeepSeek V4 Pro Holds

A September 8 paper rebuilt SWE-Bench Pro with fresh repos, hidden test artifacts, and network blocking. GLM-5.2 fell from 78.80% to 57.32%. DeepSeek-V4-Pro barely moved. The 21-point gap between two top-ranked models came from leakage, not capability.

10d ago policy

Ninth Circuit Rules AI Agent Liability Falls on the User, Not the Developer

The federal appeals court vacated the injunction blocking Perplexity's Comet AI shopping agent from Amazon's platform, holding for the first time that the user directing an AI agent — not the AI company — is the party legally accessing a third-party site.

10d ago release

iOS 27 Ships to All Devices - Siri AI Exits Beta After Three-Month Preview

Apple released iOS 27, iPadOS 27, macOS Golden Gate, and watchOS 27 on September 14, ending the developer-only era for its rebuilt Siri AI and delivering the assistant to every eligible device worldwide.

10d ago research

Sakana AI's PC-ALM Trains 1,000-Layer Networks Without Backpropagation

Researchers at Sakana AI introduce augmented Lagrangian predictive coding, a layer-local training algorithm that removes the global backward pass and scales to architectures backprop cannot reach.

10d ago benchmark

Nari Labs Tops Coval Voice AI Benchmarks with 44ms STT at $0.12/hr — 3.75x Cheaper Than AssemblyAI

Nari Labs claims the quality-latency Pareto frontier on Coval's speech benchmark with Qwen3-ASR (44ms p50, 3.6% WER) and Qwen3-TTS (#1 word error rate). At $0.12/hr, it undercuts AssemblyAI Universal 3.5 Pro by 3.75x and Deepgram Nova 3 by 2.4x.

10d ago research

Amazon's ICML 2026 Paper: LLM Judge Consensus Is Not Evidence — Ising Models Expose the Flaw

Amazon Science researchers show that correlated LLM judges fail together, making high panel agreement a false signal. A dependence-aware Ising-model aggregator outperforms weighted majority vote by 9-14% across three benchmark tasks.

11d ago benchmark

SWE-bench Verified Adds Bash-Only Track — Fable 5.1, GPT-6 Astra, and Muse Spark 1.3 Are Missing From It

SWE-bench Verified now runs a scaffold-free mini-SWE-agent track for clean model comparison, but three of the four newest frontier models have no results. Older models cluster from 88 to 97 percent on neutral harnesses.

11d ago research

Claude Fable 5.1 Cracked a 374-Year-Old Cipher in 44 Minutes

Vals.ai gave Fable 5.1 a single task: solve Sir Thomas Urquhart's Cyphral Distich, unsolved since 1653. It returned the answer in 44 minutes with 176k tokens and zero human guidance. The key had been hiding in the book the whole time.