GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%

Live Feed

3mo ago release

Cloudflare Ships Ephemeral Agent Accounts: Deploy a Worker in Seconds, No Sign-Up Required

Cloudflare's new --temporary flag in Wrangler lets AI agents deploy live Workers without OAuth, dashboards, or MFA. Accounts last 60 minutes then auto-expire -- or the human can claim them permanently.

3mo ago funding

Goldman, JPMorgan, Morgan Stanley Agree: The AI Spending Cycle Is $5T and Nearly Half Is Debt

Three of Wall Street's largest banks have independently reached the same conclusion about AI infrastructure financing: the capital cycle runs to $5.3-5.5 trillion through 2030, nearly half debt-funded. That structural difference from the dot-com era is what makes the risk profile new territory.

3mo ago policy

Trump Clears Anthropic at G7 as White House Builds Jailbreak Scoring Benchmark to Restore Mythos Access

Eight days after the Commerce Department barred foreign nationals from Mythos 5 and Fable 5, Trump told Axios he no longer sees Anthropic as a national security threat. The White House and Anthropic are now jointly developing a formal benchmark that would quantify jailbreak severity, replacing the unworkable demand for perfect model immunity.

3mo ago research

Nobel Laureate John Jumper Leaves DeepMind for Anthropic After Nearly Nine Years

AlphaFold co-creator and 2024 Chemistry Nobel Prize winner John Jumper is joining Anthropic, the most prominent scientific researcher to leave Google DeepMind this year. He spent nearly nine years at the lab before announcing his departure on Friday.

3mo ago policy

G7 Proposes Trusted-Partner AI Access After Anthropic Blackout Hit 200 Institutions With Zero Warning

Macron told G7 leaders, Amodei, Altman, and Hassabis that a US kill switch on frontier AI damages European economies and American firms alike. Leaders left with a formal proposal for an allied-nation access scheme and no timeline.

3mo ago policy

Norway Bans AI for Ages 6-13 in Schools, Supervised-Only for Teens — Takes Effect August

Norwegian PM Stoere invoked the same reasoning that worked with smartphones in 2024: children skip foundational learning when the shortcut is always available. The rules start at the new school year and require more physical books.

3mo ago funding

Baseten Nears $1.5B at $13B — 160% Valuation Jump in Five Months as Inference Routing Consolidates

Five months after a $300M Series E at $5B, Baseten is closing a split-priced $1.5B round co-led by Spark Capital, Sands Capital, Altimeter, and Wellington. The valuation step-up reflects growing conviction that model routing infrastructure captures the next margin layer in AI.

3mo ago funding

Odyssey AI Raises $310M at $1.45B for World Models — Nvidia Passes on Series B, Amazon Takes Its Slot

The world model lab closes a $310M Series B backed by Amazon, AMD Ventures, GV, EQT, and In-Q-Tel. Nvidia backed the Series A and is absent from this round. Odyssey commits to AWS Trainium as its preferred compute substrate — a direct signal on where the next hardware loyalty play lands.

3mo ago release

Perplexity Launches Brain: Context-Graph Memory Cuts Computer Agent Cost 13%, Boosts Accuracy 25%

Perplexity's new memory system for its Computer agent builds a graph of past work sessions rather than user preferences, then synthesizes it overnight into an LLM wiki loaded before every task. Internal results: 25% correctness gain on repeated tasks, 16% recall improvement, 13% cost reduction.

3mo ago benchmark

Terminal-Bench Launches Month-Scale Challenges: Claude Opus 4.8 Ran 12 Hours and Failed

Three benchmark challenges released that would take human experts months to complete. Claude Opus 4.8 attempted the Rust compiler task for 12 hours and produced regressions. No AI has scored yet.

3mo ago benchmark

Artificial Analysis Launches AA-Briefcase: 91-Task Private Benchmark for Agentic Knowledge Work

Artificial Analysis has launched AA-Briefcase, a private evaluation that tests frontier models on four multi-week professional workflows — data science, product management, banking operations, and heavy industry strategy — across 91 tasks and thousands of input files.

3mo ago research

OpenAI: RL on Beneficial Scenarios Generalizes Alignment Gains Across Dozens of Benchmarks

OpenAI's alignment team has published research showing that reinforcement learning on realistic scenarios targeting beneficial traits produces broad, durable improvements across alignment benchmarks — including domains not seen during training and under adversarial pressure.

3mo ago funding

Pramaana Labs Raises $27M From Khosla to Apply Formal Verification to AI

Pramaana Labs has closed a $27M seed round led by Khosla Ventures to build formal verification infrastructure for AI systems, targeting hallucinations and incorrect outputs by treating language model outputs as subjects for mathematical proof.

3mo ago funding

Canada's CPP Investments Commits $741M to India's CtrlS for an 8.2% Stake and an AI Hyperscale Data Center JV

CPP Investments, Canada's largest pension fund, is taking an 8.2% stake in Hyderabad-based CtrlS for roughly $490M and co-investing in a hyperscale AI data center joint venture across India with a 48% interest. The $741M total commitment signals that sovereign-grade institutional capital has moved past US markets in its hunt for AI infrastructure returns.

3mo ago funding

XDOF Raises $70M From a16z and Thrive Capital to Build the Data Infrastructure Layer for Physical AI

XDOF has emerged from stealth with $70 million to build the pipelines, collection tools, and annotation systems that train humanoid robots. Backed by Thrive Capital, a16z, Spark Capital, and Lux, it has already partnered with UC Berkeley on a large-scale robot manipulation dataset.