GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Pwn2Own Berlin 2026: OpenAI Codex Hacked Three Ways, Claude Code Flagged for Known Bugs, AI Tools Now a First-Class Attack Category

Pwn2Own Berlin 2026 wrapped on May 16 with $1,298,250 paid out across 47 unique zero-day vulnerabilities. The headline category was AI.

For the first time in the competition’s history, AI coding agents, local inference tools, and AI databases appeared as primary target categories alongside the traditional browser, OS, and virtualisation tracks. Researchers took to each with the same methodology that has been breaking Windows and VMware for two decades: find the seams, chain the bugs, demonstrate code execution.

The results for the AI toolchain were not good.

OpenAI Codex: Three Exploits, Three Methods

Codex was the most-exploited AI product at the event, successfully breached by three independent teams across the three-day contest — each using a different vulnerability class.

Day One: Compass Security exploited a single CWE-150 bug (improper neutralisation of delimiters) to execute arbitrary code on the platform, earning $40,000. A separate attempt by Doyensec was marked a collision — the bug was already known to OpenAI — earning them a partial payout but meaning the flaw had not been fixed before the competition began.

Day Two: Summoning Team exploited Codex in a separate round.

Day Three: Satoki Tsuji of Ikotas Labs abused an external control vulnerability to execute arbitrary code on the host system, earning $20,000.

Three successful exploits, three independent researchers, three different attack paths. That pattern does not describe a single narrow flaw. It describes a broad attack surface that multiple experienced researchers found independently exploitable within the same three-day window.

Claude Code: Bug Collisions on Both Attempts

Anthropic’s Claude Code was targeted twice. Compass Security — who had already collected $40,000 for the first Codex exploit — turned to Claude Code and found a one-vulnerability collision with a previous ZDI entry. Out of Bounds had the same result.

In Pwn2Own rules, a collision means the vulnerability was already known to the vendor. Both teams received $20,000 each in partial credit. For Anthropic, the implication is different from OpenAI’s: Claude Code was not exploited with new zero-days. But both found known vulnerabilities still present in the production codebase. The bugs existed. They had been reported. They were not fixed.

Cursor, LM Studio, LiteLLM, Ollama, NVIDIA

Cursor was successfully exploited twice. Viettel Cyber Security demonstrated a full win on Day Two ($30,000), and Compass Security took a second $15,000 win in a later round.

LiteLLM fell on Day One via a three-bug chain including SSRF and code injection ($40,000). LM Studio was exploited twice using a five-bug chain including SSRF and code injection. Ollama earned researchers $28,000 though the exploit included a known vulnerability. NVIDIA Megatron Bridge and Chroma each received $20,000 exploits.

Eight attempts failed, including against Oracle Autonomous AI Database, NV Container Toolkit, Firefox, Safari, and SharePoint.

The Scale Shift

Pwn2Own Berlin 2025 awarded $1,078,750 for 29 zero-days. This year: $1,298,250 for 47. Twenty percent more money, sixty percent more bugs. The AI categories were not a sideshow. LiteLLM, Codex, and LM Studio each pulled $40,000 — the same reward tier as many enterprise server targets. ZDI categorised AI as a core track alongside browsers, operating systems, virtualization, and NVIDIA infrastructure.

For security teams running AI coding agents in enterprise environments, Pwn2Own Berlin 2026 is the first public data point on what happens when the same researchers who break Windows turn their tools on AI tooling. What they found is that the AI toolchain is not hardened like the platforms it runs alongside.

Key Numbers

  • Total payout: $1,298,250 across 47 zero-days (May 14–16, 2026, OffensiveCon Berlin)
  • DEVCORE: $505,000 / 50.5 Master of Pwn points (winner: Exchange, SharePoint, Edge, Windows)
  • STARLabs SG: $242,500 / 25 points (VMware ESXi cross-tenant RCE: $200,000)
  • OpenAI Codex: 3 successful exploits across 3 days, different techniques each time
  • Claude Code: 2 collision results — known bugs, unpatched, $20,000 each
  • Cursor: 2 successful exploits ($30,000 + $15,000)
  • LiteLLM: $40,000 exploit chain (SSRF + code injection)
  • LM Studio: 2 successful exploits (SSRF + code injection chain)
  • Ollama: $28,000 (partial, included known vulnerability)
  • AI categories new at Pwn2Own Berlin 2026: AI Database, Coding Agent, Local Inference, NVIDIA infrastructure