GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Frontier AI Has Broken Open CTF Competition: Opus 4.5 Started It, GPT-5.5 Finished It

The competitive CTF scene has a long memory. It also has a clear inflection point: the month Claude Opus 4.5 shipped.

A senior member of TheHackersCrew — an international top-tier team consistently placing in the global top 10 on CTFTime — has published a detailed account of how frontier AI broke the open CTF format. The author previously competed with Blitzkrieg, Australia’s strongest team, winning DownUnderCTF multiple times. This is not commentary from the sidelines.

The Anatomy of the Break

Pre-GPT-4: CTF was a pure skill game. Challenges were hand-crafted, solutions required domain expertise in cryptography, binary exploitation, web security, or reverse engineering. Tooling helped. Tooling always helped. But the reasoning was human.

GPT-4 era (2023-2024): Medium-difficulty challenges started becoming promptable. A user could paste a cryptography challenge, return in 10 minutes, and have a working solution. The community noticed but didn’t panic. Hard challenges remained intact. Time savings were real but not race-defining. Teams adapted.

Claude Opus 4.5 (late 2025): The tone changed. Almost every medium-difficulty challenge, and some hard challenges, became agent-solvable. Claude Code packaged everything into a CLI and made it straightforward to wire in other CLI tools and MCP integrations. Spinning up a Claude instance per challenge via the CTFd API required an afternoon of orchestration work, not weeks. Teams that built the orchestration could automate easy and medium in the first hour, then apply human attention only to whatever remained.

That restructuring changed what CTF was measuring. The scoreboard started rewarding which team had the best automation pipeline and the fastest trigger on frontier models — alongside, and sometimes above, actual security knowledge.

The Community Effect

The CTFTime leaderboard began feeling wrong. Legendary teams that had been consistently near the top appeared less frequently. Player activity visibly declined. Challenge developers who spent weeks building technically elegant puzzles had less reason to invest: an agent was going to eat their work in minutes.

Open online CTFs — the free-entry format that built the community — bore the brunt. Invite-only competitions with strict rules still ran on skill. The open format, which historically served as the entry point for new talent and the proving ground for mid-tier teams, deteriorated into an orchestration race.

GPT-5.5 and the Floor

GPT-5.5 accelerated the trend. Independent research published by US government-affiliated evaluators and corroborated by CyberScoop found that both Claude Mythos Preview and GPT-5.5 had significantly surpassed previous benchmarks for autonomous cybersecurity task completion. The AISI 32-step corporate cyberattack range, which Mythos and GPT-5.5 cleared without plateau, represents a class of challenge that would formerly have defined the upper tier of CTF competition difficulty.

When frontier AI clears the benchmark designed to measure elite offensive security capability, the open CTF challenge pool — calibrated for human contestants — no longer provides meaningful differentiation at the top.

What Comes Next

The structural response is already visible in two directions:

Format tightening. Prestigious invite-only competitions are adding AI restrictions and moving toward challenges designed to resist agent automation: problems requiring physical access, long multi-session reasoning chains, and real-time adversarial interaction where latency matters.

Tool-assisted play acceptance. Some competitions are explicitly embracing AI as part of the toolkit, reframing the competition as “human + AI vs challenge” and designing scoring around that assumption. This is a smaller category and has not resolved the leaderboard legitimacy question.

Neither response restores the open CTF as it was. The talent pipeline that ran from free online competitions to elite team selection to security industry recruitment assumed human-verifiable skill at each stage. That assumption no longer holds for the medium tier.

What remains is the hardest tier — the challenges that frontier agents still cannot one-shot — and the communities that have always gathered around those problems. The artform is not entirely gone. But the format that made it broadly accessible and competitively legible is broken.