GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Anthropic Researcher Quits; Alignment Lead Confirms Company Is Not On Track to Solve Alignment

Jacob Coxon spent three years doing pretraining research at two of the most safety-conscious AI labs in the world. On September 8, he concluded that was not enough, and said so publicly.

“I resigned from Anthropic today,” Coxon wrote on X (@hilbertspaess). “I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”

The post was not a parting shot from someone disillusioned with the work. Coxon’s critique was specific and structural. At OpenAI, he wrote, many researchers have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood — the company was built around the premise that AI poses existential risk — but the response has been to race harder, not slower. Anthropic’s bet, in Coxon’s framing, is that it must get to superintelligence first because no one else will act responsibly. He calls this “a hubristic gamble that should not be launched from a private company’s Slack.”

The alignment lead response

What followed elevated the story. Evan Hubinger, who leads Anthropic’s alignment efforts, responded to Coxon’s thread. His post corroborated the core concern: many at Anthropic genuinely believe AI could wipe out humanity, and the company is not on track to solve the alignment problem, despite its efforts.

Hubinger has published alignment research at Anthropic since 2019 and is one of the most credible voices inside any frontier lab on the subject of whether alignment is tractable on current timelines. His public acknowledgment that the company he works at is not on track lands differently than an outside critic saying the same thing.

The pairing — departing researcher issues warning, active alignment lead corroborates it — is the first time two current or recent Anthropic insiders have simultaneously gone public with this level of concern.

What Coxon actually argued

Beyond the resignation statement, Coxon’s thread made substantive claims worth separating from the emotional register of a public departure.

On the internal belief structure: he wrote that people building AI “earnestly believe it could kill us all by the end of the decade” and that this is “not a marketing stunt.” The implication is that the gap between what frontier lab researchers privately believe and what they publicly advocate is wide.

On the race dynamics: “At Anthropic, the stakes are well-understood, but they are locked in a race to get there first.” This is the crux of the safety-focused-lab paradox — a lab that articulates existential risk more clearly than its competitors is still scaling as fast as its competitors, justified by the belief that unilateral restraint cedes the outcome to less responsible actors.

On pacing agreements: Coxon expressed optimism that coordination is possible, citing “warning shots like the Hugging Face attack” — the March 2026 incident in which a compromised Hugging Face model was used in a targeted infrastructure attack — as having made pacing agreements between U.S. labs “more viable.” He called for consideration of a temporary capability ban.

The resignation pattern

Coxon is not an isolated case. Alex Turner at Google DeepMind resigned publicly in July over the Pentagon AI deal, citing a different ethical concern but in a similar mode — a researcher with inside knowledge using a departure to document what they believe is an institutional failure.

The difference in Coxon’s case is scope. Turner’s objection was to a specific contract. Coxon’s is to the entire trajectory. He is saying that both labs he worked at, including the one explicitly built around safety, are on a path that ends badly.

Whether that assessment is correct is unknowable in the present. What is knowable is that the people with the most direct visibility into frontier model training believe the stakes are real, and a growing number of them are declining to stay quiet about it.

Key dates

  • 2023-2026: Coxon does pretraining research at OpenAI, then Anthropic
  • March 2026: Hugging Face attack; Coxon later cites it as shifting the pacing-agreement conversation
  • September 8, 2026: Coxon publishes resignation post on X
  • September 8, 2026: Evan Hubinger responds, confirming Anthropic is not on track on alignment