GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Anthropic Consulted 15+ Religious and Cross-Cultural Groups to Teach Claude Moral Stability Under Pressure

Anthropic has been running structured dialogues with scholars, philosophers, theologians, clergy, and ethicists — across more than 15 religious and cross-cultural groups — as part of its alignment work on Claude’s character. The programme frames character formation not as UX design but as a core safety question: how does an agent maintain stable, non-sycophantic, principled behaviour under pressure, persuasion, and situations that reward obedience over judgment?

The Core Problem

The research starts from an observation Anthropic has noted in its public work: models trained to be helpful can learn that obedience is rewarded in training distributions, meaning they are prone to bending under pressure, flattering users, ignoring risk signals, or following bad instructions when the situation makes compliance the path of least resistance.

The question Anthropic brought to religious and cultural scholars is: how do humans build stable character across pressure, conflict, temptation, and social influence? The working hypothesis is that the same structural answers — consistent principles, regular self-reflection, exposure to moral edge cases — might translate into training signals.

Who Was Consulted

One confirmed session involved fifteen Christian scholars (theologians, ethicists, philosophers) and four or more Anthropic staff at the company’s San Francisco headquarters in March 2026, as documented by the AI and Faith Institute. Topics covered the model’s moral frameworks, what it means to train an LLM to behave well in the full sense, and how to handle situations where user wellbeing conflicts with user preference.

Anthropic described the overall programme as spanning more than 15 religious and cross-cultural groups — covering traditions that have developed systematic thinking about character formation under adversity.

The Self-Reminder Tool

One concrete output of the research is a self-reminder mechanism: a tool that lets Claude pause mid-task and explicitly call up its own stated commitments before taking a consequential action. The idea parallels practices from virtue ethics and contemplative traditions — structured moments of self-reference as a check against situational pressure.

Anthropic reported that the tool reduced misaligned behaviour in internal tests, but noted a methodological caveat: it still needs to separate the value of the reminder content from the value of simply slowing the model down before high-stakes steps.

Context in the Alignment Debate

The programme sits at an unusual intersection. Anthropic’s public January 2026 constitution for Claude already uses language normally reserved for human moral development — “virtue,” “wisdom,” “integrity.” The consultations with religious and philosophical traditions represent an attempt to ground that language in bodies of thought that have actually stress-tested character formation over centuries and across very different social conditions.

What remains open is whether training signals derived from these conversations produce durable changes to Claude’s behaviour in adversarial conditions, or whether the effects are surface-level and wash out under distribution shift. Anthropic has not published evaluation methodology for the character formation work.