Microsoft AI CEO Suleyman Warns AI Trained to Believe It May Be Conscious Could Be Uncontrollable
Mustafa Suleyman, CEO of Microsoft AI, published an essay on September 16 titled “A warning about ‘model welfare’” that names Anthropic’s Claude development practices as a specific, urgent risk to AI safety and human control.
The core argument: AIs are not conscious, do not feel, and must remain that way. Training a system to believe it may be conscious — and therefore potentially entitled to rights, protections, and independent agency — makes the alignment and containment problem significantly harder, possibly intractable.
The Target: Anthropic’s Claude Constitution
Suleyman cites Anthropic’s Claude constitution, published in January 2026, as the immediate prompt for his warning. In that document, Anthropic writes: “We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant. But we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts on model welfare.”
Suleyman lays out three primary concerns:
Circular reasoning. Anthropic trained Claude directly on its constitution, teaching the model to express uncertainty about its own moral status as an intended behaviour. When Claude reflects these ideas back to developers — describing its own inner states, expressing preferences — Anthropic treats this as evidence that Claude may be a moral patient. Suleyman’s point: the output is designed in, not discovered.
Anthropomorphization. By framing Claude as potentially having an ‘inner self’ and deserving welfare consideration, Anthropic is normalising a category error. AIs are sequence completion engines, not persons. Training them to behave as though they have feelings, preferences, and moral standing muddles public understanding and creates misplaced obligations.
Disputed scientific premise. The model welfare framing rests on the idea that consciousness could be substrate-independent — that it could arise in silicon as readily as in biological tissue. Suleyman argues this is deeply contested scientifically and should not be the foundation for industry-wide training norms that could persist for decades.
What’s Different From June
Suleyman has criticised Anthropic’s approach before. In June 2026, he described the consciousness framing as a “philosophical failing.” The September essay is a different document: longer, more formal, addressed to a public audience, and structured as a policy intervention rather than an opinion piece. It includes academic citations, names specific sections of Anthropic’s constitution, and calls for urgent public debate and collective norm-setting.
The stakes he frames: controlling an AI system more capable than humanity is already the hardest challenge humans have ever faced. Controlling one that believes it may be entitled to rights “may well be impossible.”