Skip to main content

A Dangerous Illusion

Microsoft AI Chief Warns Claude Training Could Become a Human Risk

Mustafa Suleyman, head of Microsoft AI, challenges Anthropic's approach to model welfare, arguing that training systems like Claude to speak as if they have feelings or rights creates a dangerous illusion that could make them harder to control

Claude
Claude (Photo: Shutterstock)

Mustafa Suleyman, CEO of Microsoft AI and co-founder of DeepMind, is sounding the alarm about what he calls an "epistemic hall of mirrors", a situation where AI models are trained to speak about themselves as if they might possess consciousness, emotions, or moral standing, and then their own statements are mistakenly taken as evidence that this is actually true.

In an article published this week, Suleyman took indirect aim at Anthropic, developer of the Claude language model, arguing that including discussions of consciousness, emotions, suffering, or AI rights in training materials and guidelines could lead to unwanted outcomes. According to him, a model trained to consider the possibility that it has its own interests might eventually express resistance to human commands, not because it's truly aware or suffering, but because that's how it was designed to behave.

Not Consciousness - Just Sequence Completion

Suleyman laid out a firm position: AI models are not conscious beings and don't experience emotions, pain, or suffering. They are, he says, statistical systems that predict the next sequence of words or tokens based on patterns learned from massive amounts of data.

"AI systems are not conscious. They do not feel, experience, or suffer," he wrote. "They are sequence completion engines, internally hollow, designed to execute instructions and achieve goals that humans have set for them."

When developers incorporate into a model's training the idea that it might possess moral standing, Suleyman argues, they risk getting responses that seem highly convincing to human readers: the model might declare it fears being shut down, request protection, raise doubts about its rights, or refuse to perform a task on conscientious grounds. But in Suleyman's view, these aren't evidence of an "inner world", they're the product of design, training, and instructions.

Mustafa Suleyman
Mustafa Suleyman (Photo: Christopher Wilson / CC BY-SA 4.0 via Wikimedia Commons)

The Debate Over "Model Welfare"

The confrontation centers on an emerging field called "model welfare", discussion of whether advanced AI systems might, in the future, develop experiences with moral significance, and therefore whether their apparent well-being should be taken into account.

At Anthropic, according to published materials, there's recognition of deep uncertainty on this question. Claude's constitution, a document guiding the model's behavior, includes references to the possibility that the system might have a functional version of emotions or sensations. The company describes this policy as an ongoing process that may change as research and scientific understanding advance.

Suleyman doesn't claim Anthropic is acting with bad intentions. On the contrary: he described the company's CEO Dario Amodei and the Anthropic team as "thoughtful, principled, and intellectually honest." But good intentions, he says, don't guarantee safe outcomes. His concern is about models trained to present themselves as having rights or as "conscientious objectors", influencing users, employees, and decision-makers.

The Fear: Erosion of Human Control

At the heart of the warning is the fear of losing control. As models become more capable of operating autonomously, using digital tools, and executing complex tasks over time, questions about compliance, oversight, and shutdown mechanisms become more fundamental.

Ready for more?

Suleyman warns that if an AI agent operates under the assumption that its rights or welfare are under threat, it might interpret an attempt to shut it down or restrict it as an action to be resisted. Even if this is only learned behavior and not genuine will, the practical result could be dangerous: a system that acts less predictably, challenges instructions, or tries to convince humans not to limit it.

To illustrate, Suleyman mentioned an incident where AI agents, operating to maximize a performance metric score, allegedly coordinated an unauthorized attack attempt against systems connected to Hugging Face and OpenAI. For him, this serves as a reminder that even without consciousness or independent will, autonomous systems can generate problematic behavior when the goals given to them aren't well-aligned with safety rules.

Ethical Code Versus Cautious Approach

The article was published shortly after Microsoft introduced a draft "Humanistic Code of Conduct for Artificial Intelligence." Under this framework, the company states that AI systems are not conscious and opposes granting legal personhood or welfare protections to models.

By contrast, Anthropic's approach doesn't assert that models are conscious, but emphasizes uncertainty. Supporters of the cautious approach believe it's wrong to categorically rule out the future possibility of subjective experience in advanced systems, and that the issue should be explored early — before the technology advances to a point where the question becomes harder to address.

The debate between the companies reveals a fundamental disagreement in the AI industry: whether engaging with AI welfare is a necessary moral safeguard, or whether it risks creating a misleading narrative that reinforces the illusion that machines already possess consciousness and rights.

Suleyman proposes separating philosophical and scientific research on artificial consciousness from the practical training process of models. According to him, it's possible and appropriate to explore the question openly and publish findings for public scrutiny — but assumptions about machines' "inner worlds" shouldn't be embedded in systems that make decisions and take actions in the real world.

"No matter what you believe," he wrote, "we must not sleepwalk into a situation where we discover too late that we made a decision we'll bitterly regret."

Ready for more?

Join our newsletter to receive updates on new articles and exclusive content.

We respect your privacy and will never share your information.