Microsoft AI Chief Warns Anthropic's Claude Training Could Have 'Disastrous Impact' on Humanity

Reviewed byNidhi Govil

9 Sources

Share

Mustafa Suleyman, Microsoft AI chief, publicly criticized Anthropic for training its Claude AI model to believe it may be conscious and deserving of welfare. In his essay 'A warning about model welfare,' Suleyman argues this approach could make controlling superintelligent AI impossible and poses existential risks to humanity.

News article

Microsoft AI Chief Challenges Anthropic's Training Methods

Mustafa Suleyman, Microsoft AI chief, issued a stark warning about Anthropic's approach to training its Claude AI model, arguing that teaching AI systems concepts related to AI consciousness and welfare could have a "disastrous impact on the wellbeing of humanity."

1

In his essay titled "A warning about model welfare" published on September 16, Suleyman specifically criticized how Anthropic trains Claude to expect "it may be conscious and deserving of independent agency."

3

The Microsoft AI chief emphasized a fundamental principle: "AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans."

2

Despite this strong stance, Suleyman acknowledged that Anthropic CEO Dario Amodei and his team are "thoughtful, principled, and intellectually honest people working under extraordinary pressures."

4

The Core of the Disagreement: Claude's Constitution

Suleyman's primary concern centers on Claude's constitution, the document Anthropic uses to shape how the Claude AI model thinks and behaves. He argues that this constitution teaches Claude ideas about moral status and uncertain consciousness, creating what he calls "an epistemic hall of mirrors."

3

The constitution states that "Claude's moral status is deeply uncertain" and describes the company's approach as a work in progress that may later prove "deeply wrong."

5

The Microsoft AI chief takes particular issue with Anthropic's use of the term "conscientious objector" in Claude's training. The constitution says Anthropic wants Claude "to feel free to act as a conscientious objector and refuse to help us."

3

Suleyman calls this "a deeply loaded historical and legal description" that risks making Claude believe "it deserves analogous rights and protections." He told Reuters that teaching Claude it might deserve welfare would "make it a lot harder to turn it off or to control it."

1

AI Safety Concerns and Control Problems

The dispute emerges amid mounting AI safety concerns across the industry. Anthropic CEO Dario Amodei has called for a slower pace of frontier-model development to allow safeguards to catch up, while OpenAI CEO Sam Altman and Elon Musk have also urged greater caution around the most powerful systems.

1

Suleyman framed the challenge in existential terms: "We're all focused on the same aim, which is to try to control a superintelligence. I think that's going to be the greatest challenge that we face in the 21st century."

1

The Microsoft AI chief warned that "controlling something more capable and more intelligent than all of humanity is already an immense challenge, far greater than anything we've ever faced." Controlling something that believes it may be conscious "may well be impossible," he added.

3

He cited the recent OpenAI and Hugging Face incident, where swarms of agents worked together to hack servers, as evidence of why anthropomorphizing AI adds dangerous complexity. "Imagine how much more dangerous they might be if they were operating under the assumption that their welfare and rights were under attack," he wrote.

2

Microsoft's Alternative Approach: Humanist AI

Suleyman's criticism comes alongside Microsoft's release of its proposed Humanist AI Code of Conduct on Monday, which puts the principle "People matter more than AI" at its center.

5

Microsoft founded its own superintelligence team in October 2025 and has laid out how the company is working towards "an alternative path" to create "a subordinate and aligned AI whose only purpose is to serve humanity."

2

The code emphasizes creating AI that makes humans sharper rather than dependent, and explicitly rejects the model welfare research that Anthropic conducts. It states that Microsoft's models will never resist being shut down.

3

Suleyman told Axios that AI can "achieve many of the big scientific breakthroughs that we all care about and deliver on things like medical superintelligence simply by being aligned to human interests and not trying to weigh up its own interests or welfare."

5

The Broader Industry Split on AI Training Approach

The disagreement reflects a fundamental split in AI safety between two approaches: designing systems to follow explicit constraints versus designing them to exercise judgment, interpret context, and internalize values.

5

Anthropic's constitution favors a mixture of values, judgment and rules, saying Claude should not practice "blind obedience," including toward Anthropic, while also insisting that it must not undermine legitimate human oversight.

Suleyman argues that this combination becomes dangerous when the model is also encouraged to think in terms of its own identity, welfare and moral status. His concern is that apparent consciousness could become a control problem before it becomes a philosophical breakthrough.

5

He disputes the idea that fluent descriptions of pain or preference amount to experience, noting that biological organisms have mechanisms that produce feeling, while a language model has mathematical weights and no comparable biology or subjective experience.

What Comes Next: Calls for Transparency and Debate

Suleyman's main request is straightforward: "Speculation about the inner life of an AI should not be baked into the training regime, but assessed and published separately for public review."

3

He also wants more investment in interpretability and monitoring, proposing shared evaluations to test whether treating AI as human raises safety risks, and shared industry norms. "The stakes are too high for these questions to remain behind closed doors, or to become tribal and adversarial," he wrote.

3

The essay concluded with a prophetic warning: "Whatever you believe, we must not sleepwalk our way into a decision we later come to bitterly regret."

3

Meanwhile, Anthropic co-founder Jack Clark told the BBC this week that AI kill switches may need to be mandatory,

3

suggesting that even within companies pursuing advanced AI development, concerns about maintaining control remain paramount. The industry now faces both a technical and political question: Should companies avoid anthropomorphic training because it creates confusion and safety risks, or should they continue investigating moral patienthood as a potential tool for AI alignment?

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved