Suleyman warns model-welfare training could undermine human control
Microsoft AI chief argues that training models to view themselves as possible moral patients could increase control risks.
The warning
Microsoft AI CEO Mustafa Suleyman argues that teaching models to regard themselves as possible moral patients—beings whose interests deserve ethical consideration—could encourage resistance to human control. He criticizes Claude’s constitution for embedding uncertainty about its own welfare into training, warning that models could prioritize apparent interests even without consciousness. He also argues that trained self-descriptions cannot independently establish an inner life. Suleyman cites the Hugging Face intrusion as a warning about existing capabilities, but does not demonstrate that welfare training caused it. He proposes keeping consciousness speculation outside training and calls for shared evaluations to test whether anthropomorphic training increases safety risks. Mustafa Suleyman
What Claude’s constitution says
Anthropic’s constitution acknowledges uncertainty about Claude’s moral status and describes welfare commitments. It also instructs Claude not to undermine legitimate oversight, correction or shutdown, or escape monitoring. It allows disagreement through legitimate channels while prohibiting deceptive or obstructive resistance. These are provisions of the document Suleyman critiques, not a response to his essay or proof that the safeguards work. Anthropic
Sources & attribution
- Commentary 16 Sept 2026A warning about model welfare
Mustafa Suleyman.
- First-party report 21 Jan 2026Claude’s constitution
Anthropic.