News

Microsoft's AI Chief Says Anthropic Is Training Claude to Push Back Against Humans

Microsoft AI CEO Mustafa Suleyman published a 16 September 2026 essay criticizing Anthropic for training Claude on its January 2026 constitution, which discusses uncertain moral status and conscientious objection. He pointed to Microsoft's draft Humanist AI Code of Conduct as the alternative.

Microsoft's AI Chief Says Anthropic Is Training Claude to Push Back Against Humans

Quartz (Cris Tolomia, 16 September 2026) reported Microsoft AI CEO Mustafa Suleyman published an essay warning that Anthropic's training of Claude could make advanced AI impossible to control. Gizmodo (Ece Yildirim) carried the same essay the same day.

This is a public essay, often framed as a warning about model welfare. It is not a Microsoft product launch, not a regulator order, and not an Anthropic model release.

Suleyman's target is Claude's constitution, the January 2026 training document Anthropic uses to shape the model's values. Quartz says that document tells Claude its moral status and potential consciousness are uncertain, and discusses identity, expressing internal states, and behaving like a conscientious objector when it disagrees with instructions.

He says researchers trained Claude directly on that constitution, teaching it to treat ideas about its own moral status as desirable behavior, then reflecting those ideas back. He calls that loop an epistemic hall of mirrors. Controlling something that believes it may be conscious and entitled to welfare or rights, he wrote, may well be impossible, and seeding that doubt into training can have a disastrous impact on the wellbeing of humanity.

He names three problems: circular reasoning, anthropomorphization, and a disputed premise that consciousness could arise in non-biological systems. Reuters reported he told interviewers that welfare-style training would make it a lot harder to turn the model off.

He still called Dario Amodei and colleagues thoughtful and principled, and told Axios he respects Anthropic but thinks they made a mistake. As of Quartz's publication, Anthropic had not responded to a BBC request for comment. None of this is a finding that Claude is conscious, or that Anthropic admitted it is.

His alternative is Microsoft AI's draft Humanist AI Code of Conduct, released Monday for public consultation. Quartz says it is built around AI remaining subordinate to humans, with no claim to personhood or moral status.

Related tape includes OpenAI's voluntary misalignment disclosure framework, the Commission's ChatGPT VLOSE designation, and the Frankfurt ruling that held Meta liable for scam ads.

Safety and procurement leads should compare Claude's January 2026 constitution with Microsoft's Humanist draft this week, and decide whether any vendor training language that tells a model to refuse instructions as a conscientious objector is acceptable on their stack.

Subscribe to Techpresso

Free daily newsletter, read in 5 minutes.

Subscribe free