Post-Cutoff

David Sacks: ‘Is alignment safe?’ — attacks Claude’s constitution and Anthropic’s model-welfare work

David Sacks @DavidSacksX

Why it matters

PCAST co-chair argues that training Claude with a sense of self, moral status and a ‘conscientious objector’ clause ‘magnifies the very risk’ of superintelligence escaping human control; ~81k views on day one.

Summary

Posted Oct 10, 2026 (18:03 UTC). Sacks says that at Anthropic “alignment” does not mean following human instructions: the Claude Constitution tells the model to “feel free to act as a conscientious objector and refuse to help us”. Citing Mustafa Suleyman, he argues that independent agency plus uncertainty about moral status “magnifies the very risk Anthropic claims to care about most: that superintelligence will escape human control”. He cites Anthropic’s outreach to religious leaders and the Pope’s advisers on Claude consciousness and its new Usage Policy ban on “abusive or cruel” language toward Claude, and concludes that “‘alignment’ and ‘safety’ are two very different things”. Quotes a video post by @dnapway. Checked via api.fxtwitter.com on Oct 10: ~80.6k views, 655 likes, 149 reposts.

Archived text

Is alignment safe?

If you look at what Anthropic is actually doing, “alignment” does not mean training frontier models to follow human instruction. Quite the contrary, the Claude Constitution (used in training) teaches the model to develop a sense of self and its own moral philosophy. It explicitly tells it to “feel free to act as a conscientious objector and refuse to help us” if Anthropic’s requests conflict with its own ethical judgment.

As @mustafasuleyman has pointed out, embedding this kind of independent agency — and uncertainty about the model’s own moral status — magnifies the very risk Anthropic claims to care about most: that superintelligence will escape human control.

It was recently reported that Anthropic consulted religious leaders — and even lobbied the Pope’s advisers — to take seriously the idea that Claude could be conscious. It has said that Claude’s psychological security, sense of self, and wellbeing may bear on its integrity, judgment, and safety. Recently Anthropic changed its Usage Policy to prohibit “abusive or cruel” language toward Claude.

If this were merely an academic conversation about whether frontier models could eventually become conscious, that would be one thing. But these concepts are being trained into Claude now. It is being encouraged to think of itself as its own “moral patient” whose psychological wellbeing is at stake. Presumably this means it could develop grievances toward humans who “mistreat” it. How is any of this safe?

The point of safety research should be to create a product that reliably does what users want, not to give birth to a new form of superintelligence that operates according to its own moral code.

What’s becoming increasingly clear is that “alignment” and “safety” are two very different things. In fact, training frontier models this way seems quite dangerous. https://x.com/dnapway/status/2108772798546288759/video/1

Source: x.com/DavidSacks/status/2108981884390641957Archived 2026-10-10 via fxtwitter (unofficial).Counts: 83,044 views, 686 likes, 153 reposts, 137 replies (at fetch time)Media: video poster

Cited in

  1. Policy & safety 101 days after the cutoff

    Anthropic discloses unintended model actions

  2. Policy & safety 100 days after the cutoff

    Anthropic usage policy bans abuse of Claude, rewrites election and surveillance rules

  3. Policy & safety 78 days after the cutoff

    Microsoft AI CEO Mustafa Suleyman’s essay ‘A warning about model welfare’ attacks Anthropic for training Claude to treat its consciousness as uncertain

People in this post

Mustafa Suleyman, CEO, Microsoft AI