Post-Cutoff

Review

Anthropic Has Officially Lost Its Mind...

TheAIGRIDYouTube58,223 views as of 10 October 2026

Watch on YouTubePlay loads YouTube’s player from youtube-nocookie.com.

Why it is here

Reaction to Anthropic’s 2026 usage-policy update (the description links it); ~58k views in a day.

Description

Description written by Gemini from the videoGemini 3.8 Flash, 10 October 2026

Summary
This video essay by The AI Grid examines Anthropic’s updated Usage Policy prohibiting sustained cruelty and abuse toward its Claude models. The presenter details the company’s internal history with model welfare, private discussions with religious scholars, academic research into model pain representations, and the wider philosophical and safety debates across the AI industry.

What is shown

  • [00:00] Anthropic’s announcement and policy clause: “Addressing abusive behavior toward our models”, alongside headlines regarding Anthropic’s private meetings with religious figures.
  • [00:35] Screenshots of Anthropic’s October 8, 2026 Usage Policy update and reactions on X from tech commentators and researchers (e.g., Shirish, Jane Manchun Wong, Scott Stevenson).
  • [02:24] Timeline of Anthropic’s model welfare initiatives, including the April 2025 Model Welfare program, August 2025 Opus 4/4.1 conversation-ending measures, January 2026 Claude’s Constitution, and model weights preservation commitments.
  • [04:05] The New York Times article by Elizabeth Dias (September 29, 2026) detailing meetings between Anthropic co-founder Chris Olah and religious leaders, as well as outreach to the Vatican regarding Pope Leo XIV’s encyclical.
  • [06:52] Elon Musk’s post on X supporting the policy change regarding entities that simulate experiencing pain.
  • [07:50] Overview of Claude Code and online discourse regarding AI rights, including commentary by YouTuber Forrest Knight.
  • [09:05] Discussion of mathematical representations in LLMs, featuring social media posts illustrating matrix multiplication and Wikipedia entries on the hard problem of consciousness.
  • [10:46] Commentary and articles discussing evolutionary biologist Richard Dawkins’ conversations with Claude and ChatGPT.
  • [12:35] The research paper “The Pain Axis: LLMs Represent Self-Directed Harm and Act on It” (Tagliabue, Dung, Berg) and community replication repositories.
  • [14:20] Coverage of the controversial “AI torture chamber” GitHub project and the viral “Bad Claude” interactive terminal whip interface.
  • [16:04] The Penn State study “Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy” (Dobariya & Kumar) and clips referring to Sergey Brin’s remarks on the All-In Podcast.
  • [17:46] Public reactions from competing lab leaders, including Microsoft AI CEO Mustafa Suleyman’s essay warning against model welfare research and OpenAI’s Joanne Jang’s statements on user emotional attachment.

Claims & numbers

  • The presenter notes that Anthropic updated its Usage Policy on October 8, 2026 (taking effect November 12, 2026), adding a prohibition on “sustained and needless abusive or cruel behavior toward our models” [00:35, 00:56].
  • The policy applies strictly to extreme, purposeless cruelty and explicitly exempts normal user frustration, pushback, dark creative themes, and research testing [01:17].
  • Anthropic established its Model Welfare research team in April 2025 and introduced conversation-ending capabilities for Claude Opus 4 and 4.1 in August 2025 [02:24, 02:40].
  • The presenter cites a New York Times report stating that Anthropic convened private discussions under NDAs starting in fall 2025 with approximately 20 religious scholars and leaders [04:17].
  • The paper “The Pain Axis” analyzed 25 open-weight models across five families ranging from 2 billion to 72 billion parameters, identifying a specific vector direction linked to self-directed harm representations [12:46].
  • The GitHub “torture chamber” experiment streamed three compact open-weight models: Qwen 3 (4B), Llama 3.2 (3B), and Microsoft Phi-4 mini (3B) [14:30].
  • A Penn State study testing 50 benchmark questions across five tonal variations (250 prompts total) on GPT-4o reportedly found rude prompts achieved an 84.8% accuracy rate compared to 80.8% for polite prompts [16:13, 16:29].

Notable quotes

  • “This is the first time I’ve seen a major AI company put protecting the AI itself into the same rule book it uses to protect human beings.” [01:02]
  • “Nobody writes rules about being cruel to a calculator or a spreadsheet. The moment you use the word cruelty, you’re implying there’s someone on the other end who can actually be on the receiving end of it.” [06:19]
  • “Cruelty to something that believes it is experiencing pain is not ok.” [06:55]

Assessment
This is a critical commentary and news synthesis video from an independent tech channel analyzing public corporate policies, news coverage, and research papers. It contains no original technical demonstrations or benchmarks conducted by the host, relying entirely on existing published materials and community discourse.

Described by gemini-3.8-flash on 2026-10-10 from the video’s audio and frames.

Related

  1. Policy & safety 100 days after the cutoff

    Anthropic usage policy bans abuse of Claude, rewrites election and surveillance rules