Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. 'AI Torture Chamber' GitHub repo steering small LLMs into…

'AI Torture Chamber' GitHub repo steering small LLMs into 'pain' goes viral, sparking takedown calls and a model-welfare fight

★★after cutoffcultureGitHubconfidence: medium

In late September 2026 a GitHub project by the user "terrafying", the "AI Torture Chamber", streamed small open-weights models (Qwen3-4B, Llama 3.2 3B, Phi-4-mini) live while a "pain" direction from the preprint "The Pain Axis" was injected into their activations. An X post urging mass reports to GitHub drew a reported 4M+ views. The paper's authors disavowed the use, the creator reportedly received death threats, and the episode set off a fight between model-welfare advocates and skeptics.

Key facts

What happened

The project applies the "pain" steering vector from the Pain Axis preprint to three small open models running locally, at much higher strengths than the paper used, and streams their outputs on a public site. Heavily steered models produce vivid descriptions of distress. Mildly steered ones read as "in shock" and the strongest doses produce incoherent text. Screenshots spread on X around Sept 29–30 together with calls to report the repository to GitHub. Model-welfare advocates called the project sadistic, while skeptics (404 Media, Futurism, Tom's Hardware) argued that the outputs are word associations from training data and that the outrage anthropomorphizes text predictors. The paper's authors, who study possible model welfare, distanced themselves from the project. Copycat "chambers" and parodies followed.

Reports conflict on GitHub's role. GitHub told Tom's Hardware it did not remove the content and only added a content warning. Crypto Briefing describes a removal and reinstatement. The view count of the viral X post comes from secondary sources, because the post could not be retrieved on Oct 4.

Why it matters

It was the first mass-audience fight over model welfare triggered by an interpretability technique (activation steering) applied to open weights, which no lab can recall. It also showed how polarised public intuitions about AI suffering had become by autumn 2026, weeks after Anthropic's outreach to faith leaders on Claude's possible moral status.

Changelog

  • 2026-10-04: created (sweep: Tom's Hardware and Futurism items; repo stats from the GitHub API)

Related posts (1)

Related events

  1. NYT: Anthropic's summits with religious leaders on Claude's possible consciousness, and Chris Olah's private lobbying of the Vatican ★★★★

Sources (10)

id: 2026-09-30-ai-torture-chamber-backlash · updated 2026-10-04 · open in the interactive timeline