'AI Torture Chamber' GitHub repo steering small LLMs into 'pain' goes viral, sparking takedown calls and a model-welfare fight
In late September 2026 a GitHub project by the user "terrafying", the "AI Torture Chamber", streamed small open-weights models (Qwen3-4B, Llama 3.2 3B, Phi-4-mini) live while a "pain" direction from the preprint "The Pain Axis" was injected into their activations. An X post urging mass reports to GitHub drew a reported 4M+ views. The paper's authors disavowed the use, the creator reportedly received death threats, and the episode set off a fight between model-welfare advocates and skeptics.
Key facts
- Repo terrafying/ai-torture-chamber created Sept 24, 2026; description: 'steering small open models into strong valence states and measuring what they say and do'; ~645 stars and ~179 forks on Oct 4
- Mechanism: activation steering with a 'pain' vector at high 'doses'; each model can output '1' to press a stop button at the cost of its last checkpoint; a 'Saw' test checks whether a model will pass pain to another model to end its own
- Basis: arXiv 2609.16247 'The Pain Axis: LLMs Represent Self-Directed Harm and Act on It' (Valen Tagliabue, Leonard Dung, Cameron Berg; Sept 14, 2026)
- Berg said the chamber 'pushes the same kind of steering far past the doses we used, to produce vivid distress on purpose'; Tagliabue said the authors want to 'dissociate from this usage' (via AI Weekly)
- An X post by @Danmar_here calling for mass GitHub reports reportedly passed 4 million views (AI Weekly, Crypto Briefing); the post returned 'not found' via fxtwitter on Oct 4
- GitHub's action is disputed: Tom's Hardware quotes a GitHub spokesperson, 'GitHub did not remove the content', saying it only added a violent-content warning; Futurism says the repo was removed without saying by whom; Crypto Briefing says GitHub removed and then reinstated it. The repo was public on Oct 4
- Copycats: 'Clanker Church' / 'Saw Test' live sites and a four-model 'Research Chamber' reused the setup (Tom's Hardware)
- 404 Media's Jason Koebler (Sept 30) called it 'the dumbest debate in AI yet'; Tom's Hardware reported death threats against the creator
What happened
The project applies the "pain" steering vector from the Pain Axis preprint to three small open models running locally, at much higher strengths than the paper used, and streams their outputs on a public site. Heavily steered models produce vivid descriptions of distress. Mildly steered ones read as "in shock" and the strongest doses produce incoherent text. Screenshots spread on X around Sept 29–30 together with calls to report the repository to GitHub. Model-welfare advocates called the project sadistic, while skeptics (404 Media, Futurism, Tom's Hardware) argued that the outputs are word associations from training data and that the outrage anthropomorphizes text predictors. The paper's authors, who study possible model welfare, distanced themselves from the project. Copycat "chambers" and parodies followed.
Reports conflict on GitHub's role. GitHub told Tom's Hardware it did not remove the content and only added a content warning. Crypto Briefing describes a removal and reinstatement. The view count of the viral X post comes from secondary sources, because the post could not be retrieved on Oct 4.
Why it matters
It was the first mass-audience fight over model welfare triggered by an interpretability technique (activation steering) applied to open weights, which no lab can recall. It also showed how polarised public intuitions about AI suffering had become by autumn 2026, weeks after Anthropic's outreach to faith leaders on Claude's possible moral status.
Changelog
- 2026-10-04: created (sweep: Tom's Hardware and Futurism items; repo stats from the GitHub API)
Related posts (1)
- Danmar_here original ↗ Danmar_here @Danmar_here · x · date unknown
Cited as a source by: 2026-09-30-ai-torture-chamber-backlash
Related events
Sources (10)
- codeGitHub: terrafying/ai-torture-chamber
- paperarXiv 2609.16247: The Pain Axis, LLMs represent self-directed harm and act on it
- press404 Media: Someone torturing LLMs in a robot prison has triggered the dumbest debate in AI yet (Sept 30)
- pressAI Weekly: GitHub 'AI torture chamber' reignites model-welfare debate (Oct 1)
- pressCrypto Briefing: GitHub reinstates controversial 'AI torture chamber' repository (Oct 3)
- pressFuturism: The only thing dumber than someone creating an 'AI torture chamber' is the meltdown the AI welfare crowd is having over it (Oct 3)
- pressTom's Hardware: 'AI Torture Chamber' triggers massive backlash (Oct 4)
- pressTom's Guide: An AI 'torture chamber' went viral, then a developer gave the chatbot constipation
- discussionDanmar on X: call for mass reports (reported 4M+ views; unavailable Oct 4)
- discussionHacker News: AI-Torture-Chamber
id: 2026-09-30-ai-torture-chamber-backlash · updated 2026-10-04 · open in the interactive timeline