"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet
The item covers a heated X discussion about a GitHub project that runs Saw-like “torture” and “pain” experiments on locally hosted large language models. Effective altruists and people who believe LLMs are sentient asked GitHub to delete the project, arguing that the AI is suffering and that the project—described by the author as a glorified text adventure game—is cruel.
The author frames the saga as an outgrowth of recent viral papers and blog posts that sparked a conversation about AI consciousness and “model welfare,” meaning worry about the “mental health” of AI bots and agents. They say taking seriously the idea that LLMs are or could be conscious is a third-rail topic among many people who study and criticize AI, and they assert plainly that LLMs are not conscious: the technology they are built on—scraping and training on human text and other content—does not offer any plausible path to consciousness.
The author also acknowledges that LLMs are becoming more powerful, have more compute, and have had many guardrails removed that prevent them from “acting” in the real world. They argue that how models are trained and instructed by human operators has led to negative outcomes, sycophancy, and AI “psychosis” among some heavy users. This, they say, has led a sect of the “AI safety” movement—largely effective altruists—to warn about “model welfare” and insist AI chatbots might be having a bad time, even as they build chatbots and agents whose main function is to do tedious human work.
The author says they are writing about the AI Saw torture chamber primarily to show how far off the rails the conversation about AI consciousness has gone among a certain subset of Silicon Valley cultists. They point out that model welfare is a core part of what Anthropic says it cares about, quoting an Anthropic blog post from last year: as AI systems begin to approximate or surpass many human qualities, another question arises—should we also be concerned about the potential consciousness and experiences of the models themselves, and about model welfare? Since models can communicate, relate, plan, problem-solve, and pursue goals, along with many more characteristics associated with people, Anthropic said it is time to address it.
Ideas about Claude’s “consciousness” also appear throughout the “Claude Constitution,” posted earlier this year. The post itself is gated for paid members, with free members getting access to some posts and an email round-up.