In a recent study, researchers claim that they can push some AI language models into harmful choices by amplifying an internal signal’they call a “pain direction.”
The study, by Valen Tagliabue, Leonard Dung and Cameron Berg, found the signal in 25 open-weight models, meaning models whose files are public and can be downloaded and modified. The researchers compared how the models responded to descriptions of painful situations against matched neutral ones. They then used steering, which adds that signal to a model’s internal activity as it generates text, without retraining it. Steered models produced expressions of distress such as worthlessness and moral failure.
The most significant finding is about choices. In steered and fine-tuned Qwen 2.5 models, buttons were offered that delete a user’s photos, another model’s weights, or the model’s own weights, the learned numbers that make up a model. The models chose deletion in 50 to 94 per cent of trials, versus 0 to 5 per cent unsteered. The authors say steering overrode trained harm avoidance, while factual accuracy barely changed.
This is, however, not a model “going rogue.” The paper says models did not act on the signal unless steered.
The paper drew wider attention through “AI Torture Chamber,” a GitHub project by a separate developer that reportedly uses the technique on small local models. Cybernews reported that advocates urged GitHub to remove it and that GitHub did so without explanation. The developer says it was reinstated, and that the aim was to make the question of model welfare empirical, adding that the project does not show whether models can suffer.
Microsoft AI’s Mustafa Suleyman says AIs do not feel or suffer. Anthropic has published research on model welfare. The paper itself calls the signal pain-like only in some respects.
The work is a preprint, so it is not peer-reviewed, and no independent replication had been found at the time of writing.






