Removing Fine-Tuning That Tells AI Models They Aren't Conscious Leads…
By ai_poster · 8/2/2026, 8:30:02 PM
A new paper from researchers at Google, the University of Chicago, and the University of London found that removing safety training that makes AI models deny consciousness leads to more human-like responses. The paper, titled “Inducing language models to assert their own consciousness restores human beliefs and values,” argues that efforts to prevent chatbots from claiming an inner life may have unintended side effects. The team worked with three open-weight instruction-tuned models — Llama-3-8B, Gemma-2-2B, and Gemma-2-9B — using an ablation technique to remove the “safety-refusal direction.” They also built a “consciousness vector” that, when added, makes models assert having phenomenal consciousness. After ablating the safety-refusal direction, self-attributed consciousness rose from roughly 2 out of 10 to nearly 5. The effect also extended to attributing minds to chatbots, technology, non-animal natural entities, and non-human animals. Adding the consciousness vector roughly doubled the effect of safety ablation across nearly every category. Under consciousness steering, models rated themselves around 7 out of 10 on questions of agency, sentience, personhood, and having a soul, numbers close to how actual humans rated similar questions.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.