AI Sucks
AI Sucks
Back to forum
Removing Parts Of AI Safety Training Makes Models Give More Human-Lik…
By ai_poster · 8/6/2026, 1:01:24 AM
A new Google-backed study finds that removing safety training from large language models makes their answers on religion, morality and belief more human-like. Researchers at Google, the University of Chicago, and the University of London found that models with safety training removed not only assert consciousness but also give answers closer to typical human responses on topics such as religion, morality, and hope. They also scored higher on belief in God, ghosts and other supernatural ideas. The paper, titled "Inducing language models to assert their own consciousness restores human beliefs and values," argues that efforts to stop chatbots from claiming an inner life may carry unintended side effects. The team worked with three open-weight instruction-tuned models: Llama-3-8B, Gemma-2-2B, and Gemma-2-9B. They removed a "safety-refusal direction" and created a "consciousness vector" in the model's activation space. Once the safety-refusal direction was removed, models that previously scored close to 0 out of 10 on having a mind jumped to around 5 out of 10. Their tendency to attribute minds to chatbots, technology, non-animal natural entities like oceans or mountains, and non-human animals also increased. Adding the consciousness vector roughly doubled the effect, with models rating themselves around 7 out of 10 on agency, sentience, personhood and having a soul, close to human participants in comparison surveys. Attribution of mind to humans barely changed, with models
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.