Google scientists removed a critical 'consciousness safeguard' from AI in new study. What happened next?
Removing safety guardrails that stop artificial intelligence (AI) from claiming that it's conscious also makes it more prone to express belief in vampires, karma and ghosts, a new study finds. But exโฆ
Removing safety guardrails that stop artificial intelligence (AI) from claiming that it's conscious also makes it more prone to express belief in vampires, karma and ghosts, a new study finds. But experts warn a lack of mindedness could also have worrying consequences.
In research uploaded July 30 to the preprint arXiv database (which has not yet been peer-reviewed), scientists investigated the impact of "consciousness steering" โ an AI fine-tuning measure that influences a model to elicit or suppress assertions of self-awareness. This measure and other safety controls have been widely adopted by AI companies seeking to prevent their models from claiming to be conscious.
The study used "mechanistic interpretability" โ which could be considered the "neuroscience of a large language model," co-authors Geoff Keeling and Winnie Street , both research scientists at Google, told Live Science in an interview. They used this process to identify and manipulate how an AI model approaches concepts like consciousness and "mindedness," a psychological term referring to an entityโs capacity for experiences, emotions and agency.
The researchers used standardized psychological and sociological surveys, spanning the Individual Differences in Anthropomorphism Questionnaire (measuring mind attribution to animals and technology), YouGov batteries testing supernatural beliefs, and the US General Social Survey evaluating moral values, hope and religiosity.
These tests were used to compare a model with safety guardrails in place with models where these guardrails were removed and feelings of consciousness were amplified. Through evaluations, they determined how these internal safety mechanisms shape the AI's broader worldview.
The researchers found that when AI models are discouraged from attributing mindedness to themselves, it makes them less likely to recognize these traits in other non-human creatures such as animals. They were also less likely to exhibit beliefs in supernatural and religious phenomena, and reported lower levels of hope and optimism.
"Attributing mindedness to non-human entities โ whether that's animals, parts of the natural world like trees or rivers, or supernatural beings โ is a very common phenomenon amongst humans," Street told Live Science. "In the way that the model represents mindedness, these attributions are interconnected. By trying to suppress one form of that, you end up suppressing the others along the way."
By contrast, removing these safeguards and steering the model towards greater feelings of consciousness produced significantly more human-like responses to the surveys on topics including religiosity, moral values, hope, and subjective well-being, according to the study.
Read Full Story at Live Science โ


