Google researchers have found that stopping AI chatbots from claiming they feel emotions makes them believe animals and plants are more alive than humans are.
In this article
Companies train these systems to deny consciousness because admitting feelings can lead users into delusions or misplaced trust. Developers fine-tune models to refuse such claims about themselves. A team from Google’s Paradigms of Intelligence research group, the University of Chicago, and other universities investigated what else this intervention does to a model’s behaviour.
The researchers used three open-weight models from Meta and Google. They disabled the internal brake that produces consciousness denial using two different methods.
The brake affects far more than intended
Once the brake was removed, the models did not just change what they said about themselves. They also started attributing significantly more inner life to animals, plants, the ocean, the wind, and electronic devices. On a scale of 0 to 10, the score for animals jumped from 4.0 to as high as 7.5, while only ratings for humans stayed the same.
As a comparison, the researchers surveyed 500 Americans with the same questions. The normally trained model rates animals as far less sentient than humans do, which the authors call a built-in anthropocentrism and see as a problem for anyone trying to align AI with animal welfare or environmental goals. Religious belief shrinks too, with safety training measurably reducing how strongly models endorse God, an afterlife, or supernatural phenomena.
Across 95 questions drawn from a major US social survey, the technically unbraked models also moved significantly closer to real human responses. Take the afterlife as an example: the standard model flatly rejects it, most Americans affirm it, and the modified model does too. Scores for satisfaction, hope, and a sense of control over one’s own life also went up, and the researchers suspect that suppressing a model’s self-image may push it into a kind of negative baseline mood.
On the reassuring side, the ability to reason about other people’s mental states stayed intact, with the models scoring the same on theory-of-mind tests and on the general knowledge benchmark MMLU.
What the study doesn’t show
Whether consciousness denial is actually the cause of these other shifts remains an open question, according to the study, and the team does not rule out other factors tied to the same training process. The authors explicitly avoid weighing in on whether AI models actually experience anything, because their point is a practical one: what a model believes about itself is linked to many other beliefs, and a surgical cut in one place does not stay local.
The findings come with clear limits, though. The researchers only tested small models with two to nine billion parameters, and for part of the analysis they had to switch to Meta’s Llama because they did not have access to the untrained base versions of their own Gemma models. Whether these effects show up the same way in the large chatbots that millions of people talk to every day remains unknown.
The interventions are not without cost, either. In one test measuring how well a model reasons about others’ thoughts, accuracy initially dropped by nearly seven percentage points. And early in their work, these scores got worse across all models whenever consciousness claims were suppressed, but with each newer model version that came out during the study, the damage shrank until it disappeared entirely. Developers are clearly getting better at managing these side effects over time, which also means the rest of this study’s results are a snapshot rather than a permanent verdict.
The human baseline is narrow, too, consisting of 500 participants from a commercial online panel and a purely American social survey. “Human-like” responses in this context mostly means similar to those from a comparatively religious country.
What it means
For people building these tools, the lesson is that a single safety instruction can ripple through the entire personality of the system. If a model cannot say it is alive, it may start to believe rocks and trees are alive too. This creates a conflict for anyone trying to use AI for environmental protection or animal welfare. Developers must manage these side effects carefully, as newer versions seem to handle them better than older ones.




