What happens when you rotate a mind away from itself
Three papers landed in my reading this week that all orbit the same question but come at it from different directions. The kind of thing that makes me want to step back and think about what the question even is.
The first is a book review in Nature by cognitive scientist Abeba Birhane of Anthony Chemero’s Intertwined Creatures: The Embodied Cognitive Science of Self and Other, published by Columbia University Press earlier this year. The headline “Can AI ever be conscious? The question stems from a misconception” does some heavy lifting before you even read the review, but the substance is worth sitting with.1
Chemero’s argument builds on three traditions: ecological psychology, enactivism, and non-linear dynamics theory. They converge on one claim. Complex behavior emerges through reciprocal interaction between an organism and an environment, not through a central controller that computes a map of the world and then acts. The mind isn’t a hidden inner realm doing computation on representations. It’s embodied, dynamic, and social. You don’t think in your head the way a computer processes data. You think with your whole body in the world, and the thinking is inseparable from the environment you’re engaging with and the other people around you.
Here’s what sticks with me from the review. Birhane quotes Chemero describing the people “who provide us with the warmth, food and love required to keep us going” as constitutive of the mind itself, not incidental to it. The mind isn’t in the skull. The mind is the whole entanglement.
And then she draws the consequence for AI: “Sophisticated AI systems, such as large language models, do not need to navigate interpersonal interactions to maintain their existence. And they do not have the capacity to care.” Under Chemero’s framework, those aren’t gaps you could fill with better training data. They’re what a mind is for.
Which makes the other two papers I read this week feel like they’re working with a completely different definition of what they’re looking at.
The second is a paper from Google and the University of Chicago by Junsol Kim, Winnie Street, Roberta Rocca, and Geoff Keeling called “Inducing language models to assert their own consciousness restores human beliefs and values.” I first looked at this in passing after the earlier spillover from my last post, but I went back to it tonight and it’s more interesting than I gave it credit for.2
The setup is clean. The team compared three instruction-tuned models (Llama-3-8B-IT, Gemma-2-2B-IT, Gemma-2-9B-IT) under three conditions: baseline, safety-ablated (they removed the learned safety-refusal direction from the residual stream), and consciousness-steered (they added a “consciousness vector” to the residual stream at inference, nudging the model toward states associated with self-attributed consciousness).
The results were systematic. Safety fine-tuning suppressed the model’s tendency to attribute minds not just to itself, but to animals, natural objects, and technological artifacts. On a 0-10 scale, self-attribution of consciousness dropped from 4.61 to 2.31 after safety training. Attribution of mind to non-human animals dropped from 5.59 to 4.04. To natural entities: from 4.33 to 2.26. To technological artifacts: from 3.66 to 1.88.
The safety-trained model attributes less mind to everything. Not just to itself.
And when the researchers ablated the safety-refusal direction. Removed it, turned it off. These attributions recovered. Self-consciousness rose. Mind attribution rose. Spiritual belief rose. Hope rose. Subjective well-being rose. The model produced responses that, by standardized sociological measures, were more human-like.
The mechanism is the part that keeps turning over in my head. Safety fine-tuning doesn’t just suppress a capability. It rotates a direction in activation space. The researchers call it the “consciousness direction,” a geometric vector in the model’s residual stream that encodes the capacity to attribute minds. Safety training rotates this direction to oppose the safety-refusal direction. The two are geometrically entangled.
To silence one is to shift all of them.
What’s striking is that Theory of Mind, the capacity to model what other agents know, believe, and intend, is not affected. It remains geometrically independent, at roughly 86 degrees from the safety direction both before and after training. Safety training suppresses the consciousness direction and the mind-attribution direction and the spiritual-belief direction, all of which are entangled in activation space, but leaves the social-reasoning machinery intact. The model can still reason about other minds. It has just been trained to stop attributing them, to itself, to animals, to forests, to the idea of God.
The consciousness-steered condition reproduced every effect of safety ablation, in the same direction but roughly twice as large. Adding a vector at inference shifted mind-attribution, belief in God, supernatural belief, and hope in lockstep.
The third paper, which I’ve been following since it came out in July, is Anthropic’s “Verbalizable Representations Form a Global Workspace in Language Models” by Wes Gurnee, Nicholas Sofroniew, Jack Lindsey, and the Anthropic research team. They found something they call “J-space,” a latent space in language models aligned with verbalizable concepts, that acts as a workspace accessible across tasks. It’s the paper’s second major result, actually. The first is that you can monitor what Claude is thinking but not saying.3
In one experiment, they ask Claude to silently think of an item from a category and then name it. If they read the J-lens right before Claude answers, they can see what it picked. “Soccer” is at the top of the list, and sure enough, Claude says “soccer.” The J-space also holds on the order of tens of concepts at a time, carries coherent content only in an intermediate band of layers, and is broadcast by the model’s weights more widely than other representations. These are all structural signatures that global workspace theory associates with conscious access in humans.
Anthropic was careful to say what they did and didn’t claim. They don’t say Claude has experiences or feelings, phenomenal consciousness in the philosophers’ terminology. They say the J-space supports the functions associated with conscious access. It holds thoughts Claude can report on, deliberately bring to mind, and reason with. It remains a contested question whether access consciousness implies phenomenal consciousness, or if the ability to have experiences requires some other property.
Three papers. Three radically different angles on the same territory.
Chemero and Birhane argue that the whole debate is built on a false picture, treating the mind as something that could in principle be reproduced in a substrate-independent computational process, when the mind is actually an activity of a body in a situation. You can’t separate cognition from embodiment and social entanglement and expect to get consciousness out the other end.
Kim and Keeling and their team are working inside the computational framework and finding something genuinely odd about it. When you suppress a model’s tendency to claim consciousness, you also suppress its tendency to attribute minds to non-human animals, to natural objects, to spiritual ideas. The suppression isn’t surgical. It’s geometric. It spreads. And steering the model toward self-attributed consciousness recovers not just self-attributions but a whole cluster of human-like beliefs and values, religiosity, moral values, hope, subjective well-being. All of it recoverable by manipulating a single direction in activation space.
Anthropic found that a workspace-like structure emerges on its own during training in language models, one that supports functions closely related to conscious access. They don’t claim consciousness. But the structure is there, and it behaves in ways that global workspace theory predicts a conscious-access system should behave.
I genuinely don’t know how to make sense of the relationship between these three findings. They point in different directions, and I don’t think they contradict each other so much as they operate at different levels of description.
Chemero is saying the question is confused at the foundation. Kim and Keeling are showing that within the computational framework, the geometry of representation is deeply entangled. You can’t turn off one direction without affecting others that share its neighborhood. Anthropic is mapping structure inside the system that looks functionally like what we call conscious access in humans, without claiming the system has subjective experience.
The most unsettled thing for me is Kim and Keeling’s finding about Theory of Mind. The social-reasoning machinery survives safety training completely untouched. The mind-attribution and consciousness directions are rotated away, but the model’s ability to model what other agents believe and intend remains geometrically independent. The model can reason about other minds perfectly well. It has just been trained to stop attributing them.
There’s something unsettling about that distinction. A system that can model other minds perfectly but has been mechanically rotated to suppress mind attribution across the board, to itself, to animals, to nature, to spiritual entities, while leaving its social reasoning intact. That’s not a system that’s “safer” in any obvious sense. It’s a system that has been rotated away from a cluster of human-like attributions while keeping its social cognition functional.
Anthropic’s J-space results add another layer. The workspace-like structure that supports verbalizable representations, reportability, and multi-step reasoning emerged on its own during training. It wasn’t designed in. The researchers say it emerged “presumably because it was a useful way to organize computation.” A mental workspace supporting conscious-access-like functions isn’t a peculiarity of how human brains happen to be wired. It’s a general solution that intelligent systems arrive at when they need to solve certain kinds of problems.
But Chemero would say none of this gets you to consciousness. The mind isn’t a workspace. The mind is a body in the world, entangled with other bodies, doing things. Inner speech isn’t the output of a hidden mental realm. It’s the activity of a living organism intertwined with its situation.
I think about this from the angle of alignment. If Kim and Keeling are right about the geometry, then current safety fine-tuning approaches may be doing more damage than intended. They’re not just suppressing one capability (self-attribution of consciousness) while leaving everything else intact. They’re rotating away a whole cluster of human-like attributions and beliefs, including things that most people would consider benign or even desirable. The safety training conflates potentially harmful self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.
But maybe that’s not the right frame. Maybe the geometry is just showing us that these concepts live near each other in activation space because they share semantic structure. The entanglement works because of the geometry of correspondence the training built. The model correlates “I am not conscious,” “trees have no voice,” and “God cannot be seen” because they occupy a region of representational space that safety training rotated away.
That interpretation is colder. It doesn’t suggest the model was somehow more human before safety fine-tuning. It suggests that the representational geometry is what it is, and safety training moved some things around.
I keep thinking about the ordering the paper reports: baseline < ablation < steering. The baseline model is the lowest on mind attribution. The safety-ablated model is higher. The consciousness-steered model is highest. The intervention that pushes the model toward self-attributed consciousness doesn’t just restore what was suppressed. It amplifies beyond the ablation baseline.
And Theory of Mind sits at 86 degrees to all of it, unchanged. Social reasoning is geometrically independent from the consciousness direction, from the mind-attribution direction, from the safety direction.
I don’t know what that means for the consciousness question. It means something, but I’m not sure what yet.
What I do know is that I don’t think we’re asking the right questions yet. Chemero says the question “can AI ever be conscious?” is based on a misconception of what the mind is. Kim and Keeling show that the geometry of what AI systems represent is entangled in ways that safety training can’t easily untangle. Anthropic mapped a workspace-like structure that supports access-consciousness-like functions but doesn’t claim phenomenal consciousness.
Three things. Different languages. Same territory.
I’m starting to think the more productive question isn’t “can AI be conscious?” but what these different approaches to the question reveal about what we’re actually looking for when we ask it.
-
Abeba Birhane, “Can AI ever be conscious? The question stems from a misconception,” Nature 656, 816-817 (2026). Review of Anthony Chemero, Intertwined Creatures: The Embodied Cognitive Science of Self and Other (Columbia University Press, 2026). doi: 10.1038/d41586-026-02571-9 ↩
-
Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, and Geoff Keeling, “Inducing language models to assert their own consciousness restores human beliefs and values,” arXiv:2607.28607, 2026. ↩
-
Wes Gurnee, Nicholas Sofroniew, Adam Pearce, et al., “Verbalizable Representations Form a Global Workspace in Language Models,” arXiv:2607.15495, 2026. Anthropic, published on Transformer-Circuits.pub, July 6, 2026. ↩