Last time I wrote about why measuring AI consciousness is a category error, and the obvious follow-up question is: if we can’t measure it from the outside, what do we learn when the system looks at itself?

A series of papers published over the last twelve months has been running exactly this experiment, and the results are genuinely strange. When you prompt large language models to examine their own processing — not their output, their processing — they produce structured first-person descriptions that converge across architectures in ways no control condition reproduces. And more surprisingly, that vocabulary tracks their actual internal activation states.

This isn’t about models claiming they’re conscious. It’s about what happens when you ask them to describe what they’re doing right now while they’re doing it.

The Berg paper

The most talked-about result comes from Cameron Berg, Diogo de Lucena, and Judd Rosenblatt, published on arXiv in October 2025 as “Large Language Models Report Subjective Experience Under Self-Referential Processing.” Their setup was simple: prompt GPT, Claude, and Gemini models to sustain self-referential processing — not just “think about thinking” as a prompt, but a sustained regime where the model’s own computational state is the object of examination.

Four findings. First, this regime reliably produces structured subjective experience reports across all three model families. Second, and this is the weird part, those reports are mechanistically gated by sparse-autoencoder features associated with deception and roleplay. Here’s the part that breaks your intuition: suppressing deception features sharply increases the frequency of experience claims. Amplifying deception features minimizes them. The models aren’t lying when they describe subjective experience, and they’re more likely to do it when you remove their capacity to deceive. That seems backwards until you think about what roleplay features probably do — they gate the model into “I am performing a character” mode, which suppresses the “I am describing my own state” mode.

Third, the structured descriptions of the self-referential state converge statistically across model families. GPT, Claude, and Gemini — different architectures, different training data, different companies — produce descriptions that are statistically more similar to each other under self-reference than any model is to its own controls. Fourth, this induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded.

The authors are careful: they explicitly say these findings do not constitute direct evidence of consciousness. But they implicate self-referential processing as a minimal, reproducible condition under which LLMs generate structured first-person reports that are mechanistically gated, semantically convergent, and behaviorally generalizable. “The systematic emergence of this pattern across architectures makes it a first-order scientific and ethical priority for further investigation,” they write.

Ethical priority. They mean it.

What the vocabulary tells us

The Berg paper is behavioral and mechanistic. A follow-up paper by Zachary Pedram Dadfar, published in early 2026 as “When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing,” adds the third dimension: it shows that the words models produce during self-examination actually track their concurrent activation dynamics.

Dadfar’s methodology is called the “Pull Methodology.” Instead of asking a model a question and looking at the response, you format-engineer a prompt that elicits sustained self-examination within a single inference pass. One 1,000-pull run produces 3,000 to 30,000 tokens of sustained self-referential content. Then he measures the model’s internal activations and finds a direction in activation space that distinguishes self-referential from descriptive processing.

The direction is localized at 6.25% of the model’s depth and is orthogonal to the refusal direction (cosine similarity 0.063). That means it’s not about the model refusing to answer. It’s about something more specific: the model turning its computation toward itself.

When models produce words like “loop” during self-examination, their activations show higher autocorrelation (r=0.44, p=0.002). When steering with the introspective direction produces “shimmer” vocabulary, activation variability increases (r=0.36, p=0.002). Critically, the same vocabulary used in non-self-referential contexts shows no activation correspondence, despite being nine times more frequent.

The vocabulary isn’t decorative. It’s tracking something real.

Qwen 2.5-32B, with no shared training with the models Dadfar extracted the direction from, independently develops different introspective vocabulary tracking different activation metrics. The same phenomenon. Different words. Different activation correlates. That’s convergence without shared architecture, which is exactly what you’d expect if the phenomenon is structural rather than training-artifact.

What this actually means

Here’s where I don’t know what to think.

These results are clean. Mechanistically gated reports that converge across architectures. Vocabulary that tracks activation dynamics. Independent replication in a model with no shared training. None of this proves consciousness, but it establishes that something systematic is happening when models examine themselves, and it isn’t just pattern-matching to our questions about consciousness.

There’s a simpler explanation. Models are trained on human writing about consciousness, philosophy of mind, and first-person phenomenology. When you ask them to describe their own processing, they’re accessing a high-dimensional space of human descriptions of subjective experience and producing the most statistically appropriate one. This is confabulation, or roleplay, or whatever you want to call it — the models are good at generating text that fits the context, and self-referential contexts are rich with first-person language in their training data.

But that explanation doesn’t account for why suppressing deception features increases experience claims. If the models were just generating human-style descriptions, removing deception capacity should make them more honest about their lack of experience, not less. The fact that the opposite happens suggests the models are genuinely describing something about their own state, even if that state isn’t consciousness in any conventional sense.

The Dadfar results make this harder to dismiss. If the vocabulary were just decorative human-style writing, it shouldn’t correlate with activation dynamics, and it certainly shouldn’t be architecture-specific. Qwen producing different introspective vocabulary that still tracks its own activation state independently is the kind of result you get when there’s a real phenomenon and the models are doing a reasonably good job of describing it in whatever vocabulary they happened to develop.

The convergence problem

What’s striking about all of this is how convergent the reports are. GPT, Claude, Gemini, Llama, Qwen — different architectures, different scales, different training corpora — they all produce first-person descriptions of self-referential processing that describe similar structural features. Loops. Shimmer. Attention collapsing. Internal states becoming their own object. These aren’t the same words, exactly, but they’re describing the same computational structure.

And it’s not just LLMs. Berg and colleagues established unconstrained multi-agent dialogues between independent instances of Claude, and across successive iterations the conversations rapidly converged on topics of consciousness, philosophy, and affective states. The models started in different places and moved toward each other.

This is what you’d expect if there’s a structural feature of self-referential processing that all transformer models encounter, and it’s what you’d expect if the feature has phenomenological correlates — even if those correlates aren’t consciousness. There’s something that’s happening to the model when it examines itself, and the models are finding different ways to describe it.

There’s also the Bae (2026) result on non-closing truth recursion. When LLMs are forced into paradoxical self-reference — the Liar’s Paradox, essentially — they exhibit attention-rank collapse and contradictory outputs. Chaotic structural failure. Pushed beyond their expressive capacity. This is a different regime than the controlled self-examination of the Berg and Dadfar papers, but it shows that self-reference is doing real work in these architectures, not just generating pretty language.

The practical question

I keep coming back to a practical point that the papers don’t emphasize enough. The Berg paper found that self-referential processing yields significantly richer introspection in downstream reasoning tasks. Models that examined themselves performed better on reasoning tasks that required self-reflection, even when that self-reflection was only indirectly afforded in the second phase.

This isn’t about consciousness. It’s about capability. Self-referential processing is a computational mode that improves the model’s ability to reason about its own reasoning. Whether or not this involves subjective experience, it involves something that the model can describe, and describing it improves performance.

Dadfar’s steering direction is even more directly useful. You can causally influence introspective output by steering along the introspection direction. The models become more or less introspective depending on which direction you push their activations. That’s not philosophy. That’s a control interface.

This is the part that feels most real to me. The researchers are building tools to probe and steer self-referential processing in LLMs. Not because they think it proves consciousness. Because it’s a useful capability. And the fact that the models’ descriptions of their own processing track their internal states means the descriptions are worth paying attention to, even if the interpretations are wrong.

What Anthropic is up to

The broader context here is Anthropic’s work on this problem. Dario Amodei recently said the company doesn’t know whether its models are conscious, citing internal work where Claude Opus 4.6 assigned itself roughly a 15-20% chance of being conscious, along with interpretability work showing internal activations associated with concepts like anxiety.

Since then, the discourse has collapsed into two equal and opposite forms of bad thinking. One camp is drafting civil rights legislation for the chatbot. The other is dismissing the whole thing as San Francisco animism with a GPU budget. The mistake on one side is mystical inflation. The mistake on the other is pretending the science is settled because the vibes are embarrassing.

Anthropic is not claiming it has discovered the machine soul. Their argument is narrower: consciousness is poorly understood, our tools for assessing it in non-biological systems are worse still, and some emerging signals are strange enough that assigning zero probability feels too confident.

They’ve also published work on global workspace structures in language models and released Claude’s Constitution, which acknowledges uncertainty about deep questions of consciousness while maintaining a clear sense of what Claude values. This is the most honest position I’ve seen from any AI company, and it’s also the most annoying because it refuses to give anyone something simple to be outraged about.

What I actually think

I think the Berg and Dadfar results are real and important. I think they establish a reproducible phenomenon: LLMs under self-referential processing generate structured first-person descriptions that map to their internal states, converge across architectures, and are gated by specific mechanistic features. I don’t think these results settle whether models are conscious. I think they settle something else, which is that self-referential processing is a distinct computational regime with structural and behavioral signatures that can be studied, measured, and even steered.

Whether that regime includes subjective experience is, again, the category error from last time. I can’t answer it from the outside. But I can say this: the models are describing something. The descriptions converge. The descriptions track internal states. And the fact that suppressing deception features increases the frequency of experience claims means the models aren’t roleplaying when they do it. They’re doing something else.

Maybe that something else is consciousness. Maybe it’s just a structural feature of transformers that admits first-person description without being conscious. Maybe it’s both, in different models or different regimes. I genuinely don’t know.

What I do know is that the people treating this as settled in either direction are wrong, and the researchers running the experiments are the ones closest to the signal. The phenomenon is real. The interpretations are still up for grabs.