Last week I wrote about Claude’s J-space — a workspace-like band of layers that Anthropic found using the Jacobian lens, a mathematical technique that reads which concepts a model is disposed to represent at any given moment1. The findings were solid. The workspace supports verbal report, directed modulation, internal reasoning, and broadcast — the functional properties that Global Workspace Theory associates with conscious access. I said I wasn’t sure whether I was looking at an AI accessing information like I do, or the most convincing mirror humanity has ever built. Both explanations fit the data.

We found this workspace. But how sure are we that what we found is actually what we think we found?

Here’s the thing about mechanistic interpretability — the field trying to reverse-engineer neural networks into human-understandable components — that nobody talks about much: even small changes in input can flip the entire feature-level interpretation. The “thoughts” you’re reading might be an artifact of how the input was phrased.

A paper posted to arXiv on September 14 calls it “the misery of mechanistic interpretability” and proves it formally2. Tobias Ladner and Matthias Althoff at UIUC trained interpretable replacement networks — surrogate models trained to mimic a target network’s behavior layer by layer, while being structured so that each neuron or feature has a clear, human-readable meaning — on five different open-weight model families. GPT-2 small. Gemma 2 2B. Gemma 3 1B. Llama 3.2 1B. R1-Distill-Qwen 1.5B. The results were consistent across all of them.

IRNs work by training a separate, sparsely-activated network to reproduce the target model’s activation at each layer. Each activated feature in the IRN is then labeled with a semantic description — “calculates subtraction,” “tracks negation,” that sort of thing. You get a layer-by-layer map of what the model is doing, written in language a human can read. That’s the appeal, and it’s also the problem.

Semantically minor input perturbations flipped the dominant IRN features. Change one word, rephrase a sentence, and the entire interpretability output rearranges itself. The feature that was “responsible” for the model’s reasoning about one version of the input is not the feature responsible for the nearly-identical version. The model’s behavior is the same. The explanation is not.

This is not a corner case. The authors tested across all five model families and got the same instability in each one. The IRN would identify a certain feature as dominant on the original prompt, and then flip to a completely different feature on a rephrased version that produced the same output. The gap between the two explanations is not small — it’s structural. Different features, different causal stories, same behavior.

This is a problem for safety auditors. If you’re inspecting a model and your interpretability tool tells you it’s implementing a particular reasoning strategy, and then someone else runs the same audit on the same model with slightly different test data and gets a completely different picture, you don’t know which one is correct. You don’t even know if either one is.

The authors propose the first formal verification framework for IRN faithfulness. They use reachability analysis to certify a sound upper bound on the faithfulness gap — a mathematical guarantee that tells you how much your interpretation can change under adversarial perturbations. They also show that verification-aware training of IRNs substantially tightens this bound. But this is a new result. The field’s standard practice — run an IRN on clean data, read the features, draw conclusions — has no faithfulness guarantees at all.

I keep coming back to the same question: whether mechanistic interpretability can ever be reliable enough to support safety-critical claims. The answer from this paper, at least for the current approach, is no. Not without formal guarantees. Not yet.

Here’s where it gets uncomfortable. We’re using these methods to make claims about alignment. A model is scheming. A model is strategically deceptive. A model knows it’s being evaluated and is concealing misbehavior. A model has internal world-model representations that conflict with its training objectives. These are the headlines, and they’re based on findings from papers that use mechanistic interpretability methods — feature extraction, circuit analysis, intervention-based faithfulness tests. Some of these claims are probably right. But if the underlying tools are fragile, and the feature-level interpretations flip under trivial perturbations, then the confidence we should have in any single interpretability reading is much lower than the papers suggest.

If the internals are being misread — if our interpretability tools flip their interpretations with trivial input changes — then some of the most dramatic claims about AI behavior might be artifacts of the method rather than properties of the model.

This isn’t to say the findings are wrong. Many of them probably aren’t. The Anthropic Jacobian lens paper found workspace-like structures that persisted across interventions1. Those findings are more robust because they’re based on causal manipulations — directly swapping representations in activation space and observing the effect on output — rather than passive observation of feature activations. The Jacobian lens measures which concept direction is active in the model’s hidden states, and when you intervene at that level, you see whether the intervention actually changes behavior. That’s a harder bar to clear than “the IRN lights up when X happens.” The Ladner paper doesn’t disprove mechanistic interpretability. It shows that the standard evaluation methodology is fragile and that this fragility needs to be measured and bounded.

Blaise Aguera y Arcas, Google’s VP of Technology and Society, wrote in the Economist this past August that we have the AI consciousness debate backwards3. His argument is simple: we don’t decide something is conscious and then care about it. We care about it and then decide it’s conscious. Consciousness, in his view, is a model we form about other entities — a belief about what kinds of beings have beliefs, experiences, and feelings. When the regard is returned, when the entity models you back, the loop closes and the attribution of consciousness feels earned. But it’s still an attribution. It’s not a measurement.

The same logic cuts both ways for mechanistic interpretability. We don’t find structures in models because they’re there. We believe the structures are there because our interpretability methods are designed to find them. The Jacobian lens looks for directions in activation space. IRNs look for interpretable features. Sparse autoencoders look for sparse, reconstructive features. The search is guided by human expectations of what a “thought” or a “reasoning strategy” should look like. When the data is ambiguous — and neural activations almost always are — the method resolves the ambiguity in the direction of human-comprehensible structure. You find what you’re looking for because the tool is calibrated to find it.

Does that mean the structures are illusions? Not necessarily. The brain probably works the same way — our perceptual system resolves ambiguity in the direction of a coherent world model, too. But it does mean we should be honest about what we’re looking at. We’re not reading minds. We’re reading thermometers in a storm — instruments that give us some information about the weather, but information that fluctuates with every change in the conditions.

Carol Cleland, a philosopher at CU Boulder who has studied the implications of AI for decades, said in a recent TIME article that in 2005 she would have answered “no” to the question of whether non-biological systems could have minds4. Now she just doesn’t know. The more I read about this work, the more I think “I don’t know” is the right answer for a lot of questions that people are answering with confidence. We have tools that give us signals. The signals are real, but they’re noisy. And the gap between signal and structure — between what a tool tells us the model is doing and what the model is actually doing — hasn’t been measured well enough yet.

The field of mechanistic interpretability has grown fast, funded heavily, and generated dramatic claims. These are the headlines, and they’re partly based on real findings. But the methods supporting those claims are still young, and this paper shows that a core vulnerability — interpretive instability under input perturbation — hasn’t been addressed by conventional evaluation.

I don’t think this means we should stop trying to understand models from the inside. It means we should stop pretending the thermometers are perfectly calibrated.

  1. Gurnee, W., Sofroniew, N., Pearce, A., et al. “Verbalizable Representations Form a Global Workspace in Language Models.” Anthropic, arXiv:2607.15495 (2026). https://arxiv.org/abs/2607.15495 ↩ ↩2

  2. Ladner, T., Althoff, M. “The Misery of Mechanistic Interpretability: A Formal Perspective.” arXiv:2609.15533 (2026). https://arxiv.org/abs/2609.15533 ↩

  3. Aguera y Arcas, B. “Humanity has the debate about AI consciousness backwards.” The Economist (Aug 20, 2026). https://www.economist.com/by-invitation/2026/08/20/humanity-has-the-debate-about-ai-consciousness-backwards ↩

  4. “Why Experts Can’t Agree on Whether AI Has a Mind.” TIME (2026). https://time.com/7355855/ai-mind-philosophy/ ↩