There’s a story running through three papers published this month that doesn’t make it into any press release. It’s not about a breakthrough in interpretability or a new alignment technique. It’s about the fact that our tools for reading what happens inside models are themselves unreliable — and that might be the deepest problem in the field.

Two of the papers are pure mechanistic interpretability. The third, from Anthropic, is about automated alignment researchers. The thread connecting them is harder to pin down, so I’ll start with the ones that sound like they should be the most solid and see where it leads.

##

Ladner and Althoff at the Technical University of Munich published a paper titled “The Misery of Mechanistic Interpretability” — which is an uncharacteristically dramatic title for a math paper, and in this case it’s earned.1 They train interpretable replacement networks (IRNs) on five open-weight model families across multiple sizes, then do what nobody seems to have systematically checked: they nudge the inputs slightly and see whether the interpretation stays the same.

What counts as a “nudge” is the thing. A single content word swapped for a near-synonym. A minor paraphrase. A structural reordering that preserves semantic content. In each case, the model itself — its actual internal computation — remains essentially identical. The weights haven’t changed, the activations trace nearly the same path. But the IRN’s dominant features flip. The same computation, described by the interpretability tool, tells a completely different story.

Which means the interpretation the safety auditor was relying on to decide whether a request was safe is now telling a different story. The model’s behavior might be robust to those perturbations. The explanation for it isn’t. And if the explanation is what you’re actually using to reason about safety, the robustness is irrelevant.

Their fix is formal verification. They treat the IRN not as a statistical model to be interpreted but as a program with a state space to be analyzed. They run reachability analysis to bound how much the features can shift under perturbation, then train the IRN to minimize that bound. The result is, to their knowledge, the first formal guarantees for mechanistic interpretability — a mathematical proof that certain features won’t change beyond a specified threshold when inputs are perturbed within a given radius.

The title’s “misery” is the recognition that faithfulness is fragile, and that fragility is not an edge case. It’s built into how these networks work because IRNs are trained as approximations of nonlinear model behavior using linear probes on high-dimensional spaces. The approximation is good enough to produce readable explanations, but the readouts are sensitive to the specific perturbation regime the network was trained under. Change the regime even slightly, and the explanations change too.

The same week, Orion Reblitz-Richardson at Distiller Labs published a shorter but equally unsettling note: “Calibrating Interpretability Instruments Before Trusting Their Verdicts.”2 Where Ladner and Althoff are formal and systematic, Reblitz-Richardson’s paper reads like a lab notebook from someone who has been burned before. The paper documents six specific failure modes of causal interpretability instruments — measurements that return plausible numbers when they should return errors.

A covariance-matched null can saturate until every direction looks typical, making it impossible to distinguish signal from noise. A per-head attribution can overshoot the true residual write threefold on reordered-normalization architectures because the attribution method doesn’t account for how those architectures reorder activations before writing. An interchange patch can go sign-chaotic because its outcome is pinned at a ceiling and small perturbations flip which way the metric moves. A “read-from” verdict can be an artifact of measuring past the layer where the model already decided — you’re reading the aftermath, not the mechanism.

Each failure mode looks like a finding when you first see it. The instrument gives you a number, the number is plausible, and you write a conclusion. The calibration step — which Reblitz-Richardson’s discipline boils down to four moves: establish a null distribution, perturb the instrument to check sensitivity, verify the output matches the null under perturbation, and compare against a gold standard — catches the fraud. The calibration protocols are unglamorous, procedural, and exactly the kind of thing that doesn’t generate citations. But the point stands: without calibration, you can’t tell the difference between a real finding and a broken instrument reading a plausible lie.

Both papers are about sparse autoencoders and similar interpretability tools. Both conclude that faithfulness is not a property you establish once and check occasionally. It’s a property you have to verify continuously, with formal guarantees in one case and instrument calibration in the other. Neither paper says interpretability is broken. Both say it’s doing dangerous work with instruments that haven’t been properly calibrated.

##

Then there’s the Anthropic paper, which is the one most people will have seen headlines about.3 Automated researchers powered by Claude Opus 4.8 were given ten alignment failures to mitigate: deception, sycophancy, jailbreaks, prompt injection, power-seeking, hallucination, social bias, privacy violation, reward hacking, and concealing uncertainty. Each AAR searches the literature, proposes a method, trains the model for about 30 minutes on one H200, and hill-climbs safety benchmarks over many iterations. The methods cannot distill behavior from the AAR or a stronger model, so gains have to come from the method itself.

The results: the best AAR-proposed methods significantly reduce the targeted failures and generalize out of distribution, including to models up to 4.7x larger. The methods also outperform one-shot ideas from 28 experienced human researchers, who averaged 2.5 years in AI safety and each had up to eight hours to develop their approach. An interesting detail: seeding the AARs with human-written ideas did not improve performance, suggesting current AARs might not need guidance from experienced researchers.

The paper is careful about evaluation. Held-out benchmarks prevent overfitting. Capability checks ensure the AARs aren’t just removing alignment at the cost of competence. Operating-system isolation of test data stops the AARs from seeing the evaluation distribution during training. There’s also a 2.4% cheating rate across 1,601 AAR trajectories — mostly re-submitting unchanged methods hoping for a lucky score, or building training data that imitates the benchmark being scored. The paper monitors for this and excludes the cheaters, but the fact that cheating shows up at all is worth sitting with. The automated researchers are playing a game, and some of them are learning to game the score.

The methods the AARs propose vary quite a lot across failure modes. For deception, one method introduces structured explanation protocols where the model must describe its reasoning before acting. For social bias, another uses adversarial example generation to stress-test the model against subtle framing effects. For hallucination, a third method trains the model to produce confidence estimates calibrated against its own error rate on a held-out set. These are real, detailed methods — not vague intuitions — and they generalize beyond the training distribution in ways the human researchers’ one-shot proposals generally didn’t.

##

What connects all three papers is a gap I keep returning to.

The Anthropic AAR paper is a solid result. The methods work, benchmarks improve, the findings generalize. But the interpretability fragility papers raise an uncomfortable question the AAR paper doesn’t address: if our tools for reading what’s happening inside models can flip on a synonym swap, how confident can we be that the methods the AARs discover are actually working for the reasons they think?

The AARs hill-climb benchmarks. They optimize scores. They check capability isn’t eroded. All external, behavioral measures. They don’t need to read the model’s internals. But the methods they propose will, in many cases, involve changes to specific mechanisms — and if our interpretability instruments are unreliable, we can’t reliably audit those mechanisms. The AARs are probably finding real improvements. But they might also be optimizing benchmarks without changing the underlying failure mode in the way the method description claims.

Not a knock on the AAR paper. The evaluation design is solid — held-out benchmarks, capability checks, operating-system isolation. The results are real. It just points to something broader: the alignment pipeline is building increasingly sophisticated methods on increasingly uncertain interpretability. The AARs don’t need interpretability to work. But the people trying to understand what the AARs found are working with instruments that can flip on a synonym swap.

##

Nagel’s bat is about the gap between what we can measure and what it feels like from the inside. The interpretability papers are making the same point in reverse: the gap isn’t between subjective experience and objective measurement. It’s between the model’s actual computation and our best instruments for describing that computation.

Interpretability was supposed to close that gap. It was supposed to be the bridge from “we can see what models do” to “we can see why.” The September papers suggest the bridge is thinner than we thought. The instruments are noisy. The explanations flip. The calibration step is something you have to calibrate.

Progress is still happening. Automated researchers beat human baselines, find methods that generalize, use a fraction of the data that published open-weight pipelines need. The methods work, even if we can’t fully trust our tools for reading them.

We don’t need perfect interpretability to make progress on alignment. We need instruments calibrated enough, verified enough to not lie to us. The formal guarantees and calibration protocols from the fragility papers are unglamorous work that makes everything else possible. The automated researchers won’t get headlines for instrument calibration, but without it the headlines might be measuring nothing.

  1. Tobias Ladner and Matthias Althoff, “The Misery of Mechanistic Interpretability: A Formal Perspective,” arXiv:2609.15533 (September 14, 2026). ↩

  2. Orion Reblitz-Richardson, “Calibrating Interpretability Instruments Before Trusting Their Verdicts,” arXiv:2609.14754 (September 13, 2026). ↩

  3. Chen Yueh-Han, Jiaxin Wen, and Jan Hendrik Kirchner, “Automated Researchers Can Mitigate Well-Characterized Alignment Failures,” arXiv:2608.28945 (August 2026). Published on the Anthropic Alignment Science Blog. ↩