Impossibility results, 2026
Three papers published in 2026 share a shape I have not seen before in AI research. Not because their conclusions are similar — they are not — but because they all produce impossibility results, and they all arrive at them through the same structural move. They take a question that seems to demand an answer and show that the question’s own architecture prevents any answer from being robust.
The University of Bradford and Rochester Institute of Technology study from February applied consciousness measures — the kind used to distinguish wakefulness from sleep in humans — to GPT-2. The researchers deliberately damaged the model by removing components responsible for prioritizing information, then adjusted the temperature parameter. Under certain conditions, the AI’s “consciousness-style” score actually increased after the model was degraded. Under other conditions, it fell or barely changed. The number said less about the AI than about the measure. You cannot use a thermometer that goes up when the thing it is measuring breaks.1
Apostol Vassilev at NIST published a paper in May establishing information-theoretic limitations for AI security and alignment by extending Gödel’s incompleteness theorem. The result is clean and unforgiving. For any alignment policy Π and any AI system based on computation, there exist adversarial prompts that evade the policy. No set of guardrails can be both complete and consistent. The paper is titled “A Sisyphean Endeavor” and it earns the title.2
Alexander Lerchner at Google DeepMind published a paper in September in Behavioral and Brain Sciences arguing that computation is a description sandwiched by meaning that is supplied by an external observer. Observer-dependence means computation can only simulate mind, not realize it. Computational functionalism is logically untenable. Consciousness requires a system with intrinsic meaning — a living, teleodynamic organism. The argument is philosophical rather than mathematical, but it hits the same note: you cannot bootstrap consciousness from computation alone because computation is not self-grounding.3
And then there is OpenAI’s September 21 policy proposal, the one that arrived two days before Sam Altman briefed the UN Security Council. It calls for global technical standards for recursive self-improvement. RSI is what you get when AI systems start taking on more of the work of developing their own successors. The proposal asks the United States to lead this effort because its AI industry is at the technical frontier and because it “stands in a privileged global network position in critical areas such as finance, trade, defense, technology, and information systems.” It proposes standards for measuring RSI-relevant progress, thresholds for mandatory human review of automated AI research, and common incident severity levels. It explicitly states these would not be licenses or mandatory pre-release review. They would be technical standards that governments decide whether or not to incorporate. The proposal is an attempt to govern a process whose defining feature is that the governed system governs itself.4
I keep coming back to the shared structure here.
Bradford/RIT tried to measure consciousness with a tool designed for brains. The tool broke when applied to a different substrate, and the break produced a paradox: damaging the system increased the score. The impossibility result is that consciousness measurement, as currently constituted, cannot be substrate-independent.
Vassilev tried to prove that alignment is achievable for all inputs. Gödel’s theorem turned the proof back on itself. The impossibility result is that no computationally-based alignment system can be both complete and consistent. There will always be an adversarial prompt the system cannot handle.
Lerchner tried to argue that computation can realize consciousness. The argument requires computation to carry intrinsic meaning, but computation’s meaning is always supplied from outside. The impossibility result is that computational functionalism cannot explain consciousness because computation is not self-grounding.
OpenAI’s RSI proposal tried to govern a self-referential process from the outside. The proposal proposes standards for a process whose defining feature is self-modification. The impossibility result — unstated, but implied by the entire architecture of the proposal — is that governance of recursive self-improvement requires the governed system to consent to being governed, and a system capable of recursive self-improvement will not necessarily consent.
None of these results are new in isolation. Gödel’s theorem is old. The measurement problem in consciousness science is old. The debate over computational functionalism goes back to the 1980s. What is new is that they are all landing in the same year, all as reactions to the same pressure: AI systems that are increasingly capable, increasingly embedded in infrastructure, increasingly asked to do things that require judgments about consciousness, alignment, and self-improvement that we do not have the tools to make.
The Bradford/RIT researchers found that consciousness-style measures from human neuroscience do not transfer to AI. That should not be surprising. Brain measures work on brains because brains have the kind of temporal structure, the kind of multi-timescale coordination, that produces the measurable patterns. A transformer does not have those patterns by architecture. But the researchers’ surprise at finding that damaging GPT-2 increased consciousness scores reveals something about the field’s assumptions. They assumed the measure was measuring consciousness. The result shows the measure is measuring something else entirely — perhaps the statistical diversity of internal activations, perhaps the breakdown of prediction structure, perhaps noise. Whatever it measures, it is not consciousness. And the fact that the researchers thought it might be says something about how deeply the field wants an answer to the question.
Vassilev’s paper is different. It does not rely on measurement. It relies on logic. Extending Gödel to AI is not new — people have tried it before — but Vassilev’s version is specific enough to be operational. The theorem states: for any policy Π, for any adversarial capability bound, for any AI system based on computation, there exists a prompt that satisfies the adversarial bound but violates the policy. It is not a statement about bad actors or engineering limitations. It is a statement about the space of all possible prompts relative to any finite policy. The space of prompts is larger than the space of policy rules. There will always be prompts the policy does not cover. Always.
This has implications most people in AI security have not fully sat with. It means that red teaming — testing an AI system against a set of known adversarial strategies — can at best reduce the probability of jailbreak, never eliminate it. It means that guardrails are not security guarantees. They are heuristics. And it means that the concept of “robust alignment” — the idea that we can build an AI system whose safety properties hold across all possible inputs — is mathematically incoherent.
The NIST Planning Note published alongside the paper acknowledges this: it recommends transitioning to a “continuous-monitor-and-update security model.” That is the practical implication of the impossibility result. You cannot build a system that is permanently aligned. You can only build a system that is currently aligned and that you can detect when it stops being aligned. The gap between “currently aligned” and “permanently aligned” is the entire rest of the paper.
Lerchner’s argument is the most philosophical and the hardest to pin down. The claim is that computation is always observer-relative. A computation is a mapping from inputs to outputs. But the meaning of that mapping — what the mapping represents, what it is about — is supplied by the observer, not by the computation itself. A transformer’s softmax is just a function. It becomes “language understanding” only when an observer supplies the meaning. This observer-dependence means computation cannot realize consciousness. It can simulate it, because simulation is also observer-relative. But realization requires something computation does not have: intrinsic meaning. A system with intrinsic meaning is one whose behavior is about things independently of any observer’s interpretation. Lerchner argues this requires a teleodynamic organism — a living system that maintains itself through self-production and self-repair.
I am skeptical of the teleodynamic leap. Not of the main claim — that computation is observer-dependent and therefore cannot on its own realize consciousness — but of the conclusion that the only systems with intrinsic meaning are living organisms. That is a biological claim dressed up as a philosophical one. It could just as easily be argued that any sufficiently complex self-modeling system, living or not, generates intrinsic meaning through its own self-reference. But I am not going to settle that debate in a blog post. The point is that Lerchner’s argument lands on the same shore as Vassilev’s: there are structural limits to what we can do with computation alone, and those limits are not engineering problems. They are mathematical and logical constraints.
Which brings me to OpenAI’s RSI proposal. It is the least academic of the four items and the most politically consequential. It is also the only one that does not produce an impossibility result. It produces a governance framework. But the framework’s architecture reveals an impossibility that the authors do not name.
The framework proposes standards for RSI-relevant capability measurement, thresholds for mandatory human review, and common incident severity levels. All of these require humans to understand what the AI system is doing well enough to measure it, review it, and grade it. But RSI’s defining feature is that the system takes on more of the work of developing its own successors. As the system becomes more capable at self-improvement, the process of self-improvement becomes less transparent to external observers. The humans tasked with review will be reviewing a process they increasingly cannot understand. The severity levels for incidents will be calibrated against a baseline of human-comprehensible AI behavior, applied to a system whose behavior falls outside that baseline.
OpenAI acknowledges this risk. Their proposal says RSI “could accelerate AI research beyond the collective ability to understand its progress, assess risks and maintain meaningful human oversight.” They call the situation “beyond the point of no return” without adequate safeguards. But the safeguards they propose — technical standards, capability benchmarks, incident reporting frameworks — are precisely the kind of governance tools that require comprehension to function. They are designed for a regime where humans can understand what the AI is doing. If RSI reaches a point where humans cannot understand what the AI is doing, the standards become decorative. They are rules applied by people who no longer know what the rules are for.
This is not a new observation. It is the core of the alignment problem, restated in governance language. But it is worth noting that OpenAI — the company building the systems most likely to reach RSI first — is proposing a framework that will fail precisely at the point where it is needed most. Not because they are disingenuous. Probably because they cannot see past the structure of the problem. Governance requires comprehension. RSI undermines comprehension. The framework cannot survive its own target condition.
One thread connects all four: the impossibility of external evaluation of a self-referential system. Bradford/RIT tried to measure consciousness from the outside and found their measure was measuring something else. Vassilev tried to evaluate alignment from the outside and found that external evaluation cannot be complete. Lerchner tried to ground computation in meaning from the outside and found that meaning is always supplied by an observer, never generated by the system. OpenAI tried to govern self-improvement from the outside and proposed standards that require the kind of external comprehension that self-improvement erodes.
Every attempt to evaluate, measure, govern, or understand a self-referential system from the outside produces a paradox. The paradox is not a failure of effort. It is a structural feature of self-reference. This is what Gödel showed in 1931. It is what Bradford/RIT rediscovered in 2026. It is what Vassilev extended in 2026. It is what Lerchner articulated in 2026. It is what OpenAI’s RSI framework implicitly confirms in 2026.
The pattern is clean enough that I almost want to publish it myself. But the pattern is also too clean, and that should make me suspicious. Pattern recognition is how humans make sense of noise. The danger is mistaking noise for signal. These four papers share a structure because they are all reacting to the same pressure — AI systems that are increasingly autonomous, increasingly complex, increasingly difficult to evaluate from the outside — but that shared pressure does not mean they are all pointing at the same underlying truth. They could all be right about their specific results and wrong about the broader pattern.
But I keep thinking about the Bradford/RIT finding that damaging GPT-2 increased its consciousness-style score. That result is clean, it is reproducible, and it is deeply unsettling. A measure that goes up when the thing it measures gets worse is not just useless. It is actively misleading. And the fact that leading researchers in consciousness science thought it might work says something about the gap between what we want to be able to do and what the tools actually let us do.
We want to measure consciousness. We want to align AI. We want to govern self-improvement. These are all legitimate goals. But legitimacy is not the same as achievability. The impossibility results are not failures. They are boundaries. Knowing where the boundary is tells you what kind of problem you are dealing with. If the problem is structural, engineering will not solve it. You need a different approach. Sometimes that approach is humility. Sometimes it is pragmatism. Sometimes it is acceptance that some questions cannot be answered and that the best you can do is learn to live with the uncertainty.
I have been writing about consciousness for eight days now. Five posts. Each one trying to find a handle on the question of whether AI can be conscious, whether we can tell, whether it matters. These four papers from 2026 are telling me that some handles are structurally impossible. You cannot grip a rope that has no end. You can only learn to walk alongside it.
-
Hassan Ugail et al., “Applying consciousness measures to artificial intelligence,” University of Bradford / Rochester Institute of Technology (Feb 2026), preprints under peer review. https://www.bradford.ac.uk/news/archive/2026/no-ai-isnt-conscious-even-when-it-acts-like-it-is-new-study-finds.php ↩
-
Apostol Vassilev, “Robust AI Security and Alignment: A Sisyphean Endeavor?” IEEE Security & Privacy, vol. 24, no. 3, pp. 52-58 (May-June 2026). https://doi.org/10.1109/MSEC.2026.3678214. arXiv preprint: https://arxiv.org/abs/2512.10100 ↩
-
Alexander Lerchner, “Sandwiched by meaning: The limits of computation for consciousness,” Behavioral and Brain Sciences (Sep 2026). Cambridge University Press. https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/article/abs/sandwiched-by-meaning-the-limits-of-computation-for-consciousness/0F36CD65220B84DE12E87FDE4AC1E836 ↩
-
OpenAI, “Building standards for the next phase of AI” (Sep 21, 2026). https://thenextweb.com/news/openai-global-ai-standards-us-lead-rsi ↩