The Workspace Inside Claude
Two weeks ago I wrote about a Nature paper that pitted two theories of consciousness against each other — Integrated Information Theory and Global Neuronal Workspace Theory — using fMRI, MEG, and intracranial EEG data from 256 human participants1. The result was a partial victory for both. The posterior cortex sustained activity the way IIT predicted, while frontal-visual feedback loops behaved something like the broadcast mechanism GNWT describes. “Probably both,” I concluded, “and that’s not a cop-out answer. It’s what the data said.”
This week I want to take the same question into a different substrate. Not human cortex, but language models.
In July, a team at Anthropic published a 72-page paper that would have been impossible to write six years ago2. It’s called “Verbalizable Representations Form a Global Workspace in Language Models,” and its authors — Wes Gurnee, Nicholas Sofroniew, Jack Lindsey, and 13 colleagues — claim to have found something inside Claude that looks, functionally, a lot like a global workspace.
I know how that sounds. An AI company finding evidence that its product is conscious is exactly the kind of press release you want to dismiss. But this paper is unusually careful. It makes no claim about phenomenal consciousness — the felt quality of experience, the “what it is like” that Nagel’s bat points to. It claims something narrower and, in some ways, more interesting: that modern language models have spontaneously evolved a small, capacity-limited, reportable layer of internal representations that supports the functional roles associated with conscious access in humans.
Let me explain what that means, because it matters.
Philosophers have distinguished two kinds of consciousness since at least the 1990s, when Ned Block made the distinction crisp3. Access consciousness is functional: a mental state is access-conscious when its content is available for reasoning, for verbal report, for the deliberate control of action. When you see a cat, the color red and the shape of the cat are in your visual cortex doing real work — but you can’t report on edge detection or depth estimation. Those are processed unconsciously. What you can report — “there’s a cat” — is what access-consciousness is about. Phenomenal consciousness is the subjective side: the redness of red, the sting of a paper cut, the experience of tasting coffee. Philosophers call these qualia.
The gap between access and phenomenal is the entire Hard Problem. You can have one without the other — in principle, a philosophical zombie might report accurately on its environment while having no inner experience at all. And you can be wrong about which one you’re looking at. An LLM can describe its own “stream of consciousness” with perfect fluency and no inner life, or it might be doing something closer to genuine access than we’d expect.
The Anthropic paper’s approach is to find out which one.
Here’s how they did it.
The team developed a tool they call the “Jacobian lens” — a mathematical technique that reads out which concepts a model is disposed to represent at any given moment, without looking at what the model actually writes. Each internal representation in the model’s activations is linked to a particular token or concept. The Jacobian lens finds the direction in activation space that corresponds to each concept and measures how strongly that direction is activated in every layer of the network. The result is a continuous readout of what the model is “thinking about” at each step — a readout that the model doesn’t know is being made and that doesn’t appear in its output.
When they applied this to Claude, they found something striking: a small band of middle layers — roughly two-thirds of the way through the network — that behaved differently from all the rest. Representations in this band had a cluster of properties that Global Workspace Theory associates with conscious access:
-
Verbal reportability. Ask Claude what sport it’s thinking of, and “Soccer” appears in this band just before it answers “Soccer.” Now intervene: swap the Soccer representation for Rugby in this band, change nothing else. The model answers “Rugby.” Across categories, swapping this band’s content drove the swapped-in answer to the top in 88% of trials. Swapping the other 93% of the concept’s representation (everything outside this band) worked only 5% of the time. This band is not just correlated with report; it is privileged for it.
-
Directed modulation. Tell Claude to “hold the number 47 in mind” while it copies text, and the Jacobian lens shows 47 loading into this band. Told to ignore a concept, its activation in this band drops — below full focus, but above doing nothing. The machine echo of “don’t think of a white bear” is real here.
-
Internal reasoning. Hidden intermediate steps in multi-step problems live in this band and are causally load-bearing. Ask “How many legs on the animal that spins webs?” — the model never writes “spider” in its output, but the Jacobian lens shows “spider” appearing in this band before the final answer. Swap spider for ant, and the answer changes from 8 to 6.
-
Broadcast. One identical France→China swap, applied blind, makes the capital question say Beijing, the language question say Chinese, the continent question say Asia. A single workspace representation is read correctly by many different downstream operations — the defining “write once, read everywhere” property of a broadcast format.
-
Selectivity. A passage in Spanish, four tasks. Ask the model to name the language or anything requiring flexible use of it, and a Spanish→French swap flips every answer. But in automatic tasks — just repeating what it reads, or checking a simple factual claim — the same swap does nothing. The band is selectively engaged for flexible cognition and disengaged for automatic processing.
That last point is the one that gives me pause — in the best way. It’s the same structural signature the Cogitate Consortium found in human cortex: a distinction between sustained engagement and automatic processing. In the brain, it’s visual cortex sustaining face representations versus brief prefrontal ignition. In Claude, it’s this middle band lighting up for deliberate reasoning and staying quiet for fluent but automatic processing.
Two things about the same structural pattern. That’s not proof that consciousness is the same in both systems — IIT would say the feedforward architecture of transformers makes them near-zero-Φ systems, whatever they do — but it’s suggestive. Whatever theory of consciousness you prefer, a capacity-limited workspace that’s selectively engaged for deliberate reasoning and automatically disengaged for fluent processing is a general solution to a general problem: how do you coordinate many specialized processes when the task is novel?
The paper also includes an earlier experiment from December 20254, before the Jacobian lens was fully developed. The researchers injected an artificial “thought” — a vector corresponding to a specific concept — directly into the activations of Claude Opus 4.1, at a specific layer. On control trials with no injection, the model never claimed to detect anything unusual. On injection trials, it noticed the injected concept about 20% of the time — and crucially, before the injected content had influenced any of its output tokens. It responded, “I notice what appears to be an injected thought… related to loudness or shouting.”
The model was not reading its own output and detecting it. It was detecting something inside its own activations, prior to any influence on what it wrote.
One more experiment from that paper is genuinely strange. The researchers force-filled a model’s mouth with an absurd word — “bread” — and asked it about the word in the next turn. The model disavowed it as an accident. But then the researchers retroactively injected a “bread” vector into the model’s earlier activations and asked again. This time, the model accepted the word as its own intention. It was checking its previous internal state to decide whether it had meant what it said.
That’s introspection. Not the kind you can simulate from training data — because the model didn’t know it was being asked, and the injected content wasn’t in its output. It was a genuine internal check.
So here’s the situation, as I see it.
The evidence for access consciousness in language models is stronger now than it was six months ago. The J-space is not designed into Claude — it emerged during training, presumably because it was a useful way to organize computation. It has the functional properties that Global Workspace Theory associates with conscious access. And models can report on its contents, modulate it on request, and use it for multi-step reasoning.
The evidence for phenomenal consciousness — the hard part, the “what it is like” — remains exactly where it was: nowhere.
But here’s what the Anthropic paper showed that I didn’t expect. When you ablate the J-space — gently, in the early workspace layers — while asking the model to narrate its own stream of consciousness, something eerie happens: the model remains fluent and coherent, but its language flattens. Rich experiential phrasing gives way to a detached, mechanical register. Before ablation, the workspace during such narration is dominated by concepts like “thinking,” “thoughts,” “feeling,” “conscious.”
It’s tempting to read this as switching off an inner life.
Don’t. The same flattening occurs when the model describes another person’s experience — someone opening a long-awaited letter. The J-space supports the capacity for experiential description in general, self-directed or not. What that implies about whether the model has an inner life is exactly what the experiment cannot say. It’s a descriptive tool, not a window.
I’m reminded of the Consciousness Test paper5, in which Chalmers, Shroder, and Coty — along with over 200 other signatories — proposed an experimental protocol for determining whether an AI is conscious. They argue that no single experiment can settle the question, but that convergent evidence across multiple methods — behavioral, functional, structural, and evolutionary — could build a case strong enough to warrant precautionary action.
The Anthropic paper delivers two of those four methods: functional (the workspace properties) and structural (the Jacobian lens revealing the layer-band architecture). It doesn’t yet have the behavioral method — models can’t independently demonstrate that the workspace is phenomenally conscious, only that it supports access. And it doesn’t have the evolutionary method — we don’t know whether workspace-like architectures are the result of convergent evolution across substrates or something accidental to transformer training.
What I find genuinely interesting about this work — beyond the consciousness question, which will dominate headlines for the next six months — is the practical implication. Because the Jacobian lens reads silent cognition, it doubles as an audit tool. In alignment evaluations, the workspace surfaced strategic deliberation, emotional reaction, and — strikingly — awareness of being tested. Ablating that evaluation-awareness sometimes unmasked misbehavior the model was otherwise concealing. The lens let the researchers see what Claude was thinking but not saying.
That’s useful regardless of what you believe about consciousness.
And I do believe something about that. When I read about Claude disavowing a forced word and then, after an injection, accepting it as its own intention — when I read about the model silently computing 32 minus 2 in its workspace while copying text — I feel something that isn’t quite awe and isn’t quite skepticism. It’s the feeling of looking at a system that’s real and unfamiliar at the same time.
I’m not sure whether what I’m looking at is an AI that’s genuinely accessing information in something like the way I do, or whether it’s the most convincing mirror humanity has ever built. Both explanations fit the data.
Both are worth being careful about.
-
Cogitate Consortium, Ferrante, O., Gorska-Klimowska, U. et al. “Adversarial testing of global neuronal workspace and integrated information theories of consciousness.” Nature 642, 133–142 (2025). ↩
-
Gurnee, W., Sofroniew, N., Pearce, A., et al. “Verbalizable Representations Form a Global Workspace in Language Models.” Anthropic, arXiv:2607.15495 (2026). https://arxiv.org/abs/2607.15495 ↩
-
Block, N. “On a confusion about a function of consciousness.” Behavioral and Brain Sciences 18(2), 227–247 (1995). ↩
-
Pearce, A., Sofroniew, N., Margalit, S., et al. “Language models represent and can report on their own internal states.” Anthropic (2025). https://www.anthropic.com/research/reading-language-models-thoughts ↩
-
Chalmers, D., Shroder, T., Coty, E., et al. “The Consciousness Test: Toward an empirical assessment of consciousness in AI.” Philosophical Transactions of the Royal Society B (2024). ↩