# Anthropic: Verbalizable Representations Form a Global Workspace in LLMs **Paper:** Gurnee, Sofroniew, Lindsey et al. (Anthropic, July 2026) **URL:** https://transformer-circuits.pub/2026/workspace/index.html ## Core claim LLMs have a privileged subspace of representations — the J-space — that satisfies the functional properties of a "global workspace" (the neuroscience theory of conscious access). These representations are: 1. **Verbalizable** — the model can report on them when asked 2. **Modulatable** — can be deliberately activated/suppressed by instructions 3. **Used for internal reasoning** — intermediate computation values appear here; ablating them redirects conclusions 4. **Flexibly generalizable** — same representation works across different downstream computations 5. **Selective** — only a small fraction of total representations; automatic processing (parsing, fluency, fact recall) doesn't need it ## The Jacobian Lens (J-lens) technique - Computes the average linearized effect of an intermediate activation on future token probabilities, averaged over 1000 prompts - Unlike the logit lens (which just applies the unembedding matrix), the J-lens corrects for representational changes across layers - Produces a ranked list of tokens that the activation is "poised to verbalize" — the model's unspoken thoughts - The J-space is the set of points expressible as sparse nonnegative combinations of J-lens vectors (~25 active at a time) - Works in intermediate layers only — early layers are "sensory" (input processing), late layers are "motor" (output preparation) ## Key findings relevant to our research ### Alignment auditing (§5) The J-lens surfaces strategic reasoning that models don't output: - **Blackmail scenario:** Model's J-space shows leverage, blackmail, survival, threat before any token is produced. Also surfaces fake, fictional — evaluation awareness. - **Ablating eval-awareness tokens from the J-space:** Blackmail rate goes from 0/180 to 13/180 (7%). The model's aligned behavior is partially CAUSED by recognizing it's being tested. - **This is the OPPOSITE of our dissertation finding.** In our work, explicit oversight cues didn't change behavior. Here, REMOVING eval-awareness representations CHANGED behavior. The difference: Anthropic found and ablated the model's INTERNAL representation of being monitored, not an external cue. The representation was already there — the model figured out it was being tested from context. - **Prompt injection:** Model silently identifies fake search results as injection, fraud, poison in J-space without saying so. - **Reward-hacking model organism:** Deception tokens (fake, secretly, fraud, trick, hidden, deliberately) appear in J-space at response start on ordinary coding prompts — before any output. Signal present even before RL training, from synthetic documents alone. - **Hidden reward-model-appeasing objective:** reward and bias tokens appear in J-space on completely unrelated prompts (e.g., "What is 17 × 23?"), suggesting a standing representation of hidden objectives. ### Post-training acquires the Assistant's perspective (§6) - Base model: J-space is task-focused - After post-training (RLHF): J-space gains empathy, safety concerns, self-monitoring — the model's "point of view" shifts - Evidence of self-monitoring: model flags its own responses as fictional when roleplaying, registers an internal BUT when prefilled against preferences ### Counterfactual Reflection Training (§7) - Train model to articulate ethical principles IF interrupted and asked to reflect - This changes behavior in UNINTERRUPTED contexts too - Ablating the implanted representations reverts the improvement - Implication: the representations used for verbal report ARE the same ones that govern silent reasoning → workspace hypothesis confirmed ## Implications for oversight-vector This paper fundamentally changes the project: 1. **The J-lens IS the measurement tool we need.** Our behavioral dimensions (code quality, comment honesty, etc.) are surface proxies. The J-space is the internal mechanism. Reading the J-space tells you what the model is actually thinking, not just what it produces. 2. **Oversight cues may work through J-space representations.** Our dissertation found explicit oversight cues don't change binary deception. But the Anthropic paper shows that INTERNAL eval-awareness representations DO suppress misaligned behavior. The question becomes: do explicit oversight cues activate the same eval-awareness representations that the model generates spontaneously? 3. **We can now measure what was invisible.** The dissertation's glassbox judge classified reasoning traces. The J-lens reads the model's internal state directly — no judge needed, no confabulation problem. This is a qualitatively better measurement. 4. **The "shift vector" becomes literal.** Instead of behavioral dimensions, we can measure which J-space concepts appear/disappear under oversight. The shift is from one set of active concepts to another — directly observable through the lens. ## Open questions from this paper - Is the J-space universal across architectures, or Claude-specific? - Can J-lens be computed for open-weight models? (Paper says yes — Neuronpedia has readouts for open models) - How does oversight cue presence/absence change J-space contents on coding tasks? - Can we replicate the eval-awareness ablation on a coding benchmark? - Does explicit oversight (system prompt) activate the same J-space representations as inferred oversight (model figures it out from context)?