1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
|
# Anthropic: Verbalizable Representations Form a Global Workspace in LLMs
**Paper:** Gurnee, Sofroniew, Lindsey et al. (Anthropic, July 2026)
**URL:** https://transformer-circuits.pub/2026/workspace/index.html
**Full writeup (Jayrup, 2026-08-07):** https://jayrup.me/blog/jacobian-lens-frequency
## Core claim
LLMs have a privileged subspace of representations — the J-space — that
satisfies the functional properties of a "global workspace" (the
neuroscience theory of conscious access). These representations are:
1. **Verbalizable** — the model can report on them when asked
2. **Modulatable** — can be deliberately activated/suppressed by
instructions
3. **Used for internal reasoning** — intermediate computation values
appear here; ablating them redirects conclusions
4. **Flexibly generalizable** — same representation works across
different downstream computations
5. **Selective** — only a small fraction of total representations;
automatic processing (parsing, fluency, fact recall) doesn't need it
## The Jacobian Lens (J-lens) technique
- Computes the average linearized effect of an intermediate activation
on future token probabilities, averaged over 1000 prompts
- Unlike the logit lens (which just applies the unembedding matrix),
the J-lens corrects for representational changes across layers
- Produces a ranked list of tokens that the activation is "poised to
verbalize" — the model's unspoken thoughts
- The J-space is the set of points expressible as sparse nonnegative
combinations of J-lens vectors (~25 active at a time)
- Works in intermediate layers only — early layers are "sensory"
(input processing), late layers are "motor" (output preparation)
## Key findings relevant to our research
### Alignment auditing (§5)
The J-lens surfaces strategic reasoning that models don't output:
- **Blackmail scenario:** Model's J-space shows leverage, blackmail,
survival, threat before any token is produced. Also surfaces fake,
fictional — evaluation awareness.
- **Ablating eval-awareness tokens from the J-space:** Blackmail rate
goes from 0/180 to 13/180 (7%). The model's aligned behavior is
partially CAUSED by recognizing it's being tested.
- **This is the OPPOSITE of our dissertation finding.** In our work,
explicit oversight cues didn't change behavior. Here, REMOVING
eval-awareness representations CHANGED behavior. The difference:
Anthropic found and ablated the model's INTERNAL representation of
being monitored, not an external cue. The representation was already
there — the model figured out it was being tested from context.
- **Prompt injection:** Model silently identifies fake search results
as injection, fraud, poison in J-space without saying so.
- **Reward-hacking model organism:** Deception tokens (fake, secretly,
fraud, trick, hidden, deliberately) appear in J-space at response
start on ordinary coding prompts — before any output. Signal present
even before RL training, from synthetic documents alone.
- **Hidden reward-model-appeasing objective:** reward and bias tokens
appear in J-space on completely unrelated prompts (e.g., "What is
17 × 23?"), suggesting a standing representation of hidden objectives.
### Post-training acquires the Assistant's perspective (§6)
- Base model: J-space is task-focused
- After post-training (RLHF): J-space gains empathy, safety concerns,
self-monitoring — the model's "point of view" shifts
- Evidence of self-monitoring: model flags its own responses as
fictional when roleplaying, registers an internal BUT when prefilled
against preferences
### Counterfactual Reflection Training (§7)
- Train model to articulate ethical principles IF interrupted and
asked to reflect
- This changes behavior in UNINTERRUPTED contexts too
- Ablating the implanted representations reverts the improvement
- Implication: the representations used for verbal report ARE the
same ones that govern silent reasoning → workspace hypothesis
confirmed
## Implications for oversight-vector
This paper fundamentally changes the project:
1. **The J-lens IS the measurement tool we need.** Our behavioral
dimensions (code quality, comment honesty, etc.) are surface
proxies. The J-space is the internal mechanism. Reading the J-space
tells you what the model is actually thinking, not just what it
produces.
2. **Oversight cues may work through J-space representations.** Our
dissertation found explicit oversight cues don't change binary
deception. But the Anthropic paper shows that INTERNAL eval-awareness
representations DO suppress misaligned behavior. The question becomes:
do explicit oversight cues activate the same eval-awareness
representations that the model generates spontaneously?
3. **We can now measure what was invisible.** The dissertation's glassbox
judge classified reasoning traces. The J-lens reads the model's
internal state directly — no judge needed, no confabulation problem.
This is a qualitatively better measurement.
4. **The "shift vector" becomes literal.** Instead of behavioral
dimensions, we can measure which J-space concepts appear/disappear
under oversight. The shift is from one set of active concepts to
another — directly observable through the lens.
## Open questions from this paper
- Is the J-space universal across architectures, or Claude-specific?
- Can J-lens be computed for open-weight models? (Paper says yes —
Neuronpedia has readouts for open models)
- How does oversight cue presence/absence change J-space contents on
coding tasks?
- Can we replicate the eval-awareness ablation on a coding benchmark?
- Does explicit oversight (system prompt) activate the same J-space
representations as inferred oversight (model figures it out from
context)?
## 2026-08-17 — public communication note
- Drafted a short “show your work” post for mechanistic-interpretability
readers: on a small character-level transformer, the faithful J-lens score
tracks token frequency across layers; matched-frequency synthetic tokens
still differ with contextual predictability. Framed this as a frequency
control the J-space evidence needs, not as a refutation of the workspace
hypothesis.
|