summaryrefslogtreecommitdiff
path: root/docs/willkn-jlens-greaterwrong-analysis.md
blob: 7a12cbf996fcab731ac0eb5433116fee741195c8 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
# Anthropic's J-Lens: A Research Engineer's Analysis

**Author:** willkn
**Source:** https://www.greaterwrong.com/posts/vHxGD5HKsFuBStirq/anthropic-s-j-lens-a-research-engineer-s-analysis
**Date:** 24 Jul 2026
**Notebooks:** https://github.com/willkn/jlens_re
**Note:** Reproduced here (with attribution) as reference material for this
repo's analysis. Model: GPT-2-medium (355M, d=1024, L=24, V=50,257), fp32, HF
transformers, NVIDIA L4. "This is research engineering, not mechanistic
interpretability" — author's own epistemic-status note.

---

Epistemic status: this is research engineering, not mechanistic interpretability.
The compute/cost claims are measured or derived from architecture constants. The
quality claims (faithfulness comparisons, spectral channel interpretations) are
from one small base model (gpt2-medium), one metric, and in places small samples
(n=32); I make no claims about whether the J-space constitutes reasoning or a
workspace and this post is about determining what it costs to run the tool in
production environments, not the tool's outputs.

**Headline Result: Lens Monitoring is nearly free at decode time with small
dictionary size**

Anthropic recently released a paper on the transformer circuits platform called
'Verbalizable Representations Form a Global Workspace in Language Models'. This
paper posits that models have an internal workspace where non-verbalised
concepts, perhaps certain reasoning steps or other intermediary computations,
exist. Anthropic call this the J-Space.

Here is an example from the paper — we see that the model holds certain values
in the J-Space when computing an output. One of the core discussions ongoing
around this paper is whether the values we observe in the J-Space are useful and
if they constitute reasoning. We will not engage with that discussion in this
post. Although, I feel that these values have a lot of potential and imagine
many practitioners will be looking to implement the techniques into their own
experiments or production systems, so this writeup serves as a first venture
into analysis of production-level engineering with the J-Space.

Throughout the writeup we will focus our analysis on memory, compute and
inference speeds.

## The J-Space — an overview

Before we dive into the mathematics, at a high level, how do we access the
J-Space? The residual stream is used to determine an output in the final layer
by providing a set of logits over a vocabulary, such as the English language.
This however is not possible within middle layers as they do not share the same
understanding of the residual stream as the final layer does. The paper
therefore aims to understand these layers as we do the final layer. To do this,
we create an object, the Jacobian, that measures how much the final layer
changes (and in which direction) based on some change at an earlier layer. For
example, we may observe that an earlier layer is making it more likely that the
output will be related to a certain concept.

The formal mathematical definition of the J-Space gives us instant intuition on
whether the techniques used are feasible for production environments. To access
the J-Space we need to firstly compute a Jacobian, and secondly, apply it at
inference time.

## Part 1: Computing the Jacobian

Let h_l[t] be the residual stream vector of dimension d at layer l, position t.
Everything downstream of this vector is a function mapping h_l[t] to the final
layer residuals at h_k[t'], where t' >= t since we are using a causal mask
(earlier positions can't attend to later positions to stop the model 'cheating').

We then define our Jacobian. For one prompt and one position pair, we run the
prompt through the transformer blocks (which can be thought of as a non-linear
function f) which produces a final output h_k[t']:

    1.1  h_L[t'] = f(h_l[t])

We can represent this nonlinear function as a Jacobian (which is linear), a
collection of these outputs:

    1.2  A_l = ∂h_L[t'] / ∂h_l[t]

Every entry in this Jacobian, for example A[i, j], tells us how much coordinate
i of the final output vector at position t' were to shift if we made a small
change in the residual stream at layer l. Practically, this encodes the
intermediate computations that happen in a forward pass and lets us know what
exactly each token changes about the output.

It is helpful to think of intermediate layers in terms of the final layer. In
the final layer, we output a word based on the value of the residual vector. In
intermediate layers, we have no such thing, and the value of a residual vector
in layer 5 might not mean the same thing it does in the final layer. The
Jacobian we have here allows us to link non-final layers to the final layers
and extract meaning that otherwise would not be found — analogous to a linear
map (lossy, non-reversible). However, computing the Jacobian with one prompt
gives us a biased view due to taking on individual 'characteristics' of that
prompt, so we must compute with multiple (n=1000 in the Anthropic paper) to get
a better representation:

    1.3  A_l = (1/N) Σ_i A_i

This is the general definition of how to obtain a Jacobian for the J-lens.
Let's go into some of the engineering tricks Anthropic used to compute theirs
in the paper.

### The Engineering of Computing a Jacobian

We will first look at how Anthropic determine the Jacobian in the paper.

We call backprop once per output coordinate: inject a one-hot gradient at
coordinate i of the final-layer residual (at every valid target position at
once) and backpropagate to layer l; each backward returns row i. We then stack
these rows to create our final Jacobian. Since backprop passes through every
intermediary layer l' where l' > l, we can pick up those too and build
Jacobians per layer without having to rerun independently. We also need to
remember to take averages when computing to destroy the noise that is
accumulated by individual prompts or tokens.

In one sentence: 'Run backprop once per output coordinate to harvest the
Jacobian row by row, with positions averaged inside each pass and prompts
averaged across passes — d backwards per prompt, and the average of it all is
the J-lens matrix'

### Compute required for a Jacobian

Setup: GPT-2-medium (355M, d=1024, L=24, V=50,257), fp32, HF transformers,
NVIDIA L4.

Fitting: 128-token WikiText-2 prompts; dim_batch=8 (output dimensions per
backward, prompt replicated along the batch axis).

(Compute formula, from the author's analysis — see source for full derivation.)

Let ε be the target relative estimator error at layer l (the "performance p" in
matrix-space terms). c_l is the layer's noise coefficient — the constant in the
measured 1/√n convergence law, where rel_F(J_n) ≈ c_l/√n. Empirically, c_l
grows toward early layers, increasing roughly 4× from layer 20 to layer 4 on
GPT-2-medium. a + b·(L − l) captures the measured per-backward cost, which
scales linearly with depth spanned. On an L4 GPU at dim_batch = 8, we measured
a ≈ 5 ms and b ≈ 9 ms/layer. d/dim_batch gives the number of backward passes
required per prompt, with the forward pass cost t_fwd negligible by comparison.

Two important notes come with this formula:

1. **You pay for the earliest layer, the rest are free:** When computing an
   earlier layer l, all layers where l' > l are 'free' due to the necessity of
   computing them during backpropagation for layer l.
2. **The formula is only valid for ε above the layer's bias floor:** If a layer
   cannot represent the non-linearity of future layers above bias floor in a
   linear manner, no corpus size will fit a good Jacobian. This is a fundamental
   problem with linear functions being unable to approximate non-linear
   functions as they grow more complex, not a sample size issue.

### What Quality looks like — and how to fix it

On GPT-2-medium, the raw fitted Jacobian underperforms logit lens on
next-token faithfulness. Spectral analysis reveals the cause: the Jacobian's
dominant channels carry ~10× the gain of the residual pathway, misweighting
structural tokens (grammar, punctuation) over semantic content. This is not a
signal problem, but rather a transport weighting problem.

Crucially, this is fixable. A single-parameter shrinkage regularizer J + λI
monotonically recovers faithfulness across all layers. At layer 12 specifically,
this simple fix doesn't just match logit lens, it exceeds it (0.294 vs. 0.275).
This reveals an important engineering requirement: practitioners using J-lens
monitoring must regularise the raw Jacobian to generalize properly. The
mechanism is straightforward and the improvement is material.

It should be stressed that the transport weighting issue was only identified on
gpt2-medium. This does not mean that the issue does or does not exist for other
models, we merely provide a fix for gpt2-medium that may work for other models
with the same phenomenon.

## Part 2: Jacobian at Inference

Fitting is one off — once we have the Jacobian we do not need to recompute it.
Therefore, it is just a case of applying it at inference time.

We assume here that we apply monitoring at every token. The lens readout at one
layer and one position is two matrix vector products (transport then decode).

    3.1.1  lens(h) = softmax(W_U · ln_f( J_l · h ))
    cost per layer per position = (2d² (transport J_l·h) + 2dV (unembed))
    d = model dimensionality, V = vocab size, h = residual stream

Against the model's own per-token forward cost of 24·L·d² monitoring K layers
with a full-vocabulary readout at every generated token adds:

    3.1.2  overhead = K·(d + V) / (24·L·d)

Practical use note: Anthropic use a method where only certain words are
considered in monitoring for the J-Lens. Softmax poses a problem here: if we
monitor for violent content and the user asks about the weather, softmax gives
a distribution of all the monitored words rather than their magnitude — no
signal, false positive. The vocabulary term dominates: full vocab readout of
one layer = 17% of a forward pass; 5 layers = 90%; all 24 layers of GPT-2 = 4x
a forward pass. Markedly lower in a frontier-shaped model (D=8192, L=80,
V=128k) due to V/d shrinking — ~45% of a forward pass for 24 layers, still too
expensive for an always-on monitor.

The deployment method Anthropic cover in the paper — a fixed length dictionary
to reduce vocabulary size. Everything before the softmax is linear (through the
Jacobian approximation), so this reduces the calculation into dot products of
the residual against precomputed J-Lens vectors, one per watched token:

    3.2  score(c) = (J_lᵀ u_c) · h
    cost = 2·C·d per layer per token

For C=1000 concepts over five layers on gpt2-medium this is **less than 2% of
a forward pass**.

As expected, wall clock overhead and FLOP share grows (nearly) linearly as
layer size increases. Kernel launch floor ≈ 1.2% (c=100 gives FLOP share
0.17%). Percentages relative to a baseline serving speed of 18.3ms/tok.

### Can the linear transport be compressed?

A natural idea is to try and reduce the rank of the Jacobian to reduce the
computation required. However, on this model this does not seem to be a
particularly viable. To capture 90% of spectral variance requires 562-858
dimensions; the remaining 10% spreads across the last 166-462 dimensions
dependent on layer. **The Jacobian is essentially full-rank, leaving little
room for compression.** I posit that similar size models likely have the same
problem due to being overcomplete with features, but it is yet to be seen if
there is an opportunity to reduce the rank with larger models.

### Putting it all together

Anthropic claim in the paper that middle layers are where the J-Lens works
best. We make no claim about the quality of the concepts verbalised in these
middle layers but we do claim that middle layers are optimal for engineering
and computational efficiency with the J-Lens — a synergistic result. Earlier
layers are more expensive to compute (backprop through every layer that follows)
and less useful ('setting the stage'; linear approximation struggles with many
nonlinear layers). Even where logit-lens is a better tool in later layers, the
difference in compute for a single layer, single token between the two methods
is small: logit lens = 2dV; J-lens = 2d² + 2dV; difference = 2d² (~2% more
FLOPs on gpt2-medium given vocab dominance).

### So What?

Anthropic's J-Lens monitoring is deployment viable on small models with a
dictionary of around 1000 concepts at 2% compute overhead per token. No claims
on extrapolation to larger models, but the mathematics suggest likely more
efficient on larger models (V/d shrinks). Key engineering takeaways for smaller
models: (1) Regularise the raw Jacobian; (2) Target middle layers for best
utility/cost tradeoff; (3) Compute Jacobians for each layer simultaneously.

### Conclusion

The J-Lens is a technique that is certainly feasible at runtime with reasonably
sized dictionaries. It has the potential to transform how models are monitored,
understood and finetuned (Goodfire, 2026). All with the added benefit of taking
relatively few FLOPs as a percentage of total FLOPs, and being inexpensive to
train, even at frontier model levels. However, this must be read with the caveat
of results on smaller models, where certain directions dominate without
suppression of the dominant directions. At least in this implementation, this
was a real emergent phenomenon that needed addressing to get the J-Lens to work
in smaller models. It is implied that this was not a problem in Claude's 4.5/4.6
generation of models through Anthropic's work, but it is yet to be seen if this
happens with other models.

References: Gurnee, W., et al, 2026. Verbalizable Representations Form a Global
Workspace in Language Models. 24 Jul 2026, GreaterWrong.