summaryrefslogtreecommitdiff
path: root/docs/blog-jlens-frequency.md
blob: d0f7ffa47acc8e5b2fcae68524e3e4b1e2e948cb (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
# What the Jacobian Lens Actually Measures
### A small replication of Anthropic's J-lens, the token-frequency confound we found, and the bug we almost published

*This is a story about trying to look inside a language model. We found something
Anthropic didn't mention in their paper — and then we found that we'd made a
mistake, fixed it, and the thing was still there. That second part is the
stronger result.*

---

## 1. The machine that guesses words

A language model is, at its heart, a machine that guesses the next word. Show it
"the cat sat on the" and it produces a list of probabilities for what comes next:
"mat" high, "chair" high, "banana" low. Everything it "knows" is wrapped up in
that guessing.

The interesting question is: *where* does the guessing happen? A modern model
has dozens of layers, each transforming the sentence a little. Somewhere in
those layers, the model is deciding that "cat" is an animal, that "sat" is past
tense, that a location is coming. We would like to watch that happen. The
problem is that the inside of a transformer is a soup of high-dimensional
vectors, and no one has a map.

For a long time, people used the "logit lens": at each layer, take the
representation, and ask "if the model had to guess *right now*, what would it
guess?" The trouble is that representations change coordinate systems as they
travel through the layers, so early layers give you nonsense. It's like trying
to read a letter that's been translated into a language you don't know — at the
start of the chain, the translation is too rough.

## 2. Anthropic's idea: the Jacobian lens

In 2026, Anthropic published a paper — "Verbalizable Representations Form a
Global Workspace in Language Models" — introducing a smarter version: the
*Jacobian lens*. Instead of asking "what would the model guess right now?", it
asks a sharper question: *"if I nudge this representation a tiny bit, how much
does the final guess move?"*

That's what a Jacobian is: a table of "how much does each output move when each
input moves." The lens computes, for every layer, the average nudge-effect of
that layer's representation on every word in the vocabulary, averaged over a
thousand different contexts. Words whose representations are strongly "poised"
to be spoken — ready to be said, should the occasion arise — get big numbers.
Anthropic calls this collection of word-vectors the **J-space**, and they claim
it's a kind of "global workspace": a small, privileged subset of the model's
internal state that can be reported on, modulated, and used for reasoning. They
even note the resemblance to theories of consciousness, carefully, the way you
would mention a bear while making clear you are not feeding it.

The headline claim that caught our eye: **the J-space has limited capacity —
only 10 to 50 concepts are "active" at once.** A tiny privileged workspace
inside a big model. That's a strong claim. Strong claims deserve strong tests.

## 3. The itch

The moment we read the paper, something felt off. Here's the thing about token
frequencies: in any language, a handful of words ("the", "of", "and") appear
all the time, and thousands of words appear almost never. In the model's
vocabulary of 50,257 tokens, the rarest are nearly invisible.

Now, the J-lens vector for a word is a gradient — it measures how much the
model's computation tunes toward that word. And there's a mechanical quirk of
gradients through softmax: the *less* likely a word is, the *larger* the raw
gradient term can be. A gradient of log-probability contains a term that looks
like (1 - p), where p is the word's probability. Rare words have small p, so
(1 - p) is close to 1. Common words have large p, so (1 - p) is small. If the
lens is ranking words by the size of this gradient, the ranking is partly
pre-written by the frequency distribution before the model even learns
anything.

In other words: **a "privileged workspace" might just be a frequency effect
wearing a fancy hat.**

## 4. Our first attempt — and the bug three reviewers found

We set out to test this on a small model we could train ourselves: a
10.65-million-parameter character-level transformer (Karpathy's nanoGPT),
trained on Shakespeare. Small enough to run on a 4GB GPU in a few hours. Big
enough to have real layers.

Our first implementation looked reasonable. We hooked into each layer, computed
the gradient of log-probability for every character, averaged over contexts,
and — sure enough — found a strong correlation: rare characters had big
J-lens norms, common characters had small ones (r ≈ -0.65). We were excited.
We were also wrong.

Before publishing anything, we did something slightly unusual: we asked three
large independent AI models to try to tear the work apart — Gemini 3.1 Pro,
Claude Opus 4.6, and GPT-5.6. We gave them our code and our results and asked
them to find the flaws. All three, independently, found the same one:

**Our implementation was not computing Anthropic's Jacobian lens.**

Anthropic's lens computes the average Jacobian from a layer to the *final
representation* — the residual stream — and *then* reads it out through the
model's word-scoring matrix. Our code instead differentiated through the
softmax directly. That folds a frequency-dependent calibration factor — the
(1 - p) term — into the thing being averaged. Our beautiful correlation might
have been an artifact of our own measurement.

This is the part of the story we like best, because it's the part that's easy
to skip: we had built a measurement that *looked* like the paper's and wasn't.
The reviewers caught it, we fixed it, and the honest result got stronger.

## 5. The right way

We rebuilt the lens to match the paper's definition exactly. The faithful
computation is:

> For each layer ℓ, compute the average Jacobian from that layer to the final
> residual stream, over all source positions, all future positions, and many
> prompts. The J-lens vector for a word is that matrix read through the
> model's own unembedding rows.

We verified our implementation the way you verify a ruler: at the last layer,
the Jacobian from a layer to itself is the identity matrix, so the faithful
J-lens vectors *must* equal the model's word-scoring rows. Our check returned
cosine similarity 1.0000 — exactly. The ruler is correct.

## 6. What we found: frequency is everywhere

On the real trained model, all six layers, both the old (buggy) proxy and the
faithful lens, correlated with token frequency like this:

```
  Layer   proxy r   faithful r
  L0      -0.661    -0.643
  L1      -0.673    -0.668
  L2      -0.653    -0.672
  L3      -0.648    -0.685
  L4      -0.562    -0.637
  L5      -0.665    -0.606
```

The correlation survived the faithful implementation — slightly *stronger*, if
anything. The rare characters ('?', 'z', 'q', '$') sit at the top of the
J-space ranking; the common ones (space, 'e', 't', 'i') sit at the bottom. On
the paper's own quantity, the J-lens ranking is frequency-confounded. Anthropic
does not control for this anywhere in their analysis.

## 7. But not *only* frequency

Now the twist. Correlation is not causation, so we ran a cleaner test. We made
a new corpus with two brand-new characters, both at *exactly* the same
frequency (0.1%):

- `@` — appears only after the trigger "the ". The model can predict it in
  context. It is *poised to be said*.
- `#` — appears at random positions. Nothing predicts it.

Same frequency. Different structure. If the J-lens were purely a frequency
meter, the two tokens would get identical norms. Here is what three separate
training runs showed:

```
  seed   @ norm (predictable)   # norm (noise)   ratio
  0      0.0232 - 0.0246        0.0152 - 0.0154   1.51 - 1.60
  1      0.0224 - 0.0237        0.0148 - 0.0151   1.50 - 1.60
  2      0.0215 - 0.0233        0.0154 - 0.0163   1.35 - 1.51
```

The predictable token scores **~1.4-1.5x higher** than the noise token at
identical frequency, in every layer of every seed. So the lens is not a pure
frequency meter. It genuinely responds to conditional predictability — which,
honestly, is what "verbalizable" should mean. The J-lens measures *both*:
a frequency prior that is never subtracted out, and a real structure signal on
top of it.

## 8. The causal test (in progress)

We are currently running the last experiment: train three models per seed,
identical in every way, except one model gives the letter 'q' twice the
learning pressure (2x loss weight on 'q' targets — increasing its effective
frequency without corrupting the text), a control model with normal loss, and a
second control that upweights the same number of random *other* letters. If
doubling 'q's effective frequency causally shrinks its J-lens norm below both
controls, the frequency story is causal, not just correlational. Results land
within hours; this post will be updated.

## 9. What we are NOT saying

Let us be very careful here, because it would be easy to overclaim.

- We are **not** saying the J-space doesn't exist. We haven't tested
  Anthropic's actual capacity claim (which is about *occupancy* — how often
  J-lens directions are used per position — not about the rank of the word
  vectors).
- We are **not** saying the lens is useless. The synthetic-pair result shows it
  carries real structure signal.
- We are **not** saying "it's just linear algebra." Our toy models don't show
  the compression Anthropic sees in large models; that's a limitation of toy
  models, not evidence against large ones.

What we **are** saying is narrower and, we think, more durable: on the paper's
own measurement, J-lens *rankings* are strongly confounded by token frequency,
and any claim about a privileged subspace must control for frequency first.
Anthropic's paper does not. The burden of proof is on them — and it's a fair
one.

## 10. What's next

Toy scale answers the methodological question. Scale answers the real one. We
want to run the faithful lens on a real language model (V = 50K, d = 768 — the
regime where Anthropic's claims live) with proper statistical power, and to run
the occupancy test their capacity claim is actually about. That's the next
post.

## 11. How to reproduce everything

All code, data-prep scripts, experiment scripts, tests, and this analysis live
in the repository: [link to cgit]. Summary of results in `results.md`.
Reproduction steps in the README. The only requirements are a Linux machine
with Docker, a CUDA GPU (any modern card; we used a 4GB Quadro K2200), and the
`pytorch/pytorch:2.4.1-cuda11.8` image.

Run the test suite:
```
sh scripts/test.sh
```

Rebuild the main experiment from scratch:
```
# 1. train the character-level model on Shakespeare (10.65M params)
# 2. compute the faithful J-lens + old proxy, all layers:
python3 src/jlens_v3.py --checkpoint out-shakespeare-char/ckpt.pt \
    --data_dir data/shakespeare_char --layers 0,1,2,3,4,5
# 3. synthetic frequency-matched pair:
python3 src/synthetic_pair.py --step prep
python3 src/synthetic_pair.py --step train --seed 0
python3 src/synthetic_pair.py --step jlens --seed 0
python3 src/synthetic_pair.py --step summary
# 4. loss-reweighting causal test:
python3 src/loss_reweight.py --step train --mode q --seed 0
python3 src/loss_reweight.py --step summary
```

---

*Written in the spirit of the rule we keep trying to follow: the first
principle is that you must not fool yourself — and you are the easiest person
to fool.*