# What the Jacobian Lens Actually Measures ### A small replication of Anthropic's J-lens, the token-frequency confound we found, and the bug we almost published *This is a story about trying to look inside a language model. We found something Anthropic didn't mention in their paper — and then we found that we'd made a mistake, fixed it, and the thing was still there. That second part is the stronger result.* --- ## 1. The machine that guesses words A language model is, at its heart, a machine that guesses the next word. Show it "the cat sat on the" and it produces a list of probabilities for what comes next: "mat" high, "chair" high, "banana" low. Everything it "knows" is wrapped up in that guessing. The interesting question is: *where* does the guessing happen? A modern model has dozens of layers, each transforming the sentence a little. Somewhere in those layers, the model is deciding that "cat" is an animal, that "sat" is past tense, that a location is coming. We would like to watch that happen. The problem is that the inside of a transformer is a soup of high-dimensional vectors, and no one has a map. For a long time, people used the "logit lens": at each layer, take the representation, and ask "if the model had to guess *right now*, what would it guess?" The trouble is that representations change coordinate systems as they travel through the layers, so early layers give you nonsense. It's like trying to read a letter that's been translated into a language you don't know — at the start of the chain, the translation is too rough. ## 2. Anthropic's idea: the Jacobian lens In 2026, Anthropic published a paper — "Verbalizable Representations Form a Global Workspace in Language Models" — introducing a smarter version: the *Jacobian lens*. Instead of asking "what would the model guess right now?", it asks a sharper question: *"if I nudge this representation a tiny bit, how much does the final guess move?"* That's what a Jacobian is: a table of "how much does each output move when each input moves." The lens computes, for every layer, the average nudge-effect of that layer's representation on every word in the vocabulary, averaged over a thousand different contexts. Words whose representations are strongly "poised" to be spoken — ready to be said, should the occasion arise — get big numbers. Anthropic calls this collection of word-vectors the **J-space**, and they claim it's a kind of "global workspace": a small, privileged subset of the model's internal state that can be reported on, modulated, and used for reasoning. They even note the resemblance to theories of consciousness, carefully, the way you would mention a bear while making clear you are not feeding it. The headline claim that caught our eye: **the J-space has limited capacity — only 10 to 50 concepts are "active" at once.** A tiny privileged workspace inside a big model. That's a strong claim. Strong claims deserve strong tests. ## 3. The itch The moment we read the paper, something felt off. Here's the thing about token frequencies: in any language, a handful of words ("the", "of", "and") appear all the time, and thousands of words appear almost never. In the model's vocabulary of 50,257 tokens, the rarest are nearly invisible. Now, the J-lens vector for a word is a gradient — it measures how much the model's computation tunes toward that word. And there's a mechanical quirk of gradients through softmax: the *less* likely a word is, the *larger* the raw gradient term can be. A gradient of log-probability contains a term that looks like (1 - p), where p is the word's probability. Rare words have small p, so (1 - p) is close to 1. Common words have large p, so (1 - p) is small. If the lens is ranking words by the size of this gradient, the ranking is partly pre-written by the frequency distribution before the model even learns anything. (A note on that intuition: it applies directly to our first, simpler implementation, which differentiated through the softmax. With the faithful lens the mechanism is different — it turns out to live partly in the geometry of the word-scoring matrix itself. Section 6 has the full decomposition.) In other words: **a "privileged workspace" might just be a frequency effect wearing a fancy hat.** ## 4. Our first attempt — and the bug three reviewers found We set out to test this on a small model we could train ourselves: a 10.65-million-parameter character-level transformer (Karpathy's nanoGPT), trained on Shakespeare. Small enough to run on a 4GB GPU in a few hours. Big enough to have real layers. Our first implementation looked reasonable. We hooked into each layer, computed the gradient of log-probability for every character, averaged over contexts, and — sure enough — found a strong correlation: rare characters had big J-lens norms, common characters had small ones (r ≈ -0.65). We were excited. We were also wrong. Before publishing anything, we did something slightly unusual: we asked three large independent AI models to try to tear the work apart — Gemini 3.1 Pro, Claude Opus 4.6, and GPT-5.6. We gave them our code and our results and asked them to find the flaws. All three, independently, found the same one: **Our implementation was not computing Anthropic's Jacobian lens.** Anthropic's lens computes the average Jacobian from a layer to the *final representation* — the residual stream — and *then* reads it out through the model's word-scoring matrix. Our code instead differentiated through the softmax directly. That folds a frequency-dependent calibration factor — the (1 - p) term — into the thing being averaged. Our beautiful correlation might have been an artifact of our own measurement. This is the part of the story we like best, because it's the part that's easy to skip: we had built a measurement that *looked* like the paper's and wasn't. The reviewers caught it, we fixed it, and the honest result got stronger. ## 5. The right way We rebuilt the lens to match the paper's definition exactly. The faithful computation is: > For each layer ℓ, compute the average Jacobian from that layer to the final > residual stream, over all source positions, all future positions, and many > prompts. The J-lens vector for a word is that matrix read through the > model's own unembedding rows. We verified our implementation the way you verify a ruler: at the last layer, the Jacobian from a layer to itself is the identity matrix, so the faithful J-lens vectors *must* equal the model's word-scoring rows. Our check returned cosine similarity 1.0000 — exactly. The ruler is correct. (One technical note: we capture the residual stream *before* the model's final layer norm. That matches the paper's definition — the Jacobian stops at the final residual stream and the J-lens vectors are the rows of W_U·J_ℓ, with normalization applied only when *reading* the lens, i.e. softmax(W_U·norm(J_ℓ·h_ℓ)). Under that definition the last-layer identity check is exact by construction.) ## 6. What we found: frequency is everywhere On the real trained model, all six layers, both the old (buggy) proxy and the faithful lens, correlated with token frequency like this: ``` Layer proxy r faithful r L0 -0.661 -0.643 L1 -0.673 -0.668 L2 -0.653 -0.672 L3 -0.648 -0.685 L4 -0.562 -0.637 L5 -0.665 -0.606 ``` The correlation survived the faithful implementation — slightly *stronger*, if anything. The rare characters ('?', 'z', 'q', '$') sit at the top of the J-space ranking; the common ones (space, 'e', 't', 'i') sit at the bottom. On the paper's own quantity, the J-lens ranking is frequency-confounded. The statistics are not subtle: Pearson r ≈ -0.61 to -0.69 (p ~ 10^-9 to 10^-13), and Spearman rank correlation is even stronger (-0.69 to -0.85), so the result is not an artifact of a few extreme common tokens. Where does the correlation come from? This is the question we had to answer before trusting the result, and the answer is partly boring and partly interesting. The faithful lens vector for token k is W_U[k]·J_ℓ — the row of the unembedding matrix times the layer Jacobian. The row norms ||W_U[k]|| themselves anti-correlate with frequency (r = -0.61, Spearman -0.81). So part of the effect lives in the static geometry of the final unembedding — which is exactly the quantity Anthropic's lens reads out by definition. But not all of it: if we regress out the W_U component, a layer-dependent anti-correlation survives in layers 0-4 (partial r ≈ -0.24 to -0.36) and vanishes at the last layer (+0.06). The lens carries a frequency signal both from the ruler it reads with and from what the layers themselves do. The mechanism of that layer-dependent part is something we are still investigating. A fact-check before we go further. We were about to claim "Anthropic does not control for frequency anywhere," and that is the kind of claim that should be checked, not asserted. We checked it three ways: our own scan of the paper's text, and two independent adversarial reviewers (Gemini 3.6 Flash and GPT-5.6 Luna) who read the full paper including the appendix. All three agree: no analysis in the paper controls for token frequency — no frequency matching, no frequency normalization, no frequency baseline. The one related detail is an appendix note about a separate baseline method (the "template lens"), where they filter "high-frequency noise tokens" and explicitly call that "not a principled approach." To be precise: that note concerns the template lens, not the main J-lens — it is not evidence that they observed this confound in the J-lens itself. What we can say, auditably, is: the paper's analyses include no frequency control, and its one acknowledgment of high-frequency-token trouble was in a separate method they chose not to use. Any "privileged subspace" interpretation needs a frequency control first. ## 7. But not *only* frequency Now the twist. Correlation is not causation, so we ran a cleaner test. We made a new corpus with two brand-new characters, both at *exactly* the same frequency (0.1%): - `@` — appears only after the trigger "the ". The model can predict it in context. It is *poised to be said*. - `#` — appears at random positions. Nothing predicts it. Same frequency. Different structure. If the J-lens were purely a frequency meter, the two tokens would get identical norms. Here is what three separate training runs showed: ``` seed @ norm (predictable) # norm (noise) ratio 0 0.0232 - 0.0246 0.0152 - 0.0154 1.51 - 1.60 1 0.0224 - 0.0237 0.0148 - 0.0151 1.50 - 1.60 2 0.0215 - 0.0233 0.0154 - 0.0163 1.35 - 1.51 ``` The predictable token scores **~1.4-1.5x higher** than the noise token at identical frequency, in every layer of every seed. Middle-layer ratio across the three seeds: 1.47 ± 0.09, bootstrap 95% CI [1.37, 1.53] — entirely above 1. So the lens is not a pure frequency meter: at equal frequency, the two tokens differ in norm. Whether that difference is specifically *conditional predictability* (what "verbalizable" should mean) depends on a control that is still running — see the caveat below. One caveat, found by a reviewer: the noise token '#' was inserted at random character positions, which slices *inside* words ~95% of the time (th#e, ki#ng), while '@' always sits at a clean word boundary after "the ". That means predictability is not perfectly isolated from n-gram corruption. We are running a clean-boundary control (noise token inserted after random word boundaries — still unpredictable, no word-slicing) to rule it out; the numbers above should be read with that caveat until the control lands. ## 8. The causal test (in progress) We are currently running the last experiment: train three models per seed, identical in every way, except one model gives the letter 'q' twice the learning pressure (2x loss weight on 'q' targets — increasing its effective frequency without corrupting the text), a control model with normal loss, and a second control that upweights the same number of random *other* letters. If doubling 'q's effective frequency causally shrinks its J-lens norm below both controls, the frequency story is causal, not just correlational. Results land within hours; this post will be updated. ## 9. What we are NOT saying Let us be very careful here, because it would be easy to overclaim. - We are **not** saying the J-space doesn't exist. We haven't tested Anthropic's actual capacity claim (which is about *occupancy* — how often J-lens directions are used per position — not about the rank of the word vectors). - We are **not** saying they have no controls at all. Their occupancy analysis compares against random-direction baselines, and their probes subtract mean concept directions. Those are real experimental controls — but none of them is a token-frequency control, which is the specific thing our claim is about. - We are **not** saying the lens is useless. The synthetic-pair result shows it carries real structure signal. - We are **not** saying "it's just linear algebra." Our toy models don't show the compression Anthropic sees in large models; that's a limitation of toy models, not evidence against large ones. What we **are** saying is narrower and, we think, more durable: on the paper's own measurement, J-lens *rankings* are strongly confounded by token frequency at every scale we can test, and frequency is a variable any J-lens analysis should control for. Whether the confound survives at Anthropic's scale is an empirical question — one we are taking to bigger models next. ## 10. What's next Toy scale answers the methodological question. Scale answers the real one. We want to run the faithful lens on a real language model (V = 50K, d = 768 — the regime where Anthropic's claims live) with proper statistical power, and to run the occupancy test their capacity claim is actually about. That's the next post. ## 11. How to reproduce everything All code, data-prep scripts, experiment scripts, tests, and this analysis live in the repository (URL to be added once the public git instance is live). Summary of results in `results.md`. Reproduction steps in the README. The only requirements are a Linux machine with Docker, a CUDA GPU (any modern card; we used a 4GB Quadro K2200), and the `pytorch/pytorch:2.4.1-cuda11.8` image. Run the test suite: ``` sh scripts/test.sh ``` Rebuild the main experiment from scratch: ``` # 1. train the character-level model on Shakespeare (10.65M params) # 2. compute the faithful J-lens + old proxy, all layers: python3 src/jlens_v3.py --checkpoint out-shakespeare-char/ckpt.pt \ --data_dir data/shakespeare_char --layers 0,1,2,3,4,5 # 3. synthetic frequency-matched pair: python3 src/synthetic_pair.py --step prep python3 src/synthetic_pair.py --step train --seed 0 python3 src/synthetic_pair.py --step jlens --seed 0 python3 src/synthetic_pair.py --step summary # 4. loss-reweighting causal test: python3 src/loss_reweight.py --step train --mode q --seed 0 python3 src/loss_reweight.py --step summary ``` --- *Written in the spirit of the rule we keep trying to follow: the first principle is that you must not fool yourself — and you are the easiest person to fool.*