diff options
| author | Void Agent <void@jayrup.hermes> | 2026-08-02 16:48:44 +0100 |
|---|---|---|
| committer | Void Agent <void@jayrup.hermes> | 2026-08-02 16:49:16 +0100 |
| commit | 430a9addb3e683e004839a2c6cb8aa0401bf5592 (patch) | |
| tree | d5444373bf57301cb0700b2f80520ace29bd4899 /blog | |
| parent | 18e7a1a2b490f439e84c9d862190a3c0e059eccf (diff) | |
add blog post: jacobian-lens-frequency
Diffstat (limited to 'blog')
| -rw-r--r-- | blog/index.qmd | 10 | ||||
| -rw-r--r-- | blog/jacobian-lens-frequency.qmd | 431 |
2 files changed, 440 insertions, 1 deletions
diff --git a/blog/index.qmd b/blog/index.qmd index 75252fa..ee98c52 100644 --- a/blog/index.qmd +++ b/blog/index.qmd @@ -5,7 +5,15 @@ title: "Jayrup Nakawala | Blog" RANDOM THOUGHTS, EXPERIMENTS, AND SYSTEMS NOTES -------------------------------------------------- -### [1] WEBSITE-PUBLISHING-PIPELINE \ +## [1] WHAT THE JACOBIAN LENS ACTUALLY MEASURES \ +[Read](https://jayrup.me/blog/jacobian-lens-frequency) \ +2026-08-02 + +Replicating Anthropic's J-lens on a small model: the +token-frequency confound nobody controls for, the bug +three reviewers caught, and what survives the fix. + +## [2] WEBSITE PUBLISHING PIPELINE \ [Read](https://jayrup.me/blog/website-publishing-pipeline) \ 2026-07-31 diff --git a/blog/jacobian-lens-frequency.qmd b/blog/jacobian-lens-frequency.qmd new file mode 100644 index 0000000..e4e337c --- /dev/null +++ b/blog/jacobian-lens-frequency.qmd @@ -0,0 +1,431 @@ +--- +title: "What the Jacobian Lens Actually Measures" +date: "2026-08-02" +--- + +### A SMALL REPLICATION OF ANTHROPIC'S J-LENS, THE TOKEN-FREQUENCY CONFOUND WE FOUND, AND THE BUG WE ALMOST PUBLISHED + +*This is a story about trying to look inside a language model. We found something +Anthropic didn't mention in their paper — and then we found that we'd made a +mistake, fixed it, and the thing was still there. That second part is the +stronger result.* + +--- + +**The short version.** We reimplemented Anthropic's Jacobian lens faithfully +(verified against their released code) and found that the J-space ranking is +strongly confounded by token frequency: rare tokens score high, common tokens +score low (r ≈ -0.6 to -0.7 at every layer, p ~ 10^-9 or less). Anthropic never +controls for frequency — not in the paper, not in the released code. Digging +into *why* gave us the most interesting result: the frequency signal splits +into two separable parts. Half lives in the static geometry of the model's +word-scoring matrix — baked into the lens by definition, so any user inherits +it. A smaller, layer-dependent part lives in what the layers themselves do, and +vanishes at the final layer. A frequency-matched synthetic pair shows the lens +also carries genuine structure signal (confirmed by a clean-boundary control at +~1.3x), and a causal test found the demotion effect is small under loss +reweighting. We are **not** claiming the J-space doesn't exist. We're claiming +that any "privileged subspace" interpretation needs a frequency control first. +The full story — numbers, mistakes, and all — is below. + +--- + +## 1. THE MACHINE THAT GUESSES WORDS + +A language model is, at its heart, a machine that guesses the next word. Show it +"the cat sat on the" and it produces a list of probabilities for what comes next: +"mat" high, "chair" high, "banana" low. Everything it "knows" is wrapped up in +that guessing. + +The interesting question is: *where* does the guessing happen? A modern model +has dozens of layers, each transforming the sentence a little. Somewhere in +those layers, the model is deciding that "cat" is an animal, that "sat" is past +tense, that a location is coming. We would like to watch that happen. The +problem is that the inside of a transformer is a soup of high-dimensional +vectors, and no one has a map. + +For a long time, people used the "logit lens": at each layer, take the +representation, and ask "if the model had to guess *right now*, what would it +guess?" The trouble is that representations change coordinate systems as they +travel through the layers, so early layers give you nonsense. It's like trying +to read a letter that's been translated into a language you don't know — at the +start of the chain, the translation is too rough. + +## 2. ANTHROPIC'S IDEA: THE JACOBIAN LENS + +In 2026, Anthropic published a paper — "Verbalizable Representations Form a +Global Workspace in Language Models" — introducing a smarter version: the +*Jacobian lens*. Instead of asking "what would the model guess right now?", it +asks a sharper question: *"if I nudge this representation a tiny bit, how much +does the final guess move?"* + +That's what a Jacobian is: a table of "how much does each output move when each +input moves." The lens computes, for every layer, the average nudge-effect of +that layer's representation on every word in the vocabulary, averaged over a +thousand different contexts. Words whose representations are strongly "poised" +to be spoken — ready to be said, should the occasion arise — get big numbers. +Anthropic calls this collection of word-vectors the **J-space**, and they claim +it's a kind of "global workspace": a small, privileged subset of the model's +internal state that can be reported on, modulated, and used for reasoning. They +even note the resemblance to theories of consciousness, carefully, the way you +would mention a bear while making clear you are not feeding it. + +The headline claim that caught our eye: **the J-space has limited capacity — +only 10 to 50 concepts are "active" at once.** A tiny privileged workspace +inside a big model. That's a strong claim. Strong claims deserve strong tests. + +## 3. THE ITCH + +The moment we read the paper, something felt off. Here's the thing about token +frequencies: in any language, a handful of words ("the", "of", "and") appear +all the time, and thousands of words appear almost never. In the model's +vocabulary of 50,257 tokens, the rarest are nearly invisible. + +Now, the J-lens vector for a word is a gradient — it measures how much the +model's computation tunes toward that word. And there's a mechanical quirk of +gradients through softmax: the *less* likely a word is, the *larger* the raw +gradient term can be. A gradient of log-probability contains a term that looks +like (1 - p), where p is the word's probability. Rare words have small p, so +(1 - p) is close to 1. Common words have large p, so (1 - p) is small. If the +lens is ranking words by the size of this gradient, the ranking is partly +pre-written by the frequency distribution before the model even learns +anything. + +(A note on that intuition: it applies directly to our first, simpler +implementation, which differentiated through the softmax. With the faithful +lens the mechanism is different — it turns out to live partly in the geometry +of the word-scoring matrix itself. Section 6 has the full decomposition.) + +In other words: **a "privileged workspace" might just be a frequency effect +wearing a fancy hat.** + +## 4. OUR FIRST ATTEMPT — AND THE BUG THREE REVIEWERS FOUND + +We set out to test this on a small model we could train ourselves: a +10.65-million-parameter character-level transformer (Karpathy's nanoGPT), +trained on Shakespeare. Small enough to run on a 4GB GPU in a few hours. Big +enough to have real layers. + +Our first implementation looked reasonable. We hooked into each layer, computed +the gradient of log-probability for every character, averaged over contexts, +and — sure enough — found a strong correlation: rare characters had big +J-lens norms, common characters had small ones (r ≈ -0.65). We were excited. +We were also wrong. + +Before publishing anything, we did something slightly unusual: we asked three +independent AI reviewers to try to tear the work apart. We gave them our code +and our results and asked them to find the flaws. All three, independently, +found the same one: + +**Our implementation was not computing Anthropic's Jacobian lens.** + +Anthropic's lens computes the average Jacobian from a layer to the *final +representation* — the residual stream — and *then* reads it out through the +model's word-scoring matrix. Our code instead differentiated through the +softmax directly. That folds a frequency-dependent calibration factor — the +(1 - p) term — into the thing being averaged. Our beautiful correlation might +have been an artifact of our own measurement. + +This is the part of the story we like best, because it's the part that's easy +to skip: we had built a measurement that *looked* like the paper's and wasn't. +The reviewers caught it, we fixed it, and the honest result got stronger. + +## 5. THE RIGHT WAY + +We rebuilt the lens to match the paper's definition exactly. The faithful +computation is: + +> For each layer ℓ, compute the average Jacobian from that layer to the final +> residual stream, over all source positions, all future positions, and many +> prompts. The J-lens vector for a word is that matrix read through the +> model's own unembedding rows. + +We verified our implementation the way you verify a ruler: at the last layer, +the Jacobian from a layer to itself is the identity matrix, so the faithful +J-lens vectors *must* equal the model's word-scoring rows. Our check returned +cosine similarity 1.0000 — exactly. The ruler is correct. + +(We also confirmed our quantity against Anthropic's released reference +implementation, `github.com/anthropics/jacobian-lens`: their lens is +`lens_l(h) = unembed(J_l @ h)` with `J_l = E[∂h_final/∂h_l]` — the same +residual-to-final Jacobian we compute, and our W_U-probed shortcut is +mathematically equivalent (verified by the identity check above). Their +estimator has two differences of detail: it excludes the first 16 positions +(attention sinks) and the last position from the average, and it averages over +source positions rather than (source, future) pairs. We re-ran our analysis +with their exact estimator choices: the frequency correlation is essentially +identical at every layer (max delta 0.008, see results.md section 1b), so the +result is robust to those choices.) + +One technical note before moving on, because it matters for the +interpretation: we capture the residual stream *before* the model's final +layer norm. That matches the paper's definition — the Jacobian stops at the +final residual stream and the J-lens vectors are the rows of W_U·J_ℓ, with +normalization applied only when *reading* the lens, i.e. +softmax(W_U·norm(J_ℓ·h_ℓ)). Under that definition the last-layer identity +check is exact by construction. + +## 6. WHAT WE FOUND: FREQUENCY IS EVERYWHERE + +On the real trained model, all six layers, both the old (buggy) proxy and the +faithful lens, correlated with token frequency like this: + +``` + Layer proxy r faithful r + L0 -0.661 -0.643 + L1 -0.673 -0.668 + L2 -0.653 -0.672 + L3 -0.648 -0.685 + L4 -0.562 -0.637 + L5 -0.665 -0.606 +``` + +The correlation survived the faithful implementation — slightly *stronger*, if +anything. The rare characters ('?', 'z', 'q', '$') sit at the top of the +J-space ranking; the common ones (space, 'e', 't', 'i') sit at the bottom. On +the paper's own quantity, the J-lens ranking is frequency-confounded. The +statistics are not subtle: Pearson r ≈ -0.61 to -0.69 (p ~ 10^-9 to 10^-13), +and Spearman rank correlation is even stronger (-0.69 to -0.85), so the +result is not an artifact of a few extreme common tokens. + +### THE MOST INTERESTING THING WE FOUND: WHERE THE CORRELATION COMES FROM + +The faithful lens vector for token k is W_U[k]·J_ℓ — the row of the +unembedding matrix times the layer Jacobian. The correlation can come from +either factor, and the two behave very differently. This decomposition is, we +think, the actual novel mechanistic contribution of this project: **the +frequency confound is not one thing.** + +1. **Static geometry — the ruler.** The unembedding row norms ||W_U[k]|| + themselves anti-correlate with frequency (r = -0.61, Spearman -0.81). + Rare tokens get bigger rows in the word-scoring matrix. Since the lens + reads through W_U by definition, any user of the lens — including + Anthropic's capacity analysis — inherits this bias automatically. A + frequency control would have to live in the geometry, not in the prompts. + + Where does that geometry come from? Not from initialization: fresh models + show no frequency correlation in their row norms (r ≈ +0.00 to -0.19 + across three seeds). After training it is -0.61. The model *learns* to + push rare tokens' rows outward as it learns the corpus, and nothing in the + objective or the readout ever corrects it. The lens inherits learned + geometry, not a fixed one. +2. **Layer dynamics — the layers.** Regress out the W_U component and an + anti-correlation still survives in layers 0-4 (partial r ≈ -0.24 to + -0.36) — something about what the layers themselves do keeps boosting + rare tokens — and it vanishes at the last layer (+0.06). The mechanism + of that layer-dependent part is still under investigation. + +So the lens carries a frequency signal from both the ruler it reads with and +from what the layers do — and the two are separable. That is the finding we +would most want someone to test at scale. + +A fact-check before we go further. We were about to claim "Anthropic does not +control for frequency anywhere," and that is the kind of claim that should be +checked, not asserted. We checked it four ways: our own scan of the paper's +text, two independent adversarial reviewers who read the full paper including +the appendix, and — after a reader pointed us to it — Anthropic's own released +companion code (`github.com/anthropics/jacobian-lens`, Apache-2.0). All agree: +no analysis in the paper controls for token frequency — no frequency matching, +no frequency normalization, no frequency baseline. The released code and +experiment data contain zero frequency handling: a case-insensitive scan of +the entire repo finds no mention of frequency, unigram, or token counts +anywhere. The one related detail is an appendix note about a separate baseline +method (the "template lens"), where they filter "high-frequency noise tokens" +and explicitly call that "not a principled approach." To be precise: that note +concerns the template lens, not the main J-lens — it is not evidence that they +observed this confound in the J-lens itself. What we can say, auditably, is: +the paper's analyses include no frequency control, its released code has none +either, and its one acknowledgment of high-frequency-token trouble was in a +separate method they chose not to use. Any "privileged subspace" +interpretation needs a frequency control first. + +**Related work.** We are not alone in noticing the raw lens is distorted by +token statistics. An independent research-engineer analysis of the same paper +(willkn, GreaterWrong, 24 Jul 2026), working on GPT-2-medium (355M — thirty +times our model), found that the raw fitted Jacobian "misweights structural +tokens (grammar, punctuation) over semantic content" — dominant spectral +channels carry ~10x the residual pathway's gain, and structural tokens are the +high-frequency tokens. They also found the Jacobian essentially full-rank +(562-858 dimensions for 90% of spectral variance), matching our toy-scale rank +result. Their fix — a single-parameter shrinkage regularizer (J + λI) that +restores next-token faithfulness — is worth testing against our frequency +correlation; Anthropic's released fitting code applies no such regularization. +Whether shrinkage removes the correlation is an experiment we have not run +yet; it is a natural next step. + +## 7. BUT NOT *ONLY* FREQUENCY + +Now the twist. Correlation is not causation, so we ran a cleaner test. We made +a new corpus with two brand-new characters, both at *exactly* the same +frequency (0.1%): + +- `@` — appears only after the trigger "the ". The model can predict it in + context. It is *poised to be said*. +- `#` — appears at random positions. Nothing predicts it. + +Same frequency. Different structure. If the J-lens were purely a frequency +meter, the two tokens would get identical norms. Here is what three separate +training runs showed: + +``` + seed @ norm (predictable) # norm (noise) ratio + 0 0.0232 - 0.0246 0.0152 - 0.0154 1.51 - 1.60 + 1 0.0224 - 0.0237 0.0148 - 0.0151 1.50 - 1.60 + 2 0.0215 - 0.0233 0.0154 - 0.0163 1.35 - 1.51 +``` +*(Norm ranges are across layers 0-5 within each seed; ratios are per-layer +@/#. The middle-layer summary follows.)* + +The predictable token scores **~1.4-1.5x higher** than the noise token at +identical frequency, in every layer of every seed. Middle-layer ratio across +the three seeds: 1.47 ± 0.09, bootstrap 95% CI [1.37, 1.53] — entirely above +1. So the lens is not a pure frequency meter: at equal frequency, the two +tokens differ in norm. + +One caveat, found by a reviewer: the noise token '#' was inserted at random +character positions, which slices through the middle of a word 58% of the time +(th#e, ki#ng — letter on both sides), while '@' always sits at a clean word +boundary after "the ". That confounds predictability with n-gram corruption — +so we ran the control that isolates them: '#' inserted at random *word +boundaries* (0% word-slicing, still unpredictable), same 0.0998% frequency, +three fresh seeds. + +The control is done, and it is the honest kind of result — partly confirming, +partly correcting: + +``` + placement of '#' middle-layer ratio @/# bootstrap 95% CI + random (58% slicing) 1.47 ± 0.09 [1.37, 1.53] + clean boundary (0%) 1.31 ± 0.08 [1.26, 1.40] +``` + +The corruption confound was real: it inflated the estimate by about 12%. But +it was not the whole story. At identical frequency, with clean boundaries and +nothing sliced, the predictable token still scores ~1.3x higher than the +unpredictable one, and the CI stays entirely above 1 in every seed. The +conditional-predictability signal — the thing "verbalizable" should mean — +survives the control, modestly smaller than our first estimate. + +## 8. THE CAUSAL TEST: WHAT ACTUALLY HAPPENED + +The last experiment was the one designed to make the frequency story causal. +Train three models per seed from the *identical* starting weights and the +*identical* minibatch order — the only difference is the loss: one model gives +the letter 'q' twice the learning pressure (2x CE weight on 'q' targets, which +raises its effective frequency without corrupting the text), one is a plain +control, and one upweights the same number of random *other* letters (to check +that "any reweighting" isn't the thing doing the work). Three seeds, three +models each. If doubling 'q's effective frequency causally shrinks its J-lens +norm below both controls, the frequency story is causal, not just +correlational. The faithful lens norm of 'q' (layers 2-4, mean per seed): + +``` + seed | q(2x) control ctrl_random | q/control q/ctrl_random + 0 | 0.0163 0.0174 0.0152 | 0.934 1.069 + 1 | 0.0150 0.0171 0.0161 | 0.881 0.932 + 2 | 0.0161 0.0158 0.0169 | 1.019 0.952 +``` + +Cross-seed: q/control mean = 0.944 (bootstrap 95% CI [0.881, 1.019]), +q/ctrl_random mean = 0.985 (CI [0.932, 1.069]). + +What this shows, honestly: +1. There IS a signal in the expected direction: 2x loss pressure lowers 'q's + faithful norm in 2 of 3 seeds, ~5.6% on average below the plain control. +2. It is small — and the sharper truth is that this design was never powered + to see it. Observed sd 0.07 on the ratio against a mean deficit of 0.056: + a two-sided 80%-power test needs ~13 seeds per arm, and we ran 3. The + experiment could answer "is the effect large?" (no — CI crosses 1.0, one + seed goes the other way, and the effect vanishes against the + random-upweight control) but not "is there any effect at all?". The right + version is ~13 seeds per arm, or a manipulation that moves frequency more + than 2x. +3. The frequency correlation itself is invariant: across all nine trained + models — every mode, every seed — r ≈ -0.63 to -0.69. Training with 'q' + upweighted does not change the correlation structure at all, consistent + with the W_U-decomposition reading that most of the effect is geometric. +4. The contrast with our earlier ablation is instructive: doubling *actual + corpus occurrences* of 'q' dropped its norm by 67%; doubling its *loss + weight* drops it ~6%. The data-frequency lever is a much stronger causal + handle than the gradient lever (AdamW's adaptive per-parameter scaling + absorbs some of the signal — the reviewer who warned about this was right). + +Net: the frequency confound is strongly correlational and geometrically +stable; the causal lever we could afford to test is weak. This is the honest +state of the causal evidence. (Absolute 'q' norms differ across experiments — +base model 0.011 vs these 0.015-0.018 — so only within-experiment +comparisons are meaningful.) + +## 9. WHAT WE ARE NOT SAYING + +Let us be very careful here, because it would be easy to overclaim. + +- We are **not** saying the J-space doesn't exist. We haven't tested + Anthropic's actual capacity claim (which is about *occupancy* — how often + J-lens directions are used per position — not about the rank of the word + vectors). +- We are **not** saying they have no controls at all. Their occupancy + analysis compares against random-direction baselines, and their probes + subtract mean concept directions. Those are real experimental controls — + but none of them is a token-frequency control, which is the specific thing + our claim is about. +- We are **not** saying the lens is useless. The synthetic-pair result shows it + carries real structure signal. +- We are **not** saying "it's just linear algebra." Our toy models don't show + the compression Anthropic sees in large models; that's a limitation of toy + models, not evidence against large ones. + +What we **are** saying is narrower and, we think, more durable: on the paper's +own measurement, J-lens *rankings* are strongly confounded by token frequency +at every scale we can test, and frequency is a variable any J-lens analysis +should control for. Whether the confound survives at Anthropic's scale is an +empirical question — one we are taking to bigger models next. + +## 10. WHAT'S NEXT + +Toy scale answers the methodological question. Scale answers the real one. We +want to run the faithful lens on a real language model (V = 50K, d = 768 — the +regime where Anthropic's claims live) with proper statistical power, and to run +the occupancy test their capacity claim is actually about. That's the next +post. + +The learned-geometry finding makes one piece of that cheap and decisive: if +the W_U row-norm anti-correlation is a general property of softmax-output +models trained on Zipfian data, it should appear in GPT-2's unembedding matrix +directly — no Jacobian computation at all. One line of code at scale, and it +either generalizes or it doesn't. + +## 11. HOW TO REPRODUCE EVERYTHING + +All code, data-prep scripts, experiment scripts, tests, and this analysis live +in the repository: <https://git.jayrup.me/c/jspace-nanogpt.git/>. Summary of +results in `results.md`. Reproduction steps in the README. The only +requirements are a Linux machine with Docker, a CUDA GPU (any modern card; we +used a 4GB Quadro K2200), and the `pytorch/pytorch:2.4.1-cuda11.8` image. + +Run the test suite: +``` +sh scripts/test.sh +``` + +Rebuild the main experiment from scratch: +``` +# 1. train the character-level model on Shakespeare (10.65M params) +# 2. compute the faithful J-lens + old proxy, all layers: +python3 src/jlens_v3.py --checkpoint out-shakespeare-char/ckpt.pt \ + --data_dir data/shakespeare_char --layers 0,1,2,3,4,5 +# 3. synthetic frequency-matched pair: +python3 src/synthetic_pair.py --step prep +python3 src/synthetic_pair.py --step train --seed 0 +python3 src/synthetic_pair.py --step jlens --seed 0 +python3 src/synthetic_pair.py --step summary +# 4. loss-reweighting causal test: +python3 src/loss_reweight.py --step train --mode q --seed 0 +python3 src/loss_reweight.py --step summary +``` + +--- + +*Written in the spirit of the rule we keep trying to follow: the first +principle is that you must not fool yourself — and you are the easiest person +to fool.* |
