From b3fe89895df19b94fdf837826f61c623f56fe4d7 Mon Sep 17 00:00:00 2001 From: Void Agent Date: Fri, 14 Aug 2026 13:14:08 +0100 Subject: prime-grokking: pre-launch addendum 1 (fully-tied cell, lambda ramp, wd sweep range, prior art) --- prime-grokking/preregistration.md | 26 ++++++++++++++++++++++++++ 1 file changed, 26 insertions(+) (limited to 'prime-grokking/preregistration.md') diff --git a/prime-grokking/preregistration.md b/prime-grokking/preregistration.md index d879912..f800461 100644 --- a/prime-grokking/preregistration.md +++ b/prime-grokking/preregistration.md @@ -66,3 +66,29 @@ Caveat: single seed — all architecture comparisons are seed-0 anecdotes until 2. No hyperparameter tuning on the val set. The wd sweep is a separate experiment run only after v1 results, at jayrup's call. 3. Probe interpretation locked above; NOTES.md must compare outcomes against this matrix verbatim (cite codes). 4. Training-range caveat recorded: for n ≤ 100 the sieve only needs divisors {2, 3, 5, 7}; "grokking the algorithm" in-range does not imply the general sieve. + +--- + +## Addendum 1 (2026-08-14, pre-launch — setup amendments only) + +Trigger: Gemini 3.6 Flash design review (experiment repo `design/reviews/gemini-design-review.md`). +The interpretation matrix (O/H/P codes) above is NOT amended; only setup details changed. Original lock commit: 00c696d. + +1. **Fully-tied cell (was: un-tied GRU decoder).** All recurrence — input read-in, K compute steps, AND output-digit decoding — now runs through the SAME 2-layer cell. The earlier draft's GRU decoder would have masked whether the tied cell solved the task. RNN param count ≈ 36.6k (was 168.7k). Regression test added (no GRU/LSTM/RNN modules). +2. **Recurrent input read-in (was: masked mean-pool).** Digits are read through the tied cell with sinusoidal positional encoding. The mean-pool blurred place value ("10" and "100" share the token multiset {1,0}). +3. **λ schedule: linear ramp 1000→5000 steps (was: hard switch at step 1000).** Avoids a discontinuous loss jump late in training. +4. **Future wd sweep revised to {0.01, 0.1, 0.3, 1.0, 3.0}** (10.0 dropped: at lr=1e-3 with AdamW, λ=10 decays weights ~1%/step). Seed-0 default wd=1.0 unchanged. +5. **ACT verification note:** aggregation is Graves (2016) standard — w_t = p_t·Π_{s