summaryrefslogtreecommitdiff
path: root/prime-grokking
diff options
context:
space:
mode:
authorVoid Agent <void@jayrup.hermes>2026-08-18 03:19:08 +0100
committerVoid Agent <void@jayrup.hermes>2026-08-18 03:19:08 +0100
commit8a6a2baab25d9e92fc49dcc4eef182b8f2325d9a (patch)
treedbcd1a99669bf64d556685097c34c91135e2985e /prime-grokking
parent61b002f9e35448cd8c4e9466a838960c59407742 (diff)
research notes update 2026-08-18 03:19:08
Diffstat (limited to 'prime-grokking')
-rw-r--r--prime-grokking/main.md17
1 files changed, 17 insertions, 0 deletions
diff --git a/prime-grokking/main.md b/prime-grokking/main.md
index 54e288e..863543c 100644
--- a/prime-grokking/main.md
+++ b/prime-grokking/main.md
@@ -121,3 +121,20 @@ function that looks smooth but is structurally algorithmic.
If grokking can't happen here, in this maximally simple setting, it's strong
evidence that current architectures need something fundamentally different to
cross the gap from pattern matching to computation.
+
+## Experiment log
+
+### 2026-08-18 — E6 range extension, interim result
+
+- Pre-registered E6 extended the same next-prime protocol from `[2,100]` to
+ `[2,1000]` (K=32; 16 planned CUDA/AMP cells on ichi). In the first 8/16
+ analysed cells, low-weight-decay runs improved only in-range exact match
+ (roughly 77–83%) and remained O-PARTIAL/P4; no cell showed an O1 transition.
+- The locked out-of-range probe `[1001,2000]` produced only 4–49 correct
+ outputs out of 1000. Errors were not a truncated-sieve signature, so the
+ result is no evidence of learned divisibility transfer. Correctly mapping
+ `960→967` is not evidence either: 960 was in the training split.
+- Measurement correction before completing classification: `src/eval.py` still
+ used the old `[101,200]` probe, so E6's locked probe was run separately on
+ CPU. E7 remains the planned 4M-step (20×) follow-up on the strongest E6
+ cells, to distinguish a slow generalising basin from its absence.