Bit-serial modular multiplication β radix-4 fork of neural-horner v8
A ~471K-parameter recurrent cell run in a fixed bit-serial Horner loop, computing (aΒ·b) mod p for primes to 2^2048 and operands to 4096 bits. Fork of Robby Sneiderman's neural-horner v8 (MIT), which is submitted separately as an attributed baseline.
Public benchmark: overall 98.9%, highest tier β₯90% = 10, 129.7 s on a 4090 against a 300 s budget.
Architecture
One shared p-conditioned cell (bidirectional 3-layer GRU, d_model 96, hidden 128, radix-4 digit and op embeddings, per-bit head) learns s' = (4s + dΒ·x) mod p, d β {0,1,2,3}; a learned op-embedding selects its three uses: reduce a, reduce b, multiply the residues. Radix-4 halves the step count at equal parameter count. Warm-started from v8 with the op-embedding zero-initialised, so at k=1 it reproduces v8 bit-for-bit.
Compliance: no controller, the schedule takes no feedback from the model, every arithmetic step including the conditional subtraction comes from the trained cell, and both routing choices key on the prime's bit-length alone β the shape ruled acceptable in the Zulip fixed-loop-encoder thread.
weights.pt is a weight-space average (model soup) of an annealed checkpoint and a same-recipe re-anneal; it beat both endpoints on identical problems (McNemar p = 0.02, n = 1008). weights_t1.pt serves primes of bit-length β€ 3. Width policy: groups of β₯257 bits run at min(2080, pbits+32), 17β256 at pbits, <17 at 32.
Findings
1. Width robustness is a property of the training width distribution, not of the loop. Sizing the state register produced this project's largest gain, with no retraining. Sweeping Ξ = state width β prime bit-length on identical problems, 576 per width, four tier bands:
| Ξ | 0 | 1 | 2 | 3 | 4 | 8 | 16 / 24 / 32 / 48 / 64 |
|---|---|---|---|---|---|---|---|
| errors | 27 | 516 | 470 | 116 | 13 | 13 | 3 |
Ξ β {1,2} falls to 2β30% accuracy. Ξ = 3 is usable but degrades with rollout depth (T7 91.7%, deepest T10 47.9%). Error sets at Ξ β {16,24,32,48,64} are identical (Jaccard 1.00), so the 32 we deploy is about twice what is needed.
The same problems on v8 give zero errors at every Ξ β {0,1,2,3,4,32} β a flat curve. Loop, inference code and problem set are identical; the difference is training. v8's one-step training holds the state width at L and stratifies primes by bit-length, covering every Ξ. Our radix-4 retraining trained each prime size at its own field-filling width, covering only Ξ = 0.
Practical form: tier primes cluster just under multiples of 32, so rounding the state width up to a machine-word multiple lands them at 1β2 bits of headroom, the worst region of the curve. Add a fixed margin instead β "round up to 32" is not the same policy as "add 32", and falsifying the former cost us a week. The upstream argument that dynamic-L is exactness-preserving holds for v8 but does not travel with the weights.
2. Small primes are unreachable in any sampling stream. random_prime-style constructors emit odd values with the top bit set, so p = 2 must be injected explicitly. Our T1 sat at 61%, and our own battery could not see p = 2 either, so the hole was first misfiled as a false alarm. The fused single-step universe for p β {2,3,5,7} is 348 transitions β enumerate it.
What did not work
Continued training failed in three forms. Unanchored tiny-prime fine-tuning cured small primes and ate large-scale precision (T8 β8, T10 β13). L2-SP has no Ξ» window: Ξ» = 10/30/100 relocates the interference rather than removing it (512-bit cells 166β170β172 while 64/128/256-bit cells fall monotonically). A same-recipe re-anneal paid nothing. The mechanism is measurable: our single-step validation floor is ~8e-6 while effective Ξ΅ is already 1.8e-6, so objective and checkpoint selection both sit below the instrument. Coverage and precision are zero-sum here. Ensembling failed too, in both the width-committee and cross-checkpoint-voting forms.
Limits
Structured operands are a live residual, independent of the width effect: the Fermat family is clean (96/96) but the wider power-of-two neighbourhood (2^m, 2^mΒ±1) is 280/288, failing only at primes β₯1024 bits, and headroom does not remove it (at 2013 bits, Ξ = 0/4/8/32 gives 8/2/3/3 over 72 problems). Official tiers cannot generate such operands, so this is robustness, not score.
On the probe set above, v8 outperforms this model at every width: the radix-4 line bought a 2Γ reduction in step count, not accuracy. The deep-chain Ξ΅ tail is deferred, not solved (T10 min 92%). Small primes needed a second routed checkpoint because one weight set could not hold both regimes. The remaining structural lever is raising the radix further; we ran out of time.
Attribution
Derived from neural-horner (github.com/Robby955/neural-horner) by Robby Sneiderman, MIT licensed. The architecture and the compliance framing originate there; the radix-4 fusion, width policy, tiny-prime handling and the measurements above are ours.