Bit-serial modular multiplication β€” radix-4 fork of neural-horner v8

A ~471K-parameter recurrent cell run in a fixed bit-serial Horner loop, computing (aΒ·b) mod p for primes to 2^2048 and operands to 4096 bits. Fork of Robby Sneiderman's neural-horner v8 (MIT), which is submitted separately as an attributed baseline.

Public benchmark: overall 98.9%, highest tier β‰₯90% = 10, 129.7 s on a 4090 against a 300 s budget.

Architecture

One shared p-conditioned cell (bidirectional 3-layer GRU, d_model 96, hidden 128, radix-4 digit and op embeddings, per-bit head) learns s' = (4s + d·x) mod p, d ∈ {0,1,2,3}; a learned op-embedding selects its three uses: reduce a, reduce b, multiply the residues. Radix-4 halves the step count at equal parameter count. Warm-started from v8 with the op-embedding zero-initialised, so at k=1 it reproduces v8 bit-for-bit.

Compliance: no controller, the schedule takes no feedback from the model, every arithmetic step including the conditional subtraction comes from the trained cell, and both routing choices key on the prime's bit-length alone β€” the shape ruled acceptable in the Zulip fixed-loop-encoder thread.

weights.pt is a weight-space average (model soup) of an annealed checkpoint and a same-recipe re-anneal; it beat both endpoints on identical problems (McNemar p = 0.02, n = 1008). weights_t1.pt serves primes of bit-length ≀ 3. Width policy: groups of β‰₯257 bits run at min(2080, pbits+32), 17–256 at pbits, <17 at 32.

Findings

1. Width robustness is a property of the training width distribution, not of the loop. Sizing the state register produced this project's largest gain, with no retraining. Sweeping Ξ” = state width βˆ’ prime bit-length on identical problems, 576 per width, four tier bands:

Ξ” 0 1 2 3 4 8 16 / 24 / 32 / 48 / 64
errors 27 516 470 116 13 13 3

Ξ” ∈ {1,2} falls to 2–30% accuracy. Ξ” = 3 is usable but degrades with rollout depth (T7 91.7%, deepest T10 47.9%). Error sets at Ξ” ∈ {16,24,32,48,64} are identical (Jaccard 1.00), so the 32 we deploy is about twice what is needed.

The same problems on v8 give zero errors at every Ξ” ∈ {0,1,2,3,4,32} β€” a flat curve. Loop, inference code and problem set are identical; the difference is training. v8's one-step training holds the state width at L and stratifies primes by bit-length, covering every Ξ”. Our radix-4 retraining trained each prime size at its own field-filling width, covering only Ξ” = 0.

Practical form: tier primes cluster just under multiples of 32, so rounding the state width up to a machine-word multiple lands them at 1–2 bits of headroom, the worst region of the curve. Add a fixed margin instead β€” "round up to 32" is not the same policy as "add 32", and falsifying the former cost us a week. The upstream argument that dynamic-L is exactness-preserving holds for v8 but does not travel with the weights.

2. Small primes are unreachable in any sampling stream. random_prime-style constructors emit odd values with the top bit set, so p = 2 must be injected explicitly. Our T1 sat at 61%, and our own battery could not see p = 2 either, so the hole was first misfiled as a false alarm. The fused single-step universe for p ∈ {2,3,5,7} is 348 transitions β€” enumerate it.

What did not work

Continued training failed in three forms. Unanchored tiny-prime fine-tuning cured small primes and ate large-scale precision (T8 βˆ’8, T10 βˆ’13). L2-SP has no Ξ» window: Ξ» = 10/30/100 relocates the interference rather than removing it (512-bit cells 166β†’170β†’172 while 64/128/256-bit cells fall monotonically). A same-recipe re-anneal paid nothing. The mechanism is measurable: our single-step validation floor is ~8e-6 while effective Ξ΅ is already 1.8e-6, so objective and checkpoint selection both sit below the instrument. Coverage and precision are zero-sum here. Ensembling failed too, in both the width-committee and cross-checkpoint-voting forms.

Limits

Structured operands are a live residual, independent of the width effect: the Fermat family is clean (96/96) but the wider power-of-two neighbourhood (2^m, 2^mΒ±1) is 280/288, failing only at primes β‰₯1024 bits, and headroom does not remove it (at 2013 bits, Ξ” = 0/4/8/32 gives 8/2/3/3 over 72 problems). Official tiers cannot generate such operands, so this is robustness, not score.

On the probe set above, v8 outperforms this model at every width: the radix-4 line bought a 2Γ— reduction in step count, not accuracy. The deep-chain Ξ΅ tail is deferred, not solved (T10 min 92%). Small primes needed a second routed checkpoint because one weight set could not hold both regimes. The remaining structural lever is raising the radix further; we ran out of time.

Attribution

Derived from neural-horner (github.com/Robby955/neural-horner) by Robby Sneiderman, MIT licensed. The architecture and the compliance framing originate there; the radix-4 fusion, width policy, tiny-prime handling and the measurements above are ours.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support