YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Recurrent depth × context: causal-v4 exploratory experiment
Date: 2026-10-09 UTC
Question: Does recurrent depth amplify the benefit of longer context at tiny scale?
Reproducibility
Source: recurrent_ctx_interaction_v4_causal.py
Source SHA-256: d5ef19a4cc0a6a26137865553e46a707882dae111ddd52832bd826f052dc0d09
Results SHA-256: 89acd1b3f005cdcf35f33a82bd1f194d9c3046a79c3531a7fdabd5b244b8ad75
Correctness preflight: passed (target alignment, causal attention, gradients, memory, checkpoint reload).
Tracked execution: exit code 0; duration 61.64 seconds; no recorded stderr.
Byte-level TinyStories next-token prediction. A 95/5 contiguous raw-data split separates training from validation, with source-code caps of 50 million training bytes and 2 million validation bytes. Four block applications; d_model 256; four attention heads; vocabulary 256; tied token embeddings; AdamW; 2,000 updates; batch size 32; learning rate 0.0003; seed 42. Recurrent feed-forward width 2048; stacked width 128. Training uses random sampled sequences. Validation is performed in eval mode without gradients.
Recorded metrics
| Configuration | Parameters | Context | Training-token positions | Final validation loss | Final validation perplexity |
|---|---|---|---|---|---|
| recurrent_ctx64 | 1,397,504 | 64 | 4,096,000 | 1.80549 | 6.08295 |
| recurrent_ctx256 | 1,446,656 | 256 | 16,384,000 | 2.23559 | 9.35204 |
| stacked_ctx64 | 1,402,880 | 64 | 4,096,000 | 1.92167 | 6.83238 |
| stacked_ctx256 | 1,452,032 | 256 | 16,384,000 | 2.27253 | 9.70391 |
Interpretation and limitations
Recurrent configurations achieved lower validation perplexity at both context lengths in this particular run. The recurrent-minus-stacked advantage in perplexity was approximately 0.749 at context 64 and 0.352 at context 256. This is exploratory evidence, not proof of an architecture benefit or a recurrence × context interaction. Context conditions processed different training-token counts (4.096M vs 16.384M) and were evaluated at different context lengths. Recurrent and stacked variants use different feed-forward widths to approximate parameter matching, not identical architectures. The random seed was initialized once, not reset per configuration. Validation means are unweighted averages of batch losses; the last partial batch can have disproportionate influence. Only one seed was tested. No inference-speed or broad quality evaluation is implied. Previous noncausal v3 results are invalid because they permitted future-token leakage.
A stronger follow-up would match training-token exposure, evaluate all conditions on an identical held-out target set using comparable available histories, use multiple independent seeds, and report confidence intervals.