MicroMamba
MicroMamba is a small-compute test of input-dependent state-space dynamics. Each sequence contains distracting symbols, a few marked symbols, and a final query asking for one marked item by ordinal position. Solving the task requires selective storage and retrieval rather than ordinary next-token statistics.
The model uses a compact Mamba-inspired block with:
- a causal depthwise convolution;
- learned stable diagonal state dynamics;
- input-dependent discretization, input, and readout terms;
- gated residual output.
The benchmark retains two controls: a state-space model whose dynamics do not depend on the current input and a GRU with comparable scale. This is a pedagogical Mamba-inspired implementation, not a bit-exact reproduction of the official Mamba kernel or its large-scale language-model results.
Verified results
All variants trained on the same 12,000 length-48 sequences and were evaluated on 4,000 independently generated sequences at each length.
| Variant | Parameters | Length 48 | Length 96 zero-shot |
|---|---|---|---|
| Selective SSM | 4,594 | 87.85% | 87.23% |
| Fixed-dynamics SSM | 3,314 | 43.23% | 31.05% |
| GRU control | 7,146 | 69.38% | 70.00% |
On this controlled task, input-dependent state dynamics improved in-distribution accuracy by 44.63 points over fixed dynamics and 18.48 points over the larger GRU. The result is specific to this synthetic selective-memory benchmark.
Reproduce
uv run python projects/micro-mamba/train.py