TinyCeNN-LM Story AntiRepeat
Transformer-free TinyCeNN story specialization derived from vtava/TinyCeNN-LM-Sharded-MoE-Top2.
Training uses TinyStories with causal CE plus recent-token unlikelihood loss. No held-out benchmark is run in this fast workflow.
Recommended decoding: temperature 0.78, top-p 0.90, top-k 40, repetition penalty 1.18, no-repeat 4-gram.