TinyCeNN-LM Story AntiRepeat

Transformer-free TinyCeNN story specialization derived from vtava/TinyCeNN-LM-Sharded-MoE-Top2.

Training uses TinyStories with causal CE plus recent-token unlikelihood loss. No held-out benchmark is run in this fast workflow.

Recommended decoding: temperature 0.78, top-p 0.90, top-k 40, repetition penalty 1.18, no-repeat 4-gram.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support