TinyCeNN-LM Distilled v2

Transformer-free CeNN student distilled from arnir0/Tiny-LLM.

Architecture

  • Transformer layers remaining: 0
  • CeNN recurrent steps: 7
  • CeNN receptive field: 255 tokens
  • Trainable CeNN parameters: 480,192

Rigorous benchmark

  • Protocol: rigorous-v2
  • Deterministic held-out tokens: 65,536
  • Benchmark SHA256: 319c4527bb6a104bdd3772f30204fc8c32521f355f57d7c7fd387dbb058c3021
  • Cumulative distillation tokens: 40,005,632
  • Best student CE: 5.122586
  • Teacher CE: 4.234912
  • Best student PPL: 167.769
  • Teacher PPL: 69.056
  • Teacher-gap recovery: 84.42%

Load via tinycenn_lm.build_cenn_student().

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vtava/TinyCeNN-LM-Distilled-v2

Base model

arnir0/Tiny-LLM
Finetuned
(9)
this model