tiny-llm-pipeline-29m
The weights of a ~29M-parameter Chinese small model (MiniMind2-Small config: vocab 6400 / width 512 / 8 layers), from three training stages. From AEFS-Capstones / tiny-llm-pipeline.
| stage | step | notes |
|---|---|---|
pretrain |
34500 | next-token pretraining, val ppl 11.8 |
sft |
5000 | instruction tuning, holdout response ppl 6.34 (38.15 before tuning) |
dpo |
104 | DPO alignment to a "Doubao-style" voice |
The layout matches the source project, so Bundle.load() can point straight at this snapshot:
tokenizer/tokenizer.json
<stage>/ckpt.pt
<stage>/train_log.jsonl
<stage>/monitor.png # not present for dpo
ckpt.pt holds model / cfg / step / seed / tokenizer_hash / tokens / cursor, with the
optimizer stripped — so --resume is not possible, while inference is unaffected. The weights
are tensor-identical to the original training outputs.
Notes
- Generation must set
repetition_penalty(1.3-1.5); without it the DPO weights loop. - The weights are not bit-reproducible (MPS and CUDA differ in float paths).
- Training data comes from
jingyaogong/minimind_dataset; the style outline lives in the source repo underartifacts/prefs/.