Llama-3.2-1B — vanilla
æ ‡å‡† softmax 基线(attn_softmax=vanilla, attn_res_softmax_fn=vanilla).
- Base model: meta-llama/Llama-3.2-1B
- Training data: BookCorpus + Wiki40B (en)
- Steps: 1000 (warmup 100), block_size 512, effective batch 192 (6 × 4 GPU × 8 grad-accum)
- Optimizer: AdamW, lr 4e-4, linear, weight_decay 0.1
- Precision: fp16 mixed (fp32 master weights)
Research checkpoint from an outlier-efficiency study; only 1000 steps, not production-ready.
Metrics (final eval)
| metric | value |
|---|---|
| perplexity | 143.82 |
| model.norm inf-norm | 14.8 |
| max per-layer inf-norm | 311.9 |
Companion run: robinzixuan/llama-3.2-1b-oasis.
- Downloads last month
- 66
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for robinzixuan/llama-3.2-1b-vanilla
Base model
meta-llama/Llama-3.2-1B