tinymistral-477m — 477M dense base

tinymistral-477m is a 477M-parameter dense decoder-only language model — the same-total-parameter, same-everything-else companion to the sparse Mixture-of-Experts model mikecovlee/tinymixtral.

It exists to answer one controlled question: at a fixed total parameter count and fixed data, what does MoE routing buy you? tinymistral-477m (dense, 477.4M total = active) is trained on the exact same data, tokenizer, 4-segment WSD schedule, batch size and LR ladder as the 477.5M-total / 276.1M-active MoE tinymixtral. Only the FFN differs: the four 2048-wide expert FFNs are merged into one 8192-wide dense FFN per layer and the routers are removed.

This is a base (pretrained) model — it has no chat/instruction fine-tuning.

Model details

tinymistral-477m tinymixtral (MoE)
Architecture Dense SwiGLU FFN 4 routed experts, top-2, aux 1e-3
Total parameters 477,400,064 477,465,600
Active parameters 477,400,064 276,139,008
FFN intermediate 8192 (4 experts merged) 2048 (per expert)
FFN MACs / token / layer 25,165,824 12,582,912 (top-2)
Hidden size 1024 1024
Layers 16 16
Attention GQA 16 Q / 4 KV heads, head_dim 64 same
Context length 2048 2048
Positional RoPE θ = 1e6, QK-Norm same
Norm / embeddings Pre-RMSNorm (eps 1e-6), tied embeddings same
Vocab 32,000 (TinyLlama tokenizer) same
Precision float32 checkpoint (bf16 training) same
License MIT MIT

The two models differ by exactly the 16 router matrices (16 × 1024 × 4 = 65,536 parameters, 0.014%); everything else is identical. The dense model pays ~1.7× the active FLOPs per token (477.4M vs 276.1M active parameters) — this is a parameter-matched control, not a compute-matched one. The compute-matched sibling is tinymistral-276m.

Training

  • Tokens: 8.05B, split into 4 strictly disjoint segments (2.00 / 1.94 / 2.20 / 1.91B, zero repetition).
  • Schedule: WSD per segment — warmup 700 steps → constant → linear decay over the final 10%. LR ladder 5e-4 / 5e-4 / 4e-4 / 3e-4 across S1→S4.
  • Batch: 48 × 1024 tokens = 49,152 tokens/step.
  • Optimizer: AdamW β(0.9, 0.95), weight decay 0.1 (no decay on norms/embeddings), grad clip 1.0.
  • Precision: bf16 autocast + bf16 optimizer states; gradient checkpointing.
  • Seed: 42. Segment boundaries are the anneal points; AdamW momentum carries across segments.
  • Hardware: single NVIDIA RTX A5000 (Ampere, 24 GB), ~12k tokens/s.

Data mix

FineWeb-Edu 44% · DCLM (web) 20% · Cosmopedia v2 12.5% · code (OpenCodeInstruct) 12.5% · math (OpenWebMath) 6% · Wikipedia 6% (shares rounded, sum ≈ 100%).

Evaluation

Same-protocol evaluation (lm_eval 0.4.12, 0-shot, no chat template). Harness = mean of the 7 primary metrics: hellaswag acc_norm, piqa acc_norm, winogrande acc, arc_easy acc, arc_challenge acc_norm, openbookqa acc_norm, lambada acc.

Metric tinymistral-477m (dense) tinymixtral (MoE)
7-task harness 0.4013 0.3979
MMLU 5-shot 0.2399 0.2339
TruthfulQA MC1 / MC2 0.2362 / 0.4090 0.2375 / 0.4171

Per-task scores (7-task suite):

Task tinymistral-477m (dense) tinymixtral (MoE)
HellaSwag (acc_norm) 0.3414 0.335
PIQA (acc_norm) 0.6355 0.638
WinoGrande (acc) 0.5114 0.515
ARC-Easy (acc) 0.4907 0.478
ARC-Challenge (acc_norm) 0.2568 0.255
OpenBookQA (acc_norm) 0.302 0.296
LAMBADA (acc) 0.2711 0.268

Dense per-segment trajectory (7-task harness / MMLU): 0.3931 / 0.2291 → 0.3918 / 0.2386 → 0.4016 / 0.2337 → 0.4013 / 0.2399.

Harness note. Every figure on this card is computed under the one formula stated above from the campaign artifacts. The tinymistral-276m card quotes its MoE column under a slightly different formula (piqa acc) and an earlier MMLU run, so cross-card means are not bitwise comparable. Here both columns are recomputed under the identical formula.

Takeaway: at fixed total parameters and identical data, removing MoE routing does not hurt — the dense model is +0.34 pp on the 7-task harness and +0.6 pp on MMLU (six of eleven metrics up; the losses are within noise for this scale). Read together with the iso-FLOPs result (tinymistral-276m: the MoE is +0.88 pp at equal active parameters), the conclusion is that the MoE advantage is compute efficiency, not parameter efficiency: a dense model can match or beat the MoE at equal total parameters, but only by spending ~1.7× the FLOPs per token. MMLU / TruthfulQA are near chance for both models and are noise-dominated.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "mikecovlee/tinymistral-477m"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)

inputs = tok("The capital of France is", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

Requires transformers and trust_remote_code=True (custom tinymixtral architecture).

Limitations

  • Base model: no instruction tuning; not a chat model. Outputs should not be used as-is for assistant tasks.
  • Trained on only 8.05B tokens — far below modern small-model budgets (SmolLM2-360M / Qwen3-0.6B use 2–36T). Knowledge is capacity/budget-bound: MMLU ≈ chance.
  • English-centric, no safety alignment.
  • The FFN is unusually wide (8× hidden, a merged-MoE shape); standard dense models at this size use ~2.5–3.5× and more layers. This is deliberate — it keeps the comparison with tinymixtral controlled — but a textbook-shaped 477M dense would likely score slightly differently.

Family

Naming. The MoE family is published under tinymixtral; the dense ablation companions use the tinymistral spelling. Both belong to the same project.

Citation

@misc{tinymistral477m2026,
  title  = {TinyMixtral: a small Mixture-of-Experts language-model family},
  author = {Michael Lee},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/mikecovlee/tinymistral-477m}}
}

License

MIT (Copyright (C) 2026 Michael Lee).

Downloads last month
239
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train mikecovlee/tinymistral-477m

Collection including mikecovlee/tinymistral-477m