Flow-Aligned Distillation checkpoints

This private artifact repository contains the six seed-44 FAD best-validation training checkpoints used for the 3B/8B, 15%/20%/25% operating points.

These files are custom PyTorch training checkpoints, not standalone Transformers from_pretrained directories. Load them with the FAD pipeline in the companion code repository:

https://github.com/hanklin9188/Flow-Aligned-Distillation

Checkpoints

Backbone Nominal budget Stored FFN banks Seed Best validation step File
Llama-3.2-3B 15% 22 44 4,500 checkpoints/llama-3.2-3b/fad-15pct-k22-seed44.pt
Llama-3.2-3B 20% 19 44 4,500 checkpoints/llama-3.2-3b/fad-20pct-k19-seed44.pt
Llama-3.2-3B 25% 17 44 4,500 checkpoints/llama-3.2-3b/fad-25pct-k17-seed44.pt
Llama-3.1-8B 15% 25 44 7,500 checkpoints/llama-3.1-8b/fad-15pct-k25-seed44.pt
Llama-3.1-8B 20% 23 44 9,500 checkpoints/llama-3.1-8b/fad-20pct-k23-seed44.pt
Llama-3.1-8B 25% 20 44 not recorded checkpoints/llama-3.1-8b/fad-25pct-k20-seed44.pt

The 8B 25% run retains its best-validation training checkpoint but does not have a completed shared_student.pt or deploy_bundle.pt export in the source archive. It is published here as a recoverable training checkpoint and should be exported with the companion pipeline before deployment.

Training contract

  • objective: decision-token CE + FAD functional alignment
  • private adapter: rank/alpha 128/128
  • seed: 44
  • checkpoint selection: minimum validation decision_ce
  • 3B projection rank: 256
  • 8B projection rank: 512

External requirements

The checkpoints do not redistribute Meta Llama base weights, merged teacher weights, tokenizer files, datasets, or licenses. Users must obtain the matching base model and teacher artifacts separately and comply with Meta's model terms.

SHA256SUMS is generated before upload and is the integrity authority for the six checkpoint files.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NYCU-MLLab/Flow-Aligned-Distillation

Finetuned
(1496)
this model