Loop-distil: four- and eight-loop Huginn research checkpoints
Recurrent adapter/core weights from Nicholas0228/loop-distil. The code and research summary explain trajectory initialization, direct-KL distillation, the sampled-depth proposal, and all evaluation limitations.
These are partial model checkpoints. Each subdirectory contains all 37 trained recurrent adapter/core tensors as BF16 safetensors (3.05 GiB), with metadata and SHA-256 checksums. They replace the corresponding tensors in the pinned Huginn base; they are not additive deltas or LoRA adapters. Use the provided loader rather than calling AutoModelForCausalLM.from_pretrained on this repository.
Use a checkpoint
git clone https://github.com/Nicholas0228/loop-distil.git
cd loop-distil
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python scripts/generate_released.py \
--checkpoint 8loop-stage2-proposal-low-lr-100 \
--question "A box contains 12 pencils. Mia buys 3 boxes. How many pencils does she buy?"
The loader downloads tomg-group-umd/huginn-0125 at revision bb6621b65e90b6a4b9b29ef88dc83866d450470c, applies the selected weights, and sets the correct loop count. The base uses its pinned custom model code. CUDA inference was checked on an A800 40 GB. Use --cache-dir /path/with/space to choose storage, --base-model-path /path/to/pinned/base to reuse an existing snapshot, and --revision <Hub commit> to pin this release.
Generation uses the study's eight-shot GSM8K prompt, greedy decoding and a 256-token cap. --max-new-tokens changes the cap. Python users can import load_release from loop_distil.released; the returned model wrapper provides rollout with explicit loop count, seed and generation cap.
Checkpoints
All answer counts are automatic final numeric matches on the same 128 held-out GSM8K development questions. They are not a fresh official test result or a manual assessment of every solution. The frozen 32-loop teacher scores 59/128.
| ID | Stage-1 updates | Stage-2 updates | Stage-2 LR | Matches /128 |
|---|---|---|---|---|
4loop-stage1-2000 |
2,000 | β | β | 28 |
4loop-stage1-3500 |
3,500 | β | β | 38 |
4loop-stage1-4000 |
4,000 | β | β | 34 |
4loop-stage2-direct-500 |
2,000 | 500 | 2e-7 |
32 |
4loop-stage2-proposal-500 |
2,000 | 500 | 2e-7 |
31 |
8loop-stage1-1000 |
1,000 | β | β | 46 |
8loop-stage2-direct-low-lr-100 |
1,000 | 100 | 2e-7 |
48 |
8loop-stage2-proposal-low-lr-100 |
1,000 | 100 | 2e-7 |
55 |
8loop-stage2-direct-low-lr-300 |
1,000 | 300 | 2e-7 |
47 |
8loop-stage2-proposal-low-lr-300 |
1,000 | 300 | 2e-7 |
48 |
8loop-stage2-direct-high-lr-100 |
1,000 | 100 | 2e-6 |
32 |
8loop-stage2-proposal-high-lr-100 |
1,000 | 100 | 2e-6 |
8 |
Stage 1 matches student trajectory states to teacher depths 8/16/24/32 for four loops, or 4/8/β¦/32 for eight loops, using LR 2e-6. Stage 2 compares direct endpoint forward KL and the sampled-depth interpolated teacher-continuation KL. Both arms use identical cached teacher-generated answer prefixes and matched scored-token/update budgets. Four-loop stage 2 has not been rerun from the 3,500- or 4,000-update initialization.
Interpretation and limitations
- Initialization improves answer scores, but continuing to minimize MSE does not guarantee monotonic accuracy gains. The four-loop 3,500 checkpoint is a retrospectively selected peak; 4,000 is the predeclared endpoint.
- A tenfold lower stage-2 LR helps both eight-loop methods in the 100-update screen. The proposal's early 55-versus-48 result is promising but uncertain; its paired-question 95% interval versus direct KL includes zero. The later matched scores are 48 versus 47, and the four-loop comparison does not establish a proposal advantage.
- This is one training seed and a repeatedly reused development cohort. No convergence, reliable method superiority, or optimal learning rate is claimed. Loose number occurrences are not counted as extra correct answers. Numeric scoring has formatting and wrong-quantity artifacts.
- The teacher-generated training corpus contains 2,974 questions and was not filtered for teacher correctness. The models can produce incorrect or repetitive solutions.
- BF16 exports reproduce the compute weights used for evaluation. They omit FP32 masters and Adam state. They can initialize new fine-tuning but cannot exactly resume the original optimizer trajectory.
Provenance and license
The base is Huginn, pinned to the revision above. Only the adapter/core weights were trained; the frozen components come from the base. The derived weights are distributed under Apache-2.0, with the base attribution in NOTICE and the license text in LICENSE. Each checkpoint's checkpoint.json records its source master hash, training depths and updates, prefix policy, score, source report, and export checksum. catalog.json indexes the release.
Model tree for Nicholas0228/loop-distil-checkpoints
Base model
tomg-group-umd/huginn-0125