ExOPD λ=1.0 + KL-to-reference trust region, step 5/15

Intermediate checkpoint from a smoke test of a fix for a reasoning-shortcut reward hack found in On-Policy Distillation with Reward Extrapolation (ExOPD) training. Full context: ntlfi/opd_training (see SESSION_LOG.md and algos/exopd_incident_2026-08-07.md).

Not formally evaluated (only step10 and step15/final were run through the full GSM8K test set). Training-time metrics at this step: loss=0.192, kl_penalty=0.0003 (very small — most of the drift measured by the KL term happens later in the run, see step10/step15 for comparison).

Teacher: meta-llama/Meta-Llama-3-70B-Instruct (frozen). Student/base: meta-llama/Meta-Llama-3-8B-Instruct. λ=1.0 (standard on-policy distillation reduction, no reward-extrapolation reference term active), plus an added KL(π_θ‖π_ref) trust-region penalty (kl_ref_coef=0.1, not part of the published method — added to counter a reward-hacking failure mode this repo found independently) using Schulman's k3 estimator.

Downloads last month
5
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ntlfi/opd-exopd-lambda1.0-klref-step5

Finetuned
(1148)
this model

Paper for ntlfi/opd-exopd-lambda1.0-klref-step5