None defined yet.
Rethinking On-Policy Distillation of Large Language Models II: One Training Example