Diltillation of pre-distilled model

#10
by drakexp - opened

Minimax H3 is already a CFG free distilled model with 20 inference steps. Won't training a distillation adapter over a pre-distilled model affect the quality vs training on a base model?

Sign up or log in to comment