ZhengyanWan/dFlowGRPO-GenEval-KL

LoRA adapter for FUDOKI, fine-tuned with dFlowGRPO on the GenEval reward, with KL regularization. Training step: 3300.

The KL-free counterpart is dFlowGRPO-GenEval.

See WanZhengyan/dFlowGRPO for training & evaluation code.

Files

  • lora_adapter_ema

Use

huggingface-cli download ZhengyanWan/dFlowGRPO-GenEval-KL \
    --local-dir output_grpo_geneval_kl/checkpoint_0_step3300

Then run the matching evaluator from discrete_flow_grpo/.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZhengyanWan/dFlowGRPO-GenEval-KL

Adapter
(4)
this model