ZhengyanWan/dFlowGRPO-ScienceQA

LoRA adapter for FUDOKI, fine-tuned with dFlowGRPO on the ScienceQA reward. Training step: 1500.

See WanZhengyan/dFlowGRPO for training & evaluation code.

Files

  • lora_adapter
  • ema_state.pt

Use

huggingface-cli download ZhengyanWan/dFlowGRPO-ScienceQA \
    --local-dir output_grpo_u_scienceqa_kl0.01/checkpoint_0_step1500

Then run the matching evaluator from discrete_flow_grpo/.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZhengyanWan/dFlowGRPO-ScienceQA

Adapter
(4)
this model