RL-Repaired-V1

Six inference/evaluation LoRA adapters at optimizer step 125: Qwen3.5 M0-v4 4B and 9B crossed with DPO, GRPO, and PPO. Every run used the exact ordered prompts-repaired-v1 cohort and its versioned localverifier-v1 derivative.

Size DPO GRPO PPO
4B 4b/dpo/step-125 4b/grpo/step-125 4b/ppo/step-125
9B 9b/dpo/step-125 9b/grpo/step-125 9b/ppo/step-125

The adapter configs point to immutable M0-v4 bases. Optimizer states are not present in the source checkpoint directories and are therefore not included; these are loadable model-only checkpoints, not training-resume archives. See each artifact_manifest.json and the root upload_manifest.json for hashes and provenance.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support