YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
license: apache-2.0 base_model: Qwen/Qwen3-Coder-30B-A3B-Instruct library_name: transformers tags:
- rl
- grpo
- skyrl
- terminus-2 pipeline_tag: text-generation
tt-x3_kl-kl0b -- step 72 (X3 KL experiment)
GRPO checkpoint from the TaskTrove X3 (KL coefficient) sweep. Base model Qwen/Qwen3-Coder-30B-A3B-Instruct, trained on DCAgent/exp_rpt_multifile with SkyRL + Terminus-2; campaign verifier is pass_ratio shaping. KL arm runs carry a reference model (policy world size 32).
Checkpoint selection -- best RETAINED
global_step_72 is the highest-trailing-5-EMA checkpoint among the retained FSDP bank (EMA 0.1530 at step 72; step reward 0.1484; pass@8 0.3125). The in-run EMA maximum occurred around step 60 but those checkpoints were rotated out by max_ckpts_to_keep=2, and the run HF-export hook crashed at source (ModelLocatorError), so no earlier exports exist. Converted post-hoc on an 8x4 GH200 gang (fsdp_size=8, EP=4) via the checkpoint_export entrypoint.
Run status -- terminated by owner at step 73/80
Near-horizon owner stop of the X3 KL arm. Not a horizon result.
See training_logs/ for metrics.csv, report.md, reward_plot.png, rl_config.json, and the gzipped .out chain.
Training Traces
penfever/tt-x3_kl-kl0b -- 1/4 systematic subsample (every 4th trial, uniform coverage). GPFS-read-bound login node; documented, owner-approved deviation.
- Downloads last month
- 9