YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.


license: apache-2.0 base_model: Qwen/Qwen3-Coder-30B-A3B-Instruct library_name: transformers tags:

  • rl
  • grpo
  • skyrl
  • terminus-2 pipeline_tag: text-generation

tt-x3_kl-kl0b -- step 72 (X3 KL experiment)

GRPO checkpoint from the TaskTrove X3 (KL coefficient) sweep. Base model Qwen/Qwen3-Coder-30B-A3B-Instruct, trained on DCAgent/exp_rpt_multifile with SkyRL + Terminus-2; campaign verifier is pass_ratio shaping. KL arm runs carry a reference model (policy world size 32).

Checkpoint selection -- best RETAINED

global_step_72 is the highest-trailing-5-EMA checkpoint among the retained FSDP bank (EMA 0.1530 at step 72; step reward 0.1484; pass@8 0.3125). The in-run EMA maximum occurred around step 60 but those checkpoints were rotated out by max_ckpts_to_keep=2, and the run HF-export hook crashed at source (ModelLocatorError), so no earlier exports exist. Converted post-hoc on an 8x4 GH200 gang (fsdp_size=8, EP=4) via the checkpoint_export entrypoint.

Run status -- terminated by owner at step 73/80

Near-horizon owner stop of the X3 KL arm. Not a horizon result.

See training_logs/ for metrics.csv, report.md, reward_plot.png, rl_config.json, and the gzipped .out chain.

Training Traces

penfever/tt-x3_kl-kl0b -- 1/4 systematic subsample (every 4th trial, uniform coverage). GPFS-read-bound login node; documented, owner-approved deviation.

Downloads last month
9
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support