Qwen3-Coder calendar instruction-following RL — step 6

Qwen3-Coder-30B-A3B-Instruct after asynchronous RLOO training on the TaskTrove v4.9 calendar instruction-following source. This checkpoint scored 0.3046875 mean reward (39/128 passes) on the fixed holdout. The holdout was later found to overlap the training source, so this score is a selection statistic, not an uncontaminated estimate of generalization.

Native checkpoint: q3c-rl-calendar-if-v49-nodapo-lr2-r1/global_step_6.

Training Traces

The complete experiment artifacts, configs, metrics, trace samples, tracker, and contamination analysis are in qwen3coder-iris-rl-data-sweep-artifacts.

Downloads last month
446
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for penfever/qwen3coder-calendar-if-v49-lr2-step6

Finetuned
(94)
this model