Qwen3-Coder calendar agent RL — LR 2e-6, step 12

Qwen3-Coder-30B-A3B-Instruct after asynchronous RLOO training on the TaskTrove v4.9 agent-calendar source with identity-aware shaped reward. This checkpoint scored 0.460546875 mean reward (59/128 passes) on the fixed holdout. The holdout was later found to overlap the training source, so this score is a selection statistic, not an uncontaminated estimate of generalization.

Native checkpoint: q3c-rl-calendar-agent-v49-shaped-nodapo-lr2-r1/global_step_12.

Training Traces

The complete experiment artifacts, configs, metrics, trace samples, tracker, and contamination analysis are in qwen3coder-iris-rl-data-sweep-artifacts.

Downloads last month
412
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for penfever/qwen3coder-calendar-agent-v49-lr2-step12

Finetuned
(94)
this model