Qwen3-Coder calendar agent RL — LR 4e-6, step 9

Qwen3-Coder-30B-A3B-Instruct after asynchronous RLOO training on the TaskTrove v4.9 agent-calendar source with identity-aware shaped reward. This checkpoint scored 0.4609375 mean reward (59/128 passes) on the fixed holdout. The holdout was later found to overlap the training source, so this score is a selection statistic, not an uncontaminated estimate of generalization.

Native checkpoint: q3c-rl-calendar-agent-v49-shaped-nodapo-lr4-r1/global_step_9.

Training Traces

The complete experiment artifacts, configs, metrics, trace samples, tracker, and contamination analysis are in qwen3coder-iris-rl-data-sweep-artifacts.

Downloads last month
411
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for penfever/qwen3coder-calendar-agent-v49-lr4-step9

Finetuned
(94)
this model