tt-x5_gradnorm-gn0p9

RL-finetuned from Qwen/Qwen3-Coder-30B-A3B-Instruct using RLOO with FSDP2 expert-parallel training on the exp_rpt_multifile terminal-bench agentic task suite (terminus-2 Harbor harness).

Training configuration

parameter value
algorithm RLOO (n=8)
strategy FSDP2 + expert-parallel (EP=4)
max_grad_norm 0.9
learning_rate 8e-6
eps_clip low=0.2, high=0.05
loss_reduction seq_mean_token_sum_norm_global
TIS enabled (cap=2.0)
KL loss disabled (coef=0.0)
batch_size 64 groups × 8 samples
max_steps 80 (reached 66 — wall-time limited)
selected checkpoint global_step_65

Training results

metric value (step 65)
reward (mean) 0.197
pass@8 0.406
policy_entropy 0.492
ppo_clip_ratio 0.011

Training Traces

Companion trace dataset: DCAgent/tt-x5_gradnorm-gn0p9

Downloads last month
193
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for laion/tt-x5_gradnorm-gn0p9

Finetuned
(82)
this model