T2W-Qwen3.5-9B iter 24

This repository contains the full-rank Hugging Face export of training checkpoint iteration 24 from the T2W Qwen3.5-9B web-agent reinforcement-learning run.

Evaluation note

A recorded 953-task pass@2 evaluation reported 51.24% mean partial reward and 47.01% terminal pass@2 (448/953 tasks).

Evaluation metrics are checkpoint-specific measurements and are not training claims. Compare checkpoints only under the same task manifest, evaluator, sampling parameters, and pass count.

Base model and license

The checkpoint is derived from Qwen/Qwen3.5-9B. The Apache-2.0 license file is included in this repository.

Downloads last month
-
Safetensors
Model size
10B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LEONW24/T2W-Qwen3.5-9B-iter24

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(708)
this model

Collection including LEONW24/T2W-Qwen3.5-9B-iter24