T2W-Qwen3.5-9B iter 12

This repository contains the full-rank Hugging Face export of training checkpoint iteration 12 from the T2W Qwen3.5-9B web-agent reinforcement-learning run.

Evaluation note

No full-portfolio evaluation result is reported in this model card.

Evaluation metrics are checkpoint-specific measurements and are not training claims. Compare checkpoints only under the same task manifest, evaluator, sampling parameters, and pass count.

Base model and license

The checkpoint is derived from Qwen/Qwen3.5-9B. The Apache-2.0 license file is included in this repository.

Downloads last month
-
Safetensors
Model size
10B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LEONW24/T2W-Qwen3.5-9B-iter12

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(708)
this model

Collection including LEONW24/T2W-Qwen3.5-9B-iter12