Instructions to use ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop") model = AutoModelForCausalLM.from_pretrained("ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop
- SGLang
How to use ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop with Docker Model Runner:
docker model run hf.co/ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop
Qwen2.5-3B-Instruct SFT in WebShop
This repository contains a full-parameter supervised fine-tune of
Qwen/Qwen2.5-3B-Instruct
for multi-turn interaction with the WebShop environment. Improved using
Qwen.
The released weights are checkpoint global_step_160, selected because it had
the highest success rate among the evaluated SFT checkpoints on one fixed set
of 256 held-out WebShop goals. The repository also includes the 3,000
environment-verified teacher trajectories, the clean SFT messages, the exact
processed train/validation split, and checkpoint-comparison artifacts.
中文简介
这是一个面向 WebShop 多轮购物 Agent 的 Qwen2.5-3B-Instruct 全参数 SFT
模型。发布权重来自 global_step_160:它在同一组 256 条留出 WebShop
任务上取得了所有已评测 SFT checkpoint 中最高的成功率。仓库同时提供
3,000 条环境验证成功的教师轨迹、干净的 SFT 对话、实际训练使用的数据
划分,以及 baseline/各 checkpoint 的指标与对比图。
Evaluation summary
All models were evaluated greedily on the same 256 held-out goals with
max_turns=9, evaluation seed 123, and goal-generation seed 20260819.
Step 0 is the untouched Qwen2.5-3B-Instruct baseline.
| Model | Step | Success | Mean episodic/raw return | Valid-action rate | Mean actions |
|---|---|---|---|---|---|
| Baseline | 0 | 0.015625 | 0.141494 | 0.631061 | 7.7695 |
| SFT | 80 | 0.046875 | 0.290312 | 0.755315 | 6.5469 |
| SFT (released) | 160 | 0.054688 | 0.308721 | 0.754712 | 6.3398 |
| SFT | 240 | 0.031250 | 0.275714 | 0.731383 | 6.2266 |
| SFT | 320 | 0.050781 | 0.297969 | 0.770179 | 6.2734 |
| SFT | 376 | 0.046875 | 0.275796 | 0.792147 | 6.3281 |
The released checkpoint raised success from 4/256 to 14/256: an absolute gain of 3.906 percentage points on this evaluation set. Mean episodic return rose from 0.141494 to 0.308721, and valid-action rate rose by 12.365 percentage points. These are point estimates from one fixed evaluation set; no confidence interval or multi-seed significance claim is made.
Machine-readable results are available in
results/comparison_metrics.csv and
results/metrics/.
Training data
The teacher-data pipeline used qwen3.6-35b-a3b and the train split of
webshop-small. Generation was oracle-assisted for search efficiency, but
every accepted trajectory still had to pass the environment and quality gates:
- exact binary success and raw reward of
1.0; - successful episode termination and purchase;
- every action valid on the current page;
- an instruction/product semantic-consistency gate with minimum confidence
0.85; - one short rationale and exactly one action per assistant turn;
- no validation/test goals used for teacher generation.
The resulting source contains 3,000 unique successful training seeds, 284 unique target ASINs, and an average of 5.164 interaction turns per trajectory (range 3–9).
For training, one manually audited borderline seed (1344) was excluded,
examples were capped at 20 per target ASIN, and an ASIN-disjoint 90/10 split was
constructed:
| Split statistic | Value |
|---|---|
| Quality-eligible source rows | 2,999 |
| Selected rows after ASIN cap | 1,672 |
| Train rows / unique ASINs | 1,505 / 249 |
| Validation rows / unique ASINs | 167 / 35 |
| Train-validation ASIN overlap | 0 |
Data files
data/raw/sft_messages.jsonl: 3,000 clean multi-turn positive SFT examples. This is the recommended human-readable SFT source.data/raw/trajectories.jsonl: complete audit records, including environment states, available actions, rewards, semantic-gate output, and private teacher audit fields. Use it for auditing, not directly as the SFT input.data/processed/train.parquetandval.parquet: the exact balanced, ASIN-disjoint data used by the SFT run.data/processed/split_index.json: selected example IDs and split assignment.data/DATASET_MANIFEST.json: public generation and preprocessing metadata.
Only sft_messages.jsonl or the processed parquet files should be used as
positive SFT data. Rejected attempts and API logs are intentionally not
published.
Training configuration
Training used the official multi-turn FSDP SFT trainer from
verl, vendored through the RAGEN
experiment repository.
| Setting | Value |
|---|---|
| Base model | Qwen/Qwen2.5-3B-Instruct |
| Training type | Full-parameter SFT |
| Epochs / optimizer steps | 1 / 376 |
| Learning rate | 5e-6 |
| Precision | BF16 |
| Maximum sequence length | 8,192 |
| Effective global batch size | 4 |
| Hardware | 4 × NVIDIA A100-PCIE-40GB |
| Checkpoint interval | 80 optimizer steps, plus final step |
The optimizer portion took 3,650.91 seconds (4.0566 A100 GPU-hours). The full
Slurm job, including setup, checkpoint I/O, evaluation, and plotting, used an
estimated 7.6471 allocated A100 GPU-hours. See
results/gpu_time_record.md for accounting
details.
Expected interaction format
The model was trained with the following system instruction:
You are a WebShop shopping agent. Follow the shopping instruction by interacting with the current page. At every turn, choose exactly one available action. Respond with exactly <think>brief rationale</think><answer>one action</answer> and no additional text.
Assistant turns follow this structure:
<think>I should search for the requested product category.</think><answer>search[product keywords]</answer>
The answer must be an action allowed by the current WebShop page, such as
search[...] or click[...].
Loading the model
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{
"role": "system",
"content": (
"You are a WebShop shopping agent. Follow the shopping instruction "
"by interacting with the current page. At every turn, choose exactly "
"one available action. Respond with exactly <think>brief rationale"
"</think><answer>one action</answer> and no additional text."
),
},
{
"role": "user",
"content": "Shopping instruction and the current WebShop page go here.",
},
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=256, do_sample=False)
response = tokenizer.decode(
output[0, inputs.input_ids.shape[1]:],
skip_special_tokens=True,
)
print(response)
Limitations
- The absolute held-out success rate is 5.47%; this is a specialized SFT initialization for further Agent RL research, not a production shopping system.
- Checkpoint 160 was selected after comparing checkpoints on one fixed set of 256 goals. The results may contain checkpoint-selection variance and do not establish statistical significance.
- The success curve is non-monotonic, while action validity continues to improve at later checkpoints. More SFT steps are therefore not uniformly better for the selected task metric.
- The data-generation pipeline used private oracle assistance. Oracle metadata is retained only in the audit-rich trajectory file; clean conversation content was checked for private-guidance leakage.
- The model is specialized to the formatting and action space used by this WebShop setup and should not be assumed to generalize to real commerce sites.
- No GRPO training is included in these released weights; this is the SFT checkpoint intended to initialize those experiments.
License and attribution
The model is a derivative of Qwen/Qwen2.5-3B-Instruct and is distributed
under the Qwen Research License Agreement, including its non-commercial-use
restriction. See LICENSE and the upstream
license page.
The repository's previous Apache-2.0 placeholder did not match the license
shipped with the local upstream 3B checkpoint and has been corrected.
Qwen is licensed under the Qwen Research License Agreement, Copyright (c) Alibaba Cloud. All Rights Reserved. This repository contains modified model weights and must not be interpreted as an official Qwen release.
The included WebShop-derived data and evaluation artifacts are provided for research reproducibility. Users are responsible for complying with the terms of the upstream WebShop resources and any applicable dataset restrictions.
- Downloads last month
- 415
