SmolLM3-3B-SFT-WebShop

Task WebShop
Stage SFT (behavior-cloning) initialization
Base HuggingFaceTB/SmolLM3-3B
Init for this run models/smollm3-3b-agent
Source checkpoint rl/ckpts/ws_sft_div_smollm3-3b/global_step_155 (global_step 155)
Format bf16 safetensors, merged from hf-dir fp32 -> bf16 cast (per-tensor, login node)
Params 3,337,766,912
Weight drift vs. init (mean rel-L2) 4.04e-03 (max 8.66e-03, 271 / 326 tensors changed)
Reload bit-exact check True

Optimizer state is not included (inference/eval only; the fp32 FSDP shards stay on HiPerGator /blue). Load with AutoModelForCausalLM.from_pretrained(..., torch_dtype=torch.bfloat16) or vLLM.

Use the bundled chat_template.jinja — it is NOT the stock SmolLM3 template. This checkpoint was behaviour-cloned with a task-specific template that (a) defaults to /no_think, (b) does not pre-fill an empty <think></think> block into the generation prompt, and (c) replaces the long default reasoning instructions with a compact format directive. The model is expected to emit <think>…</think><action>…</action> itself. Loading the stock template instead silently changes the prompt distribution.

Apache-2.0, same as the base model.

Downloads last month
141
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wayne377/SmolLM3-3B-SFT-WebShop

Finetuned
(148)
this model