Qwen3-8B-Base-Math-TauSFT

Full BF16 Hugging Face weights exported from the final TauSFT checkpoint at step 405.

Training lineage: Qwen3-8B-Base โ†’ Qwen3-8B-Base-Math โ†’ TauSFT.

Training

  • Initialization: Math-stage checkpoint, iteration 300.
  • Supervised data: 25,956 tokenized records from Qwen3-32B Tau-bench rollout data (sft_data.jsonl).
  • Training: 405 updates, one configured epoch, global batch size 64.
  • Optimizer: Adam; learning rate 1e-5 with cosine decay to 1e-6 and 10% warmup.
  • Final reported SFT training loss: 0.3738237.

This repository contains the TauSFT stage before subsequent Tau-bench RL. Training data and optimizer states are not included.

Usage

Use a recent version of transformers with Qwen3 support, plus torch, accelerate, and safetensors.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "willhx/Qwen3-8B-Base-Math-TauSFT"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)
messages = [{"role": "user", "content": "Hello!"}]
text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(
    **inputs,
    max_new_tokens=256,
    do_sample=False,
    eos_token_id=[tokenizer.eos_token_id, tokenizer.convert_tokens_to_ids("<|im_end|>")],
    pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(outputs[0, inputs.input_ids.shape[1]:], skip_special_tokens=True))

The tokenizer and architecture configuration are preserved from Qwen3-8B-Base. For complete Tau-bench interactions, supply the corresponding retail policy, tool schemas, and environment. The example explicitly stops at the assistant turn-ending token as well as the base EOS token.

Downloads last month
263
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for willhx/Qwen3-8B-Base-Math-TauSFT

Finetuned
(1)
this model