Reinforcement Learning
Transformers
Safetensors
qwen3
text-generation
search-r1
lora
merged
text-generation-inference
Instructions to use willamazon1/sdft-search-lora-iter160 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use willamazon1/sdft-search-lora-iter160 with Transformers:
# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("willamazon1/sdft-search-lora-iter160") model = AutoModelForCausalLM.from_pretrained("willamazon1/sdft-search-lora-iter160", device_map="auto") - Notebooks
- Google Colab
- Kaggle
sdft-search-lora-iter160
Qwen3-8B SDFT checkpoint with a search-R1 RL LoRA adapter merged into the base weights (full dense model, not an adapter — load it like any Qwen3-8B).
What this is
| Base | Qwen3-8B-Base, SDFT'd on an oracle-mix (Qwen3-8B-oracle-mix-SFT-balance) |
| RL | search-R1 style multi-turn retrieval RL, trained with slime (GSPO) |
| Adapter | LoRA rank 16, alpha 32 (scaling 2.0) on q/k/v, o_proj, gate/up_proj, down_proj of all 36 layers |
| Iteration | RL step 160 |
| Merge | W += (alpha/r) * (B @ A) applied in Megatron parameter space, then converted to HF |
Provenance / honesty notes
- The adapter at this iteration is verified trained: all 144 LoRA modules have non-zero
lora_B(max |lora_B| = 1.46e-4). Earlier iterations of this project silently saved untrained adapters (lora_Bbit-exactly 0) because a TIS rejection-sampling veto zeroed the gradient for effectively every sequence; those are not published. - The merged delta is small: relative Frobenius norm ||ΔW||/||W|| per LoRA'd matrix is min 5.2e-5, median 9.6e-5, max 2.4e-4. This is "SDFT base + a real but modest RL delta", not a heavily-RL'd model. Treat it accordingly in evaluations.
- Verified after merge: 0 non-finite tensors, embedding shape (151936, 4096), greedy decoding
sane (
The capital of France is->Paris,12*8=->96). - Tokenizer/config assets come from Qwen3-8B-Base (unchanged).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "willamazon1/sdft-search-lora-iter160"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
- Downloads last month
- 20
Model tree for willamazon1/sdft-search-lora-iter160
Base model
Qwen/Qwen3-8B-Base