Instructions to use cxt7chen/qwen36-35b-sft-initial with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use cxt7chen/qwen36-35b-sft-initial with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.6-35B-A3B") model = PeftModel.from_pretrained(base_model, "cxt7chen/qwen36-35b-sft-initial") - Notebooks
- Google Colab
- Kaggle
Qwen3.6-35B-A3B SFT initial adapter
Published by cxt7chen for a small initial training experiment. This is an adapter, not the full base model. Full SWE-bench Pro V2 HARD-51, HLE, ASI-Bench and Terminal-Bench evaluation has not been completed. No benchmark improvement is claimed.
Training and limitations
Assistant-only cross-entropy SFT on 7 execution-verified SWE-smith trajectories. The independent validation repository is python-docx; training repository is python-string-similarity.
Run: sft-initial-001; 14 optimizer steps; 7 training samples and one independent validation sample.
Full trajectories were retained; only assistant tokens receive labels. Rank 8, alpha 16, dropout 0.05, 20 language full-attention q/v targets, 1,024,000 trainable parameters. Vision, routers and fused expert tensors were frozen.
Supported Linear layers used NF4 double quantization with BF16 compute; fused experts remained BF16 and normalization parameters FP32. This is not an entirely four-bit model.
Parent: None; initialized from the fixed official base.
The final checkpoint was selected by the preset last-checkpoint rule, without benchmark feedback.
Final validation cross-entropy: 0.22825974225997925. GPU reloading reproduced it exactly. One validation sample cannot support generalization or benchmark claims. Hardware: one RTX PRO 6000 96GB. Software: PyTorch 2.8.0+cu128, Transformers 5.18.0, PEFT 0.21.2, bitsandbytes 0.50.2. Absolute paths in training-config.json record the original environment; adapt those paths for reproduction.
Data and licenses
SFT trajectories derive from SWE-smith, revision 08e109b4a59eaeebf80e4675cd125d42e7ac99a4 (MIT). Selected source repositories python-string-similarity and python-docx are MIT; see UPSTREAM-NOTICES.txt. Stage2 generated responses are execution-filtered model outputs on reserved training tasks. No formal benchmark answers or scores were used for training or checkpoint selection.
The base is Qwen/Qwen3.6-35B-A3B, fixed revision 995ad96eacd98c81ed38be0c5b274b04031597b0, licensed Apache-2.0. This adapter is released under Apache-2.0; upstream notices remain applicable.
Load
Use a CUDA GPU with sufficient memory for BF16 fused experts. The tested training/reload configuration used 96GB; smaller devices have not been validated.
import torch
from transformers import AutoProcessor, BitsAndBytesConfig, Qwen3_5MoeForConditionalGeneration
from peft import PeftModel
from huggingface_hub import hf_hub_download
import importlib.util
repo = "cxt7chen/qwen36-35b-sft-initial"
# Pin the actual adapter commit SHA after upload; it is not available before publication.
adapter_revision = "REPLACE_WITH_PUBLISHED_COMMIT_SHA"
base = "Qwen/Qwen3.6-35B-A3B"
processor = AutoProcessor.from_pretrained(base, revision="995ad96eacd98c81ed38be0c5b274b04031597b0")
model = Qwen3_5MoeForConditionalGeneration.from_pretrained(
base, revision="995ad96eacd98c81ed38be0c5b274b04031597b0", dtype=torch.bfloat16, device_map={"": 0},
quantization_config=BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16))
# Inspect the included preparation helper before importing it.
helper = hf_hub_download(repo, "prepare_adapter_base.py", revision=adapter_revision)
spec = importlib.util.spec_from_file_location("adapter_preparation", helper)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
module.prepare_frozen_base(model)
model = PeftModel.from_pretrained(model, repo, revision=adapter_revision, is_trainable=False)
model.eval()
model.gradient_checkpointing_disable()
model.config.use_cache = True
# Use processor.apply_chat_template with enable_thinking=True, preserve_thinking=True.
For reproducing logged validation loss, use BF16 torch.autocast and assistant-only labels; ordinary full-text loss is a different quantity. This packaged example follows the tested loader policy, but the published remote download path still requires post-upload verification.
- Downloads last month
- 15
Model tree for cxt7chen/qwen36-35b-sft-initial
Base model
Qwen/Qwen3.6-35B-A3B