Qwen3.6-35B-A3B SFT initial adapter

Published by cxt7chen for a small initial training experiment. This is an adapter, not the full base model. Full SWE-bench Pro V2 HARD-51, HLE, ASI-Bench and Terminal-Bench evaluation has not been completed. No benchmark improvement is claimed.

Training and limitations

Assistant-only cross-entropy SFT on 7 execution-verified SWE-smith trajectories. The independent validation repository is python-docx; training repository is python-string-similarity.

Run: sft-initial-001; 14 optimizer steps; 7 training samples and one independent validation sample. Full trajectories were retained; only assistant tokens receive labels. Rank 8, alpha 16, dropout 0.05, 20 language full-attention q/v targets, 1,024,000 trainable parameters. Vision, routers and fused expert tensors were frozen. Supported Linear layers used NF4 double quantization with BF16 compute; fused experts remained BF16 and normalization parameters FP32. This is not an entirely four-bit model. Parent: None; initialized from the fixed official base. The final checkpoint was selected by the preset last-checkpoint rule, without benchmark feedback.

Final validation cross-entropy: 0.22825974225997925. GPU reloading reproduced it exactly. One validation sample cannot support generalization or benchmark claims. Hardware: one RTX PRO 6000 96GB. Software: PyTorch 2.8.0+cu128, Transformers 5.18.0, PEFT 0.21.2, bitsandbytes 0.50.2. Absolute paths in training-config.json record the original environment; adapt those paths for reproduction.

Data and licenses

SFT trajectories derive from SWE-smith, revision 08e109b4a59eaeebf80e4675cd125d42e7ac99a4 (MIT). Selected source repositories python-string-similarity and python-docx are MIT; see UPSTREAM-NOTICES.txt. Stage2 generated responses are execution-filtered model outputs on reserved training tasks. No formal benchmark answers or scores were used for training or checkpoint selection. The base is Qwen/Qwen3.6-35B-A3B, fixed revision 995ad96eacd98c81ed38be0c5b274b04031597b0, licensed Apache-2.0. This adapter is released under Apache-2.0; upstream notices remain applicable.

Load

Use a CUDA GPU with sufficient memory for BF16 fused experts. The tested training/reload configuration used 96GB; smaller devices have not been validated.

import torch
from transformers import AutoProcessor, BitsAndBytesConfig, Qwen3_5MoeForConditionalGeneration
from peft import PeftModel
from huggingface_hub import hf_hub_download
import importlib.util

repo = "cxt7chen/qwen36-35b-sft-initial"
# Pin the actual adapter commit SHA after upload; it is not available before publication.
adapter_revision = "REPLACE_WITH_PUBLISHED_COMMIT_SHA"
base = "Qwen/Qwen3.6-35B-A3B"
processor = AutoProcessor.from_pretrained(base, revision="995ad96eacd98c81ed38be0c5b274b04031597b0")
model = Qwen3_5MoeForConditionalGeneration.from_pretrained(
    base, revision="995ad96eacd98c81ed38be0c5b274b04031597b0", dtype=torch.bfloat16, device_map={"": 0},
    quantization_config=BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
        bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16))
# Inspect the included preparation helper before importing it.
helper = hf_hub_download(repo, "prepare_adapter_base.py", revision=adapter_revision)
spec = importlib.util.spec_from_file_location("adapter_preparation", helper)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
module.prepare_frozen_base(model)
model = PeftModel.from_pretrained(model, repo, revision=adapter_revision, is_trainable=False)
model.eval()
model.gradient_checkpointing_disable()
model.config.use_cache = True
# Use processor.apply_chat_template with enable_thinking=True, preserve_thinking=True.

For reproducing logged validation loss, use BF16 torch.autocast and assistant-only labels; ordinary full-text loss is a different quantity. This packaged example follows the tested loader policy, but the published remote download path still requires post-upload verification.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cxt7chen/qwen36-35b-sft-initial

Adapter
(297)
this model