Qwen3.5-2B-SFT-RU (GGUF) โ€” unofficial derivative

Unofficial Russian SFT derivative of Qwen/Qwen3.5-2B (Apache 2.0). Not affiliated with Alibaba Qwen team. Base model license (Apache 2.0) applies.

What was done

  1. SFT (LoRA r=8, MLP + attention, base frozen) on 572 Russian rows: everyday Q&A, math, 40 Python + 40 Java code tasks. 159 steps, CPU.
  2. Adapters merged into stock weights (MTP tensors preserved).
  3. Real GGUF quantization with llama.cpp: imatrix calibrated on 1200 mixed rows, then Q4_K_M (5.36 BPW, 1.22 GB). Also ships F16 (3.9 GB).
  4. Engine-verified: loads in llama-server, 15*4 โ†’ 60 at ~22 tok/s CPU.

Files

  • Qwen3.5-2B-SFT-Q4_K_M.gguf (1.22 GB) โ€” recommended for LM Studio / Ollama.
  • Qwen3.5-2B-SFT-F16.gguf (3.9 GB) โ€” full precision merged SFT.

Use (LM Studio)

Drop the .gguf into LM Studio (or pull this repo from the Models tab), context 4096 to start. Recommended sampling (Qwen3.5 non-thinking text): temperature 1.0, top_k 20, top_p 1.0, presence_penalty 2.0.

Honesty notes

  • No formal benchmarks (MMLU etc.) were run. Domain checks only: math (15*4 = 60), Russian small talk, short Python answers.
  • Small SFT (159 steps) shifts style toward the training rows; hard reasoning stays at the base 2B level. No scripts or training code distributed.
Downloads last month
127
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for RobTeam/qwen3.5-2b-ru-gguf

Finetuned
Qwen/Qwen3.5-2B
Quantized
(240)
this model