rob-1-flash-2b (GGUF) โ€” unofficial derivative

Russian chat model of the rob-1 line (flash tier). Built on Qwen/Qwen3.5-2B (Apache 2.0), not affiliated with Alibaba Qwen team. Base license applies.

What was done

  1. SFT (LoRA r=8, attention + MLP, base frozen) on 2934 clean Russian rows: everyday Q&A, math, Python/Java/Lua/C++ code, identity rows (rob-1-flash, creator robanik). Shuffled, CPU.
  2. Adapters merged into stock weights (MTP tensors preserved).
  3. GGUF via llama.cpp: F16 (3.9 GB) + Q8_0 (1.9 GB).
  4. Engine-verified in llama-server: math, Russian small talk, identity (ะœะตะฝั ะทะพะฒัƒั‚ rob-1-flash), ~15 tok/s CPU.

Files

  • rob-1-flash-2b-Q8_0.gguf (1.9 GB) โ€” recommended for LM Studio / Ollama.
  • rob-1-flash-2b-F16.gguf (3.9 GB) โ€” full precision merged.

Use (LM Studio)

Pull this repo from the Models tab or drop a .gguf in, context 4096. Sampling: temperature 1.0, top_k 20, top_p 1.0, presence_penalty 2.0.

Honesty notes

  • No formal benchmarks were run; domain spot-checks only.
  • Small SFT shifts style toward the training rows; hard reasoning stays at base 2B level. No scripts or training code distributed.
Downloads last month
3
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for RobTeam/rob-1-flash-2b

Finetuned
Qwen/Qwen3.5-2B
Quantized
(215)
this model