Mirror of unsloth/Qwen3.6-35B-A3B-GGUF (Q4_K_M + MTP) by pitcany. Mirrored 2026-09-15 as takedown protection. Apache-2.0.

Qwen3.6-35B-A3B (MoE, 35.5B params / 3.5B active) in Q4_K_M with 1 MTP speculative decoding layer merged in. Single 21 GB GGUF file.

Built by Ollama from the unsloth base quant + MTP draft module. Use with Ollama (speculative decoding) or llama.cpp.

Architecture

  • 41 transformer blocks, 256 experts (8 active), GQA 16:2
  • Context: 262,144 tokens
  • 1 nextn-predict (MTP) layer for speculative decoding

Usage

# Ollama (with MTP speculative decoding)
ollama run qwen3.6:35b-a3b-mtp-q4_K_M

# llama.cpp
llama-cli -m qwen3.6-35b-a3b-mtp-q4_K_M.gguf
Downloads last month
46
GGUF
Model size
36B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for pitcany/qwen3.6-35b-mtp-q4

Quantized
(826)
this model