MTP Issues with the Q4_K_M Quant

#21
by tanpurohit26 - opened

Running the Q4_K_M (4.27BPW) model locally but the latest fork for llama.cpp does not seem to support the MTP head from Unsloth (it worked inside Unsloth Studio itself, not on the latest llama.cpp build - version: 0.4.0-dev (build 10837, commit 5202104b5). Can anyone tell me how to make it work?

The command I am using is this -

/llama.cpp/build/bin/llama-server \ -m "/home/[]/Local Models/AtomicChat/Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64-00001-of-00033.gguf" \ --mmproj "/home/[]/Local Models/AtomicChat/Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/mmproj-Qwen3.8-Flash-Next-F16.gguf" \ --alias qwen3.8-flash-next \ -c 180000 \ -ngl 99 \ --n-cpu-moe 36 \ -fit off \ --jinja \ --flash-attn on \ --cache-type-k q8_0 --cache-type-v q8_0 \ -t 8 \ --host 127.0.0.1 --port 8081 \ --cont-batching -np 1 \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --chat-template-file /home/[]/qwen38-template.jinja

tanpurohit26 changed discussion status to closed
tanpurohit26 changed discussion status to open

llama.cpp original master haven't merge qwen 4 architecture mtp layer yet,if you want use mtp,you need go https://github.com/unslothai/llama.cpp use his branch and download one of mtp draft in https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main/MTP,need additional VRAM/RAM,
startup code etc:
$md = "C:\Users\momori\Downloads\Qwen3.8-27B-DFlash2-Q4_K_M.gguf"
$model = "E:\Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf"
Start-Process -FilePath $exe -ArgumentList @(
"-md", $md,
"-m", $model,
"-c", "80000",
"--flash-attn", "on",
"--temp", "0.6",
"--top-p", "0.95",
"--top-k", "40",
"--min-p", "0.01",
"--repeat-penalty", "1.02",
"--presence-penalty", "0.0"
"-ctk", "q4_0", "-ctv", "q4_0",
"--batch-size", "400",
"--ubatch-size", "200",
"--threads", "24",
"--api-key", "123456",
"-rea", "on",
"--jinja",
"--cache-ram", "4000",
"--parallel", "1",
"--kv-unified",
"--no-warmup",
"--spec-type", "draft-mtp",
"--spec-draft-n-max", "5",
"--spec-draft-p-min", "0.84",
"--chat-template-file", $tpl,
"--load-mode", "mmap",
"--reasoning-preserve",
"--reasoning-format", "deepseek",
"--lazy-mode", "on",
"--reasoning-effort", "medium",
)

also MTP benefits from a higher batch size value. try 2048 or even 4096 if you got the VRAM space. Try different values for spec-draft-n-max and watch the draft acceptance in the log. (higher is better, aim for at least 0.75 average acceptance)

Sign up or log in to comment