MTP support for AtomicChat Qwen3.8-Flash-Next GGUF

#12
by aydintb - opened

Hi, I'm using your Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64 GGUF with llama.cpp.

My setup is:

  • RTX 5090 Laptop GPU β€” 24 GB VRAM
  • 64 GB system RAM
  • llama.cpp b10713 / commit 557614e02

The main model loads and runs correctly at around 11 tokens/sec. I'm trying to enable the model's 4B MTP layer for speculative decoding.

I tried using a separate Qwen3.8-Flash-Next-MTP-Q8_0.gguf, but llama.cpp fails when loading the MTP model with:

tensor 'blk.0.hc_attn_norm.weight' not found

The main AtomicChat GGUF itself loads and works normally.

Does the AtomicChat quant require a particular MTP GGUF, or is there a specific MTP file, conversion, llama.cpp branch, or configuration that you recommend for these AtomicChat quants?

Also, with my 24 GB VRAM + 64 GB RAM setup, would you expect the 4B MTP layer to provide a meaningful generation-speed improvement?

Thanks!

The error tensor 'blk.0.hc_attn_norm.weight' not found looks like the draft-load path rather than these quants.

Two things that may unblock you β€” note I have read the PRs but not tested MTP myself:

--spec-type draft-mtp for qwen4exp is not in main yet. It is #27836 (open, follow-up to #27742 which added the architecture). Its companion #28097, opened today, exists specifically for this failure: it adds draft-head-only GGUF support and fixes a draft-load regression, because #27836's loader requires hc_head_norm/hc_head_down/hc_head_up and the PLE block.

A correction to #6, where someone pointed at PR 27739 β€” that one and #27793 were discarded; #27742 is the merged one.

On the underlying request: these quants ship no draft head, so even with #27836 there is nothing to pair them with. @AtomicChat, would you consider publishing an MTP/NextN draft head alongside the M64 variants? #27836 includes converter support via --mtp.

Need MTP support . thanks!

https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/blob/main/MTP/README.md

that works, here is my bash file:

#!/bin/bash

# Resolve model path relative to this script's own directory, so it works from anywhere
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
MODEL="$SCRIPT_DIR/Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64-00001-of-00033.gguf"
MMPROJ="$SCRIPT_DIR/mmproj-Qwen3.8-Flash-Next-BF16.gguf"
MPTFILE="$SCRIPT_DIR/mtp-Qwen3.8-Flash-Next-BF16.gguf"

/home/atb/Projects/llama.cpp.mtp/llama.cpp/build/bin/llama-server \
  -m "$MODEL" \
  --mmproj "$MMPROJ" \
  -ngl 99 \
  --n-cpu-moe 999 \
  -fa on \
  --parallel 1 \
  --port 8080 \
  --host 0.0.0.0 \
  --ctx-size $((32*8*1024)) \
  --no-warmup \
  --temp 0.2 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.0 \
  --presence-penalty 0.0 \
  --repeat-penalty 1.0 \
  --seed 1902 \
  --log-colors on \
  --prio 2 \
  --jinja \
  --webui-mcp-proxy \
  --spec-type draft-mtp \
  --model-draft "$MPTFILE" \
  --spec-draft-n-min 1 \
  --spec-draft-n-max 2

Compared with the original version, how many tokens does MTP improve by?

Sign up or log in to comment