llama.cpp with Victoria draft-head support
This is stock llama.cpp (upstream build b11276) plus one patch that adds the draft head (NextN/MTP) for the qwen4exp architecture. With it, the Victoria GGUF files load and can use their draft head for speculative decoding.
Unpatched llama.cpp stops with wrong number of tensors; expected 1256, got 1224 on these files. We plan to submit the patch to upstream llama.cpp. When a normal llama.cpp release supports the head, use that instead of this repo.
The source is the qwen4exp-draft-mtp branch of https://github.com/rmonsurate/llama.cpp. The zips were built by llama.cpp's own release workflow on GitHub Actions, from a branch that adds only a workflow change on top of that source. Because of that extra commit, llama-server --version reports build 11278, commit 0ed8358. The same files are on the GitHub release: https://github.com/rmonsurate/llama.cpp/releases/tag/victoria-mtp-b11276
Which file to download
| Platform | File |
|---|---|
| Windows x64, NVIDIA GPU | llama-victoria-mtp-b11276-bin-win-cuda-12.4-x64.zip or llama-victoria-mtp-b11276-bin-win-cuda-13.4-x64.zip, plus the matching cudart-llama-bin-win-cuda-*-x64.zip if you do not have the CUDA runtime installed |
| Windows x64, CPU only | llama-victoria-mtp-b11276-bin-win-cpu-x64.zip |
| Windows arm64 | llama-victoria-mtp-b11276-bin-win-cpu-arm64.zip or llama-victoria-mtp-b11276-bin-win-cuda-13.4-arm64.zip |
| Linux x64, NVIDIA GPU | llama-victoria-mtp-b11276-bin-ubuntu-cuda-12.8-x64.tar.gz or llama-victoria-mtp-b11276-bin-ubuntu-cuda-13.4-x64.tar.gz, plus the matching cudart-...tar.gz if you do not have the CUDA runtime installed |
| Linux arm64, NVIDIA GPU | llama-victoria-mtp-b11276-bin-ubuntu-cuda-13.4-arm64.tar.gz |
| Linux x64 or arm64, CPU only | llama-victoria-mtp-b11276-bin-ubuntu-x64.tar.gz or llama-victoria-mtp-b11276-bin-ubuntu-arm64.tar.gz |
| macOS, Apple Silicon (Metal) | llama-victoria-mtp-b11276-bin-macos-arm64.tar.gz |
| macOS, Intel (CPU only) | llama-victoria-mtp-b11276-bin-macos-x64.tar.gz |
For CUDA 12.4 on Windows you need a driver that supports CUDA 12.4 or newer. For 13.4 you need a CUDA 13 driver.
Quickstart
Download the smaller Victoria GGUF set (107.20 GB):
hf download rmonsurate/Victoria --include "gguf/victoria-s410-bitexact-tbl8-*" --local-dir .
Run it with the draft head on:
llama-server -m gguf/victoria-s410-bitexact-tbl8-00001-of-00003.gguf -ngl 99 -fa on -c 8192 --spec-type draft-mtp --spec-draft-n-max 3
Run it with the draft head off:
llama-server -m gguf/victoria-s410-bitexact-tbl8-00001-of-00003.gguf -ngl 99 -fa on -c 8192
With the head off, llama.cpp prints a warning for each of the 32 unused head tensors and skips them. That is expected.
On an M3 Max with 128 GB, one request at a time, the head raised generation speed from 26.8 to 27.8 tok/s (head off, two runs) to 34.3 to 38.0 tok/s (head on, --spec-draft-n-max 3, two runs). 70.4% of drafted tokens were accepted. We have not measured CUDA yet.
Three ways to get it
Download a prebuilt zip from the table above.
Clone the branch and build it:
git clone -b qwen4exp-draft-mtp https://github.com/rmonsurate/llama.cpp
cd llama.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
Add -DGGML_CUDA=ON to the first cmake command for NVIDIA GPUs. On a Mac, Metal is on by default.
- Apply the patch to your own llama.cpp checkout. The patch is
victoria-draft-mtp-b11276.patchin this repo and on the GitHub release. It was made against upstream commit 19e28a277 (b11276):
git fetch origin && git checkout 19e28a277
git am victoria-draft-mtp-b11276.patch
It may also apply to newer upstream commits.
License
llama.cpp is MIT licensed. See LICENSE in each zip.