llama.cpp with Victoria draft-head support

This is stock llama.cpp (upstream build b11276) plus one patch that adds the draft head (NextN/MTP) for the qwen4exp architecture. With it, the Victoria GGUF files load and can use their draft head for speculative decoding.

Unpatched llama.cpp stops with wrong number of tensors; expected 1256, got 1224 on these files. We plan to submit the patch to upstream llama.cpp. When a normal llama.cpp release supports the head, use that instead of this repo.

The source is the qwen4exp-draft-mtp branch of https://github.com/rmonsurate/llama.cpp. The zips were built by llama.cpp's own release workflow on GitHub Actions, from a branch that adds only a workflow change on top of that source. Because of that extra commit, llama-server --version reports build 11278, commit 0ed8358. The same files are on the GitHub release: https://github.com/rmonsurate/llama.cpp/releases/tag/victoria-mtp-b11276

Which file to download

Platform File
Windows x64, NVIDIA GPU llama-victoria-mtp-b11276-bin-win-cuda-12.4-x64.zip or llama-victoria-mtp-b11276-bin-win-cuda-13.4-x64.zip, plus the matching cudart-llama-bin-win-cuda-*-x64.zip if you do not have the CUDA runtime installed
Windows x64, CPU only llama-victoria-mtp-b11276-bin-win-cpu-x64.zip
Windows arm64 llama-victoria-mtp-b11276-bin-win-cpu-arm64.zip or llama-victoria-mtp-b11276-bin-win-cuda-13.4-arm64.zip
Linux x64, NVIDIA GPU llama-victoria-mtp-b11276-bin-ubuntu-cuda-12.8-x64.tar.gz or llama-victoria-mtp-b11276-bin-ubuntu-cuda-13.4-x64.tar.gz, plus the matching cudart-...tar.gz if you do not have the CUDA runtime installed
Linux arm64, NVIDIA GPU llama-victoria-mtp-b11276-bin-ubuntu-cuda-13.4-arm64.tar.gz
Linux x64 or arm64, CPU only llama-victoria-mtp-b11276-bin-ubuntu-x64.tar.gz or llama-victoria-mtp-b11276-bin-ubuntu-arm64.tar.gz
macOS, Apple Silicon (Metal) llama-victoria-mtp-b11276-bin-macos-arm64.tar.gz
macOS, Intel (CPU only) llama-victoria-mtp-b11276-bin-macos-x64.tar.gz

For CUDA 12.4 on Windows you need a driver that supports CUDA 12.4 or newer. For 13.4 you need a CUDA 13 driver.

Quickstart

Download the smaller Victoria GGUF set (107.20 GB):

hf download rmonsurate/Victoria --include "gguf/victoria-s410-bitexact-tbl8-*" --local-dir .

Run it with the draft head on:

llama-server -m gguf/victoria-s410-bitexact-tbl8-00001-of-00003.gguf -ngl 99 -fa on -c 8192 --spec-type draft-mtp --spec-draft-n-max 3

Run it with the draft head off:

llama-server -m gguf/victoria-s410-bitexact-tbl8-00001-of-00003.gguf -ngl 99 -fa on -c 8192

With the head off, llama.cpp prints a warning for each of the 32 unused head tensors and skips them. That is expected.

On an M3 Max with 128 GB, one request at a time, the head raised generation speed from 26.8 to 27.8 tok/s (head off, two runs) to 34.3 to 38.0 tok/s (head on, --spec-draft-n-max 3, two runs). 70.4% of drafted tokens were accepted. We have not measured CUDA yet.

Three ways to get it

  1. Download a prebuilt zip from the table above.

  2. Clone the branch and build it:

git clone -b qwen4exp-draft-mtp https://github.com/rmonsurate/llama.cpp
cd llama.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Add -DGGML_CUDA=ON to the first cmake command for NVIDIA GPUs. On a Mac, Metal is on by default.

  1. Apply the patch to your own llama.cpp checkout. The patch is victoria-draft-mtp-b11276.patch in this repo and on the GitHub release. It was made against upstream commit 19e28a277 (b11276):
git fetch origin && git checkout 19e28a277
git am victoria-draft-mtp-b11276.patch

It may also apply to newer upstream commits.

License

llama.cpp is MIT licensed. See LICENSE in each zip.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support