YuE2-3B-NVFP4 (autoregressive checkpoint, W4A16 GPTQ)

This is an NVFP4 requantisation of the autoregressive (AR) half of m-a-p/YuE2-3B. It is packaged as the Qwen3-shaped checkpoint that the official yue2_infer pipeline derives internally for its vLLM backend, and it is a drop-in replacement for that derived checkpoint. The planning (abc) and semantic-token stages run on these weights. The NAR flow-matching stage and the VAE still run on the original bf16 and fp32 weights, which you need alongside this repo.

The point of it is decode speed on bandwidth-limited hardware. On a DGX Spark (GB10, 273 GB/s unified memory) the AR stages are limited by weight bandwidth, and the bf16 AR checkpoint is 4.33 GB.

bf16 (derived by yue2_infer) this checkpoint
AR weights on disk 4.33 GB 2.31 GB
Linear layers bf16 NVFP4 W4A16: e2m1 weights, e4m3 scales per 16 values, fp32 global scale per tensor, bf16 activations
embed_tokens and lm_head (vocab 184,704) bf16 bf16, not quantised
AR decode on DGX Spark, vLLM 0.27.1, batch 1 44.8 tok/s abc, 43.0 tok/s semantic 100 to 107 tok/s abc, 95 to 97 tok/s semantic (2.3x)

A full song on the same hardware and seed took 329 s with bf16 and 213 s with this checkpoint (142 s of audio). The AR stage went from 120 s to 54 s; the NAR and VAE stages are unchanged.

What was quantised and how

Source: model.safetensors from m-a-p/YuE2-3B (sha256 starting 1d55c42c). The AR tensors were extracted exactly as yue2.fast.derive_ar_checkpoint does: 311 tensors covering attention q/k/v/o and their q/k norms, the AR MLPs, layer norms, embed_tokens and lm_head. The NAR MLPs, the llm2vae and vae2llm projections, the timestep embedder and the VAE are not in this checkpoint and are not quantised.

Tool: llm-compressor 0.13.0 with compressed-tensors 0.18.0. Recipe:

GPTQModifier(targets=["Linear"], scheme="NVFP4A16", ignore=["lm_head"],
             dampening_frac=0.01, block_size=128, actorder="static")

Calibration used sequences the bf16 model generated itself, not a text corpus, because text never exercises the codec-token activations. The set was 32 requests (16 lyric sets, each rendered under two contrasting style prompts, English and Mandarin, cot=full and cot=melody), both the abc and semantic phases, cut into 58 windows of up to 8,192 tokens: 200K tokens in total, 51% of them codec tokens. Four requests were held out for evaluation.

Format: compressed-tensors nvfp4-pack-quantized. Loads in vLLM on Blackwell (SM100, SM120, SM121) with native NVFP4 kernels, and in transformers with the weights dequantised to bf16.

Quality against bf16

Measured through vLLM 0.27.1 on the four held-out requests, with every output position teacher-forced on the bf16 model's own sampled sequences: 6,286 abc positions and 17,931 semantic positions. The first row is two bf16 runs on different engine builds, which gives the numerical floor.

Top-1 agreement is a poor guide for the semantic phase. The codec distribution is nearly flat at most positions (bf16 itself puts its argmax on the token it sampled only 22% of the time), so a small shift flips the argmax without moving much probability. The last column, total-variation distance, is the share of sampling probability that actually moved, and is the number to read.

Checkpoint abc top-1 agree abc NLL semantic top-1 agree semantic NLL semantic TV
bf16 vs bf16, two engine builds 0.999 0.141 / 0.141 0.982 3.414 / 3.414 1.2%
this checkpoint, NVFP4 W4A16 GPTQ 0.996 0.141 / 0.142 0.905 3.414 / 3.424 (+0.3%) 6.3%
FP8 W8A8, sibling repo DKmode22/YuE2-3B-FP8 0.997 0.141 / 0.141 0.942 3.414 / 3.418 (+0.1%) 3.8%
NVFP4 W4A16, AWQ 0.993 0.141 / 0.143 0.896 3.414 / 3.433 (+0.6%) 7.4%
NVFP4 W4A4, GPTQ (not published) 0.992 0.141 / 0.144 0.854 3.414 / 3.446 (+0.9%) 10.1%

Split by how confident bf16 was at each position, this checkpoint agrees with bf16 at 99.8% of positions where the bf16 top-1 probability is above 0.5, and the disagreements sit in the near-tie positions (45% of semantic positions have a top-1 probability under 0.15). A per-layer sensitivity pass found no layer that matters more than the others (all 28 within 0.008 nats of a 0.005-nat noise floor), so none is kept in bf16. Streams were checked for repetition: no repeated 24-gram on any held-out output.

W4A4 was also built and measured. It is both less accurate and slower than W4A16 at batch 1 on this hardware (vLLM on SM12x runs dense NVFP4 through a W4A16 kernel and the activation quantisation only adds work), so it is not published.

How to use

The official pipeline derives its own bf16 AR checkpoint and does not yet accept an alternative. patch_fast.py in this repo adds a YUE2_AR_CHECKPOINT environment override to yue2/fast.py. It is idempotent and changes nothing else.

pip install "yue2_infer[fast]"        # the wheel from m-a-p/YuE2-3B
python patch_fast.py "$(python -c 'import yue2.fast,os;print(os.path.abspath(yue2.fast.__file__))')"
hf download DKmode22/YuE2-3B-NVFP4 --local-dir ./yue2-3b-nvfp4
export YUE2_AR_CHECKPOINT=$PWD/yue2-3b-nvfp4
python - <<'EOF'
from yue2 import YuE2Pipeline
pipe = YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", backend="vllm", memory_budget_gib=16)
song = pipe(style="Indie folk, warm acoustic guitar, male vocal", lyrics="[Verse]\n...")
song.save("song.flac")
EOF

Notes:

  • The pipeline uses vLLM only when cfg_scale is 1 (the default for cot=full and cot=melody) and quantization is "none". With cot=off the guidance is 1.01 and the pipeline silently falls back to its torch backend, which does not use this checkpoint.
  • On DGX Spark (sm_121) the vllm==0.19.0 PyPI wheel that the package pins has no working NVFP4 path: W4A4 fails at engine init with "no kernel image", and W4A16 runs but returns wrong logits. Use a vLLM build with SM121 support (0.27.1 was used here) and widen the version check in yue2/fast.py.
  • Everything else in the pipeline is unchanged: one sequence at a time, the windowed repetition penalty, 24,576 context.

Licence and attribution

The base model is released by the M-A-P team under CC-BY-NC-4.0. This derivative carries the same licence and is for non-commercial use. Credit for the model belongs to the YuE2 authors; please cite their paper (arXiv 2503.08638). Quantisation, calibration corpus and evaluation by DKmode22 on DGX Spark hardware, September 2026.

Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DKmode22/YuE2-3B-NVFP4

Base model

m-a-p/YuE2-3B
Quantized
(8)
this model