Qwen3.8-Flash-Next NVFP4 — GGUF

A GGUF conversion of nvidia/Qwen3.8-Flash-Next-NVFP4, NVIDIA's NVFP4 quantization of Qwen/Qwen3.8-Flash-Next.

Routed experts in NVIDIA's NVFP4, everything else Q8_0, hence the name. The experts are not requantized.

This is an independent conversion. It is not made, reviewed or endorsed by NVIDIA or by Qwen.

What is different about this file and what is in it

176.944B parameters in 119.02 GiB (127,809,147,712 bytes), counted from its 1,512 tensors:

part parameters type size
routed experts (512 per layer, 48 layers) 120.796B NVFP4, with per-expert scales 63.28 GiB
dense: attention, GatedDeltaNet, shared experts, hyper-connections 4.312B Q8_0; norms F32 4.45 GiB
token embedding 0.636B Q8_0
hashed n-gram (PLE) table 51.200B Q8_0, FP8 scale restored 50.66 GiB

Per token: the 4.3B dense parameters, 10 of 512 experts per layer (2.4B), and 16 rows of the n-gram table. Not included: the MTP head (about 4B parameters) and the vision encoder, so this is a text-only model.

llama.cpp reports the file as Q8_0. Tools report the file as Q8_0: GGUF carries a single type label per file (general.file_type), which cannot express a mix.

Files

One file, as converted.

file size sha256
Qwen3.8-Flash-Next-NVFP4-Q8_0.gguf 127809147712 bytes 04124cb939a1ae968b53ce101b222a8eaf56bcb6742bae1976bede373b811a27

License

Governed by the NVIDIA Open Model License, as the source checkpoint is, and by the Qwen Community License 1.0 both texts are in this repository.

License: NVIDIA Open Model License.

Downloads last month
186
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for CompiledThoughts/Qwen3.8-Flash-Next-NVFP4-Q8_0

Quantized
(5)
this model