Text Generation
GGUF
qwen4exp
conversational

Qwen3.8 Flash Next for DwarfStar

Self-contained GGUFs for DwarfStar, with original BF16 n-grams read directly from disk. Keep the file on a local SSD. No n-gram sidecar is needed. MTP weights are included; enable them with --mtp. Requires DwarfStar's native BF16 n-gram reader. Older versions expecting --ple cannot load these files.

File File size Main/MTP weights
Qwen3.8-Flash-Next-Q2.gguf 137.10 GiB 41.73 GiB
Qwen3.8-Flash-Next-Q4.gguf 165.11 GiB 69.74 GiB

Each file contains the same 95.37 GiB BF16 n-gram table. It is not made resident or quantized. Context and runtime buffers require additional RAM.

The main/MTP tensor bytes come unchanged from Ivan Fioravanti's Q2 and Q4 releases. Q2 has imatrix-calibrated IQ2_XXS gate/up and Q2_K down routed experts, with the down rows padded from 640 to 768 inputs. Q4 has calibrated Q4_K gate/up and MXFP4 down. Other tensor formats and MTP weights are unchanged.

The n-grams are copied byte-for-byte from Qwen's original checkpoint at de4b8e4d43b917e7706784d8bb445c9af86a3540, including its padding rows. The packer checks the original hash constants and verifies every copied payload. These are quantized language-model weights, not lossless copies of the whole original model.

Original model: Qwen. Quantized main/MTP releases: Ivan Fioravanti. Native n-gram packaging and disk-only integration: Salvatore Sanfilippo / DwarfStar. The original Qwen Community License is included as LICENSE.

Downloads last month
59,052
GGUF
Model size
188B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for antirez/qwen3.8-flash-next-gguf

Quantized
(235)
this model