MiMo-V2.6 Flash MOPD GGUF (Q4, Q2) for DwarfStar

Requantized GGUF files for the MiMo-V2.6 Flash port in DwarfStar (DS4), branch kernelpool/ds4:mimo-v26. Metal only.

MOPD is Xiaomi's update of the MiMo-V2.6 Flash RL checkpoint that reduces repeated tool calls in agent sessions (Xiaomi's write-up). The checkpoint ships its routed experts in MXFP4; the lossless MXFP4 file, the DFlash drafter and the vision encoder are in kernelpool/MiMo-V2.6-Flash-MOPD-MXFP4-GGUF.

File Size Content
MiMo-V2.6-Flash-MOPD-Q4.gguf 166.2 GiB Main model, mimo2 layout: Q4_K experts (gate, up and down) requantized from the released MXFP4 experts with an importance matrix, Q8_0 attention, dense and output weights, BF16 embeddings, the three MTP blocks
MiMo-V2.6-Flash-MOPD-Q2.gguf 86.9 GiB Main model with IQ2_XXS gate/up and Q2_K down experts, requantized the same way; all other tensors as in the Q4 file
imatrix/MiMo-V2.6-Flash-MOPD-routed-moe-ds4pool.dat 0.5 GiB The routed-expert importance matrix both files were quantized with

Run

git clone -b mimo-v26 https://github.com/kernelpool/ds4 && cd ds4 && make
./ds4 -m MiMo-V2.6-Flash-MOPD-Q2.gguf --mtp
./ds4 -m MiMo-V2.6-Flash-MOPD-Q2.gguf --mtp --vision MiMo-V2.6-Flash-MOPD-Vision-F32.gguf
./ds4-server -m MiMo-V2.6-Flash-MOPD-Q2.gguf --mtp --vision MiMo-V2.6-Flash-MOPD-Vision-F32.gguf

--mtp drafts with the MTP blocks inside the main file and is the faster option; --mtp-model MiMo-V2.6-Flash-MOPD-DFlash-Q8_0.gguf uses the DFlash sidecar instead. Both verify against the target, so temperature-zero output follows plain decoding. The vision encoder and DFlash files come from the MXFP4 repository linked above and work with both files here. Resident sizes are 166.2 GiB for Q4 and 86.9 GiB for Q2; Q2 is the tier for 128 GB Macs. See docs/MIMO_V26.md.

Conversion

Written by gguf-tools/mimo26_quantize.py (--quant q4 and --quant q2) from XiaomiMiMo/MiMo-V2.6-Flash-MOPD at revision 2479e2d0029eca9a34cc7e7f55a121925f81908e. The experts are requantized from the released MXFP4 blocks using the importance matrix above, which DS4 collected on this checkpoint over its calibration prompts (gguf-tools/imatrix/dataset). The source revision is recorded in the GGUF metadata.

Quality

Scored on 100 official continuations from the Xiaomi platform with the fixture in gguf-tools/quality-testing/mimo-v2.6-flash-20260922: the Q4 file scores on par with the MXFP4 file, and the Q2 file trades some accuracy for its smaller size.

The original checkpoint is released by Xiaomi under the MIT license, which applies to these files as well.

Downloads last month
261
GGUF
Model size
310B params
Architecture
mimo2
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kernelpool/MiMo-V2.6-Flash-MOPD-GGUF

Quantized
(23)
this model

Collection including kernelpool/MiMo-V2.6-Flash-MOPD-GGUF