MiMo-V2.6 Pro MOPD GGUF for DwarfStar

GGUF files for the MiMo-V2.6 Pro MOPD port in DwarfStar (DS4), branch kernelpool/ds4:mimo-v26 (pull request: TODO link). Metal only. The model runs over tensor parallelism between two 512 GB Macs.

MOPD is Xiaomi's update of the MiMo-V2.6 Pro RL checkpoint that reduces repeated tool calls in agent sessions (Xiaomi's write-up). The main model and the DFlash drafter changed; the MTP blocks carry the same weights as the RL release. The files below replace kernelpool/MiMo-V2.6-Pro-RL-MXFP4-GGUF and use the same layout.

Files

The main GGUF is above the Hugging Face single-file limit, so it is stored in two parts and joined after download, the same way the official DwarfStar Q4 release is.

File Bytes SHA-256
MiMo-V2.6-Pro-MOPD-MXFP4.gguf.part1 480,000,000,000 10f205de98f64e656e1df1fef2394312a5fffd6e1f98ac1a7004fde8067926f0
MiMo-V2.6-Pro-MOPD-MXFP4.gguf.part2 77,146,510,560 5fe44280b42243caa7afc3e10062ebeed2081ea627738487e043119bb53aec37
MiMo-V2.6-Pro-MOPD-MXFP4.gguf (joined) 557,146,510,560 e6fd8b155d920db4e561314cca671902926c83e8997fae629b7ba42bfbd39bd4
MiMo-V2.6-Pro-MOPD-DFlash-Q8_0.gguf 2,941,587,648 58c9617f69545254cfce4a6cd8525d423d292793cbb3e3869a1f235c8ec7237c

The main file uses the mimo2 layout: the checkpoint's own MXFP4 experts repacked without requantization, Q8_0 attention, dense and output weights, BF16 embeddings, and the three MTP blocks. The DFlash sidecar uses the dflash layout plus the mask embedding and value scale DS4 needs.

hf download kernelpool/MiMo-V2.6-Pro-MOPD-MXFP4-GGUF --local-dir gguf
cd gguf
cat MiMo-V2.6-Pro-MOPD-MXFP4.gguf.part2 >> MiMo-V2.6-Pro-MOPD-MXFP4.gguf.part1
mv MiMo-V2.6-Pro-MOPD-MXFP4.gguf.part1 MiMo-V2.6-Pro-MOPD-MXFP4.gguf
rm MiMo-V2.6-Pro-MOPD-MXFP4.gguf.part2
openssl dgst -sha256 MiMo-V2.6-Pro-MOPD-MXFP4.gguf   # e6fd8b15...

Appending the second part to the first keeps the extra disk space to the size of the second part.

Run

Both Macs need the joined file locally. Each rank keeps half of the attention heads, half of the routed experts and half of the vocabulary head, about 270 GiB resident, so each Mac needs 512 GB. The link setup is in docs/DISTRIBUTED.md. Start the worker first:

git clone -b mimo-v26 https://github.com/kernelpool/ds4 && cd ds4 && make
# Machine B (worker)
./ds4 -m gguf/MiMo-V2.6-Pro-MOPD-MXFP4.gguf --ctx 8192 --mtp \
  --tensor-parallel --role worker --coordinator 192.168.0.1 9911 --transport rdma
# Machine A (coordinator)
./ds4-server -m gguf/MiMo-V2.6-Pro-MOPD-MXFP4.gguf --ctx 8192 --mtp \
  --tensor-parallel --role coordinator --listen 192.168.0.1 9911 --transport rdma

--mtp drafts with the MTP blocks inside the main file; both ranks run the draft and verify it together, so pass it to both or to neither. Greedy output follows plain decoding.

The server exposes mimo-v2.6-pro, mimo-v2.6-pro-chat (thinking off) and mimo-v2.6-pro-reasoner (thinking on).

DS4 does not run the DFlash sidecar or vision under tensor parallelism yet, so the sidecar is not used for now and no vision encoder is included. See docs/MIMO_V26.md.

Conversion

Written by gguf-tools/mimo26_quantize.py from XiaomiMiMo/MiMo-V2.6-Pro-MOPD at revision adea8e2c5373181e5a973fa1ecb343cb31af214b. The converter de-interleaves the tensor-parallel chunks of the fused QKV projection and keeps the released MXFP4 expert blocks bit for bit; the source revision is recorded in the GGUF metadata.

The original checkpoint is released by Xiaomi under the MIT license, which applies to these files as well.

Downloads last month
63
GGUF
Model size
3B params
Architecture
dflash
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kernelpool/MiMo-V2.6-Pro-MOPD-MXFP4-GGUF

Quantized
(3)
this model

Collection including kernelpool/MiMo-V2.6-Pro-MOPD-MXFP4-GGUF