MiniMax-M2.7-speedy-colibri-nvfp4

NVFP4 container for colibrì, a Rust MoE inference engine for single-box streaming-expert inference.

This repository exists so the model can be downloaded and run directly, with no conversion step. Converting the upstream checkpoint requires holding the source and the container on disk at the same time; this container is the finished artefact.

Use

huggingface-cli download Kanposer/MiniMax-M2.7-speedy-colibri-nvfp4 --local-dir MiniMax-M2.7-container
coli gen MiniMax-M2.7-container "<prompt tokens>"

What was changed

Weights were repacked into colibrì's container layout. Routed experts are NVFP4 as published upstream; the container repacks them into coalesced per-expert spans for streaming. No fine-tuning, distillation, or other modification of model behaviour was performed.

Licence and attribution

This is a derivative of nvidia/MiniMax-M2.7-NVFP4 and is distributed under the upstream licence, nvidia-software-and-model-evaluation-license (licence text). The upstream publishes its licence as a URL rather than a file, so it is linked above rather than copied here.

All rights in the original model remain with its authors. Please read and comply with the upstream licence before downloading or using these weights.

Downloads last month
26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kanposer/MiniMax-M2.7-speedy-colibri-nvfp4

Quantized
(120)
this model