MiMo-V2.6-Distill-Qwen-9B 4-bit (MLX affine)

4-bit MLX build of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (Qwen3_5ForConditionalGeneration, qwen3_5), including the vision tower for image-text-to-text.

Quantization

  • Affine, 4-bit (MLX mx.quantize): packed u32 weights plus one bf16 scale/bias pair per group of weights along the input dim.
  • Group size 64, effective 5.059 bits per weight across the whole model (the vision tower adds non-quantized parameters, raising the average).
  • Only linear layers are quantized; embeddings, norms, the patch embedder and the GatedDeltaNet convolution stay at their original precision.

Contents

config.json                     qwen3_5, text_config + vision_config, quantization {group_size 64, bits 4}
model-00001-of-00002.safetensors  language_model (4-bit, ~5.0 GB)
model-00002-of-00002.safetensors  vision_tower (~0.6 GB)
model.safetensors.index.json    weight_map
preprocessor_config.json        image processor
processor_config.json           Qwen3VLProcessor
video_preprocessor_config.json  video processor
tokenizer.json …                tokenizer + chat_template.jinja

The released checkpoint contains no MTP weights, so there is no draft head.

Usage

mlx_vlm.generate \
  --model agnosticeng/MiMo-V2.6-Distill-Qwen-9B-4bit \
  --image photo.jpg \
  --prompt "Describe this image in one sentence." \
  --max-tokens 256
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

model, processor = load("agnosticeng/MiMo-V2.6-Distill-Qwen-9B-4bit")
config = load_config("agnosticeng/MiMo-V2.6-Distill-Qwen-9B-4bit")

prompt = apply_chat_template(processor, config, "Describe this image.", num_images=1)
print(generate(model, processor, prompt, image="photo.jpg", max_tokens=256))

Text-only prompts work the same without --image.

Conversion

Built from the public bf16 original with mlx_vlm.convert -q --q-bits 4 --q-group-size 64. Requires mlx-vlm with qwen3_5 support (v0.6.16+).

Validation

Smoke-tested on Apple Silicon: text and image prompts both generate correctly (e.g. the Statue of Liberty test image is described accurately). Quantization error not independently benchmarked.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for agnosticeng/MiMo-V2.6-Distill-Qwen-9B-4bit

Finetuned
Qwen/Qwen3.5-9B
Quantized
(53)
this model