ToPo-ToPo/Qwen3.6-27B-MTP-bf16

MTP (multi-token prediction) drafter for speculative decoding with Qwen/Qwen3.6-27B on Apple Silicon (mlx-vlm).

This is not a standalone chat model — it only works bound to a Qwen3.6-27B target model.

Provenance

  • Split from the official checkpoint's built-in mtp.* weights (15 tensors) with mlx-vlm 0.6.13: python -m mlx_vlm.speculative.drafters.qwen3_5_mtp.split --model Qwen/Qwen3.6-27B --output .
  • Precision: bf16 (unquantized, split as-is)
  • Size: 829 MB, block_size: 3, model_type: qwen3_5_mtp

Usage

mlx_vlm.generate --model ToPo-ToPo/Qwen3.6-27B-mlx-4bit \
  --draft-model ToPo-ToPo/Qwen3.6-27B-MTP-bf16 --draft-kind mtp \
  --prompt "..." --max-tokens 400

Quantized targets share this drafter (under greedy decoding the drafter's argmax rarely changes with target quantization). Speedup depends on hardware, target quantization, workload and the mlx-vlm version — measure it yourself.

Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ToPo-ToPo/Qwen3.6-27B-MTP-bf16

Base model

Qwen/Qwen3.6-27B
Finetuned
(372)
this model