Qwen3.6 MTP OQE FP16
Collection
8 items • Updated
How to use wezzel98765/Qwen3.6-35B-A3B-oQ4e-fp16-mtp with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Qwen3.6-35B-A3B-oQ4e-fp16-mtp wezzel98765/Qwen3.6-35B-A3B-oQ4e-fp16-mtp
Text and vision retained
Quantised using oMLX v0.5.0.rc1 OQ Enhanced quantization (oQe) iMatrix
MTP Heads retained
FP16 is fastest on M1/M2 , but this can work on all MLX inferencing systems
This model is using the LATEST FROGGERIC chat template upgrade (Fixed jinja chat templates for Qwen 3.5 & 3.6 (v21))
4-bit
Base model
Qwen/Qwen3.6-35B-A3B