lukaskremla/Ornith-1.5-35B-A3B-MTP-4bit-MLX

This repository contains a 4-bit affine quantization of the native Multi-Token Prediction (MTP) tensors extracted from ornith-ai/Ornith-1.5-35B-A3B and converted into the standalone MLX drafter format expected by mlx-vlm.

This is not a standalone language model. It must be loaded as the draft model alongside a compatible Ornith-1.5-35B-A3B target model.

This 4-bit drafter can be used with the 2-bit, 3-bit, 4-bit, 5-bit, 6-bit and 8-bit MLX target models. Quantizations don't need to match.

You can download the vision capable or text-only Ornith 1.5 35B-A3B weights in a variety of MLX quantizations from this collection Ornith 1.5 35B-A3B MLX-Quants (Vision, Text-Only & MTP).

Hugging Face might render incorrect parameter counts for this model, it is a common display bug for MLX quants.

Credits

Converted to MLX format from ornith-ai/Ornith-1.5-35B-A3B using mlx-vlm version 0.6.17.

Downloads last month
65
Safetensors
Model size
0.8B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including lukaskremla/Ornith-1.5-35B-A3B-MTP-4bit-MLX