Ornith-1.5-35B-A3B · MLX 4-bit (oQ4e) + MTP

MLX quantization of ornith-ai/Ornith-1.5-35B-A3B — a Mixture-of-Experts model with ~3B active parameters — in 4-bit (oQ4e) with Multi-Token Prediction (MTP, speculative depth 3), built to run on Apple Silicon via a local oMLX server.

Status: weights upload in progress — this card ships first; file list and hashes will land with the weights.

Credits

All credit for the base model goes to the Ornith team (ornith-ai):

This repo is only a quantization for local Apple Silicon inference. Same MIT terms apply.

Intended use

Local inference on Apple Silicon Macs (oMLX / mlx-lm). For benchmarks and model details, see the upstream card.

Downloads last month
182
Safetensors
Model size
36B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LookUpMark/Ornith-1.5-35B-A3B-oQ4e-mtp

Quantized
(134)
this model

Collection including LookUpMark/Ornith-1.5-35B-A3B-oQ4e-mtp