Ornith-1.0-35B-MLX-3bit-MTP

An Apple-Silicon MLX build of deepreinforce-ai/Ornith-1.0-35B that combines the 3-bit multimodal backbone from mlx-community/Ornith-1.0-35B-3bit with an experimental one-layer MTP head extracted from georgeis55/Ornith-1.0-35B-MLX-oQ8-mtp.

This checkpoint is prepared for Lightning MTP speculative decoding in oMLX. The original vision encoder and VLM configuration are retained.

Technical layout

  • Architecture: Qwen3.5 MoE VLM, 40 backbone layers, 256 experts
  • Backbone: affine 3-bit MLX quantization, group size 64 (3.662 bits/weight)
  • MTP: one full-attention MoE prediction layer at language_model.mtp.*
  • MTP linears: affine 8-bit, mixed group sizes 64 and 128
  • MTP fusion projection and normalization weights: BF16
  • MTP file: model-mtp-oq8.safetensors (905,114,024 bytes)
  • MTP SHA-256: 26358039bfc747625e8d476fede0cf619401a2ca5da9a03898d9b0c282f65c40
  • Tested runtime: oMLX 0.5.3 on Apple Silicon

The weight index contains all 42 MTP tensor mappings. The configuration declares one next-token prediction layer and includes per-module quantization overrides so the 8-bit MTP tensors are not interpreted using the 3-bit backbone settings.

Download and use with oMLX

hf download wd01216-bit/Ornith-1.0-35B-MLX-3bit-MTP \
  --local-dir ~/.omlx/models/wd01216-bit/Ornith-1.0-35B-MLX-3bit-MTP

In oMLX, open the model settings and enable Lightning MTP. The model identifier will be Ornith-1.0-35B-MLX-3bit-MTP when downloaded into the directory above.

Validation

The final checkpoint was discovered and loaded successfully by oMLX 0.5.3 as a multimodal model with Native MTP active. Two short text smoke tests confirmed that the MTP decode path was used, with observed draft acceptance of 15/16 and 6/6. These tiny tests verify wiring and execution only; they are not representative benchmarks.

Important limitation

The MTP head is a verbatim graft originally sourced from a Qwopus3.6 35B model. It was not trained on Ornith hidden states, and the donor bundle was prepared around an oQ8 backbone rather than this 3-bit backbone. Acceptance rate, output quality, memory use, and speedup are workload-dependent and may be lower than those of a native jointly trained MTP head. Disable Lightning MTP if a workload shows poor acceptance or regressions.

Provenance

  • Upstream model: deepreinforce-ai/Ornith-1.0-35B
  • Quantized multimodal backbone: mlx-community/Ornith-1.0-35B-3bit
  • MTP donor bundle and conversion notes: georgeis55/Ornith-1.0-35B-MLX-oQ8-mtp
  • Adapter assembly and oMLX validation: 2026-07-27

License

The upstream Ornith model is released under the MIT license. Refer to the upstream model card and license for intended use, limitations, and attribution requirements.

Downloads last month
343
Safetensors
Model size
5B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wd01216-bit/Ornith-1.0-35B-MLX-3bit-MTP

Quantized
(171)
this model