Ornith-1.0-35B-mlx-uniform-4bit

35B (hybrid linear-attention MoE, 256 experts / 8 active) parameters — note: Hugging Face's size badge undercounts packed 4-bit MLX weights (it counts the packed uint32 tensors), so the number shown beside this repo is wrong; the figure here is the true parameter count.

MLX uniform 4-bit (affine) quant of deepreinforce-ai/Ornith-1.0-35B (MIT).

measured value
effective bits/weight 4.019 (measured; 80 of 512 quantized layers at 8-bit)
weights footprint 20.4 GB
quantized-layer bit histogram 8-bit: 80 · 4-bit: 432

Recommended sampling (measured, not vibes)

param value
temperature 0.4 (certified by a per-model temperature ladder)
top_p / top_k / min_p 0.95 / 20 / 0.0
presence_penalty 0.0
max_tokens / thinking_budget 102400 / 81920 (thinking ON)

These values were certified by an execution-gated benchmark campaign (temperature ladders with convergence gates over HumanEval+/MBPP+ and agentic harnesses) — methodology and full results: https://github.com/ivan-avramov/mlx_local_stack.

Earlier revisions of this card stated ~4.649 bpw; the measured value from the serving manifest is 4.019 — corrected 2026-08-23.

Serving: MLX (mlx-lm / mlx-vlm). Quantized on-device with mlx_lm.convert (uniform) or mlx_optiq (mixed-precision KL-sensitivity recipes).

Downloads last month
114
Safetensors
Model size
6B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for caslca/Ornith-1.0-35B-mlx-uniform-4bit

Quantized
(181)
this model