ornith-1.5-9b-affine4-vision-mlx

Compact affine 4-bit MLX quant with the parent model's BF16 vision tower. Built from Ornith-1.5-9B revision 98db59b for Apple Silicon.

Format

  • Size: 6.53 GB
  • Language trunk: 250 affine-4 modules, including recurrent inputs
  • Vision: 333 same-parent BF16 tensors
  • Context metadata: 262,144 tokens
  • Runtime: autoregressive MLX-LM/MLX-VLM
  • Native MTP: unavailable; the upstream checkpoint contains no mtp.* tensors

Exact precision and conversion records are included in BUILD_RECIPE.json and conversion_receipt.json. Immutable weight revision: 514b626.

Usage

python -m mlx_vlm.generate \
  --model Shiftedx/ornith-1.5-9b-affine4-vision-mlx \
  --image image.jpg \
  --prompt "Describe this image." \
  --max-tokens 256

Qualification

ShiftedX Bench v0.3.0 commit 3bbb0bfa01e33503163cb34ef52b4d507e456265, Apple M4 Max / 64 GiB, thinking enabled, medium reasoning, temperature 1, top-p 0.95, top-k 20, KV cache off:

Lane Result Mean decode Peak active memory
Quality 6/10 67.7 tok/s 5.60 GB
Long context 3/15 61.8 tok/s 5.60 GB
Tools 6/6 62.4 tok/s 6.70 GB
Agentic 1/2
Vision, native MLX-VLM 4/4 strict

Text lanes used MTPLX 2.7.1 in stock autoregressive mode. Vision used native MLX-VLM because MTPLX AR rejects image content. Structural loading and deterministic text/vision smokes passed. Quantization can change behavior; see the parent model card for intended use and license details.

Downloads last month
83
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shiftedx/ornith-1.5-9b-affine4-vision-mlx

Quantized
(38)
this model

Collection including Shiftedx/ornith-1.5-9b-affine4-vision-mlx