Qwen3.8-27B-Fable-Distill-mlx-uniform-4bit

27B (VLM) parameters — note: Hugging Face's size badge undercounts packed 4-bit MLX weights (it counts the packed uint32 tensors), so the number shown beside this repo is wrong; the figure here is the true parameter count.

MLX uniform 4-bit quant of TeichAI/Qwen3.8-27B-Fable-Distill.

measured value
effective bits/weight 4.0 (uniform; all 498 quantized layers at 4-bit)
weights footprint 16.05 GB
quantized-layer bit histogram 4-bit: 498

Recommended sampling (measured, not vibes)

param value
temperature 0.6 (certified by a per-model temperature ladder)
top_p / top_k / min_p 0.95 / 20 / 0.0
presence_penalty 0.0
max_tokens / thinking_budget 102400 / 81920 (thinking ON)

These values were certified by an execution-gated benchmark campaign (temperature ladders with convergence gates over HumanEval+/MBPP+ and agentic harnesses) — methodology and full results: https://github.com/ivan-avramov/mlx_local_stack.

Serving: MLX (mlx-lm / mlx-vlm). Quantized on-device with mlx_lm.convert (uniform) or mlx_optiq (mixed-precision KL-sensitivity recipes).

Downloads last month
245
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for caslca/Qwen3.8-27B-Fable-Distill-mlx-uniform-4bit

Base model

Qwen/Qwen3.8-27B
Quantized
(8)
this model

Datasets used to train caslca/Qwen3.8-27B-Fable-Distill-mlx-uniform-4bit