Qwen3.8-27B-static-mixed-4bit

27B parameters — text-only in THIS conversion (the upstream base is a VLM; the vision tower was not retained by the text-only convert path) — Hugging Face's size badge undercounts packed 4-bit MLX weights.

MLX static mixed-precision 4-bit quant of unsloth/Qwen3.8-27B. Measured weights footprint: 13.33 GB.

Recommended sampling from a per-model temperature ladder: temperature 0.4 (this recipe failed convergence screens at t0.6; a capped scan located t0.4), top_p 0.95, top_k 20, min_p 0.0, presence_penalty 0.0, thinking ON (budget 81920). Campaign methodology and results: https://github.com/ivan-avramov/mlx_local_stack.

Vision tower restored (2026-08-23)

The original conversion was language-model-only. This revision grafts the vision tower back from the upstream base (unsloth repackaging of the family release): 333 vision_tower.* tensors kept bf16 (exactly what the vision-retaining mlx_vlm convert produces for this family), +0.92 GB.

The text trunk is bit-identical to the evaluated artifact: the trunk shards are byte-copies (md5-verified), and a fixed-token forward pass through the language model produces bit-identical logits pre/post graft. Every benchmark number on this card measures exactly the weights this revision serves for text. One-image smoke passed post-graft.

Downloads last month
214
Safetensors
Model size
4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for caslca/Qwen3.8-27B-static-mixed-4bit

Base model

Qwen/Qwen3.8-27B
Quantized
(3)
this model