Qwen3.8-27B-mlx-uniform-4bit

⚠️ Identity note (2026-08-26): the language trunk of this model is bit-identical to mlx-community/Qwen3.8-27B-4bit. A full-tensor md5 sweep over both repos shows 2179 of 2180 tensors identical; the single differing tensor is vision_tower.patch_embed.proj.weight, because this copy carries the restored bf16 vision tower (see below) while the mlx-community upload keeps the original. MLX uniform 4-bit (group size 64) quantization is deterministic, so two independent conversions of the same bf16 base produce the same weights. Text results measured against either repo apply to both; prefer this copy if you want the vision path.

27B parameters, vision tower included (restored 2026-08-23, see below) — note: Hugging Face's size badge undercounts packed 4-bit MLX weights (it counts the packed uint32 tensors), so the number shown beside this repo is wrong; the figure here is the true parameter count.

MLX uniform 4-bit quant of unsloth/Qwen3.8-27B.

measured value
effective bits/weight 4.0 (uniform; all 498 quantized layers at 4-bit)
weights footprint 15.13 GB
quantized-layer bit histogram 4-bit: 498

Recommended sampling (measured, not vibes)

param value
temperature 0.6 (certified by a per-model temperature ladder)
top_p / top_k / min_p 0.95 / 20 / 0.0
presence_penalty 0.0
max_tokens / thinking_budget 102400 / 81920 (thinking ON)

These values were certified by an execution-gated benchmark campaign (temperature ladders with convergence gates over HumanEval+/MBPP+ and agentic harnesses) — methodology and full results: https://github.com/ivan-avramov/mlx_local_stack.

Serving: MLX (mlx-lm / mlx-vlm). Quantized on-device with mlx_lm.convert (uniform) or mlx_optiq (mixed-precision KL-sensitivity recipes).

Native MTP drafter available (2026-08-26)

This checkpoint ships a native multi-token-prediction sidecar (optiq/mtp.safetensors, 29 tensors, outside the weight index — inert at normal load). A standalone, servable extraction is published at caslca/Qwen3.8-27B-mlx-uniform-4bit-mtp-drafter: measured 1.46–1.58× decode at 68–76% acceptance (probe-only — no quality certification; see that card's caveats).

Vision tower restored (2026-08-23)

The original conversion was language-model-only. This revision grafts the vision tower back from the upstream base (unsloth repackaging of the family release): 333 vision_tower.* tensors kept bf16 (exactly what the vision-retaining mlx_vlm convert produces for this family), +0.92 GB.

The text trunk is bit-identical to the evaluated artifact: the trunk shards are byte-copies (md5-verified), and a fixed-token forward pass through the language model produces bit-identical logits pre/post graft. Every benchmark number on this card measures exactly the weights this revision serves for text. One-image smoke passed post-graft.

Downloads last month
595
Safetensors
Model size
5B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for caslca/Qwen3.8-27B-mlx-uniform-4bit

Base model

Qwen/Qwen3.8-27B
Quantized
(3)
this model
Quantizations
1 model