GLM-4.5-Air (Terion in-house mixed_4_6 MLX quant)

In-house MLX quantization of zai-org/GLM-4.5-Air, produced by Terion for internal evaluation and republished for the community.

  • Method: mlx_lm.convert with --quant-predicate mixed_4_6 (mixed 4/6-bit precision) plus a periodic Metal-cache-clear wrapper during the save step to avoid unified-memory pressure crashes on Apple Silicon.
  • Actual bits-per-weight: 4.809
  • Size: ~60GB
  • Base model: zai-org/GLM-4.5-Air (see base repo for architecture/license/benchmark details)

Use with mlx-lm / Rapid-MLX. This is a straightforward quantization of the original weights.

Downloads last month
-
Safetensors
Model size
107B params
Tensor type
U32
BF16
F32
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for terion-mlx/GLM-4.5-Air-mixed_4_6

Quantized
(65)
this model