mlx-community/GLM-5.3-mixed-4_5bit

This model mlx-community/GLM-5.3-mixed-4_5bit was converted to MLX format from zai-org/GLM-5.3-BF16 using mlx-lm version 0.31.3 (with PR #1410).

Note that this quant is using the GLM-5.3-BF16 as base.

Testing various quant recipes, these often start to overthink and redoing "decisions". Sadly, perplexity and KLD calculations don't always tell the full story. Getting the mix of the recipe right involves:

  • Ensuring that the model does not overthink and redo work, leading to double the token usage
  • Keeping more knowledge in the experts
  • Protect sensitive layers
  • Getting the quant to clearly understand the instructions
  • Not slowing down the processing
  • Keeping space for enough context at max reasoning

This is created for people using a single Apple Mac Studio M3 Ultra with 512 GB. The 4-bit version of GLM-5.3 fits comfortably.

You can find more similar MLX model quants for Apple Mac Studio with 512 GB at https://huggingface.co/bibproj

pip install mlx-lm

mlx_lm.generate --model mlx-community/GLM-5.3-mixed-4_5bit --prompt "Hi"

Enjoy!

Downloads last month
-
Safetensors
Model size
124B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/GLM-5.3-mixed-4_5bit

Quantized
(15)
this model