Instructions to use terion-mlx/GLM-4.5-Air-mixed_4_6 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use terion-mlx/GLM-4.5-Air-mixed_4_6 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir GLM-4.5-Air-mixed_4_6 terion-mlx/GLM-4.5-Air-mixed_4_6
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
GLM-4.5-Air (Terion in-house mixed_4_6 MLX quant)
In-house MLX quantization of zai-org/GLM-4.5-Air, produced by Terion for internal evaluation and republished for the community.
- Method:
mlx_lm.convertwith--quant-predicate mixed_4_6(mixed 4/6-bit precision) plus a periodic Metal-cache-clear wrapper during the save step to avoid unified-memory pressure crashes on Apple Silicon. - Actual bits-per-weight: 4.809
- Size: ~60GB
- Base model: zai-org/GLM-4.5-Air (see base repo for architecture/license/benchmark details)
Use with mlx-lm / Rapid-MLX. This is a straightforward quantization of the original weights.
- Downloads last month
- -
Model size
107B params
Tensor type
U32
路
BF16 路
F32 路
Hardware compatibility
Log In to add your hardware
4-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support
Model tree for terion-mlx/GLM-4.5-Air-mixed_4_6
Base model
zai-org/GLM-4.5-Air