Instructions to use avlp12/GLM-5.3-Flash-Alis-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use avlp12/GLM-5.3-Flash-Alis-MLX-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir GLM-5.3-Flash-Alis-MLX-4bit avlp12/GLM-5.3-Flash-Alis-MLX-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
GLM-5.3-Flash Alis MLX 4-bit β withdrawn / νμ
Do not download or serve these weights. The checkpoint is being rewritten.
μ΄ κ°μ€μΉλ₯Ό λ°κ±°λ μλΉνμ§ λ§μΈμ. 체ν¬ν¬μΈνΈλ₯Ό μ¬μμ± μ€μ λλ€.
English
This repository previously hosted a stock mlx-lm affine 4-bit conversion of zai-org/GLM-5.3-Flash.
That conversion used a broken mixed-bit recipe:
- the MoE router (
mlp.gate) was quantized to 8-bit β the router must not be quantized - KDA GEMMs were skipped β those attention-path GEMMs should stay as 8-bit GEMMs
Decode is stuck around 5.5 tok/s. Do not serve this checkpoint, and do not use a previously downloaded copy.
The weight files have been removed from the Hub. The recipe is being fixed and the checkpoint is being rewritten.
This was a stock affine baseline, not an ALIS or DWQ build.
The same bug affects the 4-bit, 6-bit, and 8-bit repos in this set.
νκ΅μ΄
μ΄ μ μ₯μμλ μλ zai-org/GLM-5.3-Flashμ mlx-lm affine 4-bit λ³νλ³Έμ΄ μμμ΅λλ€.
κ·Έ λ³νμ μλͺ»λ νΌν© λΉνΈ λ μνΌλ₯Ό μΌμ΅λλ€:
- MoE λΌμ°ν°(
mlp.gate)λ₯Ό 8-bitλ‘ μμννμ΅λλ€. λΌμ°ν°λ μμννλ©΄ μ λ©λλ€ - KDA GEMMμ 건λλ°μμ΅λλ€. μ΄ν μ κ²½λ‘ GEMMμ 8-bitλ‘ λμ΄μΌ ν©λλ€
λμ½λκ° μ½ 5.5 tok/sμ κ³ μ°©λ©λλ€. μλΉμ μ°μ§ λ§μΈμ. μ΄λ―Έ λ°μ λ 볡μ¬λ³Έλ μ°μ§ λ§μΈμ.
κ°μ€μΉ νμΌμ νλΈμμ μ κ±°νμ΅λλ€. λ μνΌλ₯Ό κ³ μΉ λ€ μ²΄ν¬ν¬μΈνΈλ₯Ό μ¬μμ± μ€μ λλ€.
ALIS/DWQ λΉλκ° μλλλ€. κ°μ λ²κ·Έκ° μ΄ μΈνΈμ 4-bit / 6-bit / 8-bit μ μ₯μμ λͺ¨λ μμ΅λλ€.
Status
| Weight shards | removed |
| Checkpoint | being rewritten |
| Source | zai-org/GLM-5.3-Flash |
Model tree for avlp12/GLM-5.3-Flash-Alis-MLX-4bit
Base model
zai-org/GLM-5.3-Flash