GLM-5.3-Flash Alis MLX 6-bit β€” withdrawn / 회수

Do not download or serve these weights. The checkpoint is being rewritten.

이 κ°€μ€‘μΉ˜λ₯Ό λ°›κ±°λ‚˜ μ„œλΉ™ν•˜μ§€ λ§ˆμ„Έμš”. 체크포인트λ₯Ό μž¬μž‘μ„± μ€‘μž…λ‹ˆλ‹€.

English

This repository previously hosted a stock mlx-lm affine 6-bit conversion of zai-org/GLM-5.3-Flash.

That conversion used a broken mixed-bit recipe:

  • the MoE router (mlp.gate) was quantized to 8-bit β€” the router must not be quantized
  • KDA GEMMs were skipped β€” those attention-path GEMMs should stay as 8-bit GEMMs

Decode is stuck around 5.5 tok/s. Do not serve this checkpoint, and do not use a previously downloaded copy.

The weight files have been removed from the Hub. The recipe is being fixed and the checkpoint is being rewritten.

This was a stock affine baseline, not an ALIS or DWQ build.

The same bug affects the 4-bit, 6-bit, and 8-bit repos in this set.

ν•œκ΅­μ–΄

이 μ €μž₯μ†Œμ—λŠ” μ›λž˜ zai-org/GLM-5.3-Flash의 mlx-lm affine 6-bit λ³€ν™˜λ³Έμ΄ μžˆμ—ˆμŠ΅λ‹ˆλ‹€.

κ·Έ λ³€ν™˜μ€ 잘λͺ»λœ ν˜Όν•© λΉ„νŠΈ λ ˆμ‹œν”Όλ₯Ό μΌμŠ΅λ‹ˆλ‹€:

  • MoE λΌμš°ν„°(mlp.gate)λ₯Ό 8-bit둜 μ–‘μžν™”ν–ˆμŠ΅λ‹ˆλ‹€. λΌμš°ν„°λŠ” μ–‘μžν™”ν•˜λ©΄ μ•ˆ λ©λ‹ˆλ‹€
  • KDA GEMM을 κ±΄λ„ˆλ›°μ—ˆμŠ΅λ‹ˆλ‹€. μ–΄ν…μ…˜ 경둜 GEMM은 8-bit둜 두어야 ν•©λ‹ˆλ‹€

λ””μ½”λ“œκ°€ μ•½ 5.5 tok/s에 κ³ μ°©λ©λ‹ˆλ‹€. μ„œλΉ™μ— μ“°μ§€ λ§ˆμ„Έμš”. 이미 λ°›μ•„ λ‘” 볡사본도 μ“°μ§€ λ§ˆμ„Έμš”.

κ°€μ€‘μΉ˜ νŒŒμΌμ€ ν—ˆλΈŒμ—μ„œ μ œκ±°ν–ˆμŠ΅λ‹ˆλ‹€. λ ˆμ‹œν”Όλ₯Ό 고친 λ’€ 체크포인트λ₯Ό μž¬μž‘μ„± μ€‘μž…λ‹ˆλ‹€.

ALIS/DWQ λΉŒλ“œκ°€ μ•„λ‹™λ‹ˆλ‹€. 같은 버그가 이 μ„ΈνŠΈμ˜ 4-bit / 6-bit / 8-bit μ €μž₯μ†Œμ— λͺ¨λ‘ μžˆμŠ΅λ‹ˆλ‹€.

Status

Weight shards removed
Checkpoint being rewritten
Source zai-org/GLM-5.3-Flash
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for avlp12/GLM-5.3-Flash-Alis-MLX-6bit

Finetuned
(13)
this model

Collection including avlp12/GLM-5.3-Flash-Alis-MLX-6bit