UD Q4

#1
by anvme - opened

Hello! Can u add please UD Q4 K XL?

I'm unable to do a UD_Q4_K_XL, but I have added Q4_K_M.

Correction: I looked back at how I made this for the Ling-3.0-flash ggufs, and ended up using the same methodology to create UD-Q4_K_XL.

I am uploading now. Note that this is somewhat of a "custom quant." The base quant is Q4_K_M and the recipe file overrides specific groups: expert gate/up stays Q4_K, expert down bumped to Q6_K, attention/kv/ssm bumped to Q6_K–Q8_0, the last few layers and the MTP block held at Q6_K–Q8_0, and --output-tensor-type set at Q6_K.

bloomer010 changed discussion status to closed

Sign up or log in to comment