IQ1_XXS Quants

#14
by Banaxi-Tech - opened

Hi Unsloth team, you guys can actually make models fit on everything, I want to know, would it be possible for you to make IQ1_XS, IQ1_XXS or IQ1_XXXS? I would like to run this on 96GB Unified Memory

Hi, would not suggest attempting to do that, as that will be extremely bad. I'd truthfully suggest NVIDIA Nemotron-3-Super-120B-A12B, or GPT-OSS-120B if you want quality. IQ1_XXXS of GLM-5.3-Flash is probably 50% the quality of q4_k_m on NVIDIA Nemotron-3-Super-120B-A12B.

Sign up or log in to comment