Lower variants?

#1
by JasonGWarrior - opened

I only have 16gb of ram and this model looks fun :/

JasonGWarrior changed discussion status to closed
LessThanThree AI org

Yes, they are coming. We are doing compressions now. Will try to make it within <16gb too. Will update the page soon.

kvyb changed discussion status to open
LessThanThree AI org

Hey, you can check out the 3-bit option. Or you can also use the adapter with vLLM with some 1, 2-bit base of qwen3.8 you like.

fyi, I noticed there are issues with thinking modes on lower than 8-bit checkpoints. And am working on a fix regarding that.

for 16gb cards this particular qwen model is best served at iq4_xs rather than q3km

JasonGWarrior changed discussion status to closed
JasonGWarrior changed discussion status to open

Any plans on even more lower variations?

IQ4_XS would be nice

LessThanThree AI org

for 16gb cards this particular qwen model is best served at iq4_xs rather than q3km

Thanks for the recommendation.

LessThanThree AI org

@JasonGWarrior IQ4_XS is now available. It’s about 15.1 GB. The model checkpoints have also been updated and optimised.

Let me know how it works if you try it!

Thanks for your suggestion.

Sign up or log in to comment