Any chance of Q8?

#1
by MB7977 - opened

Thank you very much for these quants. You list a Q8 in the readme, are there any plans to upload that one? I’d like to run it in the best quality possible. I’ve been running a Q6 from an early version of the MSA PR and I’m curious to see if the long context degradation can be cured a bit. Anyway, it’s a great model, thank you for these quants. 🙏

Do you happen to be running the ones made by Avar6? Those had the indexer tensors quanted to Q6, so the degradation is expected there. These ones have them in F32, so Q5 should be a lot better as well

Yes, the Avar6 Q6K quant is the one I’m running. I intend to try your Q5 but I have a ton of room still (used to running GLM 5.2 etc) so hoping someone uploads a Q8 at some point. It’s a good model but devolves pretty badly at longer contexts with the current quant I have.

Hey, I added the Q8_0 as a baseline for comparison but I expect that either bart or unsloth will be adding the "usual" quantizations including Q8_0. I'm trying to squeeze my HF storage allocation as far as it'll go, so I don't want to essentially just take up space on quants that I'm certain they'll have.

No worries at all. Hopefully Unsloth releases one. Always find your quants the best so the 5 bit may have to do. Or I’ll do it myself. 🙂

Sign up or log in to comment