Thanks!

#2
by pjsmith - opened

Tried Qwen3.8-27B-AD-Q4_K.gguf after downloading the Unsloth 4 bit UD_XL version. Only 800mb smaller (only!). But running on my 3060 12gb, went from well under 1 t/s generation (about 0.8 on avg), to ~3 t/s. Yes, clearly still miserable, but prompt processing is around 300 t/s. Goes from unusable to just about tolerable for specific tasks. 4-bit rules on these older Ampere cards, because quantisation overhead in CUDA/llama.cpp is a lot, so even the smaller models tend to run slower, despite the memory savings. Cheers.

Atomic Chat org

Hi @pjsmith !
Glad to hear it! Trying my best πŸ˜€
We all are!! Keep in touch with Atomic)

3t/s? What are you doing, offloading 90% of it to system ram?

3t/s? What are you doing, offloading 90% of it to system ram?

in dense llms,%20 offload can literally kill your token/sec.

Sign up or log in to comment