PR Merged!!

#1
by pawarshardul - opened

Waiting for your awesome AD quants now!!

Atomic Chat org

Hey @pawarshardul !
Thanks for notification, doing quants asap!! :D

Hey @pawarshardul !
Thanks for notification, doing quants asap!! :D

looks like meta has updated model repos -- can you see if they change the weights or just config and templates??

btw -- your explanation of quants making process along with imatrix creation is awesome -- loved it -- no one explained it so nicely !! thank you !! you made my day!!
Also on side note -- nemotron 3.5 lightning is here -- any plans to quantize it??

Atomic Chat org

@pawarshardul yep, I'm trying to finalize this model, and then I would probably want to dive deep into mantaining our turboquant fork, but if you insist i can 😅

take your time -- i had your muse ad-q8 running at 35t/s 550pp t/s and running quite good for a dense model -- so do it when ever you get time !! btw the ling gguf (iq4-nl) turboquant is running good too -- i am running low quant but is performing well as is deepseek v4 quant(iq2-xss)!!

Atomic Chat org

Hey @pawarshardul !
Here you go with nemotron :D
Sorry for the delay, i was thoroughly inspecting and testing if MTP works, imatrix is ok and the fact that i can't release k/i quants for it because nemotron tensor rows cannot be divided by 256 superblocks, that i/k llama.cpp quants demand. It's only an issue with llama.cpp's i/k quants. Nvidia only cares about vLLM and safetensors i suppose.

Smallest quant for nemotron would be Q2_0 - still using imatrix and all, but not i/k quant.

Anyway, ask for any question there, in the community, if you have any, or request anything that you want. I'm running benches currently for all quants btw:
https://huggingface.co/AtomicChat/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

downloading!! yes it seems nvidia only cares about vllm!!!
Btw -- you put so much valuable and detailed information in each model page -- i am in love with your quants-- boss!! Keep doing this incrediible work!!

Sign up or log in to comment