Request: NVFP4 Version

#2
by wa999 - opened

GLM-5.3-Flash-REAP50-GGUF NVFP4 Version
Thanks!

GLM-5.3-Flash-REAP50-GGUF NVFP4 Version
Thanks!

https://huggingface.co/patrickbdevaney/GLM-5.3-Flash-REAP50-NVFP4-v2

That one is fp8.
How about a 4Bit NVFP4 version?
Or even better, Uncensored REAP50 MTP 4Bit NVFP4 version🐲🔥

That one is fp8.
How about a 4Bit NVFP4 version?
Or even better, REAP50 MTP 4Bit NVFP4 version🐲🔥

I thought that it's labeled wrong, and the weights compose the same file size, or maybe it was a wrong upload/quant. I want to iterate more on fine tuned draft heads- mtp and dflash for this model. And test with techniques for the fastest decode. Then look at further quantization or minimizing memory traffic per forward pass.

In any case, more is coming related to this model

Sign up or log in to comment