gguf please?

#5
by Splarkszter - opened

I'd love to see one of your GSQ-RCO GGUF for Qwen3.6 35b-a3b.
We won't be getting an MoE this good and this small for a while.
I doubt Qwen4 is less on than 2months. So please :)

As much as I would love a gguf myself, I believe it's a challenge to make this specific 2-bit-GSQ Humming quant work in llama.cpp or convert into a gguf.
The Humming format is custom and the quant is also non-standard. They would need to port the Triton vLLM kernels and patch to CUDA C kernels and custom C/C++ types, code and patches for a ggml/llama.cpp branch supporting this 2-Bit GSQ quant.
Perhaps a better ask for the lab (please do) is IQ3_XS/IQ3_S/IQ3_M GSQ-RCO quants for MoE Qwen3.6-35B-A3B, similar to the lab's quants of the Qwen3.8-Flash-27B dense model.

Sign up or log in to comment