tooltd/GRM-3.2-Sky-MTP-GGUF

This model was converted to GGUF format from OrionLLM/GRM-3.2-Sky using llama.cpp via the ggml.ai's GGUF-my-repo space. Refer to the original model card for more details on the model.

MTP HEAD

  1. Download Qwen3.6-35B-A3B MTP-ONLY from https://huggingface.co/a4lg/Qwen3.6-35B-A3B-MTP-ONLY-GGUF

  2. Add MTP HEAD to GRM-3.2-Sky GGUF model by using my script modified. otherwise, it will cause an error. https://huggingface.co/tooltd/GRM-3.2-Sky-MTP-GGUF/blob/main/grm_mtp_graft.py

python grm_mtp_graft.py grm-3.2-sky-q8_0.gguf Qwen3.6-35B-A3B-MTP-ONLY-Q8_0.gguf grm-3.2-sky-mtp-q8_0.gguf

Draft acceptance rate is usually above 80%. Done!

Downloads last month
351
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tooltd/GRM-3.2-Sky-MTP-GGUF

Quantized
(5)
this model