grug-35b-mtp-gguf

grug rock WITH prediction head inside. Q8_0 / Q5_K_M / Q4_K_M, mtp tensors included for llama.cpp speculative decoding (needs build with qwen3_5 MTP support). draft head trained on grug data: t+2 agreement 68.1% (from-scratch head, experimental).

eye rock

caveman ask why grug blind. grug find eye.

download mmproj-grug-35b-mtp-f16.gguf beside any brain rock, then:

llama-server -m grug-35b-mtp-Q4_K_M.gguf \
  --mmproj mmproj-grug-35b-mtp-f16.gguf

eye rock same vision projector as v2.1 parent. MTP graft only touch prediction head; vision tower unchanged. need recent llama.cpp with qwen3_5_moe + MTP support.

main card: grug-35b-mtp. no-MTP rocks: grug-35b-v2-gguf. grug made by ProCreations.

Downloads last month
60
GGUF
Model size
36B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ProCreations/grug-35b-mtp-gguf

Quantized
(1)
this model