Qwen3.6-35B-A3B-MTP GGUF
Recommended way to run this model:
# max MTP draft of 2 tokens
llama-server -hf ggml-org/Qwen3.6-35B-A3B-MTP-GGUF --spec-type draft-mtp --spec-draft-n-max 2
# max MTP draft of 3 tokens
llama-server -hf ggml-org/Qwen3.6-35B-A3B-MTP-GGUF --spec-type draft-mtp --spec-draft-n-max 3
Then, access http://localhost:8080
Requires the changes from: https://github.com/ggml-org/llama.cpp/pull/22673
- Downloads last month
- 1,238
Hardware compatibility
Log In to add your hardware
8-bit
16-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support
Model tree for ggml-org/Qwen3.6-35B-A3B-MTP-GGUF
Base model
Qwen/Qwen3.6-35B-A3B