Qwen3-Coder-30B-A3B-Instruct โ€” GGUF (Q4_K_M)

Unsloth's Q4_K_M GGUF for Qwen3-Coder-30B-A3B-Instruct (a 30B-total, ~3B-active coding MoE), re-hosted with measured performance for the 1bit engine on Strix Halo.

Contents

  • Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf

Measured performance (Strix Halo, Vulkan)

pp512: 1321 tok/s ยท tg128: 76.8 tok/s

Running it

1bit serve -m Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf --device vulkan

Attribution

Downloads last month
-
GGUF
Model size
31B params
Architecture
qwen3moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for 1bit-MONSTER/Qwen3-Coder-30B-A3B-Instruct-GGUF

Quantized
(173)
this model