Qwen3-Coder-30B-A3B-Instruct - code-focused imatrix GGUF

GGUF quantizations of Qwen/Qwen3-Coder-30B-A3B-Instruct (MoE, 30.5B total / 3.3B active) produced with an importance matrix calibrated on source code only, so the quantization error is biased away from the weights that matter for code generation.

Files

File Bits Size Notes
...-IQ4_XS.gguf ~4.25 ~16 GB best quality that still mostly fits a 16 GB GPU
...-IQ3_M.gguf ~3.7 ~14 GB full offload on 16 GB VRAM with room for context
...-Q4_K_M.gguf ~4.8 ~18.6 GB highest quality here, needs partial CPU offload on 16 GB

code.imatrix is the importance matrix itself, reusable for other quantization levels.

Calibration data

~4.5 MB of real source files sampled from public repositories, covering Python, JavaScript, TypeScript, Go, Rust, C, C++, Java, shell, SQL, plus project Markdown/JSON/YAML/TOML config so formatting-heavy output stays intact (repos: requests, flask, express, gin, ripgrep, nlohmann/json, gson, sqlite, rustlings).

120 chunks x 512 tokens were used for the imatrix pass.

Reproduce

# 1. base model -> Q8_0 GGUF
python3 convert_hf_to_gguf.py ./Qwen3-Coder-30B-A3B-Instruct \
    --outtype q8_0 --outfile Qwen3-Coder-30B-A3B-Instruct-Q8_0.gguf

# 2. importance matrix on the code corpus
llama-imatrix -m Qwen3-Coder-30B-A3B-Instruct-Q8_0.gguf \
    -f code_calib.txt -o code.imatrix --chunks 120 -c 512

# 3. quantize
llama-quantize --allow-requantize --imatrix code.imatrix \
    Qwen3-Coder-30B-A3B-Instruct-Q8_0.gguf out-IQ4_XS.gguf IQ4_XS

Running on a Radeon RX 9060 (16 GB, RDNA4)

Vulkan backend is the most reliable path on RDNA4 today:

llama-server -m Qwen3-Coder-30B-A3B-Instruct-code-imatrix-IQ3_M.gguf \
    -ngl 99 -c 16384 -fa on

If VRAM runs short with IQ4_XS or Q4_K_M, keep the attention layers on the GPU and push MoE expert tensors to system RAM:

llama-server -m ...-IQ4_XS.gguf -ngl 99 --n-cpu-moe 12 -c 16384 -fa on

Only 3.3B parameters are active per token, so CPU offload of a few expert layers costs much less throughput than it would on a dense model.

Downloads last month
147
GGUF
Model size
31B params
Architecture
qwen3moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for n19875624/Qwen3-Coder-30B-A3B-Instruct-code-imatrix-GGUF

Quantized
(172)
this model