OIOXO Coder 35B (Qwen3.6-35B-A3B, Q4_K_M, with MTP head)

This is the on-device coding model used by the OIOXO IDE. It is the full-quality build of Qwen3.6-35B-A3B: nothing is pruned or distilled.

Base model Qwen/Qwen3.6-35B-A3B (Apache-2.0)
Quantization UD-Q4_K_M from unsloth/Qwen3.6-35B-A3B-MTP-GGUF, unmodified
Architecture Mixture-of-experts: 35B total parameters, about 3B active per token, 256 experts, top-8
Extras Keeps the multi-token-prediction (MTP / nextn) head for lossless self-speculative decoding
File oioxo-coder-35b-a3b-mtp-q4_k_m.gguf, 22,663,387,424 bytes

Why this build

We measured on our own hard, multi-file coding evaluation:

Build Size Score
Unpruned 4-bit (this file) 22.7 GB 43%
Unpruned 2-bit 11.5 GB 29%
Unpruned 3-bit (IQ3_XXS) 13.8 GB 18%
25% experts pruned, 3-bit 10.7 GB 21%
56% experts pruned, 2-bit 4.65 GB 17%

Every smaller variant lost a large share of its ability on real projects, so OIOXO ships only this one.

Running it

With llama.cpp b11258 or later:

llama-server -m oioxo-coder-35b-a3b-mtp-q4_k_m.gguf --jinja -fa on -c 32768 -ngl 99 --n-cpu-moe <N>
  • --n-cpu-moe keeps the experts of N layers on the CPU so the rest fits the GPU.
  • Add --spec-type draft-mtp --spec-draft-n-max 3 when the whole model fits on the GPU. With experts split to the CPU, MTP was slower in our tests.
  • On an 8 GB GPU with every layer offloaded, keep the context at 8k or less.
  • On Mac CPU-only runs, add --no-op-offload.

Credits and license

The weights are by the Qwen team; the quantization is by Unsloth. Both are redistributed unmodified under Apache-2.0.

Downloads last month
19
GGUF
Model size
0.4B params
Architecture
clip
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for payam1394/oioxo-coder-35b-a3b

Quantized
(863)
this model