OIOXO Coder 35B (Qwen3.6-35B-A3B, Q4_K_M, with MTP head)
This is the on-device coding model used by the OIOXO IDE. It is the full-quality build of Qwen3.6-35B-A3B: nothing is pruned or distilled.
| Base model | Qwen/Qwen3.6-35B-A3B (Apache-2.0) |
| Quantization | UD-Q4_K_M from unsloth/Qwen3.6-35B-A3B-MTP-GGUF, unmodified |
| Architecture | Mixture-of-experts: 35B total parameters, about 3B active per token, 256 experts, top-8 |
| Extras | Keeps the multi-token-prediction (MTP / nextn) head for lossless self-speculative decoding |
| File | oioxo-coder-35b-a3b-mtp-q4_k_m.gguf, 22,663,387,424 bytes |
Why this build
We measured on our own hard, multi-file coding evaluation:
| Build | Size | Score |
|---|---|---|
| Unpruned 4-bit (this file) | 22.7 GB | 43% |
| Unpruned 2-bit | 11.5 GB | 29% |
| Unpruned 3-bit (IQ3_XXS) | 13.8 GB | 18% |
| 25% experts pruned, 3-bit | 10.7 GB | 21% |
| 56% experts pruned, 2-bit | 4.65 GB | 17% |
Every smaller variant lost a large share of its ability on real projects, so OIOXO ships only this one.
Running it
With llama.cpp b11258 or later:
llama-server -m oioxo-coder-35b-a3b-mtp-q4_k_m.gguf --jinja -fa on -c 32768 -ngl 99 --n-cpu-moe <N>
--n-cpu-moekeeps the experts of N layers on the CPU so the rest fits the GPU.- Add
--spec-type draft-mtp --spec-draft-n-max 3when the whole model fits on the GPU. With experts split to the CPU, MTP was slower in our tests. - On an 8 GB GPU with every layer offloaded, keep the context at 8k or less.
- On Mac CPU-only runs, add
--no-op-offload.
Credits and license
The weights are by the Qwen team; the quantization is by Unsloth. Both are redistributed unmodified under Apache-2.0.
- Downloads last month
- 19
Hardware compatibility
Log In to add your hardware
4-bit
Model tree for payam1394/oioxo-coder-35b-a3b
Base model
Qwen/Qwen3.6-35B-A3B