Qwen3.6-35B-A3B-oQ4e-MTP-MLX

This is the Qwen3.6-35B-A3B model package used by Sprig in Treeish. It combines an oQ4e mixed-precision MLX quant with its embedded Multi-Token Prediction head and the chat template validated for Sprig's coding-agent workflow.

The package is ready to download and load as-is. Treeish pins an exact repository commit rather than following main.

Model

  • Base model: Qwen/Qwen3.6-35B-A3B
  • Architecture: 35B total parameters, approximately 3B active parameters
  • Quantisation: oQ4e imatrix-enhanced mixed precision
  • Default quantisation: 4-bit affine, group size 64
  • Per-tensor overrides: 5, 6 and 8-bit, with group sizes 64 and 128
  • Format: MLX safetensors
  • MTP: one embedded layer under language_model.mtp.*
  • Context length: 262,144 tokens

The package contains 2,052 indexed tensors, including 42 embedded MTP tensors. Its five weight shards contain 21,612,530,528 bytes of tensor data.

Provenance

The quantised weights, model configuration and oQ calibration report are byte-identical to Jundot/Qwen3.6-35B-A3B-oQ4e-mtp at commit 14c285372cbdb1777adea5bb49087ced0bffc0b5. The source declares oMLX 0.4.5.dev1 as the converter and Qwen/Qwen3.6-35B-A3B as the base model. Its imatrix report records 128 samples of 512 tokens from oqe_code_multilingual, with 470 of 510 entries applied. It does not identify the exact base-model commit used for conversion, so this is a curated, byte-pinned distribution rather than a byte-reproducible conversion recipe.

The source reports an 83.88% average across its MMLU, Winogrande, HumanEval and MBPP evaluation set, including 93.90% on HumanEval. See the source model card for the full methodology and comparison table.

The chat template is Froggeric's Qwen3.6 v21.3 template. It is byte-identical to archive/v21_chat_template.jinja in froggeric/Qwen-Fixed-Chat-Templates at commit 9f14778c92c3b5ed3e0738085694c0d3452802dd.

No model, tokenizer or configuration tensors were changed for this release. The release replaces the chat template and adds the licence, provenance and file manifest.

Runtime compatibility

This package is validated with Treeish's pinned MLX Swift runtime. A different runtime must support the per-tensor quantisation overrides in config.json and the embedded Qwen MTP layout.

Treeish uses this model from 24 GB of unified memory and recommends 32 GB. Headroom depends on context length, cache settings and other running applications.

Treeish validation

The release candidate was validated on a 36 GB M4 Max using the release build of Treeish's benchmark:

  • All 2,052 indexed tensors were present exactly once and assigned to the expected shard.
  • Embedded-MTP tool use at block size 2 produced a parsed search_text call with the requested query and result count.
  • The 1,066-token performance fixture generated 96.8 tokens/s without MTP and 100.9, 109.6 and 110.4 tokens/s with MTP block sizes 2, 3 and 4.
  • MTP block size 4 accepted 61 of 84 proposed draft tokens in that fixture.
  • Sprig's exact-string edit format produced 8 exact edits and 11 structurally valid edits from 12 fixtures in one trial.

These figures describe one machine and one small release fixture. They are not general model benchmarks.

Limitations

Quantisation trades some model quality for memory use and local generation speed. Applications should validate the model against their own prompts, tool format and runtime.

The model package contains no custom executable code. File sizes, SHA-256 digests and source revisions are recorded in RELEASE_MANIFEST.json.

Licence

Qwen3.6-35B-A3B is licensed under Apache 2.0. The full licence text is included in LICENSE. The Froggeric template repository also declares Apache 2.0 and is attributed above.

Downloads last month
21
Safetensors
Model size
6B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for treeish/Qwen3.6-35B-A3B-oQ4e-MTP-MLX

Quantized
(753)
this model
Quantizations
1 model