Apodex-1.1-mini-oQ4e-mtp

A 4-bit mixed-precision MLX build of apodex/Apodex-1.1-mini (35B-A3B MoE, Qwen3.5 architecture) made with the oQ quantizer built into oMLX 0.6.4, for Apple Silicon. Unlike the plain 4-bit MLX conversions, the native Multi-Token Prediction head of the source checkpoint is kept, so oMLX can run Lightning MTP speculative decoding on it.

Quantization

  • Quantizer: oMLX oQ, level 4, enhanced (oQ4e), from the bf16 checkpoint
  • Group size: 64
  • Importance matrix: oMLX oqe_code_multilingual calibration set, 523 entries, no dead experts
  • Mixed precision: MoE experts at 4 bits; attention, shared expert and the MTP head at 6 to 8 bits (the exact per-tensor map is in config.json under quantization)
  • MTP head: preserved (language_model.mtp.*, 42 tensors)
  • Size on disk: 21.6 GB, 5 safetensors shards
  • Tokenizer, chat template and generation config are unchanged from the source

oq_imatrix_report.json documents the calibration run.

Usage

Built for and tested with oMLX 0.6.4 on an M2 Max. Download it from the oMLX model page by repo id (stefanprodan/Apodex-1.1-mini-oQ4e-mtp) or place the folder under ~/.omlx/models, then enable Lightning MTP in the model settings to use the MTP head. The chat template enables thinking by default; pass chat_template_kwargs: {"enable_thinking": false} to turn it off.

Other MLX runtimes have not been tested with this build.

License

Apache 2.0, inherited from the source model. See the Apodex-1.1-mini model card for the model description, evaluation results and the technical report.

Downloads last month
-
Safetensors
Model size
6B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for stefanprodan/Apodex-1.1-mini-oQ4e-mtp

Quantized
(18)
this model