Nex-N2.5-mini-oQ5

Unofficial MLX quantization of nex-agi/Nex-N2.5-mini for Apple Silicon. The upstream model is a multimodal mixture-of-experts model; this repository contains MLX safetensors, not GGUF or PyTorch weights. I am not affiliated with Nex AGI.

What is in this repository

Property Value
Architecture Qwen3_5MoeForConditionalGeneration
Quantization affine, group size 64; 5-bit default with 6-bit and 8-bit module overrides
Weight size 23.57 GiB (25.30 GB), 5 safetensors shards
Text model 40 layers, 256 experts, 8 experts selected per token
Vision vision tensors retained in BF16
Context limit in config 262,144 tokens; usable context depends on available memory
MTP not present (mtp_num_hidden_layers: 0)

The numbers above were read from the shipped config.json, model.safetensors.index.json and safetensors headers. The five shards contain 2,010 indexed tensors. The quantization recipe is recorded in config.json so a compatible MLX loader can reconstruct the per-module precision.

This conversion has not been benchmarked against the upstream BF16 model. The upstream benchmark figures on its model card are not results for these quantized weights. Quantization can change output quality, and memory use grows with context length and cache settings.

Usage

Download the model into your oMLX model directory:

hf download suzu89/Nex-N2.5-mini-oQ5 --local-dir ~/.omlx/models/Nex-N2.5-mini-oQ5
omlx serve --model-dir ~/.omlx/models --port 8000

Use the model ID Nex-N2.5-mini-oQ5 in oMLX. This architecture includes a vision tower, so use an MLX runtime with Qwen3.5 MoE multimodal support. The 262k context value is an architecture limit, not a promise that it will fit in memory.

License and attribution

The upstream repository declares Apache License 2.0. This repository includes the license text and a notice identifying the source and the quantization change. The upstream model and its reported evaluations belong to Nex AGI.

Citation

@misc{nex-n25-mini-oq5,
  title = {Nex-N2.5-mini-oQ5: MLX quantization of Nex-N2.5-mini},
  author = {suzu89},
  year = {2026},
  url = {https://huggingface.co/suzu89/Nex-N2.5-mini-oQ5},
  note = {Unofficial quantization of nex-agi/Nex-N2.5-mini}
}
Downloads last month
-
Safetensors
Model size
35B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for suzu89/Nex-N2.5-mini-oQ5

Quantized
(36)
this model