GGUF Release: Handcrafted APEX-I-MiniPlus (3.36 BPW) + Q8_0 Vision Projector

#9
by IsValorum - opened

Hi community!

I have published a custom, handcrafted APEX-I-MiniPlus (3.36 BPW) quantization of Nex-N2.5-mini, bundled with a dedicated Q8_0 vision projector (mmproj):
πŸ‘‰ IsValorum/Nex-N2.5-mini-APEX-I-MiniPlus-GGUF

Highlights:

  • Preserved Vision Quality: Includes the dedicated mmproj-nex-agi_Nex-N2.5-mini-Q8_0.gguf ensuring near-lossless visual encoding for OCR, document parsing, and agentic screen workflows.
  • Handcrafted Hybrid Precision: Uses multi-stage imatrix calibration. Preserves critical reasoning pathways and routing weights while optimizing MoE experts down to ~14.5 GB total footprint.
  • Extreme Hardware Accessibility: Tested and verified on Unsloth Studio. Operates on 4GB VRAM laptops (~3.8 GB VRAM footprint + RAM on DDR4 3200) reaching 23-26+ tok/s generation and 300-410 tok/s prefill without MTP.
  • Enterprise / 24GB Ready: Full 256k context breakdown provided in the model card for high-end setups.

Check it out and test it with your multimodal workflows!

Sign up or log in to comment