GGUF Release: Handcrafted APEX-I-MiniPlus-V2 (3.38 BPW) + Q8_0 mmproj Vision Projector

#2
by IsValorum - opened

Hi Accio-Lab team and community!

I have created and released a custom, handcrafted APEX-I-MiniPlus-V2 (3.38 BPW) quantization of Occamy-1.0, bundled with its dedicated high-precision Q8_0 vision projector (mmproj):
πŸ‘‰ IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2-GGUF

Key Technical Highlights:

  • 100% Handcrafted IQ-Enhanced Architecture: Built with tensor-by-tensor rules using importance matrix calibration (imatrix). Sensitive layers and attention boundaries are upgraded 1:1 to non-linear IQ3_S and IQ4_NL variants without dropping bits, while MoE experts sit at IQ3_XXS (>3 BPW floor).
  • High-Precision Vision Projector: Includes mmproj-Accio-Lab_occamy-1.0-Q8_0.gguf (614 MB) ensuring near-lossless OCR, visual document parsing, and agentic UI capabilities.
  • Budget Hardware / 4GB VRAM Capable: Empirically verified in Unsloth Studio. Uses only ~3.8 GB VRAM with remainder in RAM (DDR4 3200), delivering 23 to 26+ tok/s generation and 300 to 410 tok/s prefill.
  • Full Context Scalability: Fully verified for 24GB GPUs supporting native 256k context with Q8 KV cache.

Feel free to check it out!

Sign up or log in to comment