GGUF Release: Handcrafted APEX-I-MiniPlus V2.1 (Optimized for System RAM Streaming & Massive Context)

#10
by IsValorum - opened

Hi everyone,

I have published a custom, handcrafted APEX-I-MiniPlus V2.1 GGUF release of Occamy-1.0:
IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF


Key Engineering Highlights:

  • Specially Engineered for Full or Partial System RAM Streaming: Tailored specifically for workstations and consumer setups running most or all of the model out of system RAM (DDR4 / DDR5). By utilizing linear SIMD-optimized Q3_K edge experts and upgrading shared foundation experts to Q5_K across all 40 layers, AVX2 CPU lookup stalls are completely eliminated (sustaining +24 to 28+ tok/s generation).
  • Near-VRAM Speed on Modern CPUs: Depending on your processor architecture (IPC / single-core performance) and memory bandwidth (dual-channel DDR4 or high-speed DDR5 6000+ MT/s), streaming throughput across system RAM can approach speeds remarkably close to having the entire model in VRAM.
  • Massive & Full Context Window Support (+160k to 256k): Designed to handle deep contexts without memory fragmentation or quality degradation.
  • Hybrid Offload Optimization: Includes the dedicated mmproj-Accio-Lab_occamy-1.0-Q8_0.gguf vision projector. In hybrid offload, mmproj can be loaded explicitly in GPU VRAM for instant, zero-latency visual document parsing and OCR on GPU tensor cores while the vast language/reasoning model weights stream smoothly from system RAM.
  • Zero Routing Drift & Intact Syntax: Retains uncompressed F32 router gates (blk.*.ffn_gate_inp.weight), high-precision Q6_K token output head (output.weight), and Q8_0 attention gates across all recurrent hybrid layers.
  • Full 24GB VRAM Offload: If you have 24GB+ VRAM (-ngl 99), the model runs natively at full GPU tensor core speeds.

Quick Run with llama.cpp:

# High-speed hybrid RAM streaming with GPU offload (example with 16 layers or vision in VRAM):
llama-cli -m Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF.gguf -ngl 16 -c 32768 --temp 0.6 --threads 8

Direct link to the model repository: IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF

Feedback and community benchmarks are warmly welcome! Enjoy!

Sign up or log in to comment