- Iris-mini APEX-I-MiniPlus-V2.1 GGUF
- π’ Optimization History & Transparency Notice
- β‘ Quick Navigation Index
- π¦ Model Files & Technical Specifications
- π οΈ Surgical Tensor Quantization Map (Audited from GGUF)
- ποΈ Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
- π₯ The 24GB Miracle: Full 256K Context Runs In VRAM!
- π Recommended Configuration & Setup
- π’ Optimization History & Transparency Notice
Iris-mini APEX-I-MiniPlus-V2.1 GGUF
The Definitive Frontier MoE Β· Blistering +24 to 28+ tok/s with System RAM Offload Β· Full 256K Context on 24GB Workstations
π THE DEFINITIVE SPECIFICATION IN THE 13β14 GB CEILING
This APEX-I-MiniPlus-V2.1 release represents the absolute technological limit of sparse Mixture-of-Experts quantization within the 13β14 GB envelope. Every single tensor of its 40 layers and 256 micro-experts has been mathematically audited to maximize reasoning precision, eliminate recurrence state drift, and prevent AVX2 CPU dequantization stalls.
β οΈ DO NOT CONFUSE APEX-I-MINIPLUS WITH GENERIC COMMUNITY APEX-I-MINI!
Regardless of release version (whether V1, V2, or V2.1), NEVER confuse handcrafted APEX-I-MiniPlus builds with generic community APEX-I-Mini releases:
- Generic Community APEX-I-Mini: Uniformly compresses all core MoE experts down to aggressive 2-bit
IQ2_S(dropping below the critical quality floor), leaves the sensitive token output head unarmored at 3-bitQ3_K_M, and compresses attention projections down toQ3_K. In deep reasoning models, this triggers severe perplexity spikes, syntax errors, and broken code brackets.- Handcrafted APEX-I-MiniPlus (All Editions by IsValorum): Every single MiniPlus releaseβfrom V1 and V2 to V2.1βis a custom tensor-by-tensor architecture that preserves uncompressed
F32router gates, armors the token output head in high-precisionQ6_K, safeguards attention gates inQ8_0, and keeps core reasoning experts at or above calibrated 3-bit (IQ3_XXS/IQ3_S). Even our earlier builds vastly outperform generic community APEX recipes and flat 3-bit quants.
π’ Optimization History & Transparency Notice
We maintain our previous releases publicly as a transparent engineering record of continuous optimization. Below is the exact evolutionary roadmap of our MiniPlus architectures:
| Specification | Core Experts (10β29) | Edge Experts (0β9, 30β39) | Shared Expert (shexp) |
Full Attention (L3, 7, 11, ...) | Attention Gates (30 Layers) | Output Head (output.weight) |
Routers (gate_inp) |
Size / Overhead | Real-World Impact |
|---|---|---|---|---|---|---|---|---|---|
| Generic APEX Mini | IQ2_S (2.50 bpw) |
Q3_K (only 5 layers) |
Q4_K / Q3_K |
Q3_K |
Compressed | Q3_K_M |
Compressed | Baseline (~12.5 GB) | Severe syntax errors, broken code indentation, high perplexity in <think>. |
| MiniPlus V1 | IQ3_XXS (3.06 bpw) |
Q3_K (5 layers) |
Q4_K / IQ4_NL |
Q3_K |
Q8_0 |
Q6_K |
F32 (uncompressed) |
Baseline MiniPlus (~13.56 GiB) | Lean & agile profile; runs flawlessly in system RAM; zero router drift; protects core logic. |
| π₯ MiniPlus V2.1 (CURRENT) | IQ3_XXS |
Q3_K (10 layers) |
Q5_K (All 40 layers) |
Q4_K (q/k/v) + Q6_K (output) |
Q8_0 |
Q6_K |
F32 |
< 180 MB extra over V1 (~13.74 GiB total) | Enhanced long-context stability & refined throughput; +24 to 28+ tok/s streaming under RAM offload; shared foundation experts armored across all 40 layers. |
π‘ Which Edition Should You Choose for Iris-mini? (V1 vs. V2.1)
- Both V1 and V2.1 run flawlessly with the vast majority of the model in system RAM (DDR4/DDR5): Both utilize linear, CPU-friendly dequantization that avoids AVX2 lookup stalls.
- Why choose V2.1? It provides enhanced surgical protection (
Q5_Kshared foundation experts across all 40 layers,Q8_0attention gates, andQ4_K/Q6_Kfull attention anchors) and refined throughput for only ~180 MB more, an overhead that is completely negligible when running in system RAM.- Why choose V1? If your machine has strict RAM/VRAM limits and you need the absolute leanest footprint while still vastly outperforming flat 3-bit quants and generic community APEX Mini, V1 is fantastic.
- Full GPU VRAM (
-ngl 99): Both perform identically at maximum hardware speed.π Need our leanest possible memory footprint? Explore the Iris-mini MiniPlus V1 Edition.
β‘ Quick Navigation Index
- π¦ Model Files & Technical Specifications
- π οΈ Surgical Tensor Quantization Map (Audited from GGUF)
- ποΈ Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
- π₯ The 24GB Miracle: Full 256K Context Runs In VRAM!
- π Recommended Configuration & Setup
π¦ Model Files & Technical Specifications
| File Name | File Size | Memory Footprint | BPW | Description |
|---|---|---|---|---|
Iris-mini.APEX-I-MiniPlus-V2.1.gguf |
14.75 GB (13.74 GiB) |
13.74 GiB |
3.40 BPW | Linear vector reasoning, mathematical problem solving & algorithmic synthesis MoE |
- Base Model: AllSpark-Research/Iris-mini
- Parameters: 35.2B total (approx. 2.6B to 3.2B active per token)
- Architecture: 40 layers, 256 micro-experts (8 active per token) + hybrid linear attention / DeltaNet recurrent layers
- Context Length: 262,144 tokens (native 256K)
π οΈ Surgical Tensor Quantization Map (Audited from GGUF)
The exact tensor breakdown below has been verified directly from the compiled binary weights:
| Layer Group | Sub-Component / Tensor | Qty | Precision | Engineering Rationale |
|---|---|---|---|---|
| Global Output Head | output.weight |
1 | Q6_K |
Preserves near-FP16 token classification; eliminates syntax errors, bracket drops, and hallucinations. |
| Global Embeddings | token_embd.weight |
1 | Q4_K |
High-fidelity vocabulary embedding representation. |
| All Normalizations | output_norm, attn_*_norm, ssm_norm |
171 | F32 |
100% uncompressed numerical stability across all 40 layers. |
| Expert Routers | blk.*.ffn_gate_inp, ffn_gate_inp_shexp |
80 | F32 |
100% uncompressed routing fidelity across 256 micro-experts; zero router drift. |
| Attention Gates | blk.*.attn_gate.weight (30 Hybrid Layers) |
30 | Q8_0 |
High-precision attention gating across hybrid DeltaNet recurrence layers; eliminates crosstalk. |
| Shared Foundation Experts | blk.*.ffn_{gate,down,up}_shexp (All 40 Layers) |
120 | Q5_K |
Foundation knowledge backbone active on 100% of tokens; protected in high-precision linear Q5_K. |
| Periodic Full Attention | blk.{3,7,11,...}.attn_q/k/v (10 Anchor Layers) |
30 | Q4_K |
Full quadratic attention anchor checkpoints for deep needle-in-a-haystack retrieval. |
| Periodic Full Attention | blk.{3,7,11,...}.attn_output (10 Anchor Layers) |
10 | Q6_K |
Armored attention output projection over deep context. |
| Recurrent SSM Scales | blk.*.ssm_alpha, ssm_a, ssm_conv1d, ssm_dt |
120 | F32 |
Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift. |
| Linear Attention & SSM | blk.*.attn_qkv, ssm_beta, ssm_out |
90 | Q3_K |
Linear AVX2 execution; zero SIMD CPU stalls during system RAM streaming. |
| Edge MoE Experts | Layers 0β9 & 30β39 (ffn_*_exps) |
60 | Q3_K |
Linear SIMD execution; enables +24 to 28+ tok/s streaming under system RAM offload. |
| Core MoE Experts | Layers 10β29 (ffn_*_exps) |
60 | IQ3_XXS |
Calibrated with importance matrix (imatrix) for maximum compactness in deep layers. |
ποΈ Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
Empirically verified in Unsloth Studio & llama.cpp:
| Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
|---|---|---|---|---|
| NVIDIA RTX 5080 / 5090 (Blackwell) | Full GPU (-ngl 99) |
120 β 145+ tok/s | 2,800 β 3,900+ tok/s | Blistering throughput on GDDR7 bandwidth |
| NVIDIA RTX 4090 (24GB GDDR6X) | Full GPU (-ngl 99) |
90 β 115+ tok/s | 2,000 β 2,800+ tok/s | Linear attention layers slash prefill latency |
| NVIDIA RTX 3090 (24GB GDDR6) | Full GPU (-ngl 99) |
72 β 88+ tok/s | 1,500 β 2,200+ tok/s | Full 256k native window in VRAM |
| Workstation / Laptop (DDR4 / DDR5 RAM) | Hybrid Offload (Few layers in VRAM) | 24.25 β 28.37 tok/s | 385 β 410+ tok/s | Zero AVX2 CPU stalls; fast streaming from system RAM |
- Aggressive Hybrid Offload Profile: Sustained 24.25 to 28.37 tok/s generation with reasoning enabled, even when only ~4.2 GB VRAM is available and the rest of the 13.74 GiB model streams from system RAM.
π¬ Empirical Testbed Architecture & Desktop/Server Scaling
- Empirical Benchmark Hardware: The hybrid offload and system RAM streaming figures documented above (sustaining 24.25 to 28.37 tok/s) were measured on a consumer laptop powered by an Intel 12th Gen Alder Lake architecture featuring a hybrid design of Performance Cores (P-Cores) and Efficient Cores (E-Cores) paired with dual-channel system RAM and constrained laptop power/thermal envelopes.
- Thread Scheduling & E-Core Contention: In hybrid architectures like Alder Lake, OS thread scheduling across background E-Cores and lower single-core mobile power limits introduce memory bandwidth and thread synchronization overhead during CPU dequantization.
- Dramatic Scaling on Higher-End Processors: When running on desktop or server processors (such as modern AMD Ryzen 7000 / 9000 Zen 4/5 series or high-TDP Intel desktop platforms with dedicated performance cores, large L3 caches, and high-bandwidth dual- or quad-channel DDR5 running at 6000+ MT/s), streaming generation speeds and prefill throughput will scale dramatically higher, substantially exceeding these measured mobile numbers.
π₯ The 24GB Miracle: Full 256K Context Runs In VRAM!
Iris-mini APEX-I-MiniPlus-V2.1 fits the entire 256K context window within 24GB VRAM:
| Context Length | Model Weights (Est.) | KV Cache (q8_0, 4 slots) | Compute Buffers | Total GPU VRAM (Est.) | Feasibility |
|---|---|---|---|---|---|
| 32,768 (32k) | 13.74 GiB |
0.58 GiB |
1.80 GiB |
16.12 GiB |
Full offload on 24GB; partial on 16GB |
| 65,536 (64k) | 13.74 GiB |
0.92 GiB |
1.95 GiB |
16.61 GiB |
Effortless fit on 24GB GPUs |
| 131,072 (128k) | 13.74 GiB |
1.58 GiB |
2.22 GiB |
17.54 GiB |
Effortless fit on 24GB GPUs |
| 262,144 (256k) | 13.74 GiB |
2.92 GiB |
2.80 GiB |
19.46 GiB |
π₯ FULL 256K NATIVE IN VRAM! |
Note: Leaves comfortable headroom for display drivers and compute buffers on standard 24GB GPUs (RTX 3090, RTX 4090, RTX 5090).
π Recommended Configuration & Setup
llama-server.exe \
-m Iris-mini.APEX-I-MiniPlus-V2.1.gguf \
--port 8080 \
--parallel 4 \
--flash-attn on \
--fit on \
-c 104960 \
--cache-type-k q8_0 \
--cache-type-v q8_0
- Downloads last month
- -
We're not able to determine the quantization variants.