- Paretrix Quantization Suite (v2.6.0)
- 🪜 The Mod-3 Ladder
- 🧮 THE MATRIX — measured rate knowledge
- ⚔️ Campaign engine (RCO)
- 📊 Benchmarks
- 1. ornith-ai/Ornith-1.5-9B — qwen35 GDN hybrid · untied 248k vocab
- 2. TokenRhythm/NeoHorse-1-4B — qwen3.5 fine-tune · GDN hybrid · tied
- 3. Nanbeige/Nanbeige4.2-3B — looped dense · untied 166k vocab
- 4. XHToken/Spark-X2.5-4B — dense SWA hybrid 3:1 · tied 131k vocab
- 5. webAI-Official/TwIL-LM3-Pro — granite dense · untied 128k vocab
- 6. LiquidAI/LFM2.5-2.6B — shortconv mixer · tied readout
- 7. openbmb/MiniCPM5-2B — dense llama-arch · untied 130k readout
- 8. jinaai/ReaderLM-v2 — dense qwen2 · tied 152k vocab · HTML/Markdown extractor
- 9. SeerRay-Lab/Xiaomi-OCR-0 — qwen35 GDN hybrid · tied 248k vocab
- 📉 Sub-Nano compendium — the Nano → Pico step
- Summary — 9 strict Pareto dominations
- ⚡ Speculative decoding: MTP · DSpark · DFlash
- ⚡ Quick start
- 📜 Citation & Credits
- 🪜 The Mod-3 Ladder
Paretrix Quantization Suite (v2.6.0)
Paretrix is an empirical, activation-aware quantization suite for GGUF language models and speculative decoding modules.
Pareto + Matrix → Paretrix. Each tier is priced so that no better trade exists at that compression. The suite measures real activation sensitivity (ΔKLD/MiB) per tensor class through llama-imatrix, learns rate tables from cross-architecture campaigns (THE MATRIX), and allocates bitwidths under exact budget targets: flat recipes where uniformity wins, rate-calibrated knapsack where heterogeneity pays.
Standard uniform quantization applies one bitwidth family across the whole network. Paretrix reads the model's own activation field (which classes are expensive to cut, which are nearly free, and where depth matters), then spends the budget where measurements show the greatest return.
🪜 The Mod-3 Ladder
Each tier lands on a multiple of 3 percent of the unquantized BF16 source, so the shipping name states its weight band (model-Paretrix-<Tier>-<XX>pc.gguf, ±1.5 pp). Tiers run from largest to smallest.
| Tier | Ratio | Allocation Strategy | Primary Purpose |
|---|---|---|---|
| Fidelity | 48% | Q6_K base + Q8_0 pockets biased to the late layer third (G4) |
High-fidelity archival tier. |
| Precision | 42% | Flat Q6_K + full-attention band Q8_0; the readout lever is arch-scoped (L12) |
Coding, mathematics, high-entropy reasoning. |
| Quality | 36% | Flat Q5_K_M + imatrix (L6); self re-shuffle where the twin is a stock copy |
General production tier. |
| Compact | 33% | Exact-budget DP, or flat Q4_K_M + band (Q6_K) + recurrent floors |
Daily driver for 2B–4B architectures. |
| Mini | 30% | Exact-budget DP, or L4 flat on GDN hybrids (IQ4_XS + recurrent floors + full-attention band) |
Balanced entry tier for ≥ 9B models on 8 GB VRAM. |
| Nano | 27% | Exact-budget DP (IQ3_S–IQ4_XS base, tied readout Q5_K when vocab ≥ 15% of mass) |
Compact edge deployment below stock IQ4_XS. |
| Pico | 24% | Exact-budget DP (IQ3_S–IQ4_XS base, scarce-attention band protected) |
Edge budget; default on state-rich ≥3B trunks, explicit elsewhere (PARETRIX_INCLUDE_PICO=1). |
| Femto | 21% | Exact-budget DP, deep-squeeze product (IQ3_XXS base, measured-rate squeezes) |
Deepest service rung; campaign product, default on state-rich ≥3B trunks. |
Stock twins. On topologies where a flat recipe's levers have no effect under the regex rules, the Paretrix tier is byte-identical to its stock twin (detected automatically). The suite then queues one self re-shuffle campaign anchored on the twin. Measured outcome: the re-shuffle wins below the top band and ties at it (R19).
🧮 THE MATRIX — measured rate knowledge
Three pricing levels, each overlaying the next, class by class:
- Model cache (
model-paretrix.json): this model's own measured tables, anchored on the operating point nearest the target (R20). - Arch line (
arch-paretrix/<arch>-paretrix.json): shared per-architecture knowledge, including banded mid/deep sell and buy tables, structural facts, build corrections, and witnesses. Size-normalized (rate × line_reference / model_size), so checkpoints of different scales share one line. - Universal prior (
arch-paretrix/generic-paretrix.json): the distilled six-family prior, also banded and size-aware.
The fold (Paretrix.py matrix) compiles every discoverable model cache into the arch lines:
- Sells price at the measured range maximum, which stays conservative for every cut.
- Buys price at the minimum of the range lows, which stays conservative for every gain.
- Dead buy classes price at zero (R13), so the DP keeps the measured zero instead of falling back to the MSE proxy.
⚔️ Campaign engine (RCO)
One command per rung: anchor → notch probes → measured rates → exact-budget DP → build → duel at the verdict regime → adopt.
python Paretrix.py campaign --model model-BF16.gguf --imatrix imatrix.gguf \
--anchor-from-artifact model-Q5_K_M-imx.gguf --target-mib 1062 --run
python Paretrix.py duel --model model-BF16.gguf --tier quality \
--candidate model-Paretrix-Candidate-36pc.gguf --adopt
The DP solves the multiple-choice knapsack to the exact MiB, with no constraint hyperparameter. Squeezes below the floors are priced from the measured notch rates; raises above the anchor are priced from the measured buy table, with dead buys clamped to zero. The duel binds at ctx 4096×8 and adopts on a KLD gain at or above 1×MASD (floor 0.001), the same gate the Pareto column reads.
📊 Benchmarks
All sweeps run at the verdict regime ctx 4096×8 (32,768-token span) on the corpus wiki.test.raw with Flash-Attention. The KL reference is the BF16 logits of the same model. Compare KLD only within a single regime.
Models run from largest to smallest by BF16 size, and each table runs from largest to smallest MiB.
Pareto column: ★ = on this sweep's (MiB, KLD) frontier · < X = this row is dominated by X (X is lighter and better) · ≈ X = X is lighter, with KLD inside the tie zone max(MASD, 0.001) · ≡ = byte-identical twin.
1. ornith-ai/Ornith-1.5-9B — qwen35 GDN hybrid · untied 248k vocab
BF16 17,091 MiB · target: 8+ GB VRAM (split offload) · PPL reference 8.9219 · DFlash draft: ×1.34–×2.14 net speedup.
| Model | MiB | PPL | ΔPPL | KLD | RMS Δp | top-p | Pareto |
|---|---|---|---|---|---|---|---|
| Q8_0 (stock) | 9086 | 8.9544 | +0.0325 | 0.0118 | 2.67% | 97.7% | ★ |
| Fidelity-48pc | 8232 | 8.9700 | +0.0481 | 0.0167 | 3.13% | 97.0% | ≈ Precision-42pc (−0.0011 · −854 MiB) |
| Precision-42pc | 7378 | 8.9365 | +0.0146 | 0.0156 | 3.24% | 96.6% | ★ |
| Q6_K-imx (stock) | 7018 | 8.7000† | −0.2219 | 0.0251 | 3.83% | 95.9% | ★ |
| Quality-36pc | 6234 | 8.5716† | −0.3503 | 0.0449 | 4.88% | 93.5% | ★ |
| Q5_K_M-imx (stock) | 6168 | 8.2043† | −0.7176 | 0.1013 | 7.29% | 90.2% | < Compact-33pc |
| Compact-33pc | 5787 | 8.4619† | −0.4600 | 0.0610 | 6.07% | 92.0% | ★ |
| Mini-30pc ⭐ | 5368 | 8.6772† | −0.2447 | 0.0679 | 6.42% | 91.1% | ★ |
| IQ4_XS-imx (stock) | 4956 | 9.2939 | +0.3720 | 0.0794 | 7.08% | 90.8% | ★ |
| Nano-27pc | 4586 | 8.6546† | −0.2673 | 0.1223 | 8.73% | 86.5% | ★ |
| IQ3_M-imx (stock) | 4211 | 9.1809 | +0.2590 | 0.1807 | 11.04% | 84.5% | < Pico-24pc |
| Pico-24pc | 4070 | 8.5303† | −0.3916 | 0.1540 | 10.08% | 84.7% | ★ |
| Femto-21pc | 3601 | 8.3140† | −0.6079 | 0.2524 | 12.37% | 80.6% | ★ |
Two strict dominations, the program's largest: Compact (−0.0403 KLD, −381 MiB) and Pico (−0.0268 KLD, −141 MiB). Fidelity ties Precision within the tie zone while weighing 854 MiB more. At the top band, the extra Q8_0 pockets buy nothing measurable (L16).
2. TokenRhythm/NeoHorse-1-4B — qwen3.5 fine-tune · GDN hybrid · tied
BF16 8,034 MiB · target: 4+ GB VRAM · PPL reference 9.1268 · L19 fine-tune field (7/7 dead buys at mid anchors).
| Model | MiB | PPL | ΔPPL | KLD | RMS Δp | top-p | Pareto |
|---|---|---|---|---|---|---|---|
| Q8_0 (stock) | 4275 | 9.1273 | +0.0005 | 0.0057 | 1.82% | 98.2% | ★ |
| Fidelity-48pc | 3866 | 9.1807 | +0.0539 | 0.0083 | 2.45% | 97.5% | ≈ Precision-42pc (+0.0002 · −494 MiB) |
| Precision-42pc | 3372 | 9.1828 | +0.0560 | 0.0084 | 2.82% | 97.0% | ≈ Q6_K-imx (+0.0009 · −68 MiB) |
| Q6_K-imx (stock) | 3304 | 9.1991 | +0.0724 | 0.0093 | 2.90% | 96.7% | ★ |
| Quality-36pc | 2966 | 9.2173 | +0.0905 | 0.0204 | 3.71% | 95.1% | ★ |
| Q5_K_M-imx (stock) | 2933 | 9.3403 | +0.2135 | 0.0270 | 4.32% | 94.3% | ★ |
| Compact-33pc | 2768 | 9.1505 | +0.0237 | 0.0300 | 4.71% | 93.1% | ★ |
| Mini-30pc | 2416 | 9.5442 | +0.4174 | 0.0537 | 6.02% | 90.9% | ★ |
| IQ4_XS-imx (stock) | 2398 | 9.5261 | +0.3994 | 0.0598 | 6.37% | 90.5% | ★ |
| Nano-27pc ⭐ | 2251 | 9.2769 | +0.1501 | 0.0683 | 6.85% | 89.7% | ★ |
| IQ3_M-imx (stock) | 2063 | 9.9031 | +0.7763 | 0.1480 | 10.62% | 84.8% | ★ |
| Pico-24pc | 1935 | 10.1509 | +1.0242 | 0.1570 | 10.92% | 84.6% | ★ |
| Femto-21pc | 1696 | 10.0878 | +0.9610 | 0.2534 | 13.98% | 79.3% | ★ |
Precision matches Fidelity within the tie zone at 494 MiB less. The fine-tune field flattens sells (×0.3) and leaves most mid-band buys dead (L19), so campaign value concentrates in deep sheds and self re-shuffles.
3. Nanbeige/Nanbeige4.2-3B — looped dense · untied 166k vocab
BF16 7,957 MiB · target: 6+ GB VRAM · PPL reference 29.8955 · readout pair 24.5% of mass · R3 loop-neutral depth prior active · DSpark draft: ×1.87/×2.13 net speedup.
| Model | MiB | PPL | ΔPPL | KLD | RMS Δp | top-p | Pareto |
|---|---|---|---|---|---|---|---|
| Q8_0 (stock) | 4229 | 31.5657 | +1.6702 | 0.0144 | 3.01% | 95.5% | ★ |
| Fidelity-48pc | 3821 | 31.8994 | +2.0039 | 0.0266 | 3.97% | 93.8% | ★ |
| Precision-42pc | 3266 | 32.0971 | +2.2016 | 0.0467 | 4.93% | 91.9% | ≡ Q6_K-imx |
| Q6_K-imx (stock) | 3266 | 32.0971 | +2.2016 | 0.0467 | 4.93% | 91.9% | ≡ Precision-42pc |
| Q5_K_M-imx (stock) | 2849 | 32.0553 | +2.1598 | 0.0872 | 6.58% | 88.2% | < Quality-36pc |
| Quality-36pc ⭐ | 2846 | 31.7395 | +1.8440 | 0.0714 | 5.86% | 89.2% | ★ |
| Compact-33pc | 2629 | 31.7929 | +1.8974 | 0.1115 | 7.53% | 86.5% | ★ |
| Mini-30pc | 2388 | 30.7090 | +0.8135 | 0.1841 | 9.43% | 82.8% | ★ |
| IQ4_XS-imx (stock) | 2268 | 32.7700 | +2.8745 | 0.2093 | 10.24% | 81.7% | ★ |
| Nano-27pc | 2152 | 31.3016 | +1.4061 | 0.2287 | 10.51% | 80.4% | ★ |
| IQ3_M-imx (stock) | 1985 | 33.2114 | +3.3158 | 0.4948 | 15.12% | 71.2% | < Pico-24pc |
| Pico-24pc | 1906 | 30.0424† | +0.1469 | 0.4023 | 14.04% | 73.2% | ★ |
| Femto-21pc | 1675 | 36.4632 | +6.5677 | 0.5857 | 16.35% | 68.3% | ★ |
Two strict dominations: Quality (−0.0158 KLD, −3 MiB) and Pico (−0.0925 KLD, −79 MiB). Precision is byte-identical to Q6_K-imx (detected by R19).
4. XHToken/Spark-X2.5-4B — dense SWA hybrid 3:1 · tied 131k vocab
BF16 7,849 MiB · target: 6+ GB VRAM · PPL reference 20.1134 · the verdict binds at native context (L18).
| Model | MiB | PPL | ΔPPL | KLD | RMS Δp | top-p | Pareto |
|---|---|---|---|---|---|---|---|
| Q8_0 (stock) | 4172 | 19.9740 | −0.1394 | 0.0047 | 1.60% | 97.3% | ★ |
| Fidelity-48pc | 3772 | 20.0099 | −0.1035 | 0.0115 | 2.49% | 95.3% | ★ |
| Precision-42pc | 3359 | 19.7049 | −0.4085 | 0.0128 | 2.66% | 95.0% | ★ |
| Q6_K-imx (stock) | 3223 | 19.6832 | −0.4303 | 0.0173 | 3.02% | 93.8% | ★ |
| Q5_K_M-imx (stock) | 2840 | 20.8885 | +0.7751 | 0.0571 | 5.57% | 89.4% | < Quality-36pc |
| Quality-36pc ⭐ | 2838 | 20.0193 | −0.0941 | 0.0437 | 4.70% | 90.4% | ★ |
| Compact-33pc | 2587 | 20.6036 | +0.4902 | 0.0976 | 7.16% | 86.1% | ★ |
| Mini-30pc | 2343 | 21.3848 | +1.2714 | 0.1203 | 8.06% | 84.6% | ★ |
| IQ4_XS-imx (stock) | 2266 | 23.2522 | +3.1388 | 0.1678 | 9.45% | 81.8% | ★ |
| Nano-27pc | 2123 | 22.3853 | +2.2719 | 0.2498 | 11.35% | 78.4% | ★ |
| Pico-24pc | 1982 | 19.4291† | −0.6843 | 0.2812 | 12.30% | 76.9% | ★ |
| IQ3_M-imx (stock) | 1949 | 21.6523 | +1.5389 | 0.3635 | 14.23% | 73.7% | ★ |
| Femto-21pc | 1664 | 26.5109 | +6.3975 | 0.5799 | 17.22% | 67.7% | ★ |
Quality dominates Q5_K_M-imx (−0.0134 KLD, −2 MiB). The spark2_5 corrections (readout lever L12, F32 gate pin R9, ffn_down Q4_K lever) apply to every build; the scarce 9-layer full-attention band is the family's richest buy.
5. webAI-Official/TwIL-LM3-Pro — granite dense · untied 128k vocab
BF16 6,984 MiB · target: 6+ GB VRAM · PPL reference 11.4980 · R19 reference case.
| Model | MiB | PPL | ΔPPL | KLD | RMS Δp | top-p | Pareto |
|---|---|---|---|---|---|---|---|
| Q8_0 (stock) | 3712 | 11.5097 | +0.0117 | 0.0020 | 1.33% | 97.9% | ≈ Fidelity-48pc (+0.0007 · −357 MiB) |
| Fidelity-48pc | 3355 | 11.5074 | +0.0094 | 0.0027 | 1.38% | 97.4% | ★ |
| Precision-42pc | 2867 | 11.5042 | +0.0062 | 0.0049 | 1.99% | 96.6% | ≡ Q6_K-imx |
| Q6_K-imx (stock) | 2867 | 11.5042 | +0.0062 | 0.0049 | 1.99% | 96.6% | ≡ Precision-42pc |
| Quality-36pc | 2514 | 11.6187 | +0.1207 | 0.0104 | 2.74% | 95.1% | ★ |
| Q5_K_M-imx (stock) | 2493 | 11.6684 | +0.1704 | 0.0131 | 3.33% | 94.5% | ★ |
| Compact-33pc | 2306 | 11.6594 | +0.1614 | 0.0218 | 3.90% | 93.2% | ★ |
| Mini-30pc | 2096 | 11.6689 | +0.1709 | 0.0300 | 4.59% | 91.6% | ★ |
| IQ4_XS-imx (stock) | 1937 | 11.8072 | +0.3092 | 0.0350 | 4.92% | 91.1% | ★ |
| Nano-27pc ⭐ | 1885 | 11.9953 | +0.4973 | 0.0479 | 5.64% | 89.4% | ★ |
| IQ3_M-imx (stock) | 1653 | 12.4715 | +0.9735 | 0.1052 | 8.63% | 84.7% | ★ |
R19 case: Quality-36pc and Precision-42pc started byte-identical to their stock twins. The self re-shuffle on Quality won −20.7% (0.0131 → 0.0104, 2.7× the duel gate). On Precision it tied (+0.0007, inside the gate, so the incumbent stands). The re-shuffle wins or ties by construction.
6. LiquidAI/LFM2.5-2.6B — shortconv mixer · tied readout
BF16 5,153 MiB · target: 4+ GB VRAM · PPL reference 51.4190 · DSpark draft: ×1.10 net speedup.
| Model | MiB | PPL | ΔPPL | KLD | RMS Δp | top-p | Pareto |
|---|---|---|---|---|---|---|---|
| Q8_0 (stock) | 2742 | 51.8867 | +0.4677 | 0.0032 | 1.34% | 97.4% | ★ |
| Fidelity-48pc | 2481 | 52.2324 | +0.8134 | 0.0076 | 2.09% | 96.2% | ★ |
| Precision-42pc | 2138 | 52.8312 | +1.4122 | 0.0105 | 2.44% | 95.4% | ★ |
| Q6_K-imx (stock) | 2119 | 52.7076 | +1.2886 | 0.0131 | 2.66% | 94.7% | ★ |
| Q5_K_M-imx (stock) | 1850 | 51.5220 | +0.1030 | 0.0380 | 4.39% | 91.0% | < Quality-36pc |
| Quality-36pc ⭐ | 1848 | 52.6558 | +1.2368 | 0.0309 | 4.13% | 91.6% | ★ |
| Compact-33pc | 1712 | 49.4337 | −1.9853 | 0.0696 | 6.13% | 88.0% | ★ |
| Mini-30pc | 1560 | 46.9409† | −4.4780 | 0.0995 | 7.45% | 85.9% | ★ |
| IQ4_XS-imx (stock) | 1447 | 52.6602 | +1.2412 | 0.1652 | 9.56% | 81.6% | < Nano-27pc |
| Nano-27pc | 1405 | 52.7926 | +1.3736 | 0.1536 | 8.91% | 81.8% | ★ |
| Pico-24pc | 1273 | 55.9255 | +4.5065 | 0.3597 | 13.43% | 74.6% | ★ |
| IQ3_M-imx (stock) | 1225 | 60.9679 | +9.5490 | 0.4254 | 14.31% | 72.6% | ★ |
| Femto-21pc | 1096 | 71.9524 | +20.5334 | 0.5535 | 16.29% | 69.0% | ★ |
Quality dominates Q5_K_M-imx (−0.0071 KLD, −2 MiB), and Nano dominates IQ4_XS-imx (−0.0116 KLD, −42 MiB). † PPL below the BF16 base indicates entropy collapse, not a quality gain.
7. openbmb/MiniCPM5-2B — dense llama-arch · untied 130k readout
BF16 4,806 MiB · target: 4+ GB VRAM · PPL reference 11.9212 · DSpark draft: acceptance 0.4419, ×1.89 net speedup.
| Model | MiB | PPL | ΔPPL | KLD | RMS Δp | top-p | Pareto |
|---|---|---|---|---|---|---|---|
| Q8_0 (stock) | 2556 | 11.9393 | +0.0181 | 0.0015 | 0.99% | 97.8% | ★ |
| Fidelity-48pc | 2309 | 11.9598 | +0.0386 | 0.0033 | 1.47% | 96.8% | ★ |
| Q6_K-imx (stock) | 1974 | 11.9724 | +0.0512 | 0.0059 | 1.91% | 95.9% | ≈ Precision-42pc (−0.0002 · −1 MiB) |
| Precision-42pc | 1973 | 11.9678 | +0.0466 | 0.0057 | 1.89% | 96.0% | ★ |
| Q5_K_M-imx (stock) | 1724 | 12.0905 | +0.1693 | 0.0195 | 3.58% | 92.8% | < Quality-36pc |
| Quality-36pc | 1723 | 12.1180 | +0.1968 | 0.0175 | 3.43% | 93.3% | ★ |
| Compact-33pc ⭐ | 1593 | 12.2219 | +0.3007 | 0.0427 | 5.12% | 89.2% | ★ |
| Mini-30pc | 1446 | 12.3731 | +0.4519 | 0.0672 | 6.53% | 86.7% | ★ |
| IQ4_XS-imx (stock) | 1358 | 12.3977 | +0.4765 | 0.0722 | 6.69% | 86.9% | ★ |
| Nano-27pc | 1306 | 12.6530 | +0.7318 | 0.0848 | 7.26% | 85.7% | ★ |
| IQ3_M-imx (stock) | 1170 | 13.9101 | +1.9889 | 0.2027 | 11.74% | 78.6% | < Pico-24pc |
| Pico-24pc | 1154 | 13.9155 | +1.9943 | 0.1719 | 10.52% | 79.3% | ★ |
| Femto-21pc | 1016 | 15.2794 | +3.3582 | 0.2918 | 14.25% | 73.9% | ★ |
Two strict dominations: Quality (−0.0020 KLD, −1 MiB) and Pico (−0.0308 KLD, −15 MiB). All eight rungs held through the L17 self re-shuffle wave.
8. jinaai/ReaderLM-v2 — dense qwen2 · tied 152k vocab · HTML/Markdown extractor
BF16 2,950 MiB · target: 2+ GB VRAM · PPL reference 12.3852 · FFN holds 75% of the mass.
| Model | MiB | PPL | ΔPPL | KLD | RMS Δp | top-p | Pareto |
|---|---|---|---|---|---|---|---|
| Q8_0 (stock) | 1570 | 12.4136 | +0.0284 | 0.0016 | 0.99% | 97.9% | ★ |
| Fidelity-48pc | 1422 | 12.4161 | +0.0309 | 0.0031 | 1.39% | 96.9% | ★ |
| Precision-42pc | 1214 | 12.3712† | −0.0140 | 0.0055 | 1.84% | 95.9% | ≡ Q6_K-imx |
| Q6_K-imx (stock) | 1214 | 12.3712† | −0.0140 | 0.0055 | 1.84% | 95.9% | ≡ Precision-42pc |
| Quality-36pc | 1073 | 12.4244 | +0.0392 | 0.0165 | 3.25% | 93.1% | ≡ Q5_K_M-imx |
| Q5_K_M-imx (stock) | 1073 | 12.4244 | +0.0392 | 0.0165 | 3.25% | 93.1% | ≡ Quality-36pc |
| Compact-33pc ⭐ | 979 | 12.4617 | +0.0765 | 0.0309 | 4.35% | 90.9% | ★ |
| Mini-30pc | 890 | 12.6359 | +0.2507 | 0.0553 | 5.95% | 87.7% | ≈ IQ4_XS-imx (+0.0027 · −36 MiB) |
| IQ4_XS-imx (stock) | 854 | 12.5393 | +0.1541 | 0.0580 | 5.96% | 87.6% | ★ |
| Nano-27pc | 802 | 12.6418 | +0.2566 | 0.0887 | 7.39% | 84.9% | ★ |
| IQ3_M-imx (stock) | 741 | 13.3774 | +0.9922 | 0.1706 | 10.37% | 79.6% | ★ |
Pico and Femto are omitted: the constructibility gate applies (tied vocab sits at 15% of the mass, the R6-ter boundary). Compact, Mini and Nano are campaign products. Quality and Precision are byte-identical to their stock twins (detected by R19). L2 validated: the Q4_K ffn_down floor restored monotonicity across the 75% FFN mass.
9. SeerRay-Lab/Xiaomi-OCR-0 — qwen35 GDN hybrid · tied 248k vocab
BF16 1,446 MiB · target: 2+ GB VRAM · PPL reference 28.5189 · vision via the mmproj module.
| Model | MiB | PPL | ΔPPL | KLD | RMS Δp | top-p | Pareto |
|---|---|---|---|---|---|---|---|
| Q8_0 (stock) | 774 | 28.5404 | +0.0215 | 0.0010 | 0.77% | 98.1% | ★ |
| Fidelity-48pc | 704 | 28.5178 | −0.0011 | 0.0022 | 1.12% | 97.0% | ★ |
| Precision-42pc | 611 | 28.5667 | +0.0477 | 0.0050 | 1.68% | 95.4% | ≈ Q6_K-imx (+0.0003 · −10 MiB) |
| Q6_K-imx (stock) | 601 | 28.6110 | +0.0920 | 0.0052 | 1.74% | 95.4% | ★ |
| Q5_K_M-imx (stock) | 551 | 28.9326 | +0.4136 | 0.0111 | 2.51% | 93.7% | ★ |
| Quality-36pc | 538 | 28.9461 | +0.4271 | 0.0131 | 2.72% | 92.8% | ★ |
| Compact-33pc ⭐ | 487 | 29.3475 | +0.8285 | 0.0289 | 3.70% | 90.1% | ★ |
| IQ4_XS-imx (stock) | 481 | 29.6333 | +1.1144 | 0.0425 | 4.81% | 87.7% | ★ |
| Mini-30pc | 444 | 29.8250 | +1.3060 | 0.0521 | 5.18% | 86.7% | ★ |
| IQ3_M-imx (stock) | 433 | 32.1805 | +3.6615 | 0.1050 | 8.26% | 81.0% | ★ |
Pico and Femto are omitted: the constructibility gate applies (tied vocab is 34% of the mass, and the Q6_K floor alone reaches 51% of the Nano target). Mini-30 is the family floor. Precision ≈ Q6_K-imx within the tie zone (C2). The pipeline skips this rung on small qwen35 by default; set PARETRIX_PRECISION_TWIN=0 to force it.
📉 Sub-Nano compendium — the Nano → Pico step
Families run from largest to smallest.
| Family | Architecture | Nano KLD | Pico KLD | Step | Pico verdict |
|---|---|---|---|---|---|
| Ornith-1.5-9B | GDN hybrid, untied | 0.1223 | 0.1540 | +26% | ✅ Fully usable (beats IQ3_M 0.1807) |
| NeoHorse-1-4B | GDN fine-tune, tied | 0.0683 | 0.1570 | +130% | ◐ Context variant (IQ3_M reaches 0.1480 at +128 MiB) |
| Nanbeige4.2-3B | Looped dense, untied | 0.2287 | 0.4023 | +76% | ✗ Degraded (Nano already past the 0.20 edge) |
| Spark-X2.5-4B | SWA hybrid, tied | 0.2498 | 0.2812 | +13% | ✗ Degraded (floor = Mini) |
| TwIL-LM3-Pro | Dense, untied | 0.0479 | — | — | ⊘ Ladder stops at Nano |
| LFM2.5-2.6B | Shortconv mixer, tied | 0.1536 | 0.3597 | +134% | ✗ Pico exceeds 0.20 (floor = Mini) |
| MiniCPM5-2B | Dense, untied | 0.0848 | 0.1719 | +103% | ✅ Usable (beats IQ3_M 0.2027) |
| ReaderLM-v2 | Dense, tied | 0.0887 | — | — | ⊘ Constructibility gate (floor = Nano) |
| Xiaomi-OCR-0 | GDN hybrid, tied | — | — | — | ⊘ Constructibility gate (floor = Mini) |
Summary — 9 strict Pareto dominations
Products that are both lighter and better than their stock twin at the verdict regime. Families run from largest to smallest; tiers run from largest to smallest within each family.
| Family | Product | KLD | Stock twin | Stock KLD | ΔKLD | ΔMiB |
|---|---|---|---|---|---|---|
| Ornith-1.5-9B | Compact-33pc | 0.0610 | Q5_K_M-imx | 0.1013 | −0.0403 | −381 |
| Ornith-1.5-9B | Pico-24pc | 0.1540 | IQ3_M-imx | 0.1807 | −0.0268 | −141 |
| Nanbeige4.2-3B | Quality-36pc | 0.0714 | Q5_K_M-imx | 0.0872 | −0.0158 | −3 |
| Nanbeige4.2-3B | Pico-24pc | 0.4023 | IQ3_M-imx | 0.4948 | −0.0925 | −79 |
| Spark-X2.5-4B | Quality-36pc | 0.0437 | Q5_K_M-imx | 0.0571 | −0.0134 | −2 |
| LFM2.5-2.6B | Quality-36pc | 0.0309 | Q5_K_M-imx | 0.0380 | −0.0071 | −2 |
| LFM2.5-2.6B | Nano-27pc | 0.1536 | IQ4_XS-imx | 0.1652 | −0.0116 | −42 |
| MiniCPM5-2B | Quality-36pc | 0.0175 | Q5_K_M-imx | 0.0195 | −0.0020 | −1 |
| MiniCPM5-2B | Pico-24pc | 0.1719 | IQ3_M-imx | 0.2027 | −0.0308 | −15 |
Plus four byte-identical stock twins detected (R19): Nanbeige Precision ≡ Q6_K-imx · TwIL Precision ≡ Q6_K-imx · ReaderLM Precision ≡ Q6_K-imx · ReaderLM Quality ≡ Q5_K_M-imx.
⚡ Speculative decoding: MTP · DSpark · DFlash
Paretrix quantizes speculative modules imatrix-free (Paretrix-modules.py). A draft's input is the parent's hidden state, the MTP head rides the trunk's activations, and the vision tower consumes raw pixels, so none of them exposes a corpus entry point. Speculative verification bounds draft damage to the acceptance rate and leaves output fidelity intact (L11).
| Module | Role | Palette | Naming |
|---|---|---|---|
| MTP | Multi-token prediction fusion head (extracted by 01) | Role floors + K-quants | mtp-Paretrix-<Profile>.gguf |
| DSpark | Block-diffusion speculative draft (Markov + confidence heads) | Role floors + K-quants | dspark-Paretrix-<Profile>.gguf |
| DFlash | Third-party block-diffusion draft | Role floors + K-quants | dflash-Paretrix-<Profile>.gguf |
| MMProj | CLIP vision projector | Profile grid, critical tensors pinned F32 | mmproj-Paretrix-<Profile>.gguf |
Measured acceptance invariance (L11). Draft quantization moves acceptance by no more than ±0.03: nanbeige 0.3988 → 0.4092 (Δ +0.0104) · MiniCPM5-2B 0.4419 (ledger 0.4464) · lfm2 +0.0055 · Ornith DFlash 0.2631 → 0.2684 at matched ngl. The quantized draft also runs faster: net speedups of ×1.87/×2.13 (nanbeige), ×1.89 (MiniCPM5-2B), and ×1.34–×2.14 (Ornith DFlash, 106.9 t/s on a 9B model at 8 GB VRAM). Acceptance is a property of the (draft, target) pair and the offload regime (L20), so compare acceptance only at matched --ngl.
# Acceptance battery (llama-server A/B, temperature 0, trained block auto-read)
python 13_draft-acceptance-sweep.py <working-folder> --chat --tokenizer-dir <hf-folder>
# Fuse model + modules into one deployment GGUF
python 14_gguf-module-fusion.py fused.gguf model-BF16.gguf mmproj-BF16.gguf
⚡ Quick start
# One command: convert → imatrix → full Paretrix ladder [+ classic line]
python 00_Paretrix-pipeline.py ./MyModel [--all]
# Or step by step
python 01_SAFETENSORS-to-BF16-GGUF.py ./MyModel # → Paretrix/model-BF16.gguf
python 02_BF16-GGUF-to-Q8-imatrix.py # → model-Q8_0.gguf + imatrix.gguf
python 03_BF16-GGUF-to-Paretrix.py [--all] # → full ladder + duels
# Campaign on a specific rung (anchor → probes → DP → build → duel → adopt)
python Paretrix.py campaign --model model-BF16.gguf --imatrix imatrix.gguf \
--anchor-from-artifact model-Q5_K_M-imx.gguf --target-mib 1062 --run
# Sweep every product at the verdict regime
python 11_perplexity-test.py
# Fold the model caches into the shared arch lines
python Paretrix.py matrix
Working-root convention. A model folder holds its HF repo (safetensors + sidecars) plus a Paretrix/ subfolder carrying every product: BF16/Q8/imatrix GGUFs, champions, modules, and model-paretrix.json.
<model folder>/
├─ model-00001-of-00002.safetensors · config.json · tokenizer.json …
└─ Paretrix/
├─ model-BF16.gguf · model-Q8_0.gguf · imatrix.gguf
├─ model-Paretrix-<Tier>-<XX>pc.gguf · mtp/mmproj/dspark products
└─ model-paretrix.json
Full operational detail: DOCUMENTATION.md.
📜 Citation & Credits
- Paretrix Quantization Suite: the Pareto-Matrix rate-exchange framework, with measured notch rates, exact-budget DP allocation, and a cross-architecture knowledge base (THE MATRIX).
- llama.cpp by Georgi Gerganov & ggml contributors: GGUF/GGML runtime,
llama-imatrix,llama-quantize,llama-perplexity, andllama-server. - GSQ / RCO (IST-DASLab): GSQ: Gumbel-STE quantization (arXiv:2604.18556) · RCO: relaxed knapsack allocation (arXiv:2605.00649). The exact-budget DP and Gumbel-STE manifold search build on these lines, adapted to measured rates.
- wepiqx ASHQ1: priority-queue knapsack formulation, tied-group activation hashing, and MSE scheduling.
- Empero AI: GDN state preservation (
ssm_alpha/ssm_betaat Q8_0).