Paretrix Quantization Suite (v2.6.0)

Engine Verdict regime Dominations

Paretrix is an empirical, activation-aware quantization suite for GGUF language models and speculative decoding modules.

Pareto + Matrix → Paretrix. Each tier is priced so that no better trade exists at that compression. The suite measures real activation sensitivity (ΔKLD/MiB) per tensor class through llama-imatrix, learns rate tables from cross-architecture campaigns (THE MATRIX), and allocates bitwidths under exact budget targets: flat recipes where uniformity wins, rate-calibrated knapsack where heterogeneity pays.

Standard uniform quantization applies one bitwidth family across the whole network. Paretrix reads the model's own activation field (which classes are expensive to cut, which are nearly free, and where depth matters), then spends the budget where measurements show the greatest return.


🪜 The Mod-3 Ladder

Each tier lands on a multiple of 3 percent of the unquantized BF16 source, so the shipping name states its weight band (model-Paretrix-<Tier>-<XX>pc.gguf, ±1.5 pp). Tiers run from largest to smallest.

Tier Ratio Allocation Strategy Primary Purpose
Fidelity 48% Q6_K base + Q8_0 pockets biased to the late layer third (G4) High-fidelity archival tier.
Precision 42% Flat Q6_K + full-attention band Q8_0; the readout lever is arch-scoped (L12) Coding, mathematics, high-entropy reasoning.
Quality 36% Flat Q5_K_M + imatrix (L6); self re-shuffle where the twin is a stock copy General production tier.
Compact 33% Exact-budget DP, or flat Q4_K_M + band (Q6_K) + recurrent floors Daily driver for 2B–4B architectures.
Mini 30% Exact-budget DP, or L4 flat on GDN hybrids (IQ4_XS + recurrent floors + full-attention band) Balanced entry tier for ≥ 9B models on 8 GB VRAM.
Nano 27% Exact-budget DP (IQ3_S–IQ4_XS base, tied readout Q5_K when vocab ≥ 15% of mass) Compact edge deployment below stock IQ4_XS.
Pico 24% Exact-budget DP (IQ3_S–IQ4_XS base, scarce-attention band protected) Edge budget; default on state-rich ≥3B trunks, explicit elsewhere (PARETRIX_INCLUDE_PICO=1).
Femto 21% Exact-budget DP, deep-squeeze product (IQ3_XXS base, measured-rate squeezes) Deepest service rung; campaign product, default on state-rich ≥3B trunks.

Stock twins. On topologies where a flat recipe's levers have no effect under the regex rules, the Paretrix tier is byte-identical to its stock twin (detected automatically). The suite then queues one self re-shuffle campaign anchored on the twin. Measured outcome: the re-shuffle wins below the top band and ties at it (R19).


🧮 THE MATRIX — measured rate knowledge

Three pricing levels, each overlaying the next, class by class:

  1. Model cache (model-paretrix.json): this model's own measured tables, anchored on the operating point nearest the target (R20).
  2. Arch line (arch-paretrix/<arch>-paretrix.json): shared per-architecture knowledge, including banded mid/deep sell and buy tables, structural facts, build corrections, and witnesses. Size-normalized (rate × line_reference / model_size), so checkpoints of different scales share one line.
  3. Universal prior (arch-paretrix/generic-paretrix.json): the distilled six-family prior, also banded and size-aware.

The fold (Paretrix.py matrix) compiles every discoverable model cache into the arch lines:

  • Sells price at the measured range maximum, which stays conservative for every cut.
  • Buys price at the minimum of the range lows, which stays conservative for every gain.
  • Dead buy classes price at zero (R13), so the DP keeps the measured zero instead of falling back to the MSE proxy.

⚔️ Campaign engine (RCO)

One command per rung: anchor → notch probes → measured rates → exact-budget DP → build → duel at the verdict regime → adopt.

python Paretrix.py campaign --model model-BF16.gguf --imatrix imatrix.gguf \
    --anchor-from-artifact model-Q5_K_M-imx.gguf --target-mib 1062 --run
python Paretrix.py duel --model model-BF16.gguf --tier quality \
    --candidate model-Paretrix-Candidate-36pc.gguf --adopt

The DP solves the multiple-choice knapsack to the exact MiB, with no constraint hyperparameter. Squeezes below the floors are priced from the measured notch rates; raises above the anchor are priced from the measured buy table, with dead buys clamped to zero. The duel binds at ctx 4096×8 and adopts on a KLD gain at or above 1×MASD (floor 0.001), the same gate the Pareto column reads.


📊 Benchmarks

All sweeps run at the verdict regime ctx 4096×8 (32,768-token span) on the corpus wiki.test.raw with Flash-Attention. The KL reference is the BF16 logits of the same model. Compare KLD only within a single regime.

Models run from largest to smallest by BF16 size, and each table runs from largest to smallest MiB.

Pareto column: ★ = on this sweep's (MiB, KLD) frontier · < X = this row is dominated by X (X is lighter and better) · ≈ X = X is lighter, with KLD inside the tie zone max(MASD, 0.001) · ≡ = byte-identical twin.


1. ornith-ai/Ornith-1.5-9B — qwen35 GDN hybrid · untied 248k vocab

BF16 17,091 MiB · target: 8+ GB VRAM (split offload) · PPL reference 8.9219 · DFlash draft: ×1.34–×2.14 net speedup.

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 9086 8.9544 +0.0325 0.0118 2.67% 97.7% ★
Fidelity-48pc 8232 8.9700 +0.0481 0.0167 3.13% 97.0% ≈ Precision-42pc (−0.0011 · −854 MiB)
Precision-42pc 7378 8.9365 +0.0146 0.0156 3.24% 96.6% ★
Q6_K-imx (stock) 7018 8.7000† −0.2219 0.0251 3.83% 95.9% ★
Quality-36pc 6234 8.5716† −0.3503 0.0449 4.88% 93.5% ★
Q5_K_M-imx (stock) 6168 8.2043† −0.7176 0.1013 7.29% 90.2% < Compact-33pc
Compact-33pc 5787 8.4619† −0.4600 0.0610 6.07% 92.0% ★
Mini-30pc ⭐ 5368 8.6772† −0.2447 0.0679 6.42% 91.1% ★
IQ4_XS-imx (stock) 4956 9.2939 +0.3720 0.0794 7.08% 90.8% ★
Nano-27pc 4586 8.6546† −0.2673 0.1223 8.73% 86.5% ★
IQ3_M-imx (stock) 4211 9.1809 +0.2590 0.1807 11.04% 84.5% < Pico-24pc
Pico-24pc 4070 8.5303† −0.3916 0.1540 10.08% 84.7% ★
Femto-21pc 3601 8.3140† −0.6079 0.2524 12.37% 80.6% ★

Two strict dominations, the program's largest: Compact (−0.0403 KLD, −381 MiB) and Pico (−0.0268 KLD, −141 MiB). Fidelity ties Precision within the tie zone while weighing 854 MiB more. At the top band, the extra Q8_0 pockets buy nothing measurable (L16).


2. TokenRhythm/NeoHorse-1-4B — qwen3.5 fine-tune · GDN hybrid · tied

BF16 8,034 MiB · target: 4+ GB VRAM · PPL reference 9.1268 · L19 fine-tune field (7/7 dead buys at mid anchors).

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 4275 9.1273 +0.0005 0.0057 1.82% 98.2% ★
Fidelity-48pc 3866 9.1807 +0.0539 0.0083 2.45% 97.5% ≈ Precision-42pc (+0.0002 · −494 MiB)
Precision-42pc 3372 9.1828 +0.0560 0.0084 2.82% 97.0% ≈ Q6_K-imx (+0.0009 · −68 MiB)
Q6_K-imx (stock) 3304 9.1991 +0.0724 0.0093 2.90% 96.7% ★
Quality-36pc 2966 9.2173 +0.0905 0.0204 3.71% 95.1% ★
Q5_K_M-imx (stock) 2933 9.3403 +0.2135 0.0270 4.32% 94.3% ★
Compact-33pc 2768 9.1505 +0.0237 0.0300 4.71% 93.1% ★
Mini-30pc 2416 9.5442 +0.4174 0.0537 6.02% 90.9% ★
IQ4_XS-imx (stock) 2398 9.5261 +0.3994 0.0598 6.37% 90.5% ★
Nano-27pc ⭐ 2251 9.2769 +0.1501 0.0683 6.85% 89.7% ★
IQ3_M-imx (stock) 2063 9.9031 +0.7763 0.1480 10.62% 84.8% ★
Pico-24pc 1935 10.1509 +1.0242 0.1570 10.92% 84.6% ★
Femto-21pc 1696 10.0878 +0.9610 0.2534 13.98% 79.3% ★

Precision matches Fidelity within the tie zone at 494 MiB less. The fine-tune field flattens sells (×0.3) and leaves most mid-band buys dead (L19), so campaign value concentrates in deep sheds and self re-shuffles.


3. Nanbeige/Nanbeige4.2-3B — looped dense · untied 166k vocab

BF16 7,957 MiB · target: 6+ GB VRAM · PPL reference 29.8955 · readout pair 24.5% of mass · R3 loop-neutral depth prior active · DSpark draft: ×1.87/×2.13 net speedup.

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 4229 31.5657 +1.6702 0.0144 3.01% 95.5% ★
Fidelity-48pc 3821 31.8994 +2.0039 0.0266 3.97% 93.8% ★
Precision-42pc 3266 32.0971 +2.2016 0.0467 4.93% 91.9% ≡ Q6_K-imx
Q6_K-imx (stock) 3266 32.0971 +2.2016 0.0467 4.93% 91.9% ≡ Precision-42pc
Q5_K_M-imx (stock) 2849 32.0553 +2.1598 0.0872 6.58% 88.2% < Quality-36pc
Quality-36pc ⭐ 2846 31.7395 +1.8440 0.0714 5.86% 89.2% ★
Compact-33pc 2629 31.7929 +1.8974 0.1115 7.53% 86.5% ★
Mini-30pc 2388 30.7090 +0.8135 0.1841 9.43% 82.8% ★
IQ4_XS-imx (stock) 2268 32.7700 +2.8745 0.2093 10.24% 81.7% ★
Nano-27pc 2152 31.3016 +1.4061 0.2287 10.51% 80.4% ★
IQ3_M-imx (stock) 1985 33.2114 +3.3158 0.4948 15.12% 71.2% < Pico-24pc
Pico-24pc 1906 30.0424† +0.1469 0.4023 14.04% 73.2% ★
Femto-21pc 1675 36.4632 +6.5677 0.5857 16.35% 68.3% ★

Two strict dominations: Quality (−0.0158 KLD, −3 MiB) and Pico (−0.0925 KLD, −79 MiB). Precision is byte-identical to Q6_K-imx (detected by R19).


4. XHToken/Spark-X2.5-4B — dense SWA hybrid 3:1 · tied 131k vocab

BF16 7,849 MiB · target: 6+ GB VRAM · PPL reference 20.1134 · the verdict binds at native context (L18).

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 4172 19.9740 −0.1394 0.0047 1.60% 97.3% ★
Fidelity-48pc 3772 20.0099 −0.1035 0.0115 2.49% 95.3% ★
Precision-42pc 3359 19.7049 −0.4085 0.0128 2.66% 95.0% ★
Q6_K-imx (stock) 3223 19.6832 −0.4303 0.0173 3.02% 93.8% ★
Q5_K_M-imx (stock) 2840 20.8885 +0.7751 0.0571 5.57% 89.4% < Quality-36pc
Quality-36pc ⭐ 2838 20.0193 −0.0941 0.0437 4.70% 90.4% ★
Compact-33pc 2587 20.6036 +0.4902 0.0976 7.16% 86.1% ★
Mini-30pc 2343 21.3848 +1.2714 0.1203 8.06% 84.6% ★
IQ4_XS-imx (stock) 2266 23.2522 +3.1388 0.1678 9.45% 81.8% ★
Nano-27pc 2123 22.3853 +2.2719 0.2498 11.35% 78.4% ★
Pico-24pc 1982 19.4291† −0.6843 0.2812 12.30% 76.9% ★
IQ3_M-imx (stock) 1949 21.6523 +1.5389 0.3635 14.23% 73.7% ★
Femto-21pc 1664 26.5109 +6.3975 0.5799 17.22% 67.7% ★

Quality dominates Q5_K_M-imx (−0.0134 KLD, −2 MiB). The spark2_5 corrections (readout lever L12, F32 gate pin R9, ffn_down Q4_K lever) apply to every build; the scarce 9-layer full-attention band is the family's richest buy.


5. webAI-Official/TwIL-LM3-Pro — granite dense · untied 128k vocab

BF16 6,984 MiB · target: 6+ GB VRAM · PPL reference 11.4980 · R19 reference case.

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 3712 11.5097 +0.0117 0.0020 1.33% 97.9% ≈ Fidelity-48pc (+0.0007 · −357 MiB)
Fidelity-48pc 3355 11.5074 +0.0094 0.0027 1.38% 97.4% ★
Precision-42pc 2867 11.5042 +0.0062 0.0049 1.99% 96.6% ≡ Q6_K-imx
Q6_K-imx (stock) 2867 11.5042 +0.0062 0.0049 1.99% 96.6% ≡ Precision-42pc
Quality-36pc 2514 11.6187 +0.1207 0.0104 2.74% 95.1% ★
Q5_K_M-imx (stock) 2493 11.6684 +0.1704 0.0131 3.33% 94.5% ★
Compact-33pc 2306 11.6594 +0.1614 0.0218 3.90% 93.2% ★
Mini-30pc 2096 11.6689 +0.1709 0.0300 4.59% 91.6% ★
IQ4_XS-imx (stock) 1937 11.8072 +0.3092 0.0350 4.92% 91.1% ★
Nano-27pc ⭐ 1885 11.9953 +0.4973 0.0479 5.64% 89.4% ★
IQ3_M-imx (stock) 1653 12.4715 +0.9735 0.1052 8.63% 84.7% ★

R19 case: Quality-36pc and Precision-42pc started byte-identical to their stock twins. The self re-shuffle on Quality won −20.7% (0.0131 → 0.0104, 2.7× the duel gate). On Precision it tied (+0.0007, inside the gate, so the incumbent stands). The re-shuffle wins or ties by construction.


6. LiquidAI/LFM2.5-2.6B — shortconv mixer · tied readout

BF16 5,153 MiB · target: 4+ GB VRAM · PPL reference 51.4190 · DSpark draft: ×1.10 net speedup.

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 2742 51.8867 +0.4677 0.0032 1.34% 97.4% ★
Fidelity-48pc 2481 52.2324 +0.8134 0.0076 2.09% 96.2% ★
Precision-42pc 2138 52.8312 +1.4122 0.0105 2.44% 95.4% ★
Q6_K-imx (stock) 2119 52.7076 +1.2886 0.0131 2.66% 94.7% ★
Q5_K_M-imx (stock) 1850 51.5220 +0.1030 0.0380 4.39% 91.0% < Quality-36pc
Quality-36pc ⭐ 1848 52.6558 +1.2368 0.0309 4.13% 91.6% ★
Compact-33pc 1712 49.4337 −1.9853 0.0696 6.13% 88.0% ★
Mini-30pc 1560 46.9409† −4.4780 0.0995 7.45% 85.9% ★
IQ4_XS-imx (stock) 1447 52.6602 +1.2412 0.1652 9.56% 81.6% < Nano-27pc
Nano-27pc 1405 52.7926 +1.3736 0.1536 8.91% 81.8% ★
Pico-24pc 1273 55.9255 +4.5065 0.3597 13.43% 74.6% ★
IQ3_M-imx (stock) 1225 60.9679 +9.5490 0.4254 14.31% 72.6% ★
Femto-21pc 1096 71.9524 +20.5334 0.5535 16.29% 69.0% ★

Quality dominates Q5_K_M-imx (−0.0071 KLD, −2 MiB), and Nano dominates IQ4_XS-imx (−0.0116 KLD, −42 MiB). † PPL below the BF16 base indicates entropy collapse, not a quality gain.


7. openbmb/MiniCPM5-2B — dense llama-arch · untied 130k readout

BF16 4,806 MiB · target: 4+ GB VRAM · PPL reference 11.9212 · DSpark draft: acceptance 0.4419, ×1.89 net speedup.

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 2556 11.9393 +0.0181 0.0015 0.99% 97.8% ★
Fidelity-48pc 2309 11.9598 +0.0386 0.0033 1.47% 96.8% ★
Q6_K-imx (stock) 1974 11.9724 +0.0512 0.0059 1.91% 95.9% ≈ Precision-42pc (−0.0002 · −1 MiB)
Precision-42pc 1973 11.9678 +0.0466 0.0057 1.89% 96.0% ★
Q5_K_M-imx (stock) 1724 12.0905 +0.1693 0.0195 3.58% 92.8% < Quality-36pc
Quality-36pc 1723 12.1180 +0.1968 0.0175 3.43% 93.3% ★
Compact-33pc ⭐ 1593 12.2219 +0.3007 0.0427 5.12% 89.2% ★
Mini-30pc 1446 12.3731 +0.4519 0.0672 6.53% 86.7% ★
IQ4_XS-imx (stock) 1358 12.3977 +0.4765 0.0722 6.69% 86.9% ★
Nano-27pc 1306 12.6530 +0.7318 0.0848 7.26% 85.7% ★
IQ3_M-imx (stock) 1170 13.9101 +1.9889 0.2027 11.74% 78.6% < Pico-24pc
Pico-24pc 1154 13.9155 +1.9943 0.1719 10.52% 79.3% ★
Femto-21pc 1016 15.2794 +3.3582 0.2918 14.25% 73.9% ★

Two strict dominations: Quality (−0.0020 KLD, −1 MiB) and Pico (−0.0308 KLD, −15 MiB). All eight rungs held through the L17 self re-shuffle wave.


8. jinaai/ReaderLM-v2 — dense qwen2 · tied 152k vocab · HTML/Markdown extractor

BF16 2,950 MiB · target: 2+ GB VRAM · PPL reference 12.3852 · FFN holds 75% of the mass.

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 1570 12.4136 +0.0284 0.0016 0.99% 97.9% ★
Fidelity-48pc 1422 12.4161 +0.0309 0.0031 1.39% 96.9% ★
Precision-42pc 1214 12.3712† −0.0140 0.0055 1.84% 95.9% ≡ Q6_K-imx
Q6_K-imx (stock) 1214 12.3712† −0.0140 0.0055 1.84% 95.9% ≡ Precision-42pc
Quality-36pc 1073 12.4244 +0.0392 0.0165 3.25% 93.1% ≡ Q5_K_M-imx
Q5_K_M-imx (stock) 1073 12.4244 +0.0392 0.0165 3.25% 93.1% ≡ Quality-36pc
Compact-33pc ⭐ 979 12.4617 +0.0765 0.0309 4.35% 90.9% ★
Mini-30pc 890 12.6359 +0.2507 0.0553 5.95% 87.7% ≈ IQ4_XS-imx (+0.0027 · −36 MiB)
IQ4_XS-imx (stock) 854 12.5393 +0.1541 0.0580 5.96% 87.6% ★
Nano-27pc 802 12.6418 +0.2566 0.0887 7.39% 84.9% ★
IQ3_M-imx (stock) 741 13.3774 +0.9922 0.1706 10.37% 79.6% ★

Pico and Femto are omitted: the constructibility gate applies (tied vocab sits at 15% of the mass, the R6-ter boundary). Compact, Mini and Nano are campaign products. Quality and Precision are byte-identical to their stock twins (detected by R19). L2 validated: the Q4_K ffn_down floor restored monotonicity across the 75% FFN mass.


9. SeerRay-Lab/Xiaomi-OCR-0 — qwen35 GDN hybrid · tied 248k vocab

BF16 1,446 MiB · target: 2+ GB VRAM · PPL reference 28.5189 · vision via the mmproj module.

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 774 28.5404 +0.0215 0.0010 0.77% 98.1% ★
Fidelity-48pc 704 28.5178 −0.0011 0.0022 1.12% 97.0% ★
Precision-42pc 611 28.5667 +0.0477 0.0050 1.68% 95.4% ≈ Q6_K-imx (+0.0003 · −10 MiB)
Q6_K-imx (stock) 601 28.6110 +0.0920 0.0052 1.74% 95.4% ★
Q5_K_M-imx (stock) 551 28.9326 +0.4136 0.0111 2.51% 93.7% ★
Quality-36pc 538 28.9461 +0.4271 0.0131 2.72% 92.8% ★
Compact-33pc ⭐ 487 29.3475 +0.8285 0.0289 3.70% 90.1% ★
IQ4_XS-imx (stock) 481 29.6333 +1.1144 0.0425 4.81% 87.7% ★
Mini-30pc 444 29.8250 +1.3060 0.0521 5.18% 86.7% ★
IQ3_M-imx (stock) 433 32.1805 +3.6615 0.1050 8.26% 81.0% ★

Pico and Femto are omitted: the constructibility gate applies (tied vocab is 34% of the mass, and the Q6_K floor alone reaches 51% of the Nano target). Mini-30 is the family floor. Precision ≈ Q6_K-imx within the tie zone (C2). The pipeline skips this rung on small qwen35 by default; set PARETRIX_PRECISION_TWIN=0 to force it.


📉 Sub-Nano compendium — the Nano → Pico step

Families run from largest to smallest.

Family Architecture Nano KLD Pico KLD Step Pico verdict
Ornith-1.5-9B GDN hybrid, untied 0.1223 0.1540 +26% ✅ Fully usable (beats IQ3_M 0.1807)
NeoHorse-1-4B GDN fine-tune, tied 0.0683 0.1570 +130% ◐ Context variant (IQ3_M reaches 0.1480 at +128 MiB)
Nanbeige4.2-3B Looped dense, untied 0.2287 0.4023 +76% ✗ Degraded (Nano already past the 0.20 edge)
Spark-X2.5-4B SWA hybrid, tied 0.2498 0.2812 +13% ✗ Degraded (floor = Mini)
TwIL-LM3-Pro Dense, untied 0.0479 — — ⊘ Ladder stops at Nano
LFM2.5-2.6B Shortconv mixer, tied 0.1536 0.3597 +134% ✗ Pico exceeds 0.20 (floor = Mini)
MiniCPM5-2B Dense, untied 0.0848 0.1719 +103% ✅ Usable (beats IQ3_M 0.2027)
ReaderLM-v2 Dense, tied 0.0887 — — ⊘ Constructibility gate (floor = Nano)
Xiaomi-OCR-0 GDN hybrid, tied — — — ⊘ Constructibility gate (floor = Mini)

Summary — 9 strict Pareto dominations

Products that are both lighter and better than their stock twin at the verdict regime. Families run from largest to smallest; tiers run from largest to smallest within each family.

Family Product KLD Stock twin Stock KLD ΔKLD ΔMiB
Ornith-1.5-9B Compact-33pc 0.0610 Q5_K_M-imx 0.1013 −0.0403 −381
Ornith-1.5-9B Pico-24pc 0.1540 IQ3_M-imx 0.1807 −0.0268 −141
Nanbeige4.2-3B Quality-36pc 0.0714 Q5_K_M-imx 0.0872 −0.0158 −3
Nanbeige4.2-3B Pico-24pc 0.4023 IQ3_M-imx 0.4948 −0.0925 −79
Spark-X2.5-4B Quality-36pc 0.0437 Q5_K_M-imx 0.0571 −0.0134 −2
LFM2.5-2.6B Quality-36pc 0.0309 Q5_K_M-imx 0.0380 −0.0071 −2
LFM2.5-2.6B Nano-27pc 0.1536 IQ4_XS-imx 0.1652 −0.0116 −42
MiniCPM5-2B Quality-36pc 0.0175 Q5_K_M-imx 0.0195 −0.0020 −1
MiniCPM5-2B Pico-24pc 0.1719 IQ3_M-imx 0.2027 −0.0308 −15

Plus four byte-identical stock twins detected (R19): Nanbeige Precision ≡ Q6_K-imx · TwIL Precision ≡ Q6_K-imx · ReaderLM Precision ≡ Q6_K-imx · ReaderLM Quality ≡ Q5_K_M-imx.


⚡ Speculative decoding: MTP · DSpark · DFlash

Paretrix quantizes speculative modules imatrix-free (Paretrix-modules.py). A draft's input is the parent's hidden state, the MTP head rides the trunk's activations, and the vision tower consumes raw pixels, so none of them exposes a corpus entry point. Speculative verification bounds draft damage to the acceptance rate and leaves output fidelity intact (L11).

Module Role Palette Naming
MTP Multi-token prediction fusion head (extracted by 01) Role floors + K-quants mtp-Paretrix-<Profile>.gguf
DSpark Block-diffusion speculative draft (Markov + confidence heads) Role floors + K-quants dspark-Paretrix-<Profile>.gguf
DFlash Third-party block-diffusion draft Role floors + K-quants dflash-Paretrix-<Profile>.gguf
MMProj CLIP vision projector Profile grid, critical tensors pinned F32 mmproj-Paretrix-<Profile>.gguf

Measured acceptance invariance (L11). Draft quantization moves acceptance by no more than ±0.03: nanbeige 0.3988 → 0.4092 (Δ +0.0104) · MiniCPM5-2B 0.4419 (ledger 0.4464) · lfm2 +0.0055 · Ornith DFlash 0.2631 → 0.2684 at matched ngl. The quantized draft also runs faster: net speedups of ×1.87/×2.13 (nanbeige), ×1.89 (MiniCPM5-2B), and ×1.34–×2.14 (Ornith DFlash, 106.9 t/s on a 9B model at 8 GB VRAM). Acceptance is a property of the (draft, target) pair and the offload regime (L20), so compare acceptance only at matched --ngl.

# Acceptance battery (llama-server A/B, temperature 0, trained block auto-read)
python 13_draft-acceptance-sweep.py <working-folder> --chat --tokenizer-dir <hf-folder>

# Fuse model + modules into one deployment GGUF
python 14_gguf-module-fusion.py fused.gguf model-BF16.gguf mmproj-BF16.gguf

⚡ Quick start

# One command: convert → imatrix → full Paretrix ladder [+ classic line]
python 00_Paretrix-pipeline.py ./MyModel [--all]

# Or step by step
python 01_SAFETENSORS-to-BF16-GGUF.py ./MyModel  # → Paretrix/model-BF16.gguf
python 02_BF16-GGUF-to-Q8-imatrix.py             # → model-Q8_0.gguf + imatrix.gguf
python 03_BF16-GGUF-to-Paretrix.py [--all]       # → full ladder + duels

# Campaign on a specific rung (anchor → probes → DP → build → duel → adopt)
python Paretrix.py campaign --model model-BF16.gguf --imatrix imatrix.gguf \
    --anchor-from-artifact model-Q5_K_M-imx.gguf --target-mib 1062 --run

# Sweep every product at the verdict regime
python 11_perplexity-test.py

# Fold the model caches into the shared arch lines
python Paretrix.py matrix

Working-root convention. A model folder holds its HF repo (safetensors + sidecars) plus a Paretrix/ subfolder carrying every product: BF16/Q8/imatrix GGUFs, champions, modules, and model-paretrix.json.

<model folder>/
├─ model-00001-of-00002.safetensors · config.json · tokenizer.json …
└─ Paretrix/
   ├─ model-BF16.gguf · model-Q8_0.gguf · imatrix.gguf
   ├─ model-Paretrix-<Tier>-<XX>pc.gguf · mtp/mmproj/dspark products
   └─ model-paretrix.json

Full operational detail: DOCUMENTATION.md.


📜 Citation & Credits

  • Paretrix Quantization Suite: the Pareto-Matrix rate-exchange framework, with measured notch rates, exact-budget DP allocation, and a cross-architecture knowledge base (THE MATRIX).
  • llama.cpp by Georgi Gerganov & ggml contributors: GGUF/GGML runtime, llama-imatrix, llama-quantize, llama-perplexity, and llama-server.
  • GSQ / RCO (IST-DASLab): GSQ: Gumbel-STE quantization (arXiv:2604.18556) · RCO: relaxed knapsack allocation (arXiv:2605.00649). The exact-budget DP and Gumbel-STE manifold search build on these lines, adapted to measured rates.
  • wepiqx ASHQ1: priority-queue knapsack formulation, tied-group activation hashing, and MSE scheduling.
  • Empero AI: GDN state preservation (ssm_alpha/ssm_beta at Q8_0).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for Soulfate24/Paretrix_Quantization_Suite