ICE: Quantization by Error-Propagation Class in Sparse MoE Models
Technical report and full evidence for ICE (Isolation of Compounding Error), a quantization bit-allocation method for sparse Mixture-of-Experts checkpoints.
Author: Gökhan Buz (gbuzhf)
Read the paper
What ICE is
Every quantizer in the GGUF ecosystem minimizes the same objective for every tensor: the importance-matrix-weighted error of that tensor's output, for the current token. That is the right objective for a tensor whose error dies with the token. It is the wrong objective for three other kinds, and no published method separates them.
ICE classifies every tensor by how far its error travels:
| class | mechanism | treatment |
|---|---|---|
| DISCRETE | error flips an argmax, so a different computation runs | exact (F32) |
| RECURRENT | error enters a state decay and compounds along the sequence | exact (F32) |
| CACHED | error is written to the KV cache once and re-read by every later token | near-exact (F16) |
| INSTANT | error affects this token only | this is where the budget lives |
The first three are 0.14% of the model, so freezing them is a line item rather than a trade-off. In one line: freeze what propagates, spend everything else on the library.
Headline results
Mean KL divergence against the bf16 checkpoint, WikiText-2, one harness for all files.
| tier | size | mean KLD | nearest published tier | outcome |
|---|---|---|---|---|
23G-ICE |
22.83 GB | 0.0361 | UD-Q4_K_XL 23.21 GB / 0.0380 |
0.38 GB smaller, 5.0% better |
23G-ICE |
22.83 GB | 0.0361 | APEX-I-Quality 23.84 GB / 0.0415 |
1.01 GB smaller, 13.0% better |
25G-ICE |
24.84 GB | 0.0303 | APEX-I-Balanced 26.28 GB / 0.0345 |
1.44 GB smaller, 12.2% better |
19G-ICE |
18.82 GB | 0.0608 | UD-IQ4_XS 18.68 GB / 0.0723 |
15.8% better at +0.14 GB |
Across the twelve-tier comparison, nine tiers are Pareto-optimal and three are
strictly dominated. ICE does not win at the top of the ladder: UD-Q5_K_S
and UD-Q6_K are the two best files measured, which the paper explains rather
than omits (Law 4).
The four laws
- Sparsity. A bit on the always-on core is worth
E/kbits on the expert bank. Measured 31.8 against a predicted 32. - Convexity. Error falls as
4^-b, so allocation cleverness is capped at +0.139 bpw. Two consequences: do no depth grading, and at a fixed average always pick the narrower type bracket (measured +16.6% and +7.0% penalty for widening). - Placement. Inside a fixed bracket, shallow-first is worth about −8.5%
KLD per bpw of gap and reverses below 0.47 bpw. Depth gain measured
directly as
g(t) ≈ exp(−t/9.95). - The floor is epistemic. Fitted on two independent harnesses at
k₀= 0.0205 and 0.0219, and confirmed by a direct probe at 0.018297. The best file measured is 8% above it. You cannot out-bit a wrong prior.
What is in this repository
paper/ICE_Technical_Report.md the report
appendix/original-ICE/ PRINCIPLE.md and RECIPE.md, the method as first written
appendix/recipes/ every tensor-type-file cited in the paper
appendix/measurements/ raw llama-perplexity output for every KLD quoted
appendix/SHA256SUMS.txt checksum for every file above
appendix/recipes/
| suffix | meaning |
|---|---|
_ICEbase |
the current recipe, after the Law 3 revision |
_previous |
the recipe it replaced |
_comparison |
a published UD or APEX tier, as measured |
cfg_21G-ICE_ICEbase.txt and cfg_21G-ICE_previous.txt have the same
SHA-256 (93456d45b1fc0dab...). That is not an oversight. 21G is the tier where
Law 3's condition does not hold, so the method's output is to change nothing, and
the identical checksum is the proof that nothing was changed.
Recipes for the three dominated tiers are not included: they were dropped after the Pareto analysis and none was retained. Their measured numbers are in the paper.
Reproducing a number
llama-perplexity -m <file>.gguf -f wiki.test.raw \
-c 2048 --chunks 64 --kl-divergence-base base.kld --kl-divergence
against a base.kld produced once from the bf16 checkpoint with the same corpus
and chunk count. Protocol details, including how the byte-identical control
variants are constructed, are in Appendix B.
The same recipe measured through three different paths gave 0.0345, 0.034162 and 0.034290, a spread under 1%.
Negative results
Seven are documented at the same weight as the positive ones, including the
retraction of a rule this work itself derived and shipped: bumping ffn_down
above its sibling projections, which is standard practice, measured 11.3% worse
than uniform at identical size. One of three registered predictions also failed,
and it is scored as such.
On the comparison ladders
Unsloth Dynamic 2.0 and LocalAI APEX are the work of their respective authors. They are measured here, not reproduced or modified. The comparison exists because no individual publication can provide it: each ladder is published with its own harness and its own reference, so the tiers are not comparable until someone puts them on one. The same analysis that finds ICE tiers dominating three others also finds two UD tiers to be the best files on the board and two APEX tiers to be the only options below 18.5 GB.
Citing
Cite this repository. Model cards for ICE-quantized GGUFs link here for the method description.