ICE: Quantization by Error-Propagation Class in Sparse MoE Models

Technical report and full evidence for ICE (Isolation of Compounding Error), a quantization bit-allocation method for sparse Mixture-of-Experts checkpoints.

Author: Gökhan Buz (gbuzhf)

Read the paper

paper/ICE_Technical_Report.md

What ICE is

Every quantizer in the GGUF ecosystem minimizes the same objective for every tensor: the importance-matrix-weighted error of that tensor's output, for the current token. That is the right objective for a tensor whose error dies with the token. It is the wrong objective for three other kinds, and no published method separates them.

ICE classifies every tensor by how far its error travels:

class mechanism treatment
DISCRETE error flips an argmax, so a different computation runs exact (F32)
RECURRENT error enters a state decay and compounds along the sequence exact (F32)
CACHED error is written to the KV cache once and re-read by every later token near-exact (F16)
INSTANT error affects this token only this is where the budget lives

The first three are 0.14% of the model, so freezing them is a line item rather than a trade-off. In one line: freeze what propagates, spend everything else on the library.

Headline results

Mean KL divergence against the bf16 checkpoint, WikiText-2, one harness for all files.

tier size mean KLD nearest published tier outcome
23G-ICE 22.83 GB 0.0361 UD-Q4_K_XL 23.21 GB / 0.0380 0.38 GB smaller, 5.0% better
23G-ICE 22.83 GB 0.0361 APEX-I-Quality 23.84 GB / 0.0415 1.01 GB smaller, 13.0% better
25G-ICE 24.84 GB 0.0303 APEX-I-Balanced 26.28 GB / 0.0345 1.44 GB smaller, 12.2% better
19G-ICE 18.82 GB 0.0608 UD-IQ4_XS 18.68 GB / 0.0723 15.8% better at +0.14 GB

Across the twelve-tier comparison, nine tiers are Pareto-optimal and three are strictly dominated. ICE does not win at the top of the ladder: UD-Q5_K_S and UD-Q6_K are the two best files measured, which the paper explains rather than omits (Law 4).

The four laws

  1. Sparsity. A bit on the always-on core is worth E/k bits on the expert bank. Measured 31.8 against a predicted 32.
  2. Convexity. Error falls as 4^-b, so allocation cleverness is capped at +0.139 bpw. Two consequences: do no depth grading, and at a fixed average always pick the narrower type bracket (measured +16.6% and +7.0% penalty for widening).
  3. Placement. Inside a fixed bracket, shallow-first is worth about −8.5% KLD per bpw of gap and reverses below 0.47 bpw. Depth gain measured directly as g(t) ≈ exp(−t/9.95).
  4. The floor is epistemic. Fitted on two independent harnesses at k₀ = 0.0205 and 0.0219, and confirmed by a direct probe at 0.018297. The best file measured is 8% above it. You cannot out-bit a wrong prior.

What is in this repository

paper/ICE_Technical_Report.md      the report
appendix/original-ICE/             PRINCIPLE.md and RECIPE.md, the method as first written
appendix/recipes/                  every tensor-type-file cited in the paper
appendix/measurements/             raw llama-perplexity output for every KLD quoted
appendix/SHA256SUMS.txt            checksum for every file above

appendix/recipes/

suffix meaning
_ICEbase the current recipe, after the Law 3 revision
_previous the recipe it replaced
_comparison a published UD or APEX tier, as measured

cfg_21G-ICE_ICEbase.txt and cfg_21G-ICE_previous.txt have the same SHA-256 (93456d45b1fc0dab...). That is not an oversight. 21G is the tier where Law 3's condition does not hold, so the method's output is to change nothing, and the identical checksum is the proof that nothing was changed.

Recipes for the three dominated tiers are not included: they were dropped after the Pareto analysis and none was retained. Their measured numbers are in the paper.

Reproducing a number

llama-perplexity -m <file>.gguf -f wiki.test.raw \
    -c 2048 --chunks 64 --kl-divergence-base base.kld --kl-divergence

against a base.kld produced once from the bf16 checkpoint with the same corpus and chunk count. Protocol details, including how the byte-identical control variants are constructed, are in Appendix B.

The same recipe measured through three different paths gave 0.0345, 0.034162 and 0.034290, a spread under 1%.

Negative results

Seven are documented at the same weight as the positive ones, including the retraction of a rule this work itself derived and shipped: bumping ffn_down above its sibling projections, which is standard practice, measured 11.3% worse than uniform at identical size. One of three registered predictions also failed, and it is scored as such.

On the comparison ladders

Unsloth Dynamic 2.0 and LocalAI APEX are the work of their respective authors. They are measured here, not reproduced or modified. The comparison exists because no individual publication can provide it: each ladder is published with its own harness and its own reference, so the tiers are not comparable until someone puts them on one. The same analysis that finds ICE tiers dominating three others also finds two UD tiers to be the best files on the board and two APEX tiers to be the only options below 18.5 GB.

Citing

Cite this repository. Model cards for ICE-quantized GGUFs link here for the method description.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support