K2-Horizon-0.9B — Pollard

Pollard shrank this model: 2.16 GB (f16) → 0.57 GB74% smaller, 3.8× down.

The smallest rung here; larger, higher-fidelity rungs are listed below.

format this model's size
f16 2.16 GB
Q8_0 ~1.15 GB
Q6_K 0.89 GB
Q5_K_M 0.73 GB
PollardMix (this repo's IQ3_S) 0.57 GB

Pollard builds of IFM/K2-Horizon-0.9B made with Pollard Weights — a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).

Standard GGUF — measured K/i-quant ladder (no trellis flagship yet: the k2-horizon arch isn't in ik_llama.cpp). Runs in a recent llama.cpp / Ollama / LM Studio (see Errata).

Available files (wikitext-2 test, ctx 512)

f16 reference PPL 13.01.

file PPL size Mean KLD notes
K2-Horizon-0.9B-Pollard-IQ3_S.gguf 14.50 0.57 GB smallest
K2-Horizon-0.9B-Pollard-Q5_K_M.gguf 13.27 0.73 GB recommended default
K2-Horizon-0.9B-Pollard-Q6_K.gguf 13.05 0.89 GB near-lossless

Usage

llama-cli -m K2-Horizon-0.9B-Pollard-Q5_K_M.gguf -p "Explain why the sky is blue." --temp 0.7
ollama run hf.co/PollardWeights/K2-Horizon-0.9B-Pollard

Errata

  • Requires a recent llama.cpp with k2-horizon architecture support — older builds report unknown architecture 'k2-horizon'. The arch currently ships in the MBZUAI-IFM/llama.cpp fork (branch model/K2Horizon); it runs in stock llama.cpp / Ollama / LM Studio once merged upstream.
  • Trellis (IQ*_KT) quants need ik_llama.cpp to build/run; K-quants run in any recent llama.cpp.
  • Measured allocation places bits by per-layer sensitivity under a size budget, on a Calib 3.0 imatrix.
  • Single machine; replication invited.

Built with Pollard Weights — frontier models, small hardware, no compromise.

Downloads last month
-
GGUF
Model size
1B params
Architecture
k2-horizon
Hardware compatibility
Log In to add your hardware

3-bit

5-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PollardWeights/K2-Horizon-0.9B-Pollard

Quantized
(5)
this model