K2-Horizon-3.7B — Pollard

Pollard shrank this model: 10.1 GB (f16) → 2.75 GB73% smaller, 3.7× down.

The smallest rung here; larger, higher-fidelity rungs are listed below.

format this model's size
f16 10.1 GB
Q8_0 ~5.4 GB
Q6_K ~4.1 GB
Q4_K_M ~2.9 GB
PollardMix (this repo's IQ3_S) 2.75 GB

Pollard builds of IFM/K2-Horizon-3.7B made with Pollard Weights — a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).

Standard GGUF. The k2-horizon architecture is not in stock llama.cpp yet — these files need a build of the MBZUAI-IFM/llama.cpp fork. Older builds report unknown architecture 'k2-horizon'; they will run in stock llama.cpp / Ollama / LM Studio once the arch is merged upstream.

Available files (wikitext-2 test, ctx 512)

f16 reference PPL 11.2798.

file PPL size Mean KLD notes
K2-Horizon-3.7B-Pollard-IQ3_S.gguf 12.5164 2.75 GB smallest (+1.24)
K2-Horizon-3.7B-Pollard-IQ4_XS.gguf 11.4211 3.33 GB recommended default (+0.14)
K2-Horizon-3.7B-Pollard-Q6_K.gguf 11.3603 4.16 GB near-lossless (+0.08 over f16)

Usage

llama-cli -m K2-Horizon-3.7B-Pollard-IQ4_XS.gguf -p "Explain why the sky is blue." --temp 0.7
ollama run hf.co/PollardWeights/K2-Horizon-3.7B-Pollard

Errata

  • Measured allocation places bits by per-layer sensitivity under a size budget, on a Calib 3.0 imatrix — not a uniform crush.
  • PPL was measured on these exact uploaded files, re-downloaded from this repo, on Apple Metal.
  • Building on Windows needs a one-line fix. MSVC's std::regex rejects the \p{L} escapes in K2's pre-tokenizer, so an unpatched Windows build of the fork fails to load any K2 GGUF with regex_error(error_escape); macOS is unaffected. Patch: notes/k2-horizon-msvc-regex.patch.
  • The base GGUF trips a special_eos_id is not in special_eog_ids warning from the fork. That is upstream tokenizer metadata, harmless for perplexity and generation.
  • Every rung was rebuilt from scratch on a second machine (macOS/Metal + CPU imatrix vs Windows/CUDA + GPU imatrix) and came out byte-for-byte the same size; the two machines' PPLs agree to within their error bars (Q6_K 11.3603 vs 11.3780). The allocation is reproducible, not machine-specific.

Built with Pollard Weights — frontier models, small hardware, no compromise.

Downloads last month
-
GGUF
Model size
5B params
Architecture
k2-horizon
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PollardWeights/K2-Horizon-3.7B-Pollard

Quantized
(10)
this model