Spark-X2.5-4B — Pollard

Pollard shrank this model: 8.0 GB (f16) → 1.93 GB76% smaller, 4.1× down.

The smallest rung here; larger, higher-fidelity rungs are listed below.

format this model's size
f16 8.0 GB
Q8_0 ~4.2 GB
Q6_K ~3.3 GB
Q4_K_M ~2.3 GB
PollardMix (this repo's IQ3_S) 1.93 GB

Pollard builds of XHToken/Spark-X2.5-4B made with Pollard Weights — a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).

Standard GGUF — measured K/i-quant ladder (no trellis flagship yet: the spark2_5 arch isn't in ik_llama.cpp). Runs in a recent llama.cpp / Ollama / LM Studio (see Errata).

Available files (Calib 3.0 corpus, ctx 512, 6 chunks)

f16 reference PPL 6.091.

file PPL size Mean KLD notes
Spark-X2.5-4B-Pollard-IQ3_S.gguf 6.781 1.93 GB smallest
Spark-X2.5-4B-Pollard-IQ4_XS.gguf 6.207 2.42 GB recommended default
Spark-X2.5-4B-Pollard-Q6_K.gguf 6.099 3.38 GB near-lossless

Usage

llama-cli -m Spark-X2.5-4B-Pollard-IQ3_S.gguf -p "Explain why the sky is blue." --temp 0.7
ollama run hf.co/PollardWeights/Spark-X2.5-4B-Pollard

Errata

  • Requires a recent llama.cpp (Sept 2026+, with spark2_5 architecture support) — older builds report unknown architecture 'spark2_5'. Runs in stock llama.cpp / Ollama / LM Studio once updated.
  • Trellis (IQ*_KT) quants need ik_llama.cpp to build/run; K-quants run in any recent llama.cpp.
  • Measured allocation places bits by per-layer sensitivity under a size budget.
  • Single machine; replication invited.

Built with Pollard Weights — frontier models, small hardware, no compromise.

Downloads last month
-
GGUF
Model size
4B params
Architecture
spark2_5
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PollardWeights/Spark-X2.5-4B-Pollard

Quantized
(21)
this model