Qwen2.5-7B-Instruct โ€” Pollard

Pollard shrank this model: 14.2 GB (f16) โ†’ 1.9 GB โ€” 87% smaller, 7.5ร— down, and under half the size of NVFP4 (~4.0 GB).

The 1-bit-class flagship (IQ1_KT), still beating uniform 1-bit on every metric. Want more quality? The IQ3_S / IQ4_XS / Q6_K rungs below are larger, higher-fidelity options.

format this model's size
f16 14.2 GB
Q8_0 ~8.1 GB
Q6_K ~6.3 GB
Q4_K_M / NVFP4 ~4.6 / ~4.0 GB
PollardMix (this repo's IQ1_KT) 1.9 GB

Pollard builds of Qwen2.5-7B-Instruct made with Pollard Weights. A full ladder of imatrix-guided K-quants for quality-per-byte, plus the flagship 1-bit-class mixed-precision trellis build (IQ1_KT) โ€” expert/FFN body crushed to 1-bit, attention + residual writers protected.

Standard GGUF โ€” runs in stock llama.cpp / ik_llama.cpp, Ollama, LM Studio. The IQ1_KT trellis file needs ik_llama.cpp; the K-quants run anywhere.

Available files (WikiText-2 raw, ctx 2048, 145 chunks; f16 ref PPL 6.52)

file PPL size Mean KLD notes
โ€ฆ-Q6_K.gguf 6.55 5.82 GB 0.0035 near-lossless
โ€ฆ-IQ4_XS.gguf 6.66 3.93 GB 0.024 recommended default
โ€ฆ-IQ3_S.gguf 6.96 3.26 GB 0.075 smaller
โ€ฆ-IQ1_KT.gguf 10.23 1.90 GB 0.537 flagship โ€” 1-bit mixed trellis

The IQ1_KT flagship beats a uniform 1-bit IQ1_KT baseline (PPL 11.86, Mean KLD 0.689, top-1 65.1%) on every metric at the same size class โ€” PPL โˆ’14%, Mean KLD โˆ’22%, top-1 +4.2 pts โ€” the mixed-precision "punches above its weight" build. The K-quant ladder is imatrix-guided; on a dense model that's where the bits-per-byte win lives (the measured-KL knapsack is reserved for MoE โ€” we don't claim it here).

Usage

llama-cli -m Qwen2.5-7B-Instruct-Pollard-IQ4_XS.gguf -p "Explain why the sky is blue." --temp 0.7
ollama run hf.co/PollardWeights/Qwen2.5-7B-Instruct-Pollard

Prompt format (ChatML):

<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant

Errata

  • IQ1_KT is an ik_llama.cpp trellis quant; build ik_llama.cpp for it (loads in stock llama.cpp too). K-quants run in any recent llama.cpp.
  • Chat at the 1-bit tier benefits from --repeat-penalty 1.15.
  • Single machine; replication invited.

Built with Pollard Weights โ€” frontier models, small hardware, no compromise.

Downloads last month
1,041
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for PollardWeights/Qwen2.5-7B-Instruct-Pollard

Base model

Qwen/Qwen2.5-7B
Quantized
(399)
this model