spoomplesmaxx-mockingbird-36B β€” i1-GGUF (weighted/imatrix)

Weighted/imatrix GGUF quants of spoomplesmaxx-mockingbird-36B. The importance matrix was computed on a stratified sample of the model's own training corpus β€” all ten lanes, rendered in the exact chat template the model serves with β€” not a generic calibration set. At 3–4 bit these should beat the static quants noticeably; at Q5 the difference fades.

Quant Size Notes
i1-IQ3_XXS ~14 GB smallest usable; VRAM-desperate only
i1-Q3_K_M ~18 GB the 18GB target, imatrix-weighted
i1-IQ4_XS ~19 GB best size/quality trade below Q4_K_M
i1-Q4_K_M ~22 GB recommended
i1-Q5_K_M ~26 GB closest to bf16 behavior

The seed-native chat template is embedded in the GGUF metadata.

Sampling β€” read this part

temperature 1.0 Β· top_p 0.9 Β· repeat_penalty 1.0 (OFF)

⚠ Never use repetition, presence, or frequency penalties. The template ends every message with <seed:eos>; context-wide penalties suppress that token, the model stops ending its turns, and generation degenerates into the base model's untrained Chinese vocabulary. Many frontend presets default repeat_penalty to 1.05–1.1 β€” set it back to 1.0. Use DRY or XTC if you want extra anti-repetition; both leave special tokens alone.

Usable temperature window is ~0.95–1.05: lower loops verbatim, higher frays. Full details on the main model card.

mimids 01 Β· Apache 2.0

Downloads last month
296
GGUF
Model size
36B params
Architecture
seed_oss
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for aimeri/spoomplesmaxx-mockingbird-36B-i1-GGUF

Collection including aimeri/spoomplesmaxx-mockingbird-36B-i1-GGUF