Boris-75M-Instruct-GGUF

GGUF quantizations of KSP-NMAI/Boris-75M-Instruct for llama.cpp and compatible runtimes (llama-server, LM Studio, koboldcpp, Jan).

The original safetensors weights live in the base repo — use those for finetuning or for any PyTorch-based runtime. GGUF is inference-only.

Which file should I pick?

Use Q8_0, or F16 if you want the exact reference weights.

Boris-75M is a small model, and quantization behaves differently at this scale than it does for 7B+ models. The token embedding table is a large fraction of the parameters and is kept at high precision by llama.cpp, which sets a hard floor on file size. The practical result: every file here is between 53 MB and 150 MB. Dropping from Q8_0 to IQ1_S saves you a few tens of megabytes while degrading output substantially. The aggressive quants are provided for completeness, not because they are a good trade.

Files

File Quant Size Notes
Boris-75M-Instruct-F16.gguf F16 150M Reference. Unquantized conversion of the safetensors weights.
Boris-75M-Instruct-BF16.gguf BF16 150M Reference, bfloat16.
Boris-75M-Instruct-Q8_0.gguf Q8_0 82M Effectively lossless. Recommended.
Boris-75M-Instruct-Q6_K.gguf Q6_K 78M Near-lossless.
Boris-75M-Instruct-Q5_K_M.gguf Q5_K_M 69M Very good quality.
Boris-75M-Instruct-Q5_K_S.gguf Q5_K_S 66M
Boris-75M-Instruct-Q5_1.gguf Q5_1 67M
Boris-75M-Instruct-Q5_0.gguf Q5_0 65M
Boris-75M-Instruct-Q4_K_M.gguf Q4_K_M 67M Standard 4-bit default for larger models.
Boris-75M-Instruct-Q4_K_S.gguf Q4_K_S 63M
Boris-75M-Instruct-Q4_1.gguf Q4_1 62M
Boris-75M-Instruct-Q4_0.gguf Q4_0 59M
Boris-75M-Instruct-IQ4_NL.gguf IQ4_NL 59M
Boris-75M-Instruct-IQ4_XS.gguf IQ4_XS 58M
Boris-75M-Instruct-Q3_K_L.gguf Q3_K_L 64M
Boris-75M-Instruct-Q3_K_M.gguf Q3_K_M 61M
Boris-75M-Instruct-Q3_K_S.gguf Q3_K_S 57M
Boris-75M-Instruct-IQ3_M.gguf IQ3_M 59M
Boris-75M-Instruct-IQ3_S.gguf IQ3_S 57M
Boris-75M-Instruct-IQ3_XS.gguf IQ3_XS 57M
Boris-75M-Instruct-IQ3_XXS.gguf IQ3_XXS 56M
Boris-75M-Instruct-Q2_K.gguf Q2_K 57M
Boris-75M-Instruct-Q2_K_S.gguf Q2_K_S 56M
Boris-75M-Instruct-IQ2_M.gguf IQ2_M 55M
Boris-75M-Instruct-IQ2_S.gguf IQ2_S 55M
Boris-75M-Instruct-IQ2_XS.gguf IQ2_XS 55M
Boris-75M-Instruct-IQ2_XXS.gguf IQ2_XXS 54M
Boris-75M-Instruct-IQ1_M.gguf IQ1_M 54M
Boris-75M-Instruct-IQ1_S.gguf IQ1_S 53M
Boris-75M-Instruct-TQ2_0.gguf TQ2_0 54M
Boris-75M-Instruct-TQ1_0.gguf TQ1_0 53M

All quantizations below 8-bit were produced with an importance matrix calibrated on 100 chunks of held-out data drawn from the model's own training mixture (60% fineweb-edu / 40% dclm).

Usage

# straight from the Hub
llama-server -hf KSP-NMAI/Boris-75M-Instruct-GGUF:Q8_0 --jinja

# or a local file
llama-server -m Boris-75M-Instruct-Q8_0.gguf --jinja

The Alpaca chat template is embedded in every file, so --jinja applies the correct prompt format automatically.

Prompt format

### Instruction:
{your instruction}

### Response:

Limitations

This is a very small instruction-tuned model. It will produce text that is frequently inaccurate, inconsistent, or offensive, and has received no alignment or safety tuning beyond supervised fine-tuning on Alpaca. Do not rely on it for factual information or deploy it without supervision.

License

Apache 2.0. Copyright 2026 Joseph Jones. See the base repository for the full notice.

Downloads last month
236
GGUF
Model size
77.4M params
Architecture
gpt2
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KSP-NMAI/Boris-75M-Instruct-GGUF

Quantized
(2)
this model

Dataset used to train KSP-NMAI/Boris-75M-Instruct-GGUF

Collection including KSP-NMAI/Boris-75M-Instruct-GGUF