LFM2.5‑2.6B GGUF Quantizations

Overview

This repository contains the GGUF‑formatted quantized versions of the LiquidAI LFM2.5‑2.6B model. Each variant is named according to the LFM2.5-2.6B-<quant>.gguf scheme and is optimized for different trade‑offs between size and inference speed.

IMPORTANT FIX: It was not possible to disable thinking in LiquidAI/LFM2.5-2.6B-GGUF, these quants don't have that issue.

All quantizations were generated using the imatrix quantization pipeline. Note: the token_embd.weight layer was not quantized in any of these files.

imatrix calibration file is from bartowski

Written by LFM2.5-2.6B-Q8_0: This model card is written by the Q8_0 quant.

Quantization Table

Quant File Name Size (GB)
F16 LFM2.5-2.6B-F16.gguf 5.27
Q8_0 LFM2.5-2.6B-Q8_0.gguf 2.80
Q6_K LFM2.5-2.6B-Q6_K.gguf 2.16
Q5_K_M LFM2.5-2.6B-Q5_K_M.gguf 1.89
Q4_K_M LFM2.5-2.6B-Q4_K_M.gguf 1.63
Q4_K_S LFM2.5-2.6B-Q4_K_S.gguf 1.56
IQ4_NL LFM2.5-2.6B-IQ4_NL.gguf 1.55
IQ4_XS LFM2.5-2.6B-IQ4_XS.gguf 1.48
Q3_K_S LFM2.5-2.6B-Q3_K_S.gguf 1.24
IQ3_XS LFM2.5-2.6B-IQ3_XS.gguf 1.19
Q2_K LFM2.5-2.6B-Q2_K.gguf 1.06

Notes

  • Imatrix Pipeline: All quantizations were created with the imatrix quantization tool. This method provides high‑quality compression while preserving model accuracy.
  • Token Embeddings: The token_embd.weight layer remains in full FP16 (or the original precision) and was not quantized. This ensures that embedding look‑ups remain accurate during inference.
  • KL Divergence Test: A KL‑divergence evaluation will be added to this repository shortly. Stay tuned for the updated benchmark results.

Usage

To load any of the quantized models, simply use the standard GGUF loader (e.g., llama.cpp or auto-gptq):

# Example with llama.cpp
./llama-server -m "./LFM2.5-2.6B-Q4_K_M.gguf" -c 128000

Replace the model file with any of the entries from the table above.

License

All models and quantizations are released under the LiquidAI LFM2.5‑2.6B license (see the model card for details).


For further questions or feedback, feel free to open an issue.

Downloads last month
785
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for miifanboy/LFM2.5-2.6B-GGUF

Quantized
(56)
this model