LFM2.5-2.6B-GGUF

GGUF quantizations of LiquidAI/LFM2.5-2.6B, a 2.69B-parameter dense hybrid model built for agentic, on-device workloads with a 128K-token context window and native tool calling [web:41][web:39].

Model Details

LFM2.5-2.6B combines 22 double-gated short convolution (LIV) blocks with 8 grouped-query attention (GQA) blocks across 30 total layers, an architecture selected via hardware-in-the-loop search on real edge silicon [web:40][web:44]. It was pre-trained on roughly 34 trillion tokens and post-trained through a four-stage pipeline (SFT, teacher specialization, multi-domain on-policy distillation, and agentic RL) to reliably plan, call tools, and execute multi-step tasks inside agent harnesses [web:38][web:44].

Property Value
Parameters 2.69B (dense)
Layers 30 (22 conv + 8 GQA)
Embedding dimension 2048
Context length 131,072 tokens [web:41]
Vocabulary size 128,000 tokens
Training data ~34 trillion tokens
Languages 16, including English, Arabic, Chinese, French, German, Hindi, Japanese, Korean, Russian, Spanish [web:41]
License LFM Open License v1.0 [web:50]

Files

Quantized with llama.cpp's convert_hf_to_gguf.py and llama-quantize. Tested for compatibility on an RTX 4060 Laptop (8 GB VRAM).

Quantization Size Notes
Q2_K 1.09 GB Smallest, largest quality loss
Q3_K_S 1.27 GB
Q3_K_M 1.37 GB Balanced 3-bit
Q3_K_L 1.45 GB
Q4_K_S 1.6 GB
Q4_K_M 1.67 GB Recommended default for most users
Q5_K_S 1.9 GB
Q5_K_M 1.94 GB Near-F16 quality, moderate size
Q6_K 2.22 GB Very close to F16 quality
Q8_0 2.87 GB Minimal quality loss
F16 5.4 GB Full precision, reference file

Usage

Run with llama.cpp, Ollama, LM Studio, or any GGUF-compatible inference engine:

./llama-cli -m LFM2.5-2.6B-Q4_K_M.gguf -p "Your prompt here" -n 256

The model uses a ChatML-like chat template with native tool-call tokens (<|tool_call_start|>, <|tool_call_end|>) and a Pythonic tool-call format (function_name(arg="value")) [web:40].

Recommended Quantization

For 8 GB VRAM laptops (e.g. RTX 4060 Laptop), Q4_K_M offers the best balance of quality and footprint (~1.67 GB), leaving headroom for KV cache at long context lengths. For maximum fidelity on the same hardware, Q6_K or Q8_0 still fit comfortably given the model's small base size [web:51].

License

This model inherits the LFM Open License v1.0 from the original LiquidAI/LFM2.5-2.6B release, not a permissive license like Apache 2.0 or MIT — review the terms before commercial deployment [web:50][web:51].

Credits

Original model and architecture by Liquid AI [web:38]. GGUF conversion by NANI-Nithin.

Downloads last month
301
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NANI-Nithin/LFM2.5-2.6B-GGUF

Quantized
(54)
this model