Instructions to use miifanboy/LFM2.5-2.6B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use miifanboy/LFM2.5-2.6B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use miifanboy/LFM2.5-2.6B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "miifanboy/LFM2.5-2.6B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "miifanboy/LFM2.5-2.6B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
- Ollama
How to use miifanboy/LFM2.5-2.6B-GGUF with Ollama:
ollama run hf.co/miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
- Unsloth Studio
How to use miifanboy/LFM2.5-2.6B-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for miifanboy/LFM2.5-2.6B-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for miifanboy/LFM2.5-2.6B-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for miifanboy/LFM2.5-2.6B-GGUF to start chatting
- Pi
How to use miifanboy/LFM2.5-2.6B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use miifanboy/LFM2.5-2.6B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use miifanboy/LFM2.5-2.6B-GGUF with Docker Model Runner:
docker model run hf.co/miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
- Lemonade
How to use miifanboy/LFM2.5-2.6B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.LFM2.5-2.6B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use miifanboy/LFM2.5-2.6B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default miifanboy/LFM2.5-2.6B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
LFM2.5‑2.6B GGUF Quantizations
Overview
This repository contains the GGUF‑formatted quantized versions of the LiquidAI LFM2.5‑2.6B model. Each variant is named according to the LFM2.5-2.6B-<quant>.gguf scheme and is optimized for different trade‑offs between size and inference speed.
IMPORTANT FIX: It was not possible to disable thinking in LiquidAI/LFM2.5-2.6B-GGUF, these quants don't have that issue.
All quantizations were generated using the imatrix quantization pipeline. Note: the token_embd.weight layer was not quantized in any of these files.
imatrix calibration file is from bartowski
Written by LFM2.5-2.6B-Q8_0: This model card is written by the Q8_0 quant.
Quantization Table
| Quant | File Name | Size (GB) |
|---|---|---|
| F16 | LFM2.5-2.6B-F16.gguf |
5.27 |
| Q8_0 | LFM2.5-2.6B-Q8_0.gguf |
2.80 |
| Q6_K | LFM2.5-2.6B-Q6_K.gguf |
2.16 |
| Q5_K_M | LFM2.5-2.6B-Q5_K_M.gguf |
1.89 |
| Q4_K_M | LFM2.5-2.6B-Q4_K_M.gguf |
1.63 |
| Q4_K_S | LFM2.5-2.6B-Q4_K_S.gguf |
1.56 |
| IQ4_NL | LFM2.5-2.6B-IQ4_NL.gguf |
1.55 |
| IQ4_XS | LFM2.5-2.6B-IQ4_XS.gguf |
1.48 |
| Q3_K_S | LFM2.5-2.6B-Q3_K_S.gguf |
1.24 |
| IQ3_XS | LFM2.5-2.6B-IQ3_XS.gguf |
1.19 |
| Q2_K | LFM2.5-2.6B-Q2_K.gguf |
1.06 |
Notes
- Imatrix Pipeline: All quantizations were created with the
imatrixquantization tool. This method provides high‑quality compression while preserving model accuracy. - Token Embeddings: The
token_embd.weightlayer remains in full FP16 (or the original precision) and was not quantized. This ensures that embedding look‑ups remain accurate during inference. - KL Divergence Test: A KL‑divergence evaluation will be added to this repository shortly. Stay tuned for the updated benchmark results.
Usage
To load any of the quantized models, simply use the standard GGUF loader (e.g., llama.cpp or auto-gptq):
# Example with llama.cpp
./llama-server -m "./LFM2.5-2.6B-Q4_K_M.gguf" -c 128000
Replace the model file with any of the entries from the table above.
License
All models and quantizations are released under the LiquidAI LFM2.5‑2.6B license (see the model card for details).
For further questions or feedback, feel free to open an issue.
- Downloads last month
- 785
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
16-bit