Instructions to use LiquidAI/LFM2.5-VL-3B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use LiquidAI/LFM2.5-VL-3B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use LiquidAI/LFM2.5-VL-3B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LiquidAI/LFM2.5-VL-3B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LiquidAI/LFM2.5-VL-3B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
- Ollama
How to use LiquidAI/LFM2.5-VL-3B-GGUF with Ollama:
ollama run hf.co/LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
- Unsloth Studio
How to use LiquidAI/LFM2.5-VL-3B-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for LiquidAI/LFM2.5-VL-3B-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for LiquidAI/LFM2.5-VL-3B-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for LiquidAI/LFM2.5-VL-3B-GGUF to start chatting
- Pi
How to use LiquidAI/LFM2.5-VL-3B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use LiquidAI/LFM2.5-VL-3B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use LiquidAI/LFM2.5-VL-3B-GGUF with Docker Model Runner:
docker model run hf.co/LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
- Lemonade
How to use LiquidAI/LFM2.5-VL-3B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.LFM2.5-VL-3B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use LiquidAI/LFM2.5-VL-3B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
A 3B Vision Model Running at 13 Tokens/Second — on a Phone
AI is moving from the cloud into your pocket.
I just tested Liquid AI's newly released LFM2.5-VL-3B locally on a Samsung Galaxy S26. The Q8_0 model ran through llama.cpp with Vulkan acceleration on the Qualcomm Adreno 840 GPU. I uploaded a real webpage screenshot from another device and asked the model to describe it.
It worked—and generated at 13.11 tokens per second.
That's significant for a 3B-class vision-language model running entirely on a phone. In my testing, LFM2.5-VL-3B is the most capable local vision model I've run on the S26 so far.
Liquid AI is building specifically for this world. Its focus is edge AI: capable models designed to run locally on phones, laptops, vehicles and other resource-constrained hardware instead of depending entirely on cloud inference. Its LFM family targets low memory usage, fast inference and deployment across CPUs, GPUs and NPUs.
LFM2.5-VL-3B brings vision into that strategy. It is an open-weight multimodal model capable of image understanding, OCR, document extraction, visual question answering and spatial reasoning. Liquid provides the weights publicly, including an official GGUF release for llama.cpp.
This is where open-weight AI gets interesting. A phone can now hold the model, process the image and generate useful multimodal responses locally at interactive speed.
No cloud GPU. No remote inference API. The phone is the AI computer.LFM2.5-VL-3B — Hugging Face
The image shows a man standing outdoors with a group of people behind him. The background features colorful, striped mountains with hues of red, green, yellow, and purple. The sky is overcast with clouds, and there appears to be mist or low-hanging clouds around the mountain peaks. The man in the foreground is wearing a dark jacket over a gray T-shirt that has text and graphics on it, which include the words "Western Pennsylvania Mushroom Club." He has short, light-colored hair and is wearing glasses. The overall setting suggests a hiking or outdoor gathering in a scenic, possibly volcanic, landscape.
Based on the image, there is a person standing outdoors in front of a boat. The individual is wearing a brown hoodie with "GANDER MTN." written on it, blue jeans, and brown shoes. They are holding two large fish, one in each hand, by their mouths. The fish appear to be salmon, given their size and distinctive coloration. The background includes a boat on a trailer and some grassy areas.


