LFM2.5-VL-3B-GGUF

LFM2.5-VL-3B is Liquid AI's multimodal, on-device-optimized vision-language model, built on the LFM2-VL-3B foundation with further mid- and post-training, combining the LFM2.5-2.6B language backbone with a SigLIP2 NaFlex shape-optimized 400M vision encoder that processes images at native resolution by splitting large images into non-overlapping 512×512 patches plus a resized thumbnail. It supports a 32,768-token context window across 16 languages, delivers improved grounding and object detection with natural-language queries, and introduces full-page OCR with structured layout annotation (bounding boxes, region labels like text/table/equation, LaTeX for formulas, and OTSL for tables). It's designed for single-turn, high-throughput, low-latency tasks — near-realtime object detection, batch OCR of scanned documents, or on-device translation of menus and signs — rather than long-context or reasoning-intensive work like visual web design. Compared to its LFM2-VL-3B predecessor, it shows substantial gains in screen understanding (80.7 on ScreenSpot-v2, versus far behind on tool-use benchmarks previously), grounding (87.9 on RefCOCO), multi-image reasoning (58.3 on MuirBench), and tool use (59.5 on ToolSandBox), competing closely with larger models like Gemma4-E4B while running at 228 tok/s on an Apple M5 Max and 116 tok/s on an AMD Ryzen AI Max+ 395 in under 3.3GB of memory (and even 20 tok/s on a Galaxy S26 Ultra), or reaching ~11K tokens/sec throughput on a single H100 via vLLM. It's available in native, GGUF, ONNX, and MLX formats, supports Pythonic tool calling, and is released under Liquid AI's LFM1.0 license.

Model Files

File Name Quant Type File Size File Link
LFM2.5-VL-3B.BF16.gguf BF16 5.4 GB Download
LFM2.5-VL-3B.F16.gguf F16 5.4 GB Download
LFM2.5-VL-3B.Q3_K_L.gguf Q3_K_L 1.45 GB Download
LFM2.5-VL-3B.Q3_K_M.gguf Q3_K_M 1.37 GB Download
LFM2.5-VL-3B.Q3_K_S.gguf Q3_K_S 1.27 GB Download
LFM2.5-VL-3B.Q4_K_M.gguf Q4_K_M 1.67 GB Download
LFM2.5-VL-3B.Q4_K_S.gguf Q4_K_S 1.6 GB Download
LFM2.5-VL-3B.Q5_K_M.gguf Q5_K_M 1.94 GB Download
LFM2.5-VL-3B.Q5_K_S.gguf Q5_K_S 1.9 GB Download
LFM2.5-VL-3B.Q8_0.gguf Q8_0 2.87 GB Download
LFM2.5-VL-3B.mmproj-bf16.gguf mmproj-bf16 856 MB Download
LFM2.5-VL-3B.mmproj-f16.gguf mmproj-f16 856 MB Download
LFM2.5-VL-3B.mmproj-q8_0.gguf mmproj-q8_0 583 MB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
739
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/LFM2.5-VL-3B-GGUF

Quantized
(22)
this model

Collection including prithivMLmods/LFM2.5-VL-3B-GGUF