Instructions to use litert-community/LFM2.5-VL-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/LFM2.5-VL-3B with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli litert-lm run \ --from-huggingface-repo=litert-community/LFM2.5-VL-3B \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/LFM2.5-VL-3B with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
LFM2.5-VL-3B β LiteRT-LM
LiquidAI/LFM2.5-VL-3B converted to the LiteRT-LM (.litertlm) format for on-device inference with Google's LiteRT-LM runtime.
Text + image work end-to-end on the released litert-lm 0.16.0 pip runtime β the bundle carries the vision encoder, the vision adapter and the LFM2 image-placeholder metadata, so litert-lm run β¦ --attachment photo.png just works. To our knowledge this is the first LFM2.5-VL in LiteRT form.
LFM2.5-VL-3B pairs Liquid AI's hybrid text backbone (22 gated short-convolution blocks + 8 grouped-query attention layers, 128k vocab) with a SigLIP2 vision tower (27 layers, hidden 1152) and a pixel-unshuffle projector. On this runtime an image is processed at 512Γ512 into 256 soft tokens (single image per prompt; the runtime resizes for you).
| File | Recipe | Size |
|---|---|---|
LFM2.5-VL-3B_int8.litertlm |
int8 dynamic (text linears + convs + embedding, vision tower) | 3.55 GB |
LFM2.5-VL-3B_int4.litertlm |
text int4 blockwise-32 OCTAV linears, int8 embedding + lm_head; vision tower int8 | 2.35 GB |
| Context (KV cache) | 4096 max |
| Image input | 1 per prompt, resized to 512Γ512 β 256 tokens; PNG/JPEG via --attachment |
| Backend | CPU, and GPU with litert-lm β₯ 0.16.0 (macOS verified by generation; Android OpenCL expected per the LFM2.5 family β device numbers below as measured; iOS Metal fails at engine creation for this family, tracked upstream in LiteRT-LM#3129 β use CPU on iOS) |
| Template | bundled β ChatML-style; image placeholders are inserted by the runtime's LFM2 data processor (non-thinking model) |
| Base model | LiquidAI/LFM2.5-VL-3B (LFM Open License v1.0) |
Quality
Sanity gates on the 0.16.0 pip CLI (greedy, fresh engine per question, --cache no). Vision: five deterministic synthetic fixtures (dominant color, large-text OCR, shape, counting three squares, largest word). Text: the 8-question gate used across our LiteRT conversions.
| Configuration | text 8Q | image 5Q |
|---|---|---|
| PyTorch bf16 (reference) | 7/8 | 5/5 |
| LiteRT int8, CPU | 7/8 | 5/5 |
| LiteRT int8, GPU (macOS) | 7/8 | 5/5 |
| LiteRT int4-b32, CPU | 7/8 | 5/5 |
| LiteRT int4-b32, GPU (macOS) | 7/8 | 5/5 |
All five image answers from both LiteRT variants are verbatim identical to the bf16 reference on both backends ("Red." / "Hello." / "Circle." / "Three." / "CAT."). The single text miss (rhyme completion answered "Purple.") is shared with the source model's behavior at greedy decoding, not a conversion artifact. Zero degenerate outputs across all runs.
Usage
pip install litert-lm
litert-lm run ./LFM2.5-VL-3B_int4.litertlm --prompt "Describe this image." --attachment photo.png
Text-only prompts work the same way without --attachment. --vision-backend cpu|gpu selects the vision encoder backend independently of the text backend.
Speed
litert-lm benchmark β¦ --cache no, litert-lm 0.16.0 pip, Apple M4 Max (128 GB), text path (prefill/decode; image encoding is a separate one-shot vision-encoder call at prompt time):
CPU backend:
| Variant | Prefill (256) | Prefill (1024) | Decode | TTFT |
|---|---|---|---|---|
| int8 | 156 tok/s | 387 tok/s | 36.0 tok/s | 1.67 s |
| int4 | 141 tok/s | 181 tok/s | 40.1 tok/s | 1.84 s |
GPU backend (--backend gpu; both variants verified to generate correct text and answer the image gate on GPU before quoting):
| Variant | Prefill (256) | Decode | TTFT |
|---|---|---|---|
| int8 | 1835 tok/s | 114.5 tok/s | 0.15 s |
| int4 | 1944 tok/s | 143.2 tok/s | 0.14 s |
On Android the same bundle runs GPU-accelerated. Pixel 8a (Tensor G3), litert_lm_main built from the v0.16.0 release tag, 292-token prompt, decode run to EOS (β1.1k tokens sustained), --disable_cache (worst-case load; default caching makes subsequent loads much faster), int4 file:
| Backend | Prefill (292 tok) | Decode | TTFT |
|---|---|---|---|
| GPU (OpenCL) | 91.6 tok/s | 10.8 tok/s | 3.3 s |
| CPU | 14.1 tok/s | 5.8 tok/s | 20.9 s |
The text graph delegates fully on Android OpenCL (937/937 nodes, zero rejected ops). One operational note: the GPU path writes multi-GB compile caches next to the model by default β on a nearly-full device the cache write can fail mid-initialization; run with --disable_cache or free storage first.
Conversion
Converted with the open pipeline in hf-to-litertlm (lfm_work/convert_lfm25_vl.py, litert-torch 0.9.3 --task image_text_to_text): the exact recipe, the int4 post-processing (OCTAV int4-b32 + int8 embedder + zero-scale repair + executor metadata, all vision sections preserved) and the text/image gate harnesses are in the repo's REPRODUCE.md.
- Downloads last month
- -
Model tree for litert-community/LFM2.5-VL-3B
Base model
LiquidAI/LFM2.5-VL-3B