LFM2.5-1.2B-Instruct โ€” GGUF (iPhone-optimized)

The official Q4_K_M GGUF of LiquidAI/LFM2.5-1.2B-Instruct, mirrored for on-device inference on iPhone, iPad, and Apple Silicon Mac via llama.cpp or apps that wrap it (e.g. Haplo).

Hosted by jc-builds for the Haplo ecosystem. The file is byte-identical to Liquid AI's official LFM2.5-1.2B-Instruct-Q4_K_M.gguf. Original weights ยฉ Liquid AI, redistributed under the LFM Open License v1.0 (see LICENSE).

TL;DR

A 1.2B instruct model from Liquid AI built on a hybrid architecture (short-convolution blocks mixed with a few grouped-query attention layers), which keeps the KV cache small and decoding fast on phones. Strong instruction following for its size. See the upstream model card for benchmarks.

Available quantizations

File Size Recommended use
LFM2.5-1.2B-Instruct-Q4_K_M.gguf 0.73 GB Default โ€” works on every device

Details

Parameters 1.2B
Architecture lfm2 (10 short-conv + 6 GQA layers)
Quantization Q4_K_M
Chat format ChatML
Minimum device Any iPhone that runs Haplo

How to use

Download URL:

https://huggingface.co/jc-builds/LFM2.5-1.2B-Instruct-GGUF/resolve/main/LFM2.5-1.2B-Instruct-Q4_K_M.gguf

llama.cpp

llama-cli -hf jc-builds/LFM2.5-1.2B-Instruct-GGUF:Q4_K_M

License

LFM Open License v1.0 (see LICENSE). It is based on Apache 2.0 with one addition: commercial use is licensed only while the licensee's annual revenue (including affiliates) is below US$10,000,000. Review the full terms before commercial use.

Downloads last month
53
GGUF
Model size
1B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jc-builds/LFM2.5-1.2B-Instruct-GGUF

Quantized
(101)
this model