LFM2.5-350M-IT-Extract-GGUF

GGUF build of andreagemelli/LFM2.5-350M-IT-Extract, a 350M model fine-tuned for key information extraction (KIE) from Italian forms: give it a document's text plus a field schema, get a JSON object back.

This is the file that ships inside Scrivano — a fully offline desktop app that reads Italian forms on a laptop CPU, no API key and no network. The whole point of the quantised build is that it's small enough to bundle and run locally.

Quantisation Q4_K_M
Size 219 MiB
Base (fine-tuned) model andreagemelli/LFM2.5-350M-IT-Extract
Original base LiquidAI/LFM2.5-350M
Dataset andreagemelli/xfund-kie-it
Language Italian

What it does

The model takes a system prompt describing the fields to extract (each a key + a natural-language description) and a user message containing the document text, and returns a JSON object with the requested keys, omitting any field it can't find.

It was fine-tuned on prompts averaging seven fields, so it works best with a tight schema. It learned the shape of the task — find the value, put it under the right key, close the brace — not Italian knowledge. This is a 350M model: read the output, don't trust it.

Usage

llama.cpp

# server
llama-server -hf andreagemelli/LFM2.5-350M-IT-Extract-GGUF -c 4096

# one-shot
llama-cli -hf andreagemelli/LFM2.5-350M-IT-Extract-GGUF -p "..."

llama-cpp-python

from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="andreagemelli/LFM2.5-350M-IT-Extract-GGUF",
    filename="*Q4_K_M.gguf",
    n_ctx=4096,
)

system = (
    "Extract the following fields from the document and return a JSON object. "
    "Omit any field not present.\n"
    "- nome-completo: the person's full name\n"
    "- codice-fiscale: the Italian tax code\n"
    "- data-nascita: the date of birth\n"
    "- indirizzo: the residential address"
)
document = "…text of the Italian form…"

out = llm.create_chat_completion(
    messages=[
        {"role": "system", "content": system},
        {"role": "user", "content": document},
    ],
    temperature=0.0,
)
print(out["choices"][0]["message"]["content"])

The chat template is embedded in the GGUF. Greedy decoding (temperature = 0) is recommended for deterministic extraction — that's what Scrivano ships.

Evaluation

Measured on the 50-document Italian validation split of xfund-kie-it, greedy decoding.

Model Avg F1 JSON parse failures
LiquidAI/LFM2-350M-Extract (reference) 0.2532 6 / 50
LiquidAI/LFM2.5-350M (base) 0.2138 0 / 50
andreagemelli/LFM2.5-350M-IT-Extract 0.5311 10 / 50
same, Q4_K_M GGUF 0.5307 10 / 50
same, Q4_K_M on the app's own OCR text 0.4362 10 / 50

Quantising to Q4_K_M is essentially free (0.5311 → 0.5307): the thing you ship is the thing you measured. End-to-end, OCR costs about 0.10 F1.

These are upper bounds: the schema is oracle-filtered (the model is told which fields the document contains), parse failures are excluded rather than scored zero, and the metric double-counts a wrong value. Reproduce with uv run main.py from the repo.

Limitations

  • Beta; first fine-tune on 149 training documents, so it's very sensitive to the schema descriptions — edit those before blaming the document, and prefer a tight schema.
  • Only single-page Italian extraction is tested.
  • Valid JSON is not guaranteed (10/50 validation outputs didn't parse strictly).
  • A right-looking value placed under the wrong key is not flagged.

Links

Downloads last month
-
GGUF
Model size
0.4B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for andreagemelli/LFM2.5-350M-IT-Extract-GGUF

Quantized
(1)
this model

Dataset used to train andreagemelli/LFM2.5-350M-IT-Extract-GGUF

Collection including andreagemelli/LFM2.5-350M-IT-Extract-GGUF