LightOnOCR 3 4B for LM-Kit

This repository holds LightOnOCR-3-4B packaged for LM-Kit.NET, LM-Kit's on-device AI SDK for .NET, and for LM-Kit One.

File Content
lightonocr-3-4b-Q4_K_M.lmk The model and its vision projector, both quantized to Q4_K_M, in one archive. This is the file LM-Kit downloads.
lightonocr-3-4b-F16.gguf The language model at full precision (F16).
lightonocr-3-4b-mmproj-F16.gguf The vision projector at full precision (F16).

About LightOnOCR 3

LightOnOCR 3 is LightOn's third generation of end-to-end OCR models. It comes in three sizes:

LM-Kit model ID Architecture Notes
lightonocr-3:0.8b Qwen3.5 vision-language The fastest size
lightonocr-3:1b LightOnOCR 2 (Pixtral vision encoder, Qwen3 decoder) A drop-in upgrade for LightOnOCR 2 deployments
lightonocr-3:4b Qwen3.5 vision-language The most accurate size

The models have two modes:

  • Transcription: the page is returned as clean Markdown in natural reading order, with tables in HTML and formulas in LaTeX.
  • Grounding: every block of the page is returned with a label (title, text, list, table, formula, caption, header, footer, page number, footnote, image, chart, ...) and a bounding box. Images get a short description, and charts become an HTML table of their data points.

LM-Kit maps both modes onto VlmOcr intents, so you never write a prompt:

VlmOcrIntent Mode Result
Markdown, PlainText Transcription The page text
OcrWithCoordinates Grounding The page's text blocks, each with its bounding box in source image pixels and its layout category
LayoutAnalysis Grounding Every block, figures included, with its category; the JSON payload in VlmOcrResult.NormalizedText
TableRecognition, FormulaRecognition, ChartRecognition Grounding The page's tables, formulas or chart data tables

Pages are fed at the resolution the models were trained on (a 2048 px longest edge at ImageDetail.High).

Usage

using LMKit.Document.Conversion;
using LMKit.Extraction.Ocr;
using LMKit.Model;

LM model = LM.LoadFromModelID("lightonocr-3:4b");

// Document to Markdown.
var converter = new DocumentToMarkdown(model);
string markdown = converter.Convert("report.pdf").Markdown;

// Layout analysis: positioned, categorized blocks.
var ocr = new VlmOcr(model, VlmOcrIntent.LayoutAnalysis);
var result = ocr.Run(new LMKit.Data.Attachment("page.png"));

foreach (var element in result.PageElement.TextElements)
{
    Console.WriteLine($"{element.Category} at ({element.Left:0}, {element.Top:0}): {element.Text}");
}

License

The model weights are released by LightOn under the Apache License 2.0. See the original model card for training details and the authors' evaluation.

Downloads last month
38
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for lm-kit/lightonocr-3-4b-lmk

Quantized
(11)
this model