Instructions to use lightonai/LightOnOCR-3-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lightonai/LightOnOCR-3-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="lightonai/LightOnOCR-3-1B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSeq2SeqLM processor = AutoProcessor.from_pretrained("lightonai/LightOnOCR-3-1B") model = AutoModelForSeq2SeqLM.from_pretrained("lightonai/LightOnOCR-3-1B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use lightonai/LightOnOCR-3-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lightonai/LightOnOCR-3-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lightonai/LightOnOCR-3-1B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/lightonai/LightOnOCR-3-1B
- SGLang
How to use lightonai/LightOnOCR-3-1B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "lightonai/LightOnOCR-3-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lightonai/LightOnOCR-3-1B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "lightonai/LightOnOCR-3-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lightonai/LightOnOCR-3-1B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use lightonai/LightOnOCR-3-1B with Docker Model Runner:
docker model run hf.co/lightonai/LightOnOCR-3-1B
LightOnOCR-3-1B
Same architecture as LightOnOCR-2-1B, new capabilities. LightOnOCR-3-1B keeps the architecture of LightOn's previous LightOnOCR-2 models, so switching requires no change, and adds the new grounding, image description and chart extraction features.
About LightOnOCR-3
LightOnOCR-3 is a new family of highly performant lightweight OCR models. Compared to the previous generation, they bring significant improvements in speed and transcription quality and introduce new visual understanding features: the models now output bounding box coordinates with labels for all visual elements of a document, short descriptions of images, and the numerical data of figures and charts. With these capabilities, they offer a ready-to-use, easier-to-maintain alternative to complex document understanding pipelines.
The models come in three sizes. The 1B keeps the LightOnOCR-2-1B architecture, while the 0.8B and 4B adopt the Qwen3.5 vision-language architecture, which simplifies integration with existing tools and brings a significant speed-up. All models are released under the Apache 2.0 license for research and commercial use.
Highlights
- 📝 Transcription mode: call the model with an empty prompt and get the full page text, as before. Switching from LightOnOCR-2 requires no change
- 📍 Grounding mode: call it with the
groundingprompt and every block comes back with a label and a bounding box - 🖼️ Visual understanding: images get a short description, charts become a table of their data points
- 🧠 End-to-End: one model instead of a document understanding pipeline, easier to deploy and maintain
- 🧾 Versatile: handles tables, receipts, forms, charts, multi-column layouts, handwriting, and math notation
Model Variants
| Variant | Description |
|---|---|
| LightOnOCR-3-4B | Best OCR model, recommended for most tasks |
| LightOnOCR-3-1B | LightOnOCR-2 architecture, drop-in upgrade for existing deployments |
| LightOnOCR-3-0.8B | Fast and Efficient model |
Prompt Modes
The models can be used as before in a transcription-only mode: called with an empty prompt (the image only), they output all textual elements of the page as markdown. If you are currently using our previous models, switching to LightOnOCR-3 requires no change, as this is the default behavior.
The new usage mode is to call the models with the grounding prompt. They then output the new vision features alongside the transcribed text. Each block of content is prepended with a placeholder giving its type and its bounding box, in page coordinates normalized to 0–1000:

# Company report 2025

Revenue rose 20%.

A solar-powered factory.

<table>
<tr><th>Year</th><th>Revenue</th></tr>
<tr><td>2024</td><td>€10M</td></tr>
<tr><td>2025</td><td>€12M</td></tr>
</table>
- Text blocks contain the transcribed content of paragraphs, titles and other text elements.
- Image blocks pair a bounding box with a short description, making visual content accessible to retrieval and question-answering pipelines.
- Chart blocks contain an HTML table of the data points extracted from the figure, turning visual information into structured data.
The placeholder label is one of:
| Label | Definition |
|---|---|
text |
Ordinary body text and paragraphs. |
title |
Document, section, or paragraph headings. |
list |
Bulleted or numbered-list content. |
header |
Text in the page header. |
footer |
Text in the page footer. |
page_number |
Page-number regions. |
footnote |
Footnote text. |
caption |
Captions associated with figures, tables, or other visual elements. |
formula |
Mathematical formulas and equations. |
code |
Code blocks or code-like text. |
table |
Table regions and their content. |
image |
Illustrations, photographs, and other image regions. |
chart |
Charts, graphs, and plotted visual data. |
header_image |
Image regions in page headers, such as logos. |
footer_image |
Image regions in page footers. |
aside_text |
Marginal or side text outside the main text flow. |
+ suffix |
Marks a continuation block, not a separate semantic class. |
Grounding adds about 25% output tokens over plain transcription (1,433 vs 1,158 tokens per page on average for the 4B). Other instructions are out of distribution: use the empty prompt or grounding.
Usage with Transformers
Note: LightOnOCR-3-1B shares the LightOnOCR-2 architecture and loads with the
LightOnOcrclasses available in Transformers since v5.
uv pip install transformers pillow pypdfium2
import torch
from transformers import LightOnOcrForConditionalGeneration, LightOnOcrProcessor
model_id = "lightonai/LightOnOCR-3-1B"
device = "mps" if torch.backends.mps.is_available() else "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float32 if device == "mps" else torch.bfloat16
model = LightOnOcrForConditionalGeneration.from_pretrained(model_id, dtype=dtype).to(device)
processor = LightOnOcrProcessor.from_pretrained(model_id)
url = "https://huggingface.co/datasets/hf-internal-testing/fixtures_ocr/resolve/main/SROIE-receipt.jpeg"
mode = "grounding" # or "plain"
content = [{"type": "image", "url": url}]
if mode == "grounding":
content.append({"type": "text", "text": "grounding"})
conversation = [{"role": "user", "content": content}]
inputs = processor.apply_chat_template(
conversation,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
)
inputs = {k: v.to(device=device, dtype=dtype) if v.is_floating_point() else v.to(device) for k, v in inputs.items()}
output_ids = model.generate(**inputs, max_new_tokens=1024)
generated_ids = output_ids[0, inputs["input_ids"].shape[1]:]
print(processor.decode(generated_ids, skip_special_tokens=True))
Usage with vLLM
vllm serve lightonai/LightOnOCR-3-1B \
--limit-mm-per-prompt '{"image": 1}' --mm-processor-cache-gb 0 --no-enable-prefix-caching
import base64
import io
import pypdfium2 as pdfium
import requests
ENDPOINT = "http://localhost:8000/v1/chat/completions"
MODEL = "lightonai/LightOnOCR-3-1B"
MODE = "grounding" # or "plain"
# Render the first page of a PDF at 200 DPI (scale factor = 200/72 ≈ 2.77)
pdf = pdfium.PdfDocument(requests.get("https://arxiv.org/pdf/2412.13663").content)
pil_image = pdf[0].render(scale=2.77).to_pil()
buffer = io.BytesIO()
pil_image.save(buffer, format="PNG")
image_base64 = base64.b64encode(buffer.getvalue()).decode("utf-8")
content = [{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image_base64}"}}]
if MODE == "grounding":
content.append({"type": "text", "text": "grounding"})
payload = {
"model": MODEL,
"messages": [{"role": "user", "content": content}],
"max_tokens": 4096,
"temperature": 0.2,
"top_p": 0.9,
}
response = requests.post(ENDPOINT, json=payload)
print(response.json()["choices"][0]["message"]["content"])
The LightOnOCR repository on GitHub provides a minimal client, CLI and viewer for the models served with vLLM, plus the code to reproduce our benchmarks.
Rendering and Preprocessing Tips
- Render PDFs at 200 DPI to images using a target longest dimension of 1540px
- Maintain aspect ratio to preserve text geometry
License
Apache License 2.0
Citation
@misc{lightonocr2_2026,
title = {LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR},
author = {Said Taghadouini and Adrien Cavaill\`{e}s and Baptiste Aubertin},
year = {2026},
howpublished = {\url{https://arxiv.org/abs/2601.14251}}
}
- Downloads last month
- 11