LightOnOCR-3-4B-MLX

LightOnOCR-3-4B, developed by lightonai, is the largest and most accurate model in the LightOnOCR-3 family of end-to-end OCR vision-language models, built on the Qwen3.5-4B architecture and released under the Apache 2.0 license. Recommended for most OCR tasks, it serves as a powerful, single-model alternative to complex document understanding pipelines by combining high-quality text transcription (triggered by an empty prompt) with advanced visual understanding features. When used with the grounding prompt, the model outputs labeled bounding boxes for all document elements, generates short descriptions for images, and extracts numerical data from charts into structured HTML tables. Highly versatile and optimized for seamless integration with Transformers and vLLM, it excels at processing complex layouts—including tables, receipts, forms, multi-column documents, and math notation—and performs best when documents are preprocessed at 400 DPI with a target longest dimension of 2048px.

Directory Structure

prithivMLmods/LightOnOCR-3-4B-MLX/
├── 4bit/
├── 8bit/
└── (root: bf16 base model mlx)

Use with mlx

Install the required library:

pip install -U mlx-vlm

BF16 Variant (Base Weights)

The unquantized BF16 mlx weights reside directly in the root repository directory:

CLI (Terminal)

# 1. Full-Page Transcription (Default Mode)
python -m mlx_vlm generate \
  --model prithivMLmods/LightOnOCR-3-4B-MLX \
  --max-tokens 2048 \
  --temperature 0.0 \
  --prompt "" \
  --image <path_to_image>

# 2. Visual Grounding Mode (Extract Bounding Boxes & HTML Tables)
python -m mlx_vlm generate \
  --model prithivMLmods/LightOnOCR-3-4B-MLX \
  --max-tokens 2048 \
  --temperature 0.0 \
  --prompt "grounding" \
  --image <path_to_image>

Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

model_path = "prithivMLmods/LightOnOCR-3-4B-MLX"
model, processor = load(model_path)
config = load_config(model_path)

image = ["<path_to_image>"]
# Use "" for full-page markdown transcription, or "grounding" for coordinates and structured charts
prompt = ""
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))

output = generate(
    model, 
    processor, 
    formatted_prompt, 
    image=image, 
    max_tokens=2048, 
    temperature=0.0
)
print(output.text)

8-bit Variant

Access the 8-bit quantized files using --subfolder 8bit:

CLI (Terminal)

# Full-Page Transcription
python -m mlx_vlm generate \
  --model prithivMLmods/LightOnOCR-3-4B-MLX \
  --subfolder 8bit \
  --max-tokens 2048 \
  --temperature 0.0 \
  --prompt "" \
  --image <path_to_image>

# Visual Grounding Mode
python -m mlx_vlm generate \
  --model prithivMLmods/LightOnOCR-3-4B-MLX \
  --subfolder 8bit \
  --max-tokens 2048 \
  --temperature 0.0 \
  --prompt "grounding" \
  --image <path_to_image>

Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

model_path = "prithivMLmods/LightOnOCR-3-4B-MLX"
model, processor = load(model_path, subfolder="8bit")
config = load_config(model_path, subfolder="8bit")

image = ["<path_to_image>"]
prompt = ""  # Or "grounding"
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))

output = generate(
    model, 
    processor, 
    formatted_prompt, 
    image=image, 
    max_tokens=2048, 
    temperature=0.0
)
print(output.text)

4-bit Variant

Access the 4-bit quantized files using --subfolder 4bit:

CLI (Terminal)

# Full-Page Transcription
python -m mlx_vlm generate \
  --model prithivMLmods/LightOnOCR-3-4B-MLX \
  --subfolder 4bit \
  --max-tokens 2048 \
  --temperature 0.0 \
  --prompt "" \
  --image <path_to_image>

# Visual Grounding Mode
python -m mlx_vlm generate \
  --model prithivMLmods/LightOnOCR-3-4B-MLX \
  --subfolder 4bit \
  --max-tokens 2048 \
  --temperature 0.0 \
  --prompt "grounding" \
  --image <path_to_image>

Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

model_path = "prithivMLmods/LightOnOCR-3-4B-MLX"
model, processor = load(model_path, subfolder="4bit")
config = load_config(model_path, subfolder="4bit")

image = ["<path_to_image>"]
prompt = ""  # Or "grounding"
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))

output = generate(
    model, 
    processor, 
    formatted_prompt, 
    image=image, 
    max_tokens=2048, 
    temperature=0.0
)
print(output.text)

License and Attribution

This model is based on and/or incorporates the following open-source projects and models:

This model is released under the Apache License 2.0.

Downloads last month
3
Safetensors
Model size
5B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/LightOnOCR-3-4B-MLX

Quantized
(3)
this model

Collections including prithivMLmods/LightOnOCR-3-4B-MLX