Instructions to use prithivMLmods/LightOnOCR-3-4B-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use prithivMLmods/LightOnOCR-3-4B-MLX with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("prithivMLmods/LightOnOCR-3-4B-MLX") config = load_config("prithivMLmods/LightOnOCR-3-4B-MLX") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use prithivMLmods/LightOnOCR-3-4B-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "prithivMLmods/LightOnOCR-3-4B-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prithivMLmods/LightOnOCR-3-4B-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use prithivMLmods/LightOnOCR-3-4B-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "prithivMLmods/LightOnOCR-3-4B-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prithivMLmods/LightOnOCR-3-4B-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prithivMLmods/LightOnOCR-3-4B-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "prithivMLmods/LightOnOCR-3-4B-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prithivMLmods/LightOnOCR-3-4B-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
LightOnOCR-3-4B-MLX
LightOnOCR-3-4B, developed by lightonai, is the largest and most accurate model in the LightOnOCR-3 family of end-to-end OCR vision-language models, built on the Qwen3.5-4B architecture and released under the Apache 2.0 license. Recommended for most OCR tasks, it serves as a powerful, single-model alternative to complex document understanding pipelines by combining high-quality text transcription (triggered by an empty prompt) with advanced visual understanding features. When used with the
groundingprompt, the model outputs labeled bounding boxes for all document elements, generates short descriptions for images, and extracts numerical data from charts into structured HTML tables. Highly versatile and optimized for seamless integration with Transformers and vLLM, it excels at processing complex layouts—including tables, receipts, forms, multi-column documents, and math notation—and performs best when documents are preprocessed at 400 DPI with a target longest dimension of 2048px.
Directory Structure
prithivMLmods/LightOnOCR-3-4B-MLX/
├── 4bit/
├── 8bit/
└── (root: bf16 base model mlx)
Use with mlx
Install the required library:
pip install -U mlx-vlm
BF16 Variant (Base Weights)
The unquantized BF16 mlx weights reside directly in the root repository directory:
CLI (Terminal)
# 1. Full-Page Transcription (Default Mode)
python -m mlx_vlm generate \
--model prithivMLmods/LightOnOCR-3-4B-MLX \
--max-tokens 2048 \
--temperature 0.0 \
--prompt "" \
--image <path_to_image>
# 2. Visual Grounding Mode (Extract Bounding Boxes & HTML Tables)
python -m mlx_vlm generate \
--model prithivMLmods/LightOnOCR-3-4B-MLX \
--max-tokens 2048 \
--temperature 0.0 \
--prompt "grounding" \
--image <path_to_image>
Python API
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
model_path = "prithivMLmods/LightOnOCR-3-4B-MLX"
model, processor = load(model_path)
config = load_config(model_path)
image = ["<path_to_image>"]
# Use "" for full-page markdown transcription, or "grounding" for coordinates and structured charts
prompt = ""
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))
output = generate(
model,
processor,
formatted_prompt,
image=image,
max_tokens=2048,
temperature=0.0
)
print(output.text)
8-bit Variant
Access the 8-bit quantized files using --subfolder 8bit:
CLI (Terminal)
# Full-Page Transcription
python -m mlx_vlm generate \
--model prithivMLmods/LightOnOCR-3-4B-MLX \
--subfolder 8bit \
--max-tokens 2048 \
--temperature 0.0 \
--prompt "" \
--image <path_to_image>
# Visual Grounding Mode
python -m mlx_vlm generate \
--model prithivMLmods/LightOnOCR-3-4B-MLX \
--subfolder 8bit \
--max-tokens 2048 \
--temperature 0.0 \
--prompt "grounding" \
--image <path_to_image>
Python API
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
model_path = "prithivMLmods/LightOnOCR-3-4B-MLX"
model, processor = load(model_path, subfolder="8bit")
config = load_config(model_path, subfolder="8bit")
image = ["<path_to_image>"]
prompt = "" # Or "grounding"
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))
output = generate(
model,
processor,
formatted_prompt,
image=image,
max_tokens=2048,
temperature=0.0
)
print(output.text)
4-bit Variant
Access the 4-bit quantized files using --subfolder 4bit:
CLI (Terminal)
# Full-Page Transcription
python -m mlx_vlm generate \
--model prithivMLmods/LightOnOCR-3-4B-MLX \
--subfolder 4bit \
--max-tokens 2048 \
--temperature 0.0 \
--prompt "" \
--image <path_to_image>
# Visual Grounding Mode
python -m mlx_vlm generate \
--model prithivMLmods/LightOnOCR-3-4B-MLX \
--subfolder 4bit \
--max-tokens 2048 \
--temperature 0.0 \
--prompt "grounding" \
--image <path_to_image>
Python API
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
model_path = "prithivMLmods/LightOnOCR-3-4B-MLX"
model, processor = load(model_path, subfolder="4bit")
config = load_config(model_path, subfolder="4bit")
image = ["<path_to_image>"]
prompt = "" # Or "grounding"
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))
output = generate(
model,
processor,
formatted_prompt,
image=image,
max_tokens=2048,
temperature=0.0
)
print(output.text)
License and Attribution
This model is based on and/or incorporates the following open-source projects and models:
- LightOnOCR-3-4B: https://huggingface.co/lightonai/LightOnOCR-3-4B
- mlx-vlm: https://github.com/Blaizzy/mlx-vlm
- MLX: https://github.com/ml-explore/mlx
This model is released under the Apache License 2.0.
- Downloads last month
- 3
4-bit
Model tree for prithivMLmods/LightOnOCR-3-4B-MLX
Base model
lightonai/LightOnOCR-3-4B