Instructions to use APMIC/APMIC-OCR-Parse with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use APMIC/APMIC-OCR-Parse with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="APMIC/APMIC-OCR-Parse", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("APMIC/APMIC-OCR-Parse", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use APMIC/APMIC-OCR-Parse with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "APMIC/APMIC-OCR-Parse" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "APMIC/APMIC-OCR-Parse", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/APMIC/APMIC-OCR-Parse
- SGLang
How to use APMIC/APMIC-OCR-Parse with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "APMIC/APMIC-OCR-Parse" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "APMIC/APMIC-OCR-Parse", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "APMIC/APMIC-OCR-Parse" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "APMIC/APMIC-OCR-Parse", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use APMIC/APMIC-OCR-Parse with Docker Model Runner:
docker model run hf.co/APMIC/APMIC-OCR-Parse
APMIC-OCR-Parse
Model Description
APMIC-OCR-Parse is APMIC's document-parsing model for enterprise document intelligence: it turns document images (scanned or rendered PDFs, slides, forms, reports, tables and charts) into structured, machine-readable output with text, layout classes, bounding boxes and reading order — ready for RAG indexing, data extraction and agentic workflows.
The model is based on nvidia/NVIDIA-Nemotron-Parse-2.0, a vision-encoder-decoder model (< 1B parameters) with expanded multilingual OCR including CJK scripts, handwritten-text extraction, chart-to-table parsing and improved table structure recovery.
Model Details
- Developed by: APMIC, based on NVIDIA Nemotron Parse 2.0 by NVIDIA Corporation
- Model type: NemotronParseForConditionalGeneration (Transformers,
trust_remote_code=True) — vision-encoder-decoder- Vision encoder: ViT-H based on NVIDIA C-RADIO
- Adapter: 1D convolutions and normalization layers compressing the vision latent sequence
- Decoder: mBART decoder with 10 blocks; tokenizer with 72,256 entries
- Base model: nvidia/NVIDIA-Nemotron-Parse-2.0
- Input: one RGB document image + a task prompt (recommended resolution 1024×1280 to 1664×2048)
- Output: text with semantic classes (Title, Text, Table, Chart, Picture, Caption, Page-header/footer, Footnote, Bibliography …) and bounding boxes
- Weights precision: bfloat16
- License: OpenMDW-1.1 (model and code); tokenizer under CC-BY-4.0
Usage
Transformers
Install the dependencies listed in the base model card (transformers==5.6.1, timm, open_clip_torch, einops,
beautifulsoup4, accelerate), then:
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor, GenerationConfig
model_id = "APMIC/APMIC-OCR-Parse"
device = "cuda:0"
model = AutoModel.from_pretrained(
model_id, trust_remote_code=True, torch_dtype=torch.bfloat16
).to(device).eval()
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
generation_config = GenerationConfig.from_pretrained(model_id, trust_remote_code=True)
image = Image.open("document.png")
task_prompt = "</s><s><predict_bbox><predict_classes><output_markdown><predict_no_text_in_pic>"
inputs = processor(images=[image], text=task_prompt, return_tensors="pt", add_special_tokens=False).to(device)
outputs = model.generate(**inputs, generation_config=generation_config)
generated_text = processor.batch_decode(outputs, skip_special_tokens=True)[0]
Use postprocessing.py in this repository (extract_classes_bboxes, transform_bbox_to_original,
postprocess_text) to map boxes back to the original image and to render tables as LaTeX, HTML, markdown, JSON or CSV.
Task prompts
| Purpose | Prompt |
|---|---|
| Boxes + classes + markdown text (default) | </s><s><predict_bbox><predict_classes><output_markdown><predict_no_text_in_pic> |
| Also extract text inside pictures | </s><s><predict_bbox><predict_classes><output_markdown><predict_text_in_pic> |
| Boxes + classes only | </s><s><predict_bbox><predict_classes><output_no_text><predict_no_text_in_pic> |
Serving
📘 部署指南(繁體中文): DEPLOYMENT.md 說明如何以 vLLM + OpenAI 相容代理架設服務,直接回傳乾淨的 Markdown,並包含驗證步驟與常見問題排除。
The model can be served with vLLM (v0.20–v0.26) using --trust-remote-code:
vllm serve APMIC/APMIC-OCR-Parse \
--dtype bfloat16 \
--max-num-seqs 8 \
--limit-mm-per-prompt '{"image": 1}' \
--trust-remote-code \
--port 8000
The tied-embedding runtime patch (vllm_tied_patch/), optional logits processors (logitsprocs/) and further examples
(vllm_example.py, example_with_processor.py) are included in this repository; see the
base model card for details. The weights take about 3.6 GB.
Evaluation
Results reported by NVIDIA for Nemotron Parse 2.0 (identical weights):
| Benchmark | Metric | Nemotron Parse v1.2 | Nemotron Parse 2.0 |
|---|---|---|---|
| ParseBench | Overall score | 0.5782 | 0.6391 |
| OmniDocBench Notes (Handwriting) | Text edit distance (lower is better) | 0.9739 | 0.3395 |
| IndicVisionBench | Overall ANLS character | 0.0612 | 0.7203 |
| MOSCAR (Multilingual) | Overall BoC F1 | 0.4410 | 0.9102 |
Intended Use and Limitations
Intended for
- Converting scanned or rendered document pages into structured text for RAG, search and analytics
- Table and chart extraction from reports, financial statements and forms
- Building document-understanding and training-data curation pipelines
Limitations
- The model processes one page image at a time; multi-page documents must be split into pages.
- OCR output can contain errors or hallucinated text, especially on low-resolution, dense or heavily stylized pages. Keep human review in place for high-stakes extraction.
- The model reproduces any visible text, including personal or confidential data; make sure you have the rights to process the input documents.
See the base model's subcards for further details: Bias, Explainability, Safety & Security, Privacy.
License
This model is a redistribution of NVIDIA Nemotron Parse 2.0. The model files and source code are licensed under the OpenMDW License Agreement, version 1.1 (see LICENSE); the tokenizer is licensed under CC-BY-4.0.
Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
- Downloads last month
- -
Model tree for APMIC/APMIC-OCR-Parse
Base model
nvidia/NVIDIA-Nemotron-Parse-2.0
