You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

APMIC-OCR-Parse

APMIC-logo-橫-黑 NVIDIA-NeMo

Model Description

APMIC-OCR-Parse is APMIC's document-parsing model for enterprise document intelligence: it turns document images (scanned or rendered PDFs, slides, forms, reports, tables and charts) into structured, machine-readable output with text, layout classes, bounding boxes and reading order — ready for RAG indexing, data extraction and agentic workflows.

The model is based on nvidia/NVIDIA-Nemotron-Parse-2.0, a vision-encoder-decoder model (< 1B parameters) with expanded multilingual OCR including CJK scripts, handwritten-text extraction, chart-to-table parsing and improved table structure recovery.


Model Details

  • Developed by: APMIC, based on NVIDIA Nemotron Parse 2.0 by NVIDIA Corporation
  • Model type: NemotronParseForConditionalGeneration (Transformers, trust_remote_code=True) — vision-encoder-decoder
    • Vision encoder: ViT-H based on NVIDIA C-RADIO
    • Adapter: 1D convolutions and normalization layers compressing the vision latent sequence
    • Decoder: mBART decoder with 10 blocks; tokenizer with 72,256 entries
  • Base model: nvidia/NVIDIA-Nemotron-Parse-2.0
  • Input: one RGB document image + a task prompt (recommended resolution 1024×1280 to 1664×2048)
  • Output: text with semantic classes (Title, Text, Table, Chart, Picture, Caption, Page-header/footer, Footnote, Bibliography …) and bounding boxes
  • Weights precision: bfloat16
  • License: OpenMDW-1.1 (model and code); tokenizer under CC-BY-4.0

Usage

Transformers

Install the dependencies listed in the base model card (transformers==5.6.1, timm, open_clip_torch, einops, beautifulsoup4, accelerate), then:

import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor, GenerationConfig

model_id = "APMIC/APMIC-OCR-Parse"
device = "cuda:0"

model = AutoModel.from_pretrained(
    model_id, trust_remote_code=True, torch_dtype=torch.bfloat16
).to(device).eval()
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
generation_config = GenerationConfig.from_pretrained(model_id, trust_remote_code=True)

image = Image.open("document.png")
task_prompt = "</s><s><predict_bbox><predict_classes><output_markdown><predict_no_text_in_pic>"

inputs = processor(images=[image], text=task_prompt, return_tensors="pt", add_special_tokens=False).to(device)
outputs = model.generate(**inputs, generation_config=generation_config)
generated_text = processor.batch_decode(outputs, skip_special_tokens=True)[0]

Use postprocessing.py in this repository (extract_classes_bboxes, transform_bbox_to_original, postprocess_text) to map boxes back to the original image and to render tables as LaTeX, HTML, markdown, JSON or CSV.

Task prompts

Purpose Prompt
Boxes + classes + markdown text (default) </s><s><predict_bbox><predict_classes><output_markdown><predict_no_text_in_pic>
Also extract text inside pictures </s><s><predict_bbox><predict_classes><output_markdown><predict_text_in_pic>
Boxes + classes only </s><s><predict_bbox><predict_classes><output_no_text><predict_no_text_in_pic>

Serving

📘 部署指南(繁體中文): DEPLOYMENT.md 說明如何以 vLLM + OpenAI 相容代理架設服務,直接回傳乾淨的 Markdown,並包含驗證步驟與常見問題排除。

The model can be served with vLLM (v0.20–v0.26) using --trust-remote-code:

vllm serve APMIC/APMIC-OCR-Parse \
    --dtype bfloat16 \
    --max-num-seqs 8 \
    --limit-mm-per-prompt '{"image": 1}' \
    --trust-remote-code \
    --port 8000

The tied-embedding runtime patch (vllm_tied_patch/), optional logits processors (logitsprocs/) and further examples (vllm_example.py, example_with_processor.py) are included in this repository; see the base model card for details. The weights take about 3.6 GB.


Evaluation

Results reported by NVIDIA for Nemotron Parse 2.0 (identical weights):

Benchmark Metric Nemotron Parse v1.2 Nemotron Parse 2.0
ParseBench Overall score 0.5782 0.6391
OmniDocBench Notes (Handwriting) Text edit distance (lower is better) 0.9739 0.3395
IndicVisionBench Overall ANLS character 0.0612 0.7203
MOSCAR (Multilingual) Overall BoC F1 0.4410 0.9102

Intended Use and Limitations

Intended for

  • Converting scanned or rendered document pages into structured text for RAG, search and analytics
  • Table and chart extraction from reports, financial statements and forms
  • Building document-understanding and training-data curation pipelines

Limitations

  • The model processes one page image at a time; multi-page documents must be split into pages.
  • OCR output can contain errors or hallucinated text, especially on low-resolution, dense or heavily stylized pages. Keep human review in place for high-stakes extraction.
  • The model reproduces any visible text, including personal or confidential data; make sure you have the rights to process the input documents.

See the base model's subcards for further details: Bias, Explainability, Safety & Security, Privacy.


License

This model is a redistribution of NVIDIA Nemotron Parse 2.0. The model files and source code are licensed under the OpenMDW License Agreement, version 1.1 (see LICENSE); the tokenizer is licensed under CC-BY-4.0.

Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Downloads last month
-
Safetensors
Model size
0.9B params
Tensor type
I64
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for APMIC/APMIC-OCR-Parse

Finetuned
(1)
this model