Aeddix Alpine OCR (1.2B)

Aeddix Alpine OCR turns a page image into Markdown: text in reading order, tables as HTML, charts as Markdown data tables, formulas as LaTeX, and inline styling (bold, italic, superscript, subscript) where the page has it. It is a fine-tune of MinerU2.5-Pro-2605-1.2B by OpenDataLab and keeps its architecture, prompts and two-step pipeline: the model first finds the page's blocks and their order, then reads each block.

This is a research release. It is free to use under Apache-2.0, but it was tuned against one benchmark and has the limitations listed below.

Results on ParseBench

ParseBench scores document parsers on about 2,000 real enterprise pages in five dimensions. The same weights score differently depending on the inference pipeline around them, so each number below names its pipeline. All runs are ours, on all 2,078 pages, with no failed page.

Through ParseBench's kdl_frontier_nano pipeline: 76.77

kdl_frontier_nano is the inference pipeline KoreaDeep built for KDL-Frontier-Parser-nano and contributed to ParseBench (run-llama/ParseBench#49): a layout pass, a crop of each region, one recognition pass per region and rule-based Markdown assembly, all against one vLLM endpoint. We ran it unchanged, with these weights served in place of KDL's, the way florin-parser-nano and rakedoc-nano entered the ParseBench leaderboard. The pipeline is KoreaDeep's work; the weights are ours.

Model, through kdl_frontier_nano Tables Charts Content faithfulness Semantic formatting Visual grounding Overall
MinerU2.5-Pro-2605-1.2B (base), 1 run 84.72 69.15 87.66 55.25 74.65 74.29
Aeddix Alpine OCR v0.1, mean of 3 runs 85.89 68.78 87.90 66.05 75.21 76.77
Standard deviation of the 3 runs 0.15 0.12 0.08 0.24 0.09 0.02
  • The three runs scored 76.76, 76.79 and 76.75 overall.
  • Paired on the same documents, the fine-tune adds +2.48 overall to the base model through the same pipeline, almost all of it in semantic formatting (+10.80) and tables (+1.17); charts are level (−0.37, within one standard error).
  • Run 1 ran on one NVIDIA L4 with ParseBench afb36bd; runs 2 and 3 and the base run on one NVIDIA A10 (24 GB) with ParseBench cbd57f2. Re-scoring run 1's saved outputs with cbd57f2 changes no document's score (1,741 of its 2,078 page results were kept), so the three runs are comparable.
  • Every run: vLLM 0.28.0 serving the v0.1 weights with KDL's serving settings (below), and the provider's defaults (144 dpi, 20 documents at a time, LLM normalisation off).

Through our own page server: 75.15

The page server in this repository's inference/ folder runs the model the way MinerU2.5-Pro runs (mineru-vl-utils' two-step extraction) behind ParseBench's mineru2605pro page API, and returns typed blocks with boxes as well as Markdown. Measured with ParseBench afb36bd (2026-09-29) on one NVIDIA L4 (vLLM 0.28.0, mineru-vl-utils 2.0.5, 150 dpi):

Model, through our page server Tables Charts Content faithfulness Semantic formatting Visual grounding Overall
MinerU2.5-Pro-2605-1.2B (base) 77.30 59.13 87.86 49.59 72.83 69.34
Aeddix Alpine OCR v0.1, heading and photo rules (the first release's server) 77.12 65.28 88.27 65.74 72.78 73.84
Aeddix Alpine OCR v0.1, plus the table-header rule (today's server) 83.67 65.28 88.27 65.74 72.78 75.15
  • Overall is the unweighted mean of the five headline metrics: grits_trm_composite, chart rule_pass_rate, content_faithfulness, semantic_formatting and layout_element_rule_pass_rate.
  • Paired on the same documents, the fine-tune adds +4.38 overall to the base model on this server (document bootstrap 95% interval +3.60 to +5.17), almost all of it in charts (+6.15, standard error 1.27) and semantic formatting (+15.99, standard error 1.45). That comparison is before the Markdown rules described below: the heading and photo rules add +0.12 overall, the table-header rule +1.31 more.
  • Tables, content and grounding are level with the base model within one standard error before the table-header rule.
  • Each number is one full pass; repeated passes of the same stack matched to four decimals. The table-header row is the 73.84 pass's saved blocks with the rule applied on a CPU and re-scored; the server's rule code produces that Markdown byte for byte.

Which number to read

  • 76.77 is the number for these weights on the ParseBench leaderboard's terms: a pipeline anyone can run from ParseBench's own repository.
  • 75.15 is what this repository's page server gives, with typed blocks and boxes in its output; on the same scorer (afb36bd) KDL's pipeline is 1.61 points ahead of it.
  • Leaderboard rows carry the scorer of their date. The June 2026 rows for the base model (72.78) and KDL-Frontier-Parser-nano (76.36) were scored by the maintainers' harness of that time, which credited formatting and grounding more generously than the public scorer does today. On today's public scorer the base model scores 69.34 on our server and 74.29 through KDL's pipeline; compare our numbers with those, not with 72.78.

How to run it

Through KDL's pipeline (76.77)

Serve the weights with vLLM exactly as KDL-Frontier-Parser-nano is served, then run ParseBench's kdl_frontier_nano provider against them:

vllm serve aeddix-labs/aeddix-alpine-ocr --revision v0.1 \
  --served-model-name kdl-frontier-parser-nano \
  --max-model-len 8192 --gpu-memory-utilization 0.85 \
  --max-num-seqs 24 --trust-remote-code \
  --limit-mm-per-prompt '{"image":1}'

git clone https://github.com/run-llama/ParseBench && cd ParseBench
git checkout cbd57f251c207103de14342fd1783eaa4acf9583
uv sync --extra runners
uv run parse-bench download
KDL_NANO_ENDPOINT_URL=http://localhost:8000/v1 uv run parse-bench run kdl_frontier_nano

kdl-frontier-parser-nano is the provider's default model name; to serve under another name, set KDL_NANO_MODEL to it. Inference took 76 to 77 minutes on one A10 and 96 minutes on one L4, plus a few minutes of scoring. Once ParseBench adds the aeddix_alpine_ocr_kdl pipeline (the same provider, with its own AEDDIX_ALPINE_OCR_KDL_ENDPOINT_URL), parse-bench run aeddix_alpine_ocr_kdl is the same run under this model's name.

Through our page server (75.15)

The page server in this repository's inference/ folder serves the model behind ParseBench's mineru2605pro page API and returns each page's blocks as well as its Markdown.

hf download aeddix-labs/aeddix-alpine-ocr --include "inference/*" --local-dir alpine-ocr
pip install ./alpine-ocr/inference          # vLLM 0.28.0, mineru-vl-utils 2.0.5
alpine-ocr-server --model aeddix-labs/aeddix-alpine-ocr --revision v0.1 --port 8765
import base64, requests

page = base64.b64encode(open("page.png", "rb").read()).decode()   # PDFs: render each page, 150 dpi
result = requests.post("http://127.0.0.1:8765/predict", json={"image_base64": page}).json()
print(result["markdown"])          # the page as Markdown
print(result["blocks"][:3])        # typed blocks with boxes normalised to [0, 1]

To reproduce 75.15: REVISION=v0.1 bash inference/scripts/reproduce_parsebench.sh (about 55 minutes on an L4).

The weights also work with mineru-vl-utils directly, as for the base model:

from PIL import Image
from vllm import LLM
from mineru_vl_utils import MinerUClient, MinerULogitsProcessor
from mineru_vl_utils.post_process import json2md

llm = LLM(model="aeddix-labs/aeddix-alpine-ocr", revision="v0.1", logits_processors=[MinerULogitsProcessor])
client = MinerUClient(backend="vllm-engine", vllm_llm=llm, image_analysis=True, enable_table_formula_eq_wrap=True)
print(json2md(client.two_step_extract(Image.open("page.png"))))

That path skips the server's four additions: nested-chart analysis and the three Markdown rules. Use image_analysis=True, or charts come back empty.

What the server adds

Addition What it does Effect on ParseBench
Nested charts mineru-vl-utils 2.0.5 skips any chart or image lying inside a multi-chart figure (image_block); the server analyses each panel like a standalone chart base model charts 59.13 to 64.47 (+5.34)
Heading levels a numbered title takes its number's depth (2.1 is ###); other titles are ranked by font size on the page; MinerU itself gives every heading the same label formatting +0.16 (standard error 0.11)
Photo descriptions out a block MinerU classes as a natural photo carries a description the model wrote ("Portrait of a man in a suit"), not page text; it stays out of the Markdown, and its block and box stay content +0.42 (standard error 0.11)
Table headers ParseBench reads only a table's first row as its header unless the Markdown marks more; a multi-row header becomes the table's <thead>: the rows the top row's rowspans cover, or a top row of column groups over a row with no numeric cell tables +6.54 (standard error 0.69), overall +1.31

None of them adds text or styling the model did not read. --md-rules "" and --nested-charts 0 switch them off.

How it was trained

Three steps, each from the previous one.

  1. Base: opendatalab/MinerU2.5-Pro-2605-1.2B at commit bff20d4.
  2. Supervised fine-tune: 20,000 samples (19,600 train, 400 validation), one epoch, full fine-tune of the language model and the vision-language projector with the vision tower frozen (524.8M trainable parameters). AdamW, learning rate 6e-6 with cosine decay and 10% warm-up, effective batch 16, bf16 with fp32 master weights, MinerU's own prompts and image preprocessing, loss on answer tokens only. The recipe follows the alignment stage of jina-ocr-v1 (arXiv 2609.03181): its learning rate, trainable set and data cleaning (degeneration-loop filter, duplicate images, samples under 64 tokens except charts and formatting).
  3. GRPO: two epochs of 100 steps on 1,600 prompts (783 charts, 611 formatting, 206 tables), LoRA rank 64 on the language model, merged into the weights afterwards. DAPO loss (token-level, clip 0.2 and 0.28, no KL), 8 samples per prompt at temperature 1.0, and a ReMax baseline (advantage = sampled reward minus the greedy answer's reward). Rewards are built from ParseBench's own metric code applied to the training items' known answers, never to ParseBench pages: chart data-point rules, the GriTS table metric and styled-span F-beta, each multiplied by a structure-validity term and a repetition penalty. Each epoch was kept only if no ParseBench dimension fell more than one paired standard error below the starting checkpoint.

Training data

Every training item was checked against all 2,078 ParseBench pages and removed on any overlap: a 256-bit perceptual page hash, the source file name, any shared 13-token sequence with the page text or ground truth, and for Chinese and Japanese text any shared 8-character sequence. That removed 310 of 28,600 supervised candidates and 254 of 77,350 GRPO-pool candidates, most of them real sentence overlaps with public financial filings.

Dataset Licence Used for How
ibm-granite/ChartNet, core_permissive subset only CDLA-Permissive-2.0 4,200 SFT charts, 686 GRPO prompts chart image; its data CSV as the Markdown table
docling-project/SynthChartNet CDLA-Permissive-2.0 2,000 SFT charts, 97 GRPO prompts chart image; its data table
docling-project/SynthTabNet_OTSL CDLA-Permissive-1.0 (IBM/SynthTabNet) 2,500 SFT tables, 64 GRPO prompts table image; its structure label
docling-project/FinTabNet_OTSL CDLA-Permissive-1.0 (IBM FinTabNet) 142 GRPO prompts table image (annual-report tables); reward against its structure label
docling-project/DocLayNet-v1.2 CDLA-Permissive-1.0 5,199 layout pages, 2,700 formatting crops, 901 text crops human layout boxes; bold and italic from the PDF's font names
lightonai/LightOnOCR-mix-0126 Common Crawl and Digital Corpora terms sentences for 2,500 SFT renders and 546 GRPO prompts plain prose sentences only, rendered by us with styling
HuggingFaceFW/fineweb-2, Japanese and Chinese ODC-By 1.0 sentences for 65 GRPO prompts rendered by us with Noto Sans CJK
The base model's own readings Apache-2.0 (the base model) targets of the DocLayNet crops, reading order of the layout pages self-distillation; no other model wrote any label
  • Rendered samples use DejaVu, Liberation and Noto Sans CJK fonts; the images are ours.
  • The page images in DocLayNet and FinTabNet are public documents whose copyright stays with their publishers; the datasets' licences cover their use as training data.
  • Not used, because of their licences: datasets labelled by closed models (olmOCR-mix, GPT-4.1 labels; olmOCR-synthmix, Claude labels), ChartNet's other subsets (Mistral Research Licence), and every non-commercial dataset.

Limitations

  • Charts are the weak spot, and one training attempt made them worse. A larger 65,000-sample fine-tune with 26,000 more ChartNet charts lowered chart reading by 2.25 points (standard error 1.03) against this model's SFT step: synthetic plotting-library charts did not transfer to the report charts ParseBench uses. That checkpoint was not released. What fixed the chart scores was not that data: the server's nested-chart analysis (+5.34 on the base model), the 20,000-sample fine-tune (+4.57 on the base, about half of it because the model stopped wrapping single charts in multi-chart containers) and GRPO with a chart-rule reward (+0.43 more, within noise). 85 of the 568 chart pages (15%) still score zero (the base model: 154, or 111 with the nested-chart analysis).
  • The model's own table reading did not improve (77.12 against the base model's 77.30 on our server, within noise). What lifts the table scores is the Markdown around the model: the table-header rule takes our server to 83.67 and KDL's pipeline reads 85.89, both by marking header rows; the structure score (GriTS-Con) is between 91.1 and 91.5 in all three.
  • Visual grounding did not move on our server (72.78 against 72.83); through KDL's pipeline it is 75.21 against the base model's 74.65. The model learned part of DocLayNet's layout taxonomy and now uses MinerU's ref_text, aside_text and code labels far less often (on ParseBench's pages 76, 147 and 0 blocks, against the base model's 459, 379 and 13). ParseBench maps them to Text, so its score is unaffected, but other users of the block types may notice.
  • GRPO's own effect is within noise (+0.17 overall, 95% interval −0.20 to +0.52). Most of the gain over the base model comes from the supervised step and the server.
  • Underline, strikethrough and highlight are still rarely written. The gains in formatting are bold, italic, superscript and subscript.
  • Tuned for one benchmark. Data was weighted toward ParseBench's weak dimensions, and checkpoints were selected on ParseBench scores; expect smaller gains on other document types, scans, handwriting and languages other than English. MinerU2.5-Pro's own OmniDocBench results were not re-measured for this model.
  • Our page server does not reproduce 76.77. That number needs KDL's pipeline; our server, which also returns typed blocks, scores 75.15.
  • It inherits the base model's limits: one page at a time, no cross-page table merging, and the page image must be rendered first.

Intended use

Research and evaluation of document parsing, and as a starting point for further fine-tuning. It can be used in products under Apache-2.0, but check its output on your own documents first; we measured it only on ParseBench.

Licence and attribution

  • Weights and code in this repository: Apache-2.0 (see LICENSE and NOTICE).
  • Base model: MinerU2.5-Pro-2605-1.2B by OpenDataLab, Apache-2.0; please cite it:
  • Training data: see the table above; the CDLA-Permissive datasets are by IBM, FineWeb-2 by Hugging Face (ODC-By 1.0).
  • The 76.77 pipeline: ParseBench's kdl_frontier_nano provider, KoreaDeep's pipeline for KDL-Frontier-Parser-nano, contributed to ParseBench (Apache-2.0) in run-llama/ParseBench#49. Thanks to KoreaDeep for the pipeline and to LlamaIndex for ParseBench. No KDL weights (AGPL-3.0) are used in, or distributed with, this model.
@misc{wang2026mineru25propushinglimitsdatacentric,
      title={MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale},
      author={Bin, Wang and Tianyao, He and Linke, Ouyang and Fan, Wu and Zhiyuan, Zhao and Tao, Chu and Yuan, Qu and Zhenjiang, Jin and Weijun, Zeng and Ziyang, Miao and Bangrui, Xu and Junbo, Niu and others},
      year={2026},
      eprint={2604.04771},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2604.04771},
}
Downloads last month
29
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aeddix-labs/aeddix-alpine-ocr

Finetuned
(2)
this model

Datasets used to train aeddix-labs/aeddix-alpine-ocr

Paper for aeddix-labs/aeddix-alpine-ocr

Evaluation results

  • llamaindex/ParseBench leaderboard
  • Mean View evaluation results
    source
    Pipeline name: kdl_frontier_nano, ParseBench's in-repo provider (KoreaDeep's pipeline for KDL-Frontier-Parser-nano, run-llama/ParseBench PR #49), unchanged, with these weights (tag v0.1, commit 7fa09b7) served by vLLM 0.28.0 under the provider's default model name; proposed to ParseBench as pipeline aeddix_alpine_ocr_kdl. Mean of 3 full runs (76.76, 76.79, 76.75; SD 0.02), all 2,078 pages, 0 failures: run 1 on one NVIDIA L4 with ParseBench afb36bd, runs 2 and 3 on one NVIDIA A10 24 GB with ParseBench cbd57f2 (re-scoring run 1's 1,741 kept page results with cbd57f2 changes no document's score). Provider defaults: 144 dpi, 20 concurrent documents, LLM normalization off. Unweighted mean of the five headline metrics. Same pipeline on the base model MinerU2.5-Pro-2605-1.2B: 74.29. Our own page server (inference/ in this repository, mineru2605pro API, three Markdown rules) scores 75.15 on ParseBench afb36bd.
    76.77 *
  • Table View evaluation results
    source
    kdl_frontier_nano; grits_trm_composite, 503 documents; mean of 3 runs (86.03, 85.73, 85.92); base on the same pipeline 84.72; our page server 83.67
    85.89 *
  • Chart View evaluation results
    source
    kdl_frontier_nano; rule_pass_rate, 568 documents; mean of 3 runs (68.70, 68.71, 68.92); base on the same pipeline 69.15; our page server 65.28
    68.78 *