Instructions to use dhchoi/bandit-ocr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dhchoi/bandit-ocr with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="dhchoi/bandit-ocr", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("dhchoi/bandit-ocr", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dhchoi/bandit-ocr with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dhchoi/bandit-ocr" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dhchoi/bandit-ocr", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/dhchoi/bandit-ocr
- SGLang
How to use dhchoi/bandit-ocr with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dhchoi/bandit-ocr" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dhchoi/bandit-ocr", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dhchoi/bandit-ocr" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dhchoi/bandit-ocr", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use dhchoi/bandit-ocr with Docker Model Runner:
docker model run hf.co/dhchoi/bandit-ocr
Bandit OCR
Built with dots.mocr.
Bandit OCR reads a page of a Korean historical document and returns one region per printed column, in reading order, with the characters of that column. It is a LoRA fine-tune of dots.mocr, a 1.7B-parameter document vision-language model, trained on 229,356 annotated pages of Joseon-dynasty woodblock prints and manuscripts.
On the held-out test split it reads a page at 0.0195 character error rate with a column F1 of 0.9856, against 0.1436 for the AI Hub ResNet pipeline and 0.2398 for NDLkotenOCR on the same pages.
The repository holds both forms of the same fine-tune. At the root are the merged weights, the
adapter already folded in at scale 0.75, which load and serve exactly like the base and need
no PEFT. Under lora/ is the LoRA adapter itself, for composing onto the base yourself,
changing the scale or training further. Take the root unless you know you want the adapter.
- Desktop application and project site: https://bandit.dhchoi.net
- Paper: JADH 2026, September 2026
Model details
| Base model | dots-studio/dots.mocr at revision e539fbb52280393adc081b289ec597430a0f9031 |
| Parameters | 1.7B (1.2B language decoder + 0.4B vision tower) |
| Adaptation | LoRA, rank 64, alpha 128, dropout 0.05, on the language projections and the vision tower |
| Weights | root: the adapter folded in at scale 0.75, a plain bfloat16 checkpoint, no PEFT needed |
| Adapter | lora/: the same LoRA unmerged, 580 MB, to apply at scale 0.75 yourself |
| Files | 2 safetensors shards at the root, about 6.1 GB, plus the adapter |
| Checkpoint | t3_full step 56000 of the full training run |
| Trained on | 229,356 annotated pages of Korean historical documents |
| Input | one page image, 3,136 to 11,289,600 pixels after smart_resize |
| Output | JSON array of column regions in reading order |
| Author | Donghyeok Choi, Department of History, Hong Kong Baptist University |
| Project | https://bandit.dhchoi.net |
What it returns
One JSON array, one object per printed column, in reading order:
[
{"bbox": [2205, 354, 2350, 3406], "category": "Text", "text": "ๆๆท้่ผ
ๅญซ่ๅฎๆฒๅญ็้ซๅๅผต่ณๅๅธๅ
้้ๅฐๆฌ้ปๅฑ
ๆญฃ"},
{"bbox": [2057, 366, 2193, 3430], "category": "Text", "text": "ไปปๅฝฆๅฟ ๅผตๆ้้ๆฑ็ฟผ็ญๅคงๅฐ่ฃๅๅ้่ฏ่่็ญๅฅๅๅฏถ่ฉฃ"}
]
Those are the first two of the fifteen columns of one Sillok page from the test split, 2400 x 3744 pixels, on its 2408 x 3752 input grid.
Reading order is the convention of the printed page: column slots right to left, each slot top to bottom, and the two sub-columns of an interlinear note (์ธ์ฃผ) right to left. An interlinear note is a column region like any other, at about half width, with no markup in its text.
bbox is [x1, y1, x2, y2] in the resized input grid, not in original pixels. Map it back with
the same smart_resize the processor used:
from qwen_vl_utils.vision_process import smart_resize
grid_h, grid_w = smart_resize(height, width, min_pixels=3136, max_pixels=11289600)
x1 = round(x1 * width / grid_w)
y1 = round(y1 * height / grid_h)
Usage
Transformers
The model ships the base's remote code, so trust_remote_code=True is required. It was trained and
evaluated against transformers==4.51.3; 4.52 and later change the Qwen2.5-VL processor contract
the custom DotsVLProcessor was written for.
import json
import torch
from PIL import Image
from qwen_vl_utils import process_vision_info
from qwen_vl_utils.vision_process import smart_resize
from transformers import AutoModelForCausalLM, AutoProcessor
MODEL = "dhchoi/bandit-ocr"
PROMPT = (
"Please output the layout information from the PDF image, including each layout element's "
"bbox, its category, and the corresponding text content within the bbox.\n\n"
"1. Bbox format: [x1, y1, x2, y2]\n\n"
"2. Layout Categories: The possible categories are ['Caption', 'Footnote', 'Formula', "
"'List-item', 'Page-footer', 'Page-header', 'Picture', 'Section-header', 'Table', 'Text', "
"'Title'].\n\n"
"3. Text Extraction & Formatting Rules:\n"
" - Picture: For the 'Picture' category, the text field should be omitted.\n"
" - Formula: Format its text as LaTeX.\n"
" - Table: Format its text as HTML.\n"
" - All Others (Text, Title, etc.): Format their text as Markdown.\n\n"
"4. Constraints:\n"
" - The output text must be the original text from the image, with no translation.\n"
" - All layout elements must be sorted according to human reading order.\n\n"
"5. Final Output: The entire output must be a single JSON object.\n"
)
MIN_PIXELS, MAX_PIXELS = 3136, 11289600
model = AutoModelForCausalLM.from_pretrained(
MODEL, trust_remote_code=True, torch_dtype=torch.bfloat16, device_map="auto"
)
processor = AutoProcessor.from_pretrained(MODEL, trust_remote_code=True)
path = "page.jpg"
with Image.open(path) as im:
width, height = im.size
grid_h, grid_w = smart_resize(height, width, min_pixels=MIN_PIXELS, max_pixels=MAX_PIXELS)
messages = [{"role": "user", "content": [
{"type": "image", "image": path, "min_pixels": MIN_PIXELS, "max_pixels": MAX_PIXELS},
{"type": "text", "text": PROMPT},
]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
images, videos = process_vision_info(messages)
images = [im.resize((grid_w, grid_h), Image.Resampling.LANCZOS) for im in images]
inputs = processor(text=[text], images=images, videos=videos, padding=True,
return_tensors="pt").to(model.device)
with torch.inference_mode():
out = model.generate(
**inputs,
max_new_tokens=8192,
do_sample=False,
eos_token_id=processor.tokenizer.convert_tokens_to_ids("<|endofassistant|>"),
suppress_tokens=[151643, 151672],
)
answer = processor.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0]
regions = json.loads(answer)
vLLM
vLLM 0.21.0 serves the merged weights with no conversion; it resolves DotsOCRForCausalLM from the
remote code itself.
vllm serve dhchoi/bandit-ocr \
--trust-remote-code --dtype bfloat16 --max-model-len 24576 \
--limit-mm-per-prompt '{"image": 1}' \
--mm-processor-kwargs '{"max_pixels": 11289600, "min_pixels": 3136}'
Send the image as a data URL with the prompt above, temperature 0, max_tokens 8192, and the two
decode settings of the section below: "stop_token_ids": [151673] and
"logit_bias": {"151643": -100, "151672": -100}.
The LoRA adapter
lora/ holds the same fine-tune unmerged, for composing onto the base yourself, changing
the scale or training further. The delta is scaled by 0.75 before it is applied: that scale is
part of the released model, not a training artefact, and every number below was measured with it.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained(
"dots-studio/dots.mocr", revision="e539fbb52280393adc081b289ec597430a0f9031",
trust_remote_code=True, torch_dtype=torch.bfloat16, device_map="auto",
)
model = PeftModel.from_pretrained(base, "dhchoi/bandit-ocr", subfolder="lora")
for module in model.modules(): # every LoRA layer scales its delta by scaling[name]
scaling = getattr(module, "scaling", None)
if isinstance(scaling, dict):
for name in scaling:
scaling[name] *= 0.75
model = model.merge_and_unload()
That reproduces the weights at the root of this repository. Prompt, decode and rescale boxes the same way, and set the stop id yourself, because the composed model inherits the base's generation config.
Applied at its trained scale of 1.0 the adapter is a different, measurably worse model: on the same 540-page validation set an earlier checkpoint of this run read at 0.0656 page CER at scale 1.0 against 0.0376 at scale 0.75, both under the stop rule of the day.
Decoding: stop on one token, suppress two
This matters more than any other setting here. The training target ends with <|endofassistant|>
(151673) and nothing else, but the base generation config also treats <|endoftext|> (151643) and
<|assistant|> (151672) as stop ids. The fine-tune sometimes emits <|endoftext|> in the middle of
a column at a hard glyph, and a page that stops there is truncated or fails to parse.
generation_config.jsonat the root of this repository already setseos_token_idto 151673 alone. If you compose the adapter onto the base yourself, set it there: the base's config stops on all three.- Suppress 151643 and 151672 during decoding as well (
suppress_tokensingenerate, a -100logit_biasthrough an OpenAI-compatible server).
Measured on 540 validation pages, this rule takes the released checkpoint from 0.0517 page CER with 2 pages failing to parse down to 0.0271 with none, weights unchanged. On the 16 pages of a diagnostic cohort that stopped early at a neighbouring checkpoint, micro CER fell from 0.652 to 0.027 while six control pages stayed put. Every number reported below was measured under this rule.
Evaluation
The whole held-out test split of the training corpus, 14,376 pages read through vLLM with the decode rule above. Regions are matched one to one at IoU 0.5; CER is over matched regions (region CER) and over the concatenated reading-order text of the page (page CER). The layout row scores every test page that has ground-truth geometry; the text row scores the pages that also carry ground-truth text.
| Pages | P | R | F1 | mean IoU | region CER | page CER | |
|---|---|---|---|---|---|---|---|
| layout ground truth | 13,636 | 0.9897 | 0.9805 | 0.9851 | 0.9580 | 0.0256 | 0.0235 |
| text ground truth | 12,935 | 0.9909 | 0.9804 | 0.9856 | 0.9583 | 0.0217 | 0.0195 |
By script style, on the text ground truth:
| Pages | P | R | F1 | mean IoU | region CER | page CER | |
|---|---|---|---|---|---|---|---|
| haeseo ํด์ (standard) | 5,573 | 0.9894 | 0.9713 | 0.9803 | 0.9488 | 0.0309 | 0.0252 |
| haengseo ํ์ (semi-cursive) | 1,558 | 0.9692 | 0.9779 | 0.9735 | 0.9447 | 0.0484 | 0.0634 |
| choseo ์ด์ (cursive) | 88 | 0.9555 | 0.8274 | 0.8869 | 0.9112 | 0.2272 | 0.2063 |
| yeseo ์์ (clerical) | 15 | 0.9508 | 1.0000 | 0.9748 | 0.9475 | 0.0412 | 0.0648 |
| sillok (print, no style label) | 5,701 | 0.9966 | 0.9896 | 0.9931 | 0.9681 | 0.0125 | 0.0101 |
By collection:
| Pages | P | R | F1 | mean IoU | region CER | page CER | |
|---|---|---|---|---|---|---|---|
| AI Hub 234 (๊ณ ์ ํ์ ์ธ์) | 2,578 | 0.9738 | 0.9801 | 0.9769 | 0.9485 | 0.0312 | 0.0404 |
| AI Hub 603 (๊ณ ์ ํ์ ์ธ์ OCR) | 2,391 | 0.9874 | 0.9533 | 0.9700 | 0.9414 | 0.0428 | 0.0331 |
| AI Hub 71294 (OCR ๋ฐ์ดํฐ ๊ณ ์ํ์) | 2,265 | 0.9934 | 0.9810 | 0.9872 | 0.9536 | 0.0374 | 0.0335 |
| Sillok ์กฐ์ ์์กฐ์ค๋ก | 5,701 | 0.9966 | 0.9896 | 0.9931 | 0.9681 | 0.0125 | 0.0101 |
Against the baselines, same pages, same ground truth
| System | layout F1 | page CER |
|---|---|---|
| Bandit OCR (this model) | 0.9851 | 0.0195 |
| AI Hub ResNet pipeline | 0.9595 | 0.1436 |
| NDLkotenOCR | 0.8880 | 0.2398 |
| dots.ocr, base | 0.1602 | 0.1497 |
The base dots.ocr layout F1 is a unit mismatch rather than a reading failure: the base answers with about 3 regions where the ground truth has 15, so its page CER is the comparable figure. A base dots.mocr row is not measured yet.
Out of domain
1,000 pages of Chinese historical print (HisDoc1B), a tradition the model never saw:
| Pages | P | R | F1 | mean IoU | region CER | page CER | |
|---|---|---|---|---|---|---|---|
| HisDoc1B | 1,000 | 0.6773 | 0.9512 | 0.7912 | 0.9165 | 0.2144 | 0.2241 |
Recall 0.95 says the columns are found and page CER 0.22 is the transfer number. The precision is over-segmentation: the model returns 28,983 regions against the ground truth's 20,638, and a quarter of its regions hold five characters or fewer.
Training data
229,356 pages and 3,023,103 column regions, split off a corpus of four Korean collections by book volume so that no volume straddles train and test (seed 20260828, 90/5/5):
- ์กฐ์ ์์กฐ์ค๋ก Veritable Records of the Joseon Dynasty, National Institute of Korean History: 99,764 training pages. Column boxes come from a character detector run over each page, and the characters of each column come from aligning the article's transcription against that detection.
- AI Hub ๊ณ ์ ํ์ ์ธ์ datasets 234, 603 and 71294: 129,592 training pages of woodblock prints and manuscripts in five script styles, with character-level boxes grouped into columns.
The corpus itself is not published and cannot be: the AI Hub data may not be redistributed and the Sillok images are copyrighted by the National Institute of Korean History. Only the weights are released.
The text is the printed form: Hanja, the โ article marker, โก for a glyph with no code point,
and note text. Editorial punctuation, bracket marks and collation insertions are removed.
Training procedure
One run on H200 GPUs: bfloat16, gradient checkpointing, effective batch 8 (per-device batch 1 with
gradient accumulation), AdamW, cosine schedule, warmup ratio 0.03, weight decay 0.01, seed 3407,
flash_attention_2, one page per sample at up to 11,289,600 pixels, batches grouped by length. The learning rate was 2e-4 for the first 32,000 steps and 5e-5 from there, on a
cosine schedule spanning the remaining steps; step 56000 is about 1.95 epochs over the
229,356-page training split.
Checkpoints were evaluated every 4,000 steps by generating a 540-page validation set and scoring it the way the test split is scored, and the released checkpoint is the best of that curve by page CER. The curve is flat from step 44,000 on, within 0.0036 CER of the selected point.
The LoRA delta is applied at scale 0.75. That is not a training setting: scaling the merged delta below 1.0 was measured to read better than the trained scale, and it is folded into the merged weights because a server applies no runtime scale.
Limitations
- Cursive (์ด์) is the weak style: page CER 0.2063 against 0.0252 on standard script, and the lowest recall in the corpus (0.8274). The test split holds only 88 cursive pages, so that number is itself uncertain.
- The clerical (์์) row is 15 pages. Read it as an indication, not a score.
- The Sillok layout score is not independent. The v3 Sillok column boxes were drawn by HRCenterNet, the AI Hub baseline's own detector, and this model was trained on them. A 100-page human-corrected gold set confirms the column inventory (304 boxes drawn blind against 304), so recall and F1 on Sillok are meaningful, but the mean IoU on Sillok pages carries that circularity.
- Out of domain it over-segments. On Chinese material it splits small text, headings and notes into several regions where the ground truth keeps one column, which costs precision, not reading.
- It is a column reader, not a general document parser. The prompt is the base's full layout
prompt, but the fine-tune answers with
Textcolumn regions; tables, formulas and figures are not part of the training target and the base's behaviour on them has been trained away. - No safety or content filtering of any kind. It transcribes what is on the page.
A note on the remote code
The custom modelling files are the base's, with four idempotent patches this project needs on its hosts, all of them to how attention is chosen and imported, none of them to the architecture or the weights:
- the flash-attn import is wrapped in a try/except, so the file imports on a host without it;
DotsVisionTransformerdeclares_supports_sdpaand_supports_flash_attn, which transformers checks before it will give a sub-model sdpa;- the sdpa path skips its all-true attention mask when the batch holds a single image, which is exact and avoids a mask that costs about 95 GB on a full page;
DotsVLProcessorpassesvideo_processorto the Qwen2.5-VL processor by keyword, which transformers 4.51 otherwise reads aschat_template.
config.json asks the vision tower for flash_attention_2. Without flash-attn installed it falls
back to eager attention, which is quadratic in the number of patches; set
vision_config.attn_implementation to sdpa in that case.
Licence and attribution
Built with dots.mocr.
This model is a fine-tune of dots.mocr, copyright
Xingyin Information Technology (Shanghai) Co., Ltd, used under the dots.mocr LICENSE AGREEMENT. A
copy of that agreement and the base NOTICE ship in this repository (clauses 7.1 and 7.2), the
notices in them are unchanged, and this card displays the attribution statement clause 7.3
requires. Nothing here implies endorsement by the licensor, and no name in this project is a
licensor trade name (clause 4.4). The weights are released on the terms the base is released
under; where the agreement and the metadata tag differ, the agreement is the document that
ships.
The training data is licensed separately and is not redistributed; see the section above.
Citation
@inproceedings{choi2026bandit,
title = {Bandit: A Human-in-the-Loop OCR System for East Asian Historical Documents
with Vision-Language Model Fine-tuning},
author = {Choi, Donghyeok},
booktitle = {Japanese Association for Digital Humanities (JADH) 2026},
year = {2026}
}
- Downloads last month
- 33
Model tree for dhchoi/bandit-ocr
Base model
dots-studio/dots.mocr