Qota-OCR

A text-line recognizer for documents where Korean is mixed with Chinese, Japanese and English on the same line, such as ํ•œ๊ตญ์–ด(ไธญๆ–‡) or ๆ—ฅๆœฌ่ชž(ํ•œ๊ตญ์–ด). It also reads Korean vertical text. It was trained from PP-OCRv6_small_rec with Korean added.

Highlights

  • One model reads Korean, Japanese, English and Chinese. Character F1 on SynthDoG is 0.96โ€“0.99 in all four languages.
  • It also reads scanned Korean public documents, and Korean vertical text with word spacing.
  • The weights come as safetensors for transformers (PyTorch) and as ONNX in float32, FP16, INT8 and UINT8, which run with onnxruntime alone.

Model details

  • Task: text-line recognition with CTC decoding.
  • Architecture: the same as PP-OCRv6_small_rec. PP-LCNetV4 backbone, SVTR encoder (2 blocks, width 120) and CTC head.
  • Output: 30,244 characters plus space. That is the 18,708 characters of the PP-OCRv6 dictionary plus 11,536 characters (Hangul and others) that only the korean_PP-OCRv5_mobile_rec dictionary has.
  • Parameters: 6.7M. model.safetensors is 26.8 MB and onnx/model.onnx 26.7 MB (both float32). The FP16, INT8 and UINT8 ONNX are 13.4โ€“15.4 MB.
  • Input: one cropped line, resized to height 48 with its aspect ratio kept (width up to 3,200).

Intended use

  • Printed documents where Korean is mixed with Chinese, Japanese or English: public documents, forms, books and reports, including Korean vertical text.
  • It reads line crops, so use it after a text detector.
  • Not measured: receipts, photos, handwriting and long lines. It is weak on math symbols and Greek letters, so it is not a good fit for formulas.

Performance

All numbers were measured directly. Every model except Nemotron OCR v2 used the same detector, PP-OCRv6_small_det, and only the recognizer was changed. Nemotron OCR v2 ran its own detection and recognition.

SynthDoG โ€” Korean, Japanese, English, Chinese

Character F1, order-independent, page_avg (higher = better).

Model Korean Japanese English Chinese
Qota-OCR 0.976 0.977 0.991 0.964
Nemotron OCR v2 (multilingual) 0.945 0.959 0.953 0.943
PP-OCRv6_small_rec 0.107 0.971 0.992 0.973
PP-OCRv5_server_rec 0.118 0.926 0.986 0.960
korean_PP-OCRv5_mobile_rec 0.962 0.117 0.989 0.179

OmniDocBench โ€” test split

NED, sample_avg (lower = better). Block counts are in parentheses.

Model English (7,837) Chinese (9,374) Englishโ€“Chinese mixed (1,458) Traditional Chinese (83)
Qota-OCR 0.025 0.048 0.068 0.091
PP-OCRv6_small_rec 0.022 0.040 0.035 0.034
PP-OCRv5_server_rec 0.028 0.045 0.058 0.055
korean_PP-OCRv5_mobile_rec 0.031 0.803 0.376 0.885

Usage

transformers (PyTorch)

pip install "transformers>=5.17" torch torchvision pillow opencv-python-headless huggingface_hub
import torch
from huggingface_hub import hf_hub_download
from PIL import Image
from transformers import AutoImageProcessor, AutoModelForObjectDetection, AutoModelForTextRecognition

det_proc = AutoImageProcessor.from_pretrained("PaddlePaddle/PP-OCRv6_small_det_safetensors")  # text-line detector
det = AutoModelForObjectDetection.from_pretrained("PaddlePaddle/PP-OCRv6_small_det_safetensors").eval()
proc = AutoImageProcessor.from_pretrained("atonlee/Qota-OCR")
model = AutoModelForTextRecognition.from_pretrained("atonlee/Qota-OCR").eval()

page = Image.open(hf_hub_download("atonlee/Qota-OCR", "examples/document.png")).convert("RGB")
with torch.no_grad():
    inputs = det_proc(images=page, return_tensors="pt")
    boxes = det_proc.post_process_object_detection(det(pixel_values=inputs["pixel_values"]), target_sizes=inputs["target_sizes"],
                                                   threshold=0.2, box_threshold=0.45, unclip_ratio=1.4)[0]["boxes"].tolist()
    for box in sorted(boxes, key=lambda b: (round(b[0][1] / 10), b[0][0])):  # 4 corners per line, top to bottom
        xs, ys = [x for x, _ in box], [y for _, y in box]
        line = page.crop((min(xs), min(ys), max(xs), max(ys)))
        if line.height >= 1.5 * line.width:  # vertical text
            line = line.rotate(90, expand=True)
        print(proc.post_process_text_recognition(model(**proc(images=line, return_tensors="pt")))[0]["text"])
์‹œ๋ฆฝ ๋ฌธํ™”์žฌ๋‹จ ๊ณต๊ณ  ์ œ2026-37ํ˜ธ
ใ€Œ์ „ํ†ต ํ•œ์ง€(้Ÿ“็ด™) ๊ณต์˜ˆ ํŠน๋ณ„์ „ใ€ ๊ฐœ์ตœ๋ฅผ ์•„๋ž˜์™€ ๊ฐ™์ด ์•Œ๋ฆฝ๋‹ˆ๋‹ค.
1. ๊ธฐ๊ฐ„: 2026๋…„ 10์›” 12์ผ(์›”)๋ถ€ํ„ฐ 10์›” 25์ผ(์ผ)๊นŒ์ง€, 10:00-18:00
2. ์žฅ์†Œ: ์‹œ๋ฆฝ๋ฏธ์ˆ ๊ด€ 2์ธต ์ œ3์ „์‹œ์‹ค (Gallery 3, 2F)
3. ๊ด€๋žŒ๋ฃŒ: ์„ฑ์ธ 3,000์›, ์ฒญ์†Œ๋…„ 1,500์›, 65์„ธ ์ด์ƒ ๋ฌด๋ฃŒ
4. ํ•ด์„ค: ๋งค์ผ 14:00 ํ•œ๊ตญ์–ดยทEnglish, 16:00 ๆ—ฅๆœฌ่ชžยทไธญๆ–‡
ๅฑ•็คบใฎ่ฉณใ—ใ„ใ”ๆกˆๅ†…ใฏๅ—ไป˜ใงใŠ้…ใ‚Šใ—ใฆใ„ใพใ™ใ€‚
่ฏทๅœจๅ…ฅๅฃๅค„้ข†ๅ–ไธญๆ–‡ๅฏผ่งˆๆ‰‹ๅ†Œใ€‚
5. ๋ฌธ์˜: ์ „์‹œํŒ€ 02-000-0000, exhibit@example.org
2026๋…„ 9์›” 29์ผ
์‹œ๋ฆฝ ๋ฌธํ™”์žฌ๋‹จ ์ด์‚ฌ์žฅ

ONNX

ONNX models are included in onnx/.

File Size Same text as float32 (29 lines)
onnx/model.onnx (float32) 26.7 MB 29
onnx/model_fp16.onnx 13.4 MB 28
onnx/model_int8.onnx 15.3 MB 27
onnx/model_uint8.onnx 15.4 MB 26
  • FP16 keeps weights and computation in FP16. INT8 and UINT8 quantize the MatMul layers that have weights and keep the convolutions in float32.
  • Compared on 29 lines, not on the full test sets above.
  • Input x is one cropped line in BGR, resized to height 48 with its aspect ratio kept (width 320โ€“3,200, padded with 0) and scaled to [-1, 1]. The output is per-position character probabilities; decode them with CTC using character_list in preprocessor_config.json.

Input and behavior

  • Rotate tall lines (height/width โ‰ฅ 1.5) 90 degrees counterclockwise before recognition. This is how vertical text is read.
  • Lines longer than 25 characters are read too, but were not measured separately.

Data attribution

This model used datasets from The Open AI Dataset Project (AI-Hub, S. Korea). All data information can be accessed through AI-Hub (www.aihub.or.kr).

License

Apache-2.0

Downloads last month
75
Safetensors
Model size
6.69M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support