DocScanner on-device OCR models

ONNX exports of PaddleOCR text detection and recognition models, used by DocScanner to transcribe handwritten pages entirely on the device.

DocScanner is a free, open-source tool for Bible translation field teams. Transcription has to work without a network and without per-page cost, so no cloud model is involved: a page is cut into text lines by the detector, and each line is read by the recognizer for that project's writing system.

Layout

A device downloads the detector plus exactly one recognizer β€” 13 MB for most scripts, 78 MB for Latin/CJK β€” never the whole set.

path file size sha256 (short)
detector/ det_model.onnx 4.6 MB 0c5eeee2
ppocrv6/ rec_model.onnx 73 MB 4078550d
ppocrv6/ charset.json 128 KB 46f80089
devanagari/ rec_model.onnx 7.6 MB a3d5b5fa
devanagari/ charset.json 3.6 KB 621d0073
thai/ rec_model.onnx 7.5 MB d52231fe
thai/ charset.json 3.3 KB b62ede69
arabic/ rec_model.onnx 7.6 MB a1e69c68
arabic/ charset.json 4.5 KB df730929
korean/ rec_model.onnx 13 MB 03525a1d
korean/ charset.json 81 KB e600744e

The app pins a tag of this repository in its download URL, so a given build can only ever fetch the weights it was tested against.

Provenance

Exported with paddle2onnx from the official PaddleOCR inference models, unquantized (float32):

directory source model opset
detector/ PP-OCRv5_mobile_det 17
ppocrv6/ PP-OCRv6_medium_rec 16
devanagari/ devanagari_PP-OCRv5_mobile_rec 16
thai/ th_PP-OCRv5_mobile_rec 16
arabic/ arabic_PP-OCRv5_mobile_rec 16
korean/ korean_PP-OCRv5_mobile_rec 16

charset.json is each model's character table from PaddleOCR, as a JSON array. The recognizer output has charset + 2 classes: index 0 is the CTC blank, then the table, then a space β€” the layout PaddleOCR's CTCLabelDecode expects.

Interfaces

Detector (det_model.onnx) β€” input [1, 3, H, W], BGR, long side scaled to 960 and floored to a multiple of 32, normalized with ImageNet statistics (mean 0.485/0.456/0.406, std 0.229/0.224/0.225). Output [1, 1, H, W], a per-pixel probability of text.

Recognizers (rec_model.onnx) β€” input [1, 3, 48, W], RGB, height 48 with width scaled by the line's aspect ratio and padded to a multiple of 32, normalized (x/255 - 0.5) / 0.5. Padding is left at zero after normalization (mid-grey), which is how PaddleOCR pads; padding with black instead wrecks recognition. Output [1, T, classes] logits, greedy CTC decoded.

Arabic is read right-to-left: CTC scans left to right and so emits the logically last character first. The line is reversed by grapheme cluster, keeping combining marks attached to their base letter β€” correcting this took the error on rendered Arabic from 79% to 18%.

Measured accuracy

Character error rate on handwritten sample pages, per line, after detection:

script CER note
English (neat) 5-7%
English (hard hand) ~14%
Spanish / French ~4%
Chinese 2% written in a squared-paper grid
Devanagari usable whole lines, most words legible
Thai 15-25%
Arabic 25-35% words land in the right places and reading order
Korean 15-25% several lines nearly verbatim

These are CTC models with no language model, which is deliberate: an unreadable crop comes back garbled or empty rather than as fluent invented text. A TrOCR alternative scored better on neat English (2%) but rewrote what it could not read β€” including inventing liturgical Church Slavonic from a blank strip β€” which is the wrong failure mode for a translation tool.

Not included

  • Cyrillic β€” PaddleOCR's Cyrillic recognizers manage only 76-84% CER on handwriting, so those projects use server transcription until a model trained on handwritten lines exists.
  • Hebrew β€” no PaddleOCR model.
  • Greek β€” not yet wired.

License

Apache-2.0, inherited from PaddleOCR. Please keep the attribution to the PaddleOCR project.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support