DocScanner on-device OCR models
ONNX exports of PaddleOCR text detection and recognition models, used by DocScanner to transcribe handwritten pages entirely on the device.
DocScanner is a free, open-source tool for Bible translation field teams. Transcription has to work without a network and without per-page cost, so no cloud model is involved: a page is cut into text lines by the detector, and each line is read by the recognizer for that project's writing system.
Layout
A device downloads the detector plus exactly one recognizer β 13 MB for most scripts, 78 MB for Latin/CJK β never the whole set.
| path | file | size | sha256 (short) |
|---|---|---|---|
detector/ |
det_model.onnx |
4.6 MB | 0c5eeee2 |
ppocrv6/ |
rec_model.onnx |
73 MB | 4078550d |
ppocrv6/ |
charset.json |
128 KB | 46f80089 |
devanagari/ |
rec_model.onnx |
7.6 MB | a3d5b5fa |
devanagari/ |
charset.json |
3.6 KB | 621d0073 |
thai/ |
rec_model.onnx |
7.5 MB | d52231fe |
thai/ |
charset.json |
3.3 KB | b62ede69 |
arabic/ |
rec_model.onnx |
7.6 MB | a1e69c68 |
arabic/ |
charset.json |
4.5 KB | df730929 |
korean/ |
rec_model.onnx |
13 MB | 03525a1d |
korean/ |
charset.json |
81 KB | e600744e |
The app pins a tag of this repository in its download URL, so a given build can only ever fetch the weights it was tested against.
Provenance
Exported with paddle2onnx from the official
PaddleOCR inference models, unquantized (float32):
| directory | source model | opset |
|---|---|---|
detector/ |
PP-OCRv5_mobile_det |
17 |
ppocrv6/ |
PP-OCRv6_medium_rec |
16 |
devanagari/ |
devanagari_PP-OCRv5_mobile_rec |
16 |
thai/ |
th_PP-OCRv5_mobile_rec |
16 |
arabic/ |
arabic_PP-OCRv5_mobile_rec |
16 |
korean/ |
korean_PP-OCRv5_mobile_rec |
16 |
charset.json is each model's character table from PaddleOCR, as a JSON array. The recognizer
output has charset + 2 classes: index 0 is the CTC blank, then the table, then a space β
the layout PaddleOCR's CTCLabelDecode expects.
Interfaces
Detector (det_model.onnx) β input [1, 3, H, W], BGR, long side scaled to 960 and floored
to a multiple of 32, normalized with ImageNet statistics (mean 0.485/0.456/0.406, std
0.229/0.224/0.225). Output [1, 1, H, W], a per-pixel probability of text.
Recognizers (rec_model.onnx) β input [1, 3, 48, W], RGB, height 48 with width scaled by
the line's aspect ratio and padded to a multiple of 32, normalized (x/255 - 0.5) / 0.5.
Padding is left at zero after normalization (mid-grey), which is how PaddleOCR pads; padding
with black instead wrecks recognition. Output [1, T, classes] logits, greedy CTC decoded.
Arabic is read right-to-left: CTC scans left to right and so emits the logically last character first. The line is reversed by grapheme cluster, keeping combining marks attached to their base letter β correcting this took the error on rendered Arabic from 79% to 18%.
Measured accuracy
Character error rate on handwritten sample pages, per line, after detection:
| script | CER | note |
|---|---|---|
| English (neat) | 5-7% | |
| English (hard hand) | ~14% | |
| Spanish / French | ~4% | |
| Chinese | 2% | written in a squared-paper grid |
| Devanagari | usable | whole lines, most words legible |
| Thai | 15-25% | |
| Arabic | 25-35% | words land in the right places and reading order |
| Korean | 15-25% | several lines nearly verbatim |
These are CTC models with no language model, which is deliberate: an unreadable crop comes back garbled or empty rather than as fluent invented text. A TrOCR alternative scored better on neat English (2%) but rewrote what it could not read β including inventing liturgical Church Slavonic from a blank strip β which is the wrong failure mode for a translation tool.
Not included
- Cyrillic β PaddleOCR's Cyrillic recognizers manage only 76-84% CER on handwriting, so those projects use server transcription until a model trained on handwritten lines exists.
- Hebrew β no PaddleOCR model.
- Greek β not yet wired.
License
Apache-2.0, inherited from PaddleOCR. Please keep the attribution to the PaddleOCR project.