PP-OCRv6 Hanzi IDS
Predicts the IDS (Ideographic Description Sequence) of a CJK glyph image: a character
is described by how it is built out of parts, not by which code point it is. For 明 it
returns candidates such as ⿰日月, ⿰日冃, … ranked by likelihood.
- Scope: clean, single-colour, dark-on-light renderings of one character from a computer font — the rare-character images embedded in dictionary data.
- Accuracy: 90%+ on that input, excluding characters with no meaningful decomposition
— single-component glyphs such as
木, which the model tends to split anyway. See Limitations. - No code point needed. The answer is a structure, so unencoded characters, rare variants, regional forms and gaiji (外字) are covered even where the companion Unicode model would have to force them onto the nearest encoded neighbour.
汉字结构预测:输入一张只含一个字的图片,输出最多 5 个 IDS 结构候选,例如
明→⿰日月。面向电脑字体渲染的黑字白底单字图,不依赖码位,因此对未编码字、异体、 地域字形、外字同样适用。在上述输入条件下准确率 90% 以上,本身不可拆解的字 (单部件字,如「木」)不在此列。与 Unicode OCR 模型相互独立,可交叉核验。
| Model | Output |
|---|---|
| pp-ocrv6-hanzi-ids-ocr (this repo) | structure, e.g. ⿰日月 |
| pp-ocrv6-hanzi-unicode-ocr | the character, e.g. 明 |
The two are independent. Together they cross-check each other: IDS candidates are expanded through a dictionary of legal decomposition paths and compared with the structural path of the Unicode result, so a disagreement is surfaced for review.
Files
| File | Size | SHA-256 |
|---|---|---|
model.pt |
78,774,031 B | 8eff12b078a4e8748d0d7d158b22e9df477d9f9a103ca38c35fd3de9ee30b460 |
ensemble_1.pt |
78,787,930 B | 21de62d85772220e8c31ee82fe5d1fa867864e75ac4c94ebdcf35242d59891ac |
inference_config.json |
1,149 B | 1bb3196857102e4235f5267901ceb903580682807957f3fd8e5aa3fd45cc4dbb |
vocab.json |
27,477 B | 7f0d2c4be3a852bfcfb3574af1c5a521a929d83fa7140e1ed1f1af61a7112bda |
normalization.json |
3,349,393 B | 4a83fed8522a677fed0cc3f3aa2a732d5a57a2c49211d3e4a85c5ec4058aa2b0 |
component_output_policy.json |
8,868,270 B | 047e5f83a737309a68d702d1a2488c8a963d60733ec2f44cb58598029b557e21 |
data/ids_lv1.txt |
2,161,633 B | 1f1db7d24767d7622da109ce17f9a3acdb0b61d2b86a314e971f2f2d00a8dbf3 |
hanzi_ids/ |
— | model definition and inference code |
example_usage.py |
— | minimal end-to-end inference script |
Architecture
A glyph encoder plus a grammar-constrained autoregressive decoder.
| Component | Detail |
|---|---|
| Backbone | PPLCNetV4 medium, fused form, kept as a 2-D feature map (no vertical pooling) — the backbone family used by PaddleOCR's PP-OCRv6 recognition models |
| Projection | 1×1 convolution, 768 → 256 |
| Positional | learned row (64, 128) + column (128, 128) embeddings, concatenated to 256 |
| Decoder | 4 × TransformerDecoderLayer, d_model=256, 8 heads, FFN 1024, GELU, pre-norm, causal mask |
| Output head | weight-tied to the token embedding, plus a per-token bias |
| Vocabulary | 3,459 tokens · max length 96 tokens |
| Parameters | 19,665,475 per checkpoint, FP32 |
| Checkpoints | 2, averaged at the probability level during beam search |
Vocabulary
vocab.json is a flat list; index 0 is <PAD>, 1 is <BOS>, 2 is <EOS>.
| Kind | Count | Examples |
|---|---|---|
| Special | 3 | <PAD>, <BOS>, <EOS> |
| Layout operators | 138 | ⿰, ⿱, ⿲, parameterised ⿻[b_], annotated {冃}⿵ |
| Component classes | 136 | #(-⺄𠃊), #(HN), #(HNg) |
| Annotated components | 57 | {?0㮣}⿱, {卄}⿻ |
| Leaf components | 3,176 | 一, 木, ⺀, ㇁, コ, 㐀 |
A token's arity — how many children it consumes — follows the IDS operator: ⿲/⿳ take
three, the remaining ⿰⿱⿴⿵⿶⿷⿸⿹⿺⿻ pairs take two, / take one, leaves none.
#(...) stroke programs and {...} annotations are atomic tokens carrying the full
parameter text, so a prediction round-trips losslessly.
Decoding
The decoder is never allowed to emit a syntactically invalid sequence: at every step the
logits are masked by the number of unfilled slots — <PAD>, <BOS> and any token that
would overflow the budget are forbidden, <EOS> is allowed only when zero slots remain,
and a token is allowed only if its arity fits the open slots.
Beam search runs at beam_size = 5 over the ensemble; if fewer than five distinct legal
outputs are found, the image is re-decoded at beam_size = 12. Candidates are ranked by
sequence log probability without length normalisation, matching the training objective.
If still fewer than five distinct valid sequences exist, the list is returned short —
nothing is copied or invented to pad it.
Raw beam output then passes through the canonical output policy
(component_output_policy.json): a subtree is rewritten to a more frequent component
spelling only when that spelling is at least 4× more frequent in training (operators, child
order and {...} annotations are never rewritten, and each rewrite is verified to preserve
the normalised structure); and raw sequences normalising to the same IDS string are merged
into one candidate scored by the log-sum-exp of both, their raw forms kept in
equivalent_raw_ids.
Scores are uncalibrated: canonical_log_mass is the ranking score,
mean_token_probability the geometric-mean token probability of the representative raw
path. Use them for ranking, not as accuracy estimates.
Preprocessing
One glyph per image, at 128×128. hanzi_ids.render_utils.normalize_image implements this:
- Composite RGBA/LA/palette onto white, convert to 8-bit greyscale.
- If the median of the border pixels is darker than 127, invert — dark-on-light and light-on-dark inputs normalise to the same thing.
- Crop to the ink bounding box (luminance < 200). Fewer than three ink pixels raises
BadGlyph. - Resize with LANCZOS to fit a 113×113 box (aspect preserved), paste centred on a 128×128 white canvas.
- The runner scales to
[-1, 1](/127.5 − 1) and replicates greyscale to three channels. Model input is(N, 3, 128, 128)float32.
Usage
A custom architecture means the model needs its class definition, bundled here as the
hanzi_ids package:
pip install torch pillow numpy cairosvg
python example_usage.py path/to/glyph.png
from PIL import Image
from hanzi_ids import Runner
runner = Runner(".", device="cpu")
candidates = runner.predict([Image.open("glyph.png")])[0]
for c in candidates:
print(f"{c['ids']:<24} mass={c['canonical_log_mass']:.3f} "
f"p={c['mean_token_probability']:.4f}")
Real output for font-rendered glyphs:
明 ⿰日月 mass= -0.004 p=0.9989
品 ⿱口吅 mass= -0.002 p=0.9995
字 ⿱宀子 mass= -0.004 p=0.9988
林 ⿰木木 mass= -0.000 p=1.0000
森 ⿱木林 mass= -0.054 p=0.9863
一 {?3𠄞}⿱一一 mass= -0.953 p=0.7879 <- over-decomposed
木 ⿻木人 mass= -1.344 p=0.7146 <- over-decomposed
Runner.predict returns a list of lists — up to five candidates per input image. Each
candidate carries ids (canonical IDS string), canonical_log_mass, mean_token_probability,
equivalent_raw_ids and valid (whether it parsed as a complete IDS tree). Run without an
argument, example_usage.py renders a few characters from a local font, so it doubles as a
smoke test. Use device="cuda" for GPU inference.
Training
Both checkpoints are from the same lv1 training series and are used as a pair;
inference_config.json records the provenance of each export.
| Checkpoint | Run | Init from | Step | Recorded best |
|---|---|---|---|---|
model.pt |
lv1_diverse_v4 |
lv1_compact_v2/best.pt |
10,000 | 0.8791 |
ensemble_1.pt |
lv1_synthetic_v7 |
lv1_diverse_v4/best.pt |
6,000 | 0.8751 |
Training used the marginal loss: all accepted decompositions of a glyph are treated as separate accepted output events and summed by log-sum-exp per image, instead of forcing a single contradictory label. The component vocabulary is a closed set, so enhanced IDS expressions outside it cannot be produced.
Limitations
- One glyph per image. No detection, no segmentation.
- Simple characters can be over-decomposed. A single-component glyph such as
一or木may receive a breakdown it does not need. These characters are excluded from the accuracy figure quoted above; cross-check them with the Unicode model. - Scores are uncalibrated and short candidate lists are normal.
- Output is one IDS per glyph. Nested structure uses nested operators, not multiple lines.
Data source and licences
- IDS data. Trained on
data/ids_lv1.txtfrom yi-bai/ids — "Yet another IDS (Ideographic Description Sequences) lists" by Yi Bai, MIT licensed. That repository ships three decomposition levels:lv0(all stroke differences kept),lv1(stroke-level differences merged, e.g.丶and乀in certain positions) andlv2(part of the UCV "must unifiable" variants merged). This model uses lv1;normalization.jsonand the canonical output policy derive from the same data set. The bundled copy is content-identical to the upstreamids_lv1.txt— the only difference is line endings (CRLF here, LF upstream); once normalised, its SHA-256 is51f78d7d84f52fd73e5bc155c2239b737c58ce627d446894a809f1a36547ae18. Licence text:licenses/IDS-dictionary-MIT.txt. - Encoder. A PyTorch port of PaddleOCR's
rec_lcnetv4backbone (PPLCNetV4), fine-tuned from a PaddleOCR-initialised encoder. Apache-2.0, Copyright (c) 2021 PaddlePaddle Authors. Licence text:licenses/PaddleOCR-Apache-2.0.txt. - Upstream notices are reproduced in
THIRD_PARTY_NOTICES.md. hanzi_ids/is the reference inference implementation; no model or decoding logic was altered.