PP-OCRv6 Hanzi IDS

Predicts the IDS (Ideographic Description Sequence) of a CJK glyph image: a character is described by how it is built out of parts, not by which code point it is. For 明 it returns candidates such as ⿰日月, ⿰日冃, … ranked by likelihood.

  • Scope: clean, single-colour, dark-on-light renderings of one character from a computer font — the rare-character images embedded in dictionary data.
  • Accuracy: 90%+ on that input, excluding characters with no meaningful decomposition — single-component glyphs such as 木, which the model tends to split anyway. See Limitations.
  • No code point needed. The answer is a structure, so unencoded characters, rare variants, regional forms and gaiji (外字) are covered even where the companion Unicode model would have to force them onto the nearest encoded neighbour.

汉字结构预测:输入一张只含一个字的图片,输出最多 5 个 IDS 结构候选,例如 明 → ⿰日月。面向电脑字体渲染的黑字白底单字图,不依赖码位,因此对未编码字、异体、 地域字形、外字同样适用。在上述输入条件下准确率 90% 以上,本身不可拆解的字 (单部件字,如「木」)不在此列。与 Unicode OCR 模型相互独立,可交叉核验。

Model Output
pp-ocrv6-hanzi-ids-ocr (this repo) structure, e.g. ⿰日月
pp-ocrv6-hanzi-unicode-ocr the character, e.g. 明

The two are independent. Together they cross-check each other: IDS candidates are expanded through a dictionary of legal decomposition paths and compared with the structural path of the Unicode result, so a disagreement is surfaced for review.

Files

File Size SHA-256
model.pt 78,774,031 B 8eff12b078a4e8748d0d7d158b22e9df477d9f9a103ca38c35fd3de9ee30b460
ensemble_1.pt 78,787,930 B 21de62d85772220e8c31ee82fe5d1fa867864e75ac4c94ebdcf35242d59891ac
inference_config.json 1,149 B 1bb3196857102e4235f5267901ceb903580682807957f3fd8e5aa3fd45cc4dbb
vocab.json 27,477 B 7f0d2c4be3a852bfcfb3574af1c5a521a929d83fa7140e1ed1f1af61a7112bda
normalization.json 3,349,393 B 4a83fed8522a677fed0cc3f3aa2a732d5a57a2c49211d3e4a85c5ec4058aa2b0
component_output_policy.json 8,868,270 B 047e5f83a737309a68d702d1a2488c8a963d60733ec2f44cb58598029b557e21
data/ids_lv1.txt 2,161,633 B 1f1db7d24767d7622da109ce17f9a3acdb0b61d2b86a314e971f2f2d00a8dbf3
hanzi_ids/ — model definition and inference code
example_usage.py — minimal end-to-end inference script

Architecture

A glyph encoder plus a grammar-constrained autoregressive decoder.

Component Detail
Backbone PPLCNetV4 medium, fused form, kept as a 2-D feature map (no vertical pooling) — the backbone family used by PaddleOCR's PP-OCRv6 recognition models
Projection 1×1 convolution, 768 → 256
Positional learned row (64, 128) + column (128, 128) embeddings, concatenated to 256
Decoder 4 × TransformerDecoderLayer, d_model=256, 8 heads, FFN 1024, GELU, pre-norm, causal mask
Output head weight-tied to the token embedding, plus a per-token bias
Vocabulary 3,459 tokens · max length 96 tokens
Parameters 19,665,475 per checkpoint, FP32
Checkpoints 2, averaged at the probability level during beam search

Vocabulary

vocab.json is a flat list; index 0 is <PAD>, 1 is <BOS>, 2 is <EOS>.

Kind Count Examples
Special 3 <PAD>, <BOS>, <EOS>
Layout operators 138 ⿰, ⿱, ⿲, parameterised ⿻[b_], annotated {冃}⿵
Component classes 136 #(-⺄𠃊), #(HN), #(HNg)
Annotated components 57 {?0㮣}⿱, {卄}⿻
Leaf components 3,176 一, 木, ⺀, ㇁, コ, 㐀

A token's arity — how many children it consumes — follows the IDS operator: ⿲/⿳ take three, the remaining ⿰⿱⿴⿵⿶⿷⿸⿹⿺⿻ pairs take two, ⿾/⿿ take one, leaves none. #(...) stroke programs and {...} annotations are atomic tokens carrying the full parameter text, so a prediction round-trips losslessly.

Decoding

The decoder is never allowed to emit a syntactically invalid sequence: at every step the logits are masked by the number of unfilled slots — <PAD>, <BOS> and any token that would overflow the budget are forbidden, <EOS> is allowed only when zero slots remain, and a token is allowed only if its arity fits the open slots.

Beam search runs at beam_size = 5 over the ensemble; if fewer than five distinct legal outputs are found, the image is re-decoded at beam_size = 12. Candidates are ranked by sequence log probability without length normalisation, matching the training objective. If still fewer than five distinct valid sequences exist, the list is returned short — nothing is copied or invented to pad it.

Raw beam output then passes through the canonical output policy (component_output_policy.json): a subtree is rewritten to a more frequent component spelling only when that spelling is at least 4× more frequent in training (operators, child order and {...} annotations are never rewritten, and each rewrite is verified to preserve the normalised structure); and raw sequences normalising to the same IDS string are merged into one candidate scored by the log-sum-exp of both, their raw forms kept in equivalent_raw_ids.

Scores are uncalibrated: canonical_log_mass is the ranking score, mean_token_probability the geometric-mean token probability of the representative raw path. Use them for ranking, not as accuracy estimates.

Preprocessing

One glyph per image, at 128×128. hanzi_ids.render_utils.normalize_image implements this:

  1. Composite RGBA/LA/palette onto white, convert to 8-bit greyscale.
  2. If the median of the border pixels is darker than 127, invert — dark-on-light and light-on-dark inputs normalise to the same thing.
  3. Crop to the ink bounding box (luminance < 200). Fewer than three ink pixels raises BadGlyph.
  4. Resize with LANCZOS to fit a 113×113 box (aspect preserved), paste centred on a 128×128 white canvas.
  5. The runner scales to [-1, 1] (/127.5 − 1) and replicates greyscale to three channels. Model input is (N, 3, 128, 128) float32.

Usage

A custom architecture means the model needs its class definition, bundled here as the hanzi_ids package:

pip install torch pillow numpy cairosvg
python example_usage.py path/to/glyph.png
from PIL import Image
from hanzi_ids import Runner

runner = Runner(".", device="cpu")
candidates = runner.predict([Image.open("glyph.png")])[0]

for c in candidates:
    print(f"{c['ids']:<24} mass={c['canonical_log_mass']:.3f} "
          f"p={c['mean_token_probability']:.4f}")

Real output for font-rendered glyphs:

明       ⿰日月      mass=  -0.004  p=0.9989
品       ⿱口吅      mass=  -0.002  p=0.9995
字       ⿱宀子      mass=  -0.004  p=0.9988
林       ⿰木木      mass=  -0.000  p=1.0000
森       ⿱木林      mass=  -0.054  p=0.9863

一       {?3𠄞}⿱一一  mass=  -0.953  p=0.7879   <- over-decomposed
木       ⿻木人      mass=  -1.344  p=0.7146   <- over-decomposed

Runner.predict returns a list of lists — up to five candidates per input image. Each candidate carries ids (canonical IDS string), canonical_log_mass, mean_token_probability, equivalent_raw_ids and valid (whether it parsed as a complete IDS tree). Run without an argument, example_usage.py renders a few characters from a local font, so it doubles as a smoke test. Use device="cuda" for GPU inference.

Training

Both checkpoints are from the same lv1 training series and are used as a pair; inference_config.json records the provenance of each export.

Checkpoint Run Init from Step Recorded best
model.pt lv1_diverse_v4 lv1_compact_v2/best.pt 10,000 0.8791
ensemble_1.pt lv1_synthetic_v7 lv1_diverse_v4/best.pt 6,000 0.8751

Training used the marginal loss: all accepted decompositions of a glyph are treated as separate accepted output events and summed by log-sum-exp per image, instead of forcing a single contradictory label. The component vocabulary is a closed set, so enhanced IDS expressions outside it cannot be produced.

Limitations

  • One glyph per image. No detection, no segmentation.
  • Simple characters can be over-decomposed. A single-component glyph such as 一 or 木 may receive a breakdown it does not need. These characters are excluded from the accuracy figure quoted above; cross-check them with the Unicode model.
  • Scores are uncalibrated and short candidate lists are normal.
  • Output is one IDS per glyph. Nested structure uses nested operators, not multiple lines.

Data source and licences

  • IDS data. Trained on data/ids_lv1.txt from yi-bai/ids — "Yet another IDS (Ideographic Description Sequences) lists" by Yi Bai, MIT licensed. That repository ships three decomposition levels: lv0 (all stroke differences kept), lv1 (stroke-level differences merged, e.g. 丶 and 乀 in certain positions) and lv2 (part of the UCV "must unifiable" variants merged). This model uses lv1; normalization.json and the canonical output policy derive from the same data set. The bundled copy is content-identical to the upstream ids_lv1.txt — the only difference is line endings (CRLF here, LF upstream); once normalised, its SHA-256 is 51f78d7d84f52fd73e5bc155c2239b737c58ce627d446894a809f1a36547ae18. Licence text: licenses/IDS-dictionary-MIT.txt.
  • Encoder. A PyTorch port of PaddleOCR's rec_lcnetv4 backbone (PPLCNetV4), fine-tuned from a PaddleOCR-initialised encoder. Apache-2.0, Copyright (c) 2021 PaddlePaddle Authors. Licence text: licenses/PaddleOCR-Apache-2.0.txt.
  • Upstream notices are reproduced in THIRD_PARTY_NOTICES.md.
  • hanzi_ids/ is the reference inference implementation; no model or decoding logic was altered.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support