akhet models

Models behind akhet, which reads Ancient Egyptian hieroglyphs from images: text lines, signs, Gardiner codes, reading order as Manuel de Codage, transliteration and translation, with a human in the loop at every step.

pip install "akhet[gpu,hub,translate] @ git+https://github.com/AIRI-Institute/akhet#subdirectory=pipeline"
from akhet import Akhet, text_blocks, Translator
page = Akhet.from_pretrained()("page.jpg")          # OCR, editable
blocks = text_blocks(page)                         # editable text blocks
tr = Translator.from_pretrained()
tr.transliterate(blocks)
tr.translate(blocks, "en")

Files

file role input size
lines.onnx text line detector (RT-DETR, PP-DocLayout-plus-L) page, keep-aspect resize + pad to 1024 125 MB
glyphs_row.onnx, glyphs_col.onnx hieroglyph detector (RT-DETR), one per orientation 256 x 2048 / 2048 x 256 line tile 124 MB each
direction_nonegypt.onnx line reading direction + modern text (RT-DETR) line crop, resize + pad to 1024 125 MB
sign_classifier.onnx sign embedding (DINOv2-base, ArcFace) 224 x 224 sign crop 330 MB
sign_bank.npy, sign_bank.json one prototype per Gardiner sign (980 signs) 1 MB
nllb/ fine-tuned NLLB-1.3B, fp16: hieroglyphs to transliteration to English / German. Ships tokenizer.json: load it as is, transformers 5.x would otherwise rebuild the extended vocabulary with shifted ids Gardiner text, TLA spelling 2.6 GB

The detectors output raw RT-DETR queries (boxes + class logits). The akhet package holds the exact pre- and post-processing.

Held-out results

model test result
lines 5-fold cross-validation, 620 pages F1@0.5 0.84
lines external benchmark, 84 pages F1@0.5 0.82
hieroglyphs 141 real lines AP@0.5 0.96, AP@0.75 0.86
direction + modern text 150 real lines direction accuracy 0.91, modern-text AP@0.5 0.73
sign classifier real sign crops, 3 seeds top-1 0.86, macro-F1 0.69
sign classifier signs absent from real training data top-10 0.86
translation model TLA test clauses (400 / 265 / 300) hieroglyphs to transliteration chrF 59.0 BLEU 60.9, to English via transliteration chrF 45.4 BLEU 25.1, to German chrF 51.9 BLEU 33.0

After evaluation the detectors and the classifier were retrained on all annotated data with the same recipe. These released weights are those final models. Converted to ONNX, they reproduce the original PaddleDetection and PyTorch models within 0.001 on every metric above. The fp16 translation model gives the same output as the fp32 original.

Training data

Annotated page photographs and hand copies of hieroglyphic inscriptions, open-access museum images, synthetic hieroglyph lines rendered from the Thesaurus Linguae Aegyptiae with a learned layout model, and sign crops from public Egyptian hieroglyph datasets. Sequence and translation data come from the Thesaurus Linguae Aegyptiae (CC BY-SA 4.0).

License

CC BY-NC-SA 4.0: non-commercial use only, derived work under the same license. Part of the image training data is licensed for non-commercial use, and the translation model builds on NLLB-200 (CC BY-NC 4.0).

Limitations

Trained on monumental hieroglyphic inscriptions, not hieratic. A sign outside the 980-sign inventory is mapped to the closest known sign. Transliteration and translation are drafts for a scholar to check.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support