TiBLA-RTDETR

Primary checkpoint of TiBLA (Tibetan Book Layout Analysis) β€” an RT-DETR-l detector for the page layout of modern Tibetan books (headers, text area, footers, footnotes).

  • Base model / provenance: RT-DETR-l (Ultralytics), fine-tuned on the leak-free v4 tam2col split of TiBLAD.
  • License: AGPL-3.0 (inherited from the Ultralytics RT-DETR weights).
  • Dataset: BDRC/TiBLAD
  • Paper: buda-base/papers (papers/2026-tibetan-book-layout) β€” arXiv link forthcoming
  • Code: github.com/buda-base/tibla

This checkpoint is seed 0; across five training seeds the paper reports mean F1 0.961 Β± 0.009 (this seed scores 0.959, just below the mean).

Task

A 4-class detector β€” header, text-area, footer, footnote β€” kept as four classes at training time. Evaluation folds them into a 3-class canonical scheme: header+footer are combined into one header-footer class (matched individually, merged losslessly afterwards), text-area is merged to a single page/column envelope as a post-processing step (two boxes only on genuine two-column pages), and footnote is left as-is. All numbers below are in that canonical space, on the leak-free TiBLAD v4 833-page test set, unified scorer (pycocotools bbox mAP 0.50:0.05:0.95; F1 by greedy IoUβ‰₯0.5 at the best-mean-F1 operating point).

Inference

# pip install ultralytics
from ultralytics import RTDETR

model = RTDETR("tibetan_book_layout.pt")
# recommended per-class confidence thresholds (see below); predict at the floor
res = model.predict("page.jpg", imgsz=1024, conf=0.25)[0]
TH = {0: 0.60, 1: 0.55, 2: 0.25, 3: 0.60}  # header / text-area / footnote / footer
for cls, conf, xywhn in zip(res.boxes.cls.tolist(), res.boxes.conf.tolist(),
                            res.boxes.xywhn.tolist()):
    if conf >= TH[int(cls)]:
        print(res.names[int(cls)], round(conf, 3), [round(v, 4) for v in xywhn])

A ready-made infer.py (batch, YOLO-format output) is included in this repo.

Recommended confidence thresholds (per-class max-F1 operating points): header/footer β‰ˆ 0.60, text-area β‰ˆ 0.55, footnote β‰ˆ 0.25. Footnote is deliberately kept low (recall-safe): the v4 test has only 38 footnote GT boxes, so a low threshold keeps recall near 1.0. Raising header/footer from 0.25 to 0.60 lifts precision +0.028 for a βˆ’0.014 recall cost; raising text-area from 0.25 to 0.55 (native) lifts precision +0.007 at no recall cost. If you prefer one global knob, the single best-mean-F1 confidence is 0.74 (costs β‰ˆ0.008 mean F1 vs per-class tuning).

Evaluation (TiBLAD v4, 833-page test)

metric TiBLA-RTDETR TiBLA-PP-DocLayout-L TiBLA-RFDETR
license AGPL-3.0 Apache-2.0 Apache-2.0
base model RT-DETR-l (Ultralytics) PP-DocLayout-L (PaddleOCR, RT-DETR-L) RF-DETR-L (Roboflow)
mean F1 (canonical 3-class) 0.959 0.958 0.927
  header-footer F1 0.952 0.951 0.949
  text-area F1 0.999 0.997 0.996
  footnote F1 0.925 0.925 0.835
mean AP@0.50 0.974 0.959 0.925
mean AP@[0.50:0.95] 0.786 0.781 0.667
shared-class mAP@[.50:.95] (DocLayNet-aligned) 0.650 0.641 0.604
Hidden Trespass β€” header/footer 0.008 0.003 0.020
Hidden Trespass β€” footnote 0.037 0.037 0.216
COTe (Trespass) 0.975 (0.001) 0.978 (0.000) 0.974 (0.002)
operating confidence 0.74 0.68 0.26

"operating confidence" is the single global best-mean-F1 confidence used for the reported F1.

Hidden Trespass = the missed peripheral (header/footer/footnote) ground-truth area that falls inside the predicted text-area crop; area-based, micro-averaged over the test set. Lower is better (less clutter bled into the OCR region). Formal definition in the paper.

Which checkpoint to pick

checkpoint license mean F1 shared mAP footnote HT
TiBLA-RTDETR (primary) AGPL-3.0 0.959 0.650 0.037
TiBLA-PP-DocLayout-L Apache-2.0 0.958 0.641 0.037
TiBLA-RFDETR Apache-2.0 0.927 0.604 0.216

RT-DETR-l has the top scores but its weights are AGPL-3.0 (Ultralytics). If you need a permissive license, PP-DocLayout-L matches it at Apache-2.0; RF-DETR is a lighter PyTorch-native Apache-2.0 option.

Citation

@misc{tibla2026,
  title        = {TiBLA: Tibetan Book Layout Analysis},
  author       = {Buddhist Digital Resource Center (BDRC)},
  year         = {2026},
  howpublished = {\url{https://github.com/buda-base/tibla}},
  note         = {arXiv link forthcoming}
}
Downloads last month
29
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train BDRC/TiBLA-RTDETR

Collection including BDRC/TiBLA-RTDETR