TiBLA-RFDETR

Permissive (Apache-2.0) alternative in TiBLA (Tibetan Book Layout Analysis) — a lighter, PyTorch-native RF-DETR-L detector for the page layout of modern Tibetan books (headers, text area, footers, footnotes).

Task

A 4-class detector — header, text-area, footer, footnote — kept as four classes at training time. Evaluation folds them into a 3-class canonical scheme: header+footer are combined into one header-footer class (matched individually, merged losslessly afterwards), text-area is merged to a single page/column envelope as a post-processing step (two boxes only on genuine two-column pages), and footnote is left as-is. All numbers below are in that canonical space, on the leak-free TiBLAD v4 833-page test set, unified scorer (pycocotools bbox mAP 0.50:0.05:0.95; F1 by greedy IoU≥0.5 at the best-mean-F1 operating point).

Inference

# pip install rfdetr
from rfdetr import RFDETRLarge

model = RFDETRLarge.from_checkpoint("rfdetr_tibetan_book_layout.pth")
det = model.predict("page.jpg", threshold=0.26, shape=(1024, 1024))
# checkpoint class ids are offset by 1 (id 0 = background):
#   1 header, 2 text-area, 3 footnote, 4 footer

A ready-made infer.py (batch, YOLO-format output, per-class thresholds) is included in this repo. Recommended global operating confidence: 0.26 (the best-mean-F1 point); the bundled infer.py also ships per-class max-F1 thresholds (header 0.46, text-area 0.32, footnote 0.26, footer 0.52).

Evaluation (TiBLAD v4, 833-page test)

metric TiBLA-RTDETR TiBLA-PP-DocLayout-L TiBLA-RFDETR
license AGPL-3.0 Apache-2.0 Apache-2.0
base model RT-DETR-l (Ultralytics) PP-DocLayout-L (PaddleOCR, RT-DETR-L) RF-DETR-L (Roboflow)
mean F1 (canonical 3-class) 0.959 0.958 0.927
  header-footer F1 0.952 0.951 0.949
  text-area F1 0.999 0.997 0.996
  footnote F1 0.925 0.925 0.835
mean AP@0.50 0.974 0.959 0.925
mean AP@[0.50:0.95] 0.786 0.781 0.667
shared-class mAP@[.50:.95] (DocLayNet-aligned) 0.650 0.641 0.604
Hidden Trespass — header/footer 0.008 0.003 0.020
Hidden Trespass — footnote 0.037 0.037 0.216
COTe (Trespass) 0.975 (0.001) 0.978 (0.000) 0.974 (0.002)
operating confidence 0.74 0.68 0.26

"operating confidence" is the single global best-mean-F1 confidence used for the reported F1.

Hidden Trespass = the missed peripheral (header/footer/footnote) ground-truth area that falls inside the predicted text-area crop; area-based, micro-averaged over the test set. Lower is better (less clutter bled into the OCR region). Formal definition in the paper.

Which checkpoint to pick

checkpoint license mean F1 shared mAP footnote HT
TiBLA-RTDETR (primary) AGPL-3.0 0.959 0.650 0.037
TiBLA-PP-DocLayout-L Apache-2.0 0.958 0.641 0.037
TiBLA-RFDETR Apache-2.0 0.927 0.604 0.216

RT-DETR-l has the top scores but its weights are AGPL-3.0 (Ultralytics). If you need a permissive license, PP-DocLayout-L matches it at Apache-2.0; RF-DETR is a lighter PyTorch-native Apache-2.0 option.

Citation

@misc{tibla2026,
  title        = {TiBLA: Tibetan Book Layout Analysis},
  author       = {Buddhist Digital Resource Center (BDRC)},
  year         = {2026},
  howpublished = {\url{https://github.com/buda-base/tibla}},
  note         = {arXiv link forthcoming}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train BDRC/TiBLA-RFDETR

Collection including BDRC/TiBLA-RFDETR