Instructions to use nie3e/rf-detr-large-doclay with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nie3e/rf-detr-large-doclay with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("object-detection", model="nie3e/rf-detr-large-doclay")# pip install -U transformers accelerate # Load model directly from transformers import AutoImageProcessor, AutoModelForObjectDetection processor = AutoImageProcessor.from_pretrained("nie3e/rf-detr-large-doclay") model = AutoModelForObjectDetection.from_pretrained("nie3e/rf-detr-large-doclay", device_map="auto") - Notebooks
- Google Colab
- Kaggle
rf-detr-large-doclay
This model is a fine-tuned version of Roboflow/rf-detr-large on the DocLayNet v1.2 dataset.
It achieves the following results on the DocLayNet validation split (6,489 pages):
- mAP: 0.7559
- mAP@50: 0.9218
- mAP@75: 0.8207
- mAP (small): 0.5363
- mAP (medium): 0.6524
- mAP (large): 0.8064
- mAR@1: 0.3834
- mAR@10: 0.7858
- mAR@100: 0.8386
- mAR (small): 0.7269
- mAR (medium): 0.7735
- mAR (large): 0.8964
Per-class results (epoch 40)
| Class | mAP | mAR@100 |
|---|---|---|
Caption |
0.8479 | 0.9044 |
Footnote |
0.6837 | 0.8276 |
Formula |
0.5992 | 0.7575 |
List - item |
0.8172 | 0.8821 |
Page - footer |
0.6484 | 0.7007 |
Page - header |
0.6713 | 0.7370 |
Picture |
0.8504 | 0.9191 |
Section - header |
0.6786 | 0.7643 |
Table |
0.8499 | 0.9365 |
Text |
0.8528 | 0.8944 |
Title |
0.8160 | 0.9013 |
Model description
Fine-tuned version of Roboflow/rf-detr-large for document layout analysis.
It detects the 11 layout regions annotated in DocLayNet v1.2 dataset:
{
"0": "Caption",
"1": "Footnote",
"2": "Formula",
"3": "List - item",
"4": "Page - footer",
"5": "Page - header",
"6": "Picture",
"7": "Section - header",
"8": "Table",
"9": "Text",
"10": "Title"
}
This was an exploratory run: the schedule (40 epochs, cosine, LR 5e-5) was not tuned and the published weights are the last checkpoint, not the best-scoring one.
33.6 M parameters (FP32, ~128 MiB), ViT backbone at 704 x 704, up to 300 boxes per page.
Intended uses & limitations
Usage
Requires transformers ≥ 5.9.0
import torch
from PIL import Image, ImageDraw, ImageFont
from transformers import AutoImageProcessor, AutoModelForObjectDetection
checkpoint_name = "nie3e/rf-detr-large-doclay"
device = "cuda" if torch.cuda.is_available() else "cpu"
image_processor = AutoImageProcessor.from_pretrained(checkpoint_name)
model = AutoModelForObjectDetection.from_pretrained(checkpoint_name).to(device)
model.eval()
image = Image.open("/path/to/your/image.jpg").convert("RGB")
# Inference
with torch.no_grad():
inputs = image_processor(images=[image], return_tensors="pt").to(device)
outputs = model(**inputs)
target_sizes = torch.tensor([[image.height, image.width]])
results = image_processor.post_process_object_detection(
outputs,
threshold=0.5, # change this if needed
target_sizes=target_sizes,
)[0]
# Print detections
for score, label, box in zip(
results["scores"],
results["labels"],
results["boxes"],
):
box = [round(x, 2) for x in box.tolist()]
print(
f"Detected {model.config.id2label[label.item()]} "
f"with confidence {score.item():.3f} "
f"at {box}"
)
# Show detections
draw = ImageDraw.Draw(image)
colors = {
"Caption": "#F59E0B",
"Footnote": "#8B5CF6",
"Formula": "#EC4899",
"List - item": "#14B8A6",
"Page - footer": "#64748B",
"Page - header": "#06B6D4",
"Picture": "#7C3AED",
"Section - header": "#F97316",
"Table": "#22C55E",
"Text": "#2563EB",
"Title": "#DC2626",
}
try:
font = ImageFont.truetype("DejaVuSans-Bold.ttf", 18)
except OSError:
font = ImageFont.load_default(size=18)
for score, label, box in zip(
results["scores"], results["labels"], results["boxes"]
):
box = [round(i, 2) for i in box.tolist()]
x, y, x2, y2 = tuple(box)
label_name = model.config.id2label[label.item()]
color = colors.get(label_name, "red")
draw.rectangle((x, y, x2, y2), outline=color, width=2)
text = f"{label_name} {score:.2f}"
text_bbox = draw.textbbox((0, 0), text, font=font)
text_width = text_bbox[2] - text_bbox[0]
text_height = text_bbox[3] - text_bbox[1]
padding = 4
bg_x1 = x
bg_y1 = max(0, y - text_height - 2 * padding)
bg_x2 = x + text_width + 2 * padding
bg_y2 = y
draw.rounded_rectangle(
(bg_x1, bg_y1, bg_x2, bg_y2),
radius=4,
fill=color
)
draw.text(
(x + padding, bg_y1 + padding - text_bbox[1]),
text,
font=font,
fill="white"
)
image.show()
Limitations
- may not perform well on small objects, mAP (small) 0.5363 vs 0.8064 for large
- boxes and class ids only
- no OCR
- no masks
- no table cell structure
- no reading order
- weakest classes:
Formula,Page - footer,Page - header - print-oriented corpus
Sample detections
Example 1
Example 2
Example 3
Training and evaluation data
Dataset: DocLayNet v1.2
Evaluation was run every 5 epochs on the validation split.
Data augmentation
Applied to the train split only (validation is used raw).
import albumentations as A
train_augment = A.Compose(
[
A.RandomBrightnessContrast(brightness_limit=0.2,
contrast_limit=0.2, p=0.5),
A.HueSaturationValue(hue_shift_limit=5, sat_shift_limit=10,
val_shift_limit=10, p=0.2),
A.Rotate(limit=2, border_mode=0, value=(255, 255, 255), p=0.3),
A.Affine(translate_percent=(-0.05, 0.05), scale=(0.95, 1.05),
rotate=0, p=0.3),
A.GaussNoise(p=0.15),
A.ImageCompression(quality_range=(85, 100), p=0.2),
A.Blur(blur_limit=3, p=0.15),
],
bbox_params=A.BboxParams(
format="coco",
label_fields=["category"],
clip=True,
min_area=4,
),
)
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 5e-05
- train_batch_size: 8
- eval_batch_size: 8
- seed: 42
- distributed_type: multi-GPU
- num_devices: 4
- gradient_accumulation_steps: 4
- total_train_batch_size: 128
- total_eval_batch_size: 32
- optimizer: AdamW (fused PyTorch implementation), β=(0.9, 0.999), ε=1e-8
- lr_scheduler_type: cosine
- num_epochs: 40
Training results
Metrics by epoch
| Training Loss | Epoch | Step | mAP | mAP@50 | mAP@75 | mAP (small) | mAP (medium) | mAP (large) | mAR@1 | mAR@10 | mAR@100 | mAR (small) | mAR (medium) | mAR (large) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 4.3588 | 5.0 | 2710 | 0.6338 | 0.851 | 0.6994 | 0.3613 | 0.4844 | 0.6971 | 0.3474 | 0.7152 | 0.7764 | 0.6137 | 0.6963 | 0.8376 |
| 3.7698 | 10.0 | 5420 | 0.6946 | 0.8934 | 0.7567 | 0.4623 | 0.5716 | 0.7531 | 0.3636 | 0.7466 | 0.8023 | 0.6595 | 0.7304 | 0.857 |
| 3.0678 | 15.0 | 8130 | 0.7228 | 0.9106 | 0.7939 | 0.5105 | 0.6039 | 0.7678 | 0.372 | 0.7625 | 0.816 | 0.6778 | 0.7569 | 0.8634 |
| 2.9566 | 20.0 | 10840 | 0.7331 | 0.9179 | 0.7989 | 0.5333 | 0.6353 | 0.7894 | 0.3744 | 0.7703 | 0.8224 | 0.692 | 0.7718 | 0.8866 |
| 2.6887 | 25.0 | 13550 | 0.7387 | 0.9203 | 0.7977 | 0.5085 | 0.6339 | 0.7995 | 0.3777 | 0.7717 | 0.8234 | 0.6819 | 0.7629 | 0.8916 |
| 2.5317 | 30.0 | 16260 | 0.754 | 0.9214 | 0.8241 | 0.5409 | 0.6527 | 0.8028 | 0.3824 | 0.7848 | 0.8374 | 0.7352 | 0.7763 | 0.894 |
| 2.4486 | 35.0 | 18970 | 0.7537 | 0.9221 | 0.818 | 0.5341 | 0.6548 | 0.8056 | 0.3826 | 0.784 | 0.8369 | 0.7223 | 0.7738 | 0.895 |
| 2.3271 | 40.0 | 21680 | 0.7559 | 0.9218 | 0.8207 | 0.5363 | 0.6524 | 0.8064 | 0.3834 | 0.7858 | 0.8386 | 0.7269 | 0.7735 | 0.8964 |
Per-class mAP by epoch
| Class | ep 5 | ep 10 | ep 15 | ep 20 | ep 25 | ep 30 | ep 35 | ep 40 |
|---|---|---|---|---|---|---|---|---|
Caption |
0.7177 | 0.7901 | 0.8127 | 0.8257 | 0.835 | 0.8471 | 0.8461 | 0.8479 |
Footnote |
0.4279 | 0.5381 | 0.5963 | 0.6563 | 0.6463 | 0.6695 | 0.683 | 0.6837 |
Formula |
0.5013 | 0.542 | 0.5675 | 0.592 | 0.6012 | 0.5976 | 0.5993 | 0.5992 |
List - item |
0.6846 | 0.7626 | 0.796 | 0.7913 | 0.8083 | 0.8184 | 0.8188 | 0.8172 |
Page - footer |
0.5496 | 0.5894 | 0.6052 | 0.5922 | 0.6151 | 0.6449 | 0.6437 | 0.6484 |
Page - header |
0.6504 | 0.6328 | 0.6531 | 0.6581 | 0.6346 | 0.6856 | 0.6661 | 0.6713 |
Picture |
0.771 | 0.8135 | 0.8285 | 0.8422 | 0.8499 | 0.8471 | 0.8468 | 0.8504 |
Section - header |
0.6004 | 0.6545 | 0.6631 | 0.651 | 0.6622 | 0.6841 | 0.6731 | 0.6786 |
Table |
0.7599 | 0.8092 | 0.8358 | 0.8433 | 0.8466 | 0.8482 | 0.8496 | 0.8499 |
Text |
0.7889 | 0.8251 | 0.837 | 0.8396 | 0.8365 | 0.858 | 0.8512 | 0.8528 |
Title |
0.5198 | 0.6837 | 0.7553 | 0.7721 | 0.7905 | 0.7939 | 0.8126 | 0.816 |
Per-class mAR@100 by epoch
| Class | ep 5 | ep 10 | ep 15 | ep 20 | ep 25 | ep 30 | ep 35 | ep 40 |
|---|---|---|---|---|---|---|---|---|
Caption |
0.8217 | 0.8638 | 0.8767 | 0.882 | 0.8918 | 0.9012 | 0.9011 | 0.9044 |
Footnote |
0.7244 | 0.7897 | 0.7978 | 0.8311 | 0.8179 | 0.824 | 0.8288 | 0.8276 |
Formula |
0.7119 | 0.7353 | 0.7364 | 0.7473 | 0.7555 | 0.7576 | 0.7562 | 0.7575 |
List - item |
0.8135 | 0.8471 | 0.8669 | 0.8649 | 0.8771 | 0.8851 | 0.882 | 0.8821 |
Page - footer |
0.6255 | 0.646 | 0.662 | 0.6545 | 0.6699 | 0.694 | 0.6975 | 0.7007 |
Page - header |
0.7222 | 0.6957 | 0.7081 | 0.7201 | 0.7019 | 0.7489 | 0.7334 | 0.737 |
Picture |
0.8806 | 0.8969 | 0.9086 | 0.9155 | 0.916 | 0.9169 | 0.9181 | 0.9191 |
Section - header |
0.7241 | 0.7469 | 0.7447 | 0.7457 | 0.7452 | 0.7668 | 0.7609 | 0.7643 |
Table |
0.8975 | 0.9174 | 0.9327 | 0.9349 | 0.9292 | 0.9404 | 0.9368 | 0.9365 |
Text |
0.8585 | 0.8749 | 0.8815 | 0.8841 | 0.8799 | 0.8991 | 0.8929 | 0.8944 |
Title |
0.7605 | 0.8117 | 0.8609 | 0.8666 | 0.8732 | 0.8779 | 0.898 | 0.9013 |
Framework versions
- Transformers 5.17.0
- PyTorch 2.13.0+rocm10.0.0
- Datasets 5.0.1
- Tokenizers 0.23.2
- Albumentations 2.0.8
- Accelerate 1.15.0
Citation
This model is a fine-tune of RF-DETR trained on DocLayNet. If you use it, please cite both:
@article{robinson2025rfdetr,
title = {RF-DETR: Neural Architecture Search for Real-Time Detection Transformers},
author = {Robinson, Isaac and Robicheaux, Peter and Popov, Matvei and Ramanan, Deva and Peri, Neehar},
journal = {arXiv preprint arXiv:2511.09554},
year = {2025},
url = {https://arxiv.org/abs/2511.09554}
}
@inproceedings{pfitzmann2022doclaynet,
title = {DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis},
author = {Pfitzmann, Birgit and Auer, Christoph and Dolfi, Michele and Nassar, Ahmed S. and Staar, Peter W. J.},
booktitle = {Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 22)},
year = {2022},
doi = {10.1145/3534678.3539043},
url = {https://arxiv.org/abs/2206.01062}
}
- Downloads last month
- 23
Model tree for nie3e/rf-detr-large-doclay
Base model
Roboflow/rf-detr-large