rf-detr-large-doclay

This model is a fine-tuned version of Roboflow/rf-detr-large on the DocLayNet v1.2 dataset.

It achieves the following results on the DocLayNet validation split (6,489 pages):

  • mAP: 0.7559
  • mAP@50: 0.9218
  • mAP@75: 0.8207
  • mAP (small): 0.5363
  • mAP (medium): 0.6524
  • mAP (large): 0.8064
  • mAR@1: 0.3834
  • mAR@10: 0.7858
  • mAR@100: 0.8386
  • mAR (small): 0.7269
  • mAR (medium): 0.7735
  • mAR (large): 0.8964
Per-class results (epoch 40)
Class mAP mAR@100
Caption 0.8479 0.9044
Footnote 0.6837 0.8276
Formula 0.5992 0.7575
List - item 0.8172 0.8821
Page - footer 0.6484 0.7007
Page - header 0.6713 0.7370
Picture 0.8504 0.9191
Section - header 0.6786 0.7643
Table 0.8499 0.9365
Text 0.8528 0.8944
Title 0.8160 0.9013

Model description

Fine-tuned version of Roboflow/rf-detr-large for document layout analysis.

It detects the 11 layout regions annotated in DocLayNet v1.2 dataset:

{
  "0": "Caption",
  "1": "Footnote",
  "2": "Formula",
  "3": "List - item",
  "4": "Page - footer",
  "5": "Page - header",
  "6": "Picture",
  "7": "Section - header",
  "8": "Table",
  "9": "Text",
  "10": "Title"
}

This was an exploratory run: the schedule (40 epochs, cosine, LR 5e-5) was not tuned and the published weights are the last checkpoint, not the best-scoring one.

33.6 M parameters (FP32, ~128 MiB), ViT backbone at 704 x 704, up to 300 boxes per page.

Intended uses & limitations

Usage

Requires transformers ≥ 5.9.0

import torch
from PIL import Image, ImageDraw, ImageFont
from transformers import AutoImageProcessor, AutoModelForObjectDetection

checkpoint_name = "nie3e/rf-detr-large-doclay"
device = "cuda" if torch.cuda.is_available() else "cpu"
image_processor = AutoImageProcessor.from_pretrained(checkpoint_name)
model = AutoModelForObjectDetection.from_pretrained(checkpoint_name).to(device)
model.eval()

image = Image.open("/path/to/your/image.jpg").convert("RGB")

# Inference
with torch.no_grad():
    inputs = image_processor(images=[image], return_tensors="pt").to(device)
    outputs = model(**inputs)

    target_sizes = torch.tensor([[image.height, image.width]])

    results = image_processor.post_process_object_detection(
        outputs,
        threshold=0.5,  # change this if needed
        target_sizes=target_sizes,
    )[0]

# Print detections
for score, label, box in zip(
    results["scores"],
    results["labels"],
    results["boxes"],
):
    box = [round(x, 2) for x in box.tolist()]

    print(
        f"Detected {model.config.id2label[label.item()]} "
        f"with confidence {score.item():.3f} "
        f"at {box}"
    )

# Show detections
draw = ImageDraw.Draw(image)
colors = {
    "Caption": "#F59E0B",
    "Footnote": "#8B5CF6",
    "Formula": "#EC4899",
    "List - item": "#14B8A6",
    "Page - footer": "#64748B",
    "Page - header": "#06B6D4",
    "Picture": "#7C3AED",
    "Section - header": "#F97316",
    "Table": "#22C55E",
    "Text": "#2563EB",
    "Title": "#DC2626",
}

try:
    font = ImageFont.truetype("DejaVuSans-Bold.ttf", 18)
except OSError:
    font = ImageFont.load_default(size=18)
for score, label, box in zip(
        results["scores"], results["labels"], results["boxes"]
):
    box = [round(i, 2) for i in box.tolist()]
    x, y, x2, y2 = tuple(box)

    label_name = model.config.id2label[label.item()]
    color = colors.get(label_name, "red")

    draw.rectangle((x, y, x2, y2), outline=color, width=2)

    text = f"{label_name} {score:.2f}"

    text_bbox = draw.textbbox((0, 0), text, font=font)
    text_width = text_bbox[2] - text_bbox[0]
    text_height = text_bbox[3] - text_bbox[1]

    padding = 4

    bg_x1 = x
    bg_y1 = max(0, y - text_height - 2 * padding)
    bg_x2 = x + text_width + 2 * padding
    bg_y2 = y

    draw.rounded_rectangle(
        (bg_x1, bg_y1, bg_x2, bg_y2),
        radius=4,
        fill=color
    )

    draw.text(
        (x + padding, bg_y1 + padding - text_bbox[1]),
        text,
        font=font,
        fill="white"
    )

image.show()

Limitations

  • may not perform well on small objects, mAP (small) 0.5363 vs 0.8064 for large
  • boxes and class ids only
  • no OCR
  • no masks
  • no table cell structure
  • no reading order
  • weakest classes: Formula, Page - footer, Page - header
  • print-oriented corpus

Sample detections

Example 1
Example 2
Example 3

Training and evaluation data

Dataset: DocLayNet v1.2

Evaluation was run every 5 epochs on the validation split.

Data augmentation

Applied to the train split only (validation is used raw).

import albumentations as A

train_augment = A.Compose(
    [
        A.RandomBrightnessContrast(brightness_limit=0.2,
                                   contrast_limit=0.2, p=0.5),
        A.HueSaturationValue(hue_shift_limit=5, sat_shift_limit=10,
                             val_shift_limit=10, p=0.2),
        A.Rotate(limit=2, border_mode=0, value=(255, 255, 255), p=0.3),
        A.Affine(translate_percent=(-0.05, 0.05), scale=(0.95, 1.05),
                 rotate=0, p=0.3),
        A.GaussNoise(p=0.15),
        A.ImageCompression(quality_range=(85, 100), p=0.2),
        A.Blur(blur_limit=3, p=0.15),
    ],
    bbox_params=A.BboxParams(
        format="coco",
        label_fields=["category"],
        clip=True,
        min_area=4,
    ),
)

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-05
  • train_batch_size: 8
  • eval_batch_size: 8
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 4
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 128
  • total_eval_batch_size: 32
  • optimizer: AdamW (fused PyTorch implementation), β=(0.9, 0.999), ε=1e-8
  • lr_scheduler_type: cosine
  • num_epochs: 40

Training results

Metrics by epoch

Training Loss Epoch Step mAP mAP@50 mAP@75 mAP (small) mAP (medium) mAP (large) mAR@1 mAR@10 mAR@100 mAR (small) mAR (medium) mAR (large)
4.3588 5.0 2710 0.6338 0.851 0.6994 0.3613 0.4844 0.6971 0.3474 0.7152 0.7764 0.6137 0.6963 0.8376
3.7698 10.0 5420 0.6946 0.8934 0.7567 0.4623 0.5716 0.7531 0.3636 0.7466 0.8023 0.6595 0.7304 0.857
3.0678 15.0 8130 0.7228 0.9106 0.7939 0.5105 0.6039 0.7678 0.372 0.7625 0.816 0.6778 0.7569 0.8634
2.9566 20.0 10840 0.7331 0.9179 0.7989 0.5333 0.6353 0.7894 0.3744 0.7703 0.8224 0.692 0.7718 0.8866
2.6887 25.0 13550 0.7387 0.9203 0.7977 0.5085 0.6339 0.7995 0.3777 0.7717 0.8234 0.6819 0.7629 0.8916
2.5317 30.0 16260 0.754 0.9214 0.8241 0.5409 0.6527 0.8028 0.3824 0.7848 0.8374 0.7352 0.7763 0.894
2.4486 35.0 18970 0.7537 0.9221 0.818 0.5341 0.6548 0.8056 0.3826 0.784 0.8369 0.7223 0.7738 0.895
2.3271 40.0 21680 0.7559 0.9218 0.8207 0.5363 0.6524 0.8064 0.3834 0.7858 0.8386 0.7269 0.7735 0.8964

Per-class mAP by epoch

Class ep 5 ep 10 ep 15 ep 20 ep 25 ep 30 ep 35 ep 40
Caption 0.7177 0.7901 0.8127 0.8257 0.835 0.8471 0.8461 0.8479
Footnote 0.4279 0.5381 0.5963 0.6563 0.6463 0.6695 0.683 0.6837
Formula 0.5013 0.542 0.5675 0.592 0.6012 0.5976 0.5993 0.5992
List - item 0.6846 0.7626 0.796 0.7913 0.8083 0.8184 0.8188 0.8172
Page - footer 0.5496 0.5894 0.6052 0.5922 0.6151 0.6449 0.6437 0.6484
Page - header 0.6504 0.6328 0.6531 0.6581 0.6346 0.6856 0.6661 0.6713
Picture 0.771 0.8135 0.8285 0.8422 0.8499 0.8471 0.8468 0.8504
Section - header 0.6004 0.6545 0.6631 0.651 0.6622 0.6841 0.6731 0.6786
Table 0.7599 0.8092 0.8358 0.8433 0.8466 0.8482 0.8496 0.8499
Text 0.7889 0.8251 0.837 0.8396 0.8365 0.858 0.8512 0.8528
Title 0.5198 0.6837 0.7553 0.7721 0.7905 0.7939 0.8126 0.816

Per-class mAR@100 by epoch

Class ep 5 ep 10 ep 15 ep 20 ep 25 ep 30 ep 35 ep 40
Caption 0.8217 0.8638 0.8767 0.882 0.8918 0.9012 0.9011 0.9044
Footnote 0.7244 0.7897 0.7978 0.8311 0.8179 0.824 0.8288 0.8276
Formula 0.7119 0.7353 0.7364 0.7473 0.7555 0.7576 0.7562 0.7575
List - item 0.8135 0.8471 0.8669 0.8649 0.8771 0.8851 0.882 0.8821
Page - footer 0.6255 0.646 0.662 0.6545 0.6699 0.694 0.6975 0.7007
Page - header 0.7222 0.6957 0.7081 0.7201 0.7019 0.7489 0.7334 0.737
Picture 0.8806 0.8969 0.9086 0.9155 0.916 0.9169 0.9181 0.9191
Section - header 0.7241 0.7469 0.7447 0.7457 0.7452 0.7668 0.7609 0.7643
Table 0.8975 0.9174 0.9327 0.9349 0.9292 0.9404 0.9368 0.9365
Text 0.8585 0.8749 0.8815 0.8841 0.8799 0.8991 0.8929 0.8944
Title 0.7605 0.8117 0.8609 0.8666 0.8732 0.8779 0.898 0.9013

Framework versions

  • Transformers 5.17.0
  • PyTorch 2.13.0+rocm10.0.0
  • Datasets 5.0.1
  • Tokenizers 0.23.2
  • Albumentations 2.0.8
  • Accelerate 1.15.0

Citation

This model is a fine-tune of RF-DETR trained on DocLayNet. If you use it, please cite both:

@article{robinson2025rfdetr,
  title   = {RF-DETR: Neural Architecture Search for Real-Time Detection Transformers},
  author  = {Robinson, Isaac and Robicheaux, Peter and Popov, Matvei and Ramanan, Deva and Peri, Neehar},
  journal = {arXiv preprint arXiv:2511.09554},
  year    = {2025},
  url     = {https://arxiv.org/abs/2511.09554}
}

@inproceedings{pfitzmann2022doclaynet,
  title     = {DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis},
  author    = {Pfitzmann, Birgit and Auer, Christoph and Dolfi, Michele and Nassar, Ahmed S. and Staar, Peter W. J.},
  booktitle = {Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 22)},
  year      = {2022},
  doi       = {10.1145/3534678.3539043},
  url       = {https://arxiv.org/abs/2206.01062}
}
Downloads last month
23
Safetensors
Model size
33.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nie3e/rf-detr-large-doclay

Finetuned
(6)
this model

Dataset used to train nie3e/rf-detr-large-doclay

Papers for nie3e/rf-detr-large-doclay