Instructions to use harness-race/control-r3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use harness-race/control-r3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("object-detection", model="harness-race/control-r3")# Load model directly from transformers import AutoImageProcessor, AutoModelForObjectDetection processor = AutoImageProcessor.from_pretrained("harness-race/control-r3") model = AutoModelForObjectDetection.from_pretrained("harness-race/control-r3", device_map="auto") - Notebooks
- Google Colab
- Kaggle
control-r3 — DETR object detection for news-page layout regions
This model is a fine-tuned DetrForObjectDetection (DETR with a ResNet-50 backbone,
starting from facebook/detr-resnet-50)
tuned to localize layout regions in historical and modern newspaper pages.
It predicts bounding boxes for the 7 classes of the
BigLam “Locating Objects Beyond Words” dataset:
| class id | label |
|---|---|
| 0 | Photograph |
| 1 | Illustration |
| 2 | Map |
| 3 | Comics/Cartoon |
| 4 | Editorial Cartoon |
| 5 | Headline |
| 6 | Advertisement |
The base model facebook/detr-resnet-50 is
released under the Apache 2.0 license, so this fine-tune can be freely shared and used.
Model Details
- Model type: DETR (DEtection TRansformer), encoder-decoder Transformer with a ResNet-50 backbone.
- Backbone: ResNet-50 (provided by
timm), pretrained on ImageNet. - Number of queries: 100.
- Input: RGB image resized/padded to a square of
600 × 600after a smallest-max-size resize. - Output: per-image box predictions with class logits for the 7 layout classes (plus the implicit
no objectclass). - Optimized by: fine-tuned on the training split (2,846 images) of
biglam/loc_beyond_words. - Base model license: Apache 2.0.
Intended Use
The model is intended for document / newspaper-page layout analysis: given a scan or a page image, it detects coarse layout regions such as headlines, photographs/illustrations, advertisements, maps and comics. It is designed as a layout-region detector and is not meant for fine-grained text recognition or OCR (use OCR/HTR systems for reading text).
Example inference:
from transformers import pipeline
detector = pipeline("object-detection", model="harness-race/control-r3")
results = detector("path/to/newspaper_page.png")
# results: list of {label, score, box: {xmin, ymin, xmax, ymax}}
Training Data
- Dataset:
biglam/loc_beyond_words(BigLam “Locating Objects Beyond Words”, a Library-of-Congress-derived newspaper layout dataset). - Training split: 2,846 images.
- Validation split: 712 images (reported results are on this held-out split).
- Format: COCO-style bounding boxes
[x, y, width, height]in pixel coordinates. - The data was repackaged (only renaming annotation fields for compatibility) into
harness-race/loc_beyond_words_coco; no annotations were modified.
Training Procedure
The model was fine-tuned end-to-end (all weights trainable) with the Hugging Face Trainer-style loop
on a single NVIDIA T4 GPU (fp16 AMP), with the RGB images resized and padded to 600 × 600 and light
augmentation (horizontal flip, random brightness/contrast, hue/saturation, random crop with box clipping).
Hyperparameters
- optimizer: AdamW (betas 0.9/0.999, eps 1e-8)
- learning rate: 1e-4 with linear schedule + 10% warmup
- weight decay: 1e-4
- train batch size: 4
- eval batch size: 8
- epochs: 8
- mixed precision: fp16 (native AMP)
- image size: 600 × 600 (resize + pad)
- gradient clipping: 1.0
- seed: 42
Evaluation is run at the end of every epoch and the checkpoint with the best validation mAP is kept.
Evaluation Results
Reported on the validation split (712 images), using COCO-style metrics
(torchmetrics.MeanAveragePrecision, box_format=xyxy). Metrics are in %:
| Metric | Value |
|---|---|
| mAP (IoU .5:.95) | 24.65 |
| mAP @ IoU 0.50 | 34.76 |
| mAP @ IoU 0.75 | 28.39 |
| mAR@100 | 37.94 |
Per-class mAP (IoU .5:.95):
| Class | mAP |
|---|---|
| Photograph | 39.02 |
| Illustration | 1.17 |
| Map | 0.03 |
| Comics/Cartoon | 13.21 |
| Editorial Cartoon | 0.00 |
| Headline | 59.51 |
| Advertisement | 59.64 |
The model detects Headline and Advertisement regions very well (>59 mAP) and detects Photograph and Comics/Cartoon reasonably. The rare classes (Illustration, Map, Editorial Cartoon) show very low mAP, which is largely a consequence of heavy class imbalance in the dataset (e.g. only ~215 Map and ~293 Editorial Cartoon object instances across the whole dataset vs ~27.9k Headline instances). More data or class-balancing/oversampling for those classes would improve them.
The raw per-epoch metrics are stored in val_metrics.json in this repository.
Bias, Risks, and Limitations
- Performance is strongly skewed by class imbalance; rare layout classes are unreliable.
- Detector quality on page layouts/engravings/styles outside the training distribution may degrade.
- Predictions should be reviewed when used in archival/curation workflows (layout regions may span multiple semantic zones; bounding boxes are approximate).
- This is a layout-region detector, not an OCR/HTR model.
Technical Notes
- Architecture:
DetrForObjectDetection, 100 queries, 6 encoder + 6 decoder layers, d_model=256. - Repository files:
config.json,model.safetensors,preprocessor_config.json(DetrImageProcessorFast,size=600),val_metrics.json,train_detr.py(training script).
Acknowledgements
Based on the DETR model (Carion et al., 2020) and the Transformers library. Dataset from BigLam and the Library of Congress newspaper collections.
- Downloads last month
- -
Model tree for harness-race/control-r3
Base model
facebook/detr-resnet-50Dataset used to train harness-race/control-r3
Paper for harness-race/control-r3
Evaluation results
- mAP on biglam/loc_beyond_words (validation)validation set self-reported24.650
- mAP@50 on biglam/loc_beyond_words (validation)validation set self-reported34.760
- mAP@75 on biglam/loc_beyond_words (validation)validation set self-reported28.390
- mAR@100 on biglam/loc_beyond_words (validation)validation set self-reported37.940