RF-DETR Nano Finetuned on BDD100K

Fine-tuned RF-DETR Nano object detector on the BDD100K benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.


Task Framework Base Model
mAP@50 mAP@50:95 Params
License Source

Performance

Metric Score (%)
mAP@50 56.9
mAP@50-95 31.58
Precision 80.68
Recall 64.78
F1 Score 71.86
Parameters 30.5M
FLOPs N/A (not published upstream)

Evaluation Protocol

Metrics reported in this model card are computed on the BDD100K test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).


BDD100K Model Zoo

Every model DetectionBench has trained and evaluated on BDD100K so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.

Model mAP@50 mAP@50-95 Precision Recall
YOLOv8s 57.93 33.25 75.38 51.63
YOLOv10s 57.64 33.34 75.02 52.12
YOLOv11s 57.63 33.1 74.31 52.42
RF-DETR Nano 56.9 31.58 80.68 64.78
YOLOv26n 52.25 29.23 72.56 46.87
YOLOv9t 52.04 29.46 71.34 46.72
YOLOv10n 51.95 29.31 71.58 46.62
YOLOv8n 51.67 29.09 70.95 46.59
YOLOv11n 51.63 29.06 71.68 46.34

Per-Class Performance

Class mAP@50 mAP@50-95
person 61.5 29.93
rider 49.23 25.13
car 78.84 46.87
truck 66.7 47.33
bus 67.21 50.21
train 5.43 3.79
motor 52.81 26.78
bike 52.34 25.74
traffic light 64.87 23.88
traffic sign 70.08 36.18

Evaluation Visualizations

This model was evaluated with Supervision's detection metrics, which report mAP/Precision/Recall directly but don't produce PR-curve, F1-curve, or confusion-matrix plot images the way Ultralytics' validator does. See the Performance table above for Precision/Recall/F1 and the per-class table above for the full per-class mAP breakdown.


Dataset

This model was trained on BDD100K. BDD100K is released under the BDD100K license (non-commercial research and education, registration required, no redistribution), so it is not mirrored on Hugging Face. Download it from the official site (https://www.bdd100k.com/) and see the Citation section below for the dataset's paper.

Classes

  • person
  • rider
  • car
  • truck
  • bus
  • train
  • motor
  • bike
  • traffic light
  • traffic sign

Usage

Install Dependencies

pip install rfdetr huggingface_hub

Load Model from Hugging Face

from huggingface_hub import hf_hub_download
import rfdetr

weights = hf_hub_download(
    repo_id="dronefreak/bdd100k-rfdetr-nano",
    filename="checkpoint_best_total.pth"
)

model = rfdetr.RFDETRNano(pretrain_weights=weights)

Run Inference

detections = model.predict("image.jpg", threshold=0.25)

Training Configuration

Setting Value
Dataset BDD100K
Framework RF-DETR
Training Toolkit DetectionBench
Epochs (configured max) 30
Epochs (actually trained) 29
Early Stopping Patience 8
Batch Size 16
Resolution 576
Optimizer adamw
Learning Rate 0.0001
Seed 42

Repository Contents

checkpoint_best_total.pth
metrics.csv
config.json
bdd100k_rfdetr-nano_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.md

Related Resources


Training Framework

This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.

Features include:

  • A dataset-adapter registry for converting real-world datasets into a canonical format
  • Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
  • Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
  • One-command reproducibility via versioned Hydra configs

If you find this model useful, please consider starring the repository.


Known Limitations

  • Not comparable to the official BDD100K test-server leaderboard: the official test split has no released labels, so the test split here is BDD100K's official validation set (10,000 images) and a seeded 15% slice of the official train set is held out for validation.
  • Severe class imbalance: car (55.4%), traffic sign (18.6%) and traffic light (14.5%) dominate the boxes, while rider (0.4%), motor (0.2%) and especially train (about 150 boxes in the whole dataset) are rare -- per-class accuracy on those classes is measured on very few examples and is close to noise for train.
  • Small objects: the median box covers only 0.09% of the 1280x720 frame, and traffic lights and signs are the smallest and hardest classes (medians of roughly 16 px and 21 px at native resolution), so scores on them depend heavily on input resolution.
  • Detection labels only: BDD100K's lane-marking and drivable-area annotations are dropped, so these models cover the 2D object detection task only.
  • Conditions are not broken down: the images span weather, time-of-day and scene conditions, but scores here are aggregated over all of them, and generalization outside the US road scenes BDD100K covers is untested.
  • Non-commercial data with no redistribution: BDD100K is released under the BDD100K license (non-commercial research and education, registration required), so the dataset is not mirrored on Hugging Face -- obtain it from the official site and check its terms before any use beyond research.

Citation

If you use this model in your research, please consider citing:

  1. The BDD100K dataset (see below)
  2. The original RF-DETR Nano architecture (see below)
  3. The other model architectures shown in the Model Zoo/External Comparison tables above, if you reference their results
  4. DetectionBench, the training/evaluation framework used to produce this checkpoint
@inproceedings{yu2020bdd100k,
  title={BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning},
  author={Yu, Fisher and Chen, Haofeng and Wang, Xin and Xian, Wenqi and Chen, Yingying and Liu, Fangchen and Madhavan, Vashisht and Darrell, Trevor},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
  pages={2636--2645},
  year={2020}
}
@inproceedings{robinson2026rfdetr,
  title     = {RF-DETR: Real-Time Detection Transformer},
  author    = {Robinson, Isaac and Robicheaux, Peter and Popov, Matvei and Ramanan, Deva and Peri, Neehar},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2511.09554}
}

@article{oquab2023dinov2,
  title={DINOv2: Learning Robust Visual Features without Supervision},
  author={Oquab, Maxime and Darcet, Timoth{\'e}e and Moutakanni, Theo and Vo, Huy and Szafraniec, Marc and Khalidov, Vasil and Fernandez, Pierre and Haziza, Daniel and Massa, Francisco and El-Nouby, Alaaeldin and others},
  journal={arXiv preprint arXiv:2304.07193},
  year={2023}
}

Other architectures compared against on BDD100K in this model card:

YOLOv10

@article{wang2024yolov10,
  title={YOLOv10: Real-Time End-to-End Object Detection},
  author={Wang, Ao and Chen, Hui and Liu, Lihao and Chen, Kai and Lin, Zijia and Han, Jungong and Ding, Guiguang},
  journal={arXiv preprint arXiv:2405.14458},
  year={2024}
}

YOLOv11

No official YOLO11 research paper has been published by Ultralytics; the most commonly cited independent architectural analysis is used instead:

@article{khanam2024yolov11,
  title={YOLOv11: An Overview of the Key Architectural Enhancements},
  author={Khanam, Rahima and Hussain, Muhammad},
  journal={arXiv preprint arXiv:2410.17725},
  year={2024}
}

YOLOv26

@article{jocher2026yolo26,
  title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models},
  author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat},
  journal={arXiv preprint arXiv:2606.03748},
  year={2026}
}

YOLOv8

No official YOLOv8 research paper has been published by Ultralytics; this is their own recommended software citation instead:

@software{jocher2023yolov8,
  author = {Glenn Jocher and Ayush Chaurasia and Jing Qiu},
  title = {Ultralytics YOLOv8},
  version = {8.0.0},
  year = {2023},
  url = {https://github.com/ultralytics/ultralytics},
  license = {AGPL-3.0}
}

YOLOv9

@article{wang2024yolov9,
  title={YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information},
  author={Wang, Chien-Yao and Yeh, I-Hau and Liao, Hong-Yuan Mark},
  journal={arXiv preprint arXiv:2402.13616},
  year={2024}
}
@software{Saksena_DetectionBench_2026,
  author = {Saksena, Saumya Kumaar},
  title = {DetectionBench: Reproducible Benchmarks for Modern Object Detectors on Real-World Datasets},
  url = {https://github.com/dronefreak/DetectionBench},
  year = {2026}
}
Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dronefreak/bdd100k-rfdetr-nano

Finetuned
(12)
this model

Collection including dronefreak/bdd100k-rfdetr-nano

Papers for dronefreak/bdd100k-rfdetr-nano

Evaluation results