Instructions to use dronefreak/kitti-yolo11n with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ultralytics
How to use dronefreak/kitti-yolo11n with ultralytics:
# Couldn't find a valid YOLO version tag. # Replace XX with the correct version. from ultralytics import YOLOvXX model = YOLOvXX.from_pretrained("dronefreak/kitti-yolo11n") source = 'http://images.cocodataset.org/val2017/000000039769.jpg' model.predict(source=source, save=True) - Notebooks
- Google Colab
- Kaggle
YOLO11n Finetuned on KITTI
Fine-tuned YOLO11n object detector on the KITTI benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.
Usage
Install Dependencies
pip install ultralytics huggingface_hub
Load Model from Hugging Face
from huggingface_hub import hf_hub_download
from ultralytics import YOLO
weights = hf_hub_download(
repo_id="dronefreak/kitti-yolo11n",
filename="best.pt"
)
model = YOLO(weights)
Run Inference
results = model.predict(
source="image.jpg",
conf=0.25
)
results[0].show()
Performance
Evaluated on the KITTI val split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).
| Metric | Score (%) |
|---|---|
| mAP@50 | 38.77 |
| mAP@50-95 | 23.97 |
| Precision | 52.84 |
| Recall | 40.19 |
| F1 Score | 45.65 |
| Parameters | 2.6M |
| FLOPs | 6.6B (at 640 px) |
KITTI Model Zoo
Every model DetectionBench has trained and evaluated on KITTI so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.
| Model | mAP@50 | mAP@50-95 | Precision | Recall |
|---|---|---|---|---|
| YOLO26x | 43.73 | 25.96 | 64.17 | 41.28 |
| YOLO26n | 42.54 | 25.43 | 48.62 | 42.84 |
| YOLO11s | 42.0 | 25.1 | 47.27 | 43.71 |
| YOLOv8s | 41.99 | 25.2 | 49.02 | 43.66 |
| YOLO26m | 41.91 | 25.78 | 63.34 | 39.71 |
| YOLO26s | 41.79 | 26.54 | 59.64 | 42.42 |
| YOLOv9t | 41.6 | 25.73 | 48.3 | 44.2 |
| YOLOv9s | 40.84 | 25.92 | 54.27 | 41.51 |
| YOLOv8m | 40.48 | 25.37 | 50.81 | 39.86 |
| YOLOv8n | 40.11 | 24.75 | 46.05 | 41.85 |
| YOLO11x | 39.18 | 23.6 | 46.28 | 41.69 |
| YOLO11n | 38.77 | 23.97 | 52.84 | 40.19 |
Per-Class Performance
| Class | mAP@50 | mAP@50-95 |
|---|---|---|
| Car | 90.53 | 67.3 |
| Cyclist | 43.88 | 25.04 |
| Misc | 9.56 | 4.77 |
| Pedestrian | 58.64 | 27.81 |
| Person_sitting | 4.19 | 1.28 |
| Tram | 21.32 | 11.5 |
| Truck | 44.22 | 28.73 |
| Van | 37.83 | 25.34 |
Dataset
This model was trained on KITTI. For the full dataset description, provenance, license, and citation, see the dataset card:
https://huggingface.co/datasets/dronefreak/KITTI
Classes
- Car
- Cyclist
- Misc
- Pedestrian
- Person_sitting
- Tram
- Truck
- Van
Training Configuration
| Setting | Value |
|---|---|
| Dataset | KITTI |
| Framework | Ultralytics YOLO |
| Training Toolkit | DetectionBench |
| Epochs (configured max) | 100 |
| Epochs (actually trained) | 43 |
| Early Stopping Patience | 20 |
| Batch Size | auto (Ultralytics AutoBatch) |
| Image Size | 1280 |
| Optimizer | AdamW |
| Initial Learning Rate | 0.001 |
| Seed | 0 |
Repository Contents
best.pt
results.csv
args.yaml
BoxPR_curve.png
BoxF1_curve.png
BoxP_curve.png
BoxR_curve.png
confusion_matrix.png
confusion_matrix_normalized.png
val_batch0_pred.jpg
kitti_yolo11n_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.md
Related Resources
- KITTI dataset card on Hugging Face
- DetectionBench -- reproducible benchmarks for modern object detectors on real-world datasets
- KITTI paper (CVPR 2012, doi:10.1109/CVPR.2012.6248074)
- KITTI Vision Benchmark Suite (official data source)
Training Framework
This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.
Features include:
- A dataset-adapter registry for converting real-world datasets into a canonical format
- Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
- Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
- One-command reproducibility via versioned Hydra configs
If you find this model useful, please consider starring the repository.
Known Limitations
- No official test-set labels: KITTI's real held-out test images have never had public ground truth, so this adapter (following the field-standard Chen et al. 2015 3DOP split) uses train (3,712) / valid (3,769) only -- "valid" is both the early-stopping signal and the split all metrics on this card are computed on, the same convention the wider KITTI detection literature uses.
- Severe class imbalance across 8 classes:
Caris 70.8% of all boxes, whilePerson_sittinghas only 222 instances (0.5%) across all 7,481 images -- its per-class score is measured on very few examples and should be read with caution. - Unusual native aspect ratio: frames are
1242x375 (3.3:1, not the usual 4:3/16:9), so square-letterboxed training/inference wastes canvas on padding above and below the real content; a higher input resolution (see Training Configuration) partly compensates for the resulting loss of effective resolution. - Small dataset for the task's difficulty: only 3,712 training images across a real (non-memorization-prone) street-scene detection task, so absolute scores are lower than on datasets with more training data or an easier task shape.
- Demo banner is not KITTI's own labelled data: the video above uses KITTI's official but unlabelled tracking-benchmark sequences (0000-0028.mp4) for illustration, run through the model at inference time -- it is not part of the train/valid split and the boxes shown are the model's raw predictions, not checked against ground truth.
Citation
If you use this model in your research, please consider citing the dataset and the model architecture:
@inproceedings{geiger2012kitti,
title={Are we ready for autonomous driving? The KITTI vision benchmark suite},
author={Geiger, Andreas and Lenz, Philip and Urtasun, Raquel},
booktitle={2012 IEEE Conference on Computer Vision and Pattern Recognition},
pages={3354--3361},
year={2012},
organization={IEEE},
doi={10.1109/CVPR.2012.6248074}
}
No official YOLO11 research paper has been published by Ultralytics; the most commonly cited independent architectural analysis is used instead:
@article{khanam2024yolov11,
title={YOLOv11: An Overview of the Key Architectural Enhancements},
author={Khanam, Rahima and Hussain, Muhammad},
journal={arXiv preprint arXiv:2410.17725},
year={2024}
}
- Downloads last month
- 32
Model tree for dronefreak/kitti-yolo11n
Base model
Ultralytics/YOLO11Dataset used to train dronefreak/kitti-yolo11n
Collection including dronefreak/kitti-yolo11n
Paper for dronefreak/kitti-yolo11n
Evaluation results
- mAP@50 (val split) on KITTIDetectionBench38.770
- mAP@50-95 (val split) on KITTIDetectionBench23.970
- Precision (val split) on KITTIDetectionBench52.840
- Recall (val split) on KITTIDetectionBench40.190
