YOLOv8m Finetuned on GC10-DET

Fine-tuned YOLOv8m object detector on the GC10-DET benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.

GC10-DET Detection Demo


Task Framework Base Model
mAP@50 mAP@50:95 Params
License Source

Usage

Install Dependencies

pip install ultralytics huggingface_hub

Load Model from Hugging Face

from huggingface_hub import hf_hub_download
from ultralytics import YOLO

weights = hf_hub_download(
    repo_id="dronefreak/gc10det-yolov8m",
    filename="best.pt"
)

model = YOLO(weights)

Run Inference

results = model.predict(
    source="image.jpg",
    conf=0.25
)

results[0].show()

Performance

Evaluated on the GC10-DET test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).

Metric Score (%)
mAP@50 71.88
mAP@50-95 38.8
Precision 69.77
Recall 70.84
F1 Score 70.3
Parameters 25.9M
FLOPs 78.9B (at 640 px)

GC10-DET Model Zoo

Every model DetectionBench has trained and evaluated on GC10-DET so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.

Model mAP@50 mAP@50-95 Precision Recall
RF-DETR Small 76.07 42.51 87.86 65.03
RF-DETR Medium 75.93 41.93 78.25 67.76
YOLO26s 75.77 38.15 77.16 74.07
YOLO26n 74.25 38.31 80.5 67.99
YOLO26m 73.97 36.94 75.7 68.19
YOLOv8n 73.25 38.74 67.99 70.87
YOLO11s 72.54 35.07 72.39 66.8
YOLOv8s 72.54 37.79 78.54 65.23
YOLOv8m 71.88 38.8 69.77 70.84
YOLO11n 70.44 40.09 78.93 62.64
RF-DETR Nano 70.17 38.06 77.08 71.04

Per-Class Performance

Class mAP@50 mAP@50-95
crease 30.27 14.18
crescent_gap 91.7 55.7
inclusion 28.56 9.52
oil_spot 47.6 20.49
punching_hole 89.52 50.32
rolled_pit 99.5 79.6
silk_spot 58.12 21.88
waist_folding 94.5 50.04
water_spot 96.01 56.75
welding_line 83.0 29.49

Normalized Confusion Matrix


Dataset

This model was trained on GC10-DET. For the full dataset description, provenance, license, and citation, see the dataset card:

https://huggingface.co/datasets/dronefreak/GC10-DET

Classes

  • crease
  • crescent_gap
  • inclusion
  • oil_spot
  • punching_hole
  • rolled_pit
  • silk_spot
  • waist_folding
  • water_spot
  • welding_line

Training Configuration

Setting Value
Dataset GC10-DET
Framework Ultralytics YOLO
Training Toolkit DetectionBench
Epochs (configured max) 180
Epochs (actually trained) 104
Early Stopping Patience 25
Batch Size auto (Ultralytics AutoBatch)
Image Size 640
Optimizer AdamW
Initial Learning Rate 0.001
Seed 0

Repository Contents

best.pt
results.csv
args.yaml
BoxPR_curve.png
BoxF1_curve.png
BoxP_curve.png
BoxR_curve.png
confusion_matrix.png
confusion_matrix_normalized.png
val_batch0_pred.jpg
gc10det_yolov8m_showcase.jpg
README.md

Related Resources


Training Framework

This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.

Features include:

  • A dataset-adapter registry for converting real-world datasets into a canonical format
  • Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
  • Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
  • One-command reproducibility via versioned Hydra configs

If you find this model useful, please consider starring the repository.


Known Limitations

  • No official split: GC10-DET's paper defines no train/valid/test division, so this adapter creates a deterministic seeded 80/10/10 split over the sorted-then-shuffled image list -- results are not directly comparable to a paper that uses a different split.
  • Small dataset: only 1,840 training images (2,300 total) across 10 classes, so absolute scores are more sensitive to the specific split than on the project's larger datasets, and per-class scores on the rarest classes are noisy.
  • Severe class imbalance: silk_spot is 24.8% of all boxes, while crease has only 74 instances (2.1%) across the whole dataset -- its per-class score is measured on very few examples and should be read with caution.
  • Sparse, mostly single-defect images: 1.55 instances/image on average (median 1, max 11), unlike the crowded-scene datasets (PKLot/VisDrone/UAVDT) -- this is a localization task on large, easy-to-see boxes (median 3.48% of image area) rather than a small-object or dense-detection problem.
  • Different visual domain: GC10-DET is grayscale industrial line-scan imagery of rolled steel surfaces, not a natural-scene photo -- the first industrial-inspection dataset in DetectionBench, so these results say nothing about how these checkpoints would perform on outdoor/natural-scene detection or vice versa.

Citation

If you use this model in your research, please consider citing the dataset and the model architecture:

@article{lv2020deep,
  title = {Deep Metallic Surface Defect Detection: The New Benchmark and Detection Network},
  author = {Lv, Xiaoming and Duan, Fajie and Jiang, Jia-jia and Fu, Xiao and Gan, Lin},
  journal = {Sensors},
  volume = {20},
  number = {6},
  pages = {1562},
  year = {2020},
  publisher = {MDPI},
  doi = {10.3390/s20061562}
}
No official YOLOv8 research paper has been published by Ultralytics; this is their own recommended software citation instead:

@software{jocher2023yolov8,
  author = {Glenn Jocher and Ayush Chaurasia and Jing Qiu},
  title = {Ultralytics YOLOv8},
  version = {8.0.0},
  year = {2023},
  url = {https://github.com/ultralytics/ultralytics},
  license = {AGPL-3.0}
}
Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dronefreak/gc10det-yolov8m

Finetuned
(224)
this model

Dataset used to train dronefreak/gc10det-yolov8m

Collection including dronefreak/gc10det-yolov8m

Evaluation results