face-model โ€” single-class face detector for blurring faces in video

YOLO11 (nano and small) fine-tuned from the COCO-pretrained Ultralytics weights as a one-class face detector on WIDER FACE, plus a video script that detects, tracks and pixelates (or blacks out) faces. Built for recall (privacy), not leaderboard accuracy.

Important licence notes (read before using commercially)

  • The code and weights are derived from Ultralytics YOLO11, which is AGPL-3.0. Using them in a product or network service means complying with AGPL-3.0 or obtaining an Ultralytics enterprise licence.
  • The training data, WIDER FACE, is listed as non-commercial (CC BY-NC-ND 4.0 on its Hugging Face card; the original WIDER FACE terms are for non-commercial research). Do not assume these weights are cleared for commercial use.
  • code/face_id.py (optional --keep mode) expects InsightFace buffalo_l ONNX models, which are non-commercial research only. They are not included in this repo; download them yourself.

Files

File Notes
yolo11n/face_yolo11n.pt PyTorch weights, 5.4 MB
yolo11n/face_yolo11n_fp32.onnx, _fp16.onnx ONNX, fixed 640x640, batch 1
yolo11n/face_yolo11n_int8.onnx INT8 (ONNX Runtime static QDQ). The Detect-head decode nodes are kept in float; plain INT8 returned zero detections
yolo11s/face_yolo11s.pt, _fp32.onnx, _fp16.onnx Larger, more accurate
code/ blur_video.py, train.py, convert_wider_to_yolo.py, eval_model.py, recall_at_conf.py, quantize_int8.py, face_id.py

Results (WIDER FACE val, 3,222 images / 39,112 faces, 640 px, all difficulty levels together)

Model mAP50 mAP50-95
YOLO11n .pt 0.684 0.362
YOLO11n ONNX FP32 0.683 0.362
YOLO11n ONNX FP16 0.682 0.361
YOLO11n ONNX INT8 0.667 0.348
YOLO11s .pt 0.741 0.401
YOLO11s ONNX FP16 0.739 0.399

Recall at confidence 0.25 (IoU >= 0.5), i.e. the fraction of labelled faces that get a box:

Face height (px, images stored at <=640 px) n=39,112 YOLO11n FP16 YOLO11n INT8 YOLO11s FP16
< 10 16,305 33.1% 31.6% 42.1%
10-20 10,521 72.2% 70.7% 79.1%
20-40 7,478 86.6% 85.4% 90.4%
40-80 3,167 93.7% 92.9% 95.3%
> 80 1,641 95.6% 95.0% 97.3%
Overall 61.4% 60.0% 67.9%

WIDER FACE is dominated by tiny faces (median face height about 12 px). Performance on tiny, far-away faces is the main weakness; on faces >= 20 px the nano finds about 90%. No detector guarantees every face is caught in every frame โ€” spot-check any video where missing a face matters.

Not measured: speed on phones or other hardware. Any such numbers you may see elsewhere in this project were estimates.

Training

  • Start: COCO yolo11n.pt / yolo11s.pt, fine-tuned (not from scratch), single class face.
  • Data: WIDER FACE train, converted to YOLO format (code/convert_wider_to_yolo.py): invalid boxes dropped, small faces kept, images downscaled to <= 640 px long side.
  • 640 px, AMP. Nano: 40 epochs. Small: 40 epochs (resumed from a checkpoint at epoch 34 with the schedule shortened to end at 40).
  • Extra augmentation: random motion blur and JPEG compression to imitate video frames (train.py).

Use

pip install ultralytics onnxruntime-gpu lap
python code/blur_video.py in.mp4 out.mp4 --model yolo11n/face_yolo11n_fp16.onnx -n 1            # pixelate
python code/blur_video.py in.mp4 out.mp4 --model yolo11n/face_yolo11n_fp16.onnx --mode black     # solid black

Defaults: detect every 3 frames (-n) with ByteTrack in between, confidence 0.25, 15% box padding, 8 pixel blocks across the shorter side of each face. Use -n 1 for the strictest coverage. Audio is preserved via ffmpeg. With an ONNX model use --device 0 (GPU); --device cpu makes Ultralytics try to pip-install onnxruntime.

--keep ref.jpg [...] leaves one person visible and blurs everyone else using face recognition; every other face stays blurred unless confirmed over several frames. It needs the InsightFace models mentioned above.

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support