Image Classification
ONNX
ultralytics
recaptcha
yolo11

reCAPTCHA tile classifier, 224 px (ONNX)

DannyLuna/recaptcha-classification-57k (Ultralytics YOLO11x-cls, trained at 640 px) fine-tuned at a 224 px input, for the Go solver in mmhanda/VisionAIRecaptchaSolver (go-solver/), which downloads this file on first use and checks its SHA-256.

reCAPTCHA tiles are about 100 px, so 640 px inputs upscale each one 6x. At 224 px the model does about 28x the throughput of the 640 px one on the same GPU (RTX 3060 Laptop: 134 3x3 grids/s on TensorRT fp16, against 4.8 for the 640 px model on CUDA), with click rates within noise of it on held-out data.

Model

  • File: recaptcha_classification_57k_224.onnx, SHA-256 8ab029f07247bfe6dd6a15cf527b7968120cdaf267237054defe19efecfda786
  • Input images: float32 [batch, 3, 224, 224], RGB in [0, 1]; the shortest edge resized to 224 (bilinear), then a 224x224 centre crop, as Ultralytics classification does. Dynamic batch.
  • Output output0: softmax over 14 classes, in this order: Bicycle, Bridge, Bus, Car, Chimney, Crosswalk, Hydrant, Motorcycle, Mountain, Other, Palm, Stair, Tractor, Traffic Light.
  • The input size is also in the metadata (imgsz); run the model at 224 px. At any other size its accuracy collapses without an error.

Training

From the 640 px weights, on the dataset's train split at 224 px: AdamW, lr0 5e-4 with a cosine schedule, 1 warmup epoch, batch 64, early stopping on the validation split (best epoch 13 of 19, 1.5 h on an RTX 3060). go-solver/train/train.py reproduces it.

Evaluation

The dataset's validation split leaks: 632 of its 1,474 images are byte-for-byte copies of training images and 89 more are re-encoded copies. These numbers come from the 748 that are not (train.py heldout), and from their 191 tile-sized images in particular, which are what a solver classifies. 95% intervals come from resampling the tiles.

this model (224 px) base model (640 px)
top-1, tile-sized, background excluded 91.2% [86-96] 88.8% [84-94]
target tiles over 0.7 (a dynamic click) 84.8% [78-91] 86.4% [80-92]
background tiles over 0.7 for a target 5.2% [4.4-6.1] 5.6% [4.7-6.4]
simulated static grids answered exactly 63.6% [56-73] 62.4% [55-72]

Both models share the dataset's weakness: they almost never predict "Other" (background), and about 30% of held-out background tiles score at least 0.7 as Car. The training set's "Other" class has 128 distinct images, and "Other" means "not this challenge's target", so it is noisy. Details in the solver's README, under Accuracy.

License

AGPL-3.0, as Ultralytics licenses models built with its framework (and as the file's own metadata states). The base model and the dataset are published by DannyLuna under MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 1simo/recaptcha-classification-57k-224

Quantized
(1)
this model

Dataset used to train 1simo/recaptcha-classification-57k-224