Instructions to use 1simo/recaptcha-classification-57k-224 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ultralytics
How to use 1simo/recaptcha-classification-57k-224 with ultralytics:
# Couldn't find a valid YOLO version tag. # Replace XX with the correct version. from ultralytics import YOLOvXX model = YOLOvXX.from_pretrained("1simo/recaptcha-classification-57k-224") source = 'http://images.cocodataset.org/val2017/000000039769.jpg' model.predict(source=source, save=True) - Notebooks
- Google Colab
- Kaggle
reCAPTCHA tile classifier, 224 px (ONNX)
DannyLuna/recaptcha-classification-57k
(Ultralytics YOLO11x-cls, trained at 640 px) fine-tuned at a 224 px input,
for the Go solver in mmhanda/VisionAIRecaptchaSolver (go-solver/),
which downloads this file on first use and checks its SHA-256.
reCAPTCHA tiles are about 100 px, so 640 px inputs upscale each one 6x. At 224 px the model does about 28x the throughput of the 640 px one on the same GPU (RTX 3060 Laptop: 134 3x3 grids/s on TensorRT fp16, against 4.8 for the 640 px model on CUDA), with click rates within noise of it on held-out data.
Model
- File:
recaptcha_classification_57k_224.onnx, SHA-2568ab029f07247bfe6dd6a15cf527b7968120cdaf267237054defe19efecfda786 - Input
images: float32[batch, 3, 224, 224], RGB in [0, 1]; the shortest edge resized to 224 (bilinear), then a 224x224 centre crop, as Ultralytics classification does. Dynamic batch. - Output
output0: softmax over 14 classes, in this order: Bicycle, Bridge, Bus, Car, Chimney, Crosswalk, Hydrant, Motorcycle, Mountain, Other, Palm, Stair, Tractor, Traffic Light. - The input size is also in the metadata (
imgsz); run the model at 224 px. At any other size its accuracy collapses without an error.
Training
From the 640 px weights, on the dataset's train split at 224 px: AdamW,
lr0 5e-4 with a cosine schedule, 1 warmup epoch, batch 64, early stopping on
the validation split (best epoch 13 of 19, 1.5 h on an RTX 3060).
go-solver/train/train.py reproduces it.
Evaluation
The dataset's validation split leaks: 632 of its 1,474 images are byte-for-byte
copies of training images and 89 more are re-encoded copies. These numbers
come from the 748 that are not (train.py heldout), and from their 191
tile-sized images in particular, which are what a solver classifies. 95%
intervals come from resampling the tiles.
| this model (224 px) | base model (640 px) | |
|---|---|---|
| top-1, tile-sized, background excluded | 91.2% [86-96] | 88.8% [84-94] |
| target tiles over 0.7 (a dynamic click) | 84.8% [78-91] | 86.4% [80-92] |
| background tiles over 0.7 for a target | 5.2% [4.4-6.1] | 5.6% [4.7-6.4] |
| simulated static grids answered exactly | 63.6% [56-73] | 62.4% [55-72] |
Both models share the dataset's weakness: they almost never predict "Other" (background), and about 30% of held-out background tiles score at least 0.7 as Car. The training set's "Other" class has 128 distinct images, and "Other" means "not this challenge's target", so it is noisy. Details in the solver's README, under Accuracy.
License
AGPL-3.0, as Ultralytics licenses models built with its framework (and as the file's own metadata states). The base model and the dataset are published by DannyLuna under MIT.
Model tree for 1simo/recaptcha-classification-57k-224
Base model
DannyLuna/recaptcha-classification-57k