DINO_ViTS16_IN1k
DINO helps researchers find visually similar images without needing descriptions or predefined categories.
It can be useful when exploring collections of historical photographs, artworks, or other digitized images.
Intended Use
DINO can be used to find visually related images, compare images, and explore patterns within large collections.
Unlike CLIP, it does not support searching images using text descriptions.
Limitations
Visual similarity does not necessarily mean historical or cultural similarity. Results should be checked against the research context, especially when working with specialist collections.
Technical Details
- Architecture: DINO ViT-S/16
- Original checkpoint: facebook/dino-vits16
- Training data: ImageNet-1K, self-supervised
- Embedding dimensions: 384
- Similarity: Cosine
- Input: RGB images, 224 × 224 pixels
- ONNX model size: Approximately 83 MB
The EIDORA package resizes images to 224 × 224 pixels using bilinear interpolation. This may differ from the original checkpoint's preprocessing.
Reference
Caron et al. (2021). Emerging Properties in Self-Supervised Vision Transformers. Proceedings of the IEEE/CVF International Conference on Computer Vision.
@inproceedings{Caron_2021_ICCV,
author = {Caron, Mathilde and Touvron, Hugo and Misra, Ishan and Jégou, Hervé and Mairal, Julien and Bojanowski, Piotr and Joulin, Armand},
title = {Emerging Properties in Self-Supervised Vision Transformers},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision},
pages = {9650--9660},
year = {2021}
}
- Downloads last month
- 13