See our collection for all versions of ViT.

Run ViT with Keras 3: JAX, PyTorch, or TensorFlow

GitHub Docs Collection

zeromodels/vit_tiny_patch16_384_augreg_in21k_ft_in1k

Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (arXiv:2010.11929) · HF Papers

Vision Transformer (ViT) patches an image and runs a transformer encoder. Use ViTImageClassify for logits or ViTModel for tokens / per-block features via as_backbone=True.

For more details on the model, please go to the upstream model card.

Pure-Keras 3 conversion of timm/vit_tiny_patch16_384.augreg_in21k_ft_in1k for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX.

This is an image-classification / backbone checkpoint (ViTImageClassify / ViTModel).

✨ Quick start

import os

os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from PIL import Image
from zeromodels.models.vit import ViTImageClassify, ViTModel, ViTImageProcessor

model = ViTImageClassify.from_weights("zeromodels/vit_tiny_patch16_384_augreg_in21k_ft_in1k")
processor = ViTImageProcessor.from_weights("zeromodels/vit_tiny_patch16_384_augreg_in21k_ft_in1k")

image = Image.open("your_image.jpg").convert("RGB")
pixels = processor(image)  # resize + normalize (normalization lives in the processor)
logits = model(pixels, training=False)
print(logits.shape)  # (1, num_classes)

# Feature extraction: the backbone without the classifier head
backbone = ViTModel.from_weights("zeromodels/vit_tiny_patch16_384_augreg_in21k_ft_in1k", as_backbone=True)
features = backbone(pixels, training=False)

Load any ViT variant the same way with from_weights("zeromodels/<variant>"):

Variant Hub
vit_base_patch16_224_augreg_in1k zeromodels/vit_base_patch16_224_augreg_in1k
vit_base_patch16_224_augreg_in21k zeromodels/vit_base_patch16_224_augreg_in21k
vit_base_patch16_224_augreg_in21k_ft_in1k zeromodels/vit_base_patch16_224_augreg_in21k_ft_in1k
vit_base_patch16_224_orig_in21k_ft_in1k zeromodels/vit_base_patch16_224_orig_in21k_ft_in1k
vit_base_patch16_384_augreg_in1k zeromodels/vit_base_patch16_384_augreg_in1k
vit_base_patch16_384_augreg_in21k_ft_in1k zeromodels/vit_base_patch16_384_augreg_in21k_ft_in1k
vit_base_patch16_384_orig_in21k_ft_in1k zeromodels/vit_base_patch16_384_orig_in21k_ft_in1k
vit_base_patch32_224_augreg_in1k zeromodels/vit_base_patch32_224_augreg_in1k
vit_base_patch32_224_augreg_in21k zeromodels/vit_base_patch32_224_augreg_in21k
vit_base_patch32_224_augreg_in21k_ft_in1k zeromodels/vit_base_patch32_224_augreg_in21k_ft_in1k
vit_base_patch32_384_augreg_in1k zeromodels/vit_base_patch32_384_augreg_in1k
vit_base_patch32_384_augreg_in21k_ft_in1k zeromodels/vit_base_patch32_384_augreg_in21k_ft_in1k
vit_large_patch16_224_augreg_in21k zeromodels/vit_large_patch16_224_augreg_in21k
vit_large_patch16_224_augreg_in21k_ft_in1k zeromodels/vit_large_patch16_224_augreg_in21k_ft_in1k
vit_large_patch16_384_augreg_in21k_ft_in1k zeromodels/vit_large_patch16_384_augreg_in21k_ft_in1k
vit_large_patch32_384_orig_in21k_ft_in1k zeromodels/vit_large_patch32_384_orig_in21k_ft_in1k
vit_small_patch16_224_augreg_in1k zeromodels/vit_small_patch16_224_augreg_in1k
vit_small_patch16_224_augreg_in21k zeromodels/vit_small_patch16_224_augreg_in21k
vit_small_patch16_224_augreg_in21k_ft_in1k zeromodels/vit_small_patch16_224_augreg_in21k_ft_in1k
vit_small_patch16_384_augreg_in1k zeromodels/vit_small_patch16_384_augreg_in1k
vit_small_patch16_384_augreg_in21k_ft_in1k zeromodels/vit_small_patch16_384_augreg_in21k_ft_in1k
vit_small_patch32_224_augreg_in21k zeromodels/vit_small_patch32_224_augreg_in21k
vit_small_patch32_224_augreg_in21k_ft_in1k zeromodels/vit_small_patch32_224_augreg_in21k_ft_in1k
vit_small_patch32_384_augreg_in21k_ft_in1k zeromodels/vit_small_patch32_384_augreg_in21k_ft_in1k
vit_tiny_patch16_224_augreg_in21k zeromodels/vit_tiny_patch16_224_augreg_in21k
vit_tiny_patch16_224_augreg_in21k_ft_in1k zeromodels/vit_tiny_patch16_224_augreg_in21k_ft_in1k
vit_tiny_patch16_384_augreg_in21k_ft_in1k zeromodels/vit_tiny_patch16_384_augreg_in21k_ft_in1k

Tips

  • Set KERAS_BACKEND before importing Keras / zeromodels.
  • ViTImageClassify returns class logits; ViTModel returns features (as_backbone=True for multi-scale stages).
  • See docs and Loading Weights.
  • Upstream / timm checkpoints: ViTImageClassify.from_weights("hf:timm/vit_tiny_patch16_384.augreg_in21k_ft_in1k").

Special Thanks

A huge thank you to the ViT authors and the timm / Hub communities for creating and releasing these models.

License: see YAML license (usually matches the upstream checkpoint).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zeromodels/vit_tiny_patch16_384_augreg_in21k_ft_in1k

Finetuned
(1)
this model

Collection including zeromodels/vit_tiny_patch16_384_augreg_in21k_ft_in1k

Paper for zeromodels/vit_tiny_patch16_384_augreg_in21k_ft_in1k