See our collection for all versions of BEiT.

Run BEiT with Keras 3: JAX, PyTorch, or TensorFlow

GitHub Docs Collection

zeromodels/beit-large-patch16-224-pt22k-ft22k

Paper: BEiT: BERT Pre-Training of Image Transformers (arXiv:2106.08254) · HF Papers

BEiT is a ViT-family vision transformer with a per-layer relative position bias, a learnable layer scale on each residual branch, and mean pooling of the patch tokens. Large backbone fine-tuned on ImageNet-22k (21841 classes).

For more details on the model, please go to Microsoft's original model card.

Pure-Keras 3 conversion of microsoft/beit-large-patch16-224-pt22k-ft22k for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX.

This is a image classification checkpoint (BeitImageClassify).

✨ Quick start

import os

os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from PIL import Image
from zeromodels.models.beit import BeitImageClassify, BeitModel, BeitImageProcessor

model = BeitImageClassify.from_weights("zeromodels/beit-large-patch16-224-pt22k-ft22k")
processor = BeitImageProcessor.from_weights("zeromodels/beit-large-patch16-224-pt22k-ft22k")

image = Image.open("your_image.jpg").convert("RGB")
pixels = processor(image)  # resize + normalize (normalization lives in the processor)
logits = model(pixels, training=False)
print(logits.shape)  # (1, num_classes)

# Feature extraction: the backbone without the classifier head
backbone = BeitModel.from_weights("zeromodels/beit-large-patch16-224-pt22k-ft22k", as_backbone=True)
features = backbone(pixels, training=False)

Load any BEiT variant the same way with from_weights("zeromodels/<variant>"):

Variant Hub Task
beit-base-patch16-224 zeromodels/beit-base-patch16-224 image classification
beit-large-patch16-224 zeromodels/beit-large-patch16-224 image classification
beit-large-patch16-512 zeromodels/beit-large-patch16-512 image classification
beit-base-patch16-224-pt22k-ft22k zeromodels/beit-base-patch16-224-pt22k-ft22k image classification
beit-large-patch16-224-pt22k-ft22k zeromodels/beit-large-patch16-224-pt22k-ft22k image classification
beit-base-finetuned-ade-640-640 zeromodels/beit-base-finetuned-ade-640-640 semantic segmentation
beit-large-finetuned-ade-640-640 zeromodels/beit-large-finetuned-ade-640-640 semantic segmentation

Tips

  • Set KERAS_BACKEND before importing Keras / zeromodels.
  • Normalization (0.5/0.5) is baked into the model, so pass raw [0, 255] pixels.
  • Classification uses BeitImageClassify; semantic segmentation uses BeitSemanticSegment and returns logits at a quarter of the input resolution (upsample the argmax map to the input size).
  • BeitModel.from_weights(..., as_backbone=True) returns the per-block token sequences for feature extraction.
  • See BEiT docs and Loading Weights.
  • Community / upstream safetensors still work via the hf: prefix, e.g. BeitImageClassify.from_weights("hf:microsoft/beit-large-patch16-224-pt22k-ft22k").

Special Thanks

A huge thank you to the Microsoft Research BEiT authors for creating and releasing these models.

License: Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zeromodels/beit-large-patch16-224-pt22k-ft22k

Finetuned
(6)
this model

Collection including zeromodels/beit-large-patch16-224-pt22k-ft22k

Paper for zeromodels/beit-large-patch16-224-pt22k-ft22k