Video Classification
Safetensors
lucid
vjepa2

V-JEPA 2 ViT-g/16 (384px)

https://arxiv.org/abs/2506.09985

Lucid port of facebook/vjepa2-vitg-fpc64-384/model.safetensors, converted to Lucid-native safetensors.

Available weights

Tag Params GFLOPs Size Source
FPC64_384 (default) 1034.6M — 7807.77 MB facebook

Usage

import lucid
import lucid.models as models
from lucid.models.weights import VJEPA2ViTGiant384Weights

# default tag
model = models.vjepa2_vit_giant_384(pretrained=True)

# explicit tag (enum or string)
model = models.vjepa2_vit_giant_384(weights=VJEPA2ViTGiant384Weights.FPC64_384)
model = models.vjepa2_vit_giant_384(pretrained="FPC64_384")

# preprocessing travels with the weights
weights = VJEPA2ViTGiant384Weights.FPC64_384
preprocess = weights.transforms()
# The model consumes a decoded (B, T, C, H, W) video tensor.
video = lucid.rand(1, 64, 3, 256, 256)
out = model(video)

Conversion

Converted from facebook/vjepa2-vitg-fpc64-384/model.safetensors via python -m tools.convert_weights vjepa2_vit_giant_384 --tag FPC64_384. Key mapping + numerical parity verified against the source.

License

apache-2.0 — inherited from the original weights.

Citation

Assran et al., "V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning," arXiv:2506.09985, 2025.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for lucid-dl/vjepa2-vitg-384