V-JEPA 2.1 predicted-latent to Cosmos adapter
This adapter maps two frozen V-JEPA 2.1 predicted tubelets covering four future frames into one Cosmos CV4 future latent slot. Frames 1-12 are context and frames 13-16 are predicted. Training uses latent L1 plus 0.1 cosine distance; no RGB loss is used. Upstream weights are excluded.
- V-JEPA checkpoint:
https://dl.fbaipublicfiles.com/vjepa2/vjepa2_1_vitg_384.pt - Cosmos tokenizer:
nvidia/Cosmos-0.1-Tokenizer-CV4x8x8 - Metrics:
metrics/latest.json
- Downloads last month
- 17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support