V-JEPA 2.1 predicted-latent to Cosmos adapter

This adapter maps two frozen V-JEPA 2.1 predicted tubelets covering four future frames into one Cosmos CV4 future latent slot. Frames 1-12 are context and frames 13-16 are predicted. Training uses latent L1 plus 0.1 cosine distance; no RGB loss is used. Upstream weights are excluded.

  • V-JEPA checkpoint: https://dl.fbaipublicfiles.com/vjepa2/vjepa2_1_vitg_384.pt
  • Cosmos tokenizer: nvidia/Cosmos-0.1-Tokenizer-CV4x8x8
  • Metrics: metrics/latest.json
Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support