CAST: Real-Time Motion Capture for Any Skeleton Topology

Anonymous Authors — under review.

Released checkpoints. Each file stores only the model weights (model_state) and the training step: the optimizer/scheduler state is stripped, no absolute paths are embedded, and the frozen image backbone is not included.

File Backbone Architecture Params Size
CAST_B_dinov2.pt DINOv2 ViT-L/14 CAST-B (8+8) 49,401,614 189 MiB
CAST_B_dinov3.pt DINOv3 ViT-L/16 CAST-B (8+8) 49,401,614 189 MiB
CAST_B_dinov2_no_grace.pt DINOv2 ViT-L/14 CAST-B (8+8) 48,605,961 186 MiB
CAST_B_dinov3_no_grace.pt DINOv3 ViT-L/16 CAST-B (8+8) 48,605,961 186 MiB
CAST_L_dinov3.pt DINOv3 ViT-L/16 CAST-L (16+16) 185,028,878 706 MiB

The _no_grace variants disable the GRACE correction module and are ablations of the full model. Every checkpoint was trained for 6,000 steps.

Load a checkpoint with the matching config and backbone, for example configs/experiment/CAST_B_dinov3.yaml with the DINOv3 backbone:

import torch
from utils.config_utils import instantiate_from_config, load_yaml_config

cfg = load_yaml_config("configs/experiment/CAST_B_dinov3.yaml")
model = instantiate_from_config(cfg["model"]).eval()
state = torch.load("CAST_B_dinov3.pt", map_location="cpu", weights_only=False)
model.load_trainable_state_dict(state["model_state"])

data.image_size must match the backbone: 224 for DINOv2 (14-pixel patches) and 256 for DINOv3 (16-pixel patches); both give a 16x16 token grid.

The backbone repositories and their pretrained weights are not redistributed here. The DINOv2 weights are released under CC-BY-NC-4.0 and the DINOv3 weights follow Meta's own gated licence; see the code repository for instructions.

Code and project page: https://cast-mocap.github.io/

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support