CAST: Real-Time Motion Capture for Any Skeleton Topology
Anonymous Authors — under review.
Released checkpoints. Each file stores only the model weights (model_state) and
the training step: the optimizer/scheduler state is stripped, no absolute paths
are embedded, and the frozen image backbone is not included.
| File | Backbone | Architecture | Params | Size |
|---|---|---|---|---|
CAST_B_dinov2.pt |
DINOv2 ViT-L/14 | CAST-B (8+8) | 49,401,614 | 189 MiB |
CAST_B_dinov3.pt |
DINOv3 ViT-L/16 | CAST-B (8+8) | 49,401,614 | 189 MiB |
CAST_B_dinov2_no_grace.pt |
DINOv2 ViT-L/14 | CAST-B (8+8) | 48,605,961 | 186 MiB |
CAST_B_dinov3_no_grace.pt |
DINOv3 ViT-L/16 | CAST-B (8+8) | 48,605,961 | 186 MiB |
CAST_L_dinov3.pt |
DINOv3 ViT-L/16 | CAST-L (16+16) | 185,028,878 | 706 MiB |
The _no_grace variants disable the GRACE correction module and are ablations of
the full model. Every checkpoint was trained for 6,000 steps.
Load a checkpoint with the matching config and backbone, for example
configs/experiment/CAST_B_dinov3.yaml with the DINOv3 backbone:
import torch
from utils.config_utils import instantiate_from_config, load_yaml_config
cfg = load_yaml_config("configs/experiment/CAST_B_dinov3.yaml")
model = instantiate_from_config(cfg["model"]).eval()
state = torch.load("CAST_B_dinov3.pt", map_location="cpu", weights_only=False)
model.load_trainable_state_dict(state["model_state"])
data.image_size must match the backbone: 224 for DINOv2 (14-pixel patches) and
256 for DINOv3 (16-pixel patches); both give a 16x16 token grid.
The backbone repositories and their pretrained weights are not redistributed here. The DINOv2 weights are released under CC-BY-NC-4.0 and the DINOv3 weights follow Meta's own gated licence; see the code repository for instructions.
Code and project page: https://cast-mocap.github.io/