VideoMAE_Base_K400

VideoMAE_Base_K400 is an EIDORA video-embedding adaptation of the pretrained MCG-NJU/videomae-base checkpoint.

Instead of returning the masked-reconstruction output used during VideoMAE pretraining, this package exposes a 768-dimensional representation obtained by averaging the final VideoMAE encoder tokens and applying L2 normalization. The resulting embeddings are intended to be compared with cosine similarity.

Recommended Use

This model is suitable when a spatiotemporal visual representation of short video clips is useful for analysis or comparison.

  • Clip-level visual feature extraction and video-similarity or retrieval experiments.
  • Clustering and exploratory comparison of video segments.
  • Comparing a transformer-based video representation with other video encoders on the same collection.
  • Exploratory analysis of film, broadcast, archival and other video collections when the representation is evaluated against the needs of the specific collection.

VideoMAE representations have been used in video understanding research and have also appeared in computational studies involving film analysis and video retrieval. The studies cited below use different downstream configurations or tasks, so they should be understood as evidence for the broader use of VideoMAE representations rather than validation of this exact EIDORA embedding.

Limitations

  • The upstream checkpoint was pretrained on Kinetics-400, which is centred on human actions. Its representation can therefore reflect the visual and activity distribution of that training data.
  • This package accepts exactly 16 frames per clip. Video decoding and selection of those frames happen outside the ONNX graph, so the sampling strategy can affect the resulting embedding.
  • The package produces a clip-level representation and does not directly provide a single representation for an arbitrarily long video. Longer material should be segmented or sampled before embedding.
  • This package uses the visual video stream only. Audio, subtitles, transcripts and other modalities are not included.
  • The representation was not developed specifically for archival film, newsreels or cultural collections. Its suitability should be evaluated on the target collection.
  • The upstream checkpoint is distributed under CC-BY-NC-4.0, which includes a non-commercial restriction.

Input

  • Modality: video
  • Color space: RGB
  • Number of frames: 16
  • Model input size per frame: 224 x 224
  • Runtime input name: pixel_values
  • Tensor layout: NTCHW
  • Tensor dtype: float32
  • Input value range: [0, 1]

The runtime tensor has the form:

pixel_values: float32 [batch, 16, 3, 224, 224]

Output

The model returns:

embedding: float32 [batch, 768]

The representation is produced from the final VideoMAE encoder hidden states. For a 16-frame 224 x 224 input, the encoder produces 1568 spatiotemporal tokens with 768 features per token.

EIDORA averages those final-layer token representations across the spatiotemporal token dimension and applies L2 normalization so that cosine similarity can be used directly.

The masked-reconstruction output of the original pretraining model is not exposed by this package.

Preprocessing

The preprocessing follows the processor configuration associated with MCG-NJU/videomae-base.

Before the ONNX model:

  1. Decode the video and convert frames to RGB.
  2. Select exactly 16 input frames.
  3. Resize the shorter frame side to 224 pixels using bilinear interpolation.
  4. Center crop each frame to 224 x 224.
  5. Rescale pixel values to [0, 1].

Video decoding and temporal frame selection happen outside the ONNX graph. The package does not define a single sampling policy for longer videos, so a consistent frame-selection strategy should be used when embeddings need to be comparable across a collection.

The following channel normalization is included inside the ONNX graph:

  • Mean: [0.485, 0.456, 0.406]
  • Standard deviation: [0.229, 0.224, 0.225]

This division is important for reproducing the EIDORA embedding: decoding, RGB conversion, frame selection, resizing, cropping and rescaling happen before the model, while mean/std normalization is part of the ONNX graph.

Model Source and Architecture

Exact converted checkpoint

This package was converted from:

MCG-NJU/videomae-base

Pinned Hugging Face revision:

dc740ceda42fce44faed2ea03c6d447db72f6af9

Source weight artifact:

model.safetensors

Source weight SHA256:

bc053ca2840a038b1068269a4eec06ca569689e9a1ed9376a5b2b8a111be5290

Upstream checkpoint:

https://huggingface.co/MCG-NJU/videomae-base

Official VideoMAE repository:

https://github.com/MCG-NJU/VideoMAE

Transformers implementation:

https://github.com/huggingface/transformers/tree/main/src/transformers/models/videomae

Architecture provenance

VideoMAE was introduced by Zhan Tong, Yibing Song, Jue Wang and Limin Wang in VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training (2022).

https://proceedings.neurips.cc/paper_files/paper/2022/hash/416f9cb3276121c42eebb86352a4354a-Abstract-Conference.html

VideoMAE extends masked autoencoding to video by combining a Vision Transformer-style encoder with spatiotemporal tube masking during self-supervised pretraining.

The base encoder represented by the converted checkpoint uses:

  • 12 transformer encoder layers
  • hidden size of 768
  • 12 attention heads
  • intermediate size of 3072
  • spatial patch size of 16 x 16
  • temporal tubelet size of 2
  • 16 input frames
  • input frame resolution of 224 x 224

For the standard input, this produces 1568 spatiotemporal encoder tokens, each with 768 features.

EIDORA exposes a 768-dimensional representation by averaging the final-layer encoder tokens and L2-normalizing the result. This pooling and normalization are an EIDORA embedding adaptation of the pretrained VideoMAE encoder rather than a separately released checkpoint from the original authors.

Training Data and Checkpoint Provenance

The upstream MCG-NJU/videomae-base repository identifies this checkpoint as a base-sized VideoMAE model pretrained on Kinetics-400 for 1600 epochs in a self-supervised manner.

Kinetics-400 is a large-scale video dataset organised around 400 human-action classes. For background on the dataset, see:

Joao Carreira and Andrew Zisserman (2017), Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.

https://openaccess.thecvf.com/content_cvpr_2017/html/Carreira_Quo_Vadis_Action_CVPR_2017_paper.html

The checkpoint converted here is the pretraining-only MCG-NJU/videomae-base checkpoint rather than a separately fine-tuned video-classification checkpoint.

For checkpoint provenance, this card follows the exact upstream repository, pinned revision and source-weight hash listed above.

Research Context

VideoMAE representations for video analysis

VideoMAE uses masked video modelling for self-supervised pretraining. The original work combines a high masking ratio with tube masking to learn spatiotemporal video representations from the remaining visible content.

Tong et al. evaluated VideoMAE across established video benchmarks including Kinetics-400, Something-Something V2, UCF101 and HMDB51.

Zhan Tong, Yibing Song, Jue Wang and Limin Wang (2022), VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

https://doi.org/10.52202/068431-0732

Use in film and digital-collection research

Tseng et al. (2026) include a VideoMAE-based approach in a computational humanities framework for analysing film and video. Their work uses VideoMAE for camera-setting estimation as part of a wider multimodal analysis pipeline. Their trained configuration and downstream task are not identical to the unchanged EIDORA embedding provided here.

Chiao-I Tseng et al. (2026), From shots to narratives: Expanding multimodal approaches to filmic storytelling in the digital humanities.

https://doi.org/10.1017/chr.2026.10024

Abdari, Falcon and Serra (2025) evaluate VideoMAE as one of several visual encoders for video-to-museum retrieval in the AgriMus project. Their study provides an example of VideoMAE features being evaluated in a digital-collection retrieval setting, but again does not validate this exact EIDORA embedding.

Ali Abdari, Alex Falcon and Giuseppe Serra (2025), Searching Agricultural Learning Experiences in the Metaverse via Textual and Visual Queries Within the AgriMus project.

https://doi.org/10.1007/s00799-025-00433-9

These studies motivate including VideoMAE as a spatiotemporal representation for comparison in film, video and collection-oriented analysis, but they do not imply that this exact checkpoint or embedding is optimal for every collection.

Attribution and Licensing

Original VideoMAE authors: Zhan Tong, Yibing Song, Jue Wang and Limin Wang.

The converted checkpoint is distributed upstream as MCG-NJU/videomae-base under the Creative Commons Attribution-NonCommercial 4.0 International license (CC-BY-NC-4.0).

The official VideoMAE repository states that the majority of the project is released under CC-BY-NC-4.0, with some incorporated components subject to separate license terms.

The upstream license should be reviewed before use or redistribution, particularly because it includes a non-commercial restriction.

EIDORA provides the ONNX conversion and embedding adaptation and does not claim authorship of the original VideoMAE architecture, implementation, pretrained checkpoint or Kinetics-400 training data.

References

  1. Tong, Z., Song, Y., Wang, J., & Wang, L. (2022). VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training. Advances in Neural Information Processing Systems, 35, 10078-10093.
  2. Carreira, J., & Zisserman, A. (2017). Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.
  3. Abdari, A., Falcon, A., & Serra, G. (2025). Searching Agricultural Learning Experiences in the Metaverse via Textual and Visual Queries Within the AgriMus project. International Journal on Digital Libraries, 26, 20.
  4. Tseng, C.-I. et al. (2026). From shots to narratives: Expanding multimodal approaches to filmic storytelling in the digital humanities. Computational Humanities Research, 2, e5.

Package Information

  • Package version: 0.1.0
  • ONNX opset: 17
Downloads last month
94
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including EIDORA/VideoMAE_Base_K400