You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

CARI4D CoCoNet Native MHR Overview

Description

CARI4D CoCoNet Native Momentum Human Rig (MHR) refines one human's native MHR parameters and one rigid object's pose over video and predicts one contact logit per hand. It supports 4D human-object reconstruction, motion analysis, visualization, data curation, and robotics perception without direct autonomous actuation.

CARI4D CoCoNet Native MHR was developed by NVIDIA Research.

This model is ready for commercial or non-commercial use.

License/Terms of Use

GOVERNING DOWNLOAD TERMS: Use of the model is governed by the NVIDIA Open Model Agreement.

Deployment Geography

Global.

Use Case

For developers and researchers building commercial or noncommercial computer-vision systems for 3D/4D human-object pose estimation, motion analysis, visualization, data curation, and robotics perception. The model is not intended for direct autonomous actuation or life-critical decisions.

Expected Release Date, Platform, and URL

Field Value
Expected Release Date Hugging Face 08/30/2026 via https://huggingface.co/nvidia/cari4d_commercial
Release Platform Hugging Face
Release URL https://huggingface.co/nvidia/cari4d_commercial

Model Architecture

Architecture Type: Transformer.

Network Architecture: Dual observed/initialized branches use a frozen Vision Transformer (ViT)-B/14 red-green-blue (RGB) encoder from DINOv2 (self-distillation with no labels, version 2) and a DINOv2 ViT-S/14-initialized six-channel three-dimensional Cartesian (XYZ)/mask encoder that is fine-tuned for this model. Fused features pass through two spatial-temporal attention blocks and per-frame object-pose, native-MHR, and hand-contact heads.

Number of Model Parameters: 231M measured from the step-84,000 model state after excluding registered buffers.

Input

Input Type(s): Video.

Input Format(s): MPEG-4 (MP4) red-green-blue (RGB) source video; Tensor after input materialization.

Input Parameters: Video - N-dimensional [batch, 96, channels, 224, 224]; geometry/conditioning - N-dimensional per-frame tensors.

Other Properties Related to Input: Each sample contains two float32 tensors shaped [batch, 96, 9, 224, 224]: an observed branch and an initialized-pose branch. Each branch contains RGB, normalized per-frame MHR-root-joint-centered camera-space XYZ, and human, visible-object, and full-object masks. Native MHR and rigid-object initializations provide additional conditioning.

Output

Output Type(s): Tabular.

Output Format: Tensor.

Output Parameters: Tabular output - N-dimensional per-frame tensors for object pose, native-MHR parameter blocks, and hand contact.

Other Properties Related to Output: Per-frame outputs contain rigid-object rotation and translation residuals; native-MHR global rotation, translation, 260-dimensional body pose, 108-dimensional hand, and 45-dimensional shape residuals; and two hand-contact logits. Unsupervised output blocks are frozen by the checkpoint-bound supervision contract. Inference bundles serialize the tensor dictionary as Python pickle files.

Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

Software Integration

Runtime Engine(s): PyTorch.

Supported Hardware Microarchitecture Compatibility:

  • NVIDIA Ampere

Validated Hardware: NVIDIA A100.

Supported Operating System(s): Linux.

Integration into an AI system requires use-case-specific data and unit/system testing to verify technical, functional, safety, ethical, and legal requirements before deployment.

Model Version

Field Value
Version 2026-08-25-09-35-57, step 200,000
Training configuration learning/configs/mhr-daniel-commercial-moge2-behave79-val-fp16.yml
Status Training completed

Training, Testing, and Evaluation Datasets

Training Dataset

Field Value
Link FORM-HOI
Data Modality Video, RGB, depth, masks, calibration, and 3D human/object annotations
Video Training Data Size Less than 10,000 hours; 2,126 internal Daniel human-object interaction sequences with four camera streams
Data Collection Method Hybrid: manually collected and automatic/sensors
Labeling Method Hybrid: automatic/sensors, automated processing, and manually reviewed annotations
Properties Count: 2,126 sequences and four camera streams per sequence. Modalities: RGB video, sensor/monocular depth, masks, calibration, and 3D annotations. Content: people performing object interactions, including personal data. Linguistic characteristics: Not Applicable. The internal Daniel corpus is not distributed by this repository.

Testing Dataset

Field Value
Data Collection Method Not Applicable
Labeling Method Not Applicable
Properties Count: 0. Modalities, content nature, and linguistic characteristics: Not Applicable. No separate testing dataset is defined.

Evaluation Dataset

Field Value
Link Demo instructions; published CARI4D demo bundle
Data Modality RGB video and preprocessed inference inputs, including human/object masks, estimated human initialization, and reconstructed textured object meshes
Data Collection Method Hybrid
Labeling Method Not Applicable
Properties Count: two demonstration inputs: one self-contained BEHAVE video and one in-the-wild gas-cylinder example whose source RGB video is obtained separately as documented in the README. The archive's masks, estimated human initialization, and reconstructed object meshes are inference inputs rather than ground-truth labels. No ground-truth human pose, object pose, contact, or quantitative evaluation labels are included; the separately published behave-test-gt.zip archive is not part of this evaluation dataset. This dataset is used for end-to-end pipeline execution and qualitative output inspection; no quantitative model-quality metric is computed from it. Linguistic characteristics: Not Applicable.

Inference

Acceleration Engine: PyTorch with CUDA.

Hardware Requirements (GPU Architecture, Model): NVIDIA Ampere A100 GPU validated; other targets require separate validation.

Test Hardware: NVIDIA Ampere A100 GPU.

Ethical Considerations

NVIDIA believes Trustworthy AI is a shared responsibility and has established policies and practices to support development for a wide array of AI applications. Developers must verify that this model meets requirements for the relevant industry and use case and addresses foreseeable misuse.

Users must have proper rights and permissions for all input image and video content and must apply appropriate controls for personal data, health information, and intellectual property.

For more detail, see the Model Card++ Bias, Explainability, Privacy, and Safety & Security subcards below. Report model quality, risk, security vulnerabilities, or NVIDIA AI concerns through the NVIDIA security and AI concerns form.

Bias Subcard

Field Response
Participation considerations from adversely impacted groups and protected classes in model design and testing No documented participation. Demographic and interaction coverage is uncharacterized, and the two published demos do not establish broad deployment generalization.
Measures taken to mitigate against unwanted bias No demographic-specific mitigation is currently verified.
Bias Metric (If Measured) Not measured.

Explainability Subcard

Field Response
Intended Task/Domain Computer vision: 3D human-object pose and contact estimation.
Model Type Spatiotemporal Transformer.
Intended Users Developers and researchers building commercial or noncommercial human-object reconstruction, motion-analysis, visualization, data-curation, and robotics-perception systems.
Output Tabular Tensor: per-frame rigid-object pose residuals, native-MHR parameter residuals, and two hand-contact logits.
Describe how the model works The model compares observed RGB/XYZ/mask features with rendered initialization features, fuses them across space and time, and predicts per-frame corrections and contact logits.
Adversely impacted groups tested for comparable outcomes Age, disability, ethnicity, gender, race, skin color, and body-morphology groups may be adversely impacted; comparable outcomes have not been established.
Technical Limitations and Mitigation Performance can degrade with occlusion, poor masks or depth, inaccurate initialization, unseen objects/interactions, motion blur, poor lighting, or domain shift. Contact represents proximity, not force or physical feasibility. Use overlays and quantitative metrics to inspect outputs before downstream use.
Verified to meet prescribed NVIDIA quality standards Yes
Performance Metrics Validation loss; object translation error; symmetry-aware rotation geodesic error; native-MHR vertex and joint errors.
Potential Known Risks Incorrect pose/contact estimates, privacy exposure from person videos or reconstructed motion, and unsafe downstream decisions if outputs are used without independent validation.
Terms of Use/Licensing GOVERNING DOWNLOAD TERMS: Use of the model is governed by the NVIDIA Open Model Agreement.

Privacy Subcard

Field Response
Generatable or reverse engineerable personal data? Yes; reconstructed body geometry and motion may relate to identifiable people, although the model does not synthesize source imagery.
Personal data used to create this model? Yes
Personal data context Training and evaluation videos depict people.
Categories of personal data used Visual appearance, body geometry, pose/motion, and environment context associated with people.
Was consent obtained for personal data? Yes
Are data-subject rights supported? Yes
Was data collected by NVIDIA? Yes; Daniel was collected internally and BEHAVE is externally sourced.
Is data minimized to what is required? Yes; model training uses task-required cropped/masked RGB, depth, calibration, and annotations, and this repository does not distribute the internal videos.
Can data subjects access applicable disclosures? Yes.
How often is the dataset reviewed? Before Every Release.
Was data from user interactions with the AI model used to train the model? No.
Is there provenance for all datasets used in training? Yes; source identities and preprocessing manifests are tracked.
Does data labeling comply with privacy laws? Yes
Can data-subject correction or deletion requests be applied? Yes
Applicable Privacy Policy NVIDIA Privacy Policy

Safety & Security Subcard

Field Response
Model Application Field(s) Other - Computer Vision: 3D human-object pose estimation, motion analysis, visualization, data curation, and robotics perception.
Describe the life-critical impact (if present) Not Applicable; the model is not intended for direct autonomous actuation or life-critical decisions.
Use Case Restrictions GOVERNING DOWNLOAD TERMS: Use of the model is governed by the NVIDIA Open Model Agreement.
Model and dataset restrictions Apply least-privilege access to person videos and annotations, honor dataset and third-party license restrictions, and independently validate outputs before deployment.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support