Jogg-Avatar V2V 1.3B (InfiniteTalk)

Jogg-Avatar V2V 1.3B is an audio-driven video-to-video avatar model based on Wan2.1-T2V-1.3B and the InfiniteTalk approach. It preserves source body, background, and camera motion while regenerating a face synchronized to a driving audio track.

This model card is prepared for the InfiniteTalk checkpoint release. Training, preprocessing, and inference code is available at chanjing-ai/Jogg-Avatar-V2V-1.3B.

The checkpoint directory is expected to contain:

Jogg-Avatar-V2V-1.3B/
|-- model-00001-of-00002.safetensors
|-- model-00002-of-00002.safetensors
|-- model.safetensors.index.json
`-- training_init/audio_proj.safetensors

Wan2.1-T2V-1.3B, TencentGameMate/chinese-wav2vec2-base, and the face preprocessing models are required separately. Review all upstream licenses and terms before use.

Users are responsible for obtaining consent for source videos and voices and for clearly disclosing synthetic media. Do not use this model for impersonation, fraud, harassment, or deceptive content.

Downloads last month
-
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cicada-ai/Jogg-Avatar-V2V-Infinite

Finetuned
(71)
this model