DynaConTalk
Released checkpoints and asset libraries of DynaConTalk: Wavelet-Constrained Diffusion for Long-Form and Controllable Holistic Co-Speech 3D Motion.
Code, documentation and the DynaConTalk Studio web UI: https://github.com/zhuyifeiabcd1/DynaConTalk
Project page: https://zhuyifeiabcd1.github.io/DynaConTalk/
Use
Clone the code and run bash start.sh: it downloads this repository into the code folder
(together with everything else the Studio needs) and starts the web UI. To download only these
files:
huggingface-cli download HAJIMIMANBO1/DynaConTalk --local-dir <code folder>
Files
checkpoints/
dynacontalk_edit/ editable body model: speech + keypose + root trajectory + body shape + speaker (epoch 176)
dynacontalk_speech/ speech-only body model: speech + body shape + speaker (epoch 157)
dynacontalk_face/ face model (FLAME expressions): speech + body shape + speaker (epoch 150)
trajectory_bigru/ root translation predicted from the generated body pose (Studio)
assets/
keyposes/ 186 keyposes mined from BEAT2, with covers
trajectories/ 198 root trajectories recorded in BEAT2
identities/ the 25 BEAT2 speakers: speaker id and body shape
seed.npz first frames and fallback text of every generation
Each model folder has model.ckpt (EMA weights; torch.load(..., weights_only=True) reads
it), config.yaml (the training recipe the model is built from) and stats.npz
(normalization statistics of the training data). The speech-only body model and the face model
are the ones evaluated in the paper with the EMAGE metrics on the BEAT2 test split; the Studio
uses the editable body model and the face model.
Notes
- The models were trained on BEAT2 (English, official train split) and the asset libraries were built from BEAT2 test and validation sequences; use them under the BEAT2 license.
- The SMPL-X body model is needed to run the code but is not included (its license does not allow redistribution); see the code README.