ECHO-G models
Generate full-body Unitree G1 motion from speech and text. All models output 39-channel robot motion at 30 FPS.
Code v0.2.0 · Dataset · Inference guide
Models
| Model | Weights | Guide |
|---|---|---|
| Audio + text | best.pt |
Quick start below |
| Audio only | audio_only/best.pt |
Audio-only |
| Text only | text_only/best.pt |
Text-only |
| HumanRetarget | Three checkpoints in human_retarget/ |
Inference only |
Matching configurations are in configs/; the shared FGD encoder is in evaluation/.
Use revision v0.2.0 for reproducible downloads.
Quick start
Use Python 3.10 or 3.11 with a compatible PyTorch installation:
git clone --branch v0.2.0 https://github.com/Mondo-Robotics/ECHO-G.git
cd ECHO-G
python -m pip install -e . huggingface_hub
hf download gaopusen/ECHO-G --revision v0.2.0 --local-dir weights \
--include 'best.pt' 'configs/sgdit_audio_text.yaml' 'evaluation/*' \
'SHA256SUMS' 'LICENSE*' 'NOTICE' 'README.md'
Download and extract the dataset
to data/echo-g, then generate the validation split:
echo-g-sample --checkpoint weights/best.pt \
--config configs/sgdit_audio_text.yaml --data-root data/echo-g \
--output-dir results/audio_text --split val --seeds 0 --device cuda
Predictions are saved in results/audio_text/seed_000/.
Optional file check: (cd weights && sha256sum --ignore-missing -c SHA256SUMS).
Inputs and evaluation
The released dataset includes ready-to-use frozen audio/text features. New inputs support up to 20 seconds / 600 frames and 256 text tokens; word timestamps are relative to the clip start. Longer inputs must be segmented before processing.
Other model variants · Benchmark · MuJoCo visualization
License and attribution
Project weights and documentation: CC BY-NC 4.0. Code and configurations: PolyForm Noncommercial 1.0.0. Third-party materials retain their own terms; see NOTICE.
Copyright (c) 2026 SOAR-LAB, School of Intelligence Science and Technology, Nanjing University. Developed by Dr. Hao Xu's team with sponsorship from Mondo Robotics (妙动科技). Licensing contacts: Dr. Hao Xu and Dr. Shuo Yang. Citation.