DirectS2ST
Checkpoints of DirectS2ST, the end-to-end expressive speech-to-speech translation model from Holistic Parallel Supervision for Expressive Speech-to-Speech Translation, trained on VC-Dub (professional dubbing + multilingual voice conversion).
| Folder | Direction |
|---|---|
en-de/ |
English โ German |
en-es/ |
English โ Spanish |
Each folder holds pytorch_model.bin, model_config.json and the
SentencePiece-8k text tokenizer (text_tokenizer/).
Given English source speech, the model predicts all 16 Mimi codec streams of the target speech directly: a frozen w2v-BERT 2.0 encoder, an AR temporal decoder (source text, target text and codec stream c0 with one shared head), and an NAR depth decoder (c1-c15) conditioned on a CAMPPlus speaker prefix and a source codec prompt.
Usage
Code, setup and training: https://github.com/hzted/truly-e2e-expressive-s2st-demo/tree/main/DirectS2ST
hf download TedZhangHao/DirectS2ST --include "en-es/*" --local-dir checkpoints
python inference/infer.py \
--checkpoint-dir checkpoints/en-es \
--campplus-model-root /path/to/seed-vc \
--campplus-checkpoint-path /path/to/campplus_cn_en_common.pt \
--source-wav input_en.wav --output-wav output_es.wav
Demo: https://hzted.github.io/truly-e2e-expressive-s2st-demo/
Third-party components
The checkpoints contain the frozen weights of w2v-BERT 2.0 and Mimi; their licenses apply to those parts. The CAMPPlus speaker encoder (iic/speech_campplus_sv_zh_en_16k-common_advanced) is not included and is loaded at run time.