DirectS2ST

Checkpoints of DirectS2ST, the end-to-end expressive speech-to-speech translation model from Holistic Parallel Supervision for Expressive Speech-to-Speech Translation, trained on VC-Dub (professional dubbing + multilingual voice conversion).

Folder Direction
en-de/ English โ†’ German
en-es/ English โ†’ Spanish

Each folder holds pytorch_model.bin, model_config.json and the SentencePiece-8k text tokenizer (text_tokenizer/).

Given English source speech, the model predicts all 16 Mimi codec streams of the target speech directly: a frozen w2v-BERT 2.0 encoder, an AR temporal decoder (source text, target text and codec stream c0 with one shared head), and an NAR depth decoder (c1-c15) conditioned on a CAMPPlus speaker prefix and a source codec prompt.

Usage

Code, setup and training: https://github.com/hzted/truly-e2e-expressive-s2st-demo/tree/main/DirectS2ST

hf download TedZhangHao/DirectS2ST --include "en-es/*" --local-dir checkpoints
python inference/infer.py \
  --checkpoint-dir checkpoints/en-es \
  --campplus-model-root /path/to/seed-vc \
  --campplus-checkpoint-path /path/to/campplus_cn_en_common.pt \
  --source-wav input_en.wav --output-wav output_es.wav

Demo: https://hzted.github.io/truly-e2e-expressive-s2st-demo/

Third-party components

The checkpoints contain the frozen weights of w2v-BERT 2.0 and Mimi; their licenses apply to those parts. The CAMPPlus speaker encoder (iic/speech_campplus_sv_zh_en_16k-common_advanced) is not included and is loaded at run time.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support