Joycent trained with WhisAID Medium GRL accent embeddings

This is a Joycent Mandarin accent TTS acoustic model trained using accent embeddings extracted by walston/whisaid-medium-grl. The released checkpoint is epoch 100.

Download

from huggingface_hub import hf_hub_download

checkpoint_path = hf_hub_download(
    repo_id="walston/joycent-medium-grl",
    filename="grad_100.pt",
)

Pass the downloaded checkpoint to joycent/inference_joycent.py with the --acoustic-checkpoint argument. Full synthesis also requires the Joycent vocoder and reference-audio feature extraction dependencies described in the Joycent repository.

Checkpoint

  • Epoch: 100
  • Acoustic model: Joycent / Grad-TTS
  • Accent embedding model: WhisAID Whisper Medium GRL (lambda 0.05)
  • Accent embedding dimension: 256

Citation

@misc{wang2026joycentdiffusionbasedaccenttts,
      title={Joycent: Diffusion-based Accent TTS without Accented Phone Prediction},
      author={Xintong Wang and Ye Wang},
      year={2026},
      eprint={2606.16417},
      archivePrefix={arXiv},
      primaryClass={cs.SD},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including walston/joycent-medium-grl

Paper for walston/joycent-medium-grl