HybridEmo

HybridEmo is an instruction-following multi-emotion text-to-speech model developed by post-training CosyVoice 3. It supports sequential emotion trajectories and simultaneous emotion blending within a single utterance.

Resources

Usage

Please refer to the GitHub repository for environment setup and inference instructions.

Citation

If you use HybridEmo in your research, please cite our paper:

@misc{zhou2026sequentialtrajectoriessimultaneousblending,
  title={Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS},
  author={Yan Zhou and Yun Hong and Yang Feng},
  year={2026},
  eprint={2608.30325},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2608.30325},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ICTNLP/HybridEmo

Quantized
(12)
this model

Collection including ICTNLP/HybridEmo

Paper for ICTNLP/HybridEmo