Regarding commercial use case and some doubts

#1
by Asswin04310 - opened

Hi! You previously mentioned that you wouldn't recommend this model for commercial production use. Could you please elaborate on that?

I'm particularly interested in understanding:

  • How natural and human-like does the speech actually sound in practice?
  • The documentation mentions dynamic speaking rate control. How useful is this feature in real-world applications?
  • Do you think the model is still too unstable for production deployment? If so, what are the main issues you've encountered?
  • How well does it perform for real-time streaming TTS in practice? I'm more interested in real-world experience than marketing claims (e.g., latency, stability, interruptions, audio quality under continuous streaming).

I'd really appreciate hearing your practical experience. Thanks!

Sign up or log in to comment