We use Fishaudio already and are curious about this

#6
by YellowjacketGames - opened

We make 3d games and use fish for our voiceovers by using voice cloning. what does this directly improve atop the existing fishaudio workflow?

happy to try this on some higher end GPUs if i can understand what i'm comparing against

Audio8 org

Thanks for the question! We are actually building on top of Fish’s DualAR architecture, which we think is a very strong design for voice cloning.
The main improvement is not a completely different workflow, but rather scaling efficiency: we focus on achieving comparable voice quality and cloning capability with a much smaller 0.6B model, while preserving the advantages of the DualAR architecture.
Compared with the existing FishAudio workflow using the 4.5B model, our goal is to provide a much lighter alternative with significantly lower memory usage, easier deployment, and faster inference, while keeping similar voice cloning quality.

could be useful for putting our lower-end GPUs to use around the office!

Right now we are using Fishaudio on a 48gb RTX A6000 card, so we're actually looking for speed more so than size!

Sign up or log in to comment