English

SAM’s value in this karaoke maker pipeline...

#6
by kdb8756 - opened

SAM’s value in this karaoke workflow is:

It creates the target vocal stem, which Whisper uses to produce cleaner word timestamps than it would get from the full music mix.
It creates the residual stem, removing or reducing the original singer so you can sing over the instrumental.
It helps identify vocal activity and silence, improving alignment of supplied lyrics to the performance.
It produces the audio (Our Residual without vocals) used in karaoke_singalong.mp4 video
It can isolate different singers or vocal characteristics using the Target Sound Profile.
Original song → SAM vocal/residual separation → Whisper timing → exact lyric correction/alignment → karaoke video

If you already had an excellent instrumental master and accurately timed LRC lyrics, SAM would prov

Sign up or log in to comment