Having Audio issue on this workflow only
Hi, wondering if you may be able to help. This workflow is so fast compared to other workflows, however the audio is all like wobbly and scrapy, and any dialogue I give it cuts off a bit and even tries to rapidly say it. I have other workflows that I use the same model, vaes and encoders but only this workflow seems to get this issue.
Would there be something I can adjust to fix? I have tried different settings like without any extra loras, more steps, But in all cases it seems to have the same issue
Hi, wondering if you may be able to help. This workflow is so fast compared to other workflows, however the audio is all like wobbly and scrapy, and any dialogue I give it cuts off a bit and even tries to rapidly say it. I have other workflows that I use the same model, vaes and encoders but only this workflow seems to get this issue.
Would there be something I can adjust to fix? I have tried different settings like without any extra loras, more steps, But in all cases it seems to have the same issue
Are you using all the default files and settings listed from the workflow?
Yes all default, and then tried to toggle off turbo lora, and add more steps but still the same issue.
I am on latest comfyUI desktop as well (v0.32.0+10) I might try switching to stable 0.32
Yes all default, and then tried to toggle off turbo lora, and add more steps but still the same issue.
I am on latest comfyUI desktop as well (v0.32.0+10) I might try switching to stable 0.32
Hmmm that has to be something on your install end. There is a shift node for audio but if it's happening without turbo too... how many steps and what resolution?
When turbo is removed i try with 15 steps, and I keep this settings for resolution:
Not sure if sageattention version matters but I have
Which seems to match the python, cuda, and pytorch versions that this instance has.
I'm generating some example videos to show the audio difference now, will link them in next comment
We're not using sage anymore. Using comfy attn. Try 20 and see if it's normal. If not only thing i can really think of is upgrading pytorch to something newer
Thanks, Will try and might try on a fresh instance to see if it makes a difference
EDIT: At 20 steps the audio seems better, just takes longer to generate similar to the other workflows (Editing Cause its limiting how many posts I can make)
Thanks, Will try and might try on a fresh instance to see if it makes a difference
EDIT: At 20 steps the audio seems better, just takes longer to generate similar to the other workflows (Editing Cause its limiting how many posts I can make)
yea i notice the audio tends to echo and distort when more than one person is talking. I haven't tried going to 20 steps, still at 10 step. Will report back once I do a 20 step version.
[Update] ok back from tests. I'm still getting echo distortions when there's more than one character. You guys can try a prompt like "In the other room can hear guards arguing behind the walls" and at same time "Solid Snake in the main room is quietly talking on his radio".
The guards should end up echoing and sound distorted.
I wasn't having issues with other workflow like this one. One thing to note, they were using sage attention and other speed up nodes.
https://civitai.red/models/2835250/minimax-h3-ultra-fastest-workflow-or-6gb-vram-16gb-ram

