Having Audio issue on this workflow only

#4
by MrDel - opened

Hi, wondering if you may be able to help. This workflow is so fast compared to other workflows, however the audio is all like wobbly and scrapy, and any dialogue I give it cuts off a bit and even tries to rapidly say it. I have other workflows that I use the same model, vaes and encoders but only this workflow seems to get this issue.
Would there be something I can adjust to fix? I have tried different settings like without any extra loras, more steps, But in all cases it seems to have the same issue

Hi, wondering if you may be able to help. This workflow is so fast compared to other workflows, however the audio is all like wobbly and scrapy, and any dialogue I give it cuts off a bit and even tries to rapidly say it. I have other workflows that I use the same model, vaes and encoders but only this workflow seems to get this issue.
Would there be something I can adjust to fix? I have tried different settings like without any extra loras, more steps, But in all cases it seems to have the same issue

Are you using all the default files and settings listed from the workflow?

Yes all default, and then tried to toggle off turbo lora, and add more steps but still the same issue.
I am on latest comfyUI desktop as well (v0.32.0+10) I might try switching to stable 0.32

Is the same on stable 0.32, and for reference this is my settings:

Screenshot 2026-08-13 222311

Yes all default, and then tried to toggle off turbo lora, and add more steps but still the same issue.
I am on latest comfyUI desktop as well (v0.32.0+10) I might try switching to stable 0.32

Hmmm that has to be something on your install end. There is a shift node for audio but if it's happening without turbo too... how many steps and what resolution?

When turbo is removed i try with 15 steps, and I keep this settings for resolution:

Screenshot 2026-08-13 224210

Not sure if sageattention version matters but I have
Screenshot 2026-08-13 224257
Which seems to match the python, cuda, and pytorch versions that this instance has.

I'm generating some example videos to show the audio difference now, will link them in next comment

We're not using sage anymore. Using comfy attn. Try 20 and see if it's normal. If not only thing i can really think of is upgrading pytorch to something newer

Thanks, Will try and might try on a fresh instance to see if it makes a difference

EDIT: At 20 steps the audio seems better, just takes longer to generate similar to the other workflows (Editing Cause its limiting how many posts I can make)

Thanks, Will try and might try on a fresh instance to see if it makes a difference

EDIT: At 20 steps the audio seems better, just takes longer to generate similar to the other workflows (Editing Cause its limiting how many posts I can make)

yea i notice the audio tends to echo and distort when more than one person is talking. I haven't tried going to 20 steps, still at 10 step. Will report back once I do a 20 step version.

[Update] ok back from tests. I'm still getting echo distortions when there's more than one character. You guys can try a prompt like "In the other room can hear guards arguing behind the walls" and at same time "Solid Snake in the main room is quietly talking on his radio".

The guards should end up echoing and sound distorted.

I wasn't having issues with other workflow like this one. One thing to note, they were using sage attention and other speed up nodes.
https://civitai.red/models/2835250/minimax-h3-ultra-fastest-workflow-or-6gb-vram-16gb-ram

Sign up or log in to comment