Ref2VA environmental audio feels too quiet in dialogue scenes

#82
by szczypen - opened

Hey, I’m testing H3-Base Ref2VA locally and noticed something with native audio. Sounds directly tied to visible actions work really well for me — footsteps, water, impacts, explosions, stuff like that. But passive environmental sound seems much weaker.

For example, if a scene is clearly happening on a busy street with cars moving around in the background, I’d expect the audio to feel like a real phone recording from that place. Instead I mostly get clean dialogue with a very faint generic ambience underneath.

I tried describing traffic, passing cars, engine/tire noise, putting those sounds directly into the timeline, and even asking for raw phone audio with no noise suppression, but it still stays very quiet compared to the dialogue.

Is this just how the current Ref2VA checkpoint behaves, or is there a better way to prompt this? Also curious if the hosted Context IR handles this kind of “the rest of the world is alive too” ambience differently than a manually written H3 prompt with official guidelines. And is there any meaningful audio behavior difference between Ref2VA and FL2VA here?

Not really looking for an audio-reference workaround, mostly trying to understand how far the native generation can be pushed.

Noticed same thing. I will say this word - seedance head to head comaprasion in seedance I just feel like the video was be recorded on the street as in your example because all noisy sounds are there while H3 just focus too much on dialogue and environmental sounds are (barely) hearable only while there's pause between words in dialogue.

Sign up or log in to comment