breakdown of video guide token at specific resolution ...

#6
by MatzeHali - opened

Trying to remove an actress to separate the background in a 2k plate.
Test runs in low resolution test workflow just fine with a mask input fed into the latent to zero the guide strength within the mask, otherwise Obscura still did actually put out a human in every run. The masked workflow works up towards 1344x704, any higher resolution will lead to a freshly generated video that just follows the text prompt, no apparent elements are still pulled from the reference video.
Do not know if this is a bug or to be expected within LTX2.3, just wanted to let you know.

Revisiting this, it doesn't seem to be a resolution only gate, but the point before Obscura breaks and plain video generation from text prompt starts is dependent on resolution and frame count.
So, the 1344x704 with 41frames will work, equally 768x448 with 113frames, on my machine, running LTX2.3 dev model MLX within ComfyUI an a MacBookPro M5 Max with 128GB unified memory.

After extensive testing, Obscura Remova IC-LoRA seems to break down and stop working above around 5200 latent video tokens for my use case with multiple shots and an actress in the foreground of a handheld camera shot utilizing a guide mask. So, since doing ObscuraRemova removals in chunks usually breaks temporal stability, this leads to needing to lower the resolution to get usable results.
latent_frames = (frames - 1) / 8 + 1
tokens = (width / 32) × (height / 32) × latent_frames

For 113 frames, this resulted in a 768x448 window, for another shot with 145 frames, it was 704x384 producing stable Obscura Remova Outputs.

Sign up or log in to comment