thoughts and impressions from using it so far

#38
by m8rr - opened

fl v0.1: Very powerful. It achieves 80% completion in just 1 step. Thanks to this, it cleans up images that got messed up after latent upscale very well. However, it feels too overcooked, and like it has a speed limit, it struggles with fast movements.

rf v0.1, fl 8step v1: I'm not sure, both feel similar. Fast movement is good. However, it requires a lot of steps.

fl 768 v1: Similar to fl v0.1 but a bit weaker. Also, I mainly generate vertical videos, but sometimes it outputs weird videos rotated by 90 degrees.

I am using it like this: rf v0.1 2-3 steps / latent upscale / fl v0.1 3 steps.
Since the rf part is at low resolution, the final speed is the same as standard 4 steps.
Instead of shift_video, I just added a low sigma at the end, which is very effective in preventing blurring.

In FL Unet, it feels too overcooked.
In RF Unet, it feels a bit dark.

Anyway, these are examples.

m8rr changed discussion title from houghts and impressions from using it so far to thoughts and impressions from using it so far

Hey, who's that guy on the left?

seed 1, start with 800x480, then go to 1280x768.
It's worse than lightx2v's 4-step 768 example, right?
Let's just use the 4-step 768.

The one above is rf unet, and this is fl unet.

What is that? What kind of move is that?

What's that? Is that instead of a whip?

Mind sharing a workflow?

Mind sharing a workflow?

It's included in the video, so you can just download it and drag and drop it.
By the way, boxing videos used a slightly different workflow.

this is rf unet

Mind sharing a workflow?

It's included in the video, so you can just download it and drag and drop it.

Awesome, thank you !

https://huggingface.co/lightx2v/Minimax-h3-Turbo/discussions/29#6a7d8d6af138155129ff8f22
I tried making one based on what was posted here. The horizontal video makes the person look too small, so it's not as good as I thought... I guess I really need to put in the time honestly and increase the steps, huh?

rf / fl 5s

rf / fl 10s

rf / fl vertical

Let's try without latent upscale. It's FL unet. Compare this with the FL 10 seconds above.

FL v0.1 / FL v1_768 4 steps. It takes the same amount of time as the 3/3 steps above, but the motion is slower.

RF v0.1 4 steps + FL v0.1 / FL v1_768 1 step, for a total of 5 steps. Generation time is about 20% slower, but the motion is much better. It's better to give up on latent upscale in complex scenes.

Okay, coming back to latent upscale, fl v0.1 is stronger than expected. It seems 2 steps are enough to handle the upscale artifacts. So, it's a total of 6 steps—rf_v0.1 4 / fl_v0.1 2—But the time is much faster than regular 4 steps. The reduction in attack motion is due to the different resolution, But if the higher resolution is an advantage, that could also be considered a side effect of latent upscale. But it's fast, anyway.

How you all getting around the jibberish talking that comes out nearly every shot.

The model makes good video but i cant get any audio with talking to not become a jibber jabber mess.

How you all getting around the jibberish talking that comes out nearly every shot.

It seems helpful to write according to the prompt guide. I'm not really sure about the audio quality itself. Sometimes there's a little noise at the beginning or end, and I'm not sure if it's a side effect of the upscale or an issue with the model itself.

I have tried following every prompt guide i can find. It does not help the fact that so many are throwing their hats into the pile of explanations making an already difficult prompt structure even harder.

Nothing i do fixes the audio, if you use a voice ref for the vocal timbre it talks jibberish when there is no spoken words. You can get lucky with a seed and it will work great... but change seed and you are back to jibberish. I have tried the < d> </ d> i have tried double and single quotes, i have tried inline text with no quotes... I have tried specific prompt styles where the audio is separated from the visual... nothing works.

I have tried many turbo loras, many different model quants... all are the exact same issue. I cant control the audio in ANY meaningful way.

I have noticed you have 5 sec examples. 5 seconds can usually work with a few tries. Try 10 seconds, or even 15. Its infuriating because the model makes amazing videos... i can do skipping, and basketball with just a prompt but the audio is nearly always unusable.

Sign up or log in to comment