Alibaba - H3 3-Step LoRA and Long-Form Video
https://huggingface.co/TaoLiveAIGC/TaoMate-H3
The project reportedly enables the generation of continuous, long-form videos.
Three-step LoRA streaming generation — each small chunk uses three Stage3 denoising intervals.
Someone have tested this LoRa on ComfyUI with a laptop 4090 16GB VRAM https://www.reddit.com/r/StableDiffusion/s/YlDiWx3GMt
I don't suppose that anyone has four H20s lying around that they're not using? I'd be curious to know if there's optimization from this that would work with only one GPU? Their gh repo says they used 8 GPUs and the options for inference have parameters for which GPUs of the system to use and how many GPUs to use. The description says that they generate small chunks of the output using 3 steps, but if you've got one card doing all of the work, then you're not really dividing out the inference? In other words, does it really make sense to chunk out parts of your inference task if it's all done on the same GPU? (except for maybe some kind of tiling?) In any case, I don't see how you're going to benefit, unless you wanted to tile a larger output using groups of inference steps? But then you're not getting the instant gratification of less steps.
I have been using this for 4 steps at lora strength of 0.8 and its really good at simple shots. Most motion however is a problem. dancing the hands turn to blury messes. This lora works really well for upscaling.