MiniMax H3 video VAE int8 convrot and 8-step Turbo LoRA

#16
by LabMike3D - opened

This time, I'm using the fresh, new experimental VAE from Kijai: minimax_h3_video_vae_int8_convrot.safetensors.
I also used the Turbo custom node by larryvrh ComfyUI-MiniMax-H3-Turbo. I did a quick test with the 8-step Turbo LoRA and the new VAE, and the results are better in terms of speed with slightly improved quality.

Everything was tested on an RTX 3060 12GB. Rendering a 5-second video at 864x480 resolution with Turbo at 8 steps took 4.5 minutes.

It's not as fast as LTX, but it's doable. This is just the first test, and I need more time to benchmark everything properly. This includes testing the audio and pushing MiniMax H3 with the custom prompts I use as benchmarks for every AI video model.

Cheers!

Capture

This time, I'm using the fresh, new experimental VAE from Kijai: minimax_h3_video_vae_int8_convrot.safetensors.
I also used the Turbo custom node by larryvrh ComfyUI-MiniMax-H3-Turbo. I did a quick test with the 8-step Turbo LoRA and the new VAE, and the results are better in terms of speed with slightly improved quality.

Everything was tested on an RTX 3060 12GB. Rendering a 5-second video at 864x480 resolution with Turbo at 8 steps took 4.5 minutes.

It's not as fast as LTX, but it's doable. This is just the first test, and I need more time to benchmark everything properly. This includes testing the audio and pushing MiniMax H3 with the custom prompts I use as benchmarks for every AI video model.

Cheers!

Capture

i have the same gpu (RTX 3060 12GB), but only 16gb ram, could you share with me your workflow? the speed is pretty impressive for the quality i'm seeing.

a link to 8step turbo lora? I have only 4 step lora. or is it the same?

a link to 8step turbo lora? I have only 4 step lora. or is it the same?

i didn't know an 8 step lora existed.

a link to 8step turbo lora? I have only 4 step lora. or is it the same?

oh never mind, i think he used the same lora but with 8 steps intead of 4, i'm doing the same, and sometimes i use 10, the more steps you put, the better the results generally (from my testing and understanding so far).

also a side note: the longer the video, the more of your prompt will appear and the more details you'll get since it gives the model enough time to show you what you're prompting, at least for I2V for me, i haven't experimented on RF2V yet, only once, it was great, but it kinda made the photo of the character i put in the loader image different, and by different i mean the same person but AI generated, didn't look real enough.

Why am I getting pure black videos after using minimax_h3_video_vae_int8_convrot.safetensors?

Why am I getting pure black videos after using minimax_h3_video_vae_int8_convrot.safetensors?

Update comfyui, Kijai's fix for comfyui to use this int8 vae is merged now.
image

I noticed that pairing minimax_h3_turbo_4step_ckpt850.safetensors with 4-step generation performs even better than using minimax_h3_turbo_4step.safetensors with 8step.

Please share the workflow. I think I understood but I'm confused by the wording of it maybe.

Sign up or log in to comment