Unstable turbo lora
Hello, I'm using this turbo lora of the minimax_h3_turbo_v4_step600.safetensors and sometimes the generation time is around 18 minutes to generate a 6 second video. I have the sage attention node with the sage compiler false (because when I use it on true it's slower) and I'm using 8 steps. The issue I'm having is that if I press run one time I will get great time, but the next time it will delay for over 40 minutes using the same prompt and configuration. I tried the EMA version and it makes it even slower.
What could be the reason that this is happening? I have everything updated. My VRAM isn't going to the max to think that my memory is getting full. But my concern is that the lora sometimes works properly and sometimes it doesn't.
I like this lora so much when it works. Thank you very much for your replies.
I think this is most likely due to cpu offload or disk swap?
I will add to the issue that I found that it tends to add things to the scene that weren't meant by the prompt nor references. Like randomly giving characters some jacket or scarf etc.
In comparision v1 850 steps doesn't have this issue.
EMA vs Non EMA which is better?
I think this is most likely due to cpu offload or disk swap?
I don't think that's the problem. I've been checking if I get some sort of error but I don't. At first I was using the rgthree Power Lora but I decided to change it for the normal Lora Loader. Then I saw the Minimax Turbo Lora node but it doesn't get passed from step 0. The other two get pass through but they are very slow. Right now I will test the minimax_h3_turbo_v4_step600_ema.safetensors to see if it works fine. It's weird because sometimes it works fast and sometimes it's very slow. I turned off sage attention because I made a test with the normal Lora Loader and got my results in 18 minutes which isn't bad considering that it takes around 45 minutes with the normal model without the turbo lora
I will add to the issue that I found that it tends to add things to the scene that weren't meant by the prompt nor references. Like randomly giving characters some jacket or scarf etc.
In comparision v1 850 steps doesn't have this issue.
I haven't find this problem. I guess it's a prompting problem. I use ChatGPT with the prompting guide in order to get good prompts and it works pretty fine
EMA vs Non EMA which is better?
larryvrh recommends the EMA, but I've got my best results with the non EMA. But my issue is that sometimes the lora doesn't speed up the process. Most of the time it starts with good timing, but after step 3/8 it gives me times over 50 minutes or 1 hour. Then I have to deselect and select the model, lora, etc in order to make it work. But when it works this lora is a total beast. I really love it.
Issue with the latest update of the ComfyUI‑MiniMax‑H3‑Turbo plugin
After updating to the newest version of the MiniMax‑H3‑Turbo LoRA node, I encountered the following problems:
- With
low_vramset to off (default), there is a severe VRAM usage issue. - With
low_vramset to on, the image quality degrades significantly when using theminimax_h3_turbo_v4_step600.safetensorsLoRA.
Temporary workaround
You should delete the existing MiniMax‑H3‑Turbo LoRA node from your workflow, then re‑add it until the low_vram option appears. Once it does, enable it (on). This restores the Turbo acceleration to the normal speed of the previous version. However, the downside is that you still cannot use minimax_h3_turbo_v4_step600.safetensors without quality loss.
I hope the developers can look into this issue as soon as possible. Thank you!
Issue with the latest update of the ComfyUI‑MiniMax‑H3‑Turbo plugin
After updating to the newest version of the MiniMax‑H3‑Turbo LoRA node, I encountered the following problems:
- With
low_vramset to off (default), there is a severe VRAM usage issue.- With
low_vramset to on, the image quality degrades significantly when using theminimax_h3_turbo_v4_step600.safetensorsLoRA.Temporary workaround
You should delete the existing MiniMax‑H3‑Turbo LoRA node from your workflow, then re‑add it until thelow_vramoption appears. Once it does, enable it (on). This restores the Turbo acceleration to the normal speed of the previous version. However, the downside is that you still cannot useminimax_h3_turbo_v4_step600.safetensorswithout quality loss.I hope the developers can look into this issue as soon as possible. Thank you!
Thank you for the information. I will make some tests to see if it improves speed. When you say it reduces the quality is it too much? Does the minimax_h3_turbo_v4_step600_ema.safetensors also get quality loss? I'm making some test bypassing the sage attention node but still I can't get the speed to improve. Most of the times it starts very fast, but when I reaches to step 3 the time remaing jumps from 18 minutes to over 1 hour remaining
I think I was able to make it work. I made a couple of tests and here i post the results. Note that Tests 4 and 6 where the best. I provide the screenshots of the nodes that I have connected on my workflow.
Test 1) MiniMax-H3 Turbo Lora Node (steps600_ema)(low_vram active) ---- Sage Attention (Bypass) ---- MiniMax H3 Low VRAM Attention (head chunks 4) ---- MiniMax H3 Chunk FeedForward (chunks 2, seq_threshold 4096) ---- Prompt executed in 00:22:35
Test 2) MiniMax-H3 Turbo Lora Node (steps600_ema)(low_vram bypass) ---- Sage Attention (compile false) ---- MiniMax H3 Low VRAM Attention (head chunks 4) ---- MiniMax H3 Chunk FeedForward (chunks 2, seq_threshold 4096) ---- Process interrupted in 00:29:01 (0 steps)
Test 3) MiniMax-H3 Turbo Lora Node (steps600_ema)(low_vram bypass) ---- Sage Attention (compile false) ---- MiniMax H3 Low VRAM Attention (head chunks 4) ---- MiniMax H3 Chunk FeedForward (chunks 2, seq_threshold 4096) ---- Process interrupted at 1/8 [08:59<1:02:53, 539.05s/it]
Test 4) MiniMax-H3 Turbo Lora Node (steps600_ema)(low_vram active) ---- Sage Attention (compile false) ---- MiniMax H3 Low VRAM Attention (head chunks 4) ---- MiniMax H3 Chunk FeedForward (chunks 2, seq_threshold 4096) ---- Prompt executed in 00:17:42
Test 5) MiniMax-H3 Turbo Lora Node (steps600_ema)(low_vram active) ---- Sage Attention (Bypass) ---- MiniMax H3 Low VRAM Attention (bypass) ---- MiniMax H3 Chunk FeedForward (bypass) ---- Prompt executed in 00:50:35
Test 6) MiniMax-H3 Turbo Lora Node (steps600_ema)(low_vram active) ---- Sage Attention (compile true) ---- MiniMax H3 Low VRAM Attention (head chunks 4) ---- MiniMax H3 Chunk FeedForward (chunks 2, seq_threshold 4096) ---- Prompt executed in 00:16:37
My setup:
Intel(R) Core(TM) i9-14900HX (2.20 GHz)
RAM: 96.0 GB
NVIDIA GeForce RTX 4060 Laptop GPU (8 GB)
Intel(R) UHD Graphics (128 MB)
Result:
Result:
I use beta because I like the quality it provides. I will keep making tests. If I find something that makes the generation more efficient I will let you know. If you find something that improves the efficiency please let me know.

