Severe visual degradation, ghosting and audio corruption with MiniMax-H3 Turbo LoRA on pruned INT8 ConvRot, especially at 4 steps
Severe visual degradation, ghosting and audio corruption with MiniMax-H3 Turbo LoRA on pruned INT8 ConvRot, especially at 4 steps
Hi, I am currently testing MiniMax-H3-Turbo-LoRA in ComfyUI and I am seeing very severe visual degradation after enabling the Turbo LoRA. Audio quality is also significantly degraded.
Base H3 model
MiniMax\minimax_h3_fl2va_pruned_int8_convrot.safetensors
Turbo LoRA setup
I am loading the Turbo LoRA using the MiniMax-H3 Turbo LoRA node from ComfyUI-MiniMax-H3-Turbo.
Current settings:
LoRA strength: 0.80
low_vram: OFF / bypass (sharp, more VRAM)
For the comparisons below, I kept the prompt, seed, resolution, duration, and other generation conditions the same. The main variables were whether the Turbo LoRA was enabled and the sampling step count.
Test A β Original H3, no Turbo LoRA
Base model:
minimax_h3_fl2va_pruned_int8_convrot.safetensors
Turbo LoRA: OFF
Steps: 20
Result:
Normal video quality
The intended MG / anime graphic style is preserved correctly
Large flat-color areas remain clean
Character silhouettes and edges are stable
No obvious abnormal high-frequency texture
Audio is normal
This is my current quality baseline.
Test B β Turbo LoRA, 8 steps
Turbo LoRA: ON
Strength: 0.80
Steps: 8
The result already shows noticeable degradation:
Style drift from the original clean flat MG look
Unwanted texture appears inside flat-color areas
Character contours become excessively sharpened
Shadows and outlines become much heavier
Some structures are reinterpreted by the model
Colors begin to deviate significantly from the original 20-step result
The output is still recognizable, but the visual quality is clearly worse than the original 20-step H3 generation.
Test C β Turbo LoRA, 4 steps
Turbo LoRA: ON
Strength: 0.80
Steps: 4
At 4 steps the result becomes severely corrupted.
I am seeing:
Strong multi-outline / ghosting artifacts
Repeated silhouettes along motion directions
Exploding high-frequency texture
Screen-door / grid-like artifacts
Extremely aggressive oversharpening
Incorrect texture generated inside originally solid black areas
Strong ringing around white/black edges
Severe blue/yellow oversaturation
Structural degradation of the face, arms, clothing, etc.
An oil-painted / edge-corrosion / incorrect-denoising appearance
This does not look like normal quality loss from reducing the sampling steps.
It looks more like a compatibility or sampling issue involving the Turbo LoRA, sampler, or the quantized/pruned model.
Audio issue
I am also getting significantly degraded audio when using the Turbo workflow.
Symptoms include:
Audible distortion
Increased noise
Significantly worse audio quality than the original H3 workflow
At times it sounds like the audio sampling/decoding path is behaving incorrectly
Using the exact same H3 base model without the Turbo LoRA at 20 steps produces normal audio.
This makes me suspect that the problem is related to the Turbo LoRA / Turbo sampling path rather than the base H3 checkpoint itself.
Important observation
The exact same:
minimax_h3_fl2va_pruned_int8_convrot.safetensors
works correctly when I disable the Turbo LoRA and run the normal 20-step H3 workflow.
My current quality ranking is therefore approximately:
Original H3 @ 20 steps
>
Turbo LoRA @ 8 steps
>>>
Turbo LoRA @ 4 steps
The 4-step result is currently unusable for me.
I have also already reduced the LoRA strength to:
0.80
but the artifacts remain very severe, so this does not appear to be simply an issue caused by using strength = 1.0.
Questions
Has anyone else reproduced similar behavior with this combination?
MiniMax H3
+
pruned_int8_convrot
+
MiniMax-H3 Turbo LoRA
+
ComfyUI-MiniMax-H3-Turbo
In particular:
Has pruned_int8_convrot been fully validated with the current Turbo LoRA?
Is this level of ghosting / oversharpening / screen-door artifact at 4 steps a known issue?
Is another Turbo checkpoint currently recommended for pruned INT8 ConvRot?
With the latest ComfyUI H3 changes, should I use the native H3 audio/video sampling path instead of the custom Turbo sampler?
Could the audio corruption be related to recent H3 sampling changes in ComfyUI?
What LoRA strength / step count / sampler combination is currently considered the most stable for pruned_int8_convrot?
Would testing BF16 or full INT8 ConvRot be useful to determine whether this problem is specific to the pruned model?
I can provide the following if useful:
Full workflow JSON
ComfyUI console logs
Turbo LoRA node screenshots
Sampler / Scheduler screenshots
Same-seed comparison videos for 20 / 8 / 4 steps
Audio comparison samples
Update: 6-step test
I ran another test using the exact same prompt, seed, resolution and workflow, this time with Turbo LoRA at 6 steps.
The 6-step result is less severely corrupted than the 4-step result, but it introduces another very noticeable artifact in large flat-color regions.
In particular, the blue background develops structured contouring / posterization-like patterns, including concentric curved bands, low-frequency tonal contours and other non-random texture patterns.
These structures are not present in the original 20-step H3 output.
My results now show a fairly clear progression:
20 steps / no Turbo LoRA:
clean and stable
8 steps / Turbo:
noticeable style drift and oversharpening
6 steps / Turbo:
structured contouring / posterization-like artifacts
in flat-color regions
4 steps / Turbo:
severe ghosting, screen-door artifacts,
oversharpening and structural corruption
This makes it look like the artifact severity progressively increases as the Turbo sampling step count is reduced, rather than being a single random bad generation.
Base model:
minimax_h3_fl2va_pruned_int8_convrot.safetensors
Turbo LoRA strength:
0.80
I will also inspect/save the raw PNG frame immediately after VAE Decode to rule out normal H.264 compression banding.
just updated the node to address this issue, please try it. thanks for your detailed inspection.
Update: switching to the dedicated MiniMax-H3 Turbo Sampler significantly improved the 4-step result
I ran another controlled test while keeping the same base model, Turbo LoRA, prompt, seed and 4-step configuration:
Base model:
minimax_h3_fl2va_pruned_int8_convrot.safetensors
Turbo LoRA:
ckpt850 EMA
LoRA strength:
0.80
Steps:
4
Previously, I was using the sampling path built into MiniMaxH3_KK_Director, with settings such as:
sampler = euler
scheduler = simple / beta
shift_video = 12
shift_audio = 3
At 4 steps, this still produced major problems:
Contouring / wave-like artifacts in flat-color areas
Degraded character structure and outlines
Some ghosting and excessive sharpening
Severely corrupted audio, essentially unusable
I then replaced that sampling path with the dedicated sampler provided by ComfyUI-MiniMax-H3-Turbo:
MiniMax-H3 Turbo Sampler (4-step)
while keeping the other main settings unchanged.
The result improved very noticeably.
On the video side:
The severe ghosting is mostly gone
Screen-door / exploding high-frequency artifacts are greatly reduced
Large blue flat-color regions are much cleaner
Black/white edges and character silhouettes are significantly more stable
Character structure no longer collapses like it did in my previous 4-step tests
The overall result is much closer to the intended H3 MG / graphic-animation style
The 4-step output still has some loss of detail and quality compared with the original 20-step result, but it has moved from severely corrupted / unusable to generally usable and worth further tuning.
Audio quality also improved noticeably, but it is still not fully fixed.
Compared with the previous regular / KK_Director sampling path, the distortion and noise are reduced. However, the audio is still clearly worse than my normal baseline:
20 steps / No Turbo LoRA
and I would not yet consider it fully production-ready.
So my current result can be summarized as:
4-step + regular / KK_Director sampling:
severe visual artifacts + severely degraded audio
4-step + dedicated MiniMax-H3 Turbo Sampler:
major visual improvement
noticeable audio improvement
but audio quality is still not fully restored
This test suggests that the dedicated Turbo Sampler is very important for 4-step inference, especially for video stability and audio quality.
However, with my current:
pruned_int8_convrot + ckpt850 EMA
setup, there still appears to be a remaining audio-quality issue at 4 steps.
Although there is some loss in both audio and visual quality, the speedup is very substantial, so I think this setup is especially useful for rapid iteration and pre-production validation before generating the final version. Thank you to the author for your work!
My current hardware is an RTX 4070 Ti SUPER 16GB with 32GB of system RAM. If you have any recommendations for a better model variant, LoRA, node combination, or workflow configuration that would work well with this hardware, Iβd be very interested to try them.
Forget about 4 step. 4 step is bad. Minimum is 6 steps
The current 4 steps do not work; at least 6 steps are required. In addition, the duration and resolution need to be controlled. With a duration of 5 seconds and a resolution coefficient of 0.8, performance will return to normal levels.




