MiniMax-H3 RefMod

Portable reference-conditioning files for MiniMax-H3 ref2va, as Modular Diffusers blocks. This is the diffusers-modular equivalent of the ComfyUI "RefMod": encode a reference once, save the result to a small .safetensors, and inject it straight back on later requests instead of re-encoding the media.

What it is

A ref2va request conditions on an ordered list of references. MiniMaxH3Ref2VAReferenceEncoderStep turns their pixels into the condition_latents the transformer prepends to its packed sequence. That VAE encode is deterministic (the posterior is sampled under a fixed keyframe_encode_seed) and independent of the prompt, so it is pure, reusable work. RefMod caches it.

Two blocks (both in block.py):

  • MiniMaxH3SaveRefModStep - serialize condition_latents (+ audio_condition_latents) to a .safetensors.
  • MiniMaxH3LoadRefModStep - read one back into condition_latents / audio_condition_latents, dropping in where MiniMaxH3Ref2VAReferenceEncoderStep would run.

The default block of this repo (its auto_map) is MiniMaxH3LoadRefModStep.

What it carries, and what it does not

  • It carries the VAE condition latents: one (1, C, T, H, W) tensor per image/video reference, plus one (num_audio_latents * 2, audio_channels) tensor per reference soundtrack. These are stored exactly as the live encoder leaves them (normalized, fp16-valued, in float32), so a save/load round-trip is bitwise-lossless.
  • It does not carry the Qwen3-VL conditioning. In MiniMax-H3 a reference also appears to the text encoder as a <Picture i> / <Video k> vision block, and that path reads the reference pixels and is entangled with the prompt (one conditioner call over the whole presentation), so it cannot be a prompt-independent per-identity file. This is the same trade the ComfyUI node makes, and it is why the file is small rather than the size of the media.

So there are two ways to use a RefMod:

  1. Lossless encode cache - keep passing the reference media, and swap the VAE encode for a file load. The Qwen3-VL vision block still runs, and generation is bitwise-identical to the live path. Pure speed/convenience.
  2. Media-free identity - pass the RefMod with no references. The vision block is gone from the presentation and the request conditions on the prepended latent rows alone. This is the portable "ship a small file, reuse across prompts" use.

Usage

Create a RefMod

MiniMaxH3LoadRefModStep is this repo's default block, so it is the one from_pretrained returns. MiniMaxH3SaveRefModStep also lives in block.py; grab it by downloading the module.

import importlib.util
import torch
from huggingface_hub import hf_hub_download
from diffusers.modular_pipelines.minimax_h3 import MiniMaxH3ImageReference
from diffusers.modular_pipelines.minimax_h3.encoders import MiniMaxH3Ref2VAReferenceEncoderStep

MODEL = "MiniMaxAI/MiniMax-H3"  # or your own checkpoint

# 1. Encode the reference once through the real VAE path.
encoder = MiniMaxH3Ref2VAReferenceEncoderStep().init_pipeline(MODEL)
encoder.load_components(dtype=torch.bfloat16)
refs = [MiniMaxH3ImageReference.from_file("subject.png")]
out = encoder(normalized_references=refs, output=["condition_latents", "audio_condition_latents"])

# 2. Save it. The Save block ships in this repo's block.py.
path = hf_hub_download("diffusers-modular/minimax-h3-refmod", "block.py")
spec = importlib.util.spec_from_file_location("refmod_block", path)
refmod_block = importlib.util.module_from_spec(spec)
spec.loader.exec_module(refmod_block)

save_pipe = refmod_block.MiniMaxH3SaveRefModStep().init_pipeline(MODEL)
save_pipe.load_components(dtype=torch.bfloat16)
save_pipe(
    condition_latents=out["condition_latents"],
    audio_condition_latents=out["audio_condition_latents"],
    normalized_references=refs,
    refmod_path="subject.refmod.safetensors",
    output=["refmod_path"],
)

Use a RefMod (lossless encode cache)

Swap the loader in for the vae_encoder step and keep passing the media:

import torch
from diffusers.modular_pipelines import ModularPipelineBlocks
from diffusers.modular_pipelines.minimax_h3 import MiniMaxH3Blocks, MiniMaxH3ImageReference

load_block = ModularPipelineBlocks.from_pretrained(
    "diffusers-modular/minimax-h3-refmod", trust_remote_code=True
)

blocks = MiniMaxH3Blocks()
blocks.sub_blocks["vae_encoder"] = load_block          # file load instead of VAE encode
pipe = blocks.init_pipeline("MiniMaxAI/MiniMax-H3")
pipe.load_components(dtype=torch.bfloat16)             # loads both transformer partitions; see the MiniMaxH3Blocks docs

video = pipe(
    prompt="the subject walking through a neon-lit street at night",
    references=[MiniMaxH3ImageReference.from_file("subject.png")],
    refmod_path="subject.refmod.safetensors",
    num_frames=124,
    num_inference_steps=50,
    output_type="pil",
)

File format

safetensors with:

  • tensors video.0, video.1, ... (one per image/video reference) and audio.0, audio.1, ... (one per soundtrack), in packed order.
  • a JSON metadata header: format, version, num_video, num_audio, video_shapes, audio_shapes, and reference_kinds.

Acknowledgements

This is a Modular Diffusers port of the ComfyUI "RefMod" flow:

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support