MiniMax-H3-MLX-6bit

MLX (Apple Silicon) build of the MiniMax-H3 diffusion transformer, quantized to 6-bit (group size 64).

Powered by MiniMax H3.

These files are modified. The transformer weights have been converted to MLX and quantized; they are not MiniMax's originals. Everything else about the model is unchanged.

What this is

MiniMax-H3 generates synchronized video and audio together. It is not a language model: a 33B diffusion transformer denoises video and audio latents jointly over one packed sequence, conditioned by a frozen Qwen3-VL-32B encoder, with separate video and audio VAEs. Running it needs the pipeline code, not just these weights:

git clone https://github.com/PipeNetwork/minimax-h3-mlx
cd minimax-h3-mlx && pip install -r requirements.txt
python scripts/generate.py "a red fox leaps over a mossy log" -o fox.mp4

This repository holds the transformer only. The VAEs and the text encoder come from the upstream release; the pipeline loads them directly.

Size

on disk 30.29 GB
resident during generation 16.46 GB

The gap is deliberate. ~13B of H3's 33B parameters are the per-block AdaLN projections, whose only input is the timestep embedding. For a fixed sampler schedule every modulation tensor a run needs is precomputed once into a small table, and the projections are then dropped — so they are on disk but never resident. The table scales with step count, not model size: measured at 145 MB for a 9-step schedule and 745 MB for 40 steps, against the 26 GB it replaces.

Those projections are quantized to 8-bit here. That was measured, not assumed: quantizing them shifts the modulation table by 0.25%, an order of magnitude less than the 6-bit core's own velocity error, and takes 12.2 GB off this download. (4-bit AdaLN is measurably worse — 0.77% on the table, 2.8% on its worst tensor — and is not used at any core width.)

How the widths compare

Measured with teacher forcing — one bfloat16 trajectory recorded, each variant re-predicting the velocity at those same latents, so the difference is quantization error alone rather than trajectory divergence. 20 paired observations per variant, aggregated with a paired bootstrap.

bits video rel-L2 [95% CI] audio rel-L2 video cosine
8 0.0329 [0.0277, 0.0381] 0.0130 0.99941
6 0.0611 [0.0501, 0.0728] 0.0274 0.99791
4 0.1649 [0.1324, 0.1971] 0.1016 0.98456
3 0.2842 [0.2362, 0.3358] 0.2341 0.95635

Every interval is disjoint from its neighbours, so the ranking is solid. Two things worth noting: the steepest step is 6 to 4 bits (2.7x), not at the low end; and audio degrades faster in relative terms than video (its share of the error climbs from 0.40x at 8-bit to 0.82x at 3-bit), plausibly because audio is a small fraction of the packed rows and has less redundancy to absorb it.

Why 8, 6 and 4 bits only

Velocity error ranks the widths but does not say where output stops being usable — the scheduler integrates velocity, so per-step error compounds along the trajectory. That has to be generated to be seen. The same prompt, seed and settings were rendered through each checkpoint and compared to bfloat16:

build PSNR vs bf16 correlation outcome
8-bit 27.6 dB 0.959 near-identical
4-bit 22.0 dB 0.854 cooler colour, background artifacting, subject intact
3-bit 16.3 dB 0.740 subject destroyed

At 3 bits the scene is gone — no animal, no log, just a textured field. It is built but not published. Notably it does not degrade by blurring: its per-frame variance rises (54.7 against bfloat16's 37.1) as structure is replaced by high-frequency noise, so a sharpness metric would have scored it as healthy. 2-bit is not published either; extrapolation puts it near 50% velocity error.

6-bit was not rendered separately — it is bracketed by 8-bit and 4-bit, which both pass.

Read this before choosing a quant

MiniMax has not released its sparse-attention implementation, so inference runs dense attention over tens of thousands of rows. On an M3 Ultra a single denoising step costs about 8.8 minutes for a 5-second clip (37,966 packed rows) and 1.04 hours for 15 seconds (109,318 rows).

Quantization does not change that. The bottleneck is attention FLOPs, which quantization does not reduce; the linear layers are ~42% of the work at 5 s and ~20% at 15 s, so a 4-bit build is worth roughly 1.2-1.4x end to end. Choose a quant to fit H3 on your machine, not to make it quick.

Licence

Governed by the MiniMax H3 Community License, a copy of which is included in this repository. It is not an open-source licence. Notably: redistribution must carry the agreement and mark modified files; commercial products above $20M yearly revenue need separate authorization from MiniMax; and the grant is territorially limited (worldwide, excluding the Excluded Territories defined in the agreement). By downloading these weights you accept those terms.

The MLX port code is Apache-2.0 and lives at https://github.com/PipeNetwork/minimax-h3-mlx.

Downloads last month
72
Safetensors
Model size
8B params
Tensor type
F32
·
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pipenetwork/MiniMax-H3-MLX-6bit

Finetuned
(15)
this model

Collection including pipenetwork/MiniMax-H3-MLX-6bit