MiniMax H3 · pruned INT8 ConvRot · input-major

What this repository is

An exact-layout repack of the standard MiniMax H3 FL2VA transformer from Comfy-Org/MiniMax-H3, built for H3ddle, the open-source native macOS app.

The transformer is not a standalone model. It needs the Qwen3-VL-32B INT8 text encoder, video/audio VAEs, tokenizer, and configuration files from the optimized Comfy-Org package.

Input-major checkpoint

minimax_h3_fl2va_pruned_int8_convrot_input_major.safetensors stores all 200 quantized core projections across 50 blocks as [input, output] instead of the source checkpoint's [output, input]. This lets H3ddle's Metal kernel read the weights more efficiently.

No tensor is dequantized and no value or scale is changed. The converter:

  • exactly transposes the four INT8 core projections in every block;
  • copies all 732 non-core tensors byte-for-byte;
  • preserves ConvRot metadata and per-output F32 scales;
  • adds one versioned layout marker that H3ddle validates before creating the Metal pipeline.

The resulting checkpoint is not compatible with runtimes that assume the original output-major Comfy layout. That original remains available from Comfy-Org/MiniMax-H3.

Measured performance versus the regular layout

The complete artifact was checked against the source tensor by tensor: all 200 converted matrices matched the exact transpose, all 732 unchanged tensors were byte-identical, and the native 50-layer output-major/input-major Metal smoke tests produced identical video and audio hashes.

The directly timed A/B used the same layout and kernel on a matched 512x512 Turbo generation on a 32 GiB M1 Pro; it is representative of the projection path, not a separately timed standard-model claim.

comparison regular / baseline optimized path resulting gain
transformer, output-major vs input-major 253.9 s 231.4 s 8.9% (1.10x)
complete generation, same layout A/B 278.7 s 256.5 s 8.0% (1.09x)

The generated images were identical. The standard checkpoint in this repository separately passed a real H3ddle image-quality check before publication. Actual improvement varies with canvas size and storage pressure.

H3ddle package layout

H3ddle uses this FL2VA core in its Standard + Hybrid References package. Prompt-only and keyframe generation run the complete core; ordered references add only the 43.55 MB Ref2VA AdaLN overlay published in PulpCut/MiniMax-H3-Ref2VA-Turbo-INT8-ConvRot, not a second 20.97 GB transformer. Full Ref2VA checkpoints remain available in their source repositories for users who want them.

Source and attribution

The source checkpoint is diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors at revision 014cd40f7e177756c6b2473c0d93b1c89a790dd2, with SHA-256 e889202c41dafb67b10d67b97f0d8541508036a6090af23425a5c2615d03c47a.

Licensing

Derivative of MiniMax H3 weights; the MiniMax H3 Community License Agreement applies. By downloading you agree to its terms.

File inventory

File Bytes SHA-256
minimax_h3_fl2va_pruned_int8_convrot_input_major.safetensors 20,970,379,724 5a0a3e1e73f099680896a98ab418870f27711800ac73b5f16af81724fc7e567a

Reproducibility

The layout is reproducible with Scripts/repack-h3-input-major.py.

Published by PulpCut, whose editor family shares the local-first media-generation stack implemented in H3ddle.

Contact

Open an issue in the H3ddle repository.

Downloads last month
31
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PulpCut/MiniMax-H3-INT8-ConvRot

Finetuned
(81)
this model