MiniMax H3 Pruned GGUF

This repository (Abiray/MiniMax-H3-Pruned-GGUF) provides pruned and quantized GGUF weights for the MiniMax H3 omni-modal generative model. MiniMax H3 is designed for unified multimodal context processing, capable of generating synchronized high-definition video and 32 kHz stereo audio from text, image, audio, and video inputs.


🌟 Key Highlights

  • VRAM Efficiency: Pruned architecture compressed down to 8.9 GB – 21.6 GB, bringing MiniMax H3 execution to consumer-tier GPUs.
  • Synchronized Omni-Modal Output: Simultaneous generation of video (24 FPS) and native stereo audio (32 kHz).
  • Native ComfyUI Support: Directly compatible with standard ComfyUI-GGUF workflows using the native MiniMax backend.
  • Dual Pipeline Variants: Full quant suites for both FL2VA (First/Last Frame) and Ref2VA (Omni-Reference) modes.

πŸ“‚ Repository Weights & Quantization Breakdown

Note: For optimal performance, Q4_K_M is recommended for 16 GB GPUs, while Q5_K_M is recommended for GPUs with 24 GB VRAM.

FL2VA Models (First-and-Last-Frame Mode)

Quantization File Name Size Target Hardware / Use Case
Q3_K_M MiniMax-H3-FL2VA-Pruned-Q3_K_M.gguf 8.9 GB Low VRAM (~12 GB GPUs)
Q4_K_M MiniMax-H3-FL2VA-Pruned-Q4_K_M.gguf 11.6 GB Recommended Balance (16 GB GPUs)
Q4_K_S MiniMax-H3-FL2VA-Pruned-Q4_K_S.gguf 11.6 GB Compact 4-bit K-Quant
Q5_K_M MiniMax-H3-FL2VA-Pruned-Q5_K_M.gguf 14.1 GB High Quality Balance (24 GB GPUs)
Q5_K_S MiniMax-H3-FL2VA-Pruned-Q5_K_S.gguf 14.1 GB Compact 5-bit K-Quant
Q6_K MiniMax-H3-FL2VA-Pruned-Q6_K.gguf 16.7 GB Near-Lossless Precision
Q8_0 MiniMax-H3-FL2VA-Pruned-Q8_0.gguf 21.6 GB Maximum Fidelity / Reference

Ref2VA Models (Omni-Reference Mode)

Quantization File Name Size Target Hardware / Use Case
Q3_K_M MiniMax-H3-Ref2VA-Pruned-Q3_K_M.gguf 8.9 GB Low VRAM (~12 GB GPUs)
Q4_K_M MiniMax-H3-Ref2VA-Pruned-Q4_K_M.gguf 11.6 GB Recommended Balance (16 GB GPUs)
Q4_K_S MiniMax-H3-Ref2VA-Pruned-Q4_K_S.gguf 11.6 GB Compact 4-bit K-Quant
Q5_K_M MiniMax-H3-Ref2VA-Pruned-Q5_K_M.gguf 14.1 GB High Quality Balance (24 GB GPUs)
Q5_K_S MiniMax-H3-Ref2VA-Pruned-Q5_K_S.gguf 14.1 GB Compact 5-bit K-Quant
Q6_K MiniMax-H3-Ref2VA-Pruned-Q6_K.gguf 16.7 GB Near-Lossless Precision
Q8_0 MiniMax-H3-Ref2VA-Pruned-Q8_0.gguf 21.6 GB Maximum Fidelity / Reference

βš™οΈ Quickstart & ComfyUI Deployment

Requirements

  • ComfyUI: Version v0.30.0 or higher is required for native MiniMax-H3 architecture support.
  • Extension: ComfyUI-GGUF custom node package installed.

Setup Steps

  1. Download your desired .gguf variant from the table above.
  2. Place the downloaded .gguf file into the ComfyUI/models/unet/ directory.
  3. In your ComfyUI workflow, load the model using the UnetLoaderGGUF node.

πŸ“‹ Model Variants & Input Specifications

  • H3-Base-FL2VA (First-and-Last-Frame Mode):

    • No image input: Operates as standard Text-to-Video / Text-to-Audio-Video.
    • Single image input: Generates video conditioned on the initial frame or final frame.
    • Dual image input: Interpolates seamlessly between designated first and last reference frames.
  • H3-Base-Ref2VA (Omni-Reference Mode):

    • Images: Up to 9 reference images.
    • Videos: Up to 3 video clips (2–15 seconds per clip; total combined duration $\le$ 15s).
    • Audio: Up to 3 audio clips (must be accompanied by visual input; duration 2–15s).
    • Combined Context: Maximum of 12 multi-modal input assets allowed per prompt.

πŸ“Š Output Specifications

  • Duration: 4 – 15 seconds.
  • Aspect Ratios: Flexible ratio support (16:9, 9:16, 4:3, 3:4, 1:1, 21:9).
  • Native Resolution: 768p base resolution (scalable to 2K via H3-Regenerate-2K).
  • Frame Rate: 24 FPS.
  • Audio Track: 32 kHz stereo audio output.
  • Multilingual Dialogue: Stable support across 11 primary languages (English, Chinese, French, German, Japanese, Korean, Spanish, Russian, Portuguese, Italian, Arabic).

πŸ“œ Legal & License

MiniMax H3 is released under the MiniMax H3 Community License Agreement. Please refer to the MiniMaxAI/MiniMax-H3 repository and the repository's LICENSE file for full terms and commercial usage guidelines.

Downloads last month
-
GGUF
Model size
20B params
Architecture
wan
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Abiray/MiniMax-H3-Pruned-GGUF

Quantized
(31)
this model