MiniMax-H3 Prompt Rewriter (Qwen2.5-Omni-7B + LoRA) - GGUF

This repository contains the quantized GGUF weights for MiniMax-H3 Prompt Rewriter, created by merging the LoRA adapter lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni with Qwen/Qwen2.5-Omni-7B (Thinker backbone).

Quantized using llama.cpp on Modal.com GPU infrastructure.

Available Quantizations

File Size Description Recommended For
MiniMax-H3-Prompt-Rewriter-Q4_K_M.gguf ~4.5 GB 4-bit medium quantization CPU & low-VRAM GPUs (Ollama)
MiniMax-H3-Prompt-Rewriter-Q8_0.gguf ~7.5 GB 8-bit high-fidelity quantization Near-FP16 accuracy

Quick Start with Ollama

You can run this model directly via Ollama:

ash ollama run hf.co/alibaybay/MiniMax-H3-Prompt-Rewriter-GGUF:Q4_K_M

Or using the included Modelfile: ash ollama create minimax-h3-rewriter -f ./Modelfile ollama run minimax-h3-rewriter

Prompt Rewriting Schema

The model restructures short user requests into the official MiniMax-H3 video & audio generation schema: - integrated_multimodal_description: Shot breakdown, camera motions, and visual styling. - overall_soundscape: Ambient sound, physical foley, non-verbal sound.

on_diegetic_music: Instrumentation, tempo, and background score.

Downloads last month
277
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for alibaybay/MiniMax-H3-Prompt-Rewriter-GGUF

Quantized
(24)
this model