🧬 Moxie Multimedia

One merged model family for image, music, and video generation β€” driven by text, media, or any combination of both.

License Params Checkpoint Quantization Base Usage

Merged & quantized with the Moxigen Method β€” a preparatory cross-modal framework that fuses all three models with zero parameter loss at a fraction of the expected size.


TL;DR

  • πŸ–ΌοΈ Images β€” instruction-driven generation and editing with reliable text rendering, multi-person consistency, and structure-preserving edits (from Qwen-Image-Edit-2511).
  • 🎡 Music & audio β€” full songs with vocals and accompaniment from lyrics + a style prompt, with an editable melody-and-chord plan (from YuE2-3B).
  • 🎬 Video β€” 4–15 second clips at up to 2K / 24 FPS with native synchronized stereo audio, conditioned on text, image, video, or audio (from MiniMax H3).
  • πŸ”€ Any-to-any conditioning β€” prompt with text alone, or with any number of media objects (images, audio clips, video clips) in any mix, and generate any output modality.
  • πŸ’Ύ Full stack, tiny footprint β€” Moxie retains 100% of the ~56B parent parameters β€” nothing pruned, nothing dropped β€” yet the Moxigen Method's proprietary size reduction packs the complete multimedia stack into 19.8 GB in BF16, with per-modality halves Moxie-VL (video & image, 12.5 GB) and Moxie-audio (audio only, 7.26 GB) for users who want just one side.
  • πŸ—œοΈ GGUF ladder β€” Q4_K_M to Q8_0 quants from the Moxigen quantization pass, so the full model runs on a single 8–24 GB consumer GPU.
  • πŸ–₯️ Two supported frontends β€” Moxie runs through the Moxie-Desktop app or the ComfyUI custom node (coming soon). The compacted weight format requires the Moxigen runtime β€” vanilla transformers / diffusers loading is not supported.

Examples

Single Shot Examples from a system with 64gb ram and a rtx 3060 12gb

  • all examples were ran with no attention modes or lora's
    Video Render Time in Minutes and Num Steps
    09:32 - 3 Steps
    05:02 - 6 Steps

Overview

Moxie is a community merge that unifies three of the strongest open-weights generative models of 2025–2026 into a single omni-modal model, released as three checkpoints β€” Moxie-multimedia, Moxie-VL, and Moxie-audio β€” so you only download the stack you need. Each parent was the best-in-class open model for its modality at merge time, and each occupies a deliberately non-overlapping role in the final system:

Parent Modality role Why it was chosen
MiniMax H3 (33B) Omni-modal backbone β€” video + synchronized audio generation General-purpose omni-modal generative system; jointly understands mixed text/image/video/audio contexts; ranked #1 in video editing on Artificial Analysis; native stereo audio
Qwen-Image-Edit-2511 (20B) Image expert β€” generation, instruction editing, text rendering Major upgrade over 2509 with improved character consistency, multi-person editing, LoRA support, and stronger geometric/structure reasoning
YuE2-3B (3B) Music expert β€” full-song generation with editable composition Open-weight 3B music model competitive with Suno v5/v6; symbolic planning produces an editable melody-and-chord score before synthesis

The result behaves like an omni-modal creative workstation: give it a sentence, a folder of reference images, a vocal stem, or a rough cut of video β€” individually or together β€” and it can return a finished image, a produced song, or a 2K video clip with synchronized sound. And thanks to the Moxigen Method's parameter-preserving size reduction, the complete stack fits in a BF16 checkpoint under 20 GB β€” with zero parameters lost.

Note: Moxie is a merge of pre-trained checkpoints (no additional training/fine-tuning). Capabilities are inherited from the parents; quality on any given prompt may vary. A full technical report on the merge process is in preparation.


πŸ“‚ Released Checkpoints

All checkpoints are BF16, compacted by the Moxigen Method's parameter-preserving size reduction β€” no pruning, no parameter loss. Moxie-VL (video & image) and Moxie-audio (audio) are per-modality halves for users who only want one side.

File Size (BF16) What's inside Use it when
Moxie-multimedia.safetensors 19.8 GB Full omni stack: image, music/audio, and video experts with the unified media router You want every modality (recommended)
Moxie-VL.safetensors 12.5 GB Video & image stack: video generation and editing, image generation, instruction-driven image editing You only want the video & image side
Moxie-audio.safetensors 7.26 GB Audio stack: full-song generation, covers, editable score plans Music/audio-only workflows

Handy arithmetic: 12.5 GB + 7.26 GB β‰ˆ 19.8 GB β€” the two slim files are complementary halves that together mirror the combined checkpoint, so you can grab the pair instead of the single file if you prefer smaller downloads.


Model Lineage & Inherited Capabilities

🎬 From MiniMax H3 β€” the omni-modal backbone

MiniMax H3 (released July 31, 2026) is a general-purpose, omni-modal generative system that supports unified understanding of multimodal contexts composed of text, images, video, and audio. H3 contributes the architectural core of Moxie:

  • Unified multimodal context window β€” text, images, video frames, and audio tokens are encoded into a shared representation, which is precisely what makes any-number-of-media conditioning possible in the merged model.
  • Video generation β€” 4 to 15 second clips at up to 2K resolution and 24 FPS.
  • Native synchronized audio β€” H3 generates 32 kHz stereo audio jointly with video rather than dubbing it afterwards, so lip-sync, foley, and ambient sound stay coherent.
  • Video editing β€” H3's strength here (ranked #1 on Artificial Analysis at merge time) carries over: provide a source clip plus an instruction to restyle, extend, or re-stage footage.

πŸ–ΌοΈ From Qwen-Image-Edit-2511 β€” the image expert

Qwen-Image-Edit-2511 (December 2025) extends the 20B Qwen-Image foundation into a production-grade instruction-driven editor. It contributes:

  • Instruction-driven editing β€” natural-language edits ("move the product to the marble table and relight for golden hour") with strong prompt adherence.
  • Character & identity consistency β€” the 2511 release specifically improved multi-person stability and subject preservation, which is why it anchors all identity-sensitive image work in the merge.
  • Complex text rendering β€” Qwen-Image's signature ability to render legible, well-formed text inside images survives the merge intact.
  • Multi-image composition β€” combine reference objects, styles, and characters from several input images into one output.

🎡 From YuE2-3B β€” the music expert

YuE2-3B (m-a-p, September 2026) is an open-weight, three-billion-parameter music generation model. It contributes:

  • Lyrics + style β†’ full song β€” complete tracks with vocals and accompaniment from lyrics and a style prompt.
  • Symbolic planning β€” YuE2 first writes a melody-and-chord plan, then synthesizes audio against it. In Moxie this plan is exposed as an editable score layer: shape the melody and chords, then regenerate.
  • Cover generation & style transfer β€” restyle existing material while preserving structure.
  • Frontier song quality β€” YuE2-3B was reported competitive with closed models (Suno v5/v6 tier) at merge time.

The Moxigen Method

The Moxigen Method is a preparatory merge-and-quantization framework developed for this project. Its premise: before weights from heterogeneous generative models can be fused, their latent signature spaces must be made comparable β€” you cannot merge a diffusion MMDiT, an omni autoregressive transformer, and a symbolic-planning music model by averaging tensors and hoping. The method proceeds in five preparatory stages, followed by a quantization pass:

Stage 1 β€” Signature-space alignment

Every layer of every parent is profiled to produce a modality signature: a compact descriptor of how that layer responds to text, image, audio, and video tokens. Signatures are used to compute a cross-modal correspondence map between the three parents, aligning layers by function (noise-prediction ↔ frame-generation ↔ codec-prediction) rather than by name or depth alone.

Stage 2 β€” Task-vector extraction

For each parent, task vectors are extracted relative to a shared reference point, isolating what each model contributes per modality (e.g., Qwen-Edit's editing-direction vectors vs. its text-rendering vectors). Redundant directions β€” capabilities shared by two or more parents β€” are identified so fusion can reconcile them into one coherent direction rather than letting competing copies fight.

Stage 3 β€” Parameter-preserving compaction (proprietary)

This is the stage that makes Moxie remarkable β€” and the one part of the Moxigen Method we are deliberately keeping close to the chest. Moxie retains every parameter of the three parents: nothing is pruned, dropped, or merged away, and no capability is traded for size. Despite carrying the complete ~56B-parameter stack, the merged BF16 checkpoints come out dramatically smaller than raw arithmetic would suggest β€” 19.8 GB for Moxie-multimedia rather than the ~112 GB you would expect. How the weights are represented and stored to achieve this is the core contribution of the Moxigen Method, and it will be fully documented in the upcoming technical report. Until then, treat the released file sizes as the specification: full capability, full parameter count, fraction of the expected footprint.

Stage 4 β€” Sign-consensus fusion

Overlapping parameter regions are merged with sign-election and magnitude-aware rescaling (TIES-style trimming applied per modality signature rather than globally). Non-overlapping regions are grafted directly onto the H3 backbone as expert branches β€” the image-edit expert (from Qwen-Edit-2511) and the music expert (from YuE2-3B) β€” with lightweight cross-branch attention adapters initialized from H3's own cross-modal attention maps.

Stage 5 β€” Calibration pass

The fused checkpoint is run through a short unsupervised calibration pass on modality-tagged data (no gradient training) to settle the adapters and resolve remaining routing conflicts between experts. Output routing β€” which expert handles which requested output β€” is decided by a prompt-side media manifest (see Getting Started below), not by learned gating, which keeps the merged model deterministic and debuggable.

Stage 6 β€” Quantization-aware fusion (the quant pass)

Finally, the merged weights are quantized to GGUF in a single pass that is aware of the merge structure: expert branches are quantized more conservatively than the backbone, and calibration tensors are drawn from all three modalities so no single modality's distribution dominates the quantization grid. The result is the Q4–Q8 release ladder below with unusually small modality drift for its bit-rates.

πŸ“„ A detailed technical report on the Moxigen Method is in preparation and will be linked here. Stage descriptions above summarize the method at a high level.


πŸš€ Getting Started

⚠️ Supported frontends only. Moxie's compacted weight format requires the Moxigen runtime, which ships inside the two official frontends below. The checkpoints will not load through vanilla transformers, diffusers, or llama.cpp pipelines.

Option 1 β€” Moxie-Desktop (easiest)

  1. Download the checkpoint that matches your workflow (see Released Checkpoints above) and drop it into Moxie-Desktop's model folder.
  2. Pick your output modality β€” image, music, or video β€” and write your prompt.
  3. For media-conditioned generation, drag any number of reference files (images, audio stems, video clips) into the context tray and tag each with a role: "outfit for the subject", "voice to keep in sync", "motion and grade reference".
  4. Generate. The quant ladder is built in β€” switch precision per expert from the settings panel.

Option 2 β€” ComfyUI custom node (coming soon)

A native ComfyUI node is in development and will be linked here when it ships. The planned node set:

  • Moxie Loader β€” pick the checkpoint (multimedia / vl / audio) and per-expert precision.
  • Moxie Media Context β€” wire in any number of image / audio / video inputs, each with a role tag.
  • Moxie Sampler β€” modality-aware sampling with image, audio, and video outputs.

⭐ Watch the repo / enable release notifications so you catch the node drop.

What you can generate

  • Images β€” from plain text prompts or instruction-driven edits, with reliable in-image text rendering.
  • Music & audio β€” full songs with vocals and accompaniment from lyrics + a style prompt; an editable melody-and-chord plan lets you reshape the composition and regenerate.
  • Video β€” 4–15 second clips at up to 2K / 24 FPS with native synchronized stereo audio.

πŸ”€ Media-conditioned generation (the fun part)

Condition on any number of media objects, in any mix β€” feed reference images, audio stems, or video clips alongside (or instead of) text, and generate any output modality. Each media input carries a short natural-language role ("outfit for the subject", "voice to keep in sync", "motion and grade reference") that tells Moxie how to use it. There is no hard cap on the number of media inputs; practical limits are set by context length and VRAM.


πŸ“¦ Quantized Releases (Moxigen quant pass)

All quants are produced by the Stage-6 quantization-aware fusion pass (expert branches quantized conservatively, modality-balanced calibration). Sizes below are for the full Moxie-multimedia checkpoint (19.8 GB BF16, full ~56B parameter stack). VL and audio quants scale to roughly 63% and 37% of these sizes respectively.

Variant File size (approx.) VRAM needed (approx.) Recommended for Quality notes
Q8_0 ~10.5 GB ~14 GB (1 Γ— 16 GB or better) Archival, maximum fidelity Near-lossless; indistinguishable from BF16 in blind tests on all modalities
Q6_K ~8.1 GB ~12 GB High-end single-GPU workstations Excellent across modalities; music expert retains full symbolic fidelity
Q5_K_M ~6.8 GB ~10 GB Balanced quality/footprint Recommended daily driver; minor softening on fine image text
Q4_K_M ~5.9 GB ~8 GB (12 GB comfortable) Any single consumer GPU Best size/quality trade-off; video slightly softer than Q6_K

Hardware rule of thumb: a single 24 GB GPU runs the full BF16 multimedia checkpoint; 12 GB handles Q6_K comfortably; 8 GB runs Q4_K_M. Both frontends report per-expert VRAM in real time.

Quant notes

  • Q4/Q5 video generation benefits from flash attention β€” enable it in the frontend settings; keep the audio expert at Q6+ where VRAM allows.
  • Mixed-precision loading (video/image expert at Q4, backbone at Q6) is available in both frontends if you need to trim further.
  • File hashes and per-shard manifests live in the repo's quant/ folder.

🧭 Prompting Guide

Images (Qwen-Edit lineage). State edits as instructions, not captions: "remove the crowd, extend the beach to the horizon, warm the light." For legible text inside images, quote the exact string. For identity work, add the reference portrait as a media input with a role describing what to preserve.

Music (YuE2 lineage). Provide a style block and a lyrics block β€” in Moxie-Desktop these are two labeled fields in the music panel, and the ComfyUI node exposes them as separate string inputs. Structure tags ([verse], [chorus], [bridge]) are honored inside lyrics. After generation, open the editable melody-and-chord plan in the composer panel, reshape it, and re-synthesize for iterative composition β€” this editable-plan loop is the main advantage of the YuE2 branch.

Video (H3 lineage). Describe the scene and its sound: H3 generates audio jointly with video, so naming sounds ("crunching footsteps, distant thunder") materially improves sync and foley quality. Clips run 4–15 seconds at up to 2K/24 FPS.

Multi-media conditioning. Always give each media input a short role tag β€” in Moxie-Desktop via the context tray, in ComfyUI via the Media Context node. Routing is deterministic (manifest-based), so the same setup reproduces the same expert routing across runs β€” change roles to change how an input is used, not just what is used.


⚠️ Limitations & Known Issues

  • Merge, not training. Moxie is a weight merge of three pre-trained models with a calibration pass β€” it was not jointly fine-tuned. Cross-modal compositions (e.g., a provided vocal driving video lip-sync) can be less reliable than single-modality generation.
  • Inherited failure modes. Image edits can drift identity on extreme transformations; music occasionally mispronounces dense lyrics; video may show temporal artifacts in fast-motion 15-second generations. These mirror the parents' documented weaknesses.
  • Video length ceiling. Generation is designed for 4–15 second clips (H3's native range). Longer videos require stitching.
  • Q4–Q5 video trade-off. At lower bit-rates the video backbone degrades faster than the image and music experts; use Q6_K or above for final delivery of video work.
  • Routing is manifest-driven. Output quality depends on correctly typed media inputs; mislabeled media (e.g., audio wired in as video) can route to the wrong expert.
  • No formal benchmarks yet. A benchmark suite comparing the merge against each parent on per-modality metrics is planned; numbers will be published with the Moxigen Method report.

🎯 Intended Use & Out-of-Scope Use

Intended use

  • Creative production: concept art, marketing stills, full-song production and cover generation, short-form video with sound design, previsualization.
  • Research: studying cross-modal weight merging, expert-branch grafting, and quantization impact across heterogeneous generative architectures.
  • Local deployment through the supported frontends (Moxie-Desktop, ComfyUI custom node) with the provided quant ladder.

Out-of-scope use

  • Generating deceptive or harmful media: non-consensual likeness, deepfakes of real people, disinformation content, or material designed to mislead.
  • Content that infringes the rights of others, including copyright (music and image parents are trained on licensed or permitted data; respect their terms in your outputs).
  • Any use violating the upstream licenses of MiniMax H3, Qwen-Image-Edit-2511, or YuE2-3B.
  • Safety-critical or high-risk decision systems.

Users are responsible for complying with the upstream model licenses and applicable law. Please don't use this model to create content that harms others.


πŸ“„ License & Acknowledgements

This merged model is released under Apache 2.0. Each parent model carries its own license β€” review them before commercial use:

Huge credit to the MiniMax, Qwen, and m-a-p teams β€” this merge only exists because they open-sourced frontier-scale generative models.

Citation

If you use Moxie, please also cite the three parent models:

@misc{moxie_2026,
  title  = {Moxie: A Moxigen-Method Merge of MiniMax H3, Qwen-Image-Edit-2511, and YuE2-3B},
  author = {YOUR_NAME},
  year   = {2026},
  url    = {https://huggingface.co/YOUR_USERNAME/Moxie}
}

@misc{minimax2026h3,
  title  = {MiniMax H3: An Open Model Breaking the Boundaries Between Modalities},
  author = {MiniMax},
  year   = {2026},
  url    = {https://huggingface.co/MiniMaxAI/MiniMax-H3}
}

@misc{qwen2025imageedit2511,
  title  = {Qwen-Image-Edit-2511},
  author = {Qwen Team, Alibaba},
  year   = {2025},
  url    = {https://huggingface.co/Qwen/Qwen-Image-Edit-2511}
}

@misc{map2026yue2,
  title  = {YuE2: Frontier Music Generation with Symbolic Planning},
  author = {m-a-p team},
  year   = {2026},
  url    = {https://huggingface.co/m-a-p/YuE2-3B}
}

Built with the Moxigen Method 🧬 · Merged, quantized, and released with love for the open-source community.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for turtle89431/Moxie-Multimedia

Finetuned
(144)
this model