You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Access to these weights is granted manually by the authors. Please tell us your name, affiliation and intended use.

Log in or Sign Up to review the conditions and access this model content.

mlp-moe (Kiwi-Edit 5B with a 12-expert FFN mixture-of-experts)

The Kiwi-Edit 5B DiT with the FFN of 14 of its 30 blocks (blocks 1,3,...,27) replaced by a DeepSeek-style mixture of 12 full-weight experts (top-3 routing, no shared expert, expert-level balance loss). Experts are initialised from the five category LoRA experts and a direct full-parameter fine-tune (two noised copies each) and the whole DiT is trained end to end.

Training checkpoint: step 10,480 (end of training). Optimizer state: removed (weights only). Code: https://github.com/massyzs/Video-Editor-MoE (scripts/moe/mlp-moe/)

Package contents

The weights are shipped as a tar.gz stream split into ~20 GB chunks (Hugging Face limits single files to 50 GB).

chunk size sha256
mlp-moe.tar.gz.part-00 21.47 GB ab1e1388feb391f19ac711030c4175597feb2e79f58a82508fc96cb80e7155f2
mlp-moe.tar.gz.part-01 8.30 GB f613d6eeb910071d8998d3bcccb7e1c908ab19f752c7c2a35ce384c00e11b797

Total: 2 chunks, 29.77 GB compressed / 37.14 GB extracted.

Reassemble and verify:

huggingface-cli download Massyzs/mlp-moe --local-dir mlp-moe-package      # after your access request is approved
cd mlp-moe-package
cat mlp-moe.tar.gz.part-* | tar -xzf - -C <BASE_DIR>/ckpt/                # creates <BASE_DIR>/ckpt/mlp-moe/
cp SHA256SUMS.contents <BASE_DIR>/ckpt/ && (cd <BASE_DIR>/ckpt && sha256sum -c SHA256SUMS.contents)   # optional check

MANIFEST.json lists every payload file with its size and sha256, and every chunk with its sha256 (SHA256SUMS.parts).

Extracted files:

file size sha256
mlp-moe/meta.json 373 B d0a7b3e850616b946d2ebb23b50514136502c74feb30e1c75b99265a74b07542
mlp-moe/moe_v3_trainable.safetensors 37.14 GB 82bfbe6509636b5027385ecb92c1168623bdd4fe33532b02bfe7dbe5439d8fe3

Base weights you also need

The package contains the full DiT (18.57B parameters, bf16); only the MLLM encoder (qwen/), the VAE and moe_expert_init.safetensors are needed in addition.

All frozen base components are published, already converted to the layout this code expects, in Massyzs/kiwi-edit-5b-instruct-only-videoxfun:

huggingface-cli download Massyzs/kiwi-edit-5b-instruct-only-videoxfun --local-dir <BASE_DIR>/ckpt/kiwi-edit-instruct-only-videoxfun

See BASE_WEIGHTS.md for where each component comes from (Kiwi-Edit, Wan2.2) and its checksum.

Inference

python scripts/moe/mlp-moe/moe_v3_infer.py --base_dir <BASE_DIR> --ckpt <BASE_DIR>/ckpt/mlp-moe \
    --src_video input.mp4 --prompt "Replace the red car with a blue truck" --name edited --out_dir <OUT_DIR>
# multi-step instruction: --prompts "Remove the dog ||| Convert to oil painting style"

Single 80 GB GPU. Defaults: 50 steps, 49 frames at 480x832, sigma shift 5, segmented text CFG (scale 3 on the first 60% of the steps). Output: <OUT_DIR>/<name>.mp4, lossless PNG frames in <OUT_DIR>/frames/<name>/, and per-layer expert-routing maps.

Architecture

30-block DiT (dim 3072); in blocks 1,3,...,27 the FFN is a 12-expert MoE (each expert = the original 3072->14336->3072 FFN), top-3 with renormalised gate weights, balance loss alpha 1e-4. Full state dict: 1,457 tensors, 37.1 GB bf16.

Training data

OpenVE single-instruction edits (stage A) followed by multi-instruction data (goku composite + CoinVE, stage B); 10,480 steps, effective batch 192 (8 x 24 GPUs).

License

The released weights and code are under Apache-2.0. The frozen base components (Kiwi-Edit MLLM encoder / DiT base, Wan2.2 VAE) are not redistributed here and remain subject to their own licenses.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Massyzs/mlp-moe

Finetuned
(5)
this model