Composing diffusion models through a shared latent space
Overview
InstantFusion connects heterogeneous image generators through a learned shared latent interface. Model-specific Latent AutoEncoders (LAEs) align intermediate states while preserving the information needed for reconstruction and continued denoising. The base generators remain frozen during LAE training.
The shared interface supports generation acceleration, multi-preference composition, and cross-model on-policy distillation (OPD).
Released checkpoints
This repository provides trained LAE interface weights. The base generators must be obtained separately.
| Checkpoint | Connected generators |
|---|---|
| sd3_qwen_lae.bin | SD3 and Qwen-Image |
| flux1_dev_qwen_lae.bin | FLUX.1-dev and Qwen-Image |
| flux1_schnell_qwen_lae.bin | FLUX.1-schnell and Qwen-Image |
| flux1_sd3_qwen_lae.bin | FLUX.1, SD3, and Qwen-Image |
| flux2_dev_qwen_lae.bin | FLUX.2-dev and Qwen-Image |
| flux2_klein_qwen_lae.bin | FLUX.2-Klein and Qwen-Image |
| zimage_turbo_qwen_lae.bin | Z-Image-Turbo and Qwen-Image |
Visual examples
Latent reconstruction
LAE reconstruction examples for Qwen-Image, FLUX.1, and SD3. Mean SSIM is reported over 100 images for each model.
Cross-model denoising
Intermediate states can be transferred through the shared space so that another generator can continue denoising. The examples below show cross-model generation with SD3, FLUX.1, and Qwen-Image.
For implementation details, supported inference configurations, and training instructions, visit the InstantFusion GitHub repository.