LDM-is-AE: Latent Diffusion Model is an Intrinsic Auto-Encoder for End-to-End Image Generation
Zhengqiang Zhang, Lingchen Sun, Rongyuan Wu, Qiaosi Yi, Xiangtao Kong, Chaodong Xiao, Lei Zhang
The Hong Kong Polytechnic University · OPPO Research Institute
Contents: Algorithm · Quick Start · Results · Model Weights · Citation · License
💡 LDM is an AE.
Latent diffusion is normally a two-stage pipeline: train a
VAE, then train a diffusion model in its latent space -- and inherit the VAE's bias. LDM-is-AE removes
the pipeline. We split the DiT backbone into DiT-D (the first 30 blocks) and DiT-E (the last 2 blocks) and
supervise the intermediate feature F in the image domain at every timestep. The backbone's hidden
decode->encode path then becomes an explicit intrinsic auto-encoder, trained end-to-end in a single
stage without external VAE.

Figure 1. (a) the DiT backbone performs latent → feature → latent;
(b) image-space supervision aligns the intermediate feature with the image domain at every timestep;
(c) at the zero-noise timestep (t=1) an explicit latent → image → latent path
makes the backbone an intrinsic auto-encoder, which in turn enables image → latent → image.
📐 Algorithm
Algorithm 1: Training loop of LDM-is-AE -- a single-stage, end-to-end loop
(implemented in ldm_is_ae/train.py and ldm_is_ae/denoiser.py).
Inputs: training set X, total iterations T
for i = 1, ..., T do
(x, c) = sample_batch(X), t ~ U[0, 1], z_0 ~ N(0, I)
// AE encoding
x_u = pixel_unshuffle(x, p) # patchify to the image domain
with torch.no_grad():
z_1 = dit_e(x_u, t=1, c) # image-to-latent at t = 1 (no grad)
// LDM denoising
z_t = t * z_1 + (1 - t) * z_0
F, F' = dit_d(z_t, t, c) # split output: image-aligned F and F'
F_full = gamma(t) * F + (1 - gamma(t)) * F' # auxiliary feature mixing
z1_hat = dit_e(F_full, t, c) # map the mixed feature back to latent
L_total = L_ldm(z1_hat, z_1) + w_toimg * L_toimg(F, x_u) # latent loss + image-domain supervision
L_total.backward()
optimizer.step()
end for
🚀 Quick Start
0. Clone
git clone https://github.com/PolyU-VCLab/LDMisAE.git && cd LDMisAE
1. Install
pip install -r requirements.txt # torch torchvision numpy scipy einops timm pillow
# opencv-python requests tqdm dill loguru torch-fidelity
# transformers (text-to-image only)
Two pre-trained assets are passed on the command line instead of a fixed path: the LPIPS vgg.pth via
--lpips_model_path (needed by training; the launchers forward LPIPS_MODEL_PATH), and the
torch-fidelity Inception-V3 weights via TORCH_HOME or --weights.
2. Get the weights
huggingface-cli download xtudbxk/LDMisAE LDMisAE.256.ckpt --local-dir weights
| Resolution | File | Size |
|---|---|---|
| 256×256 | LDMisAE.256.ckpt |
3.90 GB |
| 512×512 | LDMisAE.512.ckpt |
3.93 GB |
3. Sample
# CKPT = a released .ckpt, or a run directory holding checkpoint-last.pth
CKPT=weights/LDMisAE.256.ckpt IMG_SIZE=256 CFG=2.25 NUM_IMAGES=50000 bash scripts/inference.sh
CKPT=weights/LDMisAE.512.ckpt IMG_SIZE=512 NPROC=8 CFG=2.2 bash scripts/inference.sh
4. Evaluate
bash scripts/evaluate.sh <sample_dir> <ref_stats.npz> [tag]
<ref_stats.npz> holds the reference mu/sigma: ADM VIRTUAL_imagenet256_labeled.npz /
VIRTUAL_imagenet512.npz, or the JiT statistics of the matching resolution.
5. Train
# class-conditional 256x256 (JiT-H/half: DiT-E 2 blocks / DiT-D 30 blocks)
IMAGENET_PATH=/data/imagenet256 BATCH_SIZE=16 NPROC=8 bash scripts/train_256.sh
# class-conditional 512x512
IMAGENET_PATH=/data/imagenet512 BATCH_SIZE=16 NPROC=8 bash scripts/train_512.sh
# text-to-image (STAGE=1 freezes the backbone, STAGE=2 trains jointly)
TEXT_ENCODER_PATH=/models/Qwen3_1.7B BLIP3O_PATH=/data/blip3o STAGE=1 BATCH_SIZE=16 NPROC=8 bash scripts/train_t2i.sh
IMAGENET_PATH is the parent of train/ (the loader appends train/ itself).
📊 Results
Class-conditional ImageNet generation at 256×256 (Table 1) and 512×512 (Table 2). Models: Gen. = generator, AE = autoencoder, Dec. = decoder, VFM = vision foundation model; Repr.: Pixel = pixel diffusion, Fixed = a fixed latent representation, Dynamic = a dynamically evolved latent representation in training; Aux. data = external training data beyond ImageNet; Training FLOPs (×1019) is the generator-only training cost and does not include the cost of training a separate AE or VFM.
Table 1. Class-conditional ImageNet generation at 256×256.
| Method | Models | Repr. | Total params (M) | Epochs | Aux. data | Training FLOPs | FID ↓ | IS ↑ |
|---|---|---|---|---|---|---|---|---|
| Two-stage | ||||||||
| DiT-XL/2 | Gen.+AE | Fixed | 759 | 1400 | OpenImages | 45.4 | 2.27 | 278 |
| SiT-XL/2 | Gen.+AE | Fixed | 759 | 1400 | OpenImages | 45.4 | 2.06 | 270 |
| LightningDiT | Gen.+AE+VFM | Fixed | 745 | 800 | – | 19.1 | 1.35 | 295 |
| REPA-SiT | Gen.+AE+VFM | Fixed | 759 | 800 | OpenImages | 25.9 | 1.29 | 306 |
| DDT-XL/2 | Gen.+AE+VFM | Fixed | 759 | 400 | OpenImages | 19.2 | 1.26 | 311 |
| REPA-E (tuning) | Gen.+AE+VFM | Fixed | 759 | 800 | OpenImages | 57.8 | 1.12 | 303 |
| SVG-XL | Gen.+AE+VFM | Fixed | 758 | 1400 | DINOv3 | 22.8 | 1.92 | 265 |
| RAE-DiTDH | Gen.+AE+VFM | Fixed | 839 | 800 | DINOv2 | – | 1.13 | 263 |
| One-stage | ||||||||
| REPA-E (scratch) | Gen.+AE+VFM | Dynamic | 759 | 80 | – | 5.78 | 1.67 | – |
| UNITE-XL | Gen.+Dec. | Dynamic | 763 | 240 | – | 12.0 | 1.75 | 310 |
| DSD | Gen.+VFM | Dynamic | 205 | 50 | – | – | 3.35 | 255 |
| ADM-U | Gen. | Pixel | 554 | 400 | – | – | 4.59 | 187 |
| RIN | Gen. | Pixel | 410 | 480 | – | 20.5 | 3.42 | 182 |
| PixNerd | Gen.+VFM | Pixel | 700 | 160 | – | 5.49 | 2.15 | 297 |
| PixelFlow | Gen. | Pixel | 677 | 320 | – | 239 | 1.98 | 282 |
| JiT-H/16 | Gen. | Pixel | 953 | 600 | – | 14.0 | 1.86 | 303 |
| LDM-is-AE (Ours) | Gen. | Dynamic | 961 | 300 | – | 7.02 | 1.80 | 314 |
Table 2. Class-conditional ImageNet generation at 512×512.
| Method | Models | Repr. | Total params (M) | FID ↓ | IS ↑ |
|---|---|---|---|---|---|
| Two-stage | |||||
| DiT-XL/2 | Gen.+AE | Fixed | 759 | 3.04 | 241 |
| SiT-XL/2 | Gen.+AE | Fixed | 759 | 2.62 | 252 |
| REPA-SiT-XL/2 | Gen.+AE+VFM | Fixed | 759 | 2.08 | 275 |
| One-stage | |||||
| ADM-G | Gen. | Pixel | 559 | 7.72 | 173 |
| RIN | Gen. | Pixel | 320 | 3.95 | 216 |
| PixNerd-XL/16 | Gen.+VFM | Pixel | 700 | 2.84 | 246 |
| DeCo | Gen. | Pixel | 682 | 2.22 | 290 |
| JiT-H/32 | Gen. | Pixel | 956 | 1.94 | 309 |
| LDM-is-AE (Ours) | Gen. | Dynamic | 961 | 1.90 | 320 |
🤗 Model Weights
https://huggingface.co/xtudbxk/LDMisAE
| File | Resolution | Size |
|---|---|---|
LDMisAE.256.ckpt |
256×256 | 3.90 GB |
LDMisAE.512.ckpt |
512×512 | 3.93 GB |
📝 Citation
@inproceedings{ldm_is_ae_2026,
title = {LDM-is-AE: Latent Diffusion Model is an Intrinsic Auto-Encoder for End-to-End Image Generation},
author = {Zhang, Zhengqiang and Sun, Lingchen and Wu, Rongyuan and Yi, Qiaosi and
Kong, Xiangtao and Xiao, Chaodong and Zhang, Lei},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026}
}
⚖️ License
Code: Apache License 2.0 -- see LICENSE. Model weights and data are released separately and are
intended for research use. Third-party components are acknowledged in NOTICE.