LDM is an AE: z_t -> DiT-D (decoding) -> x (image domain) -> DiT-E (encoding) -> zLDM-is-AE: Latent Diffusion Model is an Intrinsic Auto-Encoder for End-to-End Image Generation

License Framework Conference Code Weights


Zhengqiang Zhang, Lingchen Sun, Rongyuan Wu, Qiaosi Yi, Xiangtao Kong, Chaodong Xiao, Lei Zhang
The Hong Kong Polytechnic University · OPPO Research Institute

LDM is an AE: z_t -> DiT-D (decoding) -> x (image domain) -> DiT-E (encoding) -> z

Contents: Algorithm · Quick Start · Results · Model Weights · Citation · License

💡 LDM is an AE.

Latent diffusion is normally a two-stage pipeline: train a VAE, then train a diffusion model in its latent space -- and inherit the VAE's bias. LDM-is-AE removes the pipeline. We split the DiT backbone into DiT-D (the first 30 blocks) and DiT-E (the last 2 blocks) and supervise the intermediate feature F in the image domain at every timestep. The backbone's hidden decode->encode path then becomes an explicit intrinsic auto-encoder, trained end-to-end in a single stage without external VAE.

LDM-is-AE: the DiT backbone is split into DiT-E / DiT-D and the intermediate feature F is aligned with PixelUnshuffle(x) in the image domain at every timestep.

Figure 1. (a) the DiT backbone performs latent → feature → latent; (b) image-space supervision aligns the intermediate feature with the image domain at every timestep; (c) at the zero-noise timestep (t=1) an explicit latent → image → latent path makes the backbone an intrinsic auto-encoder, which in turn enables image → latent → image.


📐 Algorithm

Algorithm 1: Training loop of LDM-is-AE -- a single-stage, end-to-end loop (implemented in ldm_is_ae/train.py and ldm_is_ae/denoiser.py).

Inputs: training set X, total iterations T
for i = 1, ..., T do
    (x, c) = sample_batch(X),  t ~ U[0, 1],  z_0 ~ N(0, I)

    // AE encoding
    x_u = pixel_unshuffle(x, p)                    # patchify to the image domain
    with torch.no_grad():
        z_1 = dit_e(x_u, t=1, c)                   # image-to-latent at t = 1 (no grad)

    // LDM denoising
    z_t    = t * z_1 + (1 - t) * z_0
    F, F'  = dit_d(z_t, t, c)                      # split output: image-aligned F and F'
    F_full = gamma(t) * F + (1 - gamma(t)) * F'    # auxiliary feature mixing
    z1_hat = dit_e(F_full, t, c)                   # map the mixed feature back to latent

    L_total = L_ldm(z1_hat, z_1) + w_toimg * L_toimg(F, x_u)   # latent loss + image-domain supervision
    L_total.backward()
    optimizer.step()
end for

🚀 Quick Start

0. Clone

git clone https://github.com/PolyU-VCLab/LDMisAE.git && cd LDMisAE

1. Install

pip install -r requirements.txt        # torch torchvision numpy scipy einops timm pillow
                                       # opencv-python requests tqdm dill loguru torch-fidelity
                                       # transformers (text-to-image only)

Two pre-trained assets are passed on the command line instead of a fixed path: the LPIPS vgg.pth via --lpips_model_path (needed by training; the launchers forward LPIPS_MODEL_PATH), and the torch-fidelity Inception-V3 weights via TORCH_HOME or --weights.

2. Get the weights

huggingface-cli download xtudbxk/LDMisAE LDMisAE.256.ckpt --local-dir weights
Resolution File Size
256×256 LDMisAE.256.ckpt 3.90 GB
512×512 LDMisAE.512.ckpt 3.93 GB

3. Sample

# CKPT = a released .ckpt, or a run directory holding checkpoint-last.pth
CKPT=weights/LDMisAE.256.ckpt IMG_SIZE=256 CFG=2.25 NUM_IMAGES=50000 bash scripts/inference.sh
CKPT=weights/LDMisAE.512.ckpt IMG_SIZE=512 NPROC=8 CFG=2.2 bash scripts/inference.sh

4. Evaluate

bash scripts/evaluate.sh <sample_dir> <ref_stats.npz> [tag]

<ref_stats.npz> holds the reference mu/sigma: ADM VIRTUAL_imagenet256_labeled.npz / VIRTUAL_imagenet512.npz, or the JiT statistics of the matching resolution.

5. Train

# class-conditional 256x256  (JiT-H/half: DiT-E 2 blocks / DiT-D 30 blocks)
IMAGENET_PATH=/data/imagenet256 BATCH_SIZE=16 NPROC=8 bash scripts/train_256.sh
# class-conditional 512x512
IMAGENET_PATH=/data/imagenet512 BATCH_SIZE=16 NPROC=8 bash scripts/train_512.sh
# text-to-image (STAGE=1 freezes the backbone, STAGE=2 trains jointly)
TEXT_ENCODER_PATH=/models/Qwen3_1.7B BLIP3O_PATH=/data/blip3o STAGE=1 BATCH_SIZE=16 NPROC=8 bash scripts/train_t2i.sh

IMAGENET_PATH is the parent of train/ (the loader appends train/ itself).


📊 Results

Class-conditional ImageNet generation at 256×256 (Table 1) and 512×512 (Table 2). Models: Gen. = generator, AE = autoencoder, Dec. = decoder, VFM = vision foundation model; Repr.: Pixel = pixel diffusion, Fixed = a fixed latent representation, Dynamic = a dynamically evolved latent representation in training; Aux. data = external training data beyond ImageNet; Training FLOPs (×1019) is the generator-only training cost and does not include the cost of training a separate AE or VFM.

Table 1. Class-conditional ImageNet generation at 256×256.

Method Models Repr. Total params (M) Epochs Aux. data Training FLOPs FID ↓ IS ↑
Two-stage
DiT-XL/2 Gen.+AE Fixed 759 1400 OpenImages 45.4 2.27 278
SiT-XL/2 Gen.+AE Fixed 759 1400 OpenImages 45.4 2.06 270
LightningDiT Gen.+AE+VFM Fixed 745 800 – 19.1 1.35 295
REPA-SiT Gen.+AE+VFM Fixed 759 800 OpenImages 25.9 1.29 306
DDT-XL/2 Gen.+AE+VFM Fixed 759 400 OpenImages 19.2 1.26 311
REPA-E (tuning) Gen.+AE+VFM Fixed 759 800 OpenImages 57.8 1.12 303
SVG-XL Gen.+AE+VFM Fixed 758 1400 DINOv3 22.8 1.92 265
RAE-DiTDH Gen.+AE+VFM Fixed 839 800 DINOv2 – 1.13 263
One-stage
REPA-E (scratch) Gen.+AE+VFM Dynamic 759 80 – 5.78 1.67 –
UNITE-XL Gen.+Dec. Dynamic 763 240 – 12.0 1.75 310
DSD Gen.+VFM Dynamic 205 50 – – 3.35 255
ADM-U Gen. Pixel 554 400 – – 4.59 187
RIN Gen. Pixel 410 480 – 20.5 3.42 182
PixNerd Gen.+VFM Pixel 700 160 – 5.49 2.15 297
PixelFlow Gen. Pixel 677 320 – 239 1.98 282
JiT-H/16 Gen. Pixel 953 600 – 14.0 1.86 303
LDM-is-AE (Ours) Gen. Dynamic 961 300 – 7.02 1.80 314

Table 2. Class-conditional ImageNet generation at 512×512.

Method Models Repr. Total params (M) FID ↓ IS ↑
Two-stage
DiT-XL/2 Gen.+AE Fixed 759 3.04 241
SiT-XL/2 Gen.+AE Fixed 759 2.62 252
REPA-SiT-XL/2 Gen.+AE+VFM Fixed 759 2.08 275
One-stage
ADM-G Gen. Pixel 559 7.72 173
RIN Gen. Pixel 320 3.95 216
PixNerd-XL/16 Gen.+VFM Pixel 700 2.84 246
DeCo Gen. Pixel 682 2.22 290
JiT-H/32 Gen. Pixel 956 1.94 309
LDM-is-AE (Ours) Gen. Dynamic 961 1.90 320

🤗 Model Weights

https://huggingface.co/xtudbxk/LDMisAE

File Resolution Size
LDMisAE.256.ckpt 256×256 3.90 GB
LDMisAE.512.ckpt 512×512 3.93 GB

📝 Citation

@inproceedings{ldm_is_ae_2026,
  title     = {LDM-is-AE: Latent Diffusion Model is an Intrinsic Auto-Encoder for End-to-End Image Generation},
  author    = {Zhang, Zhengqiang and Sun, Lingchen and Wu, Rongyuan and Yi, Qiaosi and
               Kong, Xiangtao and Xiao, Chaodong and Zhang, Lei},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}

⚖️ License

Code: Apache License 2.0 -- see LICENSE. Model weights and data are released separately and are intended for research use. Third-party components are acknowledged in NOTICE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support