Regularizing Layer Fusion Closes the Reconstruction-Generation Gap in Representation Autoencoders

Official checkpoints for FuseReg. FuseReg replaces heuristic encoder-layer fusion in representation autoencoders (RAEs) with training over random subsets of encoder layers, so a single pixel decoder reconstructs reliably under any layer-subset fusion and pairs better with the diffusion generator.

Paper: arXiv:2609.31620 · Project Page: FuseReg · Code: Hongyang-Du/FuseReg

Repository layout

dinov3-vitl/                      # DINOv3 ViT-L encoder, ImageNet 256x256
  decoder_k23/
    p0.00.safetensors             # reproduced RAEv2 decoder (no layer drop)
    p0.05.safetensors ... p0.95.safetensors
  decoder_k7/
    p0.60.safetensors
    p0.90.safetensors
  ditxl_k23/
    raev2_pdit0.0_ep040.safetensors
    raev2_pdit0.0_ep080.safetensors
    fusereg_pdit0.3_ep040.safetensors ... fusereg_pdit0.9_ep040.safetensors
eupe-vitb/                        # no checkpoints published yet
siglip/                           # no checkpoints published yet

The repository currently contains 16 inference checkpoints: eight DINOv3-L k=23 decoders, two DINOv3-L k=7 decoders, and six DiT-XL generators. The k=7 p=0.3, four EUPE, and four SigLIP variants have not been published. A 25-file candidate collection is tracked in the checkpoint manifest; only entries in its files array are available for download.

Checkpoint manifest records file size, SHA-256, selected EMA state, and the scope of validation. Fourteen pre-existing files have been checked against source checkpoint epoch/step and EMA tensor names/shapes. The k=23 p=0.95 source archive also matches the paper provenance SHA-256, and all 456 EMA tensors equal the published file. Both newly published k=7 decoders were checked against Drive CRC32C, round-trip tensor equality, and Hub SHA-256.

All checkpoint files contain EMA weights only in safetensors format (no optimizer, discriminator, or encoder weights). The safetensors metadata records available training provenance such as encoder, drop rate, epoch, and step; some older files do not include layers. Use the matching evaluation configuration and checkpoint manifest for fusion layers.

All numerical results below are paper-reported, not independently reproduced during checkpoint packaging. File integrity and tensor equality checks do not validate PSNR, SSIM, rFID, or gFID. Full metric reproduction requires the exact ImageNet evaluation split, encoder checkpoints, preprocessing, and latent statistics.

DINOv3 ViT-L

Pixel decoders (dinov3-vitl/decoder_k23/)

ViT decoder (hidden size 1152, 415.6M params), trained for 16 epochs on ImageNet-256 with pixel, LPIPS, and adversarial losses, over DINOv3-L layers 1-23 with random layer-drop rate p. Paper-reported reconstruction on ImageNet val (50k) under three inference-time fusions:

File p feed k=7 PSNR / SSIM / rFID feed k=23 PSNR / SSIM / rFID feed l11 PSNR / SSIM / rFID
p0.00 (RAEv2 reproduced) 0 12.54 / 0.372 / 16.108 27.10 / 0.808 / 0.189 15.81 / 0.503 / 3.392
p0.05 0.05 19.78 / 0.518 / 2.826 28.59 / 0.853 / 0.273 20.76 / 0.557 / 1.881
p0.10 0.1 20.52 / 0.550 / 1.718 28.58 / 0.852 / 0.285 21.33 / 0.586 / 1.209
p0.30 0.3 22.05 / 0.612 / 0.762 28.48 / 0.849 / 0.294 22.90 / 0.657 / 0.576
p0.50 0.5 22.93 / 0.646 / 0.655 28.43 / 0.848 / 0.294 23.98 / 0.695 / 0.522
p0.70 0.7 23.47 / 0.665 / 0.634 28.20 / 0.842 / 0.322 24.58 / 0.714 / 0.493
p0.90 0.9 23.62 / 0.671 / 0.610 27.60 / 0.827 / 0.415 24.82 / 0.722 / 0.458
p0.95 (recommended) 0.95 23.77 / 0.678 / 0.604 27.52 / 0.826 / 0.421 25.13 / 0.735 / 0.455

Pixel decoders trained on k=7 (dinov3-vitl/decoder_k7/)

Published variants: p0.60.safetensors and p0.90.safetensors. Their source checkpoint metadata identifies the seven encoder layers as [11, 13, 15, 17, 19, 21, 23]; both were saved at epoch 16. These are distinct from feeding a k=23-trained decoder with the same seven-layer subset. No new reconstruction metrics are claimed for the packaged files.

DiT-XL generators (dinov3-vitl/ditxl_k23/)

Class-conditional ImageNet-256 DiT-XL (hidden size 1440, 875.3M params) generating in the k=23 DINOv3-L latent.

File Generator p_dit Epochs gFID w/ p0.95 decoder (no guid. / IG 1.78)
raev2_pdit0.0_ep040 RAEv2 0 40
raev2_pdit0.0_ep080 RAEv2 0 80
fusereg_pdit0.3_ep040 FuseReg 0.3 40 not in paper grid
fusereg_pdit0.5_ep040 FuseReg 0.5 40 2.45 / 1.59
fusereg_pdit0.7_ep040 FuseReg 0.7 40 2.42 / 1.58
fusereg_pdit0.9_ep040 FuseReg 0.9 40 2.39 / 1.53

FuseReg DiT-XL generators are trained with random layer-drop rate p_dit over the same k=23 latent; the full p_dit x p_dec grid is in the paper appendix (DiT-XL rate sweep).

The paper reports that pairing a fixed RAEv2 k=23 generator with the FuseReg p0.95 decoder reduces unguided gFID from 3.01 to 2.21 while keeping guided gFID at 1.25 (internal guidance 1.78, 50k samples, 50 Euler steps).

Usage

from huggingface_hub import hf_hub_download
from safetensors.torch import load_file

path = hf_hub_download(
    "Hongyang-Du/FuseReg",
    "dinov3-vitl/decoder_k23/p0.95.safetensors"
)
state_dict = load_file(path)
decoder.load_state_dict(state_dict)   # decoder built from the FuseReg codebase

The full RAE checkpoint used for k=23 p=0.00 stores decoder weights under ema with a decoder. prefix. The published file contains that decoder subtree. Its source archive does not record output normalization settings; confirm the original RAE evaluation path before using it for a numerical baseline comparison. The trained drop decoders and the original RAE baseline must each use the appropriate output normalization. The optional CLS-surrogate operation in the code is an additive last-layer token mean, not a classifier-surrogate loss.

The encoder and latent normalization statistics are not included. Obtain DINOv3 ViT-L from Meta under the DINOv3 License.

Citation

@misc{du2026fuseregregularizinglayerfusion,
      title={FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders},
      author={Hongyang Du and Yunfei Xie and Junjie Ye and Jiawei Yang and Xiaoyan Cong and Haodong Zhang and Yongchao Huang and Haiyu Wu and Zongxia Li and Shihang Gui and Dawei Liu and Runhao Li and Jingcheng Ni and Chen Wei and Randall Balestriero and Yue Wang},
      year={2026},
      eprint={2609.31620},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.31620},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Hongyang-Du/FuseReg

Paper for Hongyang-Du/FuseReg