Regularizing Layer Fusion Closes the Reconstruction-Generation Gap in Representation Autoencoders
Official checkpoints for FuseReg. FuseReg replaces heuristic encoder-layer fusion in representation autoencoders (RAEs) with training over random subsets of encoder layers, so a single pixel decoder reconstructs reliably under any layer-subset fusion and pairs better with the diffusion generator.
Paper: arXiv:2609.31620 · Project Page: FuseReg · Code: Hongyang-Du/FuseReg
Repository layout
dinov3-vitl/ # DINOv3 ViT-L encoder, ImageNet 256x256
decoder_k23/
p0.00.safetensors # reproduced RAEv2 decoder (no layer drop)
p0.05.safetensors ... p0.95.safetensors
decoder_k7/
p0.60.safetensors
p0.90.safetensors
ditxl_k23/
raev2_pdit0.0_ep040.safetensors
raev2_pdit0.0_ep080.safetensors
fusereg_pdit0.3_ep040.safetensors ... fusereg_pdit0.9_ep040.safetensors
eupe-vitb/ # no checkpoints published yet
siglip/ # no checkpoints published yet
The repository currently contains 16 inference checkpoints: eight DINOv3-L
k=23 decoders, two DINOv3-L k=7 decoders, and six DiT-XL generators. The k=7
p=0.3, four EUPE, and four SigLIP variants have not been published. A 25-file
candidate collection is tracked in the checkpoint manifest; only entries in its
files array are available for download.
Checkpoint manifest records file size, SHA-256,
selected EMA state, and the scope of validation. Fourteen pre-existing files have
been checked against source checkpoint epoch/step and EMA tensor names/shapes.
The k=23 p=0.95 source archive also matches the paper provenance SHA-256, and all
456 EMA tensors equal the published file. Both newly published k=7 decoders were
checked against Drive CRC32C, round-trip tensor equality, and Hub SHA-256.
All checkpoint files contain EMA weights only in safetensors format (no optimizer,
discriminator, or encoder weights). The safetensors metadata records available training provenance such as
encoder, drop rate, epoch, and step; some older files do not include layers.
Use the matching evaluation configuration and checkpoint manifest for fusion layers.
All numerical results below are paper-reported, not independently reproduced during checkpoint packaging. File integrity and tensor equality checks do not validate PSNR, SSIM, rFID, or gFID. Full metric reproduction requires the exact ImageNet evaluation split, encoder checkpoints, preprocessing, and latent statistics.
DINOv3 ViT-L
Pixel decoders (dinov3-vitl/decoder_k23/)
ViT decoder (hidden size 1152, 415.6M params), trained for 16 epochs on ImageNet-256
with pixel, LPIPS, and adversarial losses, over DINOv3-L layers
1-23 with random layer-drop rate p. Paper-reported reconstruction on ImageNet val (50k) under three
inference-time fusions:
| File | p | feed k=7 PSNR / SSIM / rFID | feed k=23 PSNR / SSIM / rFID | feed l11 PSNR / SSIM / rFID |
|---|---|---|---|---|
p0.00 (RAEv2 reproduced) |
0 | 12.54 / 0.372 / 16.108 | 27.10 / 0.808 / 0.189 | 15.81 / 0.503 / 3.392 |
p0.05 |
0.05 | 19.78 / 0.518 / 2.826 | 28.59 / 0.853 / 0.273 | 20.76 / 0.557 / 1.881 |
p0.10 |
0.1 | 20.52 / 0.550 / 1.718 | 28.58 / 0.852 / 0.285 | 21.33 / 0.586 / 1.209 |
p0.30 |
0.3 | 22.05 / 0.612 / 0.762 | 28.48 / 0.849 / 0.294 | 22.90 / 0.657 / 0.576 |
p0.50 |
0.5 | 22.93 / 0.646 / 0.655 | 28.43 / 0.848 / 0.294 | 23.98 / 0.695 / 0.522 |
p0.70 |
0.7 | 23.47 / 0.665 / 0.634 | 28.20 / 0.842 / 0.322 | 24.58 / 0.714 / 0.493 |
p0.90 |
0.9 | 23.62 / 0.671 / 0.610 | 27.60 / 0.827 / 0.415 | 24.82 / 0.722 / 0.458 |
p0.95 (recommended) |
0.95 | 23.77 / 0.678 / 0.604 | 27.52 / 0.826 / 0.421 | 25.13 / 0.735 / 0.455 |
Pixel decoders trained on k=7 (dinov3-vitl/decoder_k7/)
Published variants: p0.60.safetensors and p0.90.safetensors. Their source
checkpoint metadata identifies the seven encoder layers as
[11, 13, 15, 17, 19, 21, 23]; both were saved at epoch 16. These are distinct
from feeding a k=23-trained decoder with the same seven-layer subset. No new
reconstruction metrics are claimed for the packaged files.
DiT-XL generators (dinov3-vitl/ditxl_k23/)
Class-conditional ImageNet-256 DiT-XL (hidden size 1440, 875.3M params) generating in the k=23 DINOv3-L latent.
| File | Generator | p_dit | Epochs | gFID w/ p0.95 decoder (no guid. / IG 1.78) |
|---|---|---|---|---|
raev2_pdit0.0_ep040 |
RAEv2 | 0 | 40 | |
raev2_pdit0.0_ep080 |
RAEv2 | 0 | 80 | |
fusereg_pdit0.3_ep040 |
FuseReg | 0.3 | 40 | not in paper grid |
fusereg_pdit0.5_ep040 |
FuseReg | 0.5 | 40 | 2.45 / 1.59 |
fusereg_pdit0.7_ep040 |
FuseReg | 0.7 | 40 | 2.42 / 1.58 |
fusereg_pdit0.9_ep040 |
FuseReg | 0.9 | 40 | 2.39 / 1.53 |
FuseReg DiT-XL generators are trained with random layer-drop rate p_dit over the same
k=23 latent; the full p_dit x p_dec grid is in the paper appendix (DiT-XL rate sweep).
The paper reports that pairing a fixed RAEv2 k=23 generator with the FuseReg p0.95 decoder reduces
unguided gFID from 3.01 to 2.21 while keeping guided gFID at 1.25 (internal guidance
1.78, 50k samples, 50 Euler steps).
Usage
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
path = hf_hub_download(
"Hongyang-Du/FuseReg",
"dinov3-vitl/decoder_k23/p0.95.safetensors"
)
state_dict = load_file(path)
decoder.load_state_dict(state_dict) # decoder built from the FuseReg codebase
The full RAE checkpoint used for k=23 p=0.00 stores decoder weights under
ema with a decoder. prefix. The published file contains that decoder subtree.
Its source archive does not record output normalization settings; confirm the
original RAE evaluation path before using it for a numerical baseline comparison.
The trained drop decoders and the original RAE baseline must each use the
appropriate output normalization. The optional CLS-surrogate operation in the
code is an additive last-layer token mean, not a classifier-surrogate loss.
The encoder and latent normalization statistics are not included. Obtain DINOv3 ViT-L from Meta under the DINOv3 License.
Citation
@misc{du2026fuseregregularizinglayerfusion,
title={FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders},
author={Hongyang Du and Yunfei Xie and Junjie Ye and Jiawei Yang and Xiaoyan Cong and Haodong Zhang and Yongchao Huang and Haiyu Wu and Zongxia Li and Shihang Gui and Dawei Liu and Runhao Li and Jingcheng Ni and Chen Wei and Randall Balestriero and Yue Wang},
year={2026},
eprint={2609.31620},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.31620},
}