Nacre v1 β provenance-clean 4Γ real-world super-resolution
A clean-room, ResShift-class generative upscaler: a 118.6M-parameter SwinUNet denoiser driving a 4-step residual-shift diffusion in the latent space of the CompVis VQ-f4 autoencoder. It takes a degraded low-resolution photo (blur, noise, JPEG, resize β the Real-ESRGAN two-stage family) and returns a 4Γ restoration with plausible fine detail.
Why it exists: every comparable generative SR model is blocked somewhere in its stack β non-commercial weights (ResShift, StableSR, SUPIR), a use-restricted base (SD2.1-based PiSA-SR / OSEDiff / SeeSR), or an undisclosed training corpus. Nacre is permissive all the way down:
| Component | Licence | Provenance |
|---|---|---|
| Code (architecture + trainer) | Apache-2.0 | written from the papers; no reference code copied |
Denoiser weights (nacre_v1.safetensors) |
Apache-2.0 | trained from scratch by Xocialize |
Autoencoder (vq_f4.safetensors) |
MIT | CompVis latent-diffusion VQ-f4, unmodified (training head removed) |
| Training images | CC0 / Public domain / CC BY family | 51,759 Wikimedia Commons images, per-file licences in ATTRIBUTION.csv |
Results β neutral benchmarks, vs the released ResShift v3
Same inputs, same degradation draws and same chain noise for both models; each run through its own autoencoder. Real-image sets are centre-cropped to 256 px LQ (β 1024 px output).
| Set | PSNR-Y β | SSIM-Y β | LPIPS β ΒΉ | CLIPIQA β | MUSIQ β |
|---|---|---|---|---|---|
| RealSR V3 Γ4 (100 real camera pairs) | 23.98 vs 23.93 | 0.705 vs 0.687 | 0.318 vs 0.364 | 0.572 vs 0.574 | 63.5 vs 59.3 |
| DIV2K-val (200 crops, synthetic damage) | 22.85 vs 22.86 | 0.579 vs 0.571 | 0.347 vs 0.356 | 0.517 vs 0.549 | 58.2 vs 56.8 |
| RealSet80 (no ground truth) | β | β | β | 0.606 vs 0.648 | 66.3 vs 64.1 |
At or above ResShift on fidelity everywhere and on MUSIQ everywhere; behind on CLIPIQA on DIV2K and RealSet by 0.03β0.04. ΒΉ LPIPS is part of Nacre's training loss, so its lead there is not independent evidence.
Usage
Inputs are RGB in [0, 1]. The network is conditioned on the LQ image in pixel space (scaled to [-1, 1]);
the LQ's latent β AE.encode(bicubic_x4(lq)) β is only the endpoint of the residual-shift chain.
Pad LQ sides to a multiple of 64 (four U-Net levels, then 8-wide shifted windows) and crop the output.
config.json carries the full model, diffusion and preprocessing parameters. Reference implementation:
nacre/ (PyTorch, Apache-2.0). Apple silicon: the Swift/MLX port xocialize/mlx-nacre-swift (an MLXEngine imageUpscale package; 131 dB parity with this reference) with MLX-layout weights at xocialize/nacre-v1-mlx.
Training
- Data: 122,021 lossless 512Β² tiles from 51,759 Wikimedia Commons photos (CC0/PD/CC BY), filtered with blockiness / HyperIQA / bytes-per-pixel floors (worst 10/5/5 % removed), β€ 4 tiles per image. Every source image was downscaled by a per-image factor in [0.15, 0.45] before tiling: native-resolution camera frames are soft at the pixel level, and a model trained on them learns to produce soft output.
- Degradations: Real-ESRGAN two-stage (blur, resize, noise, JPEG, sinc), synthesised on the fly.
- Schedule: batch 96, AdamW 5e-5 β 2e-5 cosine, EMA 0.999, latent MSE + LPIPS (weight 2Γ for the first 100k steps, 1Γ after). Selected checkpoint: step 110k (EMA), by a selection rule fixed before the candidates were scored (fidelity floor first, then realism by CLIPIQA + MUSIQ; LPIPS excluded as in-loss).
- Compute: ~300k steps on one H200 (Hugging Face Jobs), β 0.43 it/s.
Limitations
- Generative detail is invented, not recovered. On heavily degraded input the model synthesises plausible texture; it can differ from the true scene. Do not use it where exact detail matters (forensics, measurement, medical).
- Small text β same as ResShift, and a guard is required. On SA-Text (200 images, Apple Vision OCR word-F1, heaviest degradation lv3 β Γ4): Nacre 0.125 vs ResShift 0.129 vs bicubic 0.111 (Nacre β ResShift β0.004 Β± 0.014, a tie); lighter lv1: 0.229 vs 0.244 vs 0.231. Like every generative upscaler measured, it damages text that was already legible (β0.15 F1 on the 25 images bicubic already reads; ResShift β0.13, worse than bicubic on 7/25 for both). Route already-legible small text to a fidelity upscaler or mask it out.
- Trained at 4Γ only; motion blur is not in the degradation set.
Attribution
Every training image is listed in ATTRIBUTION.csv with its Commons title, URL, licence and author β 51,759 images:
22,427 CC BY 4.0 Β· 11,506 CC0 Β· 9,355 CC BY 3.0 Β· 6,229 public domain Β· and smaller CC BY 2.5 / 3.0 de / 2.0 groups.
Every CC BY image carries a named author (resolved from the file page where the original harvest truncated it).
On 2026-10-05 the current Commons licence of all 51,759 files was re-checked: no licence has changed and no file has
been removed.
- Downloads last month
- 14