Nacre v1 β€” provenance-clean 4Γ— real-world super-resolution

A clean-room, ResShift-class generative upscaler: a 118.6M-parameter SwinUNet denoiser driving a 4-step residual-shift diffusion in the latent space of the CompVis VQ-f4 autoencoder. It takes a degraded low-resolution photo (blur, noise, JPEG, resize β€” the Real-ESRGAN two-stage family) and returns a 4Γ— restoration with plausible fine detail.

Why it exists: every comparable generative SR model is blocked somewhere in its stack β€” non-commercial weights (ResShift, StableSR, SUPIR), a use-restricted base (SD2.1-based PiSA-SR / OSEDiff / SeeSR), or an undisclosed training corpus. Nacre is permissive all the way down:

Component Licence Provenance
Code (architecture + trainer) Apache-2.0 written from the papers; no reference code copied
Denoiser weights (nacre_v1.safetensors) Apache-2.0 trained from scratch by Xocialize
Autoencoder (vq_f4.safetensors) MIT CompVis latent-diffusion VQ-f4, unmodified (training head removed)
Training images CC0 / Public domain / CC BY family 51,759 Wikimedia Commons images, per-file licences in ATTRIBUTION.csv

Results β€” neutral benchmarks, vs the released ResShift v3

Same inputs, same degradation draws and same chain noise for both models; each run through its own autoencoder. Real-image sets are centre-cropped to 256 px LQ (β†’ 1024 px output).

Set PSNR-Y ↑ SSIM-Y ↑ LPIPS ↓ ΒΉ CLIPIQA ↑ MUSIQ ↑
RealSR V3 Γ—4 (100 real camera pairs) 23.98 vs 23.93 0.705 vs 0.687 0.318 vs 0.364 0.572 vs 0.574 63.5 vs 59.3
DIV2K-val (200 crops, synthetic damage) 22.85 vs 22.86 0.579 vs 0.571 0.347 vs 0.356 0.517 vs 0.549 58.2 vs 56.8
RealSet80 (no ground truth) β€” β€” β€” 0.606 vs 0.648 66.3 vs 64.1

At or above ResShift on fidelity everywhere and on MUSIQ everywhere; behind on CLIPIQA on DIV2K and RealSet by 0.03–0.04. ΒΉ LPIPS is part of Nacre's training loss, so its lead there is not independent evidence.

Usage

Inputs are RGB in [0, 1]. The network is conditioned on the LQ image in pixel space (scaled to [-1, 1]); the LQ's latent β€” AE.encode(bicubic_x4(lq)) β€” is only the endpoint of the residual-shift chain. Pad LQ sides to a multiple of 64 (four U-Net levels, then 8-wide shifted windows) and crop the output. config.json carries the full model, diffusion and preprocessing parameters. Reference implementation: nacre/ (PyTorch, Apache-2.0). Apple silicon: the Swift/MLX port xocialize/mlx-nacre-swift (an MLXEngine imageUpscale package; 131 dB parity with this reference) with MLX-layout weights at xocialize/nacre-v1-mlx.

Training

  • Data: 122,021 lossless 512Β² tiles from 51,759 Wikimedia Commons photos (CC0/PD/CC BY), filtered with blockiness / HyperIQA / bytes-per-pixel floors (worst 10/5/5 % removed), ≀ 4 tiles per image. Every source image was downscaled by a per-image factor in [0.15, 0.45] before tiling: native-resolution camera frames are soft at the pixel level, and a model trained on them learns to produce soft output.
  • Degradations: Real-ESRGAN two-stage (blur, resize, noise, JPEG, sinc), synthesised on the fly.
  • Schedule: batch 96, AdamW 5e-5 β†’ 2e-5 cosine, EMA 0.999, latent MSE + LPIPS (weight 2Γ— for the first 100k steps, 1Γ— after). Selected checkpoint: step 110k (EMA), by a selection rule fixed before the candidates were scored (fidelity floor first, then realism by CLIPIQA + MUSIQ; LPIPS excluded as in-loss).
  • Compute: ~300k steps on one H200 (Hugging Face Jobs), β‰ˆ 0.43 it/s.

Limitations

  • Generative detail is invented, not recovered. On heavily degraded input the model synthesises plausible texture; it can differ from the true scene. Do not use it where exact detail matters (forensics, measurement, medical).
  • Small text β€” same as ResShift, and a guard is required. On SA-Text (200 images, Apple Vision OCR word-F1, heaviest degradation lv3 β†’ Γ—4): Nacre 0.125 vs ResShift 0.129 vs bicubic 0.111 (Nacre βˆ’ ResShift βˆ’0.004 Β± 0.014, a tie); lighter lv1: 0.229 vs 0.244 vs 0.231. Like every generative upscaler measured, it damages text that was already legible (βˆ’0.15 F1 on the 25 images bicubic already reads; ResShift βˆ’0.13, worse than bicubic on 7/25 for both). Route already-legible small text to a fidelity upscaler or mask it out.
  • Trained at 4Γ— only; motion blur is not in the degradation set.

Attribution

Every training image is listed in ATTRIBUTION.csv with its Commons title, URL, licence and author β€” 51,759 images: 22,427 CC BY 4.0 Β· 11,506 CC0 Β· 9,355 CC BY 3.0 Β· 6,229 public domain Β· and smaller CC BY 2.5 / 3.0 de / 2.0 groups. Every CC BY image carries a named author (resolved from the file page where the original harvest truncated it). On 2026-10-05 the current Commons licence of all 51,759 files was re-checked: no licence has changed and no file has been removed.

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support