OVIE-ft β DL3DV fine-tune
In-the-Wild Monocular Pretraining for Novel View Generation
Part of the OVIE collection.
This is kyutai/ovie fine-tuned for 50k steps on DL3DV, the companion to the RealEstate10K fine-tune. The architecture is identical to the base model; only the weights differ.
It is the control that shows fine-tuning selects a domain: applying the same protocol on DL3DV markedly improves in-domain reconstruction relative to the base model, but costs distributional realism and generalization.
Metrics
| Checkpoint | DL3DV PSNR β | DL3DV SSIM β | DL3DV LPIPS β | DL3DV FID β | DL3DV MEt3R β | RealEstate10K PSNR β | RealEstate10K FID β |
|---|---|---|---|---|---|---|---|
| base OVIE (out-of-domain on DL3DV) | 14.8 | 0.369 | 0.464 | 13.6 | 0.078 | 18.8 | 6.74 |
| this checkpoint (in-domain on DL3DV) | 17.07 | 0.433 | 0.368 | 17.24 | 0.059 | 19.34 | 7.14 |
Fine-tuning raises DL3DV PSNR by 2.3 dB while raising FID from 13.6 to 17.24: the model becomes a more faithful per-view reconstructor at a slight cost to global distributional realism. Notably it retains β and on RealEstate10K slightly improves β cross-domain performance even though that benchmark is now out-of-domain (PSNR 19.34 vs 18.8), so the fine-tune sharpens in-domain fidelity without eroding generalization here.
Prefer the base OVIE for in-the-wild zero-shot use and the RealEstate10K fine-tune for indoor real-estate scenes. This checkpoint is intended for DL3DV-style captures.
Usage
import torch
from models.models import OVIEModel
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = OVIEModel.from_pretrained("kyutai/ovie-ft-dl3dv").to(device)
model.eval()
image_size = model.image_size # 256
The forward pass is identical to the base model:
with torch.no_grad():
pred = model(x=img_tensor, cam_params=cam_token) # (1, 3, 256, 256) in [0, 1]
See the base model card for the complete pose-encoding snippet, and the repository for installation and the full benchmark table.
Citation
@misc{ovie2026,
title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
year={2026},
eprint={2603.23488},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2603.23488},
}
- Downloads last month
- 35
Model tree for kyutai/ovie-ft-dl3dv
Base model
kyutai/ovie