OVIE-ft β€” DL3DV fine-tune

In-the-Wild Monocular Pretraining for Novel View Generation

Paper GitHub Collection

Part of the OVIE collection.

This is kyutai/ovie fine-tuned for 50k steps on DL3DV, the companion to the RealEstate10K fine-tune. The architecture is identical to the base model; only the weights differ.

It is the control that shows fine-tuning selects a domain: applying the same protocol on DL3DV markedly improves in-domain reconstruction relative to the base model, but costs distributional realism and generalization.

Metrics

Checkpoint DL3DV PSNR ↑ DL3DV SSIM ↑ DL3DV LPIPS ↓ DL3DV FID ↓ DL3DV MEt3R ↓ RealEstate10K PSNR ↑ RealEstate10K FID ↓
base OVIE (out-of-domain on DL3DV) 14.8 0.369 0.464 13.6 0.078 18.8 6.74
this checkpoint (in-domain on DL3DV) 17.07 0.433 0.368 17.24 0.059 19.34 7.14

Fine-tuning raises DL3DV PSNR by 2.3 dB while raising FID from 13.6 to 17.24: the model becomes a more faithful per-view reconstructor at a slight cost to global distributional realism. Notably it retains β€” and on RealEstate10K slightly improves β€” cross-domain performance even though that benchmark is now out-of-domain (PSNR 19.34 vs 18.8), so the fine-tune sharpens in-domain fidelity without eroding generalization here.

Prefer the base OVIE for in-the-wild zero-shot use and the RealEstate10K fine-tune for indoor real-estate scenes. This checkpoint is intended for DL3DV-style captures.

Usage

import torch
from models.models import OVIEModel

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model = OVIEModel.from_pretrained("kyutai/ovie-ft-dl3dv").to(device)
model.eval()
image_size = model.image_size  # 256

The forward pass is identical to the base model:

with torch.no_grad():
    pred = model(x=img_tensor, cam_params=cam_token)  # (1, 3, 256, 256) in [0, 1]

See the base model card for the complete pose-encoding snippet, and the repository for installation and the full benchmark table.

Citation

@misc{ovie2026,
      title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
      author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
      year={2026},
      eprint={2603.23488},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2603.23488},
}
Downloads last month
35
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for kyutai/ovie-ft-dl3dv

Base model

kyutai/ovie
Finetuned
(2)
this model

Collection including kyutai/ovie-ft-dl3dv

Paper for kyutai/ovie-ft-dl3dv