Depth-Anything-V2-Small — ONNX (monocular photo depth)

Depth-Anything-V2-Small (Yang et al., 2024) exported to ONNX. Relative (inverse) depth from a single photograph. All credit to the Depth Anything authors.

Mirrored for QtMeshEditor (issue #1018, epic #818 Track C4), where it depth-conditions the Stable Diffusion ControlNet texture path from a reference photo rather than only from mesh-rendered depth: qtmesh material --photo-depth, the Material Editor's "Depth from photo…" button, and the photo_depth MCP tool.

The file QtMeshEditor downloads at runtime lives in the shared fernandotonon/QtMeshEditor-models repo under depth/. This repo is the standalone model card + mirror.

Why only the Small variant

Small is Apache-2.0. Base and Large are cc-by-nc-4.0 and are deliberately NOT mirrored here or shipped by QtMeshEditor — non-commercial terms fail the project's permissive-redistribution bar (the same reason SF3D and Hunyuan3D were rejected). The exporter enforces this in code: it queries the HF model API and hard-fails on any licence outside its allow-list before downloading weights.

I/O

input   pixel_values     float32 [1,3,518,518]   ImageNet-normalised RGB (NCHW)
output  predicted_depth  float32 [1,518,518]     relative inverse depth

Larger = NEARER. QtMeshEditor normalises to 8-bit grayscale with near = bright, matching the convention its mesh-depth renderer emits, so a photo and a rendered mesh are interchangeable as ControlNet conditioning images.

Normalisation constants (upstream preprocessor_config.json): mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225].

The spatial dims are PINNED, not dynamic

Only the batch axis is dynamic. This is deliberate: a dynamic-axis export traces cleanly but drifts away from the traced resolution, because the ViT position-embedding interpolation is only partly captured. Measured against torch, max abs difference on a smooth image:

input relative error
518×518 (traced) 2.7e-06
462×462 2.3e-02
392×392 4.9e-02

An earlier dynamic export was worse still — it advertised dynamic height/width but froze the output size, returning a 518×518 map for a 462×462 input (int() casts in the HF head bake the interpolate target into a constant).

Since the reference preprocessor resizes every input to 518 anyway (size: 518, keep_aspect_ratio, ensure_multiple_of: 14), pinning costs nothing and removes a silent-wrong-answer failure mode. Resize to 518×518 before inference.

Parity

The export asserts torch-vs-onnxruntime agreement before upload, on both random noise and a smooth structured image, and refuses to write a graph containing any non-finite output:

[parity] torch vs onnxruntime
  shape (1, 518, 518)  max|diff| 4.053e-06
[parity] smooth image  max|diff| 6.914e-06

Export script: scripts/export-depth-anything-onnx.py (opset 18).

License

Apache-2.0, inherited from depth-anything/Depth-Anything-V2-Small-hf. Credit: Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for fernandotonon/QtMeshEditor-depthanything-onnx

Adapter
(4)
this model