ZipDepth Base FP16 — ExecuTorch

This repository packages the unfold-free ZipDepth Base mobile/NPU checkpoint as a static 384×384 ExecuTorch model. The depth network uses FP16 activations with FP32 serialized convolution weights and external tensors.

✨ Key Highlights

  • 2.06× faster on Android™ Vivo X300 — one-thread p50 depth-model latency is 20.431 ms, compared with 41.987 ms for FP32 under the same inference-only protocol.
  • Negligible quality change — across all 654 NYUv2 test images, δ1 changes by only −0.046 percentage points and AbsRel improves by 0.011 percentage points relative to the FP32 depth model under identical fixed-shape preprocessing and evaluation.

📦 Model Details

ZipDepth predicts affine-invariant inverse depth from one RGB image. The package accepts arbitrary image dimensions up to 4096×4096, crops the largest top-left square, resizes it to the static 384×384 depth-model input, and restores the result to the original tensor shape with zero padding.

  • Developed by: Fabio Tosi, Luca Bartolomei, Matteo Poggi, and Stefano Mattoccia
  • Model type: Zero-shot monocular relative-depth estimation
  • License: MIT
  • Base model: ZipDepth
  • Source revision: a302e543
  • Source checkpoint: zipdepth_base_npu.pth
  • ExecuTorch revision: d13c78971338b35a219e0a27f7045b29c7f4e8c1
  • Parameters: Approximately 6.1 million after inference-time fusion
  • Artifact size: 23.47 MiB, 0.27% smaller than the 23.53 MiB FP32 artifact

🚀 Get Started with the Model

🔓 Compute Flow — Early Access

The inference engine for this model package is available through the Compute Flow Early Access Program. Contact ai-early-access@arm.com to request access.

📊 Quality evaluation

Quality was evaluated on all 654 images in the NYUv2 test split. Both models received the same top-left square crop resized to 384×384. Their predictions were resized to the ground-truth shape, aligned independently in inverse-depth space using a least-squares scale and shift, and evaluated over the NYUv2 Eigen crop and 0.001–10 m depth range. This isolates the static depth models; it is not an evaluation of the dynamic prepare/restore transforms.

Metric FP32 FP16 ExecuTorch Change
AbsRel ↓ 15.796% 15.785% 0.011 percentage points lower
δ1 ↑ 77.055% 77.009% 0.046 percentage points lower

The FP16 candidate uses a different official ZipDepth checkpoint because the mobile/NPU checkpoint replaces convex unfold upsampling with a deployment-friendly learned blend. The comparison therefore measures the complete mobile graph and FP16 conversion, not FP16 rounding in isolation. On the distributed 384×384 sample, its output has 0.99645 Pearson correlation with the FP32 output.

🎯 Performance evaluation

The static depth models were profiled using the distributed samples/im0.jpg image, which is already 384×384. The dynamic prepare and restore models were therefore bypassed.

  • CPU: One Arm CPU thread.
  • Runs: One unmeasured in-process inference to initialize each runtime, followed by 5 warmup and 30 measured fresh-process runs with 1 second between processes.
  • Timing scope: The setup call is not included.
  • Runtimes: LiteRT 2.1.6 FP32 baseline and ExecuTorch d13c7897 FP16 candidate, both with the XNNPACK CPU backend.
  • Target: Android™ Vivo X300 smartphone (V2502A) on Android™ 16, USB powered at 100% battery.
Metric Android™ Vivo X300: FP32 Android™ Vivo X300: FP16 Uplift
p50 inference latency 41.987 ms 20.431 ms 2.06× faster
p90 inference latency 42.612 ms 21.342 ms 2.00× faster
p99 inference latency 43.236 ms 27.181 ms 1.59× faster

🛠️ Technical Specifications

Runtime Architecture

Component role Framework / format
Input square crop and resize ExecuTorch / PTE
ZipDepth Base inference ExecuTorch / PTE
Output resize and zero padding ExecuTorch / PTE

Precision

  • Internal depth-model activations use FP16.
  • Public input and output tensors remain FP32 to preserve the existing brick contract.
  • Dynamic image-shape transforms use FP32.
  • The fixed 24×24-to-12×12 adaptive average pool is represented as an equivalent 2×2, stride-2 average pool supported by the deployed ExecuTorch kernel set.

Input Specification

Input Data type Shape Value range
RGB image FP32 3 × H × W, with 1 ≤ H,W ≤ 4096 [0, 1]

Output Specification

Output Data type Shape Description
Relative depth FP32 H × W Affine-invariant inverse depth; padded regions are zero

Repository Contents

  • zipdepth_base_fp16.pte — static FP16 ZipDepth Base mobile/NPU depth model.
  • zipdepth_prepare_image.pte — dynamic input crop-and-resize transform.
  • zipdepth_restore_depth.pte — dynamic output resize-and-padding transform.
  • zipdepth_manifest.json — model tensor contracts and runtime configuration.
  • samples/im0.jpg — 384×384 upstream sample used for profiling.
  • metadata.yaml and benchmarks/ — model and benchmark metadata.
  • LICENSE — upstream MIT license.
  • SHA256SUMS — reproducibility checksums.

⚠️ Known Limitations

  • The square-crop package pipeline does not preserve the aspect-ratio preprocessing used for ZipDepth's published benchmark results.
  • The output is affine-invariant relative depth, not metric depth.
  • The model has no temporal-consistency mechanism and may flicker when applied independently to video frames.

🗂️ Model and Asset Origin

🔐 Checksums

Verify the checked-out files with:

shasum -a 256 -c SHA256SUMS

Citation

@inproceedings{tosi2026zipdepth,
  title     = {{ZipDepth}: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device},
  author    = {Tosi, Fabio and Bartolomei, Luca and Poggi, Matteo and Mattoccia, Stefano},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026}
}
Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support