Depth Anything V2 Small for LiteRT

LiteRT/TFLite exports of Depth Anything V2 Small, including the complete depth decoder. This model estimates a relative-depth map from a single RGB image. Its outputs are not calibrated distances in meters.

The source has a DINOv2 Small backbone and a DPT decoder, with approximately 24.8 million parameters. These files are based on the official Transformers checkpoint at revision 5426e4f0f36572d16453bbda7a8389317b1bef99.

Model files

File Weights / activations Size
depth_anything_v2_small.tflite FP32 99.53 MB
depth_anything_v2_small_wi8_afp32.tflite Dynamic INT8 weights, FP32 interface 27.73 MB

Sizes use decimal MB. File digests are in SHA256SUMS.

The INT8 variant uses AI Edge Quantizer's dynamic_wi8_afp32() recipe. Supported weight tensors are quantized; the model is not fully integer, and floating-point tensors remain. No calibration images are required. Both files have the same input and output contract.

Signature Name Type Shape / meaning
serving_default input args_0 float32 [1, 3, 518, 686], normalized RGB in NCHW order
serving_default output output_0 float32 [1, 518, 686], relative depth

Batch size and spatial dimensions are fixed. “Dynamic” describes quantization, not dynamic input shapes. Image decoding, resizing, normalization, and visualization are outside the TFLite graph.

Preprocessing

Convert the image to RGB, resize to width 686 × height 518, rescale pixel values by 1/255, and normalize each channel with mean [0.485, 0.456, 0.406] and standard deviation [0.229, 0.224, 0.225]. Use the checkpoint's bicubic image processor with the torchvision backend and keep_aspect_ratio=False, as in the example below. This is the processor backend used to prepare the validation inputs.

This export deliberately stretches the input to its fixed size. The upstream processor normally preserves aspect ratio and can produce a different width or height. preprocessor_config.json contains the upstream processor settings with the fixed size and aspect-ratio policy adjusted for these exports. Using different resizing, padding, or cropping changes the model input and may change predictions.

Run on CPU

Install ai-edge-litert, huggingface_hub, transformers, torch, torchvision, Pillow, and numpy in your Python environment. The example uses the LiteRT CompiledModel API and the Transformers torchvision image processor backend.

import numpy as np
import torch
from PIL import Image
from huggingface_hub import hf_hub_download
from transformers import AutoImageProcessor

from ai_edge_litert.compiled_model import CompiledModel
from ai_edge_litert.hardware_accelerator import HardwareAccelerator
from ai_edge_litert.options import CpuOptions, Options

repo_id = "litert-community/depth-anything-v2-small"
model_path = hf_hub_download(
    repo_id, "tflite/depth_anything_v2_small.tflite"
)
# For dynamic INT8 weights, use:
# "tflite/depth_anything_v2_small_wi8_afp32.tflite"

image = Image.open("image.jpg").convert("RGB")
processor = AutoImageProcessor.from_pretrained(
    "depth-anything/Depth-Anything-V2-Small-hf",
    revision="5426e4f0f36572d16453bbda7a8389317b1bef99",
    backend="torchvision",
)
pixel_values = processor(
    images=image,
    size={"height": 518, "width": 686},
    keep_aspect_ratio=False,
    return_tensors="np",
)["pixel_values"]
pixel_values = np.ascontiguousarray(pixel_values, dtype=np.float32)
assert pixel_values.shape == (1, 3, 518, 686)

model = CompiledModel.from_file(
    model_path,
    options=Options(
        hardware_accelerators=HardwareAccelerator.CPU,
        cpu_options=CpuOptions(num_threads=4),
    ),
)
inputs = model.create_input_buffers(0)
outputs = model.create_output_buffers(0)
inputs[0].write(pixel_values.ravel())
model.run_by_index(0, inputs, outputs)
depth = outputs[0].read(518 * 686, np.float32).reshape(1, 518, 686).copy()

# Optionally resize the relative-depth map to the original image dimensions.
depth_at_image_size = torch.nn.functional.interpolate(
    torch.from_numpy(depth).unsqueeze(1),
    size=(image.height, image.width),
    mode="bicubic",
    align_corners=False,
)[0, 0].numpy()
np.save("relative_depth.npy", depth_at_image_size)

The raw map preserves the model's output scale. Normalize it separately if you want a color visualization; per-image min/max normalization discards that scale.

Runtime considerations

Use GpuOptions(enforce_f32=True) for GPU execution, as in the validated configuration. Previous native WebGPU testing found substantially incorrect depth maps with the runtime's default GPU precision; reduced-precision GPU execution is not covered by this release's validation. Check output parity on your target device:

from ai_edge_litert.options import GpuOptions

gpu_options = Options(
    hardware_accelerators=HardwareAccelerator.GPU,
    gpu_options=GpuOptions(enforce_f32=True),
)
# Pass options=gpu_options to CompiledModel.from_file(...).

Provenance and license

The original model is by Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. See the Depth Anything V2 project and paper. This Small model is released under Apache 2.0; the source project's larger variants have different licenses.

These artifacts provide FP32 and INT8-quantized versions of the official checkpoint, with a fixed input shape and an adapted processor configuration. The WebNN TFLite publication is used as a regression baseline.

@article{depth_anything_v2,
  title={Depth Anything V2},
  author={Yang, Lihe and Kang, Bingyi and Huang, Zilong and Zhao, Zhen and
          Xu, Xiaogang and Feng, Jiashi and Zhao, Hengshuang},
  journal={arXiv:2406.09414},
  year={2024}
}
Downloads last month
89
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for litert-community/depth-anything-v2-small

Quantized
(5)
this model

Collection including litert-community/depth-anything-v2-small

Paper for litert-community/depth-anything-v2-small