Instructions to use thealper2/depth-anything-v2-base-minecraft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thealper2/depth-anything-v2-base-minecraft with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("depth-estimation", model="thealper2/depth-anything-v2-base-minecraft")# Load model directly from transformers import AutoImageProcessor, AutoModelForDepthEstimation processor = AutoImageProcessor.from_pretrained("thealper2/depth-anything-v2-base-minecraft") model = AutoModelForDepthEstimation.from_pretrained("thealper2/depth-anything-v2-base-minecraft", device_map="auto") - DepthAnythingV2
How to use thealper2/depth-anything-v2-base-minecraft with DepthAnythingV2:
# Install from https://github.com/DepthAnything/Depth-Anything-V2 # Load the model and infer depth from an image import cv2 import torch from depth_anything_v2.dpt import DepthAnythingV2 # instantiate the model model = DepthAnythingV2(encoder="<ENCODER>", features=<NUMBER_OF_FEATURES>, out_channels=<OUT_CHANNELS>) # load the weights filepath = hf_hub_download(repo_id="thealper2/depth-anything-v2-base-minecraft", filename="depth_anything_v2_<ENCODER>.pth", repo_type="model") state_dict = torch.load(filepath, map_location="cpu") model.load_state_dict(state_dict).eval() raw_img = cv2.imread("your/image/path") depth = model.infer_image(raw_img) # HxW raw depth map in numpy - Notebooks
- Google Colab
- Kaggle
Depth Anything V2 Base — Minecraft metric depth
Fine-tune of depth-anything/Depth-Anything-V2-Base-hf
on JulianBvW/Minecraft-Depth-Images.
Input: Minecraft RGB screenshot. Output: metric depth in Minecraft blocks, range (0, 95).
Model
| Architecture | DepthAnythingForDepthEstimation (DINOv2 ViT-B/14 backbone + DPT neck + depth head) |
| Parameters | 97.5M (all trainable during fine-tuning) |
| Head | depth_estimation_type="metric": sigmoid(x) * max_depth, max_depth=95 |
| Output | predicted_depth (B, H, W), float, blocks, same H×W as the input |
| Input resolution | 476×854 (H×W; native 480×854 frames, sides multiple of 14) |
| Normalisation | ImageNet mean/std (unchanged from the base processor) |
The released base checkpoint predicts affine-invariant relative inverse depth. Switching the head to the metric
variant changes only the final activation (no new parameters); the last 1×1 conv was re-initialised by a
2-parameter least-squares fit (logit(depth/max_depth) ≈ a·x + b, a=-0.1586, b=-2.0973) on training
batches before fine-tuning.
Usage
import torch
from PIL import Image
from transformers import AutoImageProcessor, AutoModelForDepthEstimation
repo = "thealper2/depth-anything-v2-base-minecraft"
processor = AutoImageProcessor.from_pretrained(repo)
model = AutoModelForDepthEstimation.from_pretrained(repo).eval()
image = Image.open("minecraft_screenshot.png").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
depth = model(**inputs).predicted_depth # (1, H', W') blocks
depth = torch.nn.functional.interpolate(depth[:, None], size=image.size[::-1],
mode="bilinear", align_corners=False)[0, 0]
The saved processor resizes with keep_aspect_ratio=True, ensure_multiple_of=14, target 476×854
(854×480 input → 854×476). Depth scale assumes the default Minecraft FOV (70°) used in the dataset.
Training data
| Source | depth_dataset.tar.gz (rev d31495a5cc1aa5a169e732939ba13b9d98bf049b): 12,000 RGB (854×480 PNG) + depth (float16 .npy) pairs |
| Depth labels | metric depth in blocks, decoded from 8-bit near/far depth shaders; quantisation step ≈0.024 (<6 blocks) to ≈0.56 (>80 blocks) |
| Far clip | label saturates at 90.5625 (sky / farther than ~90 blocks) |
| Excluded | 822 frames whose depth label is shifted by one frame relative to the RGB (detected via RGB-edge / depth-edge agreement), 169 train frames adjacent to val/test chunks |
| Split | recording runs (3000 frames) cut into 200-frame chunks, chunks assigned at random (seed 42): train 8771 / val 1039 / test 1199 frames |
Training procedure
| Loss | masked Smooth L1 on log-depth (β=0.1) + 0.1 × multi-scale (4) gradient-matching loss on the log-depth residual |
| Masking | gt ≤ 0 ignored; far-clip pixels supervised as 90.5625 |
| Augmentation | shared random crop offset + horizontal flip (RGB and depth identical); colour jitter, Gaussian blur, Gaussian noise (RGB only) |
| Optimizer | AdamW, lr 1e-05 (backbone) / 1e-04 (neck + head), weight decay 1e-04, betas (0.9, 0.999) |
| Schedule | linear warm-up 3% of steps, cosine decay to 0.01× |
| Batch | 4 × 4 gradient accumulation = 16 |
| Epochs | 10 (5480 optimizer steps); best checkpoint by val abs_rel (epoch 10) |
| Precision | bf16 autocast, gradient clipping 1.0 |
| Hardware | NVIDIA GeForce RTX 5060 Ti; training time 2.07 h; peak memory 6.6 GB |
| Software | PyTorch 2.11.0+cu128, Transformers 5.17.0 |
Evaluation
Test split: 1199 frames never used for training or model selection. Metrics are computed per image on pixels with
0.1 ≤ gt < 90.5625 blocks and averaged over images. FarRecall = fraction of far-clip (sky / >90 blocks)
pixels predicted ≥ 90.5625/1.25.
| Model | MAE ↓ | RMSE ↓ | AbsRel ↓ | SqRel ↓ | RMSElog ↓ | δ<1.25 ↑ | δ<1.25² ↑ | δ<1.25³ ↑ | FarRecall ↑ |
|---|---|---|---|---|---|---|---|---|---|
| DA-V2 Base (pretrained, aligned) | 4.840 | 11.656 | 0.4822 | 24.032 | 0.4158 | 0.6707 | 0.8347 | 0.8976 | 0.7845 |
| DA-V2 Base Minecraft (metric) | 0.627 | 2.193 | 0.0630 | 0.519 | 0.1276 | 0.9576 | 0.9849 | 0.9923 | 0.9122 |
| DA-V2 Base Minecraft (aligned) | 0.868 | 2.740 | 0.0710 | 0.645 | 0.1419 | 0.9397 | 0.9767 | 0.9876 | 0.5784 |
aligned: per-image least-squares scale and shift in disparity space fitted to the ground truth (MiDaS protocol).
Required for the base model, whose output has no metric scale; it uses the test labels and therefore favours the
baseline. metric: raw model output in blocks, no alignment.
Limitations
- Training data: overworld flight recordings only (no Nether / End, no entities / mobs), default resource pack, default FOV.
- Water and glass are transparent in the depth labels; the model predicts the depth behind them.
- Labels beyond ~90 blocks are clipped; predictions above that range are not meaningful.
- Label quantisation (8-bit shaders) limits achievable accuracy, especially at large depth.
- License follows the base model (CC-BY-NC-4.0).
- Downloads last month
- 16
Model tree for thealper2/depth-anything-v2-base-minecraft
Base model
depth-anything/Depth-Anything-V2-Base-hf