Depth Anything V2 Base — Minecraft metric depth

Fine-tune of depth-anything/Depth-Anything-V2-Base-hf on JulianBvW/Minecraft-Depth-Images. Input: Minecraft RGB screenshot. Output: metric depth in Minecraft blocks, range (0, 95).

Model

Architecture DepthAnythingForDepthEstimation (DINOv2 ViT-B/14 backbone + DPT neck + depth head)
Parameters 97.5M (all trainable during fine-tuning)
Head depth_estimation_type="metric": sigmoid(x) * max_depth, max_depth=95
Output predicted_depth (B, H, W), float, blocks, same H×W as the input
Input resolution 476×854 (H×W; native 480×854 frames, sides multiple of 14)
Normalisation ImageNet mean/std (unchanged from the base processor)

The released base checkpoint predicts affine-invariant relative inverse depth. Switching the head to the metric variant changes only the final activation (no new parameters); the last 1×1 conv was re-initialised by a 2-parameter least-squares fit (logit(depth/max_depth) ≈ a·x + b, a=-0.1586, b=-2.0973) on training batches before fine-tuning.

Usage

import torch
from PIL import Image
from transformers import AutoImageProcessor, AutoModelForDepthEstimation

repo = "thealper2/depth-anything-v2-base-minecraft"
processor = AutoImageProcessor.from_pretrained(repo)
model = AutoModelForDepthEstimation.from_pretrained(repo).eval()

image = Image.open("minecraft_screenshot.png").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
    depth = model(**inputs).predicted_depth                       # (1, H', W') blocks
depth = torch.nn.functional.interpolate(depth[:, None], size=image.size[::-1],
                                        mode="bilinear", align_corners=False)[0, 0]

The saved processor resizes with keep_aspect_ratio=True, ensure_multiple_of=14, target 476×854 (854×480 input → 854×476). Depth scale assumes the default Minecraft FOV (70°) used in the dataset.

Training data

Source depth_dataset.tar.gz (rev d31495a5cc1aa5a169e732939ba13b9d98bf049b): 12,000 RGB (854×480 PNG) + depth (float16 .npy) pairs
Depth labels metric depth in blocks, decoded from 8-bit near/far depth shaders; quantisation step ≈0.024 (<6 blocks) to ≈0.56 (>80 blocks)
Far clip label saturates at 90.5625 (sky / farther than ~90 blocks)
Excluded 822 frames whose depth label is shifted by one frame relative to the RGB (detected via RGB-edge / depth-edge agreement), 169 train frames adjacent to val/test chunks
Split recording runs (3000 frames) cut into 200-frame chunks, chunks assigned at random (seed 42): train 8771 / val 1039 / test 1199 frames

Training procedure

Loss masked Smooth L1 on log-depth (β=0.1) + 0.1 × multi-scale (4) gradient-matching loss on the log-depth residual
Masking gt ≤ 0 ignored; far-clip pixels supervised as 90.5625
Augmentation shared random crop offset + horizontal flip (RGB and depth identical); colour jitter, Gaussian blur, Gaussian noise (RGB only)
Optimizer AdamW, lr 1e-05 (backbone) / 1e-04 (neck + head), weight decay 1e-04, betas (0.9, 0.999)
Schedule linear warm-up 3% of steps, cosine decay to 0.01×
Batch 4 × 4 gradient accumulation = 16
Epochs 10 (5480 optimizer steps); best checkpoint by val abs_rel (epoch 10)
Precision bf16 autocast, gradient clipping 1.0
Hardware NVIDIA GeForce RTX 5060 Ti; training time 2.07 h; peak memory 6.6 GB
Software PyTorch 2.11.0+cu128, Transformers 5.17.0

Evaluation

Test split: 1199 frames never used for training or model selection. Metrics are computed per image on pixels with 0.1 ≤ gt < 90.5625 blocks and averaged over images. FarRecall = fraction of far-clip (sky / >90 blocks) pixels predicted ≥ 90.5625/1.25.

Model MAE ↓ RMSE ↓ AbsRel ↓ SqRel ↓ RMSElog ↓ δ<1.25 ↑ δ<1.25² ↑ δ<1.25³ ↑ FarRecall ↑
DA-V2 Base (pretrained, aligned) 4.840 11.656 0.4822 24.032 0.4158 0.6707 0.8347 0.8976 0.7845
DA-V2 Base Minecraft (metric) 0.627 2.193 0.0630 0.519 0.1276 0.9576 0.9849 0.9923 0.9122
DA-V2 Base Minecraft (aligned) 0.868 2.740 0.0710 0.645 0.1419 0.9397 0.9767 0.9876 0.5784

aligned: per-image least-squares scale and shift in disparity space fitted to the ground truth (MiDaS protocol). Required for the base model, whose output has no metric scale; it uses the test labels and therefore favours the baseline. metric: raw model output in blocks, no alignment.

Limitations

  • Training data: overworld flight recordings only (no Nether / End, no entities / mobs), default resource pack, default FOV.
  • Water and glass are transparent in the depth labels; the model predicts the depth behind them.
  • Labels beyond ~90 blocks are clipped; predictions above that range are not meaningful.
  • Label quantisation (8-bit shaders) limits achievable accuracy, especially at large depth.
  • License follows the base model (CC-BY-NC-4.0).
Downloads last month
16
Safetensors
Model size
97.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thealper2/depth-anything-v2-base-minecraft

Finetuned
(2)
this model

Dataset used to train thealper2/depth-anything-v2-base-minecraft