AltiMap height model

Single-image height estimation for aerial and satellite RGB imagery: for every pixel, the height above ground in metres (nDSM) plus an 8-class land-cover map. Built for Smart India Hackathon 2026, problem statement 26175 (ISRO, "DepthWizard"), as the model behind AltiMap, which turns one image into metric elevation GeoTIFFs and a 3D city model. Setup and usage: the repo's SETUP.md.

Model

  • Architecture: RS3DAda: DINOv2 ViT-L/14 encoder + DPT decoder with a height-regression head and a segmentation head (SynRS3D, NeurIPS 2024). Needs the model code from JTRNEO/SynRS3D.
  • Starting point: the authors' RS3DAda height weights.
  • Fine-tuning (best.pth, v1): on GAMUS (0.33 m aerial RGB with airborne-LiDAR nDSM and land-cover labels over Washington DC, Philadelphia and New York), BitFit (encoder biases) + decoder and heads, loss L1 + 0.5 x gradient L1 + 0.2 x cross-entropy, 518 px crops, bf16.
  • v2 (v2/best.pth): continues from v1 with the last 8 encoder blocks unfrozen, a height-weighted L1 (a pixel at h metres counts 1 + h/10 times), 30 % SynRS3D samples (high-rises, hills) and 50 % satellite-style degradation (1.3–6Γ— coarser, blur, haze). 6.5 h on an RTX A5000, best checkpoint at step 25,500.

Results (GAMUS test split, 500 tiles, 4-flip test-time averaging)

City RMSE (m) v1 RMSE (m) v2 RMSE (m) zero-shot RS3DAda Pearson r (v1 / v2) Building RMSE (m) (v1 / v2)
Washington DC 3.95 3.86 10.08 0.915 / 0.921 4.51 / 4.45
New York 3.95 4.09 7.17 0.855 / 0.859 3.20 / 2.95
Philadelphia 5.90 5.02 6.28 0.822 / 0.863 10.04 / 8.04
All 5.02 4.55 7.25 0.833 / 0.864 8.28 / 6.74

How AltiMap uses them

v2 wins on GAMUS but generalises worse to other imagery: against USGS airborne LiDAR on NAIP scenes it over-reads trees and ground. On buildings, though, it is right where v1 isn't. So the app combines three models by land cover:

  • v1 gives heights everywhere.
  • Building pixels take the mean of v1 and v2. Building RMSE: city 44.5 β†’ 41.9 m, hills 5.2 β†’ 4.4 m, GAMUS val 2.57 β†’ 2.40 m. The city's tallest objects rise from 66 to 104 m (LiDAR 151 m).
  • Tree pixels in extensive forest (β‰₯ 80 % canopy within 150 m) come from Meta's CHMv2 canopy model.

Limitations (measured, not guessed)

  • Out of domain, by landscape (NAIP aerial RGB vs USGS 3DEP airborne LiDAR, one pass, the app's three-model pipeline; v1 alone in brackets). Absolute DSM = these heights on Copernicus GLO-30, DEM-consistent:

    Scene nDSM RMSE (predict 0) nDSM bias DSM RMSE (GLO-30 alone)
    Dense city (0.3 m) 30.1 m (42.4) [29.8] βˆ’1.2 m [βˆ’6.8] 35.0 m (36.3) [34.7]
    Suburb (0.6 m) 4.5 m (6.5) [4.5] +1.4 m [+1.3] 3.97 m (4.02) [3.99]
    Hilly town (0.6 m) 3.5 m (6.5) [3.7] βˆ’0.4 m [βˆ’0.7] 5.02 m (5.53) [5.04]
    Forest (0.6 m) 15.7 m (24.7) [18.5] βˆ’10.8 m [βˆ’15.9] 8.75 m (8.53) [9.03]

    Very tall towers still read low (104 m vs 151 m), and forest canopy reads ~11 m low. Object heights hold up to ~1–2 m pixels and fade to flat by 5–10 m. ISRO's evaluation imagery is 0.6 m Cartosat-2S over India, outside the training domain (aerial imagery of US cities).

  • Heights are above ground. Absolute elevation needs a terrain model; AltiMap adds Copernicus GLO-30 ground for georeferenced inputs.

Usage

from huggingface_hub import hf_hub_download
ckpt = hf_hub_download("Dilavesh/altimap-height", "best.pth")  # public, no login needed

# with the AltiMap repo checked out and SynRS3D cloned into viewer/cache/SynRS3D:
from pathlib import Path
from viewer.height_model import load_model, predict
model, device = load_model(Path(ckpt), Path("viewer/cache/SynRS3D"))
ndsm_m, oem_classes = predict(model, rgb_uint8_hxwx3, device, tta=True)  # input at ~0.33 m/px

code/ holds the training and data scripts used for these weights.

License and attribution

Weights: MIT, like the RS3DAda base model. Please credit:

  • RS3DAda / SynRS3D: Song et al., SynRS3D: A Synthetic Dataset for Global 3D Semantic Understanding from Monocular Remote Sensing Imagery, NeurIPS 2024 (code and base weights MIT).
  • GAMUS (CC BY 4.0): GAMUS: A Geometry-aware Multi-modal Semantic Segmentation Benchmark for Remote Sensing Data, arXiv:2305.14914.
  • DINOv2 (Apache 2.0), Oquab et al., Meta AI.

The v2 weights were also trained on SynRS3D data, which is CC BY-NC 4.0: treat v2 as non-commercial.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Dilavesh/altimap-height

Finetuned
JTRNEO/RS3DAda
Finetuned
(1)
this model

Datasets used to train Dilavesh/altimap-height

Papers for Dilavesh/altimap-height