AltiMap height model
Single-image height estimation for aerial and satellite RGB imagery: for every pixel, the height above ground in metres (nDSM) plus an 8-class land-cover map. Built for Smart India Hackathon 2026, problem statement 26175 (ISRO, "DepthWizard"), as the model behind AltiMap, which turns one image into metric elevation GeoTIFFs and a 3D city model. Setup and usage: the repo's SETUP.md.
Model
- Architecture: RS3DAda: DINOv2 ViT-L/14 encoder + DPT decoder with a height-regression head and a segmentation head (SynRS3D, NeurIPS 2024). Needs the model code from JTRNEO/SynRS3D.
- Starting point: the authors' RS3DAda height weights.
- Fine-tuning (
best.pth, v1): on GAMUS (0.33 m aerial RGB with airborne-LiDAR nDSM and land-cover labels over Washington DC, Philadelphia and New York), BitFit (encoder biases) + decoder and heads, loss L1 + 0.5 x gradient L1 + 0.2 x cross-entropy, 518 px crops, bf16. - v2 (
v2/best.pth): continues from v1 with the last 8 encoder blocks unfrozen, a height-weighted L1 (a pixel at h metres counts 1 + h/10 times), 30 % SynRS3D samples (high-rises, hills) and 50 % satellite-style degradation (1.3β6Γ coarser, blur, haze). 6.5 h on an RTX A5000, best checkpoint at step 25,500.
Results (GAMUS test split, 500 tiles, 4-flip test-time averaging)
| City | RMSE (m) v1 | RMSE (m) v2 | RMSE (m) zero-shot RS3DAda | Pearson r (v1 / v2) | Building RMSE (m) (v1 / v2) |
|---|---|---|---|---|---|
| Washington DC | 3.95 | 3.86 | 10.08 | 0.915 / 0.921 | 4.51 / 4.45 |
| New York | 3.95 | 4.09 | 7.17 | 0.855 / 0.859 | 3.20 / 2.95 |
| Philadelphia | 5.90 | 5.02 | 6.28 | 0.822 / 0.863 | 10.04 / 8.04 |
| All | 5.02 | 4.55 | 7.25 | 0.833 / 0.864 | 8.28 / 6.74 |
How AltiMap uses them
v2 wins on GAMUS but generalises worse to other imagery: against USGS airborne LiDAR on NAIP scenes it over-reads trees and ground. On buildings, though, it is right where v1 isn't. So the app combines three models by land cover:
- v1 gives heights everywhere.
- Building pixels take the mean of v1 and v2. Building RMSE: city 44.5 β 41.9 m, hills 5.2 β 4.4 m, GAMUS val 2.57 β 2.40 m. The city's tallest objects rise from 66 to 104 m (LiDAR 151 m).
- Tree pixels in extensive forest (β₯ 80 % canopy within 150 m) come from Meta's CHMv2 canopy model.
Limitations (measured, not guessed)
Out of domain, by landscape (NAIP aerial RGB vs USGS 3DEP airborne LiDAR, one pass, the app's three-model pipeline; v1 alone in brackets). Absolute DSM = these heights on Copernicus GLO-30, DEM-consistent:
Scene nDSM RMSE (predict 0) nDSM bias DSM RMSE (GLO-30 alone) Dense city (0.3 m) 30.1 m (42.4) [29.8] β1.2 m [β6.8] 35.0 m (36.3) [34.7] Suburb (0.6 m) 4.5 m (6.5) [4.5] +1.4 m [+1.3] 3.97 m (4.02) [3.99] Hilly town (0.6 m) 3.5 m (6.5) [3.7] β0.4 m [β0.7] 5.02 m (5.53) [5.04] Forest (0.6 m) 15.7 m (24.7) [18.5] β10.8 m [β15.9] 8.75 m (8.53) [9.03] Very tall towers still read low (104 m vs 151 m), and forest canopy reads ~11 m low. Object heights hold up to ~1β2 m pixels and fade to flat by 5β10 m. ISRO's evaluation imagery is 0.6 m Cartosat-2S over India, outside the training domain (aerial imagery of US cities).
Heights are above ground. Absolute elevation needs a terrain model; AltiMap adds Copernicus GLO-30 ground for georeferenced inputs.
Usage
from huggingface_hub import hf_hub_download
ckpt = hf_hub_download("Dilavesh/altimap-height", "best.pth") # public, no login needed
# with the AltiMap repo checked out and SynRS3D cloned into viewer/cache/SynRS3D:
from pathlib import Path
from viewer.height_model import load_model, predict
model, device = load_model(Path(ckpt), Path("viewer/cache/SynRS3D"))
ndsm_m, oem_classes = predict(model, rgb_uint8_hxwx3, device, tta=True) # input at ~0.33 m/px
code/ holds the training and data scripts used for these weights.
License and attribution
Weights: MIT, like the RS3DAda base model. Please credit:
- RS3DAda / SynRS3D: Song et al., SynRS3D: A Synthetic Dataset for Global 3D Semantic Understanding from Monocular Remote Sensing Imagery, NeurIPS 2024 (code and base weights MIT).
- GAMUS (CC BY 4.0): GAMUS: A Geometry-aware Multi-modal Semantic Segmentation Benchmark for Remote Sensing Data, arXiv:2305.14914.
- DINOv2 (Apache 2.0), Oquab et al., Meta AI.
The v2 weights were also trained on SynRS3D data, which is CC BY-NC 4.0: treat v2 as non-commercial.