Scale-MAE

Model Introduction

Scale-MAE is a scale-aware masked autoencoder for multiscale geospatial imagery that learns stable remote sensing image representations through ground-sampling-distance-aware positional encoding, visible-patch encoding, and low- and high-frequency target reconstruction.

Paper: Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning
https://arxiv.org/abs/2212.14532

Model Description

Scale-MAE was proposed by research teams at NASA's Jet Propulsion Laboratory and Stanford University. The model is trained using multiscale geospatial imagery such as FMoW-RGB. The model is suitable for tasks such as remote sensing image representation learning, scene classification, and building segmentation.

Applicable Scenarios

Scenario Description
Multiscale remote sensing pre-training Use paired low-resolution and high-resolution BCHW imagery with GSD metadata.
Scene classification Perform kNN transfer evaluation through reusable CLS features.
Building segmentation Transfer scale-aware representations to building semantic-segmentation tasks such as SpaceNet and fine-tune them.
Low- and high-frequency reconstruction Use area resampling and band-pass targets to evaluate scale sensitivity.
Local quick validation Use synthetic data to check data loading, training, inference, and evaluation.
Multi-GPU training Launch distributed data-parallel training through torchrun.

Usage Instructions

1. OneCode Usage

Experience intelligent one-click AI4S programming through the OneCode online environment:

Experience intelligent one-click AI4S programming

2. Manual Installation and Usage

Hardware Requirements

  • GPU or DCU execution is recommended.
  • CPU can be used to validate the workflow with the current default small configuration.
  • DCU users need to install a DTK version matching the cluster in advance. DTK 25.04.2 or later is recommended.

Download the Model Package

hf download OneScience-Group/Scale-MAE --local-dir ./Scale-MAE
cd Scale-MAE

Install the Runtime Environment

DCU Environment

conda create -n onescience311 python=3.11 -y
conda activate onescience311
pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai

GPU Environment

conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai

Training Data Introduction

The paper uses multiscale geospatial imagery such as FMoW-RGB for pre-training and evaluates on tasks including RESISC-45, UCMerced, EuroSAT, AID, MLRSNet, and SpaceNet. Data files contain images [B,C,H,W], targets [B,C,Ht,Wt], gsd [B], and labels [B].

Synthetic data is used by default:

python scripts/fake_data.py

When using real data, do not run fake_data.py. First organize the data into the following directories and fields, and replace the files under data/ generated by the synthetic-data script:

data/train.npz
data/test.npz

Each NPZ file contains at least:

images:  float32 [N,C,input_size,input_size]
targets: float32 [N,C,target_size,target_size]
gsd:     float32 [N]
labels:  int64   [N]

Here, images is the model input, targets is the target-resolution imagery corresponding to the input scene, gsd is the ground sampling distance for each sample in meters per pixel, and labels is used for kNN feature evaluation. The channel count, input size, target size, and GSD range of real data must be consistent with conf/config.yaml and the model configuration; cropping, registration, channel organization, and numerical normalization should be completed before generating the NPZ files.

Modify input_size, target_size, channels, gsd_values, and paths in conf/config.yaml according to the real data. After completing data preparation, continue to use the unified training, inference, and evaluation commands below; if other file locations are needed, override the default paths through script arguments.

Training

Single GPU:

python scripts/train.py

Multiple GPUs:

torchrun --nproc_per_node=8 scripts/train.py

Training outputs:

result/checkpoints/scalemae.pt
result/training/metrics.json

Training outputs include a model checkpoint that can be used for subsequent inference and feature extraction, as well as training metrics reflecting changes in overall, low-frequency, and high-frequency reconstruction losses, facilitating training-state preservation and analysis of model optimization.

AdamW uses betas (0.9, 0.95) and includes gradient accumulation, AMP, warmup, and cosine decay.

Trained Weights

This repository provides weights trained on multiscale geospatial imagery in the weight/ folder. The weight files will be uploaded soon.

Inference

python scripts/inference.py

Inference results are output to:

result/output/reconstruction.npz

Evaluation and Visualization

python scripts/result.py

Evaluation and visualization outputs are saved to:

result/evaluation/metrics.json
result/evaluation/features.npy
result/evaluation/bandpass_reconstruction.png
result/evaluation/frequency_error.png
result/evaluation/gsd_reconstruction_error.png
result/evaluation/gsd_knn_accuracy.png

Evaluation results comprehensively reflect overall and low- and high-frequency reconstruction quality, scale adaptability under different GSD values, and the kNN classification capability of representation features, while reconstruction comparisons and scale-variation curves demonstrate the model's ability to process multiscale geospatial imagery. The current results are based on a small amount of synthetic data and are mainly used to confirm that the training, inference, evaluation, and visualization workflows operate normally.

OneScience Official Information

Citation and License

This repository is a reproduction of the original Scale-MAE paper.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for OneScience-Group/Scale-MAE