LEVIR-CD+ change detection with AT and COAT threshold networks
These are the weights used by the lightning-uq-box tutorial Conformal Change Detection with AT and COAT. The repository contains a binary change-detection model for LEVIR-CD+ and four threshold networks (AT and COAT, each for α = 0.1 and α = 0.2) that predict one threshold per image pair. Calibration is not included: the tutorial calibrates each threshold network on held-out images before evaluation.
Files
| file | content |
|---|---|
base_trainable.safetensors |
fusion layers, U-Net decoder and segmentation head of the base model (7,876,201 parameters) |
threshold/{at,coat}_alpha{0.1,0.2}.ckpt |
Lightning checkpoints of the trained threshold networks, before calibration, without the frozen encoder and without optimizer state |
splits.json |
file names of every split used for training, calibration and evaluation |
logs/base_lr1e-3_metrics.csv |
training log of the base model |
series_summary.csv |
test results over five calibration/test draws |
Base model
The base model is a siamese U-Net: one encoder is applied to each acquisition separately, and the
feature maps of both acquisitions are fused level by level (difference and absolute difference,
followed by a 1×1 convolution) before a U-Net decoder. The encoder is timm's
convnext_base.dinov3_lvd1689m, the ConvNeXt-B distilled from DINOv3, and it was kept frozen
during training. Its weights are not included in this repository. They are downloaded from
their original source when the model is built, and they are subject to the DINOv3 License. The
encoder was pretrained on web images (LVD-1689M), not on satellite imagery. Inputs use ImageNet
normalization.
Training used 437 image pairs of the official LEVIR-CD+ training split, one random 512 × 512 crop per pair with horizontal and vertical flips, binary cross-entropy, Adam (β = (0.5, 0.99)) with batch size 24 for 200 epochs, and a learning rate held for 100 epochs and then decayed linearly to zero. Four learning rates were trained ({1e-4, 3e-4, 1e-3, 3e-3}). The released model (1e-3) had the highest F1 on a separate set of 100 training pairs after the final epoch. This choice was made before any test image was scored. On the full official test split (348 pairs, center 512 × 512 crop, threshold 0.5), the released model has F1 0.846 and IoU 0.733. The four learning rates range from 0.8437 to 0.8461 in test F1.
Threshold networks
Both methods use the library's ThresholdPredictor, a ResNet-50 that receives the six-channel
image pair together with the base model's probability map (threshold_in_channels=6). They were
trained on the 100 training pairs that the base model did not see. AT was trained for 30 epochs
(learning rate 1e-4, batch size 24) and COAT for 60 epochs (learning rate 5e-4, batch size 64,
temperature 0.05). Load a checkpoint with the base model supplied, for example
COAT.load_from_checkpoint(path, model=base_model, pretrained_threshold_net=False, strict=False).
strict=False is required because the frozen encoder weights are omitted from the checkpoint.
Splits
Of the 637 official training pairs, 437 train the base model, 100 train the threshold networks and
100 are unused. The official test split (348 pairs) is divided into 100 calibration and 248 test
pairs, so calibration and test images come from the same population. splits.json lists the draw
used in the tutorial. The four other draws follow from numpy.random.default_rng(1000 + k)
applied to the sorted test file names.
Results over five calibration draws
Mean ± standard deviation over five random calibration/test divisions of the official test split. Coverage is the mean per-image recall of changed pixels, and the coverage gap is the mean absolute distance of each image's recall from 1 − α. The calibration controls the false-negative rate marginally, averaged over test images. It gives no per-image guarantee and does not control false positives. Around 31% of the test crops contain no change and count as fully covered.
| α | method | coverage | coverage gap | F1 | IoU | predicted foreground | threshold clipped to 0 |
|---|---|---|---|---|---|---|---|
| 0.1 | CRC | 0.905 ± 0.016 | 0.142 ± 0.009 | 0.753 ± 0.020 | 0.605 ± 0.026 | 0.061 ± 0.005 | 0 |
| 0.1 | AT | 0.899 ± 0.016 | 0.140 ± 0.010 | 0.237 ± 0.043 | 0.135 ± 0.028 | 0.281 ± 0.054 | 0.244 ± 0.056 |
| 0.1 | COAT | 0.902 ± 0.013 | 0.134 ± 0.009 | 0.558 ± 0.243 | 0.417 ± 0.225 | 0.121 ± 0.083 | 0.066 ± 0.093 |
| 0.2 | CRC | 0.785 ± 0.017 | 0.212 ± 0.002 | 0.847 ± 0.008 | 0.735 ± 0.012 | 0.038 ± 0.002 | 0 |
| 0.2 | AT | 0.786 ± 0.018 | 0.209 ± 0.005 | 0.823 ± 0.017 | 0.699 ± 0.025 | 0.039 ± 0.004 | 0.002 ± 0.002 |
| 0.2 | COAT | 0.788 ± 0.010 | 0.212 ± 0.009 | 0.806 ± 0.021 | 0.676 ± 0.030 | 0.035 ± 0.002 | 0 |
The true foreground fraction of the test crops is about 0.041. In this series the three methods reach similar coverage and similar coverage gaps. At α = 0.1, AT and COAT produce larger masks than CRC. The COAT results at α = 0.1 differ strongly between draws: in two of the five draws, the COAT network (trained with a different seed per draw) set the threshold to zero for 13% and 20% of the test crops, which gives F1 0.34 and 0.25, while the other three draws have F1 0.71 to 0.78. The checkpoints in this repository come from draw 0, which is one of the three draws without clipping.
Licenses and data
- The files in this repository are released under the Apache-2.0 license.
- The encoder weights are not redistributed here. Using them requires accepting the DINOv3 License on the timm model page.
- LEVIR-CD+ is not redistributed here. The tutorial downloads it through torchgeo, and its use is subject to the dataset's own terms.
Hashes
| file | sha256 |
|---|---|
base_trainable.safetensors |
8340f18744529e7d533135a44474a27854e1eee84f364d01b4df0e5a11906b18 |
threshold/at_alpha0.1.ckpt |
15797461860e32a3014c29234d0a319a331a3c452ef9292a90eccc4f48bf6ea5 |
threshold/at_alpha0.2.ckpt |
55817b5720efa18888d9797f380aa08b8962a618749fde8e93dedde50e466ecc |
threshold/coat_alpha0.1.ckpt |
7b5310b0e5c6a4a0880f0ab47ee4bbd53aafac11f9683e7975bd9cfe30fb9549 |
threshold/coat_alpha0.2.ckpt |
78123cd3ee41cbb3f9477800c97813d9d2462f1c314b4e79869162ca3b4468d0 |
The full base checkpoint these weights were taken from has sha256
70436fa6877bb92aa50bb94a15d69c10a68125fb3fb11ee5f9de30440b8f9736.
References
- Luo et al. (2026), Enhancing Image-Conditional Coverage in Segmentation: Adaptive Thresholding via Differentiable Miscoverage Loss, ICLR. https://openreview.net/forum?id=Gd2AiWes1J
- Angelopoulos et al. (2024), Conformal Risk Control, ICLR. https://arxiv.org/abs/2208.02814
- Shen et al. (2021), S2Looking: A Satellite Side-Looking Dataset for Building Change Detection (introduces LEVIR-CD+). https://arxiv.org/abs/2107.09244
- Chen & Shi (2020), A Spatial-Temporal Attention-Based Method and a New Dataset for Remote Sensing Image Change Detection, Remote Sensing 12(10). https://www.mdpi.com/2072-4292/12/10/1662
- Siméoni et al. (2025), DINOv3. https://arxiv.org/abs/2508.10104