S5 Transformers Models

Hugging Face–compatible checkpoints converted from the official S5 S4P pretrain weights. Each subfolder is a standalone model repo layout (config.json, model.safetensors, preprocessor, and remote code) for encoder feature extraction on optical remote sensing imagery.

Model Description

These are ViT encoders pretrained with S4P (Semi-supervised Semantic Segmentation Pre-training) on RS4P-1M. This collection currently bundles 2 converted backbone checkpoints:

  • ViT-B: ViT-Base/16, hidden size 768, 12 layers
  • ViT-L: ViT-Large/16, hidden size 1024, 24 layers

Both checkpoints use architecture: s4p_backbone and expose the s5-feature-extraction pipeline. They are encoder-only weights (not UPerNet segmentation or MoE-MDF heads).

All folders ship self-contained remote code (modeling_s5.py, processor, pipeline) and load with trust_remote_code=True.

Developed by: lianglyu / S5
Converted for Hugging Face by: BiliSakura
License (weights): Apache 2.0
Original paper: S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing (AAAI 2026 Oral)

Available checkpoints

Folder Backbone Hidden size Layers Heads Patch Image size Original file
ViT-B ViT-Base 768 12 12 16 512 vit_b_s4p.pth
ViT-L ViT-Large 1024 24 16 16 512 vit_l_s4p.pth

Original singular .pth files were converted and removed from this directory.

Usage

Processors default to do_resize: false. Pass RGB images at native resolution; ImageNet mean/std normalization is applied when enabled.

from transformers import pipeline
import numpy as np

REPO = "/path/to/S5-transformers"

pipe = pipeline(
    task="s5-feature-extraction",
    model=f"{REPO}/ViT-B",
    trust_remote_code=True,
)

image = np.random.randint(0, 255, (512, 512, 3), dtype=np.uint8)

# Global pooled features
features = pipe(image, pool=True, return_tensors=True)
print(features.shape)  # [1, 768] for ViT-B, [1, 1024] for ViT-L

# Dense feature map
featmap = pipe(image, pool=False, return_tensors=True)
print(featmap.shape)   # [1, 768, 32, 32] for ViT-B, [1, 1024, 32, 32] for ViT-L

ViT-L:

pipe = pipeline(
    task="s5-feature-extraction",
    model=f"{REPO}/ViT-L",
    trust_remote_code=True,
)
features = pipe(image, pool=True, return_tensors=True)
print(features.shape)  # [1, 1024]

To force 512×512 resize:

features = pipe(
    image,
    pool=True,
    return_tensors=True,
    image_processor_kwargs={"do_resize": True},
)

Load components directly:

from transformers import AutoModel, AutoImageProcessor

model = AutoModel.from_pretrained(f"{REPO}/ViT-B", trust_remote_code=True)
processor = AutoImageProcessor.from_pretrained(f"{REPO}/ViT-B", trust_remote_code=True)

Normalization

The bundled image processor applies ImageNet mean/std normalization by default (do_normalize=True, rescale_factor=1/255). Inputs should be RGB optical imagery.

Conversion

Checkpoints were converted with scripts/convert_s5_checkpoint.py from the official release:

python scripts/convert_s5_checkpoint.py \
  --input-path /path/to/vit_b_s4p.pth \
  --output-dir /path/to/ViT-B \
  --clean-output

For semantic segmentation heads, convert vit_*_s4p_upernet.pth or s5_vit_*_moe_mdf_seg.pth instead (task: s5-semantic-segmentation).

Dependencies

  • transformers>=4.45.0
  • torch>=2.1.0
  • safetensors
  • Pillow
  • numpy

Citation

@article{S5,
  title={S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing},
  author={Liang Lv and Di Wang and Jing Zhang and Lefei Zhang},
  journal={arXiv preprint arXiv:2508.12409},
  year={2025}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including BiliSakura/S5-transformers

Paper for BiliSakura/S5-transformers