Instructions to use BiliSakura/S5-transformers with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BiliSakura/S5-transformers with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="BiliSakura/S5-transformers")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("BiliSakura/S5-transformers", device_map="auto") - Notebooks
- Google Colab
- Kaggle
S5 Transformers Models
Hugging Face–compatible checkpoints converted from the official S5 S4P pretrain weights. Each subfolder is a standalone model repo layout (config.json, model.safetensors, preprocessor, and remote code) for encoder feature extraction on optical remote sensing imagery.
Model Description
These are ViT encoders pretrained with S4P (Semi-supervised Semantic Segmentation Pre-training) on RS4P-1M. This collection currently bundles 2 converted backbone checkpoints:
- ViT-B: ViT-Base/16, hidden size 768, 12 layers
- ViT-L: ViT-Large/16, hidden size 1024, 24 layers
Both checkpoints use architecture: s4p_backbone and expose the s5-feature-extraction pipeline. They are encoder-only weights (not UPerNet segmentation or MoE-MDF heads).
All folders ship self-contained remote code (modeling_s5.py, processor, pipeline) and load with trust_remote_code=True.
Developed by: lianglyu / S5
Converted for Hugging Face by: BiliSakura
License (weights): Apache 2.0
Original paper: S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing (AAAI 2026 Oral)
Available checkpoints
| Folder | Backbone | Hidden size | Layers | Heads | Patch | Image size | Original file |
|---|---|---|---|---|---|---|---|
ViT-B |
ViT-Base | 768 | 12 | 12 | 16 | 512 | vit_b_s4p.pth |
ViT-L |
ViT-Large | 1024 | 24 | 16 | 16 | 512 | vit_l_s4p.pth |
Original singular .pth files were converted and removed from this directory.
Usage
Processors default to do_resize: false. Pass RGB images at native resolution; ImageNet mean/std normalization is applied when enabled.
from transformers import pipeline
import numpy as np
REPO = "/path/to/S5-transformers"
pipe = pipeline(
task="s5-feature-extraction",
model=f"{REPO}/ViT-B",
trust_remote_code=True,
)
image = np.random.randint(0, 255, (512, 512, 3), dtype=np.uint8)
# Global pooled features
features = pipe(image, pool=True, return_tensors=True)
print(features.shape) # [1, 768] for ViT-B, [1, 1024] for ViT-L
# Dense feature map
featmap = pipe(image, pool=False, return_tensors=True)
print(featmap.shape) # [1, 768, 32, 32] for ViT-B, [1, 1024, 32, 32] for ViT-L
ViT-L:
pipe = pipeline(
task="s5-feature-extraction",
model=f"{REPO}/ViT-L",
trust_remote_code=True,
)
features = pipe(image, pool=True, return_tensors=True)
print(features.shape) # [1, 1024]
To force 512×512 resize:
features = pipe(
image,
pool=True,
return_tensors=True,
image_processor_kwargs={"do_resize": True},
)
Load components directly:
from transformers import AutoModel, AutoImageProcessor
model = AutoModel.from_pretrained(f"{REPO}/ViT-B", trust_remote_code=True)
processor = AutoImageProcessor.from_pretrained(f"{REPO}/ViT-B", trust_remote_code=True)
Normalization
The bundled image processor applies ImageNet mean/std normalization by default (do_normalize=True, rescale_factor=1/255). Inputs should be RGB optical imagery.
Conversion
Checkpoints were converted with scripts/convert_s5_checkpoint.py from the official release:
python scripts/convert_s5_checkpoint.py \
--input-path /path/to/vit_b_s4p.pth \
--output-dir /path/to/ViT-B \
--clean-output
For semantic segmentation heads, convert vit_*_s4p_upernet.pth or s5_vit_*_moe_mdf_seg.pth instead (task: s5-semantic-segmentation).
Dependencies
transformers>=4.45.0torch>=2.1.0safetensorsPillownumpy
Citation
@article{S5,
title={S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing},
author={Liang Lv and Di Wang and Jing Zhang and Lefei Zhang},
journal={arXiv preprint arXiv:2508.12409},
year={2025}
}