SkyVLM

A 99.4M-parameter vision-language model for satellite imagery, trained from scratch on a single 8 GB laptop GPU. It writes a structured caption such as a satellite image of <main object>, surrounded by <object>; <object>.

Code, training scripts, a Streamlit demo and full results: https://github.com/thorOdinson16/SkyVLM

Model

image ─► ViT (25.7M; masked-image-modeling pretrained, then CLIP-aligned)
          └─ 196 patch tokens ─► 2-layer MLP projector (1.05M) ─► LM (72.6M decoder) ─► caption
Component Parameters
Vision encoder (ViT, 8 layers, 512 dim, 16×16 patches, 224 px) 25.7M
MLP projector (512→1024→512) 1.05M
Language model (20 layers, RoPE, RMSNorm, 16k vocabulary) 72.6M
Total 99.4M

The language model was pretrained on ~1B tokens (90% FineWeb-Edu, 10% SkyScript captions), the encoder on 200k unlabelled SkyScript images and then aligned to captions with a CLIP-style loss, and the whole model was fine-tuned on 200k SkyScript image-caption pairs (projector first, then everything).

Usage

This is a custom PyTorch architecture, not a transformers model. Download the repo and run:

pip install torch torchvision safetensors sentencepiece pandas pillow
python inference.py my_image1.jpg my_image2.png

or from Python:

from PIL import Image
from inference import load, caption

model, tok, cfg = load(".")                      # repo directory
print(caption(model, tok, cfg, [Image.open("my_image.jpg")])[0])

Images are resized to 224×224 and decoded greedily. The output follows the SkyScript caption template (OpenStreetMap-style tags).

Results

Official SkyScript test captions, all 30k images, greedy decoding:

Model CIDEr-D BLEU-4 ROUGE-L Main object exact Surrounding-object F1
SkyVLM (this model) 1.90 0.475 0.643 21.3% 0.386
Same recipe without CLIP alignment 1.75 0.464 0.632 19.6% 0.372
Same recipe, random-init ViT (first 5k test images) 0.69 0.323 0.530 5.9% 0.264

Shuffled or blank images drop CIDEr-D to 0.09 / 0.04, so the model uses the image.

Image-text retrieval on 1,000 test pairs, ranking captions by log p(caption | image) (image→text uses a PMI correction):

Model Image→text R@1 Text→image R@1
SkyVLM 36.5 34.7
SkyCLIP ViT-B/32 (external reference) 8.6 6.8

SkyCLIP is a general remote-sensing model while SkyVLM is a specialist trained on these templated captions, so this is a reference point rather than a claim of a better model.

Vision encoder

RESISC45 linear probe (45 scene classes, 6,300 test images, frozen features; no RESISC45 data seen in training):

Encoder Test accuracy Macro-F1
Random ViT (untrained) 43.5% 0.425
MIM-pretrained (100 epochs, 200k unlabelled images) 78.6% 0.785
MIM + CLIP alignment (this model's encoder) 88.4% 0.885
ImageNet ViT-S/16 (supervised, ~14M labelled images) 88.5% 0.885

Language model

FineWeb-Edu validation perplexity 21.7 (72.6M parameters, 983M training tokens). SkyScript caption validation loss bottomed at 0.76 and ended at 0.94 as the model began to memorize the captions (10% of each batch, ~18 passes).

Sanity checks

Main-object category (building, road, railway, ...) is right 61.5% of the time versus 21% for always guessing the most common one (measured on the variant without CLIP alignment); all outputs follow the caption template.

Single training run per condition; no variance estimates.

Limitations

  • Trained only on SkyScript's templated, OpenStreetMap-derived captions: it will not follow free-form prompts or answer questions.
  • Typical errors are the right category with the wrong specifics (an apartment block described as an office). Exact main-object match is 21%, and references have several valid descriptions per image, so n-gram and exact-match scores understate quality.
  • Images are ~200 px at 224×224 input; fine detail is limited.
  • Not evaluated outside SkyScript imagery, on other regions' sensors, or for any safety-critical use.

Data and license

Apache-2.0. Trained on SkyScript (MIT), which pairs Google Earth Engine imagery with OpenStreetMap-derived captions. OpenStreetMap data is © OpenStreetMap contributors (ODbL); imagery remains subject to its providers' terms. FineWeb-Edu (ODC-By) was used for language-model pretraining.

Downloads last month
34
Safetensors
Model size
99.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train AbhiDS16/SkyVLM