SkyVLM
A 99.4M-parameter vision-language model for satellite imagery, trained from scratch on a single 8 GB laptop GPU.
It writes a structured caption such as
a satellite image of <main object>, surrounded by <object>; <object>.
Code, training scripts, a Streamlit demo and full results: https://github.com/thorOdinson16/SkyVLM
Model
image ─► ViT (25.7M; masked-image-modeling pretrained, then CLIP-aligned)
└─ 196 patch tokens ─► 2-layer MLP projector (1.05M) ─► LM (72.6M decoder) ─► caption
| Component | Parameters |
|---|---|
| Vision encoder (ViT, 8 layers, 512 dim, 16×16 patches, 224 px) | 25.7M |
| MLP projector (512→1024→512) | 1.05M |
| Language model (20 layers, RoPE, RMSNorm, 16k vocabulary) | 72.6M |
| Total | 99.4M |
The language model was pretrained on ~1B tokens (90% FineWeb-Edu, 10% SkyScript captions), the encoder on 200k unlabelled SkyScript images and then aligned to captions with a CLIP-style loss, and the whole model was fine-tuned on 200k SkyScript image-caption pairs (projector first, then everything).
Usage
This is a custom PyTorch architecture, not a transformers model. Download the repo and run:
pip install torch torchvision safetensors sentencepiece pandas pillow
python inference.py my_image1.jpg my_image2.png
or from Python:
from PIL import Image
from inference import load, caption
model, tok, cfg = load(".") # repo directory
print(caption(model, tok, cfg, [Image.open("my_image.jpg")])[0])
Images are resized to 224×224 and decoded greedily. The output follows the SkyScript caption template (OpenStreetMap-style tags).
Results
Official SkyScript test captions, all 30k images, greedy decoding:
| Model | CIDEr-D | BLEU-4 | ROUGE-L | Main object exact | Surrounding-object F1 |
|---|---|---|---|---|---|
| SkyVLM (this model) | 1.90 | 0.475 | 0.643 | 21.3% | 0.386 |
| Same recipe without CLIP alignment | 1.75 | 0.464 | 0.632 | 19.6% | 0.372 |
| Same recipe, random-init ViT (first 5k test images) | 0.69 | 0.323 | 0.530 | 5.9% | 0.264 |
Shuffled or blank images drop CIDEr-D to 0.09 / 0.04, so the model uses the image.
Image-text retrieval on 1,000 test pairs, ranking captions by log p(caption | image) (image→text uses a PMI correction):
| Model | Image→text R@1 | Text→image R@1 |
|---|---|---|
| SkyVLM | 36.5 | 34.7 |
| SkyCLIP ViT-B/32 (external reference) | 8.6 | 6.8 |
SkyCLIP is a general remote-sensing model while SkyVLM is a specialist trained on these templated captions, so this is a reference point rather than a claim of a better model.
Vision encoder
RESISC45 linear probe (45 scene classes, 6,300 test images, frozen features; no RESISC45 data seen in training):
| Encoder | Test accuracy | Macro-F1 |
|---|---|---|
| Random ViT (untrained) | 43.5% | 0.425 |
| MIM-pretrained (100 epochs, 200k unlabelled images) | 78.6% | 0.785 |
| MIM + CLIP alignment (this model's encoder) | 88.4% | 0.885 |
| ImageNet ViT-S/16 (supervised, ~14M labelled images) | 88.5% | 0.885 |
Language model
FineWeb-Edu validation perplexity 21.7 (72.6M parameters, 983M training tokens). SkyScript caption validation loss bottomed at 0.76 and ended at 0.94 as the model began to memorize the captions (10% of each batch, ~18 passes).
Sanity checks
Main-object category (building, road, railway, ...) is right 61.5% of the time versus 21% for always guessing the most common one (measured on the variant without CLIP alignment); all outputs follow the caption template.
Single training run per condition; no variance estimates.
Limitations
- Trained only on SkyScript's templated, OpenStreetMap-derived captions: it will not follow free-form prompts or answer questions.
- Typical errors are the right category with the wrong specifics (an apartment block described as an office). Exact main-object match is 21%, and references have several valid descriptions per image, so n-gram and exact-match scores understate quality.
- Images are ~200 px at 224×224 input; fine detail is limited.
- Not evaluated outside SkyScript imagery, on other regions' sensors, or for any safety-critical use.
Data and license
Apache-2.0. Trained on SkyScript (MIT), which pairs Google Earth Engine imagery with OpenStreetMap-derived captions. OpenStreetMap data is © OpenStreetMap contributors (ODbL); imagery remains subject to its providers' terms. FineWeb-Edu (ODC-By) was used for language-model pretraining.
- Downloads last month
- 34