Instructions to use impresso-project/image-classification-siglip2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use impresso-project/image-classification-siglip2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-classification", model="impresso-project/image-classification-siglip2") pipe("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/hub/parrots.png")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForImageClassification processor = AutoProcessor.from_pretrained("impresso-project/image-classification-siglip2") model = AutoModelForImageClassification.from_pretrained("impresso-project/image-classification-siglip2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Impresso newspaper image classifier (multi-facet L0-L3)
Classifies image regions cropped from historical newspapers along four independent facets,
with one shared google/siglip2-so400m-patch14-384 backbone and one linear
classifier of 28 outputs: each facet is a fixed slice of the logits (one linear head per
facet, stored side by side). Part of the Impresso project.
- L0: visual content: Image, Not Image
- L1: technique: Not photograph, Photograph
- L2: communication goal: Advertising, Decorative, Entertainment, Informative or Illustrative
- L3: type: Caricature or Editorial cartoon or Humoristic drawing, Comic strip, Game, Graph, Human representations - Fashion visual, Human representations - Portrait, Human representations - Scene, Illustrated story, Map - Geological map, Map - Geopolitical map, Map - Physical map or Roadmap, Map - Plan, Map - Weather map, Non-figurative visual content, Object, Ornament or Illustrated Title, Other, Scenery or Landscape, Technical drawing, Weather infographic
Every facet is decided by argmax, including L0. L1-L3 apply only to crops that L0
classifies as Image. They were trained on image rows only, so their output on a
non-image is meaningless and must be discarded.
Test results
Held-out test split of impresso-project/newspaper-image-classification. Each facet
is scored on its own eligible rows: L1-L3 on gold Image rows.
| facet | classes | accuracy | macro-F1 |
|---|---|---|---|
| L0: visual content | 2 | 0.9945 | 0.9895 |
| L1: technique | 2 | 0.9647 | 0.9646 |
| L2: communication goal | 4 | 0.9163 | 0.8709 |
| L3: type | 20 | 0.8092 | 0.8082 |
L0 as a filter: image recall (retention) 0.9974, not-image precision 0.9858, not-image recall 0.9789, AUROC 0.9991.
Classes with no test support (not evaluated): l3/Map - Geological map, l3/Map - Plan.
Training
Recipe
- What trains: backbone frozen; only the 28-way classifier trains (32,284 parameters).
- Classifier input: mean of all output tokens (SigLIP's attention-pooling head is not used).
- Optimizer: AdamW, learning rate 0.001, weight decay 0.05.
- Schedule: cosine schedule with 5% warmup, 20 epochs, batch 64, fp16 mixed precision.
- Loss: sum of per-facet cross-entropies (λ = 1), L1-L3 masked on
Not Imagerows; label smoothing 0. - Data: augmentation
none, samplingnone, seed 42.
The full resolved config is in resolved_config.yaml.
Benchmark of the recipe
The recipe was chosen on a multi-seed benchmark: trained on the train split only, best
validation epoch kept, seeds 42, 123, 456 (wandb group siglip2-b-hf-bs64-lr1e3). Mean ± std over seeds:
| metric | validation | test |
|---|---|---|
| mean macro-F1 (L1-L3) | 0.8584 ± 0.0015 | 0.8638 ± 0.0022 |
| mean accuracy (L1-L3) | 0.8984 ± 0.0004 | 0.8905 ± 0.0018 |
| L0 macro-F1 | 0.9867 ± 0.0007 | 0.9895 ± 0.0000 |
| L1 macro-F1 | 0.9673 ± 0.0013 | 0.9659 ± 0.0013 |
| L2 macro-F1 | 0.8434 ± 0.0009 | 0.8635 ± 0.0026 |
| L3 macro-F1 | 0.7645 ± 0.0057 | 0.7619 ± 0.0067 |
This model
- Final run: train + validation splits merged (7812 rows), a fixed 20 epochs (the budget from the multi-seed benchmark), no model selection on held-out data. Test was untouched until this evaluation.
- Data:
impresso-project/newspaper-image-classification@f92e9c1b383b0404b61630e507f58932bb66d70b(configdefault). - Code: impresso-newspaper-image-classification @
045f1da. - Weights SHA-256:
b253c6d42810d299eef4634fd1f3c113d49ab8bb7fa178ac48229b4995f99ed2.
Input contract
Aspect-preserving LANCZOS resize and center padding to a 384x384
square (pad colour [0, 0, 0]), rescale to [0, 1], then normalize with
mean [0.5, 0.5, 0.5] / std [0.5, 0.5, 0.5]. The stock image processor does not pad, so pad first
(pad_to_square below, identical to the training code); on a padded square its resize is a no-op.
Usage
A standard transformers image classifier: no custom code, no trust_remote_code.
import torch
from PIL import Image
from transformers import AutoImageProcessor, AutoModelForImageClassification
repo = "impresso-project/image-classification-siglip2"
model = AutoModelForImageClassification.from_pretrained(repo).eval()
processor = AutoImageProcessor.from_pretrained(repo)
cfg = model.config
prep = cfg.impresso_preprocessing
def pad_to_square(img, size, color):
# Aspect-preserving LANCZOS resize, then center padding (the training contract).
scale = min(size / img.width, size / img.height)
w, h = max(1, int(img.width * scale)), max(1, int(img.height * scale))
out = Image.new("RGB", (size, size), tuple(color))
out.paste(img.resize((w, h), Image.Resampling.LANCZOS), ((size - w) // 2, (size - h) // 2))
return out
img = pad_to_square(Image.open("crop.jpg").convert("RGB"), prep["image_size"], prep["pad_color"])
with torch.no_grad():
logits = model(**processor(img, return_tensors="pt")).logits[0]
# One softmax per facet over its slice of the 28 logits ({'l0': [0, 2], 'l1': [2, 4], 'l2': [4, 8], 'l3': [8, 28]}).
pred = {}
for facet, (a, b) in cfg.facet_slices.items():
probs = logits[a:b].softmax(-1)
pred[facet] = (cfg.facets[facet][int(probs.argmax())], float(probs.max()))
if pred["l0"][0] != "Image":
pred = {"l0": pred["l0"]} # L1-L3 are meaningless for a non-image
Do not use pipeline("image-classification"): it applies one softmax over all 28 labels.
For Image / Not Image filtering only, read the l0 slice: every facet comes from the same
backbone pass, so skipping the others saves nothing.
Files
model.safetensors (weights), config.json (architecture, label vocabularies facets and their
logit slices facet_slices, input contract impresso_preprocessing), preprocessor_config.json,
resolved_config.yaml (training config + data/code revisions), test_results.json,
model.safetensors.sha256.
Limitations
- Trained on Impresso's annotated newspaper crops; quality on other collections, periods or scan pipelines is not measured.
- Small classes have few test examples, so their per-class scores are noisy; classes without test support are not evaluated at all.
- L2 (communication goal) labels can be inherited from L3 correspondences, so agreement with them does not independently validate those rules.
- The per-facet softmax probabilities are not calibrated (no temperature scaling): use them to rank, not as true probabilities.
License
The weights follow the backbone licence (apache-2.0). Rights on the training images are
described on the dataset card.
- Downloads last month
- 31
Model tree for impresso-project/image-classification-siglip2
Base model
google/siglip2-so400m-patch14-384Evaluation results
- Test mean macro-F1 over L1/L2/L3 on impresso-project/newspaper-image-classificationtest set self-reported0.881