DistillPath-IS16-HOpt0

A 22M ViT-S/16 pathology tile encoder distilled from H-optimus-0 (1.1B ViT-g/14) into an ImageNet-21k pretrained ViT-S/16 student using backbone-token distillation on 6,000 public TCGA slides.

This is the ImageNet-initialized variant. For the stronger kaiko-initialized variant, see DistillPath-KS16-HOpt0.

Paper: DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance Ramon Kaspar, Andrey Ignatov, Valentina Boeva. ETH Zurich. Published at the ECCV 2026 Workshop on Medical Foundation Models and Benchmarks (MedFM-Bench).

Model details

Property Value
Architecture ViT-S/16 (vit_small_patch16_224 in timm)
Parameters 21.7M
Feature dimension 384
Input size 224 x 224
Normalization mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225)
Student initialization ImageNet-21k ViT-S/16
Teacher H-optimus-0 (1.1B, ViT-g/14, Apache 2.0)
Training data 6,000 TCGA H&E whole-slide images, 32 cohorts
Training steps 50,000 (batch size 256)

Benchmark results

Benchmark DistillPath-IS16-HOpt0 IN21K baseline H-optimus-0 teacher
EVA mean (7 tasks) 0.768 0.729 0.803
HEST mean (9 tasks) 0.363 0.311 0.415
PLISM score 0.526 0.383 0.480

See the paper for per-task results.

Usage

Load directly from the Hub with timm:

import timm

model = timm.create_model(
    "hf_hub:RamonK/DistillPath-IS16-HOpt0",
    pretrained=True,
    num_classes=0,
)
model.eval()

Or load manually:

import timm
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file

model = timm.create_model("vit_small_patch16_224", pretrained=False, num_classes=0)
path = hf_hub_download("RamonK/DistillPath-IS16-HOpt0", "model.safetensors")
state_dict = load_file(path)
model.load_state_dict(state_dict, strict=True)
model.eval()

This model uses ImageNet normalization: mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225).

from torchvision import transforms

transform = transforms.Compose([
    transforms.Resize(224),
    transforms.CenterCrop(224),
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])

Distillation recipe

The recipe reads only the teacher's final class and patch tokens (no teacher pretraining heads required):

  • Class-token loss: cosine distance + RKD (relational knowledge distillation)
  • Patch-token loss: cosine distance after bicubic grid resizing (teacher 16x16 to student 14x14)
  • Optimizer: AdamW, lr=1e-4, weight decay 0.05, cosine decay, 500 warmup steps
  • Projector: DINO-style MLP (384 to 2048 to 2048 to 256 to d_t), discarded after training

Full details in the paper and the DistillPath repository.

License

This model is released under the Apache 2.0 License. Both the student (ImageNet-21k ViT-S/16) and teacher (H-optimus-0) are Apache 2.0, so this model is fully permissive, including for commercial use.

Citation

@inproceedings{kaspar2026distillpath,
  title     = {DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance},
  author    = {Kaspar, Ramon and Ignatov, Andrey and Boeva, Valentina},
  booktitle = {Medical Foundation Models and Benchmarks (MedFM-Bench), ECCV 2026},
  year      = {2026}
}
Downloads last month
4
Safetensors
Model size
21.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including RamonK/DistillPath-IS16-HOpt0