DistillPath-KS16-UNI2h

A 22M ViT-S/16 pathology tile encoder distilled from UNI2-h (681M ViT-H/14) into the kaiko ViT-S/16 student using backbone-token distillation on 6,000 public TCGA slides.

For the ImageNet-initialized variant, see DistillPath-IS16-UNI2h.

Paper: DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance Ramon Kaspar, Andrey Ignatov, Valentina Boeva. ETH Zurich. Published at the ECCV 2026 Workshop on Medical Foundation Models and Benchmarks (MedFM-Bench).

Model details

Property Value
Architecture ViT-S/16 (vit_small_patch16_224 in timm)
Parameters 21.7M
Feature dimension 384
Input size 224 x 224
Normalization mean=(0.5, 0.5, 0.5), std=(0.5, 0.5, 0.5)
Student initialization kaiko ViT-S/16 (pathology-pretrained)
Teacher UNI2-h (681M, ViT-H/14, CC BY-NC-ND 4.0)
Training data 6,000 TCGA H&E whole-slide images, 32 cohorts
Training steps 50,000 (batch size 256, approx. 24-29 GPU-hours on 1x RTX 4090)

Benchmark results

Benchmark DistillPath-KS16-UNI2h kaiko baseline UNI2-h teacher
EVA mean (7 tasks) 0.772 0.764 0.806
HEST mean (9 tasks) 0.375 0.349 0.414
PLISM score 0.484 0.307 0.333

See the paper for per-task results and comparisons with all four DistillPath variants.

Usage

Load directly from the Hub with timm:

import timm

model = timm.create_model(
    "hf_hub:RamonK/DistillPath-KS16-UNI2h",
    pretrained=True,
    num_classes=0,
)
model.eval()

Or load manually:

import timm
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file

model = timm.create_model("vit_small_patch16_224", pretrained=False, num_classes=0)
path = hf_hub_download("RamonK/DistillPath-KS16-UNI2h", "model.safetensors")
state_dict = load_file(path)
model.load_state_dict(state_dict, strict=True)
model.eval()

This model uses normalization mean=(0.5, 0.5, 0.5) and std=(0.5, 0.5, 0.5), inherited from the kaiko student. Apply this normalization to input tiles before inference:

from torchvision import transforms

transform = transforms.Compose([
    transforms.Resize(224),
    transforms.CenterCrop(224),
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]),
])

Distillation recipe

The recipe reads only the teacher's final class and patch tokens (no teacher pretraining heads required):

  • Class-token loss: cosine distance + RKD (relational knowledge distillation)
  • Patch-token loss: cosine distance after bicubic grid resizing (teacher 16x16 to student 14x14)
  • Optimizer: AdamW, lr=1e-4, weight decay 0.05, cosine decay, 500 warmup steps
  • Projector: DINO-style MLP (384 to 2048 to 2048 to 256 to d_t), discarded after training

Full details in the paper and the DistillPath repository.

License

This model is released under the Kaiko Non-Commercial Public License, inherited from the kaiko ViT-S/16 student weights. The UNI2-h teacher is released under CC BY-NC-ND 4.0. The distillation process used UNI2-h only to generate supervisory outputs; the released student contains no UNI2-h weights. This model is intended solely for non-commercial academic research.

Citation

@inproceedings{kaspar2026distillpath,
  title     = {DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance},
  author    = {Kaspar, Ramon and Ignatov, Andrey and Boeva, Valentina},
  booktitle = {Medical Foundation Models and Benchmarks (MedFM-Bench), ECCV 2026},
  year      = {2026}
}
Downloads last month
4
Safetensors
Model size
21.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including RamonK/DistillPath-KS16-UNI2h