Instructions to use cloudkites/her2-reader with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- timm
How to use cloudkites/her2-reader with timm:
import timm model = timm.create_model("hf-hub:cloudkites/her2-reader", pretrained=True) - Notebooks
- Google Colab
- Kaggle
HER2 reader: a small ViT-S/16 encoder and HER2 slide classifier for routine H&E
Model weights and output data behind the paper "A small HER2 reader for routine H&E slides: real-world deployment" (Thang Tran, Lan Dang, 2026; arXiv link to follow). We publish them so that anyone can check our numbers, rerun our evaluation and build on our encoder.
Research use only. This is not a medical device and must not be used for clinical decisions. Our reader proposes a probability that a breast cancer is HER2-positive from its H&E slide, for a pathologist who reads it beside IHC and ISH, never instead of them.
What is in this repository
| Path | What it holds |
|---|---|
encoders/ |
ViT-S/16 tile encoders (22M parameters) at every training stage, with timm parameter names |
checkpoints/ |
the same models exactly as trained, in our original layout (including the self-distillation heads), for audit |
reader/her2-jepa2-t2y-s4/ |
the deployed reader: encoder, five CLAM-SB fold heads, and a manifest with architecture, tile recipe, operating point and validation |
data/predictions/ |
per-slide predictions, by seed: TCGA-BRCA out of fold, Yale-HER2 external, TCGA-BRCA + CPTAC-BRCA under the published protocol |
data/paper/ |
data behind the paper's tables and figures |
CHECKSUMS.sha256 |
SHA-256 of every file |
Encoders
File (encoders/) |
Stage | TCGA-to-Yale AUC* |
|---|---|---|
stage1_ssl-vits16-final.safetensors |
self-supervised pretraining (DINOv2 + iBOT + KoLeo) on 98.6M TCGA tiles | 0.589 |
stage2_distilled-hoptimus0-step25000.safetensors |
distilled from H-optimus-0 (CLS + 4x4 patch targets), step 25,000 | 0.762 |
stage3a_jepa-round1.safetensors |
one JEPA refinement round | 0.798 |
stage3b_jepa-round2_deployed-encoder.safetensors |
second JEPA round: the deployed encoder | 0.812 |
r12a_distilled-1um-step25000.safetensors |
stage 3b distilled again at 1.0 um/px (not adopted) | 0.813** |
r12b_distilled-1um-jepa6000.safetensors |
r12a plus one JEPA round (not adopted) | 0.812** |
* CLAM-SB reader trained on 490 TCGA-BRCA slides, tested on 192 Yale-HER2 slides, mean of 10 seeds (our protocol). ** Under the published reader's protocol (TCGA + CPTAC training, all tiles, 40 attention-MIL models); the deployed encoder scores 0.809 there.
Load an encoder (timm)
import timm
from safetensors.torch import load_file
enc = timm.models.vision_transformer.VisionTransformer(
img_size=224, patch_size=16, embed_dim=384, depth=12, num_heads=6, mlp_ratio=4,
qkv_bias=True, init_values=1.0, class_token=True, no_embed_class=False, reg_tokens=0,
global_pool="token", fc_norm=False, num_classes=0)
enc.load_state_dict(load_file("encoders/stage3b_jepa-round2_deployed-encoder.safetensors"), strict=True)
enc.eval()
# input: float32 [N, 3, 224, 224], RGB, ImageNet mean (0.485, 0.456, 0.406) and std (0.229, 0.224, 0.225)
# output: [N, 384], the class token after the final LayerNorm
Tiles for our reader: 128 um squares read at 0.5 um/px (the coarsest pyramid level at or below 0.5 um/px), resized to
224 px by nearest neighbour, tissue only; up to 400 tiles per slide, taken at an even stride across the section. The
manifest in reader/ gives every step.
The deployed reader
Five CLAM-SB heads (5-fold cross-validation on 490 TCGA-BRCA slides, seed 4) read one bag of 400 tile vectors. The reader's probability is the mean of the five heads' softmax probabilities; at or above 0.27228 it proposes "positive". That threshold comes from TCGA-BRCA's out-of-fold predictions, and a laboratory should refit it on its own slides, because scores shift between hospitals.
| Test set (never trained on) | AUC |
|---|---|
| Yale-HER2, 192 slides (5-model ensemble, seed 4) | 0.815 [0.753, 0.876] |
| BCNB, 1,058 core biopsies (mean of 10 seeds) | 0.628 |
| HEROHE test split, 150 slides (mean of 10 seeds) | 0.746 |
| HEROHE IHC 2+ cases, 85 slides (amplified vs not) | 0.761 |
At its threshold on Yale-HER2: sensitivity 0.72, specificity 0.79. On a 12-core CPU server it reads a slide in 5.2 to 5.6 s.
Limitations
- Its 490 TCGA-BRCA training slides were size-biased: negatives carry less tissue than positives. This inflates the Yale AUC by about 0.04; on representative training sets our readers reach about 0.81.
- Probabilities run low at hospitals other than TCGA's; set the threshold per laboratory.
- On equivocal IHC 2+ cases (AUC 0.76 on HEROHE) it can help order and prioritise ISH. It does not replace ISH.
- The best published reader (H-optimus-0, 1.1B parameters, Yale AUC 0.907) remains well ahead. The gap lies in the encoder.
Training data and attributions
- Encoder: TCGA diagnostic slides (NCI Genomic Data Commons, open access). Distillation teacher: H-optimus-0
(Bioptimus), Apache-2.0. See
NOTICE. - Reader heads: 490 TCGA-BRCA slides with HER2 status from GDC clinical data.
- Evaluation cohorts: Yale-HER2 (TCIA HER2-TUMOR-ROIS, CC BY 4.0), CPTAC-BRCA (open), HEROHE (CC BY-NC-ND 3.0, evaluation only) and BCNB (non-commercial, evaluation only). Per-slide data for HEROHE and BCNB are not published here; only aggregate numbers appear in the paper.
Licences
- Model weights (
encoders/,checkpoints/,reader/): CC BY-NC 4.0 (LICENSE-weights.md). - Data (
data/): CC BY 4.0 (LICENSE-data.md). - The distillation teacher's Apache-2.0 notice:
NOTICE.
Citation
@misc{tran2026her2reader,
title = {A small HER2 reader for routine H\&E slides: real-world deployment},
author = {Tran, Thang and Dang, Lan},
year = {2026},
note = {Model weights and data: https://huggingface.co/cloudkites/her2-reader}
}
- Downloads last month
- -