HER2 reader: a small ViT-S/16 encoder and HER2 slide classifier for routine H&E

Model weights and output data behind the paper "A small HER2 reader for routine H&E slides: real-world deployment" (Thang Tran, Lan Dang, 2026; arXiv link to follow). We publish them so that anyone can check our numbers, rerun our evaluation and build on our encoder.

Research use only. This is not a medical device and must not be used for clinical decisions. Our reader proposes a probability that a breast cancer is HER2-positive from its H&E slide, for a pathologist who reads it beside IHC and ISH, never instead of them.

What is in this repository

Path What it holds
encoders/ ViT-S/16 tile encoders (22M parameters) at every training stage, with timm parameter names
checkpoints/ the same models exactly as trained, in our original layout (including the self-distillation heads), for audit
reader/her2-jepa2-t2y-s4/ the deployed reader: encoder, five CLAM-SB fold heads, and a manifest with architecture, tile recipe, operating point and validation
data/predictions/ per-slide predictions, by seed: TCGA-BRCA out of fold, Yale-HER2 external, TCGA-BRCA + CPTAC-BRCA under the published protocol
data/paper/ data behind the paper's tables and figures
CHECKSUMS.sha256 SHA-256 of every file

Encoders

File (encoders/) Stage TCGA-to-Yale AUC*
stage1_ssl-vits16-final.safetensors self-supervised pretraining (DINOv2 + iBOT + KoLeo) on 98.6M TCGA tiles 0.589
stage2_distilled-hoptimus0-step25000.safetensors distilled from H-optimus-0 (CLS + 4x4 patch targets), step 25,000 0.762
stage3a_jepa-round1.safetensors one JEPA refinement round 0.798
stage3b_jepa-round2_deployed-encoder.safetensors second JEPA round: the deployed encoder 0.812
r12a_distilled-1um-step25000.safetensors stage 3b distilled again at 1.0 um/px (not adopted) 0.813**
r12b_distilled-1um-jepa6000.safetensors r12a plus one JEPA round (not adopted) 0.812**

* CLAM-SB reader trained on 490 TCGA-BRCA slides, tested on 192 Yale-HER2 slides, mean of 10 seeds (our protocol). ** Under the published reader's protocol (TCGA + CPTAC training, all tiles, 40 attention-MIL models); the deployed encoder scores 0.809 there.

Load an encoder (timm)

import timm
from safetensors.torch import load_file

enc = timm.models.vision_transformer.VisionTransformer(
    img_size=224, patch_size=16, embed_dim=384, depth=12, num_heads=6, mlp_ratio=4,
    qkv_bias=True, init_values=1.0, class_token=True, no_embed_class=False, reg_tokens=0,
    global_pool="token", fc_norm=False, num_classes=0)
enc.load_state_dict(load_file("encoders/stage3b_jepa-round2_deployed-encoder.safetensors"), strict=True)
enc.eval()
# input: float32 [N, 3, 224, 224], RGB, ImageNet mean (0.485, 0.456, 0.406) and std (0.229, 0.224, 0.225)
# output: [N, 384], the class token after the final LayerNorm

Tiles for our reader: 128 um squares read at 0.5 um/px (the coarsest pyramid level at or below 0.5 um/px), resized to 224 px by nearest neighbour, tissue only; up to 400 tiles per slide, taken at an even stride across the section. The manifest in reader/ gives every step.

The deployed reader

Five CLAM-SB heads (5-fold cross-validation on 490 TCGA-BRCA slides, seed 4) read one bag of 400 tile vectors. The reader's probability is the mean of the five heads' softmax probabilities; at or above 0.27228 it proposes "positive". That threshold comes from TCGA-BRCA's out-of-fold predictions, and a laboratory should refit it on its own slides, because scores shift between hospitals.

Test set (never trained on) AUC
Yale-HER2, 192 slides (5-model ensemble, seed 4) 0.815 [0.753, 0.876]
BCNB, 1,058 core biopsies (mean of 10 seeds) 0.628
HEROHE test split, 150 slides (mean of 10 seeds) 0.746
HEROHE IHC 2+ cases, 85 slides (amplified vs not) 0.761

At its threshold on Yale-HER2: sensitivity 0.72, specificity 0.79. On a 12-core CPU server it reads a slide in 5.2 to 5.6 s.

Limitations

  • Its 490 TCGA-BRCA training slides were size-biased: negatives carry less tissue than positives. This inflates the Yale AUC by about 0.04; on representative training sets our readers reach about 0.81.
  • Probabilities run low at hospitals other than TCGA's; set the threshold per laboratory.
  • On equivocal IHC 2+ cases (AUC 0.76 on HEROHE) it can help order and prioritise ISH. It does not replace ISH.
  • The best published reader (H-optimus-0, 1.1B parameters, Yale AUC 0.907) remains well ahead. The gap lies in the encoder.

Training data and attributions

  • Encoder: TCGA diagnostic slides (NCI Genomic Data Commons, open access). Distillation teacher: H-optimus-0 (Bioptimus), Apache-2.0. See NOTICE.
  • Reader heads: 490 TCGA-BRCA slides with HER2 status from GDC clinical data.
  • Evaluation cohorts: Yale-HER2 (TCIA HER2-TUMOR-ROIS, CC BY 4.0), CPTAC-BRCA (open), HEROHE (CC BY-NC-ND 3.0, evaluation only) and BCNB (non-commercial, evaluation only). Per-slide data for HEROHE and BCNB are not published here; only aggregate numbers appear in the paper.

Licences

  • Model weights (encoders/, checkpoints/, reader/): CC BY-NC 4.0 (LICENSE-weights.md).
  • Data (data/): CC BY 4.0 (LICENSE-data.md).
  • The distillation teacher's Apache-2.0 notice: NOTICE.

Citation

@misc{tran2026her2reader,
  title  = {A small HER2 reader for routine H\&E slides: real-world deployment},
  author = {Tran, Thang and Dang, Lan},
  year   = {2026},
  note   = {Model weights and data: https://huggingface.co/cloudkites/her2-reader}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support