EndoVLM ViT-B/16
EndoVLM is an endoscopy vision-language pre-training model trained on 348K unordered endoscopic image sets paired with comprehensive gastrointestinal clinical reports, via Anatomy-Guided Sparse Pooling (AGSP), Progressive Semantic-Aware Alignment (PSAA), and Semantic-Concentrated Masked Autoencoder (SC-MAE).
This repository hosts the released ViT-B/16 checkpoint of EndoVLM. Code is available at github.com/Scatteredrain/EndoVLM.
Model Details
| Vision backbone | DINOv3 ViT-B/16 (vit_base_patch16, 4 storage tokens) |
| Text encoder | BiomedCLIP PubMedBERT (microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224) |
| Checkpoint file | endovlm_vitb16.pth |
| Size | 895,158,199 bytes (~854 MB, fp32, model weights only) |
Usage
pip install torch open-clip-torch timm transformers
git clone https://github.com/Scatteredrain/EndoVLM.git
cd EndoVLM
# download endovlm_vitb16.pth into pretrained/
import torch
import torch.nn.functional as F
import open_clip
from endovlm import build_endovlm
# Build the model, then load the released weights
model = build_endovlm(model_type='vit_base_patch16')
msg = model.load_state_dict(
torch.load('pretrained/endovlm_vitb16.pth', map_location='cpu', weights_only=False)['model'],
strict=False,
)
model = model.cuda().eval()
tokenizer = open_clip.get_tokenizer('hf-hub:microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224')
# Fine-grained (anatomy-level) image features
latent, _, _ = model.forward_encoder(image_tensor, 0.0)
cls_tokens = latent[:, 0]
patch_tokens = latent[:, 1 + model.encoder.n_storage_tokens:]
feat = model.image_feat_projection(torch.cat([cls_tokens, patch_tokens.mean(dim=1)], dim=1))
image_features = F.normalize(model.image_projection_fg(feat), p=2, dim=-1)
# Text features
text_features = F.normalize(
model.forward_text(tokenizer(["An endoscopic image of pylorus."])), p=2, dim=-1
)
A ready-to-run zero-shot demo (upper-GI anatomy recognition on bundled samples) is provided in inference.ipynb in the GitHub repository.
See the paper and GitHub repository for details.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support