EndoVLM ViT-B/16

EndoVLM is an endoscopy vision-language pre-training model trained on 348K unordered endoscopic image sets paired with comprehensive gastrointestinal clinical reports, via Anatomy-Guided Sparse Pooling (AGSP), Progressive Semantic-Aware Alignment (PSAA), and Semantic-Concentrated Masked Autoencoder (SC-MAE).

This repository hosts the released ViT-B/16 checkpoint of EndoVLM. Code is available at github.com/Scatteredrain/EndoVLM.

Model Details

Vision backbone DINOv3 ViT-B/16 (vit_base_patch16, 4 storage tokens)
Text encoder BiomedCLIP PubMedBERT (microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224)
Checkpoint file endovlm_vitb16.pth
Size 895,158,199 bytes (~854 MB, fp32, model weights only)

Usage

pip install torch open-clip-torch timm transformers
git clone https://github.com/Scatteredrain/EndoVLM.git
cd EndoVLM
# download endovlm_vitb16.pth into pretrained/
import torch
import torch.nn.functional as F
import open_clip
from endovlm import build_endovlm

# Build the model, then load the released weights
model = build_endovlm(model_type='vit_base_patch16')
msg = model.load_state_dict(
    torch.load('pretrained/endovlm_vitb16.pth', map_location='cpu', weights_only=False)['model'],
    strict=False,
)
model = model.cuda().eval()

tokenizer = open_clip.get_tokenizer('hf-hub:microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224')

# Fine-grained (anatomy-level) image features
latent, _, _ = model.forward_encoder(image_tensor, 0.0)
cls_tokens = latent[:, 0]
patch_tokens = latent[:, 1 + model.encoder.n_storage_tokens:]
feat = model.image_feat_projection(torch.cat([cls_tokens, patch_tokens.mean(dim=1)], dim=1))
image_features = F.normalize(model.image_projection_fg(feat), p=2, dim=-1)

# Text features
text_features = F.normalize(
    model.forward_text(tokenizer(["An endoscopic image of pylorus."])), p=2, dim=-1
)

A ready-to-run zero-shot demo (upper-GI anatomy recognition on bundled samples) is provided in inference.ipynb in the GitHub repository.

See the paper and GitHub repository for details.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support