Transformers
Safetensors
surgical-video
spatio-temporal-grounding
medical-vision-language-model
eccv-2026
Instructions to use linzher/RefineRank with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use linzher/RefineRank with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("linzher/RefineRank", device_map="auto") - Notebooks
- Google Colab
- Kaggle
RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding
Official checkpoints for the ECCV 2026 MedVidU Workshop paper RefineRank.
RefineRank couples two frozen backbones β a MedVLM (Qwen2.5-VL
architecture) and GroundingDINO β with a compact 1.25M-parameter trainable
module, RefineNet (QueryConditionedProposalAdapter). RefineNet uses
MedVLM language and regional features to predict coordinate corrections and
box-quality scores for GroundingDINO proposals; a parameter-free decoder then
returns the highest-scoring original or refined box.
- Code, training & inference guide: https://github.com/linzhe001/RefineRank
- Headline result: 0.421 STG mIoU on the archived MedVidBench Community leaderboard snapshot (27 July 2026) β the best STG mIoU among the ten ranking metrics on that snapshot.
- Controlled evaluation (video-separated split over CholecTrack20 / CoPESD / EgoSurgery): STG mIoU 0.2719 β 0.4534 over the frozen MedVLM + GroundingDINO baseline.
Contents
This repo hosts the complete checkpoints/ tree expected by the code
repository β three flat folders, each holding its core files directly:
checkpoints/
βββ vlm/ # frozen MedVLM, HF format (~16 GB)
β βββ config.json, generation_config.json, tokenizer*, preprocessor_config.json, ...
β βββ model-00001..00004-of-00004.safetensors
βββ grounding_dino/
β βββ groundingdino_swinb_cogcoor.pth # frozen GroundingDINO SwinB (~895 MB)
βββ refinenet/ # trained RefineNet, this work (~5 MB)
βββ proposal_adapter_full.pt # SHA-256: 932e479bβ¦463d9b
βββ deployment_manifest.json
refinenet/proposal_adapter_full.ptis the exact checkpoint behind the paper's MedVidBench submission (run_iter132_submission). SHA-256:932e479b854c3d5fbafee25a3fcf9e6481e864fe98a6867502d6ebff39463d9b.vlm/andgrounding_dino/are third-party frozen weights (uAI-NEXUS-MedVLM by UII-AI and GroundingDINO by IDEA-Research), mirrored here for one-stop reproducibility. They are never fine-tuned by RefineRank; please follow their original licenses and cite the original works.
Usage
pip install "huggingface_hub[hf_transfer]" # hf_transfer optional, faster
hf download linzher/RefineRank --local-dir . # restores the checkpoints/ tree
git clone https://github.com/linzhe001/RefineRank
cd RefineRank
pip install -r requirements.txt
# place the downloaded checkpoints/ next to interface.py, then:
python interface.py predict # auto-discovers checkpoints/refinenet/
Citation
@inproceedings{jiang2026refinerank,
title = {RefineRank: Joint Box Refinement and Ranking for Surgical
Spatio-Temporal Grounding},
author = {Jiang, Linzhe and Huang, Jiayuan and Zhang, Changhao and
Jiang, Chunyang and Mao, Zhehua and
Garcia-Peraza-Herrera, Luis C. and Hoque, Mobarak I.},
booktitle = {ECCV Workshops (MedVidU)},
year = {2026}
}
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support