OmniRAS models and datasets
Shared source footage: please also cite SurgeNet
YT-Chole Triplets and the YT-Chole shared-video subset of OmniRAS-PR
(OmniRAS-PR/yt-chole-shared-videos/) use robotic-cholecystectomy footage from
the Surgical YouTube collection released with the SurgeNet paper by
Jaspers et al. These two tasks share source-video batches; their clip
boundaries, labels, and training/validation assignments differ. The larger
OmniRAS-PR subset, labeled used-in-paper/, is a separate source collection.
SurgeNet provides the underlying public source footage. The custom phase and tool–verb–target annotations, task-specific clips, and benchmark splits in this release are part of the OmniRAS work. If you use YT-Chole Triplets or the shared-video phase subset, please cite SurgeNet as well as OmniRAS.
Tim J. M. Jaspers et al. Scaling up self-supervised learning for improved surgical foundation models. Medical Image Analysis, 108:103873, 2026. Paper · Official repository and citation.
@article{JASPERS2026103873,
title = {Scaling up self-supervised learning for improved surgical foundation models},
author = {Tim J.M. Jaspers and Ronald L.P.D. de Jong and Yiping Li and Carolus H.J. Kusters and Franciscus H.A. Bakker and Romy C. van Jaarsveld and Gino M. Kuiper and Richard van Hillegersberg and Jelle P. Ruurda and Willem M. Brinkman and Josien P.W. Pluim and Peter H.N. de With and Marcel Breeuwer and Yasmina Al Khalil and Fons van der Sommen},
journal = {Medical Image Analysis},
volume = {108},
pages = {103873},
year = {2026},
doi = {10.1016/j.media.2025.103873},
url = {https://www.sciencedirect.com/science/article/pii/S1361841525004190}
}
Dataset release
Dataset files are organized as:
OmniRAS-PR/used-in-paper/: primary phase contexts and microclips.OmniRAS-PR/yt-chole-shared-videos/: shared-video phase clips.YT-Chole-Triplets/: tool–verb–target clips.
The annotation archive contains labels and train/validation splits. Dataset downloads require separate manual approval.
Model checkpoints
Surgical V-JEPA 2.1 checkpoints. Both frozen backbones and all nine downstream probes listed below are available in this repository with manual access approval.
Two frozen (raw CPT) backbones plus a set of fine-tuned downstream-probe checkpoints.
For each probe/variant, the checkpoint from the seed with the best documented task metric (mAP, IVT mAP, F1@10, or AP) was selected.
How to use — what you need to run each thing
Each probe checkpoint stores the task head plus the encoder. The suffix (_ft4/_ft8/_full) describes the fine-tuning recipe — how many top blocks were unfrozen during training — not what is saved to disk. Almost all probes save the complete 48-block encoder and are therefore self-contained.
Self-contained probes (all except
esad_ft4). The file holds the full fine-tuned encoder (blocks 0–47) + the task head. Load it alone — no separate backbone needed — regardless of whether the recipe was_ft4,_ft8, or_full.The one exception:
esad_ft4. This file stores only the last 4 fine-tuned blocks (44–47) + head, not the frozen base. To run it you must also load the frozen 2B backbone (frozen_backbone_2B/) for blocks 0–43, then overlay the probe's top blocks + head.Frozen backbones (
frozen_backbone_1B/,frozen_backbone_2B/) are the raw pretrained encoders. On their own they emit embeddings only (no task head) — use them to train your own head on a new task, or as the base foresad_ft4.
All probes were fine-tuned on the 2B backbone; the 1B backbone is provided for independent use.
The "Needs backbone?" column in the table below states this per file (verified by inspecting each checkpoint's stored encoder blocks).
Frozen backbones
| Path | Model | Epoch | Params | Source run |
|---|---|---|---|---|
| frozen_backbone_1B/OmniRAS_1B_frozen.pth.tar | ViT-g (1B) | e19 | ~1B | surg_2_1_vitg384_cleandata/vitg384_n16g12_weak/e19.pth.tar |
| frozen_backbone_2B/OmniRAS_2B_frozen.pth.tar | ViT-G (2B) | e199 | ~2B | daos_shakeout/vitG384_prod_37M_n256_8766188/e199.pth.tar |
The 1B backbone is the vitg384_cleandata e19 checkpoint (40 blocks, width 1408). The 2B backbone is the prod_37M e199 checkpoint (48 blocks, width 1664; 200 epochs, 6000 steps total). Checkpoint identities were verified from their stored encoder tensors.
Downstream probes (fine-tuned on the 2B backbone)
Suffix convention: _ft4 = last-4-block unfreeze, _ft8 = last-8-block unfreeze, _full = full unfreeze.
| Path | Probe | Variant | Seed | Metric (best seed) | Needs backbone? | Source run |
|---|---|---|---|---|---|---|
| OmniRAS_probes/OmniRAS_grasp_phase_ft4.pt | GraSP Phase | last4 unfreeze | s1 | 84.27 mAP | no (self-contained) | grasp_unfreeze_384/prod37m_e199_last4_s1 |
| OmniRAS_probes/OmniRAS_grasp_step_ft4.pt | GraSP Step (21-cls) | last4 unfreeze, head d768/dr0.0 | s0 | 59.35 mAP | no (self-contained) | grasp_unfreeze_384/step_prod37m_e199_ft_last4_d768dr00_s0 |
| OmniRAS_probes/OmniRAS_triplet_full.pt | Triplet (SITL/YT-Chole, CholecT80-style) | full (48/48) unfreeze, bs=2 | s1 | 41.17 IVT mAP | no (self-contained) | triplet/tripft_full_prod37M_e199_bs2_s1 |
| OmniRAS_probes/OmniRAS_triplet_ft8.pt | Triplet | last8 unfreeze | s0 | 41.18 IVT mAP | no (self-contained) | triplet/tripft_last8_prod37M_e199_s0 |
| OmniRAS_probes/OmniRAS_triplet_ft4.pt | Triplet | last4 unfreeze | s1 | 37.99 IVT mAP | no (self-contained) | triplet/tripft_prod37M_e199_s1 |
| OmniRAS_probes/OmniRAS_sar_rarp50_full.pt | SAR-RARP50 (action seg.) | full unfreeze | s2, ep19 | 91.82 F1@10 | no (self-contained) | sarft_sweep/prod37M_e199_full_s2 |
| OmniRAS_probes/OmniRAS_sar_rarp50_ft4.pt | SAR-RARP50 | last4 unfreeze | s0, ep20 | ~90.9 F1@10 | no (self-contained) | sarft_sweep/prod37M_e199 |
| OmniRAS_probes/OmniRAS_esad_ft4.pt | ESAD (detection) | last4 unfreeze + augmentation, SWA-smoothed | s1 | 0.2585 AP_mean | yes (2B) | esad/probes/ft_runs/esad_double_prod37M_e199_ft_last4_swa5peak_smoothed |
| OmniRAS_probes/OmniRAS_ft4_OmniRAS-PR.pt | Anonymized surgical phase-recognition probe (PR) | last4 unfreeze | — | 62.29 mAP | no (self-contained) | — |
Checkpoint inventory verified and completed 2026-10-06. Use SHA256SUMS to verify downloaded checkpoint files.
Please cite OmniRAS
If you use anything from this repository or release—including datasets, annotations, splits, code, model weights, or checkpoints—please cite the OmniRAS paper:
Leonardo Borgioli, Neil Getty, et al. OmniRAS: Standardizing Foundation Model Training and Evaluation in Robot-Assisted Surgery. arXiv:2608.31048, 2026. Paper.
@misc{borgioli2026omniras,
title = {{OmniRAS}: Standardizing Foundation Model Training and Evaluation in Robot-Assisted Surgery},
author = {Leonardo Borgioli and Neil Getty and Wenli Xiu and Jessica Cassiani and Alvaro Ducas and Carlos Agustin Orda and Hira Waris and Fangfang Xia and Rick Stevens and Pier Cristoforo Giulianotti and Milos Zefran},
year = {2026},
eprint = {2608.31048},
archivePrefix = {arXiv},
primaryClass = {eess.IV},
doi = {10.48550/arXiv.2608.31048},
url = {https://arxiv.org/abs/2608.31048}
}
