Luna — Retinal Foundation Model with Hybrid Clinical & Anatomical Supervision

Luna is a multimodal, multi-task retinal foundation model for high-resolution color fundus photographs (CFPs). Instead of pixel reconstruction, it is pre-trained by predicting masked clinical tokens (disease label, age, sex, eye laterality) from the image, while a UNETR decoder jointly supervises vessel / optic-disc segmentation.

Architecturally it is a single-stream VL-BERT: a SigLIP-2 ViT-Base/16 (512px) vision encoder, a BERT-Base masked-LM initialized from PubMedBERT, and a UNETR head for anatomy.

Paper: N/A

Code: https://github.com/loopback-kr/Luna

Model Details

Attribute Value
Architecture SigLIP-2 ViT-Base/16 vision encoder + BERT-Base masked-LM (single-stream VL-BERT fusion, 128 compressed vision tokens) + UNETR segmentation decoder
Parameters 243M total; ≈93.5M in the SigLIP-2 ViT-Base vision encoder that downstream tasks reuse
Input resolution 512 × 512 RGB (CLAHE-enhanced)
Vision encoder init. google/siglip2-base-patch16-512
Text encoder init. NeuML/pubmedbert-base-embeddings tokenizer, extended with [LABEL], [AGE], [SEX], [DIR] special tokens
Pre-training objective Masked clinical-token prediction (label / age / sex / direction) + vessel & optic-disc Dice loss, not pixel-level MAE reconstruction
Pre-training data 361,519 CFPs (52K public + 309K private institutional images)
License MIT

Prompt templates used during pre-training: "A CFP image of {CLS}", "age is {CLS} years old", "gender is {CLS}", "this is {CLS} direction eye".

Files in this repository

This is the full accelerate training state at epoch 399:

File Contents
model.safetensors / pytorch_model.bin Model weights (243,122,497 parameters, fp32)
optimizer.bin, scheduler.bin, scaler.pt AdamW / LR-scheduler / GradScaler state, for exact resumption
random_states_{0..3}.pkl RNG states of the four training processes

Usage

Full training, fine-tuning and evaluation code lives in the GitHub repository — see its README.md for the Upstream training, Downstream training, Zero-Shot and Segmentation sections. Checkpoint paths there expect pytorch_model.bin; passing the containing directory instead resumes optimizer and scheduler state as well.

Training Data

* = private institutional data, not available externally.

Pre-training datasets

Dataset Reference
LAG Li, L. et al. Attention Based Glaucoma Detection: A Large-Scale Database and CNN Model. CVPR (2019).
ODIR Ocular Disease Intelligent Recognition (ODIR-2019). Peking University Grand Challenge (2019).
PARAGUAY Castillo Benítez, V.E. et al. Dataset from fundus images for the study of diabetic retinopathy. Data in Brief 36, 107068 (2021).
G1020 Bajwa, M.N. et al. G1020: A Benchmark Retinal Fundus Image Dataset for Computer-Aided Glaucoma Detection. IJCNN (2020).
FUND-OCT Hassan, T., Akram, M.U., Werghi, N. & Nazir, N. A hybrid convolutional framework for the automated extraction of retinal lesions and lesion-influenced grading of human retinal pathology. IEEE J. Biomed. Health Inform. 25(1) (2020).
Drishti-GS1 Sivaswamy, J. et al. Drishti-GS: Retinal Image Dataset for Optic Nerve Head Segmentation. ISBI 53–56 (2014).
HRF Budai, A. et al. Robust Vessel Segmentation in Fundus Images. Int. J. Biomed. Imaging 2013, 154860 (2013).
ORIGA Zhang, Z. et al. ORIGA(-light): An Online Retinal Fundus Image Database for Glaucoma Analysis and Research. EMBC 3065–3068 (2010).
OIA-DDR Li, T. et al. Diagnostic Assessment of Deep Learning Algorithms for Diabetic Retinopathy Screening. Information Sciences 501, 511–522 (2019).
SUSTech-SYSU Lin, L. et al. The SUSTech-SYSU dataset for automated exudate detection and diabetic retinopathy grading. Sci. Data 7, 409 (2020).
JICHI Takahashi, H. et al. Applying artificial intelligence to disease staging: Deep learning for improved staging of diabetic retinopathy. PLoS ONE 12, e0179790 (2017).
CHAKSU Kumar J H, R. et al. Chákṣu: A glaucoma specific fundus image database. Sci. Data 10, 70 (2023).
DR1 & DR2 Pires, R. et al. Advancing Bag-of-Visual-Words Representations for Lesion Classification in Retinal Images. PLoS ONE 9, e96814 (2014).
DeepDRiD Liu, R. et al. DeepDRiD: Diabetic Retinopathy—Grading and Image Quality Estimation Challenge. Patterns 3, 100512 (2022).
AIROGS-light-V2 Kiefer, R. Glaucoma Dataset: EyePACS-AIROGS-light-V2. Kaggle (2024). doi:10.34740/KAGGLE/DSV/7802508
FIDVS Jin, K. et al. Fundus Image Dataset for Vessel Segmentation. Kaggle (2025). doi:10.34740/KAGGLE/DS/7319687
JustRAIGS Madadi, Y. et al. JustRAIGS: Justified Referral in AI Glaucoma Screening Challenge. IEEE Trans. Med. Imaging (2025).
SMDG Kiefer, R. SMDG, A Standardized Fundus Glaucoma Dataset. Kaggle (2023). doi:10.34740/KAGGLE/DS/2329670
AMC health-screen* Institutional cohort, Asan Medical Center — not externally published.
AMC clinic* Institutional cohort, Asan Medical Center — not externally published.

Downstream evaluation datasets

Dataset Reference
APTOS 2019 Karthik, Maggie & Dane, S. APTOS 2019 Blindness Detection. Kaggle (2019).
IDRiD Porwal, P. et al. Indian Diabetic Retinopathy Image Dataset (IDRiD): A Database for Diabetic Retinopathy Screening Research. Data 3, 25 (2018).
MESSIDOR Decencière, E. et al. Feedback on a Publicly Distributed Image Database: The MESSIDOR Database. Image Anal. Stereol. 231–234 (2014).
PAPILA Kovalyk, O. et al. PAPILA: Dataset with fundus images and clinical data of both eyes of the same patient for glaucoma assessment. Sci. Data 9, 291 (2022).
Retina jr2ngb. Retina Cataract Dataset. Kaggle (2016).
JSIEC Cen, L.-P. et al. Automatic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks. Nat. Commun. 12, 4828 (2021).
BRSET Nakayama, L.F. et al. A Brazilian Multilabel Ophthalmological Dataset (BRSET). PhysioNet (2024).
ARIA Farnell, D.J.J. et al. Enhancement of blood vessels in digital fundus photographs via the application of multiscale line operators. J. Franklin Inst. 345, 748–765 (2008).
MAPLES-DR Lepetit-Aimon, G. et al. MAPLES-DR: MESSIDOR Anatomical and Pathological Labels for Explainable Screening of Diabetic Retinopathy. Sci. Data 11, 914 (2024).
AMC health-screen subset* Institutional cohort, Asan Medical Center — systemic/anthropometric & lifestyle biomarker fine-tuning.
AMC CAC cohort* Institutional cohort, Asan Medical Center — coronary artery calcium classification.

Citation

@article{lim_luna_2026,
  title   = {Multimodal Foundation Model for High-Resolution Fundus Photographs Incorporating
             Clinical Metadata via Prediction of Demographics and Anatomical Structures},
  author  = {Lim, Hyunseok and Kim, Junseok and Oh, Joonseo and Kim, Kanghyun and Lim, Jongsoo
             and Jeong, Jinhoon and Jeong, Hy and Kim, Yoonjeon and Kim, Namkug},
  year    = {2026},
}

License

Released under the MIT License — Copyright (c) 2026 Hyunseok Lim. The third-party datasets listed above carry their own licenses and access terms, and the Asan Medical Center institutional cohorts are not redistributable.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support