Luna — Retinal Foundation Model with Hybrid Clinical & Anatomical Supervision
Luna is a multimodal, multi-task retinal foundation model for high-resolution color fundus photographs (CFPs). Instead of pixel reconstruction, it is pre-trained by predicting masked clinical tokens (disease label, age, sex, eye laterality) from the image, while a UNETR decoder jointly supervises vessel / optic-disc segmentation.
Architecturally it is a single-stream VL-BERT: a SigLIP-2 ViT-Base/16 (512px) vision encoder, a BERT-Base masked-LM initialized from PubMedBERT, and a UNETR head for anatomy.
Paper: N/A
Model Details
| Attribute | Value |
|---|---|
| Architecture | SigLIP-2 ViT-Base/16 vision encoder + BERT-Base masked-LM (single-stream VL-BERT fusion, 128 compressed vision tokens) + UNETR segmentation decoder |
| Parameters | 243M total; ≈93.5M in the SigLIP-2 ViT-Base vision encoder that downstream tasks reuse |
| Input resolution | 512 × 512 RGB (CLAHE-enhanced) |
| Vision encoder init. | google/siglip2-base-patch16-512 |
| Text encoder init. | NeuML/pubmedbert-base-embeddings tokenizer, extended with [LABEL], [AGE], [SEX], [DIR] special tokens |
| Pre-training objective | Masked clinical-token prediction (label / age / sex / direction) + vessel & optic-disc Dice loss, not pixel-level MAE reconstruction |
| Pre-training data | 361,519 CFPs (52K public + 309K private institutional images) |
| License | MIT |
Prompt templates used during pre-training: "A CFP image of {CLS}", "age is {CLS} years old",
"gender is {CLS}", "this is {CLS} direction eye".
Files in this repository
This is the full accelerate training state at epoch 399:
| File | Contents |
|---|---|
model.safetensors / pytorch_model.bin |
Model weights (243,122,497 parameters, fp32) |
optimizer.bin, scheduler.bin, scaler.pt |
AdamW / LR-scheduler / GradScaler state, for exact resumption |
random_states_{0..3}.pkl |
RNG states of the four training processes |
Usage
Full training, fine-tuning and evaluation code lives in the
GitHub repository — see its README.md for the Upstream training,
Downstream training, Zero-Shot and Segmentation sections. Checkpoint paths there expect
pytorch_model.bin; passing the containing directory instead resumes optimizer and scheduler state as well.
Training Data
* = private institutional data, not available externally.
Pre-training datasets
| Dataset | Reference |
|---|---|
| LAG | Li, L. et al. Attention Based Glaucoma Detection: A Large-Scale Database and CNN Model. CVPR (2019). |
| ODIR | Ocular Disease Intelligent Recognition (ODIR-2019). Peking University Grand Challenge (2019). |
| PARAGUAY | Castillo Benítez, V.E. et al. Dataset from fundus images for the study of diabetic retinopathy. Data in Brief 36, 107068 (2021). |
| G1020 | Bajwa, M.N. et al. G1020: A Benchmark Retinal Fundus Image Dataset for Computer-Aided Glaucoma Detection. IJCNN (2020). |
| FUND-OCT | Hassan, T., Akram, M.U., Werghi, N. & Nazir, N. A hybrid convolutional framework for the automated extraction of retinal lesions and lesion-influenced grading of human retinal pathology. IEEE J. Biomed. Health Inform. 25(1) (2020). |
| Drishti-GS1 | Sivaswamy, J. et al. Drishti-GS: Retinal Image Dataset for Optic Nerve Head Segmentation. ISBI 53–56 (2014). |
| HRF | Budai, A. et al. Robust Vessel Segmentation in Fundus Images. Int. J. Biomed. Imaging 2013, 154860 (2013). |
| ORIGA | Zhang, Z. et al. ORIGA(-light): An Online Retinal Fundus Image Database for Glaucoma Analysis and Research. EMBC 3065–3068 (2010). |
| OIA-DDR | Li, T. et al. Diagnostic Assessment of Deep Learning Algorithms for Diabetic Retinopathy Screening. Information Sciences 501, 511–522 (2019). |
| SUSTech-SYSU | Lin, L. et al. The SUSTech-SYSU dataset for automated exudate detection and diabetic retinopathy grading. Sci. Data 7, 409 (2020). |
| JICHI | Takahashi, H. et al. Applying artificial intelligence to disease staging: Deep learning for improved staging of diabetic retinopathy. PLoS ONE 12, e0179790 (2017). |
| CHAKSU | Kumar J H, R. et al. Chákṣu: A glaucoma specific fundus image database. Sci. Data 10, 70 (2023). |
| DR1 & DR2 | Pires, R. et al. Advancing Bag-of-Visual-Words Representations for Lesion Classification in Retinal Images. PLoS ONE 9, e96814 (2014). |
| DeepDRiD | Liu, R. et al. DeepDRiD: Diabetic Retinopathy—Grading and Image Quality Estimation Challenge. Patterns 3, 100512 (2022). |
| AIROGS-light-V2 | Kiefer, R. Glaucoma Dataset: EyePACS-AIROGS-light-V2. Kaggle (2024). doi:10.34740/KAGGLE/DSV/7802508 |
| FIDVS | Jin, K. et al. Fundus Image Dataset for Vessel Segmentation. Kaggle (2025). doi:10.34740/KAGGLE/DS/7319687 |
| JustRAIGS | Madadi, Y. et al. JustRAIGS: Justified Referral in AI Glaucoma Screening Challenge. IEEE Trans. Med. Imaging (2025). |
| SMDG | Kiefer, R. SMDG, A Standardized Fundus Glaucoma Dataset. Kaggle (2023). doi:10.34740/KAGGLE/DS/2329670 |
| AMC health-screen* | Institutional cohort, Asan Medical Center — not externally published. |
| AMC clinic* | Institutional cohort, Asan Medical Center — not externally published. |
Downstream evaluation datasets
| Dataset | Reference |
|---|---|
| APTOS 2019 | Karthik, Maggie & Dane, S. APTOS 2019 Blindness Detection. Kaggle (2019). |
| IDRiD | Porwal, P. et al. Indian Diabetic Retinopathy Image Dataset (IDRiD): A Database for Diabetic Retinopathy Screening Research. Data 3, 25 (2018). |
| MESSIDOR | Decencière, E. et al. Feedback on a Publicly Distributed Image Database: The MESSIDOR Database. Image Anal. Stereol. 231–234 (2014). |
| PAPILA | Kovalyk, O. et al. PAPILA: Dataset with fundus images and clinical data of both eyes of the same patient for glaucoma assessment. Sci. Data 9, 291 (2022). |
| Retina | jr2ngb. Retina Cataract Dataset. Kaggle (2016). |
| JSIEC | Cen, L.-P. et al. Automatic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks. Nat. Commun. 12, 4828 (2021). |
| BRSET | Nakayama, L.F. et al. A Brazilian Multilabel Ophthalmological Dataset (BRSET). PhysioNet (2024). |
| ARIA | Farnell, D.J.J. et al. Enhancement of blood vessels in digital fundus photographs via the application of multiscale line operators. J. Franklin Inst. 345, 748–765 (2008). |
| MAPLES-DR | Lepetit-Aimon, G. et al. MAPLES-DR: MESSIDOR Anatomical and Pathological Labels for Explainable Screening of Diabetic Retinopathy. Sci. Data 11, 914 (2024). |
| AMC health-screen subset* | Institutional cohort, Asan Medical Center — systemic/anthropometric & lifestyle biomarker fine-tuning. |
| AMC CAC cohort* | Institutional cohort, Asan Medical Center — coronary artery calcium classification. |
Citation
@article{lim_luna_2026,
title = {Multimodal Foundation Model for High-Resolution Fundus Photographs Incorporating
Clinical Metadata via Prediction of Demographics and Anatomical Structures},
author = {Lim, Hyunseok and Kim, Junseok and Oh, Joonseo and Kim, Kanghyun and Lim, Jongsoo
and Jeong, Jinhoon and Jeong, Hy and Kim, Yoonjeon and Kim, Namkug},
year = {2026},
}
License
Released under the MIT License — Copyright (c) 2026 Hyunseok Lim. The third-party datasets listed above carry their own licenses and access terms, and the Asan Medical Center institutional cohorts are not redistributable.