RG-DermNet: A Multimodal Attention-Based Model with Residual Block Usage for Skin Lesion Classification
Skin cancer accounts for nearly one-third of all diagnosed tumors worldwide, making early and accurate recognitioncritical for improving patient outcomes. In this work, we propose RG-DermNet, a multimodal deep learning framework that integrates skin lesion images with structured clinical metadata through a residual gated-attention (RG-ATT) fusion mechanism. The architecture combines CNN- and Transformer-based visual backbones with a lightweight one-hot encoding pipeline for metadata, enabling effective cross-modal interaction. The proposed model is evaluated using a patient-wise cross-validation protocol across four dermatological datasets with heterogeneous metadata. On PAD-UFES-20, using Caformer-B36 as the visual backbone, RG-DermNet achieves an accuracy of 0.75 ± 0.05, balanced accuracy of 0.78 ± 0.03, F1-score of 0.77 ± 0.04, and AUC of 0.95 ± 0.01, outperforming existing multimodal baselines under the same evaluation setting. In addition, a SHAP-based analysis provides insights into the contribution of clinical metadata to the model’s predictions, supporting both performance gains and interpretability.
Github
https://github.com/wyctorfogos/rg-dermnet
Demo
https://huggingface.co/spaces/wyctorfogos/GradCAMPlusPlus_SkinLesion