Abstract
Background Our aim is to demonstrate that multimodal deep learning can enhance the accuracy of classifying skin lesions using both images and textual descriptions (e.g., demographics, clinical information) compared to a classifier that learn only on images.
Methods
We used the HAM10000 and ISIC 2017 datasets in our study containing 10,000 and 2,750 skin lesion images, respectively. We combined the images with patients’ data (e.g., sex, age, lesion location) for training and evaluating a multimodal deep learning classification model. The dataset was split into 70% for training the model, 20% for the validation set, and 10% for the testing set. We compared the multimodal model’s performance to well-known deep learning models that only use images for classification.
Results
We used accuracy and area under the curve (AUC) receiver operating characteristic (ROC) as the metrics to compare the models’ performance. Our multimodal model outperformed the competitors and achieved the best results. Our model’s accuracy and AUCROC was 0.9411 and 0.9426, respectively, on HAM10000. On ISIC 2017, our model’s accuracy and AUCROC was 0.7971 and 0.8253, respectively.
Conclusion
Our study showed that a multimodal deep learning model can outperform traditional deep learning models for skin lesion classification on the HAM10000 and ISIC 2017 datasets. Our approach can enable primary care clinicians to screen for skin cancer in patients (residing in areas lacking access to expert dermatologists) with higher accuracy and reliability.
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This project was funded by the Translational Research Informing Useful and Meaningful Precision Health (TRIUMPH) grant from the University of Missouri-Columbia.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
This study used only publicly available dataset. The HAM10000 and ISIC2017 datasets.
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.
Yes
Footnotes
This is the uppdated version
Data Availability
The HAM10000 is available at https://doi.org/10.7910/DVN/DBW86T; the ISIC 2017 dataset is available at https://challenge.isic-archive.com/data/#2017.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.