Accurate Skin Lesion Classification Using Multimodal Learning on the HAM10000 and ISIC 2017 Datasets

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

Background Our aim is to demonstrate that multimodal deep learning can enhance the accuracy of classifying skin lesions using both images and textual descriptions (e.g., demographics, clinical information) compared to a classifier that learn only on images. Methods We used the HAM10000 and ISIC 2017 datasets in our study containing 10,000 and 2,750 skin lesion images, respectively. We combined the images with patients’ data (e.g., sex, age, lesion location) for training and evaluating a multimodal deep learning classification model. The dataset was split into 70% for training the model, 20% for the validation set, and 10% for the testing set. We compared the multimodal model’s performance to well-known deep learning models that only use images for classification. Results We used accuracy and area under the curve (AUC) receiver operating characteristic (ROC) as the metrics to compare the models’ performance. Our multimodal model outperformed the competitors and achieved the best results. Our model’s accuracy and AUCROC was 0.9411 and 0.9426, respectively, on HAM10000. On ISIC 2017, our model’s accuracy and AUCROC was 0.7971 and 0.8253, respectively. Conclusion Our study showed that a multimodal deep learning model can outperform traditional deep learning models for skin lesion classification on the HAM10000 and ISIC 2017 datasets. Our approach can enable primary care clinicians to screen for skin cancer in patients (residing in areas lacking access to expert dermatologists) with higher accuracy and reliability.
Full text 3,371 characters · extracted from oa-doi-fallback · 4 sections · click to expand

Abstract

Background Our aim is to demonstrate that multimodal deep learning can enhance the accuracy of classifying skin lesions using both images and textual descriptions (e.g., demographics, clinical information) compared to a classifier that learn only on images.

Methods

We used the HAM10000 and ISIC 2017 datasets in our study containing 10,000 and 2,750 skin lesion images, respectively. We combined the images with patients’ data (e.g., sex, age, lesion location) for training and evaluating a multimodal deep learning classification model. The dataset was split into 70% for training the model, 20% for the validation set, and 10% for the testing set. We compared the multimodal model’s performance to well-known deep learning models that only use images for classification.

Results

We used accuracy and area under the curve (AUC) receiver operating characteristic (ROC) as the metrics to compare the models’ performance. Our multimodal model outperformed the competitors and achieved the best results. Our model’s accuracy and AUCROC was 0.9411 and 0.9426, respectively, on HAM10000. On ISIC 2017, our model’s accuracy and AUCROC was 0.7971 and 0.8253, respectively.

Conclusion

Our study showed that a multimodal deep learning model can outperform traditional deep learning models for skin lesion classification on the HAM10000 and ISIC 2017 datasets. Our approach can enable primary care clinicians to screen for skin cancer in patients (residing in areas lacking access to expert dermatologists) with higher accuracy and reliability. Competing Interest Statement The authors have declared no competing interest. Funding Statement This project was funded by the Translational Research Informing Useful and Meaningful Precision Health (TRIUMPH) grant from the University of Missouri-Columbia. Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study used only publicly available dataset. The HAM10000 and ISIC2017 datasets. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Footnotes This is the uppdated version Data Availability The HAM10000 is available at https://doi.org/10.7910/DVN/DBW86T; the ISIC 2017 dataset is available at https://challenge.isic-archive.com/data/#2017.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00