A Novel Clinically Explainable Vision Transformer for OCT-Based Retinal Disease Classification: Integrating UniMIE Enhancement and Grad-CAM Interpretability | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article A Novel Clinically Explainable Vision Transformer for OCT-Based Retinal Disease Classification: Integrating UniMIE Enhancement and Grad-CAM Interpretability Vishal Upmanu, Jaya Singh, Pranshu Saxena, Jagendra Singh, Shilpa Srivastava, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6510478/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Explainable and precise Optical Coherence Tomography (OCT) image classification plays an essential role in early retinal disease detection and follow-up for conditions like Choroidal Neovascularization (CNV), Diabetic Macular Edema (DME), and Drusen. Conventionally applied deep learning models, including transformers and convolutional neural networks, yield state-of-the-art classification results but suffer from lacking interpretability within clinical practice and the difficulty in subtle differentiation among diseases. This paper suggests a clinically interpretable vision transformer (ViT) model, combining Universal Medical Image Enhancement (UniMIE)-based image enhancement, hierarchical ViT feature extraction, and Gradient-weighted Class Activation Mapping (Grad-CAM) based visualization to enhance both classification accuracy and interpretability. The Proposed ViT model is tested on UCSD and Mendeley OCT datasets, which has a top accuracy of 98.84% in 5-fold cross-validation, outperforming existing convolutional neural based and transformer-based methods. The model also attains an AUC-ROC value of 99.45%, showing better discriminative ability in CNV, DME, Drusen, and Normal classes. An extensive hyperparameter tuning approach optimized the dropout rate, encoder depth, and learning rate to improve accuracy and generalization. Grad-CAM visualizations also add clinical interpretability, where decision-critical retinal areas are pointed out, ensuring predictions to be consistent with pathological features noticed by ophthalmologists. Comparative analysis against current deep learning models reaffirms that the proposed ViT Model offers top-class performance without sacrificing the ability to solve primary shortcomings like deficiency in fine-grained classification, overfitting, and interpretability. Health sciences/Diseases/Eye diseases Health sciences/Signs and symptoms/Eye manifestations Optical Coherence Tomography Vision Transformer UniMIE Grad-CAM Explainability Retinal Disease Classification Deep Learning Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6510478","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":475932639,"identity":"39730b54-74e3-4de0-baf4-6e1bec97499f","order_by":0,"name":"Vishal Upmanu","email":"","orcid":"","institution":"R D Engineering College","correspondingAuthor":false,"prefix":"","firstName":"Vishal","middleName":"","lastName":"Upmanu","suffix":""},{"id":475932640,"identity":"62d955b5-6200-4799-9305-76d84887ea30","order_by":1,"name":"Jaya Singh","email":"","orcid":"","institution":"ABES Engineering College","correspondingAuthor":false,"prefix":"","firstName":"Jaya","middleName":"","lastName":"Singh","suffix":""},{"id":475932641,"identity":"ccc969ba-4c6d-45cf-b4fd-f387300f3158","order_by":2,"name":"Pranshu Saxena","email":"","orcid":"","institution":"Bennett University","correspondingAuthor":false,"prefix":"","firstName":"Pranshu","middleName":"","lastName":"Saxena","suffix":""},{"id":475932642,"identity":"b27d9409-b16e-4c48-9b00-aec122ebe4c9","order_by":3,"name":"Jagendra Singh","email":"","orcid":"","institution":"Bennett University","correspondingAuthor":false,"prefix":"","firstName":"Jagendra","middleName":"","lastName":"Singh","suffix":""},{"id":475932643,"identity":"2bce5ffd-b06c-4129-9ff4-59dbf60d1121","order_by":4,"name":"Shilpa Srivastava","email":"","orcid":"","institution":"Christ University","correspondingAuthor":false,"prefix":"","firstName":"Shilpa","middleName":"","lastName":"Srivastava","suffix":""},{"id":475932644,"identity":"d52cefa5-86fb-40d4-b214-f92d36b09644","order_by":5,"name":"Aprna Tripathi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABSElEQVRIie3Rv2vCQBQH8BcCcTnjeiFq/oUXAlrB2n8lcpAuhdotYyCgS9u5xX8iU6Fb5MAuWldLHAShXTpYujgE2vNXqYd07nBfuOMddx8ejwNQUfnPMcTSAZu62HdZF3MALToKcE20GDCQiP8HgQ0BLl0cIaX+zet8Bfmlacdvi05nUij1ikM7hAenXigOlj40K0l6QOjsqe5eAza65aEb32GmU24G1ggy9zE2GfUh8CQC08CgBBAN6rsxEQQ4qVkRZFrCCQrC2xJxBLHyDTn/EORZd3bkTBBv5cOXTFAQe9vlYt0l1XFH2oLURJdUJu5saNhl9NAoj676BJnucpM1IsxYIoY68ZF594ekmnUN6z2sotPvJZ8kb7HqZDx4icLsNJmM+XQZtiq30vh0/y10e2YAZP9RP8UR8qtobV/CnqioqKiowDez4m53TRmi6gAAAABJRU5ErkJggg==","orcid":"","institution":"Manipal University Jaipur","correspondingAuthor":true,"prefix":"","firstName":"Aprna","middleName":"","lastName":"Tripathi","suffix":""}],"badges":[],"createdAt":"2025-04-23 08:23:31","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6510478/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6510478/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":87700028,"identity":"af9ce070-a3f0-422b-96bb-0d8fd12e1465","added_by":"auto","created_at":"2025-07-28 07:09:49","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1042859,"visible":true,"origin":"","legend":"","description":"","filename":"SmartAttendanceSystemAbstractshilpamay07.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6510478/v1_covered_28f7dffe-3301-46d9-b5e1-7cc86b391cbe.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Novel Clinically Explainable Vision Transformer for OCT-Based Retinal Disease Classification: Integrating UniMIE Enhancement and Grad-CAM Interpretability","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Optical Coherence Tomography, Vision Transformer, UniMIE, Grad-CAM, Explainability, Retinal Disease Classification, Deep Learning","lastPublishedDoi":"10.21203/rs.3.rs-6510478/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6510478/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eExplainable and precise Optical Coherence Tomography (OCT) image classification plays an essential role in early retinal disease detection and follow-up for conditions like Choroidal Neovascularization (CNV), Diabetic Macular Edema (DME), and Drusen. Conventionally applied deep learning models, including transformers and convolutional neural networks, yield state-of-the-art classification results but suffer from lacking interpretability within clinical practice and the difficulty in subtle differentiation among diseases. This paper suggests a clinically interpretable vision transformer (ViT) model, combining Universal Medical Image Enhancement (UniMIE)-based image enhancement, hierarchical ViT feature extraction, and Gradient-weighted Class Activation Mapping (Grad-CAM) based visualization to enhance both classification accuracy and interpretability. The Proposed ViT model is tested on UCSD and Mendeley OCT datasets, which has a top accuracy of 98.84% in 5-fold cross-validation, outperforming existing convolutional neural based and transformer-based methods. The model also attains an AUC-ROC value of 99.45%, showing better discriminative ability in CNV, DME, Drusen, and Normal classes. An extensive hyperparameter tuning approach optimized the dropout rate, encoder depth, and learning rate to improve accuracy and generalization. Grad-CAM visualizations also add clinical interpretability, where decision-critical retinal areas are pointed out, ensuring predictions to be consistent with pathological features noticed by ophthalmologists. Comparative analysis against current deep learning models reaffirms that the proposed ViT Model offers top-class performance without sacrificing the ability to solve primary shortcomings like deficiency in fine-grained classification, overfitting, and interpretability.\u003c/p\u003e","manuscriptTitle":"A Novel Clinically Explainable Vision Transformer for OCT-Based Retinal Disease Classification: Integrating UniMIE Enhancement and Grad-CAM Interpretability","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-25 17:47:08","doi":"10.21203/rs.3.rs-6510478/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9cbded2a-2113-4a0a-b2b7-5a52faf16745","owner":[],"postedDate":"June 25th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":50539131,"name":"Health sciences/Diseases/Eye diseases"},{"id":50539132,"name":"Health sciences/Signs and symptoms/Eye manifestations"}],"tags":[],"updatedAt":"2025-07-28T07:08:48+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-25 17:47:08","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6510478","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6510478","identity":"rs-6510478","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.