ERBVS: Enhanced Retinal Blood Vessel Segmentation using Multiple Modalities and Attention Mechanisms with Adversarial Training and Ensemble Deep Learning Operations | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article ERBVS: Enhanced Retinal Blood Vessel Segmentation using Multiple Modalities and Attention Mechanisms with Adversarial Training and Ensemble Deep Learning Operations Komal Umare Thool This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4553609/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract It would, therefore, require highly advanced prediction tools to enhance early diagnosis and preemptive mechanisms for all these burgeoning diseases. Fast and correct disease prediction and pre-emption have huge potential for changing clinical outcome and ensuring timely and effective interventions that reduce morbidity and mortality. Current predictive models, instrumental as they are, have been found faltering in precision, recall, accuracy, and timeliness. Such delays and inaccuracies often miss the therapeutic window or lead to misguided clinical decisions. In this work, we present a novel model that aims to quite dramatically improve the process of segmentation and classification. Our approach embeds Attention Mechanisms with Adversarial Training and Ensemble Deep Learning Operations, together with a multimodal approach, which places it substantially higher across several metrics. This improves the precision, accuracy, recall, and AUC by 8.5%, 8.3%, 4.9%, and 3.9%, respectively, for segmentation and classification, while reducing the classification delay by 5.9% in different situations. Not only does our model handle the intrinsic limitations of current methods, but it also shows flexibility for a wide range of clinical applications. The compelling improvements in classification and preemption metrics strengthen its potential to make a sea change in the disease prediction framework for establishing optimum patient outcomes and efficient scenarios of healthcare delivery. Retinal Image Segmentation Attention Mechanisms Adversarial Training Ensemble Deep Learning Diabetic Retinopathy Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Introduction Modern medicine rests on the concept of early and correct diagnosis. Correct diagnosis, if made on time, can greatly decrease morbidity and mortality by preventing the development of a complication. Yet, despite this reality and the enormous progress in science and technology and research, the penalty for delayed diagnosis and lost therapeutic windows is still borne by humanity. In these cases, predictive tools would be relevant and possible; therefore, progressive and accurate models are realized to inform timely clinical decisions. Digital health innovation has opened the vast potential of 'medical data' toward the creation of an era of truly personalized medicine if it can be correctly analyzed and interpreted. The human health is so intricate and dynamic that it requires a predictive model, not only attuned to the subtleties in disease progression but one which also recognizes subtle cues indicative of possible health anomalies in the future. Traditional methods, having served as essential foundational blocks, fall short of catering to the evolving demands of modern medicine. Their limitations are multiplied in situations where real-time decisions are imperatively needed, and even the smallest delays could add up to finish in less-than-desired patient results. Real-time impact of a sophisticated predictive model is multifaceted. Clinically, it ensures timely interventions to patients, thereby significantly reducing the progression to severe disease stages. It economically reduces the exorbitant costs associated with late-stage treatments and long hospital stays. Such robust predictive tools can also be used for epidemic control, resource allocation, and formulation of policies from a global health perspective. Here, one can see what value an improved model brings to the table: driving a healthcare revolution in delivery and outcomes regarding betterment of health for patients, and helping medical professionals in the field to always stay one step ahead with their therapeutic interventions. In these regards, a model that corrects the intrinsic deficiencies of the methods at hand, by rendering flexibility and accuracy under highly variable clinical settings, is overdue. In this work, a proposed game-changing tool in such diseases characterization and pre-emption is offered for the greater healthcare environment. Motivation & Objectives : Motivation : The more complex and heterogeneous diseases are, the more robust medical models are required to accurately locate their beginnings and progress. The dominant predictive models, instrumental in many ways, display certain lacunae that may prove very dangerous in a clinical setting. Mistimed diagnoses might miss out on treatment opportunities, while false alarms may initiate therapies that are unnecessary, both of which mean undue physical, emotional, and financial pressure on patients. This pressure to upgrade predictive tools comes with an added force—the medical community is moving toward an approach of personalized medicine. With each patient carrying a unique genetic makeup and life history, he embodies an individual medical enigma, and the tools which we implement on the inputs from these patients have to be agile while resolving these complex tasks in real-time scenarios. This many-fold dilemma gets further increased by the global health scenarios. The pandemics of the contemporary era have shown that not only individual early diagnosis and prevention act at the community, national, and global levels, but it also works vice versa process. The broader social and economic implications of outbreaks and the evident need for faster and more accurate predictive tools raise the motivation for this study process. Objectives: • Improvement in Accuracy: The model will be improved for better prediction accuracy, so as to identify diseases with high specificity, reducing false positives. • Improved Classification and Preemption: Strengthen model skills toward the correct classification of diseases and the preemption of possible health complications, ensuring timely clinical interventions for different scenarios. •Reduced Delays: Since time is of essence in a medical environment, this becomes one of the major objectives—to reduce time lags in prediction and offer quicker insight to healthcare professionals. • Across clinical scenarios, adaptability: The model has to be versatile and handle all kinds of situations viewed against the wide spectrum of diseases and their infinite manifestations. • Real-time implementation: Not only should it be capable of prediction, but it is also intended to be integrated into a clinical setup in real-time. That means healthcare professionals use this insight instantly for various use cases. Briefly, it will be an effort with one of the finest human aspirations—bridging gaps in disease prediction and preemption. These objectives stand before us, like guiding stars, ensuring that the outcomes really make a difference amidst the pulsing needs of modern medical landscapes. Review of existing models used for identification of Diabetic Retinopathy Segmentation of the retinal vessels, especially in rapidly changing medical imaging, has seen incredible leaps with the incorporation of advanced deep learning models. The literature shows concentrated effort towards improving the accuracy and efficiency of these models, evidenced by a lot of works on different issues related to retinal vessel segmentation. Deari et al. [ 1 ] went a step further into providing a deep learning framework that reviews block attention and switchable normalization in ensuring an improved segmentation process. The new technique in this case is radically different from the previous ones; hence, it provides more accurate and finer results. Again, Li et al. [ 2 ] came up with a globally – transformer and dual local-attention network. The model applied a deep–shallow hierarchical feature fusion approach in this case. This approach underlines the trend toward the application of complex architectures that could handle the intrinsic complexity of retotal vessel structures. Moreover, innovations are not only limited to the introduction of new network architectures. Kande et al. [ 3 ] proposed another variant of the U-Net model, called MSR U-Net, toward the same application. In this model, it is an extension of a broadly accepted standard in medical image segmentation, and it is such a proposal that proves the continuous development and optimization of already existing models. Concurrently, Li et al. [ 4 ] developed the dual-path progressive fusion network, DPF-Net, which gave a novel perspective in terms of the integration of multiple pathways to facilitate improved segmentation. On the other side of those very complex models, Aurangzeb et al. [ 5 ] and Arsalan et al. [ 6 ] develop an efficient and lightweight model for retinal vessel segmentation, with critical needs in computational efficiency, especially when resources are limited. This approach is more relevant to medical settings, where the processing of retinal images requires quickness and reliability. At the same time, semi-supervised learning paradigms are considered. For example, Shen et al. [ 7 ] have proposed an expert-guided knowledge distillation for vessel segmentation in a semi-supervised setting. Since the model proposed by them is close to being on one side point of view with full supervision and unsupervised on the other, it may be said to be positioned between these two major approaches. For example, the applications of AI in diagnostic systems, such as on glaucoma and diabetic retinopathy, have been systematically developed. Aurangzeb et al. [ 8 ] describe this holistic view of AI-enabled diagnostics, showing how advances in retinal vessel segmentation have very far-reaching implications. Further, Monemian and Rabbani [ 9 ] presented a computationally efficient method for extracting red lesions in retinal fundus images, hence showing the necessity for being concerned with only some of the features in retinal imaging. Shen et al. [ 10 ] and Xu et al. [ 11 ] further contributed with unified semi-supervised learning frameworks and automatic arteriole-venule segmentation methods, respectively. Other important aspects of retinal imaging that have been done to give greater understanding of the structural makeup of the retina are the reconstruction and visualization of macula vasculature investigated by Al-Hinnawi et al. [ 12 ]. In the same way, Sundar and Sumathy [ 13 ] worked on the classification of levels of diabetic retinopathy disease underpinning the critical role played by topological features and graph neural networks in medical diagnostics. The protection of sparse retinal templates, addressed by Sadeghpour et al. [ 14 ], underlines the increasing concern for security and privacy of medical data in the age of AI. This becomes relevant when medicine is increasingly reliant on digital technologies. Another work by Zhang et al. [ 15 ] was related to the detection of microaneurysms within retinal fundus images; for this purpose, they utilized a hierarchical pyramid network called T-Net to improve the accuracy of detection. Finally, Pang et al. [ 16 ] went beyond traditional CNNs in searching for more inherent symmetries in medical image segmentation. The present study puts one of the most important developments ongoing in this area in a nutshell that continues with continuous exploration of novel techniques, methodologies, and approaches to enhance medical image analysis with respect to accuracy, efficiency, etc. Ali et al. [ 17 ] proposed a hybrid convolutional neural network model for the automatic classification of diabetic retinopathy using fundus images. This model epitomizes how traditional image processing can be combined with deep learning to ensure an improved level of accuracy and efficiency. Again, Raiaan et al. [ 18 ] made contributions toward a lightweight, effective deep learning model representing a high-accuracy approach underlining the Spanner ability of the model in handling most images of diabetic retinopathy. The technique stands very special in balancing computational efficiency and performance. Jagadesh et al. [ 19 ] suggested a framework that integrated the IC2T model for segmentation with Rock Hyrax swarm-based coordination attention mechanism for classification. Their approach has been able to depict the potentials of using bio-inspired algorithms in the enhancement of medical image analysis with improved accuracy. Soni et al. [ 20 ], on the other hand, applied an IoT-based federated learning approach for the classification of hypertensive retinopathy lesions, which opened new avenues toward applying distributed learning systems within the medical domain. Vadduri and Kuppusamy [ 21 ] focused on the improvement of ocular healthcare by proposing a deep learning-based multi-class diabetically eye disease segmentation and classification approach. Their work decidedly brings to the fore the growing importance of deep learning in the different medical imaging tasks. Nazih et al. [ 22 ] applied the vision transformer model for severity prediction in diabetic retinopathy against fundus photography-based retina images, which brings out the increasing interest in transformer models in medical imaging. Naz et al. [ 23 ] proposed an ensembled deep convolutional generative adversarial network for grading imbalanced diabetic retinopathy recognition. The authors have tried to surmount the problem of class imbalance generally observed in most of the medical image datasets. Wong, Juwono, and Apriono [ 24 ] used transfer learning with optimization of parameters simultaneously along with a feature-weighted ECOC ensemble for diabetic retinopathy detection and grading. Feng et al. [ 25 ] dealt with the grading of diabetic retinopathy images using a graph neural network, a novel approach that brings out the importance of topological features in the context of medical imaging. Siebert, Graßhoff, and Rostalski [ 26 ] examined uncertainty analysis of deep kernel learning methods for diabetic retinopathy grading, placing much importance on the reliability and confidence levels expected of medical diagnoses. Kukkar et al. [ 27 ] have used IoMT systems socially implemented for the optimization of deep learning model parameters for the diabetic retinopathy classification problem. Very much, the study draws a gap between social systems and medical imaging, reflecting interdisciplinarity in modern research. Palaniswamy and Vellingiri [ 28 ] have shown that active internet of things and deep learning techniques could be applied to diagnose Diabetic Retinopathy using retinal fundus images, proving the use of technology in health. Liu and Chi [ 29 ] proposed a cross-lesion attention network for the accurate grading of diabetic retinopathy using fundus images, with regard to lesion type specificity. Mohan et al. [ 30 ] proposed DRFL, a federated learning framework for grading diabetic retinopathy using fundus images; it lays much emphasis on collaborative learning in a distributed environment. Hou et al. [ 31 ] proposed an image quality assessment-driven collaborative learning framework for image enhancement and classification toward diabetic retinopathy grading. That meant the image quality would impact the accuracy of diagnosis in the study. Nur-A-Alam et al. [ 32 ] proposed a faster RCNN-based detection based on fused features from retina images, which is evidence for the ongoing improvements in object detection models under medical imaging. Afterward, Kumar et al. [ 33 ] utilized a DL-UNet empowered auto-encoder-decoder for fundus image analysis, redefining retinal lesion segmentation and demonstrating the flexibility of neural networks in complex segmentation tasks. Zang et al. [ 34 ] worked on interpretable diabetic retinopathy diagnosis by means of biomarker activation maps, which interestingly combined diagnostic accuracy with the requirement for interpretability within AI systems. Lu et al. [ 35 ] presented a new perspective toward transferred multi- to mono-modal generation for the diagnosis of lupus retinopathy, which was an emerging interest in cross-modal medical image analysis. At last, Zhang et al. [ 36 ] investigated ADD-Net-based image quality assessment of diabetic retinopathy. This sets forth the critical image quality in the effectiveness of deep learning models. Radha and Karuna [ 37 ] presented a depthwise parallel attention UNet for segmentation of the retinal vessels, which proved the efficiency of attention mechanisms in enhancing the accuracy of segmentation models. The innovation is one more step to develop deep learning models specified for medical imaging. Pereira et al. [ 38 ] worked on the effective detection of fundus lesions using YOLOR-CSP architecture and slicing aided hyper inference. Their methodology couples state-of-the-art deep learning architectures with advanced inference strategies and represents the continuous change of methodologies in this area. Hussain [ 39 ] investigated exudate detection through the integration of retinal-based affine mapping with the design flow mechanism, contributing to lightweight architecture designs. This research answers a pressing need: that efficient and resource-light models for medical imaging are provided, especially in settings constrained by computational resources. Yin et al. [ 40 ] presented a dual-branch U-Net architecture for the segmentation of retinal lesions in fundus images. This further underlined the flexibility of the U-Net architecture across a wide range of medical imaging applications. Bar-David et al. [ 41 ] suggested elastic deformation of OCT images from diabetic macular edema patients for training deep learning models. Their work is focused on the preprocessing aspects of medical images and places much attention on how high the influence of image manipulation is during the training and results of deep learning models. Bhargav and Puhan [ 42 ] proposed a new contra-harmonic correlative attention loss for the segmentation of microaneurysms in fundus images, and this majorly brings out the part of the loss functions playing in the optimization involved in the segmentation tasks. Abdallah et al. [ 43 ] proposed a noise-estimation-based isotropic diffusion approach that corrects the methods of noise reduction, quite essential in clear and accurate analysis for the segmentation of retinal blood vessels. Singh et al. [ 44 ] proposed a deep-learning-based system which has proved to be efficient and automatic in blood vessel segmentation from retinal fundus images, hence showing the trend towards the use of deep learning in automating the complex tasks of medical image segmentation. Kumar and Singh [ 45 ] worked on the prediction of diseases in the retina by segmenting blood vessels and classifying them using ensemble-based deep learning approaches. This work helped demonstrate how segmentation acts at the joint of disease prediction, further powered by ensemble learning for high diagnostic accuracy. Upadhyay et al. [ 46 ] investigated learning multi-scale deep fusion for the extraction of retinal blood vessels in fundus images and proved that multi-scale approaches can become very efficient in the characterization of a wide range of vessel details. Kumar and Singh [ 47 ] proposed a segmentation approach based on the generalized nextreme value probability distribution function-based matched filter, allowing the interface of statistical methods with deep learning to improve segmentation. Mahapatra et al. [ 48 ] proposed a mean global based on hysteresis thresholding for the segmentation of retinal blood vessels using enhanced homomorphic filtering. This is perhaps one of the traditional image processing methods that extracts the most essential elements from the new thresholding techniques. Saranya et al. [ 49 ] researched blood vessel segmentation in retinal fundus images for proliferative diabetic retinopathy screening, calling attention to its application in specific disease screening. Gu, Tian, and Oh [ 50 ] contributed self-distillation and implicit neural representation-based retinal vessel segmentation. This is a very original technique that uses advanced neural network training techniques with implicit representations. Based on this review, it can be observed that several challenges consistently emerge across the different models: Data Diversity : The vast and varied nature of medical data poses integration and processing challenges. Scalability : Many models show robust results in controlled or pilot studies but falter when scaled to larger, diverse patient populations. Interpretability : The need for models that are both high-performing and interpretable remains a significant concern, especially in clinical settings. As the volume and variety of medical data grow, the drive to develop an optimal predictive model remains at the forefront of medical research process. This work aims to synthesize the learnings from past models and propose advancements for the future scopes. 3. Proposed design of an augmented Bioinspired Multidomain feature nextraction & selection model for Diabetic retinopathy Severity estimation via Ensemble learning process This paper proposes multiple threads of advanced technology, including modalities such as specialized convolutional layers that go really deep into the intricacies of retinal imagery scans. Adversarial training empowers this model further by having generative and discriminative networks engage in an intricate dance that pushes the model toward ever higher accuracy and greater robustness. This is complemented by ensemble deep learning operations—a symphony of various learning techniques harmoniously together decoding the most intricate patterns of DR. Therefore, the amalgamation would catapult the performance of the model in terms of accuracy, precision, and speed and manifest remarkable adaptability to diversity and volume in datasets. ERBVS is not a model with awesome powers of analysis, but it represents the torchbearer of innovation that ushers in a new epoch into the diagnostic capability segment pertaining to healthcare. In this work, an intricate marriage of the UNet architecture with an attention mechanism is proposed for the segmentation of the retinal vessels. The result will be a sophisticated model delineating not only the intricate vasculature of the retina but underlining salient features necessary for a detailed analysis. First, the images collected regarding the retina are fed into the model process. Let these images be represented by I, where I∈R(H×W×C), where H, W, and C are the image height, image width, and the number of channels respectively. The first stage of the model details passing I through the UNet's contracting path formulated as a series of convolutional operations. Let Cn ( I ) represent the n th convolutional layer applied to the image, defined via Eq. 1, $$Cn\left(I\right)=\sigma \left(Wn*I+bn\right)\dots \left(1\right)$$ Where, ∗ represents the convolution operation, Wn represents the weights, bn the bias, and σ the Rectilinear Unit activation process. The output of each convolutional layer is subsequently pooled to reduce dimensionality, enhancing the model’s focus on prominent features. This pooling operation P is expressed via Eq. 2, $$Pn=Max\left(Cn\left(I\right)\right)\dots \left(2\right)$$ After passing through the contracting path, the image representation Pn enters the expansive path of the UNet process. This path includes a series of up-convolutions and concatenations with the corresponding cropped feature map from the contracting paths. The up-convolution is represented via Eq. 3, $$Un=UpConv\left(Pn\right)\dots \left(3\right)$$ Where, UpConv represents the up-convolution operation process. The concatenation process, which merges the up-convolved feature map with the cropped feature map from the contracting path is expressed via Eq. 4, $$Mn=Concat\left(Un,Cc\left(n\right)\right)\dots \left(4\right)$$ Where, Cc ( n ) is the cropped feature map from the contracting path corresponding to the an th layer of the expansive paths. At this juncture, the Attention Mechanism comes into play, refining the feature map Mn by focusing on pertinent areas while suppressing less relevant regions. The Attention Mechanism, A ( Mn ), is formulated via Eq. 5, $$A\left(Mn\right)=SoftMax\left(Wa1*Mn\right)\odot \left(Wa2*Mn\right)\dots \left(5\right)$$ Where, Wa 1 and Wa 2 are the weights of the attention layers, and ⊙ represents element-wise multiplications. The softmax ensures that the attention weights sum up to one, focusing the model’s ‘attention’ on specific regions of the image sets. The refined feature map, ow imbued with focused attention, is then passed through additional convolutional layers in the expansive paths. This process enhances the feature details and is expressed via Eq. 6, $$C{n}^{{\prime }}=\sigma \left(W{n}^{{\prime }}*A\left(Mn\right)+b{n}^{{\prime }}\right)\dots \left(6\right)$$ Finally, the output of the last layer in the expansive path is subjected to a sigmoid activation function to obtain the segmented vessel image S , which is mathematically represented via Eq. 7, $$S=sigmoid\left(Clas{t}^{{\prime }}\right)\dots \left(7\right)$$ The sigmoid function ensures that the output values are ormalized between 0 and 1, which is essential for image segmentation tasks. Thus, the proposed model operates through a series of meticulously designed convolutional operations, pooling, up-convolutions, and concatenations, all under the vigilant guidance of an Attention Mechanism process. This intricate process ensures that the model not only captures the complex patterns in the retinal images but also focuses on critical aspects for precise vessel segmentation, resulting in output that is both accurate and detailed for different scenarios. As shown in Fig. 1.3 , the model’s ability to emphasize key features while suppressing the irrelevant ones makes it an exceptionally powerful tool for the segmentation of retinal vessels for different scenarios. These segmented images are further processed for integrated Projected Gradient Descent with a Bidirectional Long Short-Term Memory network for the classification of sets of segmented images. This fusion proposes an effective and subtle approach in the field of medical image analysis. Being able to leverage the iterative optimization power of PGD and the sequential data processing capability from BiLSTM, this combination makes it very good at complex image classification tasks. S denotes the images represented after segmentation, considered as input to this model process. These segmented images are rich in contextual and spatial information, which is very essential for an accurate classification process arising from the previous UNet with Attention Mechanism stage. First of all, these images will be flattened and normalized to create a suitable input format for the PGD algorithm process. Let F(S) be the flattened image vector, where F: R^{H \times W \times C} \rightarrow R^{n}, with H, W, and C being the dimensions of the segmented image, and n being the length of the flattened vector sets. The PGD algorithm is employed to iteratively refine the image features, enhancing their discriminative properties. The iterative update rule for PGD is estimated via Eq. 8, $$V\left(k+1\right)=ProjS\left(V\left(k\right)+\alpha \nabla L\left(F\left(S\right),V\left(k\right)\right)\right)\dots . \left(8\right)$$ Where, V(k) is the feature vector at iteration k , α is the learning rate, ∇L is the gradient of the loss function L with respect to V(k) , and ProjS is the projection operator that ensures V(k + 1) remains within a feasible set S for different scans. Once the iterative feature refinement via PGD is complete, the resulting feature vectors are fed into the BiLSTM network process. The BiLSTM processes the data in both forward and backward scopes, capturing temporal dependencies in both scopes. The forward and backward hidden states at timestamp t , represented as \(h(t,f)\) and \(h(t,b)\) , are updated via equations 9 & 10, $$h\left(t,f\right)=LSTM\left(h\left(t-1\right),V\left(t\right)\right)\dots \left(9\right)$$ $$h\left(t,b\right)=LSTM\left(h\left(t+1\right),V\left(t\right)\right)\dots \left(10\right)$$ Where, V(t) is the input feature vector at time t , and LSTM represents the LSTM operation process. At the heart of the LSTM unit are several gates and states that work in unison to regulate the flow of information sets. These include the forget gate, input gate, and output gate, along with the cell state and hidden states. For the single LSTM unit at timestamp t , where h(t − 1) is the hidden state from the previous timestamp, and Vt is the current input vector sets. The dynamics of the LSTM unit are governed by the following process, Forget Gate : The forget gate decides what information to discard from the cell states. Wf and bf are the weight matrix and bias for the forget gate, respectively, and σ represents the sigmoid function, which are fused via Eq. 11 to form the forget gate, $$f\left(t\right)=\sigma \left(Wf\cdot \left[h\left(t-1\right),V\left(t\right)\right]+bf\right)\dots \left(11\right)$$ Input Gate : The input gate determines which values will be updated in the cell states. Wi and bi are the weight matrix and bias for the input gate, which is represented via Eq. 12, $$i\left(t\right)=\sigma \left(Wi\cdot \left[h\left(t-1\right),V\left(t\right)\right]+bi\right)\dots \left(12\right)$$ Cell State Candidate : This creates a vector of candidate values that could be added to the cell state, and is represented via Eq. 13, $$C\sim\left(t\right)=\varvec{t}\varvec{a}\varvec{n}\varvec{h}\left(WC\cdot \left[h\left(t-1\right),V\left(t\right)\right]+bC\right)\dots \left(13\right)$$ Cell State Update : The cell state is updated by combining the past state C(t − 1) and the new candidate values, modulated by the forget and input gates via Eq. 14, $$C\left(t\right)=f\left(t\right)*C\left(t-1\right)+i\left(t\right)*C\sim\left(t\right)\dots \left(14\right)$$ Output Gate : The output gate controls the output from the LSTM units. Wo and bo are the weight matrix and bias for the output gate, which is estimated via Eq. 15, $$o\left(t\right)=\sigma \left(Wo\cdot \left[h\left(t-1\right),V\left(t\right)\right]+bo\right)\dots \left(15\right)$$ Hidden State Update : The hidden state for the current timestamp t is obtained by filtering the cell state through the output gate via Eq. 16, $$h\left(t\right)=o\left(t\right)*\varvec{t}\varvec{a}\varvec{n}\varvec{h}\left(C\left(t\right)\right)\dots \left(16\right)$$ Based on this, the model estimates forward and backward hidden states, which are then concatenated to form a comprehensive feature representation for each timestamp via Eq. 17, $$H\left(t\right)=h\left(t,f\right)\oplus h\left(t,b\right)\dots \left(17\right)$$ Where, ⊕ represents concatenation process. The concatenated hidden states Ht are then passed through a fully connected layer followed by a softmax activation to obtain the classification probabilities for each of the image sets. The classification layer is described via Eq. 18, $$y\left(t\right)=softmax\left(Wc*Ht+bc\right)\dots \left(18\right)$$ Where Wc and bc are weights and bias of the classification layer, respectively, and y(t) is the output vector representing classification probabilities. The classification results obtained for each segmented image, thus, would comprise the probability of the image belonging to various DR types. These are then aggregated to form the complete classification results for the whole set of segmented images and provide an overall view of the types of DR represented in both the dataset and the samples. This will also validate the results of the classification by applying Ensemble Deep Learning methodology in an innovative way, which merges Naive Bayes, k-Nearest Neighbors, Support Vector Machine, Logistic Regression, and Multilayer Perceptron for processing the segmented images and blood reports collected. Only very few such ensemble models exist where traditional machine learning algorithms meet advanced deep learning techniques to carve out a pathway for comprehensive and accurate medical diagnostics. The model has two very different inputs: images that are segmented and blood reports. Consider XI to be the set of features that are to be extracted from the segmented images and XB, a set of features derived from blood reports. In more preliminary steps, both types of inputs will be dealt with separately by the model in order to extract specific characteristics from each of the domains. Finally, features are extracted from the segmented images, FI(XI), and from the blood reports, FB(XB), through equations 19 & 20 as shown below, $$FI\left(XI\right)=\sigma \left(Wn*I+bn\right)\dots \left(19\right)$$ $$F\left(B\right)\left[k\right]=\sum _{t=0}^{T-1}B\left[t\right]\cdot {e}^{-\frac{2\pi ikt}{T}}\dots \left(20\right)$$ Where, B represents the blood report signal, where \(B\in RT\) and T represents length of the signals. These feature nextractors leverage deep learning techniques to distill essential features from the data, preparing them for subsequent analysis. The ensemble model then applies multiple classifiers to these features. Each classifier Ci , where i represents Naive Bayes, kNN, SVM, Logistic Regression, or MLP, processes the features and outputs a prediction via Eq. 21, $$Pi=Ci\left(FI\left(XI\right)\oplus FB\left(XB\right)\right)\dots \left(21\right)$$ Where, ⊕ represents the concatenation of the feature sets from the image and blood report domains. The ensemble approach employs a weighted voting system to integrate the outputs from each of the classifiers. The final classification decision, D , is determined by aggregating these predictions, taking into account the weights assigned to each classifier wi via Eq. 22, $$D=argmax\left(\sum _{i}w\left(i\right)\cdot P\left(i\right)\right)\dots \left(22\right)$$ The weights wi are optimized during the training phase to reflect each classifier’s contribution to the overall performance of the ensemble process. To further refine the model's predictions, a feedback loop is incorporated in this process. This loop involves assessing the confidence level of the classification results and, if this confidence is below a certain threshold, the model reiterates this process with adjusted parameters for different operations. Let \(Conf\left(D\right)\) represent the confidence score of the decision which is estimated via Eq. 23, and θ be the threshold levels. $$Conf\left(D\right)=max\left(P1,P2,\dots ,Pn\right)\dots \left(23\right)$$ Where, \(Pc\) represents the probability of class c being the correct classification, and an is the number of classes. For lower confidence classes, the model parameters are tuned via Eq. 24, \(p{i}^{{\prime }}=pi+\varDelta pi\dots \left(24\right)\) Where, p i ′ is the adjusted parameter, and Δ pi is the change applied to the original parameter p i sets. Incorporating this feedback mechanism ensures that the model dynamically adapts to the intricacies of the data, enhancing the accuracy of the results. The final output of the ensemble model is a set of validated classification results, taking into account both the segmented retinal images and the corresponding blood reports. This dual-input approach enables the model to cross validate its findings, ensuring a higher degree of reliability and precision in the classification of medical conditions. Efficiency of this model was evaluated for different scenarios, and compared with existing methods in terms of different performance metrics in the next section of this text. Result Analysis & Comparisons This will be a new frontier in medical imaging, particularly in the detection and segmentation of DR image scans. The model represents the fusion of multimodalities together with attention mechanisms, further strengthened by adversarial training and ensemble deep learning operations. It excelled in adaptability and robustness, able to efficiently deal with diversified and large data with remarkable precision and accuracy. Advanced AI techniques embody the ingenuity of the ERBVS model through attention mechanisms that focus on key features in retinal images, on one hand, and adversarial training, which sharpens its ability to tell between DR types on the other. The ensemble learning part at the core of this model fuses a diversity of deep learning strategies, guaranteeing in-depth learning from complex data patterns. Not only is there a considerable enhancement of the accuracy in DR diagnosis in this model, but it also reduces false negatives and false positives, which are very important in medical diagnostics. Sensitivity, specificity, precision, recall, and reduced processing delay have been some of the key factors in the making of this ERBVS model, which definitely holds promise to be one of the most pioneering solutions for changing the screening and diagnosis of DR forever, hence carving a niche for AI in healthcare scenarios. In this work, an experimental setting has been designed that will carefully evaluate the performance of the proposed ERBVS model in the classification and segmentation of retinal images with respect to the different types of DR. The setting will be comprehensive with different datasets that ensure the robustness of the evaluation. Datasets Used : Indian Diabetic Retinopathy Dataset : Accessed from Kaggle ( https://www.kaggle.com/datasets/aaryapatel98/indian-diabetic-retinopathy-image-dataset ), this dataset includes high-resolution fundus images specifically representing the Indian demographic. It provides a diverse set of images, aiding in assessing the model's performance across different ethnic backgrounds. Deep Diabetic Retinopathy Dataset : Available on Kaggle ( https://www.kaggle.com/c/diabetic-retinopathy-detection ), this dataset consists of a large collection of high-quality retinal images. It is used to evaluate the model's scalability and performance in handling nextensive and varied data. TensorFlow Dataset Samples : Sourced from the TensorFlow Datasets Catalog ( https://www.tensorflow.org/datasets/catalog/diabetic_retinopathy_detection ), this dataset offers a standardized set of images and is utilized to benchmark the model against established datasets in the field. APTOS Dataset Samples : Available via Academic Torrents ( https://academictorrents.com/details/d8653db45e7f111dc2c1b595bdac7ccf695efcfd ), this dataset includes images from the APTOS Blindness Detection competition, providing a range of images with varying degrees of DR severity. AGAR300 Dataset Samples : Accessed from IEEE DataPort ( https://ieee-dataport.org/open-access/diabetic-retinopathy-fundus-image-datasetagar300 ), this dataset comprises 300 high-quality fundus images specifically curated for DR research. It offers a focused dataset for detailed analysis. Experimental Parameters : Image Preprocessing : Images are resized to a standard dimension of 256x256 pixels. Histogram equalization and contrast adjustment are performed to enhance image quality levels. Model Training : The ERBVS model is trained using a 70:30 split for each dataset, with 70% of the data used for training and 30% for validation process. Augmentation Techniques : Data augmentation, including rotations, flips, and zooms, is applied to increase the diversity of the training data samples. This assists in generating larger data samples for training, testing & validation operations. Learning Rate : An initial learning rate of 0.001 is employed, with a decay rate of 0.1 every 10 epochs. Optimizer : Adam optimizer is utilized for its efficiency in handling sparse gradients and adaptive learning rate capabilities. Loss Functions : A combination of cross-entropy and dice loss functions is used to optimize the segmentation and classification tasks. Batch Size : A batch size of 32 is chosen to balance between computational efficiency and model performance levels. Epochs : The model is trained for 50 epochs, providing a balance between sufficient learning time and avoiding overfitting scenarios. Hardware Specifications : Training is conducted on a system equipped with an NVIDIA Tesla P100 GPU, ensuring high computational power for processing large datasets & samples. Metrics Evaluated : Precision, Recall, Accuracy, Specificity, AUC, and MMSE are the primary metrics used to evaluate the model's performance. A comparative analysis is conducted with existing models MSRNet, DPFNet, and EDGAN to benchmark the ERBVS model's effectiveness. For each dataset, specific configurations are adjusted to account for inherent dataset characteristics like image quality, resolution, and distribution of DR types. Based on the experimental setup, Precision, Accuracy, Recall, Latency, AUC, and Specificity ratings are used to assess the process's effectiveness. Equations 25, 26, 27, and 28 were used to determine these parameters for the entire number of NI images, and the results were compared with those from MSR UNet [ 3 ], DPFNet [ 4 ] and Ensembled Deep Convolutional Generative Adversarial Network (EDGAN) [ 23 ], which use similar prediction techniques. $$P=\frac{1}{NI}\sum _{i=1}^{NI}\frac{tp\left(i\right)}{tp\left(i\right)+fn\left(i\right)}\dots \left(25\right)$$ $$A=\frac{1}{NI}\sum _{i=1}^{NI}\frac{tp\left(i\right)+tn\left(i\right)}{\begin{array}{c}tp\left(i\right)+tn\left(i\right)+\\ fp\left(i\right)+fn\left(i\right)\end{array}}\dots \left(26\right)$$ $$R=\frac{1}{NI}\sum _{i=1}^{NI}\frac{tp\left(i\right)}{\begin{array}{c}tp\left(i\right)+tn\left(i\right)+\\ fp\left(i\right)+fn\left(i\right)\end{array}}\dots \left(27\right)$$ $$d=\frac{1}{NI}\sum _{i=1}^{NI}ts(complete, i)-ts(start, i)\dots \left(28\right)$$ Where fp and fn represent the correct and incorrect counts for classifying inputs into incorrect classes, tp and tn represent the number of samples that are correctly classified into a given class and an incorrect class, respectively, and ts(complete) & ts(start) represent the timestamps for when the segmentation and classification processes were finished and started, respectively, for different evaluations. Using this, the Precision for segmentation was estimated w.r.t. Number of Test Samples (NTS), and can be observed from Fig. 3 as follows, For small NTS, like 5k, the performance of ERBVS is way ahead of the rest: it is at 93.18% _precision, against EDGAN at 88.23%, DPFNet at 82.34%, and MSRNet at 75.02%. That means ERBVS does pretty well with small datasets, which usually suffer from the scarcity of data. For instance, as the NTS increases to 8k and 10k, the ERBVS slightly decreases and then increases in precision, thereby indicating its adaptability and learning efficiency even with escalating data complexity. In the middle range of NTS, from 12k through 26k, ERBVS modulated between higher ranges of precision and peaked at 96.99% for 22k NTS, whereas the other models were very variable. This demonstrates ERBVS's more remarkable ability to process larger and more diverse datasets, which is important in real-world applications. Of particular note is the performance of ERBVS at 19k and 22k NTS, where the gap opened against others by it is dramatic, indicative of its robustness to complex image sets. While the NTS continued to increase further up to 29,000 and higher, ERBVS gave a very high precision with an exceptional peak at 99.80% at an NTS of 43,000, outclassing its closest competitor, EDGAN, which showed a maximum of 93.38% at 29,000 NTS. Such high performance can only be attributed to the fact that ERBVS fuses multimodal information using mechanisms of attention and adversarial training to develop one that perfectly differentiates the relevant features from the rest in such wide and diversified datasets. The highest precision obtained for ERBVS lies in the 43k and 48k NTS, both of which arrive at ear-perfect precision scores of 99.80% and 99.40%, respectively. This outstanding performance is a consequence of the fact that it is brisk in ensemble deep learning operations, particularly in complex scenarios or high-accuracy applications. The ERBVS model therefore has better performance compared to the other analyzed NTSs, more especially in big datasets and advanced features, such as attentional mechanisms and ensemble operations that bring about high-precision levels—hence, quite an outstanding model for the segmentation of blood vessels in the retina, especially in scenarios where high accuracy is critical. The greater accuracy of ERBVS not only provides more reliable segmentation but also improves its applicability in the clinical setting where accurate diagnosis is the prime requirement for an effective treatment planning process. Figure 4 : Accuracy levels achieved during the identification of DR classes using such techniques. At the beginning, when the NTS is small at 5k, the ERBVS accuracy is remarkably high, reaching 95.18%, far surpassing EDGAN with 88.23%, DPFNet with 87.34%, and MSRNet with 75.02%. This high accuracy in smaller datasets suggests the superior ability of ERBVS to generalize from limited data, which should be a pertinent factor in medical imaging where not every dataset could be large. While the NTS increased up to 8k and 10k, ERBVS managed to stay quite strong with an observable slight dip at 10k with 86.51%. This could reflect the adaptability of the model and the learning curve with respect to a slightly higher dataset size. Particularly, on 10k NTS, DPFNet showed an accuracy of 87.09%, thus proving that probably these models have different strengths with a view to various sizes of the dataset. At mid-range NTS (12k to 26k), ERBVS generally has very high accuracy, peaking at 19k NTS with 99.75% and another time at 22k NTS with 99.99%. This exceptionally high accuracy, mostly at 22k NTS, further underscores the competence of this model to have more complex and large datasets, probably due to advanced fusion of multiple modalities along with attention mechanisms. At 25k and 26k NTS, though showing reduced accuracy compared to its optimum, ERBVS was beyond the other models. This behavior can be ascribed to the model's sensitivity to growing data variability or higher complexity inherent in larger image sets. As these sets continue to grow further from 29 to 60k NTS, the accuracy graph for ERBVS does not necessarily take on a consistent pattern, peaking significantly at 31k with an accuracy of 94.83 percent and spiking at 60k with an accuracy of 99.50 percent. It could be that this fluctuation in accuracy is a consequence of how the model handles large datasets with more diversified, complex characteristics of data. At any rate, high accuracy is maintained at certain points—for instance, at 60k NTS—a result that showcases the strength and effectiveness of ERBVS in large-scale scenarios. Similar to Fig. 4 , Fig. 5 shows the categorization recall as follows, For small NTS (5k) at an early stage, ERBVS has a very high recall of 92.48%, outperforming other models significantly. It tells that it has exceptionally outstanding ability in identifying positive DR cases, which is of central importance in all cases of disease diagnosis at their early stages. In contrast, when using an 8k NTS, the recall of ERBVS drastically drops to 82.09%, slightly behind MSRNet's 85.37%. Such fluctuation could also allude to the challenges models face in maintaining consistency because of the different dataset size for scenarios. For instance, as the NTS increases, it bounces back with an ERBVS of 85.27% and 86.54%, showing the strength and flexibility in recall rates for different use cases. This recovery in performance may be reflective of its learning prowess and robustness while handling large amounts and increasing levels of complexity in data. On the middle-range area of 14k through 26k NTS, there is an interesting performance interplay. In general, ERBVS held a higher recall compared to EDGAN and MSRNet, peaking remarkably at 16k NTS with 92.92%. That is, ERBVS is extremely accurate at correctly classifying the positive DR cases within a moderately large dataset—another principal factor necessary for use within a clinical setting. There is also another interesting aspect at higher values of NTS: from 29k to 60k, ERBVS's recall demonstrates a very steep uptrend, reaching its peak at 41k (93.84%), approaching nearly perfect recall at 46k NTS with 99.88%. These very good performances on large NTS suggest that sophisticated algorithms, most likely fusion between multiple modalities and attention mechanisms of ERBVS, make them highly capable of processing large and complex datasets & samples. Independently compared against DPFNet and MSRNet across several NTS, in most cases, ERBVS retains most of its competitive advantage, more so on larger datasets. It gives better recall rates under these conditions, thus showing the potential for clinical application in real-world scenarios where large-scale screening of various situations is done. Proposed models are potential for effective usage in clinical diagnosis for DR for the reason that it can attain pretty high proficiency in recall, especially with bigger datasets. This goes to prove the robustness and level of reliability resultant from its ability to keep on picking all positive cases across different dataset sizes, with a special focus on larger datasets. These are very important considerations in the area of medical imaging, where missing a positive case is costly in every sense of the word, lending support to the possible impact of this study on improving the process of patient care and disease management. Similar to Fig. 5 , Fig. 6 shows the categorization delay as follows, Right from the beginning, on a smaller NTS of 5k, it is quite outstanding—the lowest delay, 52.03 ms, was offered by ERBVS, against 63.84 ms offered by MSRNet, 79.64 ms by DPFNet, and 78.09 ms by EDGAN. This faster speed on smaller datasets gives the potential of our ELRBC network for applications that demand faster results. For example, in emergency blinded diagnosis cases or high-throughput screening scenarios, some initial results are required immediately before conducting further analysis with temporal instance sets. It continues to post efficiency with delays of 58.55 ms and 57.49 ms for the 8k and 10k NTS, respectively. Even as the delay slightly increases compared to the 5k NTS, ERBVS still stays ahead of the other models. This consistent performance in processing speed with increasing dataset sizes underlines the robust computational design of the model, probably assisted by advanced ensemble deep learning operations. Within the middle-range NTS from 12k to 26k, ERBVS generally kept a rather efficient processing time, with remarkable performances at 14k, 55.09 ms, and 19k, 52.24 ms. This means that the model could have clinical applications without large delays on larger datasets. However, if NTS is further increased from 29k to 60k, there is still a trivial increase in the average delay for ERBVS. The maximum average delay is 63.95 ms for 22k NTS and 64.49 ms for 46k NTS. Though it is there, the increase consistently stays in front with ERBVS, again proving that it is much more capable of dealing efficiently with computational load, even in large-scale datasets. The trend noted in ERBVS's delay across the different NTS brings into clear perspective how the model balances between its accuracy and efficiency in the sense that, while remaining very competitive at speed, it does not compromise on accuracy or recall, as has already been noted in previous comparisons. This balance is important in real clinical scenarios in which accuracy and time are two most vital components for the success of any case related to the management of disease and effective patient care. Similarly, the MMSE for these evaluations can be observed in Fig. 7 , For instance, using a smaller NTS of 5k, ERBVS always performs in a superior manner with the lowest MMSE of 0.1317, thus showing that it is more accurate in segmenting images correctly compared to MSRNet's 0.1528, DPFNet's 0.1456, and EDGAN's 0.1555. This large degree of accuracy in small datasets is quite imperative, especially in medical imaging, where accurate segmentation for even a few images can be very vital. The ERBVS keeps a low MMSE at 0.1290 and 0.1352 for 8k and 10k, respectively, when the NTS increases, portraying that it can be quite robust in handling slightly higher datasets without significant loss of accuracy in segmentation. This supports evidence of its superior algorithms that might have been improved by its ensemble deep learning with attention mechanisms. At low- to mid-range NTS, ERBVS will generally maintain a lower MMSE than the other models, with slight increases, such as at 14k, 0.1378, and 25k, 0.1483. This could be tied to the adaptability of the model and its response due to the growing complexity and variability of larger image sets. Noticeably, with further increases in NTS from 29k to 60k, ERBVS shows a changing MMSE, peaking at some points; for example, at 50k NTS with 0.1455. However, it always stands competitive and as often as not outperforms the remaining models. It follows that ERBVS was able to maintain segmentation accuracy in large-scale datasets. For instance, the average MMSE for ERBVS demonstrated relative performance across different numbers of time sections that just about explained its better overall performance in terms of segmentation accuracy. While it did exhibit some fluctuation within its MMSE with an increase in NTS, these variations always remained at a relatively low range that underscored the effectiveness of this model in producing accurate segmentations consistently for different scenarios. The AUC levels can be observed from Fig. 7 as follows, First, for a relatively small NTS of 5k, ERBVS has an AUC of 90.72%, against EDGAN at 90.40%, DPFNet at 87.39%, and MSRNet at 86.73%. It can then be interpreted that this is evidence of the effective distinction capability of ERBVS among different DR types, even in smaller datasets, a very important factor in disease early and accurate detection. In the case of 8k and 10k NTS, ERBVS floats around with a competitive AUC of 86.28% and 89.39%, respectively. This clearly shows its adaptability and reliability for classification on any dataset size. At 8k NTS of EDGAN, it went ahead of ERBVS with an AUC of 91.55%; this, however, shows that on some certain characteristics of the dataset, different models would have varied strengths. For ERBVS, as the NTS increases in the mid-range—from 12,000 to 26,000—its AUC performance changes, dipping at 25,000 NTS to 84.46%. This could be how the model responds to an increase in complexity and variability that naturally comes with a greater number of images. At any rate, ERBVS still has fairly competitive AUC values, since it is able to classify effectively in moderately large datasets. Though the NTS further increases to 29k, 60k, ERBVS still shows an enormous improvement in AUC values to its peak at 96.92% with the 31k NTS, while peaking at 96.95% and 97.09% with 34k and 60k NTS, respectively. High AUC values on these large NTS indicate that very large and complex datasets can be handled more deftly by advanced algorithms and ensemble learning techniques of ERBVS—and this has established the far-flung clinical diagnosis application base. Comparing these with DPFNet, MSRNet, and EDGAN for various NTS, one can see that ERBVS generally performs very well and sometimes oscillates. Clearly, its more competitive AUC scores indicate that in a real-world clinical application with large-scale screening and classification requirements in different use cases, it has great potential to actually apply. While, the specificity levels can be observed from Fig. 9 as follows, In this case, at a smaller NTS of 5k, ERBVS has great specificity of 82.03%. This is comparable to EDGAN's 83.11%, which is still higher than that of DPFNet. Often, this result illustrates that ERBVS is strong in classifying negative patients with events, which means that it detects diseases at their early stages to avoid overdiagnosis. By extension, on the enlarged NTS of 8k and 10k, ERBVS gives an upward trend of specificity to about 85.76% and 89.25%, respectively. This increase highlights the growing accuracy of ERBVS in large datasets, which is indicative of the system's robustness and capacity to handle a diversity of data without miss-classifying healthy cases as diseased. In mid-range NTS from 12k to 26k, ERBVS' specificity changes but still remains very competitive. Especially for the 16k NTS, it has a very outstanding specificity of 95.47%, showing its great ability to avoid false positives in moderately large datasets. This performance is critical in clinical environments, wherein false positives may be dear. As NTS is increased even further, from 29,000 to 60,000, ERBVS holds up extremely well, peaking quite notably at 36,000 with an almost perfect score of 99.86%. The deviations from the observed high performance in this unusual performance at larger NTS are absolutely riveting, suggesting that ERBVS does truly have advanced algorithms, probably with attention mechanisms and ensemble deep learning operations working in large-scale and complex datasets and samples. Comparing these results against DPFNet, MSRNet, EDGAN, across a range of NTS, obviously, in most cases, ERBVS usually retains this heady performance when it comes to specificity. Its potential usefulness in wide clinical diagnostic applications is underlined, call it an ability to recognize on-DR cases correctly across all these consistently, specifically in large datasets. This means that strong performance in specificity across different dataset sizes is expressed through ERBVS, making it a potentially very useful tool in clinical diagnostics for DR. Its reliability and consistency for identifying negative cases, particularly in large datasets, puts it as a potential high-impact model in improving the care of these patients by reducing the risk of unnecessary treatments or interventions due to false-positive diagnoses. Conclusion and Future Scopes The paper contributes much in the field of medical imaging, particularly for the diagnosis and management of diabetic retinopathy. In the paper, a new model called ERBVS was proposed that delivered superior performance in parameters such as precision, recall, accuracy, specificity, AUC, MMSE, and processing delay against already developed models like MSRNet, DPFNet, and EDGAN using advanced techniques of Artificial Intelligence. Key findings included that this model showed an exceptional result on different dataset sizes, thus its adaptability and robustness in segmentation accuracy and classification efficacy. Another part of this, element of scalability, is the prominent performance of the ERBVS in larger datasets. This work also found high precision and recall rates to help in ensuring the accurate identification of DR, which helps reduce the number of false negatives that are very important in preventing delayed or missed diagnosis. The specificity analysis further consolidated the fact that ERBVS is competent in accurately differentiating on-DR cases, which reduces false positives and excess associated interventions. Moreover, this reduction in processing delay points toward the potential of the model in delivering timely diagnostics, a very critical factor in clinical decision-making process. Impact of This Work These better results in the detection and segmentation of DR using the ERBVS model show several important ramifications: 1. Higher Diagnostic Accuracy: Precise diagnosis of DR could mean a much better outcome for patients. The key to effective management of the disease is essentially hinged on early and accurate detection. 2. Less Burden of Health Care System: ERBVS shares a significant role in reducing false positives and false negatives, hence adding up to the efficacy of the health care system by fewer unwarranted treatments and repeated follow-up checks. 3. Mass Screening Scalability: The proficiency of the model in handling bulk data makes it an ideal candidate for mass screening programs, especially needed in countries where diabetes prevalence is high. 4. Across Diverse Populations Customization: Due to its tested effectiveness on a wide range of datasets and its performance on datasets specific to Indian demographics, ERBVS shows promise for customization and application across different ethnic groups. Future Scope There is much scope for future research and development in several ways as envisaged herein: Ophthalmic Devices Integration: ERBVS could be directly integrated into ophthalmic imaging devices to ease diagnosis by providing real-time analysis during patient examinations. Other Ocular Diseases Detection: This model will have wider applications in its ability to detect other ocular conditions for it to become one of the umbrella diagnostic tools of ophthalmology. ERBVS in Telemedicine Applications: This could become very important in telemedicine, particularly in the remote or underserved parts of the country where access to ophthalmologists is limited. • Individualized Treatment Plans: Later versions may incorporate predictive analytics that point toward therapy tailored according to the severity and progression of DR. • Interdisciplinary Applications: Tapping into the application potential off this model in other areas of medical imaging opens new frontiers in AI-driven diagnostics. In a nutshell, ERBVS pushes the bounds on how AI could really revolutionize medical diagnostics. Its development not only points to improving solutions in the present for DR but also opens doorways to wider applications in healthcare, contributing to more efficient, accurate, and accessible medical services. Declarations Acknowledgment Competing Interests: The author declares no competing interests. Authors' Contribution Statement: The sole author conceived, conducted, and analyzed the research, and drafted the manuscript independently in clinical scenarios. Ethical and Informed Consent for Data Used: Plagiarism has been strictly avoided and proper acknowledgment has been given to all sources used. Any errors identified will be promptly corrected ensuring the integrity of the work. Data Availability and Access : The dataset used in this paper is openly available References S. Deari, I. Oksuz and S. Ulukaya, "Block Attention and Switchable Normalization Based Deep Learning Framework for Segmentation of Retinal Vessels," in IEEE Access, vol. 11, pp. 38263-38274, 2023, doi: 10.1109/ACCESS.2023.3265729. Y. Li et al., "Global Transformer and Dual Local Attention Network via Deep-Shallow Hierarchical Feature Fusion for Retinal Vessel Segmentation," in IEEE Transactions on Cybernetics, vol. 53, o. 9, pp. 5826-5839, Sept. 2023, doi: 10.1109/TCYB.2022.3194099. G. B. Kande et al., "MSR U-Net: An Improved U-Net Model for Retinal Blood Vessel Segmentation," in IEEE Access, vol. 12, pp. 534-551, 2024, doi: 10.1109/ACCESS.2023.3347196. J. Li, G. Gao, L. Yang, G. Bian and Y. Liu, "DPF-Net: A Dual-Path Progressive Fusion Network for Retinal Vessel Segmentation," in IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1-17, 2023, Art o. 2517817, doi: 10.1109/TIM.2023.3277946. K. Aurangzeb, R. S. Alharthi, S. I. Haider and M. Alhussein, "An Efficient and Light Weight Deep Learning Model for Accurate Retinal Vessels Segmentation," in IEEE Access, vol. 11, pp. 23107-23118, 2023, doi: 10.1109/ACCESS.2022.3217782. M. Arsalan, T. M. Khan, S. S. Naqvi, M. Nawaz and I. Razzak, "Prompt Deep Light-Weight Vessel Segmentation Network (PLVS-Net)," in IEEE/ACM Transactions on Computational Biology and Bioinformatics, vol. 20, o. 2, pp. 1363-1371, 1 March-April 2023, doi: 10.1109/TCBB.2022.3211936. N. Shen, T. Xu, S. Huang, F. Mu and J. Li, "Expert-Guided Knowledge Distillation for Semi-Supervised Vessel Segmentation," in IEEE Journal of Biomedical and Health Informatics, vol. 27, o. 11, pp. 5542-5553, Nov. 2023, doi: 10.1109/JBHI.2023.3312338. K. Aurangzeb, R. S. Alharthi, S. I. Haider and M. Alhussein, "Systematic Development of AI-Enabled Diagnostic Systems for Glaucoma and Diabetic Retinopathy," in IEEE Access, vol. 11, pp. 105069-105081, 2023, doi: 10.1109/ACCESS.2023.3317348. M. Monemian and H. Rabbani, "A Computationally Efficient Red-Lesion Extraction Method for Retinal Fundus Images," in IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1-13, 2023, Art o. 5001613, doi: 10.1109/TIM.2022.3229712. N. Shen et al., "SCANet: A Unified Semi-Supervised Learning Framework for Vessel Segmentation," in IEEE Transactions on Medical Imaging, vol. 42, o. 9, pp. 2476-2489, Sept. 2023, doi: 10.1109/TMI.2022.3193150. X. Xu et al., "AV-casNet: Fully Automatic Arteriole-Venule Segmentation and Differentiation in OCT Angiography," in IEEE Transactions on Medical Imaging, vol. 42, o. 2, pp. 481-492, Feb. 2023, doi: 10.1109/TMI.2022.3214291. A. -R. M. Al-Hinnawi, A. BaniMustafa, M. Al-Latayfeh and M. Tavakoli, "Reconstruction and Visualization of 5μm Sectional Coronal Views for Macula Vasculature in OptoVue OCTA," in IEEE Access, vol. 11, pp. 28280-28293, 2023, doi: 10.1109/ACCESS.2023.3257720. S. Sundar and S. Sumathy, "Classification of Diabetic Retinopathy Disease Levels by Extracting Topological Features Using Graph Neural Networks," in IEEE Access, vol. 11, pp. 51435-51444, 2023, doi: 10.1109/ACCESS.2023.3279393. M. Sadeghpour, A. Arakala, S. A. Davis and K. J. Horadam, "Protection of Sparse Retinal Templates Using Cohort-Based Dissimilarity Vectors," in IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 5, o. 2, pp. 233-243, April 2023, doi: 10.1109/TBIOM.2023.3239866. X. Zhang et al., "T-Net: Hierarchical Pyramid Network for Microaneurysm Detection in Retinal Fundus Image," in IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1-13, 2023, Art o. 5019613, doi: 10.1109/TIM.2023.3286003. S. Pang et al., "Beyond CNNs: Exploiting Further Inherent Symmetries in Medical Image Segmentation," in IEEE Transactions on Cybernetics, vol. 53, o. 11, pp. 6776-6787, Nov. 2023, doi: 10.1109/TCYB.2022.3195447. G. Ali, A. Dastgir, M. W. Iqbal, M. Anwar and M. Faheem, "A Hybrid Convolutional Neural Network Model for Automatic Diabetic Retinopathy Classification From Fundus Images," in IEEE Journal of Translational Engineering in Health and Medicine, vol. 11, pp. 341-350, 2023, doi: 10.1109/JTEHM.2023.3282104. M. A. K. Raiaan et al., "A Lightweight Robust Deep Learning Model Gained High Accuracy in Classifying a Wide Range of Diabetic Retinopathy Images," in IEEE Access, vol. 11, pp. 42361-42388, 2023, doi: 10.1109/ACCESS.2023.3272228. B. N. Jagadesh, M. G. Karthik, D. Siri, S. K. K. Shareef, S. V. Mantena and R. Vatambeti, "Segmentation Using the IC2T Model and Classification of Diabetic Retinopathy Using the Rock Hyrax Swarm-Based Coordination Attention Mechanism," in IEEE Access, vol. 11, pp. 124441-124458, 2023, doi: 10.1109/ACCESS.2023.3330436. M. Soni et al., "IoT-Based Federated Learning Model for Hypertensive Retinopathy Lesions Classification," in IEEE Transactions on Computational Social Systems, vol. 10, o. 4, pp. 1722-1731, Aug. 2023, doi: 10.1109/TCSS.2022.3213507. M. Vadduri and P. Kuppusamy, "Enhancing Ocular Healthcare: Deep Learning-Based Multi-Class Diabetic Eye Disease Segmentation and Classification," in IEEE Access, vol. 11, pp. 137881-137898, 2023, doi: 10.1109/ACCESS.2023.3339574. W. Nazih, A. O. Aseeri, O. Y. Atallah and S. El-Sappagh, "Vision Transformer Model for Predicting the Severity of Diabetic Retinopathy in Fundus Photography-Based Retina Images," in IEEE Access, vol. 11, pp. 117546-117561, 2023, doi: 10.1109/ACCESS.2023.3326528. H. Naz et al., "Ensembled Deep Convolutional Generative Adversarial Network for Grading Imbalanced Diabetic Retinopathy Recognition," in IEEE Access, vol. 11, pp. 120554-120568, 2023, doi: 10.1109/ACCESS.2023.3327900. W. K. Wong, F. H. Juwono and C. Apriono, "Diabetic Retinopathy Detection and Grading: A Transfer Learning Approach Using Simultaneous Parameter Optimization and Feature-Weighted ECOC Ensemble," in IEEE Access, vol. 11, pp. 83004-83016, 2023, doi: 10.1109/ACCESS.2023.3301618. M. Feng, J. Wang, K. Wen and J. Sun, "Grading of Diabetic Retinopathy Images Based on Graph Neural Network," in IEEE Access, vol. 11, pp. 98391-98401, 2023, doi: 10.1109/ACCESS.2023.3312709. M. Siebert, J. Graßhoff and P. Rostalski, "Uncertainty Analysis of Deep Kernel Learning Methods on Diabetic Retinopathy Grading," in IEEE Access, vol. 11, pp. 146173-146184, 2023, doi: 10.1109/ACCESS.2023.3343642. A. Kukkar et al., "Optimizing Deep Learning Model Parameters Using Socially Implemented IoMT Systems for Diabetic Retinopathy Classification Problem," in IEEE Transactions on Computational Social Systems, vol. 10, o. 4, pp. 1654-1665, Aug. 2023, doi: 10.1109/TCSS.2022.3213369. T. Palaniswamy and M. Vellingiri, "Internet of Things and Deep Learning Enabled Diabetic Retinopathy Diagnosis Using Retinal Fundus Images," in IEEE Access, vol. 11, pp. 27590-27601, 2023, doi: 10.1109/ACCESS.2023.3257988. X. Liu and W. Chi, "A Cross-Lesion Attention Network for Accurate Diabetic Retinopathy Grading With Fundus Images," in IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1-12, 2023, Art o. 5029312, doi: 10.1109/TIM.2023.3322497. N. J. Mohan, R. Murugan, T. Goel and P. Roy, "DRFL: Federated Learning in Diabetic Retinopathy Grading Using Fundus Images," in IEEE Transactions on Parallel and Distributed Systems, vol. 34, o. 6, pp. 1789-1801, June 2023, doi: 10.1109/TPDS.2023.3264473. Q. Hou, P. Cao, L. Jia, L. Chen, J. Yang and O. R. Zaiane, "Image Quality Assessment Guided Collaborative Learning of Image Enhancement and Classification for Diabetic Retinopathy Grading," in IEEE Journal of Biomedical and Health Informatics, vol. 27, o. 3, pp. 1455-1466, March 2023, doi: 10.1109/JBHI.2022.3231276. M. Nur-A-Alam, M. M. K. Nasir, M. Ahsan, M. A. Based, J. Haider and S. Palani, "A Faster RCNN-Based Diabetic Retinopathy Detection Method Using Fused Features From Retina Images," in IEEE Access, vol. 11, pp. 124331-124349, 2023, doi: 10.1109/ACCESS.2023.3330104. B. N. Kumar, T. R. Mahesh, G. Geetha and S. Guluwadi, "Redefining Retinal Lesion Segmentation: A Quantum Leap With DL-UNet Enhanced Auto Encoder-Decoder for Fundus Image Analysis," in IEEE Access, vol. 11, pp. 70853-70864, 2023, doi: 10.1109/ACCESS.2023.3294443. P. Zang et al., "Interpretable Diabetic Retinopathy Diagnosis Based on Biomarker Activation Map," in IEEE Transactions on Biomedical Engineering, vol. 71, o. 1, pp. 14-25, Jan. 2024, doi: 10.1109/TBME.2023.3290541. R. Liu et al., "TMM-Nets: Transferred Multi- to Mono-Modal Generation for Lupus Retinopathy Diagnosis," in IEEE Transactions on Medical Imaging, vol. 42, o. 4, pp. 1083-1094, April 2023, doi: 10.1109/TMI.2022.3223683. P. Zhang et al., "Image Quality Assessment of Diabetic Retinopathy Based on ADD-Net," in IEEE Access, vol. 11, pp. 105130-105139, 2023, doi: 10.1109/ACCESS.2023.3318876. K. Radha and Y. Karuna, "Modified Depthwise Parallel Attention UNet for Retinal Vessel Segmentation," in IEEE Access, vol. 11, pp. 102572-102588, 2023, doi: 10.1109/ACCESS.2023.3317176. A. Pereira, C. Santos, M. Aguiar, D. Welfer, M. Dias and M. Ribeiro, "Improved Detection of Fundus Lesions Using YOLOR-CSP Architecture and Slicing Aided Hyper Inference," in IEEE Latin America Transactions, vol. 21, o. 7, pp. 806-813, July 2023, doi: 10.1109/TLA.2023.10244179. M. Hussain, "Exudate Detection: Integrating Retinal-Based Affine Mapping and Design Flow Mechanism to Develop Lightweight Architectures," in IEEE Access, vol. 11, pp. 125185-125203, 2023, doi: 10.1109/ACCESS.2023.3328386. M. Yin et al., "Dual-Branch U-Net Architecture for Retinal Lesions Segmentation on Fundus Image," in IEEE Access, vol. 11, pp. 130451-130465, 2023, doi: 10.1109/ACCESS.2023.3333364. D. Bar-David et al., "Elastic Deformation of Optical Coherence Tomography Images of Diabetic Macular Edema for Deep-Learning Models Training: How Far to Go?," in IEEE Journal of Translational Engineering in Health and Medicine, vol. 11, pp. 487-494, 2023, doi: 10.1109/JTEHM.2023.3294904. P. R. Bhargav and N. B. Puhan, "Novel Contraharmonic Correlative Attention Loss for Microaneurysm Segmentation in Fundus Images," in IEEE Sensors Letters, vol. 7, o. 7, pp. 1-4, July 2023, Art o. 7003504, doi: 10.1109/LSENS.2023.3290597. Abdallah, M.B., Azar, A.T., Guedri, H. e t al. Correction to: Noise-estimation-based isotropic diffusion approach for retinal blood vessel segmentation. Neural Comput & Applic 34 , 2499 (2022). https://doi.org/10.1007/s00521-021-06819-5 Singh, L.K., Khanna, M., Thawkar, S. e t al. Deep-learning based system for effective and automatic blood vessel segmentation from Retinal fundus images. Multimed Tools Appl (2023). https://doi.org/10.1007/s11042-023-15348-3 Kumar, K.S., Singh, N.P. Retinal disease prediction through blood vessel segmentation and classification using ensemble-based deep learning approaches. Neural Comput & Applic 35 , 12495–12511 (2023). https://doi.org/10.1007/s00521-023-08402-6 Upadhyay, K., Agrawal, M. & Vashist, P. Learning multi-scale deep fusion for retinal blood vessel nextraction in fundus images. Vis Comput 39 , 4445–4457 (2023). https://doi.org/10.1007/s00371-022-02600-4 Kumar, K.S., Singh, N.P. Segmentation of retinal blood vessel using generalized nextreme value probability distribution function(pdf)-based matched filter approach. Pattern Anal Applic 26 , 307–332 (2023). https://doi.org/10.1007/s10044-022-01108-w Mahapatra, S., Jena, U.R. & Dash, S. Mean global based on hysteresis thresholding for retinal blood vessel segmentation using enhanced homomorphic filtering. Multimed Tools Appl 81 , 41911–41928 (2022). https://doi.org/10.1007/s11042-022-13517-4 Saranya, P., Prabakaran, S., Kumar, R. e t al. Blood vessel segmentation in retinal fundus images for proliferative diabetic retinopathy screening using deep learning. Vis Comput 38 , 977–992 (2022). https://doi.org/10.1007/s00371-021-02062-0 Gu, J., Tian, F. & Oh, IS. Retinal vessel segmentation based on self-distillation and implicit neural representation. Appl Intell 53 , 15027–15044 (2023). https://doi.org/10.1007/s10489-022-04252-2 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4553609","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":329214104,"identity":"526100df-f544-4ffa-9b99-de4378a6f873","order_by":0,"name":"Komal Umare Thool","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+UlEQVRIiWNgGAWjYBACCRDxoMIGwvsAxGzsxGhJOJMG5jDOAGlhJkZLYtthMIeZB0wS0CLZ3p34IYHtcOLa9rMPH9v82ibPx8zA+OFjDm4t0jxnN0sk8KQnbjuTbmyc23fbsI2ZgVly5jbcWuQkcjdIJEhYJ247kMYmndtzmxGohY2ZF7+WzT8SDJgTt51/xv7bsue2PUEt0hK52yQSEpwTt91IA4bVj9uJBLVI9pzdZpFwIM14241nzJK9DbeT25gZm/H6ReJ47+YbH//ZyG47n8b44cef27bz25sPfviIRwsqYGwDkw3EqgeBP6QoHgWjYBSMgpECAFjTU7II10YnAAAAAElFTkSuQmCC","orcid":"","institution":"Ramdeobaba University Nagpur","correspondingAuthor":true,"prefix":"","firstName":"Komal","middleName":"Umare","lastName":"Thool","suffix":""}],"badges":[],"createdAt":"2024-06-09 11:38:22","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4553609/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4553609/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":61181972,"identity":"5b9cd208-e6e2-4155-8b66-743822fb408b","added_by":"auto","created_at":"2024-07-26 16:52:56","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":147409,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eDesign of the proposed model for segmentation \u0026amp; classification operations\u003c/em\u003e\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/bf7f8a0de60e4f0f56e00fe2.png"},{"id":61183014,"identity":"f8857cae-673c-44d2-ab89-886e8e367db5","added_by":"auto","created_at":"2024-07-26 17:00:56","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":103865,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFigure 1.2. Overall flow of the proposed segmentation \u0026amp; classification process\u003c/em\u003e\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/d2fdf8a1270c15e81c4171f7.png"},{"id":61183016,"identity":"664b2fb1-91ae-48b9-83e2-89b1ccf3fd9f","added_by":"auto","created_at":"2024-07-26 17:00:56","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":527252,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFigure 1.3. Flow of the segmentation process (a. Input Images, b. Ground Truth, c. CNN Segmented Image, d. Coarse Segmentation, e. Fine Segmentation, f. Final Segmentation Results)\u003c/em\u003e\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/5d6e5f1ca0bbd4cb51070193.png"},{"id":61181969,"identity":"9d583b7a-0457-4840-af79-ae5d8122fe22","added_by":"auto","created_at":"2024-07-26 16:52:56","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":49420,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFigure 3. Precision Levels for Different Number of Image Sets\u003c/em\u003e\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/8a25e806d5b057710470613d.png"},{"id":61181970,"identity":"54315e5d-9763-4514-bf63-ba83b337575a","added_by":"auto","created_at":"2024-07-26 16:52:56","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":48857,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFigure 4. Accuracy for classification of image samples into DR Types\u003c/em\u003e\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/2bfd12a88da960a42ce15667.png"},{"id":61181973,"identity":"60092d0f-0bda-4c14-8a95-7e796ab69a07","added_by":"auto","created_at":"2024-07-26 16:52:56","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":14048,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFigure 5. Recall for classification of image samples into DR Types\u003c/em\u003e\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/3ffcacdc0db0331505f2405f.png"},{"id":61181976,"identity":"b77c9c80-81fb-48e7-8e00-d1206028bd24","added_by":"auto","created_at":"2024-07-26 16:52:56","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":14680,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFigure 6. Delay for classification of image samples into DR Types\u003c/em\u003e\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/e7f313178451fc03c32ab3fa.png"},{"id":61183015,"identity":"2077162e-aae9-4672-ba22-e32a97240042","added_by":"auto","created_at":"2024-07-26 17:00:56","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":53889,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFigure 7. MMSE to segment the image samples\u003c/em\u003e\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/88018947b211f18d542fa29a.png"},{"id":61181978,"identity":"789e06e9-7b85-4488-87b4-e53efd7d5c63","added_by":"auto","created_at":"2024-07-26 16:52:56","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":11124,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFigure 8. AUC for classification of image samples into DR Types\u003c/em\u003e\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/dbe6233c8f065b13120501a2.png"},{"id":61181974,"identity":"b4a0328b-7300-407a-9cb1-a982da02fcb7","added_by":"auto","created_at":"2024-07-26 16:52:56","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":55014,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFigure 9. Specificity for classification of image samples into DR Types\u003c/em\u003e\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/75076a99d28faf1a9c4ad410.png"},{"id":65711602,"identity":"d1a2ec67-79aa-4c29-9d71-677f720974e0","added_by":"auto","created_at":"2024-10-01 14:47:14","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1546308,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4553609/v1/4db01be2-9dbd-43db-9bf7-1e8f39c66ae9.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"ERBVS: Enhanced Retinal Blood Vessel Segmentation using Multiple Modalities and Attention Mechanisms with Adversarial Training and Ensemble Deep Learning Operations","fulltext":[{"header":"Introduction","content":"\u003cp\u003eModern medicine rests on the concept of early and correct diagnosis. Correct diagnosis, if made on time, can greatly decrease morbidity and mortality by preventing the development of a complication. Yet, despite this reality and the enormous progress in science and technology and research, the penalty for delayed diagnosis and lost therapeutic windows is still borne by humanity. In these cases, predictive tools would be relevant and possible; therefore, progressive and accurate models are realized to inform timely clinical decisions. Digital health innovation has opened the vast potential of \u0026apos;medical data\u0026apos; toward the creation of an era of truly personalized medicine if it can be correctly analyzed and interpreted. The human health is so intricate and dynamic that it requires a predictive model, not only attuned to the subtleties in disease progression but one which also recognizes subtle cues indicative of possible health anomalies in the future. Traditional methods, having served as essential foundational blocks, fall short of catering to the evolving demands of modern medicine. Their limitations are multiplied in situations where real-time decisions are imperatively needed, and even the smallest delays could add up to finish in less-than-desired patient results. Real-time impact of a sophisticated predictive model is multifaceted. Clinically, it ensures timely interventions to patients, thereby significantly reducing the progression to severe disease stages. It economically reduces the exorbitant costs associated with late-stage treatments and long hospital stays. Such robust predictive tools can also be used for epidemic control, resource allocation, and formulation of policies from a global health perspective. Here, one can see what value an improved model brings to the table: driving a healthcare revolution in delivery and outcomes regarding betterment of health for patients, and helping medical professionals in the field to always stay one step ahead with their therapeutic interventions. In these regards, a model that corrects the intrinsic deficiencies of the methods at hand, by rendering flexibility and accuracy under highly variable clinical settings, is overdue. In this work, a proposed game-changing tool in such diseases characterization and pre-emption is offered for the greater healthcare environment.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMotivation \u0026amp; Objectives\u003c/strong\u003e:\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMotivation\u003c/strong\u003e:\u003c/p\u003e\n\u003cp\u003eThe more complex and heterogeneous diseases are, the more robust medical models are required to accurately locate their beginnings and progress. The dominant predictive models, instrumental in many ways, display certain lacunae that may prove very dangerous in a clinical setting. Mistimed diagnoses might miss out on treatment opportunities, while false alarms may initiate therapies that are unnecessary, both of which mean undue physical, emotional, and financial pressure on patients. This pressure to upgrade predictive tools comes with an added force\u0026mdash;the medical community is moving toward an approach of personalized medicine. With each patient carrying a unique genetic makeup and life history, he embodies an individual medical enigma, and the tools which we implement on the inputs from these patients have to be agile while resolving these complex tasks in real-time scenarios. This many-fold dilemma gets further increased by the global health scenarios. The pandemics of the contemporary era have shown that not only individual early diagnosis and prevention act at the community, national, and global levels, but it also works vice versa process. The broader social and economic implications of outbreaks and the evident need for faster and more accurate predictive tools raise the motivation for this study process.\u003c/p\u003e\n\u003cp\u003eObjectives:\u003c/p\u003e\n\u003cp\u003e\u0026bull; Improvement in Accuracy: The model will be improved for better prediction accuracy, so as to identify diseases with high specificity, reducing false positives.\u003c/p\u003e\n\u003cp\u003e\u0026bull; Improved Classification and Preemption: Strengthen model skills toward the correct classification of diseases and the preemption of possible health complications, ensuring timely clinical interventions for different scenarios.\u003c/p\u003e\n\u003cp\u003e\u0026bull;Reduced Delays: Since time is of essence in a medical environment, this becomes one of the major objectives\u0026mdash;to reduce time lags in prediction and offer quicker insight to healthcare professionals.\u003c/p\u003e\n\u003cp\u003e\u0026bull; Across clinical scenarios, adaptability: The model has to be versatile and handle all kinds of situations viewed against the wide spectrum of diseases and their infinite manifestations.\u003c/p\u003e\n\u003cp\u003e\u0026bull; Real-time implementation: Not only should it be capable of prediction, but it is also intended to be integrated into a clinical setup in real-time. That means healthcare professionals use this insight instantly for various use cases.\u003c/p\u003e\n\u003cp\u003eBriefly, it will be an effort with one of the finest human aspirations\u0026mdash;bridging gaps in disease prediction and preemption. These objectives stand before us, like guiding stars, ensuring that the outcomes really make a difference amidst the pulsing needs of modern medical landscapes.\u003c/p\u003e"},{"header":"Review of existing models used for identification of Diabetic Retinopathy","content":"\u003cp\u003eSegmentation of the retinal vessels, especially in rapidly changing medical imaging, has seen incredible leaps with the incorporation of advanced deep learning models. The literature shows concentrated effort towards improving the accuracy and efficiency of these models, evidenced by a lot of works on different issues related to retinal vessel segmentation. Deari et al. [\u003cspan\u003e1\u003c/span\u003e] went a step further into providing a deep learning framework that reviews block attention and switchable normalization in ensuring an improved segmentation process. The new technique in this case is radically different from the previous ones; hence, it provides more accurate and finer results. Again, Li et al. [\u003cspan\u003e2\u003c/span\u003e] came up with a globally \u0026ndash; transformer and dual local-attention network. The model applied a deep\u0026ndash;shallow hierarchical feature fusion approach in this case. This approach underlines the trend toward the application of complex architectures that could handle the intrinsic complexity of retotal vessel structures. Moreover, innovations are not only limited to the introduction of new network architectures. Kande et al. [\u003cspan\u003e3\u003c/span\u003e] proposed another variant of the U-Net model, called MSR U-Net, toward the same application. In this model, it is an extension of a broadly accepted standard in medical image segmentation, and it is such a proposal that proves the continuous development and optimization of already existing models. Concurrently, Li et al. [\u003cspan\u003e4\u003c/span\u003e] developed the dual-path progressive fusion network, DPF-Net, which gave a novel perspective in terms of the integration of multiple pathways to facilitate improved segmentation. On the other side of those very complex models, Aurangzeb et al. [\u003cspan\u003e5\u003c/span\u003e] and Arsalan et al. [\u003cspan\u003e6\u003c/span\u003e] develop an efficient and lightweight model for retinal vessel segmentation, with critical needs in computational efficiency, especially when resources are limited. This approach is more relevant to medical settings, where the processing of retinal images requires quickness and reliability. At the same time, semi-supervised learning paradigms are considered. For example, Shen et al. [\u003cspan\u003e7\u003c/span\u003e] have proposed an expert-guided knowledge distillation for vessel segmentation in a semi-supervised setting. Since the model proposed by them is close to being on one side point of view with full supervision and unsupervised on the other, it may be said to be positioned between these two major approaches. For example, the applications of AI in diagnostic systems, such as on glaucoma and diabetic retinopathy, have been systematically developed. Aurangzeb et al. [\u003cspan\u003e8\u003c/span\u003e] describe this holistic view of AI-enabled diagnostics, showing how advances in retinal vessel segmentation have very far-reaching implications. Further, Monemian and Rabbani [\u003cspan\u003e9\u003c/span\u003e] presented a computationally efficient method for extracting red lesions in retinal fundus images, hence showing the necessity for being concerned with only some of the features in retinal imaging. Shen et al. [\u003cspan\u003e10\u003c/span\u003e] and Xu et al. [\u003cspan\u003e11\u003c/span\u003e] further contributed with unified semi-supervised learning frameworks and automatic arteriole-venule segmentation methods, respectively. Other important aspects of retinal imaging that have been done to give greater understanding of the structural makeup of the retina are the reconstruction and visualization of macula vasculature investigated by Al-Hinnawi et al. [\u003cspan\u003e12\u003c/span\u003e]. In the same way, Sundar and Sumathy [\u003cspan\u003e13\u003c/span\u003e] worked on the classification of levels of diabetic retinopathy disease underpinning the critical role played by topological features and graph neural networks in medical diagnostics. The protection of sparse retinal templates, addressed by Sadeghpour et al. [\u003cspan\u003e14\u003c/span\u003e], underlines the increasing concern for security and privacy of medical data in the age of AI. This becomes relevant when medicine is increasingly reliant on digital technologies. Another work by Zhang et al. [\u003cspan\u003e15\u003c/span\u003e] was related to the detection of microaneurysms within retinal fundus images; for this purpose, they utilized a hierarchical pyramid network called T-Net to improve the accuracy of detection. Finally, Pang et al. [\u003cspan\u003e16\u003c/span\u003e] went beyond traditional CNNs in searching for more inherent symmetries in medical image segmentation. The present study puts one of the most important developments ongoing in this area in a nutshell that continues with continuous exploration of novel techniques, methodologies, and approaches to enhance medical image analysis with respect to accuracy, efficiency, etc. Ali et al. [\u003cspan\u003e17\u003c/span\u003e] proposed a hybrid convolutional neural network model for the automatic classification of diabetic retinopathy using fundus images. This model epitomizes how traditional image processing can be combined with deep learning to ensure an improved level of accuracy and efficiency. Again, Raiaan et al. [\u003cspan\u003e18\u003c/span\u003e] made contributions toward a lightweight, effective deep learning model representing a high-accuracy approach underlining the Spanner ability of the model in handling most images of diabetic retinopathy. The technique stands very special in balancing computational efficiency and performance. Jagadesh et al. [\u003cspan\u003e19\u003c/span\u003e] suggested a framework that integrated the IC2T model for segmentation with Rock Hyrax swarm-based coordination attention mechanism for classification. Their approach has been able to depict the potentials of using bio-inspired algorithms in the enhancement of medical image analysis with improved accuracy. Soni et al. [\u003cspan\u003e20\u003c/span\u003e], on the other hand, applied an IoT-based federated learning approach for the classification of hypertensive retinopathy lesions, which opened new avenues toward applying distributed learning systems within the medical domain. Vadduri and Kuppusamy [\u003cspan\u003e21\u003c/span\u003e] focused on the improvement of ocular healthcare by proposing a deep learning-based multi-class diabetically eye disease segmentation and classification approach. Their work decidedly brings to the fore the growing importance of deep learning in the different medical imaging tasks. Nazih et al. [\u003cspan\u003e22\u003c/span\u003e] applied the vision transformer model for severity prediction in diabetic retinopathy against fundus photography-based retina images, which brings out the increasing interest in transformer models in medical imaging. Naz et al. [\u003cspan\u003e23\u003c/span\u003e] proposed an ensembled deep convolutional generative adversarial network for grading imbalanced diabetic retinopathy recognition. The authors have tried to surmount the problem of class imbalance generally observed in most of the medical image datasets. Wong, Juwono, and Apriono [\u003cspan\u003e24\u003c/span\u003e] used transfer learning with optimization of parameters simultaneously along with a feature-weighted ECOC ensemble for diabetic retinopathy detection and grading. Feng et al. [\u003cspan\u003e25\u003c/span\u003e] dealt with the grading of diabetic retinopathy images using a graph neural network, a novel approach that brings out the importance of topological features in the context of medical imaging. Siebert, Gra\u0026szlig;hoff, and Rostalski [\u003cspan\u003e26\u003c/span\u003e] examined uncertainty analysis of deep kernel learning methods for diabetic retinopathy grading, placing much importance on the reliability and confidence levels expected of medical diagnoses. Kukkar et al. [\u003cspan\u003e27\u003c/span\u003e] have used IoMT systems socially implemented for the optimization of deep learning model parameters for the diabetic retinopathy classification problem. Very much, the study draws a gap between social systems and medical imaging, reflecting interdisciplinarity in modern research. Palaniswamy and Vellingiri [\u003cspan\u003e28\u003c/span\u003e] have shown that active internet of things and deep learning techniques could be applied to diagnose Diabetic Retinopathy using retinal fundus images, proving the use of technology in health. Liu and Chi [\u003cspan\u003e29\u003c/span\u003e] proposed a cross-lesion attention network for the accurate grading of diabetic retinopathy using fundus images, with regard to lesion type specificity. Mohan et al. [\u003cspan\u003e30\u003c/span\u003e] proposed DRFL, a federated learning framework for grading diabetic retinopathy using fundus images; it lays much emphasis on collaborative learning in a distributed environment. Hou et al. [\u003cspan\u003e31\u003c/span\u003e] proposed an image quality assessment-driven collaborative learning framework for image enhancement and classification toward diabetic retinopathy grading. That meant the image quality would impact the accuracy of diagnosis in the study. Nur-A-Alam et al. [\u003cspan\u003e32\u003c/span\u003e] proposed a faster RCNN-based detection based on fused features from retina images, which is evidence for the ongoing improvements in object detection models under medical imaging. Afterward, Kumar et al. [\u003cspan\u003e33\u003c/span\u003e] utilized a DL-UNet empowered auto-encoder-decoder for fundus image analysis, redefining retinal lesion segmentation and demonstrating the flexibility of neural networks in complex segmentation tasks. Zang et al. [\u003cspan\u003e34\u003c/span\u003e] worked on interpretable diabetic retinopathy diagnosis by means of biomarker activation maps, which interestingly combined diagnostic accuracy with the requirement for interpretability within AI systems. Lu et al. [\u003cspan\u003e35\u003c/span\u003e] presented a new perspective toward transferred multi- to mono-modal generation for the diagnosis of lupus retinopathy, which was an emerging interest in cross-modal medical image analysis. At last, Zhang et al. [\u003cspan\u003e36\u003c/span\u003e] investigated ADD-Net-based image quality assessment of diabetic retinopathy. This sets forth the critical image quality in the effectiveness of deep learning models. Radha and Karuna [\u003cspan\u003e37\u003c/span\u003e] presented a depthwise parallel attention UNet for segmentation of the retinal vessels, which proved the efficiency of attention mechanisms in enhancing the accuracy of segmentation models. The innovation is one more step to develop deep learning models specified for medical imaging. Pereira et al. [\u003cspan\u003e38\u003c/span\u003e] worked on the effective detection of fundus lesions using YOLOR-CSP architecture and slicing aided hyper inference. Their methodology couples state-of-the-art deep learning architectures with advanced inference strategies and represents the continuous change of methodologies in this area. Hussain [\u003cspan\u003e39\u003c/span\u003e] investigated exudate detection through the integration of retinal-based affine mapping with the design flow mechanism, contributing to lightweight architecture designs. This research answers a pressing need: that efficient and resource-light models for medical imaging are provided, especially in settings constrained by computational resources. Yin et al. [\u003cspan\u003e40\u003c/span\u003e] presented a dual-branch U-Net architecture for the segmentation of retinal lesions in fundus images. This further underlined the flexibility of the U-Net architecture across a wide range of medical imaging applications. Bar-David et al. [\u003cspan\u003e41\u003c/span\u003e] suggested elastic deformation of OCT images from diabetic macular edema patients for training deep learning models. Their work is focused on the preprocessing aspects of medical images and places much attention on how high the influence of image manipulation is during the training and results of deep learning models. Bhargav and Puhan [\u003cspan\u003e42\u003c/span\u003e] proposed a new contra-harmonic correlative attention loss for the segmentation of microaneurysms in fundus images, and this majorly brings out the part of the loss functions playing in the optimization involved in the segmentation tasks. Abdallah et al. [\u003cspan\u003e43\u003c/span\u003e] proposed a noise-estimation-based isotropic diffusion approach that corrects the methods of noise reduction, quite essential in clear and accurate analysis for the segmentation of retinal blood vessels. Singh et al. [\u003cspan\u003e44\u003c/span\u003e] proposed a deep-learning-based system which has proved to be efficient and automatic in blood vessel segmentation from retinal fundus images, hence showing the trend towards the use of deep learning in automating the complex tasks of medical image segmentation. Kumar and Singh [\u003cspan\u003e45\u003c/span\u003e] worked on the prediction of diseases in the retina by segmenting blood vessels and classifying them using ensemble-based deep learning approaches. This work helped demonstrate how segmentation acts at the joint of disease prediction, further powered by ensemble learning for high diagnostic accuracy. Upadhyay et al. [\u003cspan\u003e46\u003c/span\u003e] investigated learning multi-scale deep fusion for the extraction of retinal blood vessels in fundus images and proved that multi-scale approaches can become very efficient in the characterization of a wide range of vessel details. Kumar and Singh [\u003cspan\u003e47\u003c/span\u003e] proposed a segmentation approach based on the generalized nextreme value probability distribution function-based matched filter, allowing the interface of statistical methods with deep learning to improve segmentation. Mahapatra et al. [\u003cspan\u003e48\u003c/span\u003e] proposed a mean global based on hysteresis thresholding for the segmentation of retinal blood vessels using enhanced homomorphic filtering. This is perhaps one of the traditional image processing methods that extracts the most essential elements from the new thresholding techniques. Saranya et al. [\u003cspan\u003e49\u003c/span\u003e] researched blood vessel segmentation in retinal fundus images for proliferative diabetic retinopathy screening, calling attention to its application in specific disease screening. Gu, Tian, and Oh [\u003cspan\u003e50\u003c/span\u003e] contributed self-distillation and implicit neural representation-based retinal vessel segmentation. This is a very original technique that uses advanced neural network training techniques with implicit representations.\u003c/p\u003e\n\u003cp\u003eBased on this review, it can be observed that several challenges consistently emerge across the different models:\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eData Diversity\u003c/strong\u003e: The vast and varied nature of medical data poses integration and processing challenges.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eScalability\u003c/strong\u003e: Many models show robust results in controlled or pilot studies but falter when scaled to larger, diverse patient populations.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eInterpretability\u003c/strong\u003e: The need for models that are both high-performing and interpretable remains a significant concern, especially in clinical settings.\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eAs the volume and variety of medical data grow, the drive to develop an optimal predictive model remains at the forefront of medical research process. This work aims to synthesize the learnings from past models and propose advancements for the future scopes.\u003c/p\u003e\n\u003cp\u003e\u003cspan\u003e\u003cstrong\u003e3. Proposed design of an augmented Bioinspired Multidomain feature nextraction \u0026amp; selection model for Diabetic retinopathy Severity estimation via Ensemble learning process\u003c/strong\u003e\u003cbr\u003e\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003eThis paper proposes multiple threads of advanced technology, including modalities such as specialized convolutional layers that go really deep into the intricacies of retinal imagery scans. Adversarial training empowers this model further by having generative and discriminative networks engage in an intricate dance that pushes the model toward ever higher accuracy and greater robustness. This is complemented by ensemble deep learning operations\u0026mdash;a symphony of various learning techniques harmoniously together decoding the most intricate patterns of DR. Therefore, the amalgamation would catapult the performance of the model in terms of accuracy, precision, and speed and manifest remarkable adaptability to diversity and volume in datasets. ERBVS is not a model with awesome powers of analysis, but it represents the torchbearer of innovation that ushers in a new epoch into the diagnostic capability segment pertaining to healthcare.\u003c/p\u003e\n\u003cp\u003eIn this work, an intricate marriage of the UNet architecture with an attention mechanism is proposed for the segmentation of the retinal vessels. The result will be a sophisticated model delineating not only the intricate vasculature of the retina but underlining salient features necessary for a detailed analysis. First, the images collected regarding the retina are fed into the model process. Let these images be represented by I, where I\u0026isin;R(H\u0026times;W\u0026times;C), where H, W, and C are the image height, image width, and the number of channels respectively. The first stage of the model details passing I through the UNet\u0026apos;s contracting path formulated as a series of convolutional operations. Let \u003cem\u003eCn\u003c/em\u003e(\u003cem\u003eI\u003c/em\u003e) represent the n\u003cem\u003eth\u003c/em\u003e convolutional layer applied to the image, defined via Eq.\u0026nbsp;1,\u003c/p\u003e\n\u003cdiv id=\"Equa\"\u003e\n \u003cdiv id=\"FileID_Equa\" name=\"EquationSource\"\u003e$$Cn\\left(I\\right)=\\sigma \\left(Wn*I+bn\\right)\\dots \\left(1\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, \u0026lowast; represents the convolution operation, \u003cem\u003eWn\u003c/em\u003e represents the weights, \u003cem\u003ebn\u003c/em\u003e the bias, and \u003cem\u003e\u0026sigma;\u003c/em\u003e the Rectilinear Unit activation process. The output of each convolutional layer is subsequently pooled to reduce dimensionality, enhancing the model\u0026rsquo;s focus on prominent features. This pooling operation \u003cem\u003eP\u003c/em\u003e is expressed via Eq.\u0026nbsp;2,\u003c/p\u003e\n\u003cdiv id=\"Equb\"\u003e\n \u003cdiv id=\"FileID_Equb\" name=\"EquationSource\"\u003e$$Pn=Max\\left(Cn\\left(I\\right)\\right)\\dots \\left(2\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eAfter passing through the contracting path, the image representation \u003cem\u003ePn\u003c/em\u003e enters the expansive path of the UNet process. This path includes a series of up-convolutions and concatenations with the corresponding cropped feature map from the contracting paths.\u003c/p\u003e\n\u003cp\u003eThe up-convolution is represented via Eq.\u0026nbsp;3,\u003c/p\u003e\n\u003cdiv id=\"Equc\"\u003e\n \u003cdiv id=\"FileID_Equc\" name=\"EquationSource\"\u003e$$Un=UpConv\\left(Pn\\right)\\dots \\left(3\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, UpConv represents the up-convolution operation process. The concatenation process, which merges the up-convolved feature map with the cropped feature map from the contracting path is expressed via Eq.\u0026nbsp;4,\u003c/p\u003e\n\u003cdiv id=\"Equd\"\u003e\n \u003cdiv id=\"FileID_Equd\" name=\"EquationSource\"\u003e$$Mn=Concat\\left(Un,Cc\\left(n\\right)\\right)\\dots \\left(4\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, \u003cem\u003eCc\u003c/em\u003e(\u003cem\u003en\u003c/em\u003e) is the cropped feature map from the contracting path corresponding to the an\u003cem\u003eth\u003c/em\u003e layer of the expansive paths.\u003c/p\u003e\n\u003cp\u003eAt this juncture, the Attention Mechanism comes into play, refining the feature map \u003cem\u003eMn\u003c/em\u003e by focusing on pertinent areas while suppressing less relevant regions. The Attention Mechanism, \u003cem\u003eA\u003c/em\u003e(\u003cem\u003eMn\u003c/em\u003e), is formulated via Eq.\u0026nbsp;5,\u003c/p\u003e\n\u003cdiv id=\"Eque\"\u003e\n \u003cdiv id=\"FileID_Eque\" name=\"EquationSource\"\u003e$$A\\left(Mn\\right)=SoftMax\\left(Wa1*Mn\\right)\\odot \\left(Wa2*Mn\\right)\\dots \\left(5\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, \u003cem\u003eWa\u003c/em\u003e1 and \u003cem\u003eWa\u003c/em\u003e2 are the weights of the attention layers, and ⊙ represents element-wise multiplications. The softmax ensures that the attention weights sum up to one, focusing the model\u0026rsquo;s \u0026lsquo;attention\u0026rsquo; on specific regions of the image sets. The refined feature map, ow imbued with focused attention, is then passed through additional convolutional layers in the expansive paths. This process enhances the feature details and is expressed via Eq.\u0026nbsp;6,\u003c/p\u003e\n\u003cdiv id=\"Equf\"\u003e\n \u003cdiv id=\"FileID_Equf\" name=\"EquationSource\"\u003e$$C{n}^{{\\prime }}=\\sigma \\left(W{n}^{{\\prime }}*A\\left(Mn\\right)+b{n}^{{\\prime }}\\right)\\dots \\left(6\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eFinally, the output of the last layer in the expansive path is subjected to a sigmoid activation function to obtain the segmented vessel image \u003cem\u003eS\u003c/em\u003e, which is mathematically represented via Eq.\u0026nbsp;7,\u003c/p\u003e\n\u003cdiv id=\"Equg\"\u003e\n \u003cdiv id=\"FileID_Equg\" name=\"EquationSource\"\u003e$$S=sigmoid\\left(Clas{t}^{{\\prime }}\\right)\\dots \\left(7\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eThe sigmoid function ensures that the output values are ormalized between 0 and 1, which is essential for image segmentation tasks. Thus, the proposed model operates through a series of meticulously designed convolutional operations, pooling, up-convolutions, and concatenations, all under the vigilant guidance of an Attention Mechanism process. This intricate process ensures that the model not only captures the complex patterns in the retinal images but also focuses on critical aspects for precise vessel segmentation, resulting in output that is both accurate and detailed for different scenarios. As shown in Fig.\u0026nbsp;\u003cspan\u003e1.3\u003c/span\u003e, the model\u0026rsquo;s ability to emphasize key features while suppressing the irrelevant ones makes it an exceptionally powerful tool for the segmentation of retinal vessels for different scenarios.\u003c/p\u003e\n\u003cp\u003eThese segmented images are further processed for integrated Projected Gradient Descent with a Bidirectional Long Short-Term Memory network for the classification of sets of segmented images. This fusion proposes an effective and subtle approach in the field of medical image analysis. Being able to leverage the iterative optimization power of PGD and the sequential data processing capability from BiLSTM, this combination makes it very good at complex image classification tasks. S denotes the images represented after segmentation, considered as input to this model process. These segmented images are rich in contextual and spatial information, which is very essential for an accurate classification process arising from the previous UNet with Attention Mechanism stage. First of all, these images will be flattened and normalized to create a suitable input format for the PGD algorithm process. Let F(S) be the flattened image vector, where F: R^{H \\times W \\times C} \\rightarrow R^{n}, with H, W, and C being the dimensions of the segmented image, and n being the length of the flattened vector sets. The PGD algorithm is employed to iteratively refine the image features, enhancing their discriminative properties. The iterative update rule for PGD is estimated via Eq.\u0026nbsp;8,\u003c/p\u003e\n\u003cdiv id=\"Equh\"\u003e\n \u003cdiv id=\"FileID_Equh\" name=\"EquationSource\"\u003e$$V\\left(k+1\\right)=ProjS\\left(V\\left(k\\right)+\\alpha \\nabla L\\left(F\\left(S\\right),V\\left(k\\right)\\right)\\right)\\dots . \\left(8\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, \u003cem\u003eV(k)\u003c/em\u003e is the feature vector at iteration \u003cem\u003ek\u003c/em\u003e, \u003cem\u003e\u0026alpha;\u003c/em\u003e is the learning rate, \u0026nabla;L is the gradient of the loss function L with respect to \u003cem\u003eV(k)\u003c/em\u003e, and ProjS is the projection operator that ensures \u003cem\u003eV(k\u003c/em\u003e\u0026thinsp;+\u0026thinsp;1) remains within a feasible set S for different scans. Once the iterative feature refinement via PGD is complete, the resulting feature vectors are fed into the BiLSTM network process. The BiLSTM processes the data in both forward and backward scopes, capturing temporal dependencies in both scopes. The forward and backward hidden states at timestamp \u003cem\u003et\u003c/em\u003e, represented as \u003cspan\u003e\u003cspan\u003e\\(h(t,f)\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan\u003e\u003cspan\u003e\\(h(t,b)\\)\u003c/span\u003e\u003c/span\u003e, are updated via equations 9 \u0026amp; 10,\u003c/p\u003e\n\u003cdiv id=\"Equi\"\u003e\n \u003cdiv id=\"FileID_Equi\" name=\"EquationSource\"\u003e$$h\\left(t,f\\right)=LSTM\\left(h\\left(t-1\\right),V\\left(t\\right)\\right)\\dots \\left(9\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Equj\"\u003e\n \u003cdiv id=\"FileID_Equj\" name=\"EquationSource\"\u003e$$h\\left(t,b\\right)=LSTM\\left(h\\left(t+1\\right),V\\left(t\\right)\\right)\\dots \\left(10\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, \u003cem\u003eV(t)\u003c/em\u003e is the input feature vector at time \u003cem\u003et\u003c/em\u003e, and LSTM represents the LSTM operation process. At the heart of the LSTM unit are several gates and states that work in unison to regulate the flow of information sets. These include the forget gate, input gate, and output gate, along with the cell state and hidden states. For the single LSTM unit at timestamp \u003cem\u003et\u003c/em\u003e, where \u003cem\u003eh(t\u003c/em\u003e\u0026thinsp;\u0026minus;\u0026thinsp;1) is the hidden state from the previous timestamp, and \u003cem\u003eVt\u003c/em\u003e is the current input vector sets. The dynamics of the LSTM unit are governed by the following process,\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eForget Gate\u003c/strong\u003e: The forget gate decides what information to discard from the cell states. \u003cem\u003eWf\u003c/em\u003e and \u003cem\u003ebf\u003c/em\u003e are the weight matrix and bias for the forget gate, respectively, and \u003cem\u003e\u0026sigma;\u003c/em\u003e represents the sigmoid function, which are fused via Eq.\u0026nbsp;11 to form the forget gate,\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cdiv id=\"Equk\"\u003e\n \u003cdiv id=\"FileID_Equk\" name=\"EquationSource\"\u003e$$f\\left(t\\right)=\\sigma \\left(Wf\\cdot \\left[h\\left(t-1\\right),V\\left(t\\right)\\right]+bf\\right)\\dots \\left(11\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eInput Gate\u003c/strong\u003e: The input gate determines which values will be updated in the cell states. \u003cem\u003eWi\u003c/em\u003e and \u003cem\u003ebi\u003c/em\u003e are the weight matrix and bias for the input gate, which is represented via Eq.\u0026nbsp;12,\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cdiv id=\"Equl\"\u003e\n \u003cdiv id=\"FileID_Equl\" name=\"EquationSource\"\u003e$$i\\left(t\\right)=\\sigma \\left(Wi\\cdot \\left[h\\left(t-1\\right),V\\left(t\\right)\\right]+bi\\right)\\dots \\left(12\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eCell State Candidate\u003c/strong\u003e: This creates a vector of candidate values that could be added to the cell state, and is represented via Eq.\u0026nbsp;13,\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cdiv id=\"Equm\"\u003e\n \u003cdiv id=\"FileID_Equm\" name=\"EquationSource\"\u003e$$C\\sim\\left(t\\right)=\\varvec{t}\\varvec{a}\\varvec{n}\\varvec{h}\\left(WC\\cdot \\left[h\\left(t-1\\right),V\\left(t\\right)\\right]+bC\\right)\\dots \\left(13\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eCell State Update\u003c/strong\u003e: The cell state is updated by combining the past state \u003cem\u003eC(t\u003c/em\u003e\u0026thinsp;\u0026minus;\u0026thinsp;1) and the new candidate values, modulated by the forget and input gates via Eq.\u0026nbsp;14,\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cdiv id=\"Equn\"\u003e\n \u003cdiv id=\"FileID_Equn\" name=\"EquationSource\"\u003e$$C\\left(t\\right)=f\\left(t\\right)*C\\left(t-1\\right)+i\\left(t\\right)*C\\sim\\left(t\\right)\\dots \\left(14\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eOutput Gate\u003c/strong\u003e: The output gate controls the output from the LSTM units. \u003cem\u003eWo\u003c/em\u003e and \u003cem\u003ebo\u003c/em\u003e are the weight matrix and bias for the output gate, which is estimated via Eq.\u0026nbsp;15,\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cdiv id=\"Equo\"\u003e\n \u003cdiv id=\"FileID_Equo\" name=\"EquationSource\"\u003e$$o\\left(t\\right)=\\sigma \\left(Wo\\cdot \\left[h\\left(t-1\\right),V\\left(t\\right)\\right]+bo\\right)\\dots \\left(15\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eHidden State Update\u003c/strong\u003e: The hidden state for the current timestamp \u003cem\u003et\u003c/em\u003e is obtained by filtering the cell state through the output gate via Eq.\u0026nbsp;16,\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cdiv id=\"Equp\"\u003e\n \u003cdiv id=\"FileID_Equp\" name=\"EquationSource\"\u003e$$h\\left(t\\right)=o\\left(t\\right)*\\varvec{t}\\varvec{a}\\varvec{n}\\varvec{h}\\left(C\\left(t\\right)\\right)\\dots \\left(16\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eBased on this, the model estimates forward and backward hidden states, which are then concatenated to form a comprehensive feature representation for each timestamp via Eq.\u0026nbsp;17,\u003c/p\u003e\n\u003cdiv id=\"Equq\"\u003e\n \u003cdiv id=\"FileID_Equq\" name=\"EquationSource\"\u003e$$H\\left(t\\right)=h\\left(t,f\\right)\\oplus h\\left(t,b\\right)\\dots \\left(17\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, \u0026oplus; represents concatenation process. The concatenated hidden states \u003cem\u003eHt\u003c/em\u003e are then passed through a fully connected layer followed by a softmax activation to obtain the classification probabilities for each of the image sets. The classification layer is described via Eq.\u0026nbsp;18,\u003c/p\u003e\n\u003cdiv id=\"Equr\"\u003e\n \u003cdiv id=\"FileID_Equr\" name=\"EquationSource\"\u003e$$y\\left(t\\right)=softmax\\left(Wc*Ht+bc\\right)\\dots \\left(18\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere Wc and bc are weights and bias of the classification layer, respectively, and y(t) is the output vector representing classification probabilities. The classification results obtained for each segmented image, thus, would comprise the probability of the image belonging to various DR types. These are then aggregated to form the complete classification results for the whole set of segmented images and provide an overall view of the types of DR represented in both the dataset and the samples.\u003c/p\u003e\n\u003cp\u003eThis will also validate the results of the classification by applying Ensemble Deep Learning methodology in an innovative way, which merges Naive Bayes, k-Nearest Neighbors, Support Vector Machine, Logistic Regression, and Multilayer Perceptron for processing the segmented images and blood reports collected. Only very few such ensemble models exist where traditional machine learning algorithms meet advanced deep learning techniques to carve out a pathway for comprehensive and accurate medical diagnostics.\u003c/p\u003e\n\u003cp\u003eThe model has two very different inputs: images that are segmented and blood reports. Consider XI to be the set of features that are to be extracted from the segmented images and XB, a set of features derived from blood reports. In more preliminary steps, both types of inputs will be dealt with separately by the model in order to extract specific characteristics from each of the domains. Finally, features are extracted from the segmented images, FI(XI), and from the blood reports, FB(XB), through equations 19 \u0026amp; 20 as shown below,\u003c/p\u003e\n\u003cdiv id=\"Equs\"\u003e\n \u003cdiv id=\"FileID_Equs\" name=\"EquationSource\"\u003e$$FI\\left(XI\\right)=\\sigma \\left(Wn*I+bn\\right)\\dots \\left(19\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Equt\"\u003e\n \u003cdiv id=\"FileID_Equt\" name=\"EquationSource\"\u003e$$F\\left(B\\right)\\left[k\\right]=\\sum _{t=0}^{T-1}B\\left[t\\right]\\cdot {e}^{-\\frac{2\\pi ikt}{T}}\\dots \\left(20\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, \u003cem\u003eB\u003c/em\u003e represents the blood report signal, where \u003cspan\u003e\u003cspan\u003e\\(B\\in RT\\)\u003c/span\u003e\u003c/span\u003e and \u003cem\u003eT\u003c/em\u003e represents length of the signals. These feature nextractors leverage deep learning techniques to distill essential features from the data, preparing them for subsequent analysis. The ensemble model then applies multiple classifiers to these features. Each classifier \u003cem\u003eCi\u003c/em\u003e, where i represents Naive Bayes, kNN, SVM, Logistic Regression, or MLP, processes the features and outputs a prediction via Eq.\u0026nbsp;21,\u003c/p\u003e\n\u003cdiv id=\"Equu\"\u003e\n \u003cdiv id=\"FileID_Equu\" name=\"EquationSource\"\u003e$$Pi=Ci\\left(FI\\left(XI\\right)\\oplus FB\\left(XB\\right)\\right)\\dots \\left(21\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, \u0026oplus; represents the concatenation of the feature sets from the image and blood report domains. The ensemble approach employs a weighted voting system to integrate the outputs from each of the classifiers. The final classification decision, \u003cem\u003eD\u003c/em\u003e, is determined by aggregating these predictions, taking into account the weights assigned to each classifier \u003cem\u003ewi\u003c/em\u003e via Eq.\u0026nbsp;22,\u003c/p\u003e\n\u003cdiv id=\"Equv\"\u003e\n \u003cdiv id=\"FileID_Equv\" name=\"EquationSource\"\u003e$$D=argmax\\left(\\sum _{i}w\\left(i\\right)\\cdot P\\left(i\\right)\\right)\\dots \\left(22\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eThe weights \u003cem\u003ewi\u003c/em\u003e are optimized during the training phase to reflect each classifier\u0026rsquo;s contribution to the overall performance of the ensemble process.\u003c/p\u003e\n\u003cp\u003eTo further refine the model\u0026apos;s predictions, a feedback loop is incorporated in this process. This loop involves assessing the confidence level of the classification results and, if this confidence is below a certain threshold, the model reiterates this process with adjusted parameters for different operations. Let \u003cspan\u003e\u003cspan\u003e\\(Conf\\left(D\\right)\\)\u003c/span\u003e\u003c/span\u003e represent the confidence score of the decision which is estimated via Eq.\u0026nbsp;23, and \u003cem\u003e\u0026theta;\u003c/em\u003e be the threshold levels.\u003c/p\u003e\n\u003cdiv id=\"Equw\"\u003e\n \u003cdiv id=\"FileID_Equw\" name=\"EquationSource\"\u003e$$Conf\\left(D\\right)=max\\left(P1,P2,\\dots ,Pn\\right)\\dots \\left(23\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, \u003cspan\u003e\u003cspan\u003e\\(Pc\\)\u003c/span\u003e\u003c/span\u003e represents the probability of class \u003cem\u003ec\u003c/em\u003e being the correct classification, and an is the number of classes. For lower confidence classes, the model parameters are tuned via Eq.\u0026nbsp;24,\u003c/p\u003e\n\u003cp\u003e\u003cspan\u003e\u0026nbsp;\u003cspan\u003e\\(p{i}^{{\\prime }}=pi+\\varDelta pi\\dots \\left(24\\right)\\)\u003c/span\u003e\u0026nbsp;\u003c/span\u003eWhere, p\u003cem\u003ei\u003c/em\u003e\u0026prime; is the adjusted parameter, and \u0026Delta;\u003cem\u003epi\u003c/em\u003e is the change applied to the original parameter p\u003cem\u003ei\u003c/em\u003e sets. Incorporating this feedback mechanism ensures that the model dynamically adapts to the intricacies of the data, enhancing the accuracy of the results. The final output of the ensemble model is a set of validated classification results, taking into account both the segmented retinal images and the corresponding blood reports. This dual-input approach enables the model to cross validate its findings, ensuring a higher degree of reliability and precision in the classification of medical conditions. Efficiency of this model was evaluated for different scenarios, and compared with existing methods in terms of different performance metrics in the next section of this text.\u003c/p\u003e"},{"header":"Result Analysis \u0026 Comparisons","content":"\u003cp\u003eThis will be a new frontier in medical imaging, particularly in the detection and segmentation of DR image scans. The model represents the fusion of multimodalities together with attention mechanisms, further strengthened by adversarial training and ensemble deep learning operations. It excelled in adaptability and robustness, able to efficiently deal with diversified and large data with remarkable precision and accuracy. Advanced AI techniques embody the ingenuity of the ERBVS model through attention mechanisms that focus on key features in retinal images, on one hand, and adversarial training, which sharpens its ability to tell between DR types on the other. The ensemble learning part at the core of this model fuses a diversity of deep learning strategies, guaranteeing in-depth learning from complex data patterns. Not only is there a considerable enhancement of the accuracy in DR diagnosis in this model, but it also reduces false negatives and false positives, which are very important in medical diagnostics. Sensitivity, specificity, precision, recall, and reduced processing delay have been some of the key factors in the making of this ERBVS model, which definitely holds promise to be one of the most pioneering solutions for changing the screening and diagnosis of DR forever, hence carving a niche for AI in healthcare scenarios.\u003c/p\u003e \u003cp\u003eIn this work, an experimental setting has been designed that will carefully evaluate the performance of the proposed ERBVS model in the classification and segmentation of retinal images with respect to the different types of DR. The setting will be comprehensive with different datasets that ensure the robustness of the evaluation.\u003c/p\u003e \u003cp\u003e \u003cb\u003eDatasets Used\u003c/b\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eIndian Diabetic Retinopathy Dataset\u003c/b\u003e: Accessed from Kaggle (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.kaggle.com/datasets/aaryapatel98/indian-diabetic-retinopathy-image-dataset\u003c/span\u003e\u003cspan address=\"https://www.kaggle.com/datasets/aaryapatel98/indian-diabetic-retinopathy-image-dataset\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), this dataset includes high-resolution fundus images specifically representing the Indian demographic. It provides a diverse set of images, aiding in assessing the model's performance across different ethnic backgrounds.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eDeep Diabetic Retinopathy Dataset\u003c/b\u003e: Available on Kaggle (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.kaggle.com/c/diabetic-retinopathy-detection\u003c/span\u003e\u003cspan address=\"https://www.kaggle.com/c/diabetic-retinopathy-detection\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), this dataset consists of a large collection of high-quality retinal images. It is used to evaluate the model's scalability and performance in handling nextensive and varied data.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eTensorFlow Dataset Samples\u003c/b\u003e: Sourced from the TensorFlow Datasets Catalog (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.tensorflow.org/datasets/catalog/diabetic_retinopathy_detection\u003c/span\u003e\u003cspan address=\"https://www.tensorflow.org/datasets/catalog/diabetic_retinopathy_detection\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), this dataset offers a standardized set of images and is utilized to benchmark the model against established datasets in the field.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eAPTOS Dataset Samples\u003c/b\u003e: Available via Academic Torrents (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://academictorrents.com/details/d8653db45e7f111dc2c1b595bdac7ccf695efcfd\u003c/span\u003e\u003cspan address=\"https://academictorrents.com/details/d8653db45e7f111dc2c1b595bdac7ccf695efcfd\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), this dataset includes images from the APTOS Blindness Detection competition, providing a range of images with varying degrees of DR severity.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eAGAR300 Dataset Samples\u003c/b\u003e: Accessed from IEEE DataPort (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieee-dataport.org/open-access/diabetic-retinopathy-fundus-image-datasetagar300\u003c/span\u003e\u003cspan address=\"https://ieee-dataport.org/open-access/diabetic-retinopathy-fundus-image-datasetagar300\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), this dataset comprises 300 high-quality fundus images specifically curated for DR research. It offers a focused dataset for detailed analysis.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eExperimental Parameters\u003c/b\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eImage Preprocessing\u003c/b\u003e: Images are resized to a standard dimension of 256x256 pixels. Histogram equalization and contrast adjustment are performed to enhance image quality levels.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eModel Training\u003c/b\u003e: The ERBVS model is trained using a 70:30 split for each dataset, with 70% of the data used for training and 30% for validation process.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eAugmentation Techniques\u003c/b\u003e: Data augmentation, including rotations, flips, and zooms, is applied to increase the diversity of the training data samples. This assists in generating larger data samples for training, testing \u0026amp; validation operations.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eLearning Rate\u003c/b\u003e: An initial learning rate of 0.001 is employed, with a decay rate of 0.1 every 10 epochs.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eOptimizer\u003c/b\u003e: Adam optimizer is utilized for its efficiency in handling sparse gradients and adaptive learning rate capabilities.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eLoss Functions\u003c/b\u003e: A combination of cross-entropy and dice loss functions is used to optimize the segmentation and classification tasks.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eBatch Size\u003c/b\u003e: A batch size of 32 is chosen to balance between computational efficiency and model performance levels.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eEpochs\u003c/b\u003e: The model is trained for 50 epochs, providing a balance between sufficient learning time and avoiding overfitting scenarios.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eHardware Specifications\u003c/b\u003e: Training is conducted on a system equipped with an NVIDIA Tesla P100 GPU, ensuring high computational power for processing large datasets \u0026amp; samples.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eMetrics Evaluated\u003c/b\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003ePrecision, Recall, Accuracy, Specificity, AUC, and MMSE are the primary metrics used to evaluate the model's performance.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eA comparative analysis is conducted with existing models MSRNet, DPFNet, and EDGAN to benchmark the ERBVS model's effectiveness.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eFor each dataset, specific configurations are adjusted to account for inherent dataset characteristics like image quality, resolution, and distribution of DR types. Based on the experimental setup, Precision, Accuracy, Recall, Latency, AUC, and Specificity ratings are used to assess the process's effectiveness. Equations\u0026nbsp;25, 26, 27, and 28 were used to determine these parameters for the entire number of NI images, and the results were compared with those from MSR UNet [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], DPFNet [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] and Ensembled Deep Convolutional Generative Adversarial Network (EDGAN) [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e], which use similar prediction techniques.\u003cdiv id=\"Equx\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equx\" name=\"EquationSource\"\u003e\n$$P=\\frac{1}{NI}\\sum _{i=1}^{NI}\\frac{tp\\left(i\\right)}{tp\\left(i\\right)+fn\\left(i\\right)}\\dots \\left(25\\right)$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equy\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equy\" name=\"EquationSource\"\u003e\n$$A=\\frac{1}{NI}\\sum _{i=1}^{NI}\\frac{tp\\left(i\\right)+tn\\left(i\\right)}{\\begin{array}{c}tp\\left(i\\right)+tn\\left(i\\right)+\\\\ fp\\left(i\\right)+fn\\left(i\\right)\\end{array}}\\dots \\left(26\\right)$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equz\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equz\" name=\"EquationSource\"\u003e\n$$R=\\frac{1}{NI}\\sum _{i=1}^{NI}\\frac{tp\\left(i\\right)}{\\begin{array}{c}tp\\left(i\\right)+tn\\left(i\\right)+\\\\ fp\\left(i\\right)+fn\\left(i\\right)\\end{array}}\\dots \\left(27\\right)$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equaa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equaa\" name=\"EquationSource\"\u003e\n$$d=\\frac{1}{NI}\\sum _{i=1}^{NI}ts(complete, i)-ts(start, i)\\dots \\left(28\\right)$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere fp and fn represent the correct and incorrect counts for classifying inputs into incorrect classes, tp and tn represent the number of samples that are correctly classified into a given class and an incorrect class, respectively, and ts(complete) \u0026amp; ts(start) represent the timestamps for when the segmentation and classification processes were finished and started, respectively, for different evaluations. Using this, the Precision for segmentation was estimated w.r.t. Number of Test Samples (NTS), and can be observed from Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e3\u003c/span\u003e as follows,\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFor small NTS, like 5k, the performance of ERBVS is way ahead of the rest: it is at 93.18% _precision, against EDGAN at 88.23%, DPFNet at 82.34%, and MSRNet at 75.02%. That means ERBVS does pretty well with small datasets, which usually suffer from the scarcity of data. For instance, as the NTS increases to 8k and 10k, the ERBVS slightly decreases and then increases in precision, thereby indicating its adaptability and learning efficiency even with escalating data complexity. In the middle range of NTS, from 12k through 26k, ERBVS modulated between higher ranges of precision and peaked at 96.99% for 22k NTS, whereas the other models were very variable. This demonstrates ERBVS's more remarkable ability to process larger and more diverse datasets, which is important in real-world applications. Of particular note is the performance of ERBVS at 19k and 22k NTS, where the gap opened against others by it is dramatic, indicative of its robustness to complex image sets. While the NTS continued to increase further up to 29,000 and higher, ERBVS gave a very high precision with an exceptional peak at 99.80% at an NTS of 43,000, outclassing its closest competitor, EDGAN, which showed a maximum of 93.38% at 29,000 NTS. Such high performance can only be attributed to the fact that ERBVS fuses multimodal information using mechanisms of attention and adversarial training to develop one that perfectly differentiates the relevant features from the rest in such wide and diversified datasets. The highest precision obtained for ERBVS lies in the 43k and 48k NTS, both of which arrive at ear-perfect precision scores of 99.80% and 99.40%, respectively. This outstanding performance is a consequence of the fact that it is brisk in ensemble deep learning operations, particularly in complex scenarios or high-accuracy applications. The ERBVS model therefore has better performance compared to the other analyzed NTSs, more especially in big datasets and advanced features, such as attentional mechanisms and ensemble operations that bring about high-precision levels\u0026mdash;hence, quite an outstanding model for the segmentation of blood vessels in the retina, especially in scenarios where high accuracy is critical. The greater accuracy of ERBVS not only provides more reliable segmentation but also improves its applicability in the clinical setting where accurate diagnosis is the prime requirement for an effective treatment planning process. Figure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003e: Accuracy levels achieved during the identification of DR classes using such techniques.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAt the beginning, when the NTS is small at 5k, the ERBVS accuracy is remarkably high, reaching 95.18%, far surpassing EDGAN with 88.23%, DPFNet with 87.34%, and MSRNet with 75.02%. This high accuracy in smaller datasets suggests the superior ability of ERBVS to generalize from limited data, which should be a pertinent factor in medical imaging where not every dataset could be large. While the NTS increased up to 8k and 10k, ERBVS managed to stay quite strong with an observable slight dip at 10k with 86.51%. This could reflect the adaptability of the model and the learning curve with respect to a slightly higher dataset size. Particularly, on 10k NTS, DPFNet showed an accuracy of 87.09%, thus proving that probably these models have different strengths with a view to various sizes of the dataset. At mid-range NTS (12k to 26k), ERBVS generally has very high accuracy, peaking at 19k NTS with 99.75% and another time at 22k NTS with 99.99%. This exceptionally high accuracy, mostly at 22k NTS, further underscores the competence of this model to have more complex and large datasets, probably due to advanced fusion of multiple modalities along with attention mechanisms. At 25k and 26k NTS, though showing reduced accuracy compared to its optimum, ERBVS was beyond the other models. This behavior can be ascribed to the model's sensitivity to growing data variability or higher complexity inherent in larger image sets. As these sets continue to grow further from 29 to 60k NTS, the accuracy graph for ERBVS does not necessarily take on a consistent pattern, peaking significantly at 31k with an accuracy of 94.83 percent and spiking at 60k with an accuracy of 99.50 percent. It could be that this fluctuation in accuracy is a consequence of how the model handles large datasets with more diversified, complex characteristics of data. At any rate, high accuracy is maintained at certain points\u0026mdash;for instance, at 60k NTS\u0026mdash;a result that showcases the strength and effectiveness of ERBVS in large-scale scenarios. Similar to Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e5\u003c/span\u003e shows the categorization recall as follows,\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFor small NTS (5k) at an early stage, ERBVS has a very high recall of 92.48%, outperforming other models significantly. It tells that it has exceptionally outstanding ability in identifying positive DR cases, which is of central importance in all cases of disease diagnosis at their early stages. In contrast, when using an 8k NTS, the recall of ERBVS drastically drops to 82.09%, slightly behind MSRNet's 85.37%. Such fluctuation could also allude to the challenges models face in maintaining consistency because of the different dataset size for scenarios. For instance, as the NTS increases, it bounces back with an ERBVS of 85.27% and 86.54%, showing the strength and flexibility in recall rates for different use cases. This recovery in performance may be reflective of its learning prowess and robustness while handling large amounts and increasing levels of complexity in data. On the middle-range area of 14k through 26k NTS, there is an interesting performance interplay. In general, ERBVS held a higher recall compared to EDGAN and MSRNet, peaking remarkably at 16k NTS with 92.92%. That is, ERBVS is extremely accurate at correctly classifying the positive DR cases within a moderately large dataset\u0026mdash;another principal factor necessary for use within a clinical setting. There is also another interesting aspect at higher values of NTS: from 29k to 60k, ERBVS's recall demonstrates a very steep uptrend, reaching its peak at 41k (93.84%), approaching nearly perfect recall at 46k NTS with 99.88%. These very good performances on large NTS suggest that sophisticated algorithms, most likely fusion between multiple modalities and attention mechanisms of ERBVS, make them highly capable of processing large and complex datasets \u0026amp; samples. Independently compared against DPFNet and MSRNet across several NTS, in most cases, ERBVS retains most of its competitive advantage, more so on larger datasets. It gives better recall rates under these conditions, thus showing the potential for clinical application in real-world scenarios where large-scale screening of various situations is done. Proposed models are potential for effective usage in clinical diagnosis for DR for the reason that it can attain pretty high proficiency in recall, especially with bigger datasets. This goes to prove the robustness and level of reliability resultant from its ability to keep on picking all positive cases across different dataset sizes, with a special focus on larger datasets. These are very important considerations in the area of medical imaging, where missing a positive case is costly in every sense of the word, lending support to the possible impact of this study on improving the process of patient care and disease management. Similar to Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e5\u003c/span\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e6\u003c/span\u003e shows the categorization delay as follows,\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eRight from the beginning, on a smaller NTS of 5k, it is quite outstanding\u0026mdash;the lowest delay, 52.03 ms, was offered by ERBVS, against 63.84 ms offered by MSRNet, 79.64 ms by DPFNet, and 78.09 ms by EDGAN. This faster speed on smaller datasets gives the potential of our ELRBC network for applications that demand faster results. For example, in emergency blinded diagnosis cases or high-throughput screening scenarios, some initial results are required immediately before conducting further analysis with temporal instance sets. It continues to post efficiency with delays of 58.55 ms and 57.49 ms for the 8k and 10k NTS, respectively. Even as the delay slightly increases compared to the 5k NTS, ERBVS still stays ahead of the other models. This consistent performance in processing speed with increasing dataset sizes underlines the robust computational design of the model, probably assisted by advanced ensemble deep learning operations. Within the middle-range NTS from 12k to 26k, ERBVS generally kept a rather efficient processing time, with remarkable performances at 14k, 55.09 ms, and 19k, 52.24 ms. This means that the model could have clinical applications without large delays on larger datasets. However, if NTS is further increased from 29k to 60k, there is still a trivial increase in the average delay for ERBVS. The maximum average delay is 63.95 ms for 22k NTS and 64.49 ms for 46k NTS. Though it is there, the increase consistently stays in front with ERBVS, again proving that it is much more capable of dealing efficiently with computational load, even in large-scale datasets. The trend noted in ERBVS's delay across the different NTS brings into clear perspective how the model balances between its accuracy and efficiency in the sense that, while remaining very competitive at speed, it does not compromise on accuracy or recall, as has already been noted in previous comparisons. This balance is important in real clinical scenarios in which accuracy and time are two most vital components for the success of any case related to the management of disease and effective patient care. Similarly, the MMSE for these evaluations can be observed in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e7\u003c/span\u003e,\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFor instance, using a smaller NTS of 5k, ERBVS always performs in a superior manner with the lowest MMSE of 0.1317, thus showing that it is more accurate in segmenting images correctly compared to MSRNet's 0.1528, DPFNet's 0.1456, and EDGAN's 0.1555. This large degree of accuracy in small datasets is quite imperative, especially in medical imaging, where accurate segmentation for even a few images can be very vital. The ERBVS keeps a low MMSE at 0.1290 and 0.1352 for 8k and 10k, respectively, when the NTS increases, portraying that it can be quite robust in handling slightly higher datasets without significant loss of accuracy in segmentation. This supports evidence of its superior algorithms that might have been improved by its ensemble deep learning with attention mechanisms. At low- to mid-range NTS, ERBVS will generally maintain a lower MMSE than the other models, with slight increases, such as at 14k, 0.1378, and 25k, 0.1483. This could be tied to the adaptability of the model and its response due to the growing complexity and variability of larger image sets. Noticeably, with further increases in NTS from 29k to 60k, ERBVS shows a changing MMSE, peaking at some points; for example, at 50k NTS with 0.1455. However, it always stands competitive and as often as not outperforms the remaining models. It follows that ERBVS was able to maintain segmentation accuracy in large-scale datasets. For instance, the average MMSE for ERBVS demonstrated relative performance across different numbers of time sections that just about explained its better overall performance in terms of segmentation accuracy. While it did exhibit some fluctuation within its MMSE with an increase in NTS, these variations always remained at a relatively low range that underscored the effectiveness of this model in producing accurate segmentations consistently for different scenarios. The AUC levels can be observed from Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e7\u003c/span\u003e as follows,\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFirst, for a relatively small NTS of 5k, ERBVS has an AUC of 90.72%, against EDGAN at 90.40%, DPFNet at 87.39%, and MSRNet at 86.73%. It can then be interpreted that this is evidence of the effective distinction capability of ERBVS among different DR types, even in smaller datasets, a very important factor in disease early and accurate detection. In the case of 8k and 10k NTS, ERBVS floats around with a competitive AUC of 86.28% and 89.39%, respectively. This clearly shows its adaptability and reliability for classification on any dataset size. At 8k NTS of EDGAN, it went ahead of ERBVS with an AUC of 91.55%; this, however, shows that on some certain characteristics of the dataset, different models would have varied strengths. For ERBVS, as the NTS increases in the mid-range\u0026mdash;from 12,000 to 26,000\u0026mdash;its AUC performance changes, dipping at 25,000 NTS to 84.46%. This could be how the model responds to an increase in complexity and variability that naturally comes with a greater number of images. At any rate, ERBVS still has fairly competitive AUC values, since it is able to classify effectively in moderately large datasets. Though the NTS further increases to 29k, 60k, ERBVS still shows an enormous improvement in AUC values to its peak at 96.92% with the 31k NTS, while peaking at 96.95% and 97.09% with 34k and 60k NTS, respectively. High AUC values on these large NTS indicate that very large and complex datasets can be handled more deftly by advanced algorithms and ensemble learning techniques of ERBVS\u0026mdash;and this has established the far-flung clinical diagnosis application base. Comparing these with DPFNet, MSRNet, and EDGAN for various NTS, one can see that ERBVS generally performs very well and sometimes oscillates. Clearly, its more competitive AUC scores indicate that in a real-world clinical application with large-scale screening and classification requirements in different use cases, it has great potential to actually apply. While, the specificity levels can be observed from Fig.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e9\u003c/span\u003e as follows,\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn this case, at a smaller NTS of 5k, ERBVS has great specificity of 82.03%. This is comparable to EDGAN's 83.11%, which is still higher than that of DPFNet. Often, this result illustrates that ERBVS is strong in classifying negative patients with events, which means that it detects diseases at their early stages to avoid overdiagnosis. By extension, on the enlarged NTS of 8k and 10k, ERBVS gives an upward trend of specificity to about 85.76% and 89.25%, respectively. This increase highlights the growing accuracy of ERBVS in large datasets, which is indicative of the system's robustness and capacity to handle a diversity of data without miss-classifying healthy cases as diseased. In mid-range NTS from 12k to 26k, ERBVS' specificity changes but still remains very competitive. Especially for the 16k NTS, it has a very outstanding specificity of 95.47%, showing its great ability to avoid false positives in moderately large datasets. This performance is critical in clinical environments, wherein false positives may be dear. As NTS is increased even further, from 29,000 to 60,000, ERBVS holds up extremely well, peaking quite notably at 36,000 with an almost perfect score of 99.86%. The deviations from the observed high performance in this unusual performance at larger NTS are absolutely riveting, suggesting that ERBVS does truly have advanced algorithms, probably with attention mechanisms and ensemble deep learning operations working in large-scale and complex datasets and samples. Comparing these results against DPFNet, MSRNet, EDGAN, across a range of NTS, obviously, in most cases, ERBVS usually retains this heady performance when it comes to specificity. Its potential usefulness in wide clinical diagnostic applications is underlined, call it an ability to recognize on-DR cases correctly across all these consistently, specifically in large datasets. This means that strong performance in specificity across different dataset sizes is expressed through ERBVS, making it a potentially very useful tool in clinical diagnostics for DR. Its reliability and consistency for identifying negative cases, particularly in large datasets, puts it as a potential high-impact model in improving the care of these patients by reducing the risk of unnecessary treatments or interventions due to false-positive diagnoses.\u003c/p\u003e"},{"header":"Conclusion and Future Scopes","content":"\u003cp\u003eThe paper contributes much in the field of medical imaging, particularly for the diagnosis and management of diabetic retinopathy. In the paper, a new model called ERBVS was proposed that delivered superior performance in parameters such as precision, recall, accuracy, specificity, AUC, MMSE, and processing delay against already developed models like MSRNet, DPFNet, and EDGAN using advanced techniques of Artificial Intelligence. Key findings included that this model showed an exceptional result on different dataset sizes, thus its adaptability and robustness in segmentation accuracy and classification efficacy. Another part of this, element of scalability, is the prominent performance of the ERBVS in larger datasets. This work also found high precision and recall rates to help in ensuring the accurate identification of DR, which helps reduce the number of false negatives that are very important in preventing delayed or missed diagnosis. The specificity analysis further consolidated the fact that ERBVS is competent in accurately differentiating on-DR cases, which reduces false positives and excess associated interventions. Moreover, this reduction in processing delay points toward the potential of the model in delivering timely diagnostics, a very critical factor in clinical decision-making process.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eImpact of This Work\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThese better results in the detection and segmentation of DR using the ERBVS model show several important ramifications:\u003c/p\u003e\n\n\u003cp\u003e\u003cspan\u003e1. Higher Diagnostic Accuracy: Precise diagnosis of DR could mean a much better outcome for patients. The key to effective management of the disease is essentially hinged on early and accurate detection.\u003cbr\u003e\u003c/span\u003e\u003cspan\u003e2. Less Burden of Health Care System: ERBVS shares a significant role in reducing false positives and false negatives, hence adding up to the efficacy of the health care system by fewer unwarranted treatments and repeated follow-up checks.\u003cbr\u003e\u003c/span\u003e\u003cspan\u003e3. Mass Screening Scalability: The proficiency of the model in handling bulk data makes it an ideal candidate for mass screening programs, especially needed in countries where diabetes prevalence is high.\u003cbr\u003e\u003c/span\u003e\u003cspan\u003e4. Across Diverse Populations Customization: Due to its tested effectiveness on a wide range of datasets and its performance on datasets specific to Indian demographics, ERBVS shows promise for customization and application across different ethnic groups.\u003cbr\u003e\u003c/span\u003e\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eFuture Scope\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThere is much scope for future research and development in several ways as envisaged herein:\u003c/p\u003e\n\u003cp\u003eOphthalmic Devices Integration: ERBVS could be directly integrated into ophthalmic imaging devices to ease diagnosis by providing real-time analysis during patient examinations.\u003c/p\u003e\n\u003cp\u003eOther Ocular Diseases Detection: This model will have wider applications in its ability to detect other ocular conditions for it to become one of the umbrella diagnostic tools of ophthalmology.\u003c/p\u003e\n\u003cp\u003eERBVS in Telemedicine Applications: This could become very important in telemedicine, particularly in the remote or underserved parts of the country where access to ophthalmologists is limited.\u003c/p\u003e\n\n\u003cp\u003e\u0026bull; Individualized Treatment Plans: Later versions may incorporate predictive analytics that point toward therapy tailored according to the severity and progression of DR.\u003c/p\u003e\n\u003cp\u003e\u0026bull; Interdisciplinary Applications: Tapping into the application potential off this model in other areas of medical imaging opens new frontiers in AI-driven diagnostics.\u003c/p\u003e\n\n\u003cp\u003eIn a nutshell, ERBVS pushes the bounds on how AI could really revolutionize medical diagnostics. Its development not only points to improving solutions in the present for DR but also opens doorways to wider applications in healthcare, contributing to more efficient, accurate, and accessible medical services.\u003c/p\u003e\n"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgment\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests:\u003c/strong\u003e The author declares no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; Contribution Statement:\u0026nbsp;\u003c/strong\u003eThe sole author conceived, conducted, and analyzed the research, and drafted the manuscript independently in clinical scenarios.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical and Informed Consent for Data Used:\u0026nbsp;\u003c/strong\u003ePlagiarism has been strictly avoided and proper acknowledgment has been given to all sources used. Any errors identified will be promptly corrected ensuring the integrity of the work.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability and Access\u003c/strong\u003e: The dataset used in this paper is openly available\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eS. Deari, I. Oksuz and S. Ulukaya, \u0026quot;Block Attention and Switchable Normalization Based Deep Learning Framework for Segmentation of Retinal Vessels,\u0026quot; in IEEE Access, vol. 11, pp. 38263-38274, 2023, doi: 10.1109/ACCESS.2023.3265729.\u003c/li\u003e\n\u003cli\u003eY. Li et al., \u0026quot;Global Transformer and Dual Local Attention Network via Deep-Shallow Hierarchical Feature Fusion for Retinal Vessel Segmentation,\u0026quot; in IEEE Transactions on Cybernetics, vol. 53, o. 9, pp. 5826-5839, Sept. 2023, doi: 10.1109/TCYB.2022.3194099.\u003c/li\u003e\n\u003cli\u003eG. B. Kande et al., \u0026quot;MSR U-Net: An Improved U-Net Model for Retinal Blood Vessel Segmentation,\u0026quot; in IEEE Access, vol. 12, pp. 534-551, 2024, doi: 10.1109/ACCESS.2023.3347196.\u003c/li\u003e\n\u003cli\u003eJ. Li, G. Gao, L. Yang, G. Bian and Y. Liu, \u0026quot;DPF-Net: A Dual-Path Progressive Fusion Network for Retinal Vessel Segmentation,\u0026quot; in IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1-17, 2023, Art o. 2517817, doi: 10.1109/TIM.2023.3277946.\u003c/li\u003e\n\u003cli\u003eK. Aurangzeb, R. S. Alharthi, S. I. Haider and M. Alhussein, \u0026quot;An Efficient and Light Weight Deep Learning Model for Accurate Retinal Vessels Segmentation,\u0026quot; in IEEE Access, vol. 11, pp. 23107-23118, 2023, doi: 10.1109/ACCESS.2022.3217782.\u003c/li\u003e\n\u003cli\u003eM. Arsalan, T. M. Khan, S. S. Naqvi, M. Nawaz and I. Razzak, \u0026quot;Prompt Deep Light-Weight Vessel Segmentation Network (PLVS-Net),\u0026quot; in IEEE/ACM Transactions on Computational Biology and Bioinformatics, vol. 20, o. 2, pp. 1363-1371, 1 March-April 2023, doi: 10.1109/TCBB.2022.3211936.\u003c/li\u003e\n\u003cli\u003eN. Shen, T. Xu, S. Huang, F. Mu and J. Li, \u0026quot;Expert-Guided Knowledge Distillation for Semi-Supervised Vessel Segmentation,\u0026quot; in IEEE Journal of Biomedical and Health Informatics, vol. 27, o. 11, pp. 5542-5553, Nov. 2023, doi: 10.1109/JBHI.2023.3312338.\u003c/li\u003e\n\u003cli\u003eK. Aurangzeb, R. S. Alharthi, S. I. Haider and M. Alhussein, \u0026quot;Systematic Development of AI-Enabled Diagnostic Systems for Glaucoma and Diabetic Retinopathy,\u0026quot; in IEEE Access, vol. 11, pp. 105069-105081, 2023, doi: 10.1109/ACCESS.2023.3317348.\u003c/li\u003e\n\u003cli\u003eM. Monemian and H. Rabbani, \u0026quot;A Computationally Efficient Red-Lesion Extraction Method for Retinal Fundus Images,\u0026quot; in IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1-13, 2023, Art o. 5001613, doi: 10.1109/TIM.2022.3229712.\u003c/li\u003e\n\u003cli\u003eN. Shen et al., \u0026quot;SCANet: A Unified Semi-Supervised Learning Framework for Vessel Segmentation,\u0026quot; in IEEE Transactions on Medical Imaging, vol. 42, o. 9, pp. 2476-2489, Sept. 2023, doi: 10.1109/TMI.2022.3193150.\u003c/li\u003e\n\u003cli\u003eX. Xu et al., \u0026quot;AV-casNet: Fully Automatic Arteriole-Venule Segmentation and Differentiation in OCT Angiography,\u0026quot; in IEEE Transactions on Medical Imaging, vol. 42, o. 2, pp. 481-492, Feb. 2023, doi: 10.1109/TMI.2022.3214291.\u003c/li\u003e\n\u003cli\u003eA. -R. M. Al-Hinnawi, A. BaniMustafa, M. Al-Latayfeh and M. Tavakoli, \u0026quot;Reconstruction and Visualization of 5\u0026mu;m Sectional Coronal Views for Macula Vasculature in OptoVue OCTA,\u0026quot; in IEEE Access, vol. 11, pp. 28280-28293, 2023, doi: 10.1109/ACCESS.2023.3257720.\u003c/li\u003e\n\u003cli\u003eS. Sundar and S. Sumathy, \u0026quot;Classification of Diabetic Retinopathy Disease Levels by Extracting Topological Features Using Graph Neural Networks,\u0026quot; in IEEE Access, vol. 11, pp. 51435-51444, 2023, doi: 10.1109/ACCESS.2023.3279393.\u003c/li\u003e\n\u003cli\u003eM. Sadeghpour, A. Arakala, S. A. Davis and K. J. Horadam, \u0026quot;Protection of Sparse Retinal Templates Using Cohort-Based Dissimilarity Vectors,\u0026quot; in IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 5, o. 2, pp. 233-243, April 2023, doi: 10.1109/TBIOM.2023.3239866.\u003c/li\u003e\n\u003cli\u003eX. Zhang et al., \u0026quot;T-Net: Hierarchical Pyramid Network for Microaneurysm Detection in Retinal Fundus Image,\u0026quot; in IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1-13, 2023, Art o. 5019613, doi: 10.1109/TIM.2023.3286003.\u003c/li\u003e\n\u003cli\u003eS. Pang et al., \u0026quot;Beyond CNNs: Exploiting Further Inherent Symmetries in Medical Image Segmentation,\u0026quot; in IEEE Transactions on Cybernetics, vol. 53, o. 11, pp. 6776-6787, Nov. 2023, doi: 10.1109/TCYB.2022.3195447.\u003c/li\u003e\n\u003cli\u003eG. Ali, A. Dastgir, M. W. Iqbal, M. Anwar and M. Faheem, \u0026quot;A Hybrid Convolutional Neural Network Model for Automatic Diabetic Retinopathy Classification From Fundus Images,\u0026quot; in IEEE Journal of Translational Engineering in Health and Medicine, vol. 11, pp. 341-350, 2023, doi: 10.1109/JTEHM.2023.3282104.\u003c/li\u003e\n\u003cli\u003eM. A. K. Raiaan et al., \u0026quot;A Lightweight Robust Deep Learning Model Gained High Accuracy in Classifying a Wide Range of Diabetic Retinopathy Images,\u0026quot; in IEEE Access, vol. 11, pp. 42361-42388, 2023, doi: 10.1109/ACCESS.2023.3272228.\u003c/li\u003e\n\u003cli\u003eB. N. Jagadesh, M. G. Karthik, D. Siri, S. K. K. Shareef, S. V. Mantena and R. Vatambeti, \u0026quot;Segmentation Using the IC2T Model and Classification of Diabetic Retinopathy Using the Rock Hyrax Swarm-Based Coordination Attention Mechanism,\u0026quot; in IEEE Access, vol. 11, pp. 124441-124458, 2023, doi: 10.1109/ACCESS.2023.3330436.\u003c/li\u003e\n\u003cli\u003eM. Soni et al., \u0026quot;IoT-Based Federated Learning Model for Hypertensive Retinopathy Lesions Classification,\u0026quot; in IEEE Transactions on Computational Social Systems, vol. 10, o. 4, pp. 1722-1731, Aug. 2023, doi: 10.1109/TCSS.2022.3213507.\u003c/li\u003e\n\u003cli\u003eM. Vadduri and P. Kuppusamy, \u0026quot;Enhancing Ocular Healthcare: Deep Learning-Based Multi-Class Diabetic Eye Disease Segmentation and Classification,\u0026quot; in IEEE Access, vol. 11, pp. 137881-137898, 2023, doi: 10.1109/ACCESS.2023.3339574.\u003c/li\u003e\n\u003cli\u003eW. Nazih, A. O. Aseeri, O. Y. Atallah and S. El-Sappagh, \u0026quot;Vision Transformer Model for Predicting the Severity of Diabetic Retinopathy in Fundus Photography-Based Retina Images,\u0026quot; in IEEE Access, vol. 11, pp. 117546-117561, 2023, doi: 10.1109/ACCESS.2023.3326528.\u003c/li\u003e\n\u003cli\u003eH. Naz et al., \u0026quot;Ensembled Deep Convolutional Generative Adversarial Network for Grading Imbalanced Diabetic Retinopathy Recognition,\u0026quot; in IEEE Access, vol. 11, pp. 120554-120568, 2023, doi: 10.1109/ACCESS.2023.3327900.\u003c/li\u003e\n\u003cli\u003eW. K. Wong, F. H. Juwono and C. Apriono, \u0026quot;Diabetic Retinopathy Detection and Grading: A Transfer Learning Approach Using Simultaneous Parameter Optimization and Feature-Weighted ECOC Ensemble,\u0026quot; in IEEE Access, vol. 11, pp. 83004-83016, 2023, doi: 10.1109/ACCESS.2023.3301618.\u003c/li\u003e\n\u003cli\u003eM. Feng, J. Wang, K. Wen and J. Sun, \u0026quot;Grading of Diabetic Retinopathy Images Based on Graph Neural Network,\u0026quot; in IEEE Access, vol. 11, pp. 98391-98401, 2023, doi: 10.1109/ACCESS.2023.3312709.\u003c/li\u003e\n\u003cli\u003eM. Siebert, J. Gra\u0026szlig;hoff and P. Rostalski, \u0026quot;Uncertainty Analysis of Deep Kernel Learning Methods on Diabetic Retinopathy Grading,\u0026quot; in IEEE Access, vol. 11, pp. 146173-146184, 2023, doi: 10.1109/ACCESS.2023.3343642.\u003c/li\u003e\n\u003cli\u003eA. Kukkar et al., \u0026quot;Optimizing Deep Learning Model Parameters Using Socially Implemented IoMT Systems for Diabetic Retinopathy Classification Problem,\u0026quot; in IEEE Transactions on Computational Social Systems, vol. 10, o. 4, pp. 1654-1665, Aug. 2023, doi: 10.1109/TCSS.2022.3213369.\u003c/li\u003e\n\u003cli\u003eT. Palaniswamy and M. Vellingiri, \u0026quot;Internet of Things and Deep Learning Enabled Diabetic Retinopathy Diagnosis Using Retinal Fundus Images,\u0026quot; in IEEE Access, vol. 11, pp. 27590-27601, 2023, doi: 10.1109/ACCESS.2023.3257988.\u003c/li\u003e\n\u003cli\u003eX. Liu and W. Chi, \u0026quot;A Cross-Lesion Attention Network for Accurate Diabetic Retinopathy Grading With Fundus Images,\u0026quot; in IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1-12, 2023, Art o. 5029312, doi: 10.1109/TIM.2023.3322497.\u003c/li\u003e\n\u003cli\u003eN. J. Mohan, R. Murugan, T. Goel and P. Roy, \u0026quot;DRFL: Federated Learning in Diabetic Retinopathy Grading Using Fundus Images,\u0026quot; in IEEE Transactions on Parallel and Distributed Systems, vol. 34, o. 6, pp. 1789-1801, June 2023, doi: 10.1109/TPDS.2023.3264473.\u003c/li\u003e\n\u003cli\u003eQ. Hou, P. Cao, L. Jia, L. Chen, J. Yang and O. R. Zaiane, \u0026quot;Image Quality Assessment Guided Collaborative Learning of Image Enhancement and Classification for Diabetic Retinopathy Grading,\u0026quot; in IEEE Journal of Biomedical and Health Informatics, vol. 27, o. 3, pp. 1455-1466, March 2023, doi: 10.1109/JBHI.2022.3231276.\u003c/li\u003e\n\u003cli\u003eM. Nur-A-Alam, M. M. K. Nasir, M. Ahsan, M. A. Based, J. Haider and S. Palani, \u0026quot;A Faster RCNN-Based Diabetic Retinopathy Detection Method Using Fused Features From Retina Images,\u0026quot; in IEEE Access, vol. 11, pp. 124331-124349, 2023, doi: 10.1109/ACCESS.2023.3330104.\u003c/li\u003e\n\u003cli\u003eB. N. Kumar, T. R. Mahesh, G. Geetha and S. Guluwadi, \u0026quot;Redefining Retinal Lesion Segmentation: A Quantum Leap With DL-UNet Enhanced Auto Encoder-Decoder for Fundus Image Analysis,\u0026quot; in IEEE Access, vol. 11, pp. 70853-70864, 2023, doi: 10.1109/ACCESS.2023.3294443.\u003c/li\u003e\n\u003cli\u003eP. Zang et al., \u0026quot;Interpretable Diabetic Retinopathy Diagnosis Based on Biomarker Activation Map,\u0026quot; in IEEE Transactions on Biomedical Engineering, vol. 71, o. 1, pp. 14-25, Jan. 2024, doi: 10.1109/TBME.2023.3290541.\u003c/li\u003e\n\u003cli\u003eR. Liu et al., \u0026quot;TMM-Nets: Transferred Multi- to Mono-Modal Generation for Lupus Retinopathy Diagnosis,\u0026quot; in IEEE Transactions on Medical Imaging, vol. 42, o. 4, pp. 1083-1094, April 2023, doi: 10.1109/TMI.2022.3223683.\u003c/li\u003e\n\u003cli\u003eP. Zhang et al., \u0026quot;Image Quality Assessment of Diabetic Retinopathy Based on ADD-Net,\u0026quot; in IEEE Access, vol. 11, pp. 105130-105139, 2023, doi: 10.1109/ACCESS.2023.3318876.\u003c/li\u003e\n\u003cli\u003eK. Radha and Y. Karuna, \u0026quot;Modified Depthwise Parallel Attention UNet for Retinal Vessel Segmentation,\u0026quot; in IEEE Access, vol. 11, pp. 102572-102588, 2023, doi: 10.1109/ACCESS.2023.3317176.\u003c/li\u003e\n\u003cli\u003eA. Pereira, C. Santos, M. Aguiar, D. Welfer, M. Dias and M. Ribeiro, \u0026quot;Improved Detection of Fundus Lesions Using YOLOR-CSP Architecture and Slicing Aided Hyper Inference,\u0026quot; in IEEE Latin America Transactions, vol. 21, o. 7, pp. 806-813, July 2023, doi: 10.1109/TLA.2023.10244179.\u003c/li\u003e\n\u003cli\u003eM. Hussain, \u0026quot;Exudate Detection: Integrating Retinal-Based Affine Mapping and Design Flow Mechanism to Develop Lightweight Architectures,\u0026quot; in IEEE Access, vol. 11, pp. 125185-125203, 2023, doi: 10.1109/ACCESS.2023.3328386.\u003c/li\u003e\n\u003cli\u003eM. Yin et al., \u0026quot;Dual-Branch U-Net Architecture for Retinal Lesions Segmentation on Fundus Image,\u0026quot; in IEEE Access, vol. 11, pp. 130451-130465, 2023, doi: 10.1109/ACCESS.2023.3333364.\u003c/li\u003e\n\u003cli\u003eD. Bar-David et al., \u0026quot;Elastic Deformation of Optical Coherence Tomography Images of Diabetic Macular Edema for Deep-Learning Models Training: How Far to Go?,\u0026quot; in IEEE Journal of Translational Engineering in Health and Medicine, vol. 11, pp. 487-494, 2023, doi: 10.1109/JTEHM.2023.3294904.\u003c/li\u003e\n\u003cli\u003eP. R. Bhargav and N. B. Puhan, \u0026quot;Novel Contraharmonic Correlative Attention Loss for Microaneurysm Segmentation in Fundus Images,\u0026quot; in IEEE Sensors Letters, vol. 7, o. 7, pp. 1-4, July 2023, Art o. 7003504, doi: 10.1109/LSENS.2023.3290597.\u003c/li\u003e\n\u003cli\u003eAbdallah, M.B., Azar, A.T., Guedri, H. e\u003cem\u003et al.\u003c/em\u003e Correction to: Noise-estimation-based isotropic diffusion approach for retinal blood vessel segmentation. \u003cem\u003eNeural Comput \u0026amp; Applic\u003c/em\u003e\u003cstrong\u003e34\u003c/strong\u003e, 2499 (2022). https://doi.org/10.1007/s00521-021-06819-5\u003c/li\u003e\n\u003cli\u003eSingh, L.K., Khanna, M., Thawkar, S. e\u003cem\u003et al.\u003c/em\u003e Deep-learning based system for effective and automatic blood vessel segmentation from Retinal fundus images. \u003cem\u003eMultimed Tools Appl\u003c/em\u003e (2023). https://doi.org/10.1007/s11042-023-15348-3\u003c/li\u003e\n\u003cli\u003eKumar, K.S., Singh, N.P. Retinal disease prediction through blood vessel segmentation and classification using ensemble-based deep learning approaches. \u003cem\u003eNeural Comput \u0026amp; Applic\u003c/em\u003e\u003cstrong\u003e35\u003c/strong\u003e, 12495\u0026ndash;12511 (2023). https://doi.org/10.1007/s00521-023-08402-6\u003c/li\u003e\n\u003cli\u003eUpadhyay, K., Agrawal, M. \u0026amp; Vashist, P. Learning multi-scale deep fusion for retinal blood vessel nextraction in fundus images. \u003cem\u003eVis Comput\u003c/em\u003e\u003cstrong\u003e39\u003c/strong\u003e, 4445\u0026ndash;4457 (2023). https://doi.org/10.1007/s00371-022-02600-4\u003c/li\u003e\n\u003cli\u003eKumar, K.S., Singh, N.P. Segmentation of retinal blood vessel using generalized nextreme value probability distribution function(pdf)-based matched filter approach. \u003cem\u003ePattern Anal Applic\u003c/em\u003e\u003cstrong\u003e26\u003c/strong\u003e, 307\u0026ndash;332 (2023). https://doi.org/10.1007/s10044-022-01108-w\u003c/li\u003e\n\u003cli\u003eMahapatra, S., Jena, U.R. \u0026amp; Dash, S. Mean global based on hysteresis thresholding for retinal blood vessel segmentation using enhanced homomorphic filtering. \u003cem\u003eMultimed Tools Appl\u003c/em\u003e\u003cstrong\u003e81\u003c/strong\u003e, 41911\u0026ndash;41928 (2022). https://doi.org/10.1007/s11042-022-13517-4\u003c/li\u003e\n\u003cli\u003eSaranya, P., Prabakaran, S., Kumar, R. e\u003cem\u003et al.\u003c/em\u003e Blood vessel segmentation in retinal fundus images for proliferative diabetic retinopathy screening using deep learning. \u003cem\u003eVis Comput\u003c/em\u003e\u003cstrong\u003e38\u003c/strong\u003e, 977\u0026ndash;992 (2022). https://doi.org/10.1007/s00371-021-02062-0\u003c/li\u003e\n\u003cli\u003eGu, J., Tian, F. \u0026amp; Oh, IS. Retinal vessel segmentation based on self-distillation and implicit neural representation. \u003cem\u003eAppl Intell\u003c/em\u003e\u003cstrong\u003e53\u003c/strong\u003e, 15027\u0026ndash;15044 (2023). https://doi.org/10.1007/s10489-022-04252-2 \u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Retinal Image Segmentation, Attention Mechanisms, Adversarial Training, Ensemble Deep Learning, Diabetic Retinopathy","lastPublishedDoi":"10.21203/rs.3.rs-4553609/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4553609/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIt would, therefore, require highly advanced prediction tools to enhance early diagnosis and preemptive mechanisms for all these burgeoning diseases. Fast and correct disease prediction and pre-emption have huge potential for changing clinical outcome and ensuring timely and effective interventions that reduce morbidity and mortality. Current predictive models, instrumental as they are, have been found faltering in precision, recall, accuracy, and timeliness. Such delays and inaccuracies often miss the therapeutic window or lead to misguided clinical decisions. In this work, we present a novel model that aims to quite dramatically improve the process of segmentation and classification. Our approach embeds Attention Mechanisms with Adversarial Training and Ensemble Deep Learning Operations, together with a multimodal approach, which places it substantially higher across several metrics. This improves the precision, accuracy, recall, and AUC by 8.5%, 8.3%, 4.9%, and 3.9%, respectively, for segmentation and classification, while reducing the classification delay by 5.9% in different situations. Not only does our model handle the intrinsic limitations of current methods, but it also shows flexibility for a wide range of clinical applications. The compelling improvements in classification and preemption metrics strengthen its potential to make a sea change in the disease prediction framework for establishing optimum patient outcomes and efficient scenarios of healthcare delivery.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e","manuscriptTitle":"ERBVS: Enhanced Retinal Blood Vessel Segmentation using Multiple Modalities and Attention Mechanisms with Adversarial Training and Ensemble Deep Learning Operations","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-07-26 16:52:51","doi":"10.21203/rs.3.rs-4553609/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"78731d5f-353c-45a7-96cf-2c19198575b2","owner":[],"postedDate":"July 26th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-10-01T14:39:02+00:00","versionOfRecord":[],"versionCreatedAt":"2024-07-26 16:52:51","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4553609","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4553609","identity":"rs-4553609","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.