Quality controlling in capsule gastroduodenoscopy with less annotation via self-supervised learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Quality controlling in capsule gastroduodenoscopy with less annotation via self-supervised learning Yaqiong Zhang, Kai Zhang, Meijia Wang, Peng Bai, Shengqiang Wang, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6324281/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background It is possible to control the quality of capsule endoscopic images using artificial intelligence (AI), but it requires a great deal of time for labeling. Methods SimCLR (a simple framework for contrastive learning of visual representations), is capable of acquiring the inherent image representation with minimal annotation, but the feasibility is not studied. 62850 images were collected to train models. In internal cross-validation (more training data and less testing data) and reversed cross-validation (less training data and more testing data). Random forest and Xgboost (eXtreme Gradient Boosting) were used to finish the quality controlling after SimCLR extracting the features from images. Results SimCLR reported that the mean AUROC (Area Under the Receiver Operating Characteristic) curve exceeded 0.98 and 0.97. Moreover, Xgboost surpassed supervised CNN (Convolutional Neural Network). Extra 18636 pictures were gathered and the AUROC of SimCLR surpassed 0.93 (95% CI 0.9271–0.9548), which is close to supervised CNN (Convolutional Neural Network) (0.9645) in cross validation. Moreover, the AUROC of SimCLR surpass 0.96, which is better than supervised CNN (0.8374) in reversed cross validation. Conclusions Through SimCLR, the capsule endoscopic image quality control task can be completed with a performance similar to or better than that of supervised learning with fewer annotations. SimCLR quality controlling esophagogastroduodenoscopy Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Many digestive tract disorders 1 , 2 can be diagnosed with esophagogastroduodenoscopy, a common procedure in the department of gastroenterology 3 . Capsule endoscopes 4 , may generate a higher number of images than traditional endoscopes when using the automatic capture mode throughout the examination. Gastroenterologists should routinely offer massive endoscopy evaluations and documentation in their daily clinical work. Moreover, unqualified images (for example, blurry or overexposed images) are inevitable in the whole examination. Ensuring image quality 5 is a crucial and time-consuming responsibility for gastroenterologists, as it is a key component before providing automated diagnostic reports. Utilizing the vast collection of endoscopic images, artificial intelligence tools have the potential to alleviate physicians' tedious and repetitive tasks. In recent years, DL (deep learning) 6 , representated by convolutional neural networks (CNNs) 7 , 8 , have shown great potential in medical imaging field 9 , 10 in recent years, but training them requires a significant number of annotated images, placing a heavy burden on doctors in terms of time and effort. Self-supervised learning 11 , 12 is a emerging machine learning idea which applied a proxy task to learn the intrinsic representation of original dataset so that the target task can be completed under the assistance of the model from the proxy task with less annotation. In this study, we utilized SimCLR (a simple framework for contrastive learning of visual representations 13 ), a self-supervised algorithm, to perform the image quality control task on capsule endoscopic images. Furthermore, we evaluated how SimCLR performed in comparison to a supervised CNN using identical algorithmic hyperparameters setting (e.g., learning rate, batch size). The performance of SimCLR is close to supervised learning in both cross validation (more training and less validation data) and reversed cross validation (less training and more validation data). An illustration of overall experimental designs and the validation policy is displayed in Fig. 1 . Methods Approval of study design and ethical considerations It was conducted at Friendship Hospital in Beijing and Minhang Hospital in Shanghai. Both hospitals' Medical Ethics Committees (2022-P1-equipment-018-01and 2023-309-01X) gave their approval to the study protocol. Written consent was obtained from all patients for all datasets. Dataset collection and preprocessing Friendship Hospital retrospectively collected capsule endoscopic images for training and internal cross or reversed cross validation from March to October 2022. Images were obtained using a capsule endoscope (Pillbot C10000SI, Jiangsu CITRON Bio Technology Co. Ltd., Jiaxing, China) that is available for purchase and Friendship Hospital's imaging protocol. Capsule endoscopic images were collected by Friendship Hospital from November 2022 to February 2023 in a prospective manner. All images are from upper gastrointestinal tract and do not include any lesion. Black borders were removed from all images to remove the personal information of patient. Image labeling All images are saved in bitmap or jpeg format within image database and the quality of each image was labeled by a team consisted of healthcare professionals. The images of capsule endoscopic examinations that were blurred, overexposed, or covered in food were considered unqualified. Two gastroenterologists (YQ Zhang, P Bai) with a minimum of five-year of endoscope management experience firstly categorized all unqualified images and all remaining images as qualified images. Experienced senior doctors (P Li) resolved conflicts between two physicians. Development of deep learning models In order to develop and validate models, the retrospective dataset was used to randomly split patients into training and cross-validation or reversed cross-validation datasets, following four fold cross validation manner 14 . A cross validation technique that is not dependent on the subject will be employed to ensure that images from a single patient are not simultaneously divided into training and validation datasets 15 . The training workflow of SimCLR is shown in Fig. 2 , First, we trained SimCLR, which applies two types of image argumentation for one image and maximize the similarity between them. Furthermore, SimCLR minimizes the similarity between two images augmented from two different images and maximizes the similarity between two images augmented from one images (proxy task). In current study, colorjetter 16 and horizonal flip were selected as two types of data augmentation to train SimCLR. Then the backbone (ResNet101) 17 in SimCLR was used to extract the intrinsic features of images so that the features were used to train a model (random forest and Xgboost (eXtreme Gradient Boosting)) for our target task (image quality controlling). The performance of the model for target task can be a indirect reflection of the performance of SimCLR. We compared the performance of SimCLR using both supervised learning cross validation and reversed cross validation, as depicted in Fig. 1 . Random forest 18 and Xgboost 19 were selected to complete target task in both cross validation and reversed cross validation manners. The size of training dataset is large so that the training process costs much time, so we adopted principal component analysis (PCA 20 ) to reduce the dimension of the features. Furthermore, we evaluated the effectiveness of ResNet-101 21 . Due to the imbalance in the training dataset, we implemented a policy of assigning weights to different classes during the training process 22 . If the softmax layer (last layer) is given input corresponding to the original training dataset, \(\:\left\{\right({x}_{1},{y}_{1}),({x}_{2},{y}_{2})...({x}_{n},{y}_{n}\left)\right\},{x}_{i}\in\:{R}^{N},{y}_{i}\in\:\left\{\text{1,0}\right\}\) , then the loss is represented as Eq. 1 23 . $$\:L\left(w\right)=-\frac{1}{{\sum\:}_{i=1}^{m}{\sum\:}_{j=1}^{k}Class\_weigℎt\left\{{y}_{i}=j\right\}}\left[{\sum\:}_{i=1}^{m}{\sum\:}_{j=1}^{k}Class\_weigℎt\left\{{y}_{i}=j\right\}\ast\:{log}\frac{{e}^{{w}_{j}^{T}{x}_{i}}}{{\sum\:}_{s=1}^{k}{e}^{{w}_{s}^{T}{x}_{i}}}\right]+\frac{\lambda\:}{2}{\sum\:}_{i=1}^{k}{\sum\:}_{j=1}^{n}{w}_{ij}^{2}$$ 1 where, m, w, n and k represent the sample size in mini-batch, the weight in network undergoing training, the number of neurons of the layer in the layer prior to the softmax layer and the number of classes respectively. \(\:Class\_weigℎt\) and \(\:\frac{\lambda\:}{2}{\sum\:}_{i=1}^{k}{\sum\:}_{j=1}^{n}{w}_{ij}^{2}\) indicate the sample's weight for the class label and a penalization term to prevent overfitting. Balancing the effect between majority and minority classes could be achieved through this policy 24 . The best performing model from internal cross validation will be evaluated using a prospective validation dataset. All models were trained or tested using images resized to 512x512 in the workstation equipped with four NVIDIA RTX A4000 GPUs and PyTorch 1.8. We utilized a stochastic gradient descent (SGD) algorithm 25 to optimize the models. Learning rate was 0.01. Furthermore, there were 32 in each batch. Performance metrics The statistical analyses towards all performance metrics were conducted using MATLAB R2016a and Python 3.7.3. We evaluated the effectiveness of the DL models by analyzing accuracy, sensitivity, specificity, ROC (Area Under the Receiver Operating Characteristic) curves, and precision-recall curves 21 . Point estimation was utilized in the internal cross validation dataset to determine the confidence interval 26 , 27 . On the other hand, the confidence intervals (95% CI) for prospective validation dataset are calculated using the Wilson method. Results Characteristics of datasets A total of 81486 capsule endoscopic images from 677 patients were utilized to develop, validate and prospective validate the effectiveness of SimCLR and CNN (Convolutional Neural Network). Table 1 displays the basic demographic data of regular and capsule endoscopic pictures gathered from Friendship hospital. Table 1 The characteristics of all datasets in current study. Capsule endoscopic images Developing and cross validation dataset Prospective validation dataset No. of subjects 322 355 No. of images 62850 18636 Age (mean) 39.4 41.9 Sex Male: 182 Female: 140 Male: 229 Female: 126 Performance in internal cross validation The foundational supervised CNN (ResNet101) was also used to train for identifying unqualified images with supervised manner. Figure 3 (top two rows) displays their performance in internal cross validation. It shows that SimCLR's performance was comparable to supervised learning with the average AUROC difference between ResNet101 and random forest models (PCA dimension = 20, 30, 50; class weight = 30) being 0.0036, 0.0026, 0.0022, respectively. Similarly, the average AUROC difference between ResNet101 and random forest models (PCA dimension = 20, 30, 50; class weight = 50) was 0.0036, 0.0026, 0.0022, respectively (Table 2 ). Xgboost outperformed random forest, with the average AUROC difference between ResNet101 and Xgboost models (with PCA dimensions of 20, 30, 50) being 0.0014, 0.0001, and 0.0014, as shown in Table 2 . Xgboost, which was closer to the performance of supervised learning, performed better than random forest. Table 2 displays additional metrics, indicating that the supervised learning model achieved the highest accuracy, with SimCLR closely to the performance of supervised learning. The main difference lies in the sensitivity of the two models. Table 2 The performance of SimCLR and supervised CNN in internal cross validation (Mean value ± standard deviation [95% confidence interval]). Model Class weight PCA dimension Accuracy Sensitivity Specificity AUROC Random forest [1,30] 20 0.9506 ± 0.0220 [0.9075–0.9937] 0.8629 ± 0.1236 [0.6207-1.0000] 0.9678 ± 0.0414 [0.8867-1.0000] 0.9802 ± 0.0188 [0.9434-1.0000] 30 0.9505 ± 0.0236 [0.9043–0.9967] 0.8676 ± 0.1068 [0.6582-1.0000] 0.9679 ± 0.0410 [0.8875-1.0000] 0.9810 ± 0.0178 [0.9461-1.0000] 50 0.9420 ± 0.0289 [0.8854–0.9987] 0.8467 ± 0.1133 [0.6247-1.0000] 0.9690 ± 0.0406 [0.8895-1.0000] 0.9816 ± 0.0174 [0.9475-1.0000] [1,50] 20 0.9501 ± 0.0223 [0.9064–0.9937] 0.8621 ± 0.1215 [0.6239-1.0000] 0.9679 ± 0.0413 [0.8870-1.0000] 0.9802 ± 0.0187 [0.9436-1.0000] 30 0.9480 ± 0.0245 [0.8999–0.9961] 0.8624 ± 0.1081 [0.6506-1.0000] 0.9688 ± 0.0418 [0.8869-1.0000] 0.9810 ± 0.0176 [0.9465-1.0000] 50 0.9402 ± 0.0316 [0.8783-1.0000] 0.8442 ± 0.1114 [0.6258-1.0000] 0.9689 ± 0.0406 [0.8894-1.0000] 0.9816 ± 0.0172 [0.9479-1.0000] Xgboost N/A 20 0.9656 ± 0.0202 [0.9261-1.0000] 0.9104 ± 0.1114 [0.6920-1.0000] 0.9648 ± 0.0417 [0.8830-1.0000] 0.9824 ± 0.0144 [0.9542-1.0000] 30 0.9685 ± 0.0203 [0.9288-1.0000] 0.9265 ± 0.0898 [0.7506-1.0000] 0.9652 ± 0.0417 [0.8835-1.0000] 0.9837 ± 0.0149 [0.9545-1.0000] 50 0.9707 ± 0.0206 [0.9304-1.0000] 0.9302 ± 0.0936 [0.7467-1.0000] 0.9654 ± 0.0419 [0.8832-1.0000] 0.9852 ± 0.0137 [0.9583-1.0000] ResNet101 (supervised) N/A 0.9734 ± 0.0208 [0.9327-1.0000] 0.9598 ± 0.0583 [0.8455-1.0000] 0.9610 ± 0.0428 [0.8771-1.0000] 0.9838 ± 0.0134 [0.9575-1.0000] Performance in prospective validation (cross validation) The performance of SimCLR and ResNet101 was shown in Fig. 3 (bottom two rows), which shows SimCLR was closer to the performance of supervised learning, the difference is not so significant than internal cross validation in accuracy and sensitivity. Compared to supervised CNN, random forest models with PCA dimensions of 20, 30, and 50 and class weight of 50 had AUROC differences of 0.0294, 0.0141, and 0.0137. In Table 3 , the AUROC gap between random forest models with PCA dimensions of 20, 30, and 50 and supervised CNN was 0.0332, 0.0140, and 0.0181, correspondingly. Xgboost outperformed random forest, with the average AUROC difference between ResNet101 and Xgboost (with PCA dimensions of 20, 30, and 50) being 0.0073, 0.0051, and 0.0004, as shown in Table 3 . Even the accuracies of random forest and Xgboost models are better than supervised learning CNN, especially, Xgboost performed better than CNN in all metrics with a proper PCA dimension (30). Table 3 The performance of SimCLR and CNN in prospective validation (cross validation). Model Class weight PCA dimension Accuracy Sensitivity Specificity AUROC Random forest [1,30] 20 0.9367 [0.9324–0.9407] 0.8052 [0.7394–0.8420] 0.9375 [0.9332–0.9415] 0.9351 [0.9308–0.9391] 30 0.9370 [0.9327–0.9410] 0.8182 [0.7515–0.8547] 0.9378 [0.9335–0.9418] 0.9504 [0.9461–0.9544] 50 0.9417 [0.9374–0.9457] 0.7662 [0.7032–0.8039] 0.9429 [0.9386–0.9469] 0.9508 [0.9465–0.9548] [1,50] 20 0.9368 [0.9325–0.9408] 0.8052 [0.7394–0.8420] 0.9376 [0.9333–0.9416] 0.9313 [0.9271–0.9353] 30 0.9371 [0.9328–0.9411] 0.8182 [0.7515–0.8547] 0.9379 [0.9336–0.9419] 0.9505 [0.9462–0.9545] 50 0.9374 [0.9331–0.9414] 0.8182 [0.7515–0.8547] 0.9382 [0.9339–0.9422] 0.9474 [0.9431–0.9514] Xgboost N/A 20 0.9359 [0.9316–0.9399] 0.8701 [0.7996–0.9054] 0.9363 [0.9320–0.9403] 0.9572 [0.9529–0.9612] 30 0.9361 [0.9320–0.9399] 0.8831 [0.8675–0.8957] 0.9364 [0.9321–0.9404] 0.9696 [0.9654–0.9735] 50 0.9350 [0.9307–0.9390] 0.8831 [0.8117–0.9181] 0.9353 [0.9310–0.9393] 0.9641 [0.9598–0.9681] ResNet101 (supervised) N/A 0.9333 [0.9290–0.9373] 0.8831 [0.8117–0.9181] 0.9336 [0.9293–0.9376] 0.9645 [0.9602–0.9685] Performance in internal reversed cross validation In order to test whether SimCLR can complete the same task with less annotation, we used reversed cross validation (less training data and more testing data) to fairly compare the performance of SimCLR and supervised CNN. Figure 4 (top row) displays the results of their performance in internal reversed cross validation. The comparison of SimCLR's performance to supervised learning showed that the average AUROC difference between ResNet101 and the random forest model was 0.0235, while the average AUROC difference between ResNet101 and the Xgboost model was 0.0125 (Table 2 ). Table 4 displays additional metrics, indicating that while the accuracy of supervised ResNet101 is the highest, the AUROC values of random forest and Xgboost models are similar to that of supervised learning ResNet101. Table 4 The performance of SimCLR and supervised CNN in internal reversed cross validation (Mean value ± standard deviation [95% confidence interval]). Method Accuracy Sensitivity Specificity AUROC Random forest 0.8409 ± 0.2113 [0.4268-1.0000] 0.6862 ± 0.4388 [0.0000–1.0000] 0.9796 ± 0.0178 [0.9448-1.0000] 0.9824 ± 0.0096 [0.9635-1.0000] Xgboost 0.8489 ± 0.2107 [0.4359-1.0000] 0.7093 ± 0.4390 [0.0000–1.0000] 0.9769 ± 0.0177 [0.9422-1.0000] 0.9714 ± 0.0305 [0.9116-1.0000] ResNet101 (supervised) 0.9048 ± 0.1305 [0.6489-1.0000] 0.8521 ± 0.2323 [0.3969-1.0000] 0.9533 ± 0.0333 [0.8881-1.0000] 0.9589 ± 0.0395 [0.8816-1.0000] Performance in prospective validation (reversed cross validation) The performance of SimCLR and ResNet101 is shown in Fig. 4 (bottom row), which shows SimCLR was better than the performance of supervised learning. The AUROC discrepancy between the random forest and the supervised CNN was 0.1256. The AUROC discrepancy between the Xgboost and the supervised CNN was 0.1353. Table 5 displays additional metrics, indicating that SimCLR performed exceptionally well and the image features extracted can be utilized for downstream tasks. Moreover, Xgboost model is better than random forest model and supervised learning CNN. Table 5 The performance of SimCLR and CNN in prospective validation (reversed cross validation). Method Accuracy Sensitivity Specificity AUROC Random forest 0.9445 [0.9402–0.9485] 0.8312 [0.7635–0.8674] 0.9452 [0.9409–0.9492] 0.9630 [0.9587–0.9670] Xgboost 0.9390 [0.9347–0.9430] 0.8831 [0.8117–0.9181] 0.9394 [0.9351–0.9434] 0.9727 [0.9684–0.9767] ResNet101 (supervised) 0.8924 [0.8883–0.8963] 0.7273 [0.6672–0.7658] 0.8934 [0.8892–0.8973] 0.8374 [0.8334–0.8412] Discussion The constant need for annotations places a significant strain on healthcare workers and doctors. Self-supervised learning (SimCLR) can complete image quality controlling task with less annotation for capsule endoscopic images. SimCLR outperformed supervised learning in both cross validation and reversed cross validation, and also achieved better results than supervised CNN using fewer annotated data. The annotation burden of medical staffs and physicians can be alleviated largely. Sharib Ali and colleagues 28 introduced a framework that evaluates the quality of endoscopic images through the identification and segmentation of artifacts like blurriness, reflections, color intensity, air bubbles, contrast, and other imperfections artefacts. Four object detection algorithms were compared, including YOLOv3, RetinaNet, YOLOv3-spp, and Faster-RCNN, with YOLOv3-spp demonstrating the highest overall performance. Additionally, they evaluated the performance of four different semantic segmentation models (ResNet-UNet, DeepLabv3+, FCN8, pspNet) in segmenting all artifacts, with DeepLabv3 + performing best. Finally, they created a rating system that is highly correlated with three experts, with a correlation exceeding 0.6. He et al. 29 conducted a study where 3704 endoscopic images were gathered from 211 patients to train Inception-v3, ResNet-50, VGG-16-bn, DenseNet-121 and VGG-11-bn models in order to classify the images as either qualified or unqualified. According to VGG-11-bn, unqualified images had the highest accuracy of 67.6%, while qualified images had the highest accuracy of 80.9%. Lianlian Wu et al. 30 trained a deep CNN using 12220 in vitro, 25222 in vivo, and 16760 unqualified endoscopic images to determine if the endoscopic image contained the unrelated content (out of body) during esophagogastroduodenoscopy. The result shows that the accuracies for each class are higher than 70%. In addition, our team also utilized transfer learning (domain adaptation) from standard endoscopic images to capsule endoscopic images in order to manage the image quality. The findings indicated that domain adaptation could successfully accomplish this task using annotations from just a single domain, Moreover, the performance surpassed the supervised learning 31 . We also compared CNN with vision transformer (ViT) in quality controlling for capsule and conventional endoscopic images. The result indicated that CNN outperformed ViT (Vision Transformer) in completing this task 32 . By limiting the selection to capsule endoscope, there are some drawbacks, like not taking into account the distinctions between different brands. Due to our failure to classify all eligible conditions, the distribution of images in actual clinical situations may not be uniform. Our image quality control model can be seamlessly integrated into the workflow of the auto image reading system, whether video or image analysis is conducted in real-time or during post-processing. It is expected that the examining process would capture a large number of images, so the unqualified images might be sifted out in order to save more time. A new model can be trained for a healthcare facility by incorporating data from a new source (healthcare center) and leveraging the existing research model as a pretrained model. Conclusions Utilizing SimCLR, intrinsic characteristics can be extracted from standard endoscopic images, enabling machine learning algorithms to identify and filter out inadequate images with minimal manual labeling. This not only reduces the daily tasks of doctors, but also decreases the amount of annotations needed for medical artificial intelligence. Declarations Code and Data Availability: Access to the Python scripts necessary for the analysis can be found at https://github.com/Hugo0512/ASI4EGD, while the source code for SimCLR can be found at https://github.com/Spijkervet/SimCLR. The information and resources in this research can be obtained from the corresponding author upon request. Acknowledgements: None Fundings : The project received funding from various sources including the General Program of National Natural Science Foundation of China (8217100675), University Department Construction Project in Endoscope Center Area (2020MWDXK03); Shanghai Key Medical College Construction Project (ZK2019C05); Minhang District Leading Talent Project (2021LJRC03); Natural Science Basic Research Plan in Shaanxi Province of China (2022JQ-175), Scientific Research Plan of Shaanxi Education Department (22JK0303) and National Major Scientific Instruments and Equipments Development Project of National Natural Science Foundation of China (82027801). The sponsor or funding organization had no role in the design or conduct of this research. Author Contributions : Y.Q.Z. helped shape the study's framework. F.H. and P.L. critically reviewed the manuscript. Y.Q.Z., and K.Z. conducted the research and performed the literature review. T.M., Y.Q.Z., and P.L. gathered the information and categorized it. Y.Q.Z., K.Z., P.B., S.Q.W., and M.J.W. assisted in developing the statistical analysis and interpreting the data. Y.Q.Z., P.B. and K.Z. drafted the manuscript. M.J.W., and P.L. provided research funding. Y.Q.Z., K.Z., and F.H. managed the research project. All authors had access to all the raw datasets and the corresponding authors (T.M., F.H., and P.L.) confirming the data.The final manuscript was reviewed and approved by all authors. Consent for publication: : The authors affirm that all content and resources included in the manuscript are authentic. Ethics approval and consent to participat : This study complied with the Declaration of Helsinki. The research was approved by the Medical Ethics Committee of Friendship hospital (2022-P1-equipment-018-01) and Minhang hospital (2023-009-01X). The informed consent was signed by all patients. Competing interests : None Declaration of competing interest: No conflicting relationship exists for any author. References Shorten C, Khoshgoftaar TM. A survey on image data augmentation for deep learning. J big data. 2019;6:1–48. Zhang M, et al. Differential diagnosis for esophageal protruded lesions using a deep convolution neural network in endoscopic images. Gastrointest Endosc. 2021;93:1261–72. Chen D, et al. Comparing blind spots of unsedated ultrafine, sedated, and unsedated conventional gastroscopy with and without artificial intelligence: a prospective, single-blind, 3-parallel-group, randomized, single-center trial. Gastrointest Endosc. 2020;91:332–9. e333. Liao Z, Zou W, Li Z-S. Clinical application of magnetically controlled capsule gastroscopy in gastric disease diagnosis: recent advances. Sci China Life Sci. 2018;61:1304–9. Chow LS, Paramesran R. Review of medical image quality assessment. Biomed Signal Process Control. 2016;27:145–54. Zhang S, et al. The applications of artificial intelligence in digestive system neoplasms: a review. Health Data Sci. 2023;3:0005. Sharma P, Hassan C. Artificial intelligence and deep learning for upper gastrointestinal neoplasia. Gastroenterology. 2022;162:1056–66. Luo H, et al. Real-time artificial intelligence for detection of upper gastrointestinal cancer by endoscopy: a multicentre, case-control, diagnostic study. Lancet Oncol. 2019;20:1645–54. Luo J, et al. A deep learning method to assist with chronic atrophic gastritis diagnosis using white light images. Dig Liver Disease. 2022;54:1513–9. Wang P, et al. Development and validation of a deep-learning algorithm for the detection of polyps during colonoscopy. Nat biomedical Eng. 2018;2:741–8. Dahmardeh M, Dastjerdi M, Mazal H, Köstler H, H., Sandoghdar V. Self-supervised machine learning pushes the sensitivity limit in label-free detection of single proteins below 10 kDa. Nat Methods, 1–6 (2023). Hendrycks D, Mazeika M, Kadavath S, Song D. Using self-supervised learning can improve model robustness and uncertainty. Adv Neural Inf Process Syst 32 (2019). Chen T, Kornblith S, Norouzi M, Hinton G. in International conference on machine learning. 1597–1607 (PMLR). Zhou W-D, et al. Deep Learning for Automatic Detection of Recurrent Retinal Detachment after Surgery Using Ultra-Widefield Fundus Images: A Single‐Center Study. Adv Intell Syst. 2022;4:2200067. Hui S, et al. Noninvasive identification of Benign and malignant eyelid tumors using clinical images via deep learning system. J Big Data. 2022;9:84. Zhang M, et al. Computerized assisted evaluation system for canine cardiomegaly via key points detection with deep learning. Prev Vet Med. 2021;193:105399. Zhang H, et al. Validation of the relationship between iris color and uveal melanoma using artificial intelligence with multiple paths in a large Chinese population. Front Cell Dev Biology. 2021;9:713209. Zhang K, et al. Prediction of postoperative complications of pediatric cataract patients using data mining. J translational Med. 2019;17:1–10. Gangil T, et al. Predicting clinical outcomes of radiotherapy for head and neck squamous cell carcinoma patients using machine learning algorithms. J Big Data. 2022;9:1–19. Daffertshofer A, Lamoth CJ, Meijer OG, Beek P. J. PCA in studying coordination and variability: a tutorial. Clin Biomech Elsevier Ltd. 2004;19:415–28. Li W, et al. Dense anatomical annotation of slit-lamp images improves the performance of deep learning for the diagnosis of ophthalmic disorders. Nat Biomedical Eng. 2020;4:767–77. Sampath V, Maurtua I, Martin A, J. J., Gutierrez A. A survey on generative adversarial networks for imbalance problems in computer vision tasks. J big Data. 2021;8:1–59. Zhang R, et al. Automatic retinoblastoma screening and surveillance using deep learning. Br J Cancer. 2023;129:466–74. 10.1038/s41416-023-02320-z . Zhang K, et al. An interpretable and expandable deep learning diagnostic system for multiple ocular diseases: qualitative study. J Med Internet Res. 2018;20:e11144. Zinkevich M, Weimer M, Li L, Smola A. Parallelized stochastic gradient descent. Adv Neural Inf Process Syst 23 (2010). Zhang K, et al. A human-in-the-loop deep learning paradigm for synergic visual evaluation in children. Neural Netw. 2020;122:163–73. Schmidt H, Spieker AJ, Luo T, Szymczak JE, Grande D. in JAMA Health Forum e212932–212932 (American Medical Association). Ali S, et al. A deep learning framework for quality assessment and restoration in video endoscopy. Med Image Anal. 2021;68:101900. He Q, et al. Deep learning-based anatomical site classification for upper gastrointestinal endoscopy. Int J Comput Assist Radiol Surg. 2020;15:1085–94. Wu L, et al. Randomised controlled trial of WISENSE, a real-time quality improving system for monitoring blind spots during esophagogastroduodenoscopy. Gut. 2019;68:2161–9. Zhang Y et al. Deep transfer learning from ordinary to capsule esophagogastroduodenoscopy for image quality controlling. Eng Rep n/a, e12776, https://doi.org/10.1002/eng2.12776 Zhang K, et al. Anatomical sites identification in both ordinary and capsule gastroduodenoscopy via deep learning. Biomed Signal Process Control. 2024;90:105911. https://doi.org/10.1016/j.bspc.2023.105911 . Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6324281","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":452802553,"identity":"329cd714-0382-4245-a8e9-f9e801f868bf","order_by":0,"name":"Yaqiong Zhang","email":"","orcid":"","institution":"Shaanxi Provincial People's Hospital","correspondingAuthor":false,"prefix":"","firstName":"Yaqiong","middleName":"","lastName":"Zhang","suffix":""},{"id":452802554,"identity":"7e381746-159c-4100-966a-31b5300de9d6","order_by":1,"name":"Kai Zhang","email":"","orcid":"","institution":"Zhejiang Citron Robotics Technology (Group) Co., Ltd","correspondingAuthor":false,"prefix":"","firstName":"Kai","middleName":"","lastName":"Zhang","suffix":""},{"id":452802555,"identity":"16212ecc-d391-42d7-a7d2-577139c240cd","order_by":2,"name":"Meijia Wang","email":"","orcid":"","institution":"Shaanxi University of Science \u0026 Technology, Xi'an Weiyang University Park","correspondingAuthor":false,"prefix":"","firstName":"Meijia","middleName":"","lastName":"Wang","suffix":""},{"id":452802556,"identity":"28b349de-db1e-487f-b592-50f7de1f5bc8","order_by":3,"name":"Peng Bai","email":"","orcid":"","institution":"Affiliated Hospital of Shanxi University of Chinese Medicine, Shanxi University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Peng","middleName":"","lastName":"Bai","suffix":""},{"id":452802557,"identity":"50289e32-66f6-4961-bbc0-08bdbd40aaad","order_by":4,"name":"Shengqiang Wang","email":"","orcid":"","institution":"Shaanxi University of Science \u0026 Technology, Xi'an Weiyang University Park","correspondingAuthor":false,"prefix":"","firstName":"Shengqiang","middleName":"","lastName":"Wang","suffix":""},{"id":452802558,"identity":"a384360e-c39d-4be8-ac8f-f835814d670f","order_by":5,"name":"Ting Ma","email":"","orcid":"","institution":"Zhejiang Citron Robotics Technology (Group) Co., Ltd","correspondingAuthor":false,"prefix":"","firstName":"Ting","middleName":"","lastName":"Ma","suffix":""},{"id":452802559,"identity":"f4c0c86d-4327-4a4d-9943-45bcd774cb1a","order_by":6,"name":"Feng Hu","email":"","orcid":"","institution":"Zhejiang Citron Robotics Technology (Group) Co., Ltd","correspondingAuthor":false,"prefix":"","firstName":"Feng","middleName":"","lastName":"Hu","suffix":""},{"id":452802560,"identity":"0796e397-8e98-481a-841b-bb590f078eab","order_by":7,"name":"Peng Li","email":"","orcid":"","institution":"Capital Medical University","correspondingAuthor":false,"prefix":"","firstName":"Peng","middleName":"","lastName":"Li","suffix":""},{"id":452802561,"identity":"82174213-24e5-45f6-9063-756198657553","order_by":8,"name":"Guisheng Liu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5klEQVRIie3PMWvCQBTA8YPAuTyM4wuIfoVrh+AQ8Ku8I5ApQscMDoGKGdT2qzg6Xhq46dI5Q4eEQnc3XcQ6t/Ti1uF+0xvuz7vHmOP8Q3zw2rXyEoF/Gyhb2pMhGE+0PBkH29tgtD2ZYMqDlr9FovkeupXX42NQayRQINRCZzLnzC82ZLnlJZ4RfkCg3pNGHsYMTb23bFEPDYkvGJZ52EjDmcCFJUESSFQBq1j4JNdenyR9RFIVjDSErF8COhYyTyDY8hjJaLDeMi2ey+6cR3N/+lkeT9ly4he7v5Mf4L7njuM4zq+uJBdNNqeDPlkAAAAASUVORK5CYII=","orcid":"","institution":"Shaanxi Provincial People's Hospital","correspondingAuthor":true,"prefix":"","firstName":"Guisheng","middleName":"","lastName":"Liu","suffix":""}],"badges":[],"createdAt":"2025-03-28 02:53:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6324281/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6324281/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":82360142,"identity":"fce8dffc-fea1-49d2-8e2f-7bc41f4c2f4e","added_by":"auto","created_at":"2025-05-09 11:37:59","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":431691,"visible":true,"origin":"","legend":"\u003cp\u003eThe framework for image quality controlling and the experiment settings.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6324281/v1/8c16a05024f53bd353464efb.jpeg"},{"id":82357321,"identity":"35cc1d54-9dbc-488f-a7a4-4c1bdd2c5c80","added_by":"auto","created_at":"2025-05-09 11:21:59","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":321002,"visible":true,"origin":"","legend":"\u003cp\u003eThe working flow of SimCLR.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6324281/v1/e330054b8545767108fa6e35.jpeg"},{"id":82359361,"identity":"76bc18a6-822a-4379-bb38-4e665c549d95","added_by":"auto","created_at":"2025-05-09 11:29:59","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":44566,"visible":true,"origin":"","legend":"\u003cp\u003eThe ROC and PR curves of SimCLR and ResNet101 in internal cross validation and prospective validation.\u003c/p\u003e","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6324281/v1/085288175624566cfe442da3.png"},{"id":82359360,"identity":"10dc621e-379c-439a-a361-3081b2a27385","added_by":"auto","created_at":"2025-05-09 11:29:59","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":366088,"visible":true,"origin":"","legend":"\u003cp\u003eThe ROC and PR curves of SimCLR and supervised ResNet101 in reversed cross validation and prospective validation.\u003c/p\u003e","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6324281/v1/3b020f8b4225a2ee6f9b24aa.jpeg"},{"id":104694040,"identity":"e5bed21b-fafd-46f9-9b15-0df2ba7ef39c","added_by":"auto","created_at":"2026-03-16 06:57:39","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2109248,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6324281/v1/b09d498a-db56-414e-bbf4-2f0f35353f68.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Quality controlling in capsule gastroduodenoscopy with less annotation via self-supervised learning","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMany digestive tract disorders\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e can be diagnosed with esophagogastroduodenoscopy, a common procedure in the department of gastroenterology\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Capsule endoscopes\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e, may generate a higher number of images than traditional endoscopes when using the automatic capture mode throughout the examination. Gastroenterologists should routinely offer massive endoscopy evaluations and documentation in their daily clinical work. Moreover, unqualified images (for example, blurry or overexposed images) are inevitable in the whole examination.\u003c/p\u003e \u003cp\u003eEnsuring image quality\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e is a crucial and time-consuming responsibility for gastroenterologists, as it is a key component before providing automated diagnostic reports. Utilizing the vast collection of endoscopic images, artificial intelligence tools have the potential to alleviate physicians' tedious and repetitive tasks. In recent years, DL (deep learning)\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, representated by convolutional neural networks (CNNs)\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e, have shown great potential in medical imaging field\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e in recent years, but training them requires a significant number of annotated images, placing a heavy burden on doctors in terms of time and effort. Self-supervised learning\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e,\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e is a emerging machine learning idea which applied a proxy task to learn the intrinsic representation of original dataset so that the target task can be completed under the assistance of the model from the proxy task with less annotation.\u003c/p\u003e \u003cp\u003eIn this study, we utilized SimCLR (a simple framework for contrastive learning of visual representations\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e), a self-supervised algorithm, to perform the image quality control task on capsule endoscopic images. Furthermore, we evaluated how SimCLR performed in comparison to a supervised CNN using identical algorithmic hyperparameters setting (e.g., learning rate, batch size). The performance of SimCLR is close to supervised learning in both cross validation (more training and less validation data) and reversed cross validation (less training and more validation data). An illustration of overall experimental designs and the validation policy is displayed in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eApproval of study design and ethical considerations\u003c/h2\u003e \u003cp\u003eIt was conducted at Friendship Hospital in Beijing and Minhang Hospital in Shanghai. Both hospitals' Medical Ethics Committees (2022-P1-equipment-018-01and 2023-309-01X) gave their approval to the study protocol. Written consent was obtained from all patients for all datasets.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eDataset collection and preprocessing\u003c/h3\u003e\n\u003cp\u003eFriendship Hospital retrospectively collected capsule endoscopic images for training and internal cross or reversed cross validation from March to October 2022. Images were obtained using a capsule endoscope (Pillbot C10000SI, Jiangsu CITRON Bio Technology Co. Ltd., Jiaxing, China) that is available for purchase and Friendship Hospital's imaging protocol. Capsule endoscopic images were collected by Friendship Hospital from November 2022 to February 2023 in a prospective manner. All images are from upper gastrointestinal tract and do not include any lesion. Black borders were removed from all images to remove the personal information of patient.\u003c/p\u003e\n\u003ch3\u003eImage labeling\u003c/h3\u003e\n\u003cp\u003eAll images are saved in bitmap or jpeg format within image database and the quality of each image was labeled by a team consisted of healthcare professionals. The images of capsule endoscopic examinations that were blurred, overexposed, or covered in food were considered unqualified. Two gastroenterologists (YQ Zhang, P Bai) with a minimum of five-year of endoscope management experience firstly categorized all unqualified images and all remaining images as qualified images. Experienced senior doctors (P Li) resolved conflicts between two physicians.\u003c/p\u003e\n\u003ch3\u003eDevelopment of deep learning models\u003c/h3\u003e\n\u003cp\u003eIn order to develop and validate models, the retrospective dataset was used to randomly split patients into training and cross-validation or reversed cross-validation datasets, following four fold cross validation manner\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. A cross validation technique that is not dependent on the subject will be employed to ensure that images from a single patient are not simultaneously divided into training and validation datasets\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. The training workflow of SimCLR is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, First, we trained SimCLR, which applies two types of image argumentation for one image and maximize the similarity between them. Furthermore, SimCLR minimizes the similarity between two images augmented from two different images and maximizes the similarity between two images augmented from one images (proxy task). In current study, colorjetter\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e and horizonal flip were selected as two types of data augmentation to train SimCLR. Then the backbone (ResNet101)\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e in SimCLR was used to extract the intrinsic features of images so that the features were used to train a model (random forest and Xgboost (eXtreme Gradient Boosting)) for our target task (image quality controlling). The performance of the model for target task can be a indirect reflection of the performance of SimCLR. We compared the performance of SimCLR using both supervised learning cross validation and reversed cross validation, as depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Random forest\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e and Xgboost\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e were selected to complete target task in both cross validation and reversed cross validation manners. The size of training dataset is large so that the training process costs much time, so we adopted principal component analysis (PCA\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e) to reduce the dimension of the features. Furthermore, we evaluated the effectiveness of ResNet-101\u003csup\u003e21\u003c/sup\u003e. Due to the imbalance in the training dataset, we implemented a policy of assigning weights to different classes during the training process\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. If the softmax layer (last layer) is given input corresponding to the original training dataset, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left\\{\\right({x}_{1},{y}_{1}),({x}_{2},{y}_{2})...({x}_{n},{y}_{n}\\left)\\right\\},{x}_{i}\\in\\:{R}^{N},{y}_{i}\\in\\:\\left\\{\\text{1,0}\\right\\}\\)\u003c/span\u003e\u003c/span\u003e, then the loss is represented as Eq.\u0026nbsp;1\u003csup\u003e23\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv id=\"Equ1\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$\\:L\\left(w\\right)=-\\frac{1}{{\\sum\\:}_{i=1}^{m}{\\sum\\:}_{j=1}^{k}Class\\_weigℎt\\left\\{{y}_{i}=j\\right\\}}\\left[{\\sum\\:}_{i=1}^{m}{\\sum\\:}_{j=1}^{k}Class\\_weigℎt\\left\\{{y}_{i}=j\\right\\}\\ast\\:{log}\\frac{{e}^{{w}_{j}^{T}{x}_{i}}}{{\\sum\\:}_{s=1}^{k}{e}^{{w}_{s}^{T}{x}_{i}}}\\right]+\\frac{\\lambda\\:}{2}{\\sum\\:}_{i=1}^{k}{\\sum\\:}_{j=1}^{n}{w}_{ij}^{2}$$\u003c/div\u003e \u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003ewhere, \u003cem\u003em, w, n\u003c/em\u003e and \u003cem\u003ek\u003c/em\u003e represent the sample size in mini-batch, the weight in network undergoing training, the number of neurons of the layer in the layer prior to the softmax layer and the number of classes respectively. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:Class\\_weigℎt\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\frac{\\lambda\\:}{2}{\\sum\\:}_{i=1}^{k}{\\sum\\:}_{j=1}^{n}{w}_{ij}^{2}\\)\u003c/span\u003e\u003c/span\u003e indicate the sample's weight for the class label and a penalization term to prevent overfitting. Balancing the effect between majority and minority classes could be achieved through this policy\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. The best performing model from internal cross validation will be evaluated using a prospective validation dataset.\u003c/p\u003e \u003cp\u003eAll models were trained or tested using images resized to 512x512 in the workstation equipped with four NVIDIA RTX A4000 GPUs and PyTorch 1.8. We utilized a stochastic gradient descent (SGD) algorithm\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e to optimize the models. Learning rate was 0.01. Furthermore, there were 32 in each batch.\u003c/p\u003e\n\u003ch3\u003ePerformance metrics\u003c/h3\u003e\n\u003cp\u003eThe statistical analyses towards all performance metrics were conducted using MATLAB R2016a and Python 3.7.3. We evaluated the effectiveness of the DL models by analyzing accuracy, sensitivity, specificity, ROC (Area Under the Receiver Operating Characteristic) curves, and precision-recall curves\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. Point estimation was utilized in the internal cross validation dataset to determine the confidence interval\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e,\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. On the other hand, the confidence intervals (95% CI) for prospective validation dataset are calculated using the Wilson method.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eCharacteristics of datasets\u003c/h2\u003e \u003cp\u003eA total of 81486 capsule endoscopic images from 677 patients were utilized to develop, validate and prospective validate the effectiveness of SimCLR and CNN (Convolutional Neural Network). Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e displays the basic demographic data of regular and capsule endoscopic pictures gathered from Friendship hospital.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eThe characteristics of all datasets in current study.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eCapsule endoscopic images\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDeveloping and cross validation dataset\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eProspective validation dataset\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo. of subjects\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e322\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e355\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo. of images\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e62850\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e18636\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge (mean)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e39.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e41.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSex\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMale: 182\u003c/p\u003e \u003cp\u003eFemale: 140\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMale: 229\u003c/p\u003e \u003cp\u003eFemale: 126\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePerformance in internal cross validation\u003c/h3\u003e\n\u003cp\u003eThe foundational supervised CNN (ResNet101) was also used to train for identifying unqualified images with supervised manner. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e (top two rows) displays their performance in internal cross validation. It shows that SimCLR's performance was comparable to supervised learning with the average AUROC difference between ResNet101 and random forest models (PCA dimension\u0026thinsp;=\u0026thinsp;20, 30, 50; class weight\u0026thinsp;=\u0026thinsp;30) being 0.0036, 0.0026, 0.0022, respectively. Similarly, the average AUROC difference between ResNet101 and random forest models (PCA dimension\u0026thinsp;=\u0026thinsp;20, 30, 50; class weight\u0026thinsp;=\u0026thinsp;50) was 0.0036, 0.0026, 0.0022, respectively (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Xgboost outperformed random forest, with the average AUROC difference between ResNet101 and Xgboost models (with PCA dimensions of 20, 30, 50) being 0.0014, 0.0001, and 0.0014, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. Xgboost, which was closer to the performance of supervised learning, performed better than random forest. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e displays additional metrics, indicating that the supervised learning model achieved the highest accuracy, with SimCLR closely to the performance of supervised learning. The main difference lies in the sensitivity of the two models.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eThe performance of SimCLR and supervised CNN in internal cross validation (Mean value\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation [95% confidence interval]).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eClass weight\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePCA dimension\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eSensitivity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSpecificity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eAUROC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003eRandom\u003c/p\u003e \u003cp\u003eforest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e[1,30]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9506\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0220\u003c/p\u003e \u003cp\u003e[0.9075\u0026ndash;0.9937]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8629\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1236\u003c/p\u003e \u003cp\u003e[0.6207-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9678\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0414\u003c/p\u003e \u003cp\u003e[0.8867-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9802\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0188\u003c/p\u003e \u003cp\u003e[0.9434-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9505\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0236\u003c/p\u003e \u003cp\u003e[0.9043\u0026ndash;0.9967]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8676\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1068\u003c/p\u003e \u003cp\u003e[0.6582-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9679\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0410\u003c/p\u003e \u003cp\u003e[0.8875-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9810\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0178\u003c/p\u003e \u003cp\u003e[0.9461-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9420\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0289\u003c/p\u003e \u003cp\u003e[0.8854\u0026ndash;0.9987]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8467\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1133\u003c/p\u003e \u003cp\u003e[0.6247-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9690\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0406\u003c/p\u003e \u003cp\u003e[0.8895-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9816\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0174\u003c/p\u003e \u003cp\u003e[0.9475-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e[1,50]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9501\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0223\u003c/p\u003e \u003cp\u003e[0.9064\u0026ndash;0.9937]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8621\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1215\u003c/p\u003e \u003cp\u003e[0.6239-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9679\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0413\u003c/p\u003e \u003cp\u003e[0.8870-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9802\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0187\u003c/p\u003e \u003cp\u003e[0.9436-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9480\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0245\u003c/p\u003e \u003cp\u003e[0.8999\u0026ndash;0.9961]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8624\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1081\u003c/p\u003e \u003cp\u003e[0.6506-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9688\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0418\u003c/p\u003e \u003cp\u003e[0.8869-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9810\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0176\u003c/p\u003e \u003cp\u003e[0.9465-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9402\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0316\u003c/p\u003e \u003cp\u003e[0.8783-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8442\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1114\u003c/p\u003e \u003cp\u003e[0.6258-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9689\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0406\u003c/p\u003e \u003cp\u003e[0.8894-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9816\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0172\u003c/p\u003e \u003cp\u003e[0.9479-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eXgboost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eN/A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9656\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0202\u003c/p\u003e \u003cp\u003e[0.9261-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9104\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1114\u003c/p\u003e \u003cp\u003e[0.6920-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9648\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0417\u003c/p\u003e \u003cp\u003e[0.8830-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9824\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0144\u003c/p\u003e \u003cp\u003e[0.9542-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9685\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0203\u003c/p\u003e \u003cp\u003e[0.9288-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9265\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0898\u003c/p\u003e \u003cp\u003e[0.7506-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9652\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0417\u003c/p\u003e \u003cp\u003e[0.8835-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9837\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0149\u003c/p\u003e \u003cp\u003e[0.9545-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9707\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0206\u003c/p\u003e \u003cp\u003e[0.9304-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9302\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0936\u003c/p\u003e \u003cp\u003e[0.7467-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9654\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0419\u003c/p\u003e \u003cp\u003e[0.8832-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9852\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0137\u003c/p\u003e \u003cp\u003e[0.9583-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eResNet101 (supervised)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eN/A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9734\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0208\u003c/p\u003e \u003cp\u003e[0.9327-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9598\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0583\u003c/p\u003e \u003cp\u003e[0.8455-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9610\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0428\u003c/p\u003e \u003cp\u003e[0.8771-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9838\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0134\u003c/p\u003e \u003cp\u003e[0.9575-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003ePerformance in prospective validation (cross validation)\u003c/h2\u003e \u003cp\u003eThe performance of SimCLR and ResNet101 was shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e (bottom two rows), which shows SimCLR was closer to the performance of supervised learning, the difference is not so significant than internal cross validation in accuracy and sensitivity. Compared to supervised CNN, random forest models with PCA dimensions of 20, 30, and 50 and class weight of 50 had AUROC differences of 0.0294, 0.0141, and 0.0137. In Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, the AUROC gap between random forest models with PCA dimensions of 20, 30, and 50 and supervised CNN was 0.0332, 0.0140, and 0.0181, correspondingly. Xgboost outperformed random forest, with the average AUROC difference between ResNet101 and Xgboost (with PCA dimensions of 20, 30, and 50) being 0.0073, 0.0051, and 0.0004, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Even the accuracies of random forest and Xgboost models are better than supervised learning CNN, especially, Xgboost performed better than CNN in all metrics with a proper PCA dimension (30).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eThe performance of SimCLR and CNN in prospective validation (cross validation).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eClass weight\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePCA dimension\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eSensitivity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSpecificity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eAUROC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003eRandom\u003c/p\u003e \u003cp\u003eforest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e[1,30]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9367\u003c/p\u003e \u003cp\u003e[0.9324\u0026ndash;0.9407]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8052\u003c/p\u003e \u003cp\u003e[0.7394\u0026ndash;0.8420]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9375\u003c/p\u003e \u003cp\u003e[0.9332\u0026ndash;0.9415]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9351\u003c/p\u003e \u003cp\u003e[0.9308\u0026ndash;0.9391]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9370\u003c/p\u003e \u003cp\u003e[0.9327\u0026ndash;0.9410]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8182\u003c/p\u003e \u003cp\u003e[0.7515\u0026ndash;0.8547]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9378\u003c/p\u003e \u003cp\u003e[0.9335\u0026ndash;0.9418]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9504\u003c/p\u003e \u003cp\u003e[0.9461\u0026ndash;0.9544]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9417\u003c/p\u003e \u003cp\u003e[0.9374\u0026ndash;0.9457]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.7662\u003c/p\u003e \u003cp\u003e[0.7032\u0026ndash;0.8039]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9429\u003c/p\u003e \u003cp\u003e[0.9386\u0026ndash;0.9469]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9508\u003c/p\u003e \u003cp\u003e[0.9465\u0026ndash;0.9548]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e[1,50]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9368\u003c/p\u003e \u003cp\u003e[0.9325\u0026ndash;0.9408]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8052\u003c/p\u003e \u003cp\u003e[0.7394\u0026ndash;0.8420]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9376\u003c/p\u003e \u003cp\u003e[0.9333\u0026ndash;0.9416]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9313\u003c/p\u003e \u003cp\u003e[0.9271\u0026ndash;0.9353]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9371\u003c/p\u003e \u003cp\u003e[0.9328\u0026ndash;0.9411]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8182\u003c/p\u003e \u003cp\u003e[0.7515\u0026ndash;0.8547]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9379\u003c/p\u003e \u003cp\u003e[0.9336\u0026ndash;0.9419]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9505\u003c/p\u003e \u003cp\u003e[0.9462\u0026ndash;0.9545]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9374\u003c/p\u003e \u003cp\u003e[0.9331\u0026ndash;0.9414]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8182\u003c/p\u003e \u003cp\u003e[0.7515\u0026ndash;0.8547]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9382\u003c/p\u003e \u003cp\u003e[0.9339\u0026ndash;0.9422]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9474\u003c/p\u003e \u003cp\u003e[0.9431\u0026ndash;0.9514]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eXgboost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eN/A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9359\u003c/p\u003e \u003cp\u003e[0.9316\u0026ndash;0.9399]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8701\u003c/p\u003e \u003cp\u003e[0.7996\u0026ndash;0.9054]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9363\u003c/p\u003e \u003cp\u003e[0.9320\u0026ndash;0.9403]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9572\u003c/p\u003e \u003cp\u003e[0.9529\u0026ndash;0.9612]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9361\u003c/p\u003e \u003cp\u003e[0.9320\u0026ndash;0.9399]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8831\u003c/p\u003e \u003cp\u003e[0.8675\u0026ndash;0.8957]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9364\u003c/p\u003e \u003cp\u003e[0.9321\u0026ndash;0.9404]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9696\u003c/p\u003e \u003cp\u003e[0.9654\u0026ndash;0.9735]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9350\u003c/p\u003e \u003cp\u003e[0.9307\u0026ndash;0.9390]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8831\u003c/p\u003e \u003cp\u003e[0.8117\u0026ndash;0.9181]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9353\u003c/p\u003e \u003cp\u003e[0.9310\u0026ndash;0.9393]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9641\u003c/p\u003e \u003cp\u003e[0.9598\u0026ndash;0.9681]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eResNet101 (supervised)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eN/A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9333\u003c/p\u003e \u003cp\u003e[0.9290\u0026ndash;0.9373]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8831\u003c/p\u003e \u003cp\u003e[0.8117\u0026ndash;0.9181]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9336\u003c/p\u003e \u003cp\u003e[0.9293\u0026ndash;0.9376]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.9645\u003c/p\u003e \u003cp\u003e[0.9602\u0026ndash;0.9685]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003ePerformance in internal reversed cross validation\u003c/h2\u003e \u003cp\u003eIn order to test whether SimCLR can complete the same task with less annotation, we used reversed cross validation (less training data and more testing data) to fairly compare the performance of SimCLR and supervised CNN. Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e (top row) displays the results of their performance in internal reversed cross validation. The comparison of SimCLR's performance to supervised learning showed that the average AUROC difference between ResNet101 and the random forest model was 0.0235, while the average AUROC difference between ResNet101 and the Xgboost model was 0.0125 (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e displays additional metrics, indicating that while the accuracy of supervised ResNet101 is the highest, the AUROC values of random forest and Xgboost models are similar to that of supervised learning ResNet101.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eThe performance of SimCLR and supervised CNN in internal reversed cross validation (Mean value\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation [95% confidence interval]).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSensitivity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSpecificity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAUROC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.8409\u0026thinsp;\u0026plusmn;\u0026thinsp;0.2113\u003c/p\u003e \u003cp\u003e[0.4268-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.6862\u0026thinsp;\u0026plusmn;\u0026thinsp;0.4388\u003c/p\u003e \u003cp\u003e[0.0000\u0026ndash;1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9796\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0178\u003c/p\u003e \u003cp\u003e[0.9448-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9824\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0096\u003c/p\u003e \u003cp\u003e[0.9635-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXgboost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.8489\u0026thinsp;\u0026plusmn;\u0026thinsp;0.2107\u003c/p\u003e \u003cp\u003e[0.4359-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.7093\u0026thinsp;\u0026plusmn;\u0026thinsp;0.4390\u003c/p\u003e \u003cp\u003e[0.0000\u0026ndash;1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9769\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0177\u003c/p\u003e \u003cp\u003e[0.9422-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9714\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0305\u003c/p\u003e \u003cp\u003e[0.9116-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eResNet101 (supervised)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.9048\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1305\u003c/p\u003e \u003cp\u003e[0.6489-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.8521\u0026thinsp;\u0026plusmn;\u0026thinsp;0.2323\u003c/p\u003e \u003cp\u003e[0.3969-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9533\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0333\u003c/p\u003e \u003cp\u003e[0.8881-1.0000]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9589\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0395\u003c/p\u003e \u003cp\u003e[0.8816-1.0000]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003ePerformance in prospective validation (reversed cross validation)\u003c/h2\u003e \u003cp\u003eThe performance of SimCLR and ResNet101 is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e (bottom row), which shows SimCLR was better than the performance of supervised learning. The AUROC discrepancy between the random forest and the supervised CNN was 0.1256. The AUROC discrepancy between the Xgboost and the supervised CNN was 0.1353. Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e displays additional metrics, indicating that SimCLR performed exceptionally well and the image features extracted can be utilized for downstream tasks. Moreover, Xgboost model is better than random forest model and supervised learning CNN.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eThe performance of SimCLR and CNN in prospective validation (reversed cross validation).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSensitivity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSpecificity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAUROC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.9445\u003c/p\u003e \u003cp\u003e[0.9402\u0026ndash;0.9485]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.8312\u003c/p\u003e \u003cp\u003e[0.7635\u0026ndash;0.8674]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9452\u003c/p\u003e \u003cp\u003e[0.9409\u0026ndash;0.9492]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9630\u003c/p\u003e \u003cp\u003e[0.9587\u0026ndash;0.9670]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXgboost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.9390\u003c/p\u003e \u003cp\u003e[0.9347\u0026ndash;0.9430]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.8831\u003c/p\u003e \u003cp\u003e[0.8117\u0026ndash;0.9181]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9394\u003c/p\u003e \u003cp\u003e[0.9351\u0026ndash;0.9434]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9727\u003c/p\u003e \u003cp\u003e[0.9684\u0026ndash;0.9767]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eResNet101 (supervised)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.8924\u003c/p\u003e \u003cp\u003e[0.8883\u0026ndash;0.8963]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.7273\u003c/p\u003e \u003cp\u003e[0.6672\u0026ndash;0.7658]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.8934\u003c/p\u003e \u003cp\u003e[0.8892\u0026ndash;0.8973]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.8374\u003c/p\u003e \u003cp\u003e[0.8334\u0026ndash;0.8412]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe constant need for annotations places a significant strain on healthcare workers and doctors. Self-supervised learning (SimCLR) can complete image quality controlling task with less annotation for capsule endoscopic images. SimCLR outperformed supervised learning in both cross validation and reversed cross validation, and also achieved better results than supervised CNN using fewer annotated data. The annotation burden of medical staffs and physicians can be alleviated largely.\u003c/p\u003e \u003cp\u003eSharib Ali and colleagues\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e introduced a framework that evaluates the quality of endoscopic images through the identification and segmentation of artifacts like blurriness, reflections, color intensity, air bubbles, contrast, and other imperfections artefacts. Four object detection algorithms were compared, including YOLOv3, RetinaNet, YOLOv3-spp, and Faster-RCNN, with YOLOv3-spp demonstrating the highest overall performance. Additionally, they evaluated the performance of four different semantic segmentation models (ResNet-UNet, DeepLabv3+, FCN8, pspNet) in segmenting all artifacts, with DeepLabv3\u0026thinsp;+\u0026thinsp;performing best. Finally, they created a rating system that is highly correlated with three experts, with a correlation exceeding 0.6. He et al.\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e conducted a study where 3704 endoscopic images were gathered from 211 patients to train Inception-v3, ResNet-50, VGG-16-bn, DenseNet-121 and VGG-11-bn models in order to classify the images as either qualified or unqualified. According to VGG-11-bn, unqualified images had the highest accuracy of 67.6%, while qualified images had the highest accuracy of 80.9%. Lianlian Wu et al.\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e trained a deep CNN using 12220 in vitro, 25222 in vivo, and 16760 unqualified endoscopic images to determine if the endoscopic image contained the unrelated content (out of body) during esophagogastroduodenoscopy. The result shows that the accuracies for each class are higher than 70%. In addition, our team also utilized transfer learning (domain adaptation) from standard endoscopic images to capsule endoscopic images in order to manage the image quality. The findings indicated that domain adaptation could successfully accomplish this task using annotations from just a single domain, Moreover, the performance surpassed the supervised learning\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. We also compared CNN with vision transformer (ViT) in quality controlling for capsule and conventional endoscopic images. The result indicated that CNN outperformed ViT (Vision Transformer) in completing this task\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eBy limiting the selection to capsule endoscope, there are some drawbacks, like not taking into account the distinctions between different brands. Due to our failure to classify all eligible conditions, the distribution of images in actual clinical situations may not be uniform.\u003c/p\u003e \u003cp\u003eOur image quality control model can be seamlessly integrated into the workflow of the auto image reading system, whether video or image analysis is conducted in real-time or during post-processing. It is expected that the examining process would capture a large number of images, so the unqualified images might be sifted out in order to save more time. A new model can be trained for a healthcare facility by incorporating data from a new source (healthcare center) and leveraging the existing research model as a pretrained model.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eUtilizing SimCLR, intrinsic characteristics can be extracted from standard endoscopic images, enabling machine learning algorithms to identify and filter out inadequate images with minimal manual labeling. This not only reduces the daily tasks of doctors, but also decreases the amount of annotations needed for medical artificial intelligence.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch3\u003eCode and Data Availability:\u0026nbsp;\u003c/h3\u003e\n\u003cp\u003eAccess to the Python scripts necessary for the analysis can be found at https://github.com/Hugo0512/ASI4EGD, while the source code for SimCLR can be found at https://github.com/Spijkervet/SimCLR. The information and resources in this research can be obtained from the corresponding author upon request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements:\u003c/strong\u003e None\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFundings\u003c/strong\u003e:\u0026nbsp;The project received funding from various sources including the General Program of National Natural Science Foundation of China (8217100675), University Department Construction Project in Endoscope Center Area (2020MWDXK03); Shanghai Key Medical College Construction Project (ZK2019C05); Minhang District Leading Talent Project (2021LJRC03); Natural Science Basic Research Plan in Shaanxi Province of China (2022JQ-175), Scientific Research Plan of Shaanxi Education Department (22JK0303) and National Major Scientific Instruments and Equipments Development Project of National Natural Science Foundation of China (82027801). \u003c/p\u003e\u003cp\u003eThe sponsor or funding organization had no role in the design or conduct of this research.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e: Y.Q.Z. helped shape the study\u0026apos;s framework. F.H. and P.L. critically reviewed the manuscript. Y.Q.Z., and K.Z. conducted the research and performed the literature review. T.M., Y.Q.Z., and P.L. gathered the information and categorized it. Y.Q.Z., K.Z., P.B., S.Q.W., and M.J.W. assisted in developing the statistical analysis and interpreting the data. Y.Q.Z., P.B. and K.Z. drafted the manuscript. M.J.W., and P.L. provided research funding. Y.Q.Z., K.Z., and F.H. managed the research project. All authors had access to all the raw datasets and the corresponding authors (T.M., F.H., and P.L.) confirming the data.The final manuscript was reviewed and approved by all authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication:\u003c/strong\u003e : The authors affirm that all content and resources included in the manuscript are authentic.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participat\u003c/strong\u003e: This study complied with the Declaration of Helsinki. The research was approved by the Medical Ethics Committee of Friendship hospital (2022-P1-equipment-018-01) and Minhang hospital (2023-009-01X). \u0026nbsp;The informed consent was signed by all patients.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e: None\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDeclaration of competing interest:\u0026nbsp;\u003c/strong\u003eNo conflicting relationship exists for any author.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eShorten C, Khoshgoftaar TM. A survey on image data augmentation for deep learning. J big data. 2019;6:1\u0026ndash;48.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang M, et al. Differential diagnosis for esophageal protruded lesions using a deep convolution neural network in endoscopic images. Gastrointest Endosc. 2021;93:1261\u0026ndash;72.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen D, et al. Comparing blind spots of unsedated ultrafine, sedated, and unsedated conventional gastroscopy with and without artificial intelligence: a prospective, single-blind, 3-parallel-group, randomized, single-center trial. Gastrointest Endosc. 2020;91:332\u0026ndash;9. e333.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiao Z, Zou W, Li Z-S. Clinical application of magnetically controlled capsule gastroscopy in gastric disease diagnosis: recent advances. Sci China Life Sci. 2018;61:1304\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChow LS, Paramesran R. Review of medical image quality assessment. Biomed Signal Process Control. 2016;27:145\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang S, et al. The applications of artificial intelligence in digestive system neoplasms: a review. Health Data Sci. 2023;3:0005.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSharma P, Hassan C. Artificial intelligence and deep learning for upper gastrointestinal neoplasia. Gastroenterology. 2022;162:1056\u0026ndash;66.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuo H, et al. Real-time artificial intelligence for detection of upper gastrointestinal cancer by endoscopy: a multicentre, case-control, diagnostic study. Lancet Oncol. 2019;20:1645\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuo J, et al. A deep learning method to assist with chronic atrophic gastritis diagnosis using white light images. Dig Liver Disease. 2022;54:1513\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang P, et al. Development and validation of a deep-learning algorithm for the detection of polyps during colonoscopy. Nat biomedical Eng. 2018;2:741\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDahmardeh M, Dastjerdi M, Mazal H, K\u0026ouml;stler H, H., Sandoghdar V. Self-supervised machine learning pushes the sensitivity limit in label-free detection of single proteins below 10 kDa. Nat Methods, 1\u0026ndash;6 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHendrycks D, Mazeika M, Kadavath S, Song D. Using self-supervised learning can improve model robustness and uncertainty. Adv Neural Inf Process Syst 32 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen T, Kornblith S, Norouzi M, Hinton G. in \u003cem\u003eInternational conference on machine learning.\u003c/em\u003e 1597\u0026ndash;1607 (PMLR).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou W-D, et al. Deep Learning for Automatic Detection of Recurrent Retinal Detachment after Surgery Using Ultra-Widefield Fundus Images: A Single‐Center Study. Adv Intell Syst. 2022;4:2200067.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHui S, et al. Noninvasive identification of Benign and malignant eyelid tumors using clinical images via deep learning system. J Big Data. 2022;9:84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang M, et al. Computerized assisted evaluation system for canine cardiomegaly via key points detection with deep learning. Prev Vet Med. 2021;193:105399.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang H, et al. Validation of the relationship between iris color and uveal melanoma using artificial intelligence with multiple paths in a large Chinese population. Front Cell Dev Biology. 2021;9:713209.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang K, et al. Prediction of postoperative complications of pediatric cataract patients using data mining. J translational Med. 2019;17:1\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGangil T, et al. Predicting clinical outcomes of radiotherapy for head and neck squamous cell carcinoma patients using machine learning algorithms. J Big Data. 2022;9:1\u0026ndash;19.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDaffertshofer A, Lamoth CJ, Meijer OG, Beek P. J. PCA in studying coordination and variability: a tutorial. Clin Biomech Elsevier Ltd. 2004;19:415\u0026ndash;28.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi W, et al. Dense anatomical annotation of slit-lamp images improves the performance of deep learning for the diagnosis of ophthalmic disorders. Nat Biomedical Eng. 2020;4:767\u0026ndash;77.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSampath V, Maurtua I, Martin A, J. J., Gutierrez A. A survey on generative adversarial networks for imbalance problems in computer vision tasks. J big Data. 2021;8:1\u0026ndash;59.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang R, et al. Automatic retinoblastoma screening and surveillance using deep learning. Br J Cancer. 2023;129:466\u0026ndash;74. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41416-023-02320-z\u003c/span\u003e\u003cspan address=\"10.1038/s41416-023-02320-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang K, et al. An interpretable and expandable deep learning diagnostic system for multiple ocular diseases: qualitative study. J Med Internet Res. 2018;20:e11144.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZinkevich M, Weimer M, Li L, Smola A. Parallelized stochastic gradient descent. Adv Neural Inf Process Syst 23 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang K, et al. A human-in-the-loop deep learning paradigm for synergic visual evaluation in children. Neural Netw. 2020;122:163\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchmidt H, Spieker AJ, Luo T, Szymczak JE, Grande D. in JAMA Health Forum e212932\u0026ndash;212932 (American Medical Association).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAli S, et al. A deep learning framework for quality assessment and restoration in video endoscopy. Med Image Anal. 2021;68:101900.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe Q, et al. Deep learning-based anatomical site classification for upper gastrointestinal endoscopy. Int J Comput Assist Radiol Surg. 2020;15:1085\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu L, et al. Randomised controlled trial of WISENSE, a real-time quality improving system for monitoring blind spots during esophagogastroduodenoscopy. Gut. 2019;68:2161\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Y et al. Deep transfer learning from ordinary to capsule esophagogastroduodenoscopy for image quality controlling. Eng Rep n/a, e12776, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/eng2.12776\u003c/span\u003e\u003cspan address=\"10.1002/eng2.12776\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang K, et al. Anatomical sites identification in both ordinary and capsule gastroduodenoscopy via deep learning. Biomed Signal Process Control. 2024;90:105911. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.bspc.2023.105911\u003c/span\u003e\u003cspan address=\"10.1016/j.bspc.2023.105911\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"SimCLR, quality controlling, esophagogastroduodenoscopy","lastPublishedDoi":"10.21203/rs.3.rs-6324281/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6324281/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eIt is possible to control the quality of capsule endoscopic images using artificial intelligence (AI), but it requires a great deal of time for labeling.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eSimCLR (a simple framework for contrastive learning of visual representations), is capable of acquiring the inherent image representation with minimal annotation, but the feasibility is not studied. 62850 images were collected to train models. In internal cross-validation (more training data and less testing data) and reversed cross-validation (less training data and more testing data). Random forest and Xgboost (eXtreme Gradient Boosting) were used to finish the quality controlling after SimCLR extracting the features from images.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eSimCLR reported that the mean AUROC (Area Under the Receiver Operating Characteristic) curve exceeded 0.98 and 0.97. Moreover, Xgboost surpassed supervised CNN (Convolutional Neural Network). Extra 18636 pictures were gathered and the AUROC of SimCLR surpassed 0.93 (95% CI 0.9271\u0026ndash;0.9548), which is close to supervised CNN (Convolutional Neural Network) (0.9645) in cross validation. Moreover, the AUROC of SimCLR surpass 0.96, which is better than supervised CNN (0.8374) in reversed cross validation.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eThrough SimCLR, the capsule endoscopic image quality control task can be completed with a performance similar to or better than that of supervised learning with fewer annotations.\u003c/p\u003e","manuscriptTitle":"Quality controlling in capsule gastroduodenoscopy with less annotation via self-supervised learning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-09 11:21:54","doi":"10.21203/rs.3.rs-6324281/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"fcfcad24-ebd9-433d-a3c7-7b6ba176ec6a","owner":[],"postedDate":"May 9th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-03-16T06:56:02+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-09 11:21:54","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6324281","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6324281","identity":"rs-6324281","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.