Training certified detectives to track down the intrinsic shortcuts in COVID-19 chest x-ray data sets

preprint OA: gold CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Deep learning faces a significant challenge wherein the trained models often underperform when used with external test data sets. This issue has been attributed to spurious correlations between irrelevant features in the input data and corresponding labels. This study uses the classification of COVID-19 from chest x-ray radiographs as an example to demonstrate that the image contrast and sharpness, which are characteristics of a chest radiograph dependent on data acquisition systems and imaging parameters, can be intrinsic shortcuts that impair the model’s generalizability. The study proposes training certified shortcut detective models that meet a set of qualification criteria which can then identify these intrinsic shortcuts in a curated data set.
Full text 92,669 characters · extracted from preprint-html · click to expand
Training certified detectives to track down the intrinsic shortcuts in COVID-19 chest x-ray data sets | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Training certified detectives to track down the intrinsic shortcuts in COVID-19 chest x-ray data sets Ran Zhang, Dalton Griner, John W. Garrett, Zhihua Qi, Guang-Hong Chen This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2818347/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 04 Aug, 2023 Read the published version in Scientific Reports → Version 1 posted 8 You are reading this latest preprint version Abstract Deep learning faces a significant challenge wherein the trained models often underperform when used with external test data sets. This issue has been attributed to spurious correlations between irrelevant features in the input data and corresponding labels. This study uses the classification of COVID-19 from chest x-ray radiographs as an example to demonstrate that the image contrast and sharpness, which are characteristics of a chest radiograph dependent on data acquisition systems and imaging parameters, can be intrinsic shortcuts that impair the model’s generalizability. The study proposes training certified shortcut detective models that meet a set of qualification criteria which can then identify these intrinsic shortcuts in a curated data set. Figures Figure 1 Figure 2 Figure 3 Introduction Deep learning has been incredibly successful in object detection, image classification, and natural language processing over the past ten years due to its ability to learn complex features from data. However, despite its success on benchmark datasets, there are limitations and practical issues when using these models in real-world scenarios. A major challenge is poor generalizability, where performance significantly drops when applied to external datasets 1 – 3 . This limits the translation and deployment of deep learning models for high-stakes tasks, such as in healthcare applications. This lesson was evident during the COVID-19 pandemic, where many machine learning models were developed, but very few performed well on real-world clinical tests 4 – 6 . The concept of shortcut learning has recently been explored in deep learning studies 7 . It has been discovered that poor model generalizability can be attributed to shortcut learning when the training dataset has hidden shortcuts, meaning there are spurious correlations between irrelevant image features and the corresponding training labels. This causes models to quickly pick up these spurious correlations instead of the desired image features, establishing incorrect connections between input image data and output labels. For instance, early studies have shown that deep learning models can differentiate chest x-rays from different hospitals and patient groups 1 . This suggests that different data sources and patient characteristics like gender, age, and race could also become shortcuts. To illustrate the issue of shortcut learning in real-world clinical scenarios, let's take the example of COVID-19 classification using chest x-ray radiographs (CXRs). DeGrave et al. 8 discovered that if COVID-19 positive and negative training data were collected from two different sources, the model would only learn the source label as a shortcut for prediction. As a result, the model would not have the desired prediction power in real-world clinical scenarios. The authors also found that the trained model used extrinsic image features, such as lead markers, for prediction, even though these markers only indicated the orientation of patients in x-ray image acquisitions and did not correspond to any disease features. Some suggestions have been made to remove these extrinsic shortcuts by a segmentation step 9 . However, even with these markers removed from the training dataset, some other shortcuts still existed in the segmented lung tissue-only training dataset 10 . As a result, the nearly perfect performance of the trained model 9 is still not generalizable to real-world clinical datasets. The remaining shortcuts may be attributed to the inherent defining features of CXRs, such as image contrast and sharpness, which can vary from hospital to hospital due to different types of imaging systems, generations of x-ray imaging equipment, hardware components, image post-processing methods used by vendors, and imaging protocols used by technologists. All these factors can impact the digital representation of the acquired image data in terms of variations in image contrast and sharpness. Figure 1 demonstrates these variations. Unlike other shortcuts that have been previously studied, such as age, gender, race, and markers, which are extrinsic and can be removed through careful data collection and cleaning, contrast- and sharpness-related shortcuts are more difficult to detect and mitigate. This is because desired image features are also represented as image contrast and spatial correlations, which are similar to the features of contrast- and sharpness-related shortcuts. This entanglement between desired image features and shortcut features makes studying contrast and sharpness-related shortcuts particularly challenging. Consequently, these shortcuts are referred to as intrinsic shortcuts in this work. In order to develop an effective strategy for mitigating contrast- and sharpness-related shortcuts, it is necessary to first develop reliable methods for detecting their presence and severity within a carefully curated dataset. While post-hoc model interpretability methods, such as class activation maps 11 and expected gradient 12 , have been developed to identify relevant image features used by trained deep learning models for prediction, these methods are unable to detect intrinsic shortcuts within a curated training dataset prior to model training. Furthermore, studies have suggested that these methods may not be effective in diagnosing poor generalization performance of the model 13 . In this paper, we present a novel approach for detecting contrast- and sharpness-related intrinsic shortcuts using certified shortcut detective models. Our approach involves establishing qualification standards for suspected intrinsic shortcuts, designing a training curriculum for training the shortcut detectives to detect these shortcuts, performing certification tests on the trained detectives, and finally deploying them to curated datasets to examine the suspected shortcuts. We applied this approach to the available COVID-19 datasets to assess their quality. Our results demonstrate the effectiveness of this approach in detecting and mitigating intrinsic shortcuts. Methods Datasets Figure 2 provides an overview of the datasets utilized in this study. The MIMIC-CXR dataset served as the training data for the shortcut detectives, while the HF-train dataset, a privately curated COVID-19 chest x-ray dataset, was utilized for certification tests of the trained detectives. The trained shortcut detectives were then applied to a variety of public and private chest x-ray datasets. Specific information about each dataset is provided below. Figure 2. An overview of the datasets used in this work. MIMIC dataset MIMIC-CXR 14 chest x-ray dataset consists of 377,110 CXRs from 65,379 patients presenting to the Beth Israel Deaconess Medical Center Emergency Department between 2011–2016. In this work, 46,894 frontal-view (AP/PA) normal CXRs (cases with “No Finding” labels) were used to train shortcut detectives. HF dataset This is a privately curated COVID-19 CXR dataset from patients presenting to the Henry Ford Health between March 1, 2020, and October 31, 2020. The COVID-19-positive and COVID-19-negative cohorts are collected within the same time range, from the same hospitals, and labeled by their most recent RT-PCR test result seven days before or after the imaging study. For model training and internal testing, two data partitions are generated: HF-train consisted of 8,733 COVID-19-positive CXRs from 4,383 patients and 16,584 COVID-19-negative CXRs from 8,733 patients; HF-test consisted of 695 COVID-19-positive CXRs from 526 patients and 8,878 COVID-19-negative CXRs from 6,081 patients. BIMCV dataset This is a public COVID-19 CXR dataset collected in Spain 15 . This dataset was collected from 11 hospitals in the Valencian Region, Spain, between February and April 2020. After data curation, the dataset consisted of 4,169 COVID-19-positive CXRs from 2,663 patients and 5,050 COVID-19-negative CXRs from 3,710 patients. UW dataset This is a privately curated COVID-19 CXR dataset. It includes consecutive patient cases from the University of Wisconsin Hospitals and Clinics (UW Health) from March 2020 to September 2021. The dataset comprised 1,025 COVID-19-positive CXRs from 658 patients and 8,774 COVID-19-negative CXRs from 5,953 patients. MIDRC dataset This is a large, multi-institution public COVID-19 CXR dataset curated and released by the Medical Imaging & Data Resource Center (MIDRC). A total of 6,453 COVID-19-positive CXRs from 5,199 patients and 20,072 COVID-19-negative CXRs from 9,947 patients were pulled from the MIDRC Data Commons ( https://data.midrc.org/ , date accessed: December 7th, 2022). COVIDx dataset This is a public COVID-19 CXR dataset released by the COVID-Net Open Source Initiative 16 . A total of 29,986 CXRs from 16,648 patients are included in the training dataset ( COVIDx-train ), and 400 CXRs are included in the test dataset (COVIDx-test). (Data downloaded from https://www.kaggle.com/datasets/andyczhao/covidx-cxr2?select=competition_test , date accessed: December 7th, 2022). RoentGen-MIMIC dataset This dataset contains 943 synthetic CXRs generated by the RoentGen model 17 and 1,000 real “No Finding” CXRs from the MIMIC dataset. The RoentGen model, which is trained using the MIMIC dataset, is able to generate visually convincing synthetic CXRs with different pathologies. To generate the synthetic CXRs used in this work, a text prompt of “No finding” was used as the input, only frontal view CXRs are included. Training and certification of shortcut detectives Overview of the proposed framework An overview of the shortcut detective training and certification process is shown in Fig. 3 . To begin, 46,894 normal CXRs from the MIMIC dataset were selected and randomly divided into two equal groups: 23,447 of the CXRs were assigned as the positive class (“1”) while the rest were assigned as the negative class (“0”). Then, to construct the training dataset for the shortcut detective, the image contrast or the image sharpness of the positive class were adjusted using the approach shown in Appendix A1. Since only normal CXRs were included, there were no disease-specific features present that could be used to distinguish between the two classes. Therefore, if a model was able to differentiate between the two classes, it would be because the model had learned the corresponding global image contrast or sharpness characteristics, rather than disease features. Subsequently, shortcut detectives (neural network models for binary classification) were trained using the constructed training datasets with contrast or sharpness shortcuts. Details of the training are discussed in the following section and in Appendix A2. To assess the efficacy of the shortcut detectives on COVID-19 CXR dataset, two types of examinations are necessary. Firstly, when the shortcut detective is deployed on a COVID-CXR dataset without the corresponding shortcut, it should not be able to differentiate between the COVID-positive and COVID-negative classes. Hence, an Area Under the Receiver Operating Characteristics curve (AUC) close to 0.5 is expected. This examination is crucial to ensure that the image features utilized by the shortcut detectives are not interwoven with the original imaging task. Secondly, when the shortcut detective is applied to a COVID-CXR dataset with a known shortcut, it should demonstrate superior classification performance, ideally with an AUC close to 1. For the first examination, the HF-train dataset is utilized. As the positive and negative cohorts are gathered within the same timeframe and from the same hospitals, it is expected that no contrast and sharpness shortcut exists. This has also been corroborated recently 18 , where a model trained on this dataset demonstrated consistent test performance on various external COVID-19 clinical test datasets. For the second examination, known shortcuts are integrated into the COVID-positive class or COVID-negative class of the HF-train dataset using the same procedures outlined in Appendix A1. Finally, if the trained shortcut detectives pass the two exams, i.e. AUC of close to 0.5 on the shortcut-free dataset and AUC of close to 1.0 on the shortcut-present dataset, they are referred to as certified shortcut detectives. Image preprocessing and model architecture For the MIMIC, HF, BIMCV, UW Health and MIDRC datasets, the original DICOM images are converted to 8-bit png format with a size of 224-by-224 using the default window level and window width. For the COVIDx dataset, images are directly resized to 224-by-224. To train the shortcut detectives, five different model architectures (Table 1 ) that are broadly used for image classification with state-of-the-art performance on the ImageNet 19 classification tasks are investigated in this work. Although we cannot exhaust all the possible model architectures for this purpose, the models we investigated include classic and modern Convolutional Neural Networks (CNN) and the recently introduced Swin Transformer. These models vary in architectural design and complexity (number of model parameters and floating-point operations, FLOPs). For each model architecture, an ensemble of five individually trained models with different training-validation splits are used. More technical details on the model training are shown in Appendix A2. Table 1 List of different model architectures studied in this work Model Number of parameters FLOPs ImageNet Accuracy Year developed VGG-16 20 138 M 31 B 73.4% 2014 DenseNet-121 21 8 M 5.7 B 74.4% 2017 EfficientNet 22 54 M 24 B 85.1% 2020 Swin Transformer 23 88 M 15 B 83.6% 2021 ConvNeXt 24 89 M 15 B 84.1% 2022 Note: ImageNet accuracy data obtained from: https://pytorch.org/vision/stable/models.html#classification Deployment of the shortcut detectives Certified shortcut detectives are deployed to detect shortcuts in real-world datasets, including BIMCV, UW, MIDRC, COVIDx, and RoentGen-MIMIC. We also trained two COVID-19 classification models using HF-train dataset and COVIDx-train dataset and compared their generalizability using internal and external tests. Statistics The 95% confidence intervals (CI) for the AUC were calculated using the statistical software R (version 4.0.0) with the pROC package. CIs were calculated using the bootstrap method with 2000 bootstrap replicates. Results Certification of the shortcut detectives As shown in Table 2 , the trained detectives successfully passed the two exams. Note that an AUC of 0.0 is simply due to the assignment of class labels, which is equivalent to an AUC of 1.0. Both indicate perfect classification performance. It is also shown in Table 2 that all five neural network architectures achieve a similar performance level. This result aligns well with our understanding that contrast and sharpness shortcuts are intrinsic to the dataset and the task. Shortcut detection in real-world datasets Using the certified shortcut detectives, we investigated several curated COVID-CXR datasets for possible shortcuts. UW Health is a single-site, privately curated dataset; BIMCV and MIDRC are both multi-institutional public datasets; COVIDx is the first open access COVID-CXR dataset made available by the experts in the computer science community. Similar to the detective certification process, for a COVID-CXR dataset where COVID-19-positive cases are assigned a label “1” and negative cases are given a label “0”, if the shortcut detective can differentiate the two classes (AUC significantly deviates from 0.50), it indicates the existence of a corresponding shortcut. For the RoentGen-MIMIC dataset, real CXRs are labelled as “1” and synthetic CXRs are labelled as “0”. The results presented in Table 3 were obtained using shortcut detectives based on the DenseNet model. The results confirm the presence of image sharpness and contrast shortcuts in the COVIDx dataset, which can be exploited by models trained on such data and compromise their generalizability in real clinical settings. Conversely, the other three datasets that were curated by medical professionals exhibit no such shortcuts. We conducted a performance comparison between two models trained on COVIDx and Henry Ford datasets, respectively, which are of similar size, but the former has both sharpness and contrast shortcuts as identified by the shortcut detectives. The results displayed in Table 4 indicate that the COVIDx model exhibits poor generalization performance, as evidenced by the large AUC gap between internal and external tests. In contrast, the HF model exhibits consistent performance on both internal and external tests. Additionally, as shown in Table 3 , some contrast and sharpness differences were also detected in the RoentGen-MIMIC dataset. While the generated synthetic CXRs appear visually realistic, caution must be exercised when using them for AI model development due to the potential learning of shortcuts caused by the inherent contrast and sharpness differences between real and synthetic data. Table 2 Certification of the shortcut detectives VGG DENSE Eff Swin Conv ADA(S) shortcut detective Exam 1: HF-train 0.49 [0.48,0.50] 0.56 [0.56,0.57] 0.56 [0.55,0.57] 0.53 [0.52,0.53] 0.54 [0.53,0.54] Exam 2a: With ADA-S(+) 1.00 [1.00,1.00] 1.00 [1.00,1.00] 1.00 [1.00,1.00] 1.00 [1.00,1.00] 1.00 [1.00,1.00] Exam 2b: With ADA-S(-) 0.00 [0.00,0.00] 0.00 [0.00,0.00] 0.00 [0.00,0.00] 0.00 [0.00,0.00] 0.00 [0.00,0.00] ADA(C) shortcut detective Exam 1: HF-train 0.47 [0.46,0.48] 0.48 [0.47,0.49] 0.47 [0.46,0.48] 0.49 [0.48,0.49] 0.48 [0.47,0.48] Exam 2a: With ADA-C(+) 1.00 [1.00,1.00] 1.00 [1.00,1.00] 1.00 [1.00,1.00] 1.00 [1.00,1.00] 1.00 [1.00,1.00] Exam 2b: With ADA-C(-) 0.00 [0.00,0.00] 0.00 [0.00,0.00] 0.00 [0.00,0.00] 0.00 [0.00,0.00] 0.00 [0.00,0.00] Note: (+) means the shortcut is added to the COVID-positive CXRs; (-) means the shortcut is added to the COVID-negative CXRs. Table 3 Shortcut detection on real-world COVID-CXR datasets (DenseNet model result) COVIDx RoentGen-MIMIC UW BIMCV MIDRC ADA(S) shortcut 0.84 [0.83,0.84] 0.37 [0.34,0.39] 0.50 [0.48,0.52] 0.47 [0.45,0.48] 0.45 [0.45,0.46] ADA(C) shortcut 0.81 [0.80,0.81] 0.05 [0.04, 0.06] 0.53 [0.51,0.55] 0.56 [0.55,0.57] 0.53 [0.52,0.54] Table 4 Performance of two trained models Internal UW BIMCV MIDRC COVIDx 1.00 [1.00, 1.00] 0.60 [0.58,0.62] 0.61 [0.60,0.62] 0.57 [0.56,0.58] HF 0.78 [0.76,0.80] 0.77 [0.76,0.79] 0.79 [0.78,0.80] 0.75 [0.74,0.75] Discussion Shortcut learning has been a topic of interest in the machine learning community, particularly in computer vision (CV) and natural language processing (NLP). Researchers have explored shortcut learning behavior from different perspectives, such as underspecification 3 , shortcut learning in various NLP tasks 25 , and mitigation strategies for domain-knowledge agnostic models 26 . Notably, it has been observed that convolutional neural networks tend to rely on content with high spatial frequency or strong local correlations to establish connections between input and labels in CV and NLP 27 , 28 . However, it remains unclear whether these observations can be extended to medical diagnosis tasks, where clinical datasets have distinct characteristics from those in ImageNet. Therefore, studying shortcut learning for well-defined, clinically relevant tasks using real-world clinical datasets is crucial for medical AI applications. In this study, we demonstrated that acquisition-dependent attributes (ADAs), such as image contrast and sharpness differences arising from the entire image generation pipeline, can serve as intrinsic shortcuts during clinical diagnosis learning tasks. Inadequate quality control procedures during data collection can allow these shortcuts to inadvertently infiltrate the curated dataset. If shortcuts contaminate the dataset, neural networks can easily exploit them during training, thereby impairing their ability to generalize to other real-world datasets. Thus, it is imperative to identify possible shortcuts in the training dataset prior to model development. In this study, we present a methodical framework for training and validating shortcut detectives for chest X-ray classification, with emphasis on image contrast and sharpness - two essential intrinsic characteristics of chest X-ray images. However, it should be noted that these are not the only possible shortcuts that may exist in chest X-ray datasets. If other intrinsic shortcuts are suspected in a collected dataset, the general framework presented in this work can be utilized to construct similar shortcut detectives and identify alleged intrinsic shortcuts. Once the intrinsic shortcuts are identified, it is essential to develop strategies to mitigate their impact on the learned models. One possible approach is to develop standardization and normalization techniques for image contrast and sharpness to adjust these attributes without affecting the disease features. Alternatively, examining proven intrinsic shortcut-free datasets, such as the baseline dataset (Henry Ford Health) and the three additional datasets (UW Health, BIMCV, and MIDRC) shown to be free of intrinsic shortcuts in this study, can provide further insight on how to avoid these shortcuts in the data curation process. However, it is worth noting that this is a limitation of the present work, and future research should investigate the development of mitigation strategies for the identified intrinsic shortcuts. Declarations Data Availability Statement BIMCV(https://bimcv.cipf.es/bimcv-projects/bimcv-covid19/), MIDRC (https://data.midrc.org/) and MIMIC-CXR (https://physionet.org/content/mimic-cxr/2.0.0/) datasets are publicly available. Henry Ford Dataset and UW Health Dataset are private datasets with protected patient information. The python script used for model training and evaluation; weights of all trained models can be found on the GitHub repository: https://github.com/uw-ctgroup/shortcut. Acknowledgments: This study was partially supported by NIH Grants 3U01EB021183-W4, R01HL153594, R01EB032474, and a grant from the Wisconsin Partnership Program. The imaging and associated clinical data downloaded from MIDRC (The Medical Imaging Data Resource Center) and used for research in this [publication/press release] were made possible by the National Institute of Biomedical Imaging and Bioengineering (NIBIB) of the National Institutes of Health under contracts 75N92020C00008 and 75N92020C00021. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Author Contributions Statement: R.Z. and G.H.C. conceived the study and designed the experiments. J.W.G. and Z.Q. curated two private clinical datasets for tests. R.Z. prepared the datasets for model training and testing. D.G. and R.Z. wrote the software code. R.Z. and G.H.C. performed the statistical analysis. G.H.C. supervised the study. R.Z. and G.H.C. wrote the manuscript. All authors read and edited the manuscript. References Zech, J. R. et al. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS medicine 15 , e1002683 (2018). Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J. & Song, D. Natural adversarial examples. in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 15262–15271 (2021). D’Amour, A. et al. Underspecification presents challenges for credibility in modern machine learning. arXiv preprint arXiv:2011.03395 (2020). Wynants, L. et al. Prediction models for diagnosis and prognosis of covid-19: systematic review and critical appraisal. bmj 369 , (2020). Roberts, M. et al. Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nature Machine Intelligence 3 , 199–217 (2021). Born, J. et al. On the role of artificial intelligence in medical imaging of COVID-19. Patterns 2 , 100269 (2021). Geirhos, R. et al. Shortcut learning in deep neural networks. Nature Machine Intelligence 2 , 665–673 (2020). DeGrave, A. J., Janizek, J. D. & Lee, S.-I. AI for radiographic COVID-19 detection selects shortcuts over signal. Nature Machine Intelligence 3 , 610–619 (2021). Oh, Y., Park, S. & Ye, J. C. Deep learning COVID-19 features on CXR using limited training data sets. IEEE transactions on medical imaging 39 , 2688–2700 (2020). Teixeira, L. O. et al. Impact of lung segmentation on the diagnosis and explanation of COVID-19 in chest X-ray images. Sensors 21, 7116 (2021). Selvaraju, R. R. et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. in Proceedings of the IEEE international conference on computer vision 618–626 (2017). Erion, G., Janizek, J. D., Sturmfels, P., Lundberg, S. M. & Lee, S.-I. Improving performance of deep learning models with axiomatic attribution priors and expected gradients. Nature machine intelligence 3 , 620–631 (2021). Viviano, J. D., Simpson, B., Dutil, F., Bengio, Y. & Cohen, J. P. Saliency is a possible red herring when diagnosing poor generalization. arXiv preprint arXiv:1910.00199 (2019). Johnson, A. E. et al. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data 6 , 1–8 (2019). Vayá, M. D. L. I. et al. Bimcv covid-19+: a large annotated dataset of rx and ct images from covid-19 patients. arXiv preprint arXiv:2006.01174 (2020). Wang, L., Lin, Z. Q. & Wong, A. Covid-net: A tailored deep convolutional neural network design for detection of covid-19 cases from chest x-ray images. Scientific Reports 10 , 1–12 (2020). Chambon, P. et al. RoentGen: Vision-Language Foundation Model for Chest X-ray Generation. Preprint at https://doi.org/10.48550/arXiv.2211.12737 (2022). Zhang, R. et al. A Generalizable Artificial Intelligence Model for COVID-19 Classification Task Using Chest X-ray Radiographs: Evaluated Over Four Clinical Datasets with 15,097 Patients. arXiv preprint arXiv:2210.02189 (2022). Deng, J. et al. ImageNet: A large-scale hierarchical image database. in 2009 IEEE Conference on Computer Vision and Pattern Recognition 248–255 (2009). doi:10.1109/CVPR.2009.5206848. Simonyan, K. & Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014). Huang, G., Liu, Z., Van Der Maaten, L. & Weinberger, K. Q. Densely connected convolutional networks. in Proceedings of the IEEE conference on computer vision and pattern recognition 4700–4708 (2017). Tan, M. & Le, Q. V. EfficientNetV2: Smaller Models and Faster Training. Preprint at https://doi.org/10.48550/arXiv.2104.00298 (2021). Liu, Z. et al. Swin transformer: Hierarchical vision transformer using shifted windows. in Proceedings of the IEEE/CVF International Conference on Computer Vision 10012–10022 (2021). Liu, Z. et al. A convnet for the 2020s. in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 11976–11986 (2022). Niven, T. & Kao, H.-Y. Probing neural network comprehension of natural language arguments. arXiv preprint arXiv:1907.07355 (2019). Du, M. et al. Towards interpreting and mitigating shortcut learning behavior of NLU models. arXiv preprint arXiv:2103.06922 (2021). Wang, H., Wu, X., Huang, Z. & Xing, E. P. High-frequency component helps explain the generalization of convolutional neural networks. in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 8684–8694 (2020). Jo, J. & Bengio, Y. Measuring the tendency of cnns to learn surface statistical regularities. arXiv preprint arXiv:1711.11561 (2017). Additional Declarations No competing interests reported. Supplementary Files Supplement.docx Cite Share Download PDF Status: Published Journal Publication published 04 Aug, 2023 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Major revision 04 Jul, 2023 Reviews received at journal 21 Jun, 2023 Reviewers agreed at journal 11 Jun, 2023 Reviewers invited by journal 02 Jun, 2023 Editor assigned by journal 02 Jun, 2023 Editor invited by journal 26 Apr, 2023 Submission checks completed at journal 26 Apr, 2023 First submitted to journal 14 Apr, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2818347","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":195312298,"identity":"d1bdd872-7db1-4302-8f33-5cad8ce0b1d1","order_by":0,"name":"Ran Zhang","email":"","orcid":"","institution":"University of Wisconsin–Madison","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ran","middleName":"","lastName":"Zhang","suffix":""},{"id":195312299,"identity":"9af9baf4-a1c9-4d96-969e-a67bec5aa3aa","order_by":1,"name":"Dalton Griner","email":"","orcid":"","institution":"University of Wisconsin–Madison","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Dalton","middleName":"","lastName":"Griner","suffix":""},{"id":195312300,"identity":"7da7c946-c690-46b0-bb5c-324b0d85012f","order_by":2,"name":"John W. Garrett","email":"","orcid":"","institution":"University of Wisconsin–Madison","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"John","middleName":"W.","lastName":"Garrett","suffix":""},{"id":195312301,"identity":"ee8175ff-25b0-48c3-99d6-c90065977f43","order_by":3,"name":"Zhihua Qi","email":"","orcid":"","institution":"Henry Ford Health","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Zhihua","middleName":"","lastName":"Qi","suffix":""},{"id":195312302,"identity":"10c8f8b6-00b7-4ea5-bd2a-2da2fce8cdc5","order_by":4,"name":"Guang-Hong Chen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAuklEQVRIiWNgGAWjYFACHgYGxn82DGxgDhuxWhjY0kjXchjKIUaLvPvZg48LeM4n9vEffsDwoewwYS2GZ/KSjWdI3E5sk0gzYJxxjhgtM3jMpHkMbue2SfAwMPO2EafF/DdPwrncNv4zDMx/idEiL8Fjxsxz4EBuG0MOAzMjMVoMeHKMpXkbkutBfjnYcy6dCFvazxh+5m2wM5bvP/zwwY8yayJsOYDEOYBDEZotDUQpGwWjYBSMghENAKKmNHzWWC+VAAAAAElFTkSuQmCC","orcid":"","institution":"University of Wisconsin–Madison","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Guang-Hong","middleName":"","lastName":"Chen","suffix":""}],"badges":[],"createdAt":"2023-04-14 17:59:03","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2818347/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2818347/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-023-39855-3","type":"published","date":"2023-08-04T21:51:23+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":36429942,"identity":"7b4c1d67-37d9-4a4f-9e78-1bfe954325f7","added_by":"auto","created_at":"2023-04-28 13:58:03","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":659484,"visible":true,"origin":"","legend":"\u003cp\u003eAcquisition Dependent Attributes (ADA) in chest x-ray images.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-2818347/v1/88721fa13fd67a1c834a2f4e.png"},{"id":36429941,"identity":"18f4b956-ba49-4062-9ecd-89f4d5773f47","added_by":"auto","created_at":"2023-04-28 13:58:03","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":57237,"visible":true,"origin":"","legend":"\u003cp\u003eAn overview of the datasets used in this work.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-2818347/v1/2ac81104e211ef5885fc3e8d.png"},{"id":36429944,"identity":"16c300fa-2f68-4a57-8257-ac7bca8c761f","added_by":"auto","created_at":"2023-04-28 13:58:03","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":680416,"visible":true,"origin":"","legend":"\u003cp\u003eA framework to train and certify shortcut detectives.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-2818347/v1/7e910e3f9edb62e4bfd560f8.png"},{"id":44734671,"identity":"ef7a341f-c450-4b2e-89f7-b1a618afed40","added_by":"auto","created_at":"2023-10-16 22:19:41","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1126883,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2818347/v1/b1eb3d23-0eb5-47e4-968c-19c7fd69bc0d.pdf"},{"id":36429943,"identity":"423c0233-7b93-4840-bf40-6c81cef5620e","added_by":"auto","created_at":"2023-04-28 13:58:03","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":572886,"visible":true,"origin":"","legend":"","description":"","filename":"Supplement.docx","url":"https://assets-eu.researchsquare.com/files/rs-2818347/v1/4750d1c882801dd0a228d5b9.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Training certified detectives to track down the intrinsic shortcuts in COVID-19 chest x-ray data sets","fulltext":[{"header":"Introduction","content":"\u003cp\u003eDeep learning has been incredibly successful in object detection, image classification, and natural language processing over the past ten years due to its ability to learn complex features from data. However, despite its success on benchmark datasets, there are limitations and practical issues when using these models in real-world scenarios. A major challenge is poor generalizability, where performance significantly drops when applied to external datasets\u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. This limits the translation and deployment of deep learning models for high-stakes tasks, such as in healthcare applications. This lesson was evident during the COVID-19 pandemic, where many machine learning models were developed, but very few performed well on real-world clinical tests\u003csup\u003e\u003cspan additionalcitationids=\"CR5\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe concept of shortcut learning has recently been explored in deep learning studies\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. It has been discovered that poor model generalizability can be attributed to shortcut learning when the training dataset has hidden shortcuts, meaning there are spurious correlations between irrelevant image features and the corresponding training labels. This causes models to quickly pick up these spurious correlations instead of the desired image features, establishing incorrect connections between input image data and output labels. For instance, early studies have shown that deep learning models can differentiate chest x-rays from different hospitals and patient groups\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. This suggests that different data sources and patient characteristics like gender, age, and race could also become shortcuts.\u003c/p\u003e \u003cp\u003eTo illustrate the issue of shortcut learning in real-world clinical scenarios, let's take the example of COVID-19 classification using chest x-ray radiographs (CXRs). DeGrave et al. \u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e discovered that if COVID-19 positive and negative training data were collected from two different sources, the model would only learn the source label as a shortcut for prediction. As a result, the model would not have the desired prediction power in real-world clinical scenarios. The authors also found that the trained model used extrinsic image features, such as lead markers, for prediction, even though these markers only indicated the orientation of patients in x-ray image acquisitions and did not correspond to any disease features. Some suggestions have been made to remove these extrinsic shortcuts by a segmentation step\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. However, even with these markers removed from the training dataset, some other shortcuts still existed in the segmented lung tissue-only training dataset\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. As a result, the nearly perfect performance of the trained model\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e is still not generalizable to real-world clinical datasets. The remaining shortcuts may be attributed to the inherent defining features of CXRs, such as image contrast and sharpness, which can vary from hospital to hospital due to different types of imaging systems, generations of x-ray imaging equipment, hardware components, image post-processing methods used by vendors, and imaging protocols used by technologists. All these factors can impact the digital representation of the acquired image data in terms of variations in image contrast and sharpness. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e demonstrates these variations.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eUnlike other shortcuts that have been previously studied, such as age, gender, race, and markers, which are extrinsic and can be removed through careful data collection and cleaning, contrast- and sharpness-related shortcuts are more difficult to detect and mitigate. This is because desired image features are also represented as image contrast and spatial correlations, which are similar to the features of contrast- and sharpness-related shortcuts. This entanglement between desired image features and shortcut features makes studying contrast and sharpness-related shortcuts particularly challenging. Consequently, these shortcuts are referred to as intrinsic shortcuts in this work.\u003c/p\u003e \u003cp\u003eIn order to develop an effective strategy for mitigating contrast- and sharpness-related shortcuts, it is necessary to first develop reliable methods for detecting their presence and severity within a carefully curated dataset. While post-hoc model interpretability methods, such as class activation maps\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e and expected gradient\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e, have been developed to identify relevant image features used by trained deep learning models for prediction, these methods are unable to detect intrinsic shortcuts within a curated training dataset prior to model training. Furthermore, studies have suggested that these methods may not be effective in diagnosing poor generalization performance of the model\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn this paper, we present a novel approach for detecting contrast- and sharpness-related intrinsic shortcuts using certified shortcut detective models. Our approach involves establishing qualification standards for suspected intrinsic shortcuts, designing a training curriculum for training the shortcut detectives to detect these shortcuts, performing certification tests on the trained detectives, and finally deploying them to curated datasets to examine the suspected shortcuts. We applied this approach to the available COVID-19 datasets to assess their quality. Our results demonstrate the effectiveness of this approach in detecting and mitigating intrinsic shortcuts.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eDatasets\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure 2 provides an overview of the datasets utilized in this study. The MIMIC-CXR dataset served as the training data for the shortcut detectives, while the HF-train dataset, a privately curated COVID-19 chest x-ray dataset, was utilized for certification tests of the trained detectives. The trained shortcut detectives were then applied to a variety of public and private chest x-ray datasets. Specific information about each dataset is provided below.\u003c/p\u003e \u003cp\u003eFigure 2. An overview of the datasets used in this work.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eMIMIC dataset\u003c/h2\u003e \u003cp\u003eMIMIC-CXR\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e chest x-ray dataset consists of 377,110 CXRs from 65,379 patients presenting to the Beth Israel Deaconess Medical Center Emergency Department between 2011\u0026ndash;2016. In this work, 46,894 frontal-view (AP/PA) normal CXRs (cases with \u0026ldquo;No Finding\u0026rdquo; labels) were used to train shortcut detectives.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eHF dataset\u003c/h2\u003e \u003cp\u003eThis is a privately curated COVID-19 CXR dataset from patients presenting to the Henry Ford Health between March 1, 2020, and October 31, 2020. The COVID-19-positive and COVID-19-negative cohorts are collected within the same time range, from the same hospitals, and labeled by their most recent RT-PCR test result seven days before or after the imaging study. For model training and internal testing, two data partitions are generated: \u003cb\u003eHF-train\u003c/b\u003e consisted of 8,733 COVID-19-positive CXRs from 4,383 patients and 16,584 COVID-19-negative CXRs from 8,733 patients; \u003cb\u003eHF-test\u003c/b\u003e consisted of 695 COVID-19-positive CXRs from 526 patients and 8,878 COVID-19-negative CXRs from 6,081 patients.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eBIMCV dataset\u003c/h2\u003e \u003cp\u003eThis is a public COVID-19 CXR dataset collected in Spain\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. This dataset was collected from 11 hospitals in the Valencian Region, Spain, between February and April 2020. After data curation, the dataset consisted of 4,169 COVID-19-positive CXRs from 2,663 patients and 5,050 COVID-19-negative CXRs from 3,710 patients.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eUW dataset\u003c/h2\u003e \u003cp\u003eThis is a privately curated COVID-19 CXR dataset. It includes consecutive patient cases from the University of Wisconsin Hospitals and Clinics (UW Health) from March 2020 to September 2021. The dataset comprised 1,025 COVID-19-positive CXRs from 658 patients and 8,774 COVID-19-negative CXRs from 5,953 patients.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eMIDRC dataset\u003c/h2\u003e \u003cp\u003eThis is a large, multi-institution public COVID-19 CXR dataset curated and released by the Medical Imaging \u0026amp; Data Resource Center (MIDRC). A total of 6,453 COVID-19-positive CXRs from 5,199 patients and 20,072 COVID-19-negative CXRs from 9,947 patients were pulled from the MIDRC Data Commons (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://data.midrc.org/\u003c/span\u003e\u003cspan address=\"https://data.midrc.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, date accessed: December 7th, 2022).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eCOVIDx dataset\u003c/h2\u003e \u003cp\u003eThis is a public COVID-19 CXR dataset released by the COVID-Net Open Source Initiative\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. A total of 29,986 CXRs from 16,648 patients are included in the training dataset (\u003cb\u003eCOVIDx-train\u003c/b\u003e), and 400 CXRs are included in the test dataset (COVIDx-test). (Data downloaded from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.kaggle.com/datasets/andyczhao/covidx-cxr2?select=competition_test\u003c/span\u003e\u003cspan address=\"https://www.kaggle.com/datasets/andyczhao/covidx-cxr2?select=competition_test\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, date accessed: December 7th, 2022).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eRoentGen-MIMIC dataset\u003c/h2\u003e \u003cp\u003eThis dataset contains 943 synthetic CXRs generated by the RoentGen model\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e and 1,000 real \u0026ldquo;No Finding\u0026rdquo; CXRs from the MIMIC dataset. The RoentGen model, which is trained using the MIMIC dataset, is able to generate visually convincing synthetic CXRs with different pathologies. To generate the synthetic CXRs used in this work, a text prompt of \u0026ldquo;No finding\u0026rdquo; was used as the input, only frontal view CXRs are included.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eTraining and certification of shortcut detectives\u003c/h3\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eOverview of the proposed framework\u003c/h2\u003e \u003cp\u003eAn overview of the shortcut detective training and certification process is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo begin, 46,894 normal CXRs from the MIMIC dataset were selected and randomly divided into two equal groups: 23,447 of the CXRs were assigned as the positive class (\u0026ldquo;1\u0026rdquo;) while the rest were assigned as the negative class (\u0026ldquo;0\u0026rdquo;).\u003c/p\u003e \u003cp\u003eThen, to construct the training dataset for the shortcut detective, the image contrast or the image sharpness of the positive class were adjusted using the approach shown in Appendix A1. Since only normal CXRs were included, there were no disease-specific features present that could be used to distinguish between the two classes. Therefore, if a model was able to differentiate between the two classes, it would be because the model had learned the corresponding global image contrast or sharpness characteristics, rather than disease features.\u003c/p\u003e \u003cp\u003eSubsequently, shortcut detectives (neural network models for binary classification) were trained using the constructed training datasets with contrast or sharpness shortcuts. Details of the training are discussed in the following section and in Appendix A2.\u003c/p\u003e \u003cp\u003eTo assess the efficacy of the shortcut detectives on COVID-19 CXR dataset, two types of examinations are necessary. Firstly, when the shortcut detective is deployed on a COVID-CXR dataset without the corresponding shortcut, it should not be able to differentiate between the COVID-positive and COVID-negative classes. Hence, an Area Under the Receiver Operating Characteristics curve (AUC) close to 0.5 is expected. This examination is crucial to ensure that the image features utilized by the shortcut detectives are not interwoven with the original imaging task. Secondly, when the shortcut detective is applied to a COVID-CXR dataset with a known shortcut, it should demonstrate superior classification performance, ideally with an AUC close to 1. For the first examination, the HF-train dataset is utilized. As the positive and negative cohorts are gathered within the same timeframe and from the same hospitals, it is expected that no contrast and sharpness shortcut exists. This has also been corroborated recently\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e, where a model trained on this dataset demonstrated consistent test performance on various external COVID-19 clinical test datasets. For the second examination, known shortcuts are integrated into the COVID-positive class or COVID-negative class of the HF-train dataset using the same procedures outlined in Appendix A1.\u003c/p\u003e \u003cp\u003eFinally, if the trained shortcut detectives pass the two exams, i.e. AUC of close to 0.5 on the shortcut-free dataset and AUC of close to 1.0 on the shortcut-present dataset, they are referred to as certified shortcut detectives.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eImage preprocessing and model architecture\u003c/h2\u003e \u003cp\u003eFor the MIMIC, HF, BIMCV, UW Health and MIDRC datasets, the original DICOM images are converted to 8-bit png format with a size of 224-by-224 using the default window level and window width. For the COVIDx dataset, images are directly resized to 224-by-224.\u003c/p\u003e \u003cp\u003eTo train the shortcut detectives, five different model architectures (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e) that are broadly used for image classification with state-of-the-art performance on the ImageNet\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e classification tasks are investigated in this work. Although we cannot exhaust all the possible model architectures for this purpose, the models we investigated include classic and modern Convolutional Neural Networks (CNN) and the recently introduced Swin Transformer. These models vary in architectural design and complexity (number of model parameters and floating-point operations, FLOPs). For each model architecture, an ensemble of five individually trained models with different training-validation splits are used. More technical details on the model training are shown in Appendix A2.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eList of different model architectures studied in this work\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of parameters\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFLOPs\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eImageNet Accuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eYear developed\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVGG-16\u003csup\u003e20\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e138 M\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e31 B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e73.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2014\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDenseNet-121\u003csup\u003e21\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8 M\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5.7 B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e74.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2017\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEfficientNet\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e54 M\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e24 B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e85.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2020\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSwin Transformer\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e88 M\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e15 B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e83.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2021\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eConvNeXt\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e89 M\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e15 B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e84.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2022\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003eNote: ImageNet accuracy data obtained from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pytorch.org/vision/stable/models.html#classification\u003c/span\u003e\u003cspan address=\"https://pytorch.org/vision/stable/models.html#classification\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eDeployment of the shortcut detectives\u003c/h2\u003e \u003cp\u003eCertified shortcut detectives are deployed to detect shortcuts in real-world datasets, including BIMCV, UW, MIDRC, COVIDx, and RoentGen-MIMIC. We also trained two COVID-19 classification models using \u003cb\u003eHF-train\u003c/b\u003e dataset and \u003cb\u003eCOVIDx-train\u003c/b\u003e dataset and compared their generalizability using internal and external tests.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eStatistics\u003c/h2\u003e \u003cp\u003eThe 95% confidence intervals (CI) for the AUC were calculated using the statistical software R (version 4.0.0) with the pROC package. CIs were calculated using the bootstrap method with 2000 bootstrap replicates.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eCertification of the shortcut detectives\u003c/h2\u003e \u003cp\u003eAs shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, the trained detectives successfully passed the two exams. Note that an AUC of 0.0 is simply due to the assignment of class labels, which is equivalent to an AUC of 1.0. Both indicate perfect classification performance. It is also shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e that all five neural network architectures achieve a similar performance level. This result aligns well with our understanding that contrast and sharpness shortcuts are intrinsic to the dataset and the task.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eShortcut detection in real-world datasets\u003c/h2\u003e \u003cp\u003eUsing the certified shortcut detectives, we investigated several curated COVID-CXR datasets for possible shortcuts. UW Health is a single-site, privately curated dataset; BIMCV and MIDRC are both multi-institutional public datasets; COVIDx is the first open access COVID-CXR dataset made available by the experts in the computer science community.\u003c/p\u003e \u003cp\u003eSimilar to the detective certification process, for a COVID-CXR dataset where COVID-19-positive cases are assigned a label \u0026ldquo;1\u0026rdquo; and negative cases are given a label \u0026ldquo;0\u0026rdquo;, if the shortcut detective can differentiate the two classes (AUC significantly deviates from 0.50), it indicates the existence of a corresponding shortcut. For the RoentGen-MIMIC dataset, real CXRs are labelled as \u0026ldquo;1\u0026rdquo; and synthetic CXRs are labelled as \u0026ldquo;0\u0026rdquo;.\u003c/p\u003e \u003cp\u003eThe results presented in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e were obtained using shortcut detectives based on the DenseNet model. The results confirm the presence of image sharpness and contrast shortcuts in the COVIDx dataset, which can be exploited by models trained on such data and compromise their generalizability in real clinical settings. Conversely, the other three datasets that were curated by medical professionals exhibit no such shortcuts. We conducted a performance comparison between two models trained on COVIDx and Henry Ford datasets, respectively, which are of similar size, but the former has both sharpness and contrast shortcuts as identified by the shortcut detectives. The results displayed in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e indicate that the COVIDx model exhibits poor generalization performance, as evidenced by the large AUC gap between internal and external tests. In contrast, the HF model exhibits consistent performance on both internal and external tests. Additionally, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, some contrast and sharpness differences were also detected in the RoentGen-MIMIC dataset. While the generated synthetic CXRs appear visually realistic, caution must be exercised when using them for AI model development due to the potential learning of shortcuts caused by the inherent contrast and sharpness differences between real and synthetic data.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCertification of the shortcut detectives\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVGG\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDENSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEff\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eSwin\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eConv\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"6\" nameend=\"c6\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eADA(S) shortcut detective\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExam 1:\u003c/p\u003e \u003cp\u003eHF-train\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.49\u003c/b\u003e [0.48,0.50]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.56\u003c/b\u003e [0.56,0.57]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.56\u003c/b\u003e [0.55,0.57]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.53\u003c/b\u003e [0.52,0.53]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.54\u003c/b\u003e [0.53,0.54]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExam 2a:\u003c/p\u003e \u003cp\u003eWith ADA-S(+)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExam 2b:\u003c/p\u003e \u003cp\u003eWith ADA-S(-)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"6\" nameend=\"c6\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eADA(C) shortcut detective\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExam 1:\u003c/p\u003e \u003cp\u003eHF-train\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.47\u003c/b\u003e [0.46,0.48]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.48\u003c/b\u003e [0.47,0.49]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.47\u003c/b\u003e [0.46,0.48]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.49\u003c/b\u003e [0.48,0.49]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.48\u003c/b\u003e [0.47,0.48]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExam 2a:\u003c/p\u003e \u003cp\u003eWith ADA-C(+)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.00 [1.00,1.00]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExam 2b:\u003c/p\u003e \u003cp\u003eWith ADA-C(-)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.00\u003c/b\u003e [0.00,0.00]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"6\" nameend=\"c6\" namest=\"c1\"\u003e \u003cp\u003eNote: (+) means the shortcut is added to the COVID-positive CXRs; (-) means the shortcut is added to the COVID-negative CXRs.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eShortcut detection on real-world COVID-CXR datasets (DenseNet model result)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCOVIDx\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRoentGen-MIMIC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUW\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eBIMCV\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eMIDRC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eADA(S) shortcut\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.84 [0.83,0.84]\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.37\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e[0.34,0.39]\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.50 [0.48,0.52]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.47 [0.45,0.48]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.45 [0.45,0.46]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eADA(C) shortcut\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.81 [0.80,0.81]\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.05\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e[0.04, 0.06]\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.53 [0.51,0.55]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.56 [0.55,0.57]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.53 [0.52,0.54]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of two trained models\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eInternal\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUW\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBIMCV\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMIDRC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOVIDx\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003cp\u003e[1.00, 1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.60 [0.58,0.62]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.61\u003c/p\u003e \u003cp\u003e[0.60,0.62]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.57 [0.56,0.58]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.78 [0.76,0.80]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.77 [0.76,0.79]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003cp\u003e[0.78,0.80]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.75 [0.74,0.75]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eShortcut learning has been a topic of interest in the machine learning community, particularly in computer vision (CV) and natural language processing (NLP). Researchers have explored shortcut learning behavior from different perspectives, such as underspecification\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e, shortcut learning in various NLP tasks\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e, and mitigation strategies for domain-knowledge agnostic models\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. Notably, it has been observed that convolutional neural networks tend to rely on content with high spatial frequency or strong local correlations to establish connections between input and labels in CV and NLP\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e,\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. However, it remains unclear whether these observations can be extended to medical diagnosis tasks, where clinical datasets have distinct characteristics from those in ImageNet. Therefore, studying shortcut learning for well-defined, clinically relevant tasks using real-world clinical datasets is crucial for medical AI applications.\u003c/p\u003e \u003cp\u003eIn this study, we demonstrated that acquisition-dependent attributes (ADAs), such as image contrast and sharpness differences arising from the entire image generation pipeline, can serve as intrinsic shortcuts during clinical diagnosis learning tasks. Inadequate quality control procedures during data collection can allow these shortcuts to inadvertently infiltrate the curated dataset. If shortcuts contaminate the dataset, neural networks can easily exploit them during training, thereby impairing their ability to generalize to other real-world datasets.\u003c/p\u003e \u003cp\u003eThus, it is imperative to identify possible shortcuts in the training dataset prior to model development. In this study, we present a methodical framework for training and validating shortcut detectives for chest X-ray classification, with emphasis on image contrast and sharpness - two essential intrinsic characteristics of chest X-ray images. However, it should be noted that these are not the only possible shortcuts that may exist in chest X-ray datasets. If other intrinsic shortcuts are suspected in a collected dataset, the general framework presented in this work can be utilized to construct similar shortcut detectives and identify alleged intrinsic shortcuts.\u003c/p\u003e \u003cp\u003eOnce the intrinsic shortcuts are identified, it is essential to develop strategies to mitigate their impact on the learned models. One possible approach is to develop standardization and normalization techniques for image contrast and sharpness to adjust these attributes without affecting the disease features. Alternatively, examining proven intrinsic shortcut-free datasets, such as the baseline dataset (Henry Ford Health) and the three additional datasets (UW Health, BIMCV, and MIDRC) shown to be free of intrinsic shortcuts in this study, can provide further insight on how to avoid these shortcuts in the data curation process. However, it is worth noting that this is a limitation of the present work, and future research should investigate the development of mitigation strategies for the identified intrinsic shortcuts.\u003c/p\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData Availability Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBIMCV(https://bimcv.cipf.es/bimcv-projects/bimcv-covid19/), MIDRC (https://data.midrc.org/) and MIMIC-CXR (https://physionet.org/content/mimic-cxr/2.0.0/) datasets are publicly available. Henry Ford Dataset and UW Health Dataset are private datasets with protected patient information. The python script used for model training and evaluation; weights of all trained models can be found on the GitHub repository: https://github.com/uw-ctgroup/shortcut.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments:\u003c/strong\u003e This study was partially supported by NIH Grants 3U01EB021183-W4, R01HL153594, R01EB032474, and a grant from the Wisconsin Partnership Program.\u003c/p\u003e\n\u003cp\u003eThe imaging and associated clinical data downloaded from MIDRC (The Medical Imaging Data Resource Center) and used for research in this [publication/press release] were made possible by the National Institute of Biomedical Imaging and Bioengineering (NIBIB) of the National Institutes of Health under contracts 75N92020C00008 and 75N92020C00021. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions Statement:\u003c/strong\u003e R.Z. and G.H.C. conceived the study and designed the experiments. J.W.G. and Z.Q. curated two private clinical datasets for tests. R.Z. prepared the datasets for model training and testing. D.G. and R.Z. wrote the software code. R.Z. and G.H.C. performed the statistical analysis. G.H.C. supervised the study. R.Z. and G.H.C. wrote the manuscript. All authors read and edited the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eZech, J. R. \u003cem\u003eet al.\u003c/em\u003e Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. \u003cem\u003ePLoS medicine\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, e1002683 (2018).\u003c/li\u003e\n\u003cli\u003eHendrycks, D., Zhao, K., Basart, S., Steinhardt, J. \u0026amp; Song, D. Natural adversarial examples. in \u003cem\u003eProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition\u003c/em\u003e 15262\u0026ndash;15271 (2021).\u003c/li\u003e\n\u003cli\u003eD\u0026rsquo;Amour, A. \u003cem\u003eet al.\u003c/em\u003e Underspecification presents challenges for credibility in modern machine learning. \u003cem\u003earXiv preprint arXiv:2011.03395\u003c/em\u003e (2020).\u003c/li\u003e\n\u003cli\u003eWynants, L. \u003cem\u003eet al.\u003c/em\u003e Prediction models for diagnosis and prognosis of covid-19: systematic review and critical appraisal. \u003cem\u003ebmj\u003c/em\u003e \u003cstrong\u003e369\u003c/strong\u003e, (2020).\u003c/li\u003e\n\u003cli\u003eRoberts, M. \u003cem\u003eet al.\u003c/em\u003e Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. \u003cem\u003eNature Machine Intelligence\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 199\u0026ndash;217 (2021).\u003c/li\u003e\n\u003cli\u003eBorn, J. \u003cem\u003eet al.\u003c/em\u003e On the role of artificial intelligence in medical imaging of COVID-19. \u003cem\u003ePatterns\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, 100269 (2021).\u003c/li\u003e\n\u003cli\u003eGeirhos, R. \u003cem\u003eet al.\u003c/em\u003e Shortcut learning in deep neural networks. \u003cem\u003eNature Machine Intelligence\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, 665\u0026ndash;673 (2020).\u003c/li\u003e\n\u003cli\u003eDeGrave, A. J., Janizek, J. D. \u0026amp; Lee, S.-I. AI for radiographic COVID-19 detection selects shortcuts over signal. \u003cem\u003eNature Machine Intelligence\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 610\u0026ndash;619 (2021).\u003c/li\u003e\n\u003cli\u003eOh, Y., Park, S. \u0026amp; Ye, J. C. Deep learning COVID-19 features on CXR using limited training data sets. \u003cem\u003eIEEE transactions on medical imaging\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, 2688\u0026ndash;2700 (2020).\u003c/li\u003e\n\u003cli\u003eTeixeira, L. O. et al. Impact of lung segmentation on the diagnosis and explanation of COVID-19 in chest X-ray images. Sensors 21, 7116 (2021).\u003c/li\u003e\n\u003cli\u003eSelvaraju, R. R. et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. in Proceedings of the IEEE international conference on computer vision 618\u0026ndash;626 (2017).\u003c/li\u003e\n\u003cli\u003eErion, G., Janizek, J. D., Sturmfels, P., Lundberg, S. M. \u0026amp; Lee, S.-I. Improving performance of deep learning models with axiomatic attribution priors and expected gradients. \u003cem\u003eNature machine intelligence\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 620\u0026ndash;631 (2021).\u003c/li\u003e\n\u003cli\u003eViviano, J. D., Simpson, B., Dutil, F., Bengio, Y. \u0026amp; Cohen, J. P. Saliency is a possible red herring when diagnosing poor generalization. \u003cem\u003earXiv preprint arXiv:1910.00199\u003c/em\u003e (2019).\u003c/li\u003e\n\u003cli\u003eJohnson, A. E. \u003cem\u003eet al.\u003c/em\u003e MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. \u003cem\u003eScientific data\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, 1\u0026ndash;8 (2019).\u003c/li\u003e\n\u003cli\u003eVay\u0026aacute;, M. D. L. I. \u003cem\u003eet al.\u003c/em\u003e Bimcv covid-19+: a large annotated dataset of rx and ct images from covid-19 patients. \u003cem\u003earXiv preprint arXiv:2006.01174\u003c/em\u003e (2020).\u003c/li\u003e\n\u003cli\u003eWang, L., Lin, Z. Q. \u0026amp; Wong, A. Covid-net: A tailored deep convolutional neural network design for detection of covid-19 cases from chest x-ray images. \u003cem\u003eScientific Reports\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 1\u0026ndash;12 (2020).\u003c/li\u003e\n\u003cli\u003eChambon, P. \u003cem\u003eet al.\u003c/em\u003e RoentGen: Vision-Language Foundation Model for Chest X-ray Generation. Preprint at https://doi.org/10.48550/arXiv.2211.12737 (2022).\u003c/li\u003e\n\u003cli\u003eZhang, R. \u003cem\u003eet al.\u003c/em\u003e A Generalizable Artificial Intelligence Model for COVID-19 Classification Task Using Chest X-ray Radiographs: Evaluated Over Four Clinical Datasets with 15,097 Patients. \u003cem\u003earXiv preprint arXiv:2210.02189\u003c/em\u003e (2022).\u003c/li\u003e\n\u003cli\u003eDeng, J. \u003cem\u003eet al.\u003c/em\u003e ImageNet: A large-scale hierarchical image database. in \u003cem\u003e2009 IEEE Conference on Computer Vision and Pattern Recognition\u003c/em\u003e 248\u0026ndash;255 (2009). doi:10.1109/CVPR.2009.5206848.\u003c/li\u003e\n\u003cli\u003eSimonyan, K. \u0026amp; Zisserman, A. Very deep convolutional networks for large-scale image recognition. \u003cem\u003earXiv preprint arXiv:1409.1556\u003c/em\u003e (2014).\u003c/li\u003e\n\u003cli\u003eHuang, G., Liu, Z., Van Der Maaten, L. \u0026amp; Weinberger, K. Q. Densely connected convolutional networks. in \u003cem\u003eProceedings of the IEEE conference on computer vision and pattern recognition\u003c/em\u003e 4700\u0026ndash;4708 (2017).\u003c/li\u003e\n\u003cli\u003eTan, M. \u0026amp; Le, Q. V. EfficientNetV2: Smaller Models and Faster Training. Preprint at https://doi.org/10.48550/arXiv.2104.00298 (2021).\u003c/li\u003e\n\u003cli\u003eLiu, Z. \u003cem\u003eet al.\u003c/em\u003e Swin transformer: Hierarchical vision transformer using shifted windows. in \u003cem\u003eProceedings of the IEEE/CVF International Conference on Computer Vision\u003c/em\u003e 10012\u0026ndash;10022 (2021).\u003c/li\u003e\n\u003cli\u003eLiu, Z. \u003cem\u003eet al.\u003c/em\u003e A convnet for the 2020s. in \u003cem\u003eProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition\u003c/em\u003e 11976\u0026ndash;11986 (2022).\u003c/li\u003e\n\u003cli\u003eNiven, T. \u0026amp; Kao, H.-Y. Probing neural network comprehension of natural language arguments. \u003cem\u003earXiv preprint arXiv:1907.07355\u003c/em\u003e (2019).\u003c/li\u003e\n\u003cli\u003eDu, M. \u003cem\u003eet al.\u003c/em\u003e Towards interpreting and mitigating shortcut learning behavior of NLU models. \u003cem\u003earXiv preprint arXiv:2103.06922\u003c/em\u003e (2021).\u003c/li\u003e\n\u003cli\u003eWang, H., Wu, X., Huang, Z. \u0026amp; Xing, E. P. High-frequency component helps explain the generalization of convolutional neural networks. in \u003cem\u003eProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition\u003c/em\u003e 8684\u0026ndash;8694 (2020).\u003c/li\u003e\n\u003cli\u003eJo, J. \u0026amp; Bengio, Y. Measuring the tendency of cnns to learn surface statistical regularities. \u003cem\u003earXiv preprint arXiv:1711.11561\u003c/em\u003e (2017).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-2818347/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2818347/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDeep learning faces a significant challenge wherein the trained models often underperform when used with external test data sets. This issue has been attributed to spurious correlations between irrelevant features in the input data and corresponding labels. This study uses the classification of COVID-19 from chest x-ray radiographs as an example to demonstrate that the image contrast and sharpness, which are characteristics of a chest radiograph dependent on data acquisition systems and imaging parameters, can be intrinsic shortcuts that impair the model\u0026rsquo;s generalizability. The study proposes training certified shortcut detective models that meet a set of qualification criteria which can then identify these intrinsic shortcuts in a curated data set.\u003c/p\u003e","manuscriptTitle":"Training certified detectives to track down the intrinsic shortcuts in COVID-19 chest x-ray data sets","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-04-28 13:57:58","doi":"10.21203/rs.3.rs-2818347/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-07-04T04:04:01+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-06-21T19:52:59+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"a632f382-7e7f-4010-896e-9fe4115ae438","date":"2023-06-12T00:24:34+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-06-02T22:25:55+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-06-02T22:22:31+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2023-04-26T12:09:57+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-04-26T12:02:51+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2023-04-14T17:44:31+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"89c0f5e7-6b69-461a-8086-be9e9683ae18","owner":[],"postedDate":"April 28th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-10-16T22:01:53+00:00","versionOfRecord":{"articleIdentity":"rs-2818347","link":"https://doi.org/10.1038/s41598-023-39855-3","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2023-08-04 21:51:23","publishedOnDateReadable":"August 4th, 2023"},"versionCreatedAt":"2023-04-28 13:57:58","video":"","vorDoi":"10.1038/s41598-023-39855-3","vorDoiUrl":"https://doi.org/10.1038/s41598-023-39855-3","workflowStages":[]},"version":"v1","identity":"rs-2818347","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2818347","identity":"rs-2818347","version":["v1"]},"buildId":"ehx78VzkSd0WSzXnipQa-","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-4.0