Towards Accurate Detection of Axial Spondyloarthritis by Using Deep Learning to Capture Sacroiliac Joints on Plain Radiographs. | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research article Towards Accurate Detection of Axial Spondyloarthritis by Using Deep Learning to Capture Sacroiliac Joints on Plain Radiographs. Nils Friedrich Grauhan, Keno Kyrill Bressem, Yves Nicolas Manzoni, and 8 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-379664/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Well-informed decisions about how to best treat patients with axial spondyloarthritis (SpA) regularly include an evaluation of the sacroiliac joints (SIJ) on plain radiographs. However, grading radiographic findings correctly has proven to be a considerable challenge to expert readers as well as to state-of-the-art convolutional neural networks (CNNs). A method to reduce image information to the clinically relevant core would undoubtedly lead to more accurate results. We, therefore, trained a CNN only to detect SIJs on radiographs and evaluated its potential as a preprocessing pipeline in the automated classification of SpA. Materials and Methods We employed a CNN of the RetinaNet architecture, which was trained on a total of 423 plain radiographs of the sacroiliac joints (SIJs). Images were taken from two completely independent datasets. Training and tuning were performed on image data from the Patients With Axial Spondyloarthritis (PROOF) study and testing was executed using images from the German Spondyloarthritis Inception Cohort (GESPIC). Performance was evaluated by manual review and standard object detection metrics from PASCAL and Microsoft COCO. Results The CNN produced excellent results in detecting SIJs on the tuning (n =106) and on the holdout dataset (n =140). Object detection metrics for the tuning data were AP = 0.996 and mAP = 0.538; values for the independent holdout data were AP = 0.981 and mAP = 0.515. Conclusions The developed CNN was highly accurate in detecting SIJs on radiographs. Such a model could increase the reliability of deep learning-based algorithms in detecting and grading SpA. Health Economics & Outcomes Research Axial spondyloarthritis sacroiliac joint radiograph deep learning convolutional neural network object detection Figures Figure 1 Figure 2 Figure 3 Introduction Plain radiographs of the sacroiliac joints together with magnetic resonance imaging (MRI) play an important role in detecting axial spondyloarthritis (SpA) (1, 2). However, accuracy particularly in grading pathological findings in radiographs according to the modified New York criteria can differ significantly between observers (3, 4). Deep learning-based methods using convolutional neural networks (CNN) have recently demonstrated precision levels of human experts in reading medical images (5-7). They may therefore aid diagnosis of SpA as they have proven to excel when confronted with a specific task (8). Such an approach was recently pursued by training a CNN to detect definite radiographic sacroiliitis on plain radiographs (9). The results demonstrated a high accuracy of the developed CNN. While these are excellent results and comparable to the accuracy of domain experts, performance may still have been compromised: Firstly, computational capacity put a limit on the resolution of the radiographs included in the study, resulting in a significant downsizing. Although downsizing is a common practice in deep learning this entails the risk of losing crucial image information (10). Secondly, studies have shown that a CNN may not base its prediction solely on clinically relevant features (11). An excellent way to visually track a model’s decision is to use Gradient-weighted Class Activation Mappings (Grad-CAMs) (12). Grad-CAM highlights regions that have a strong influence on the model’s decision. In reviewing images that were incorrectly predicted in the holdout dataset of our previous study, several images were found where the model had based its decision on aspects unrelated to SpA features. The aim of this study was therefore to train and validate a CNN in the detection and segmentation of both sacroiliac joints on plain radiographs by drawing bounding boxes around the region of interest. If proven to be accurate, this CNN could be integrated in a future preprocessing pipeline augmenting the actual detection or classification of SpA no longer restrained by low image resolution or irrelevant information. To test the generalizability of our results, image data for training and testing were taken from two completely independent datasets. Methods Data In this study, we used plain radiographs retrieved from two completely separate image collections: a) Patients With Axial Spondyloarthritis: Multic o untry Registry of Clinical Characteristics (PROOF) is a continuous study involving medical facilities in 29 countries. Overall 2170 adult patients diagnosed with SpA (non-radiographic or radiographic) ≤12 months before study enrolment and fulfilling the ASAS classification criteria for SpA are included (13). B) German Spondyloarthritis Inception Cohort (GESPIC) is a multicenter inception cohort study conducted in Germany that includes 646 patients with SpA (14). Pre-processing and Segmentation Both PROOF and GESPIC datasets contain radiographs of sacroiliac joints in DICOM format (Digital Imaging and Communications in Medicine). All grey-scale values were adjusted to a unified representation state using the Horos Project DICOM Viewer (version 4.0.0, www.horosproject.org). Afterward, all images were converted to the Tagged Image File Format (TIFF). Annotations for training, tuning and holdout datasets were done using the COCO Annotator (https://github.com/jsbroks/coco-annotator). Bounding boxes were placed by one expert radiologist (JLV) separately around each of the sacroiliac joints (SIJ) on every radiograph. A total of 529 annotated images from the PROOF dataset was available and then randomly split into a training dataset (80%, n=423) and a tuning dataset (20%, n=106). A total of 140 annotated images from GESPIC was available and taken to form the holdout dataset (see flowchart in Figure 1). Model Training We used a RetinaNet with a ResNet-50 backbone, pre-trained in ImageNet, specifically suited for object detection (15). Model training was performed using IceVision (https://airctic.com, version 0.5.2), which is an object-detection framework built on top of Python (https://www.python.org, version 3.8), as well as the Fastai (https://fast.ai, version 2.2.5) and PyTorch (https://pytorch.org, version 1.7.0) libraries. Training was executed on a dedicated Ubuntu 18.04 Workstation with two Nvidia GeForce RTX 2080ti Graphic cards. All input images were resized and cropped to 512 x 512 px. Random transformations were used to augment training data. During model training, an early stopping approach was utilized to avoid overfitting. As we used a pre-trained ResNet50 as a backbone, we disabled weight updates for the backbone during the first 53 training epochs. Then we trained for 27 more epochs, updating all weights of the model. The model with the best performance on the tuning data was exported and used for inference on the holdout data. Statistical Analysis Model performance was evaluated using two separate approaches: first, in a swift manual review, two expert readers (NFG, YNM) independently checked the results of the model on the holdout data and counted all images in which the model either partially or completely failed to detect the SIJs (see Table 1). Second, average precision (AP) and mean average precision (mAP) were calculated based on Intersection over Union (IoU). This metric is commonly used in object detection competitions such as the PASCAL Visual Object Classes 2012 (VOC 2012) challenge (16) or in the Microsoft Common objects in context (COCO) challenge (17). IoU is an expression for the overlap of the predicted area and the actual image (ground truth) divided by the area enclosed by both (see Figure 2). A standard requirement to identify a prediction as a true positive is a ratio of 0.5 (17). Results In a manual review of the inferred segmentations in the tuning data, both observers found no instances and only one instance in the holdout data in which the model had only partially captured the SIJ. In the tuning dataset, one instance was found in which the model had missed the SIJ, and seven instances in the holdout dataset. All remaining 105 (tuning data, 99.1%) and 132 (holdout data, 94.3%) instances were counted as fully captured by both observers. There was no disagreement between the two observers in assigning each image to the above categories. Object detection metrics calculated from IoU were at 0.538 for mean average precision and 0.996 for average precision in the tuning dataset. In the holdout dataset, values were 0.515 for mean average precision and 0.981 for average precision. The results are summarized in Table 1 . Discussion Our model showed excellent results both in the manual review and displayed by high AP and mAP values. In addition, we would like to highlight that the manual review revealed only one instance of a partial capture of the SIJ. This underlines the fact that the model was very accurate and that a further increase in AP and mAP values would not necessarily reflect a more precise capture of the actual SIJ. Also, it should be noted that in the majority of cases where the model failed on holdout data radiographs included X-ray protective devices commonly absent in all the training and tuning data (see examples in Figure 3). All of this suggests that such a model could very well serve as a reliable pre-stage placed ahead of a second CNN charged with detecting specific features relevant to SpA. Only a few studies have tried to use deep learning approaches to either detect SpA in a simple binary task or to grade them according to the modified New York Criteria. The study by Türk et al. is the only one known to us that relied on conventional radiographs to grade SpA (18). In their work authors report that while their model performed well compared to human observers accuracy still dropped when asked to differentiate between early stages of SpA (grade 0 and grade 1). Notably, no segmentation of the included images was performed. Therefore, we are confident that isolating SIJs on radiographs before predicting grades of SpA will boost model performance, especially in detecting these early stages. We base this judgment on the possibility to reduce the need for downsizing images that contain rather subtle image information relevant to a correct prediction. By automatically cropping bounding boxes containing the SIJ and resizing the snippet to its original high resolution, we could achieve an efficient pipeline layout. Such an efficient pipeline is very important in grading SpA as the anatomy of SIJs is complex and features are often difficult to detect even for expert readers (19, 20). Also, lowering image resolution to facilitate model training can lead to unnecessary bias (10). Deep learning models are known to be susceptible to various confounders in medical image data, such as the gender or age of the participants. (21, 22). Such confounders pose a serious problem as they can lead to overestimating the generalizability of the results (10). Furthermore, the frequently described “black box nature” of deep learning-based algorithms carries the risk that further systematic errors may be overlooked (23, 24). We are convinced that reducing image information to the clinically relevant core will help to decrease these errors. An important prerequisite to verify the generalizability of the results is to check the performance of a CNN on holdout data from external datasets (25). As with our previous study, we have chosen two completely independent datasets for training and testing. Detection metrics used to evaluate the tuning and the holdout data (AP and mAP) were very close and therefore imply a robust performance of our model. Beyond the application of detecting SpA, our study promises to be valuable in the broader field of digital applications in Rheumatology. Radiographic diagnostics in Rheumatology often rely on very subtle changes in radiographs such as small bony erosions, soft tissue swelling, or joint space narrowing. A similar approach could, therefore, potentially enhance any automated detection of more delicate findings on plain radiographs. Conclusions Convolutional neural networks are highly accurate in detecting sacroiliac joints on plain radiographs and may therefore aid in a more focused image analysis of these regions. Abbreviations AP: Average precision; ASAS: Assessment of Spondylarthritis international Society; CNN: Convolutional neural network; COCO: Microsoft Common objects in context; DICOM: Digital Imaging and Communications in Medicine; GESPIC: German Spondyloarthritis Inception Cohort; IoU: Intersection over union; mAP: mean Average precision; MRI: Magnetic Resonance Imaging; PASCAL VOC 2012: Pascal Visual Object Classes Challenge 2012; PROOF: P atients With Axial Spondyloarthritis (study); SIJ: Sacroiliac joint; SpA: axial Spondyloarthritis Declarations Ethical Approval and Consent to participate Both PROOF and GESPIC cohorts were approved by the local ethics committees of each study center in accordance with the local laws and regulations and is being conducted in accordance with the Declaration of Helsinki and Good Clinical Practice. GESPIC was additionally approved by a central ethics committee of the coordinating center. Written informed consent to participate was obtained from all patients. Consent for publication Not applicable. Availability of supporting data The data that support the findings of this study are available from the corresponding author, N.F.G., upon reasonable request. Competing Interests: N.F.G, K.K.B., L.C.A., and Y.N.M. have nothing to disclose. J.L.V reports non-financial support from Bayer, non-financial support from Guerbet, non-financial support from Medtronic, personal fees and non-financial support from Merit Medical, outside the submitted work. S.M.N. reports grants from the German Research Foundation (Deutsche Forschungsgemeinschaft), personal fees from Bracco Imaging, personal fees from Canon Medical Systems, personal fees from Guerbet, personal fees from Teleflex Medical GmbH, personal fees from Bayer Vital GmbH, outside the submitted work. H.H. reports personal fees from Pfizer, personal fees from Janssen, personal fees from Novartis, personal fees from Roche, personal fees from MSD, outside the submitted work. F.P. reports personal fees from AbbVie, personal fees from AMGEN, personal fees from BMS, personal fees from Celgene, from MSD, grants and personal fees from Novartis, personal fees from Pfizer, from Roche, personal fees from UCB, outside the submitted work. M.R. received honoraria and/or consulting fees from AbbVie, BMS, Celgene, Janssen, Eli Lilly, MSD, Novartis, Pfizer, Roche, UCB Pharma. D.P. reports grants and personal fees from AbbVie, during the conduct of the study; grants and personal fees from AbbVie, personal fees from BMS, personal fees from Celgene, grants and personal fees from Lilly, grants and personal fees from MSD, grants and personal fees from Novartis, grants and personal fees from Pfizer, personal fees from Roche, personal fees from UCB, outside the submitted work. V.R.R. reports personal fees from Abbvie, personal fees from Novartis, outside the submitted work. Funding: GESPIC was initially supported by the German Federal Ministry of Education and Research (Bundesministerium für Bildung und Forschung, BMBF). As scheduled, BMBF was reduced in 2005 and stopped in 2007; thereafter, complementary financial support was obtained also from Abbott, Amgen, Centocor, Schering–Plough, and Wyeth. Starting in 2010, the core GESPIC cohort was supported by AbbVie. The PROOF study is funded by AbbVie. We thank AbbVie for allowing us to use the PROOF dataset for the aim of the current study. Author’s contributions: NFG and JLV: Data preparation, execution of model training and testing, statistical analysis, drafting and finalizing the manuscript. KKB, YNM, DP, SMN and LCA: Essential contributions to the concept of the study, editing and correcting the manuscript. VRR, FP, HH, and MR: Editing and correcting the manuscript. Authors’ information 1 Charité – Universitätsmedizin Berlin, Department of Radiology, Berlin, Germany 2 Berlin Institute of Health (BIH), Berlin, Germany 3 Charité – Universitätsmedizin Berlin, Department of Gastroenterology, Infectiology and Rheumatology, Berlin, Germany 4 Department of Internal Medicine and Rheumatology, Klinikum Bielefeld Rosenhöhe, Bielefeld, Germany Acknowledgements: LCA is participant in the BIH-Charité Clinician Scientist Program funded by the Charité –Universitätsmedizin Berlin and the Berlin Institute of Health. KKB is participant in the BIH-Charité Digital Clinician Scientist Program funded by the Charité –Universitätsmedizin Berlin and the Berlin Institute of Health. References Slobodin G, Hussein H, Rosner I, Eshed I. Sacroiliitis–early diagnosis is key. Journal of inflammation research. 2018;11:339. Heuft-Dorenbosch L, Landewe R, Weijers R, Wanders A, Houben H, Van Der Linden S, et al. Combining information obtained from magnetic resonance imaging and conventional radiographs to detect sacroiliitis in patients with recent onset inflammatory back pain. Ann Rheum Dis. 2006;65(6):804–8. van den Berg R, Lenczner G, Feydy A, van der Heijde D, Reijnierse M, Saraux A, et al. Agreement between clinical practice and trained central reading in reading of sacroiliac joints on plain pelvic radiographs: results from the DESIR cohort. Arthritis rheumatology. 2014;66(9):2403–11. Weiss PF, Xiao R, Brandon TG, Biko DM, Maksymowych WP, Lambert RG, et al. Radiographs in screening for sacroiliitis in children: what is the value? Arthritis research therapy. 2018;20(1):141. Esteva A, Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017;542(7639):115–8. Bressem KK, Adams LC, Erxleben C, Hamm B, Niehues SM, Vahldiek JL. Comparing different deep learning architectures for classification of chest radiographs. Scientific reports. 2020;10(1):1–16. Urakawa T, Tanaka Y, Goto S, Matsuzawa H, Watanabe K, Endo N. Detecting intertrochanteric hip fractures with orthopedist-level accuracy using a deep convolutional neural network. Skeletal Radiol. 2019;48(2):239–44. Chea P, Mandell JC. Current applications and future directions of deep learning in musculoskeletal radiology. Skeletal radiology. 2020:1–15. Bressem KK, Vahldiek JL, Adams LC, Niehues SM, Haibel H, Rodriguez VR, et al. Detecting radiographic sacroiliitis using deep learning with expert-level accuracy in axial spondyloarthritis. medRxiv. 2020. Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Confounding variables can degrade generalization performance of radiological deep learning models. arXiv preprint arXiv:180700431. 2018. Ferrari E, Retico A, Bacciu D. Measuring the effects of confounders in medical supervised classification problems: the Confounding Index (CI). Artif Intell Med. 2020;103:101804. Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D, editors. Grad-cam: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE international conference on computer vision; 2017. Poddubnyy D, Inman RD, Sieper J, Ganz F, Hojnik M. Region-Specific Differences in Clinical Presentation of Patients with Axial Spondyloarthritis – Results from a Large Multinational Cohort Study [abstract]. Arthritis & Rheumatology. 2018;70. Rudwaleit M, Haibel H, Baraliakos X, Listing J, Märker-Hermann E, Zeidler H, et al. The early disease stage in axial spondylarthritis: results from the German Spondyloarthritis Inception Cohort. Arthritis Rheumatism: Official Journal of the American College of Rheumatology. 2009;60(3):717–27. Lin T-Y, Goyal P, Girshick R, He K, Dollár P, editors. Focal loss for dense object detection. Proceedings of the IEEE international conference on computer vision; 2017. Everingham M, Van Gool L, Williams CK, Winn J, Zisserman A. The pascal visual object classes (voc) challenge. Int J Comput Vision. 2010;88(2):303–38. Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D et al, editors. Microsoft coco: Common objects in context. European conference on computer vision; 2014: Springer. Türk S, Cukur T, Hepdurgun C, Dar S, Karabulut A, Tamsel I et al, editors. Deep Learning Algorithm for Sacroiliac Joint Evaluation: Grading with Radiographs2020: European Congress of Radiology 2020. Spoorenberg A, De Vlam K, van der Linden S, Dougados M, Mielants H, van de Tempel H, et al. Radiological scoring methods in ankylosing spondylitis. Reliability and change over 1 and 2 years. The Journal of Rheumatology. 2004;31(1):125–32. Van Tubergen A, Heuft-Dorenbosch L, Schulpen G, Landewe R, Wijers R, Van Der Heijde D, et al. Radiographic assessment of sacroiliitis by radiologists and rheumatologists: does training improve quality? Annals of the rheumatic diseases. 2003;62(6):519–25. Rao A, Monteiro JM, Mourao-Miranda J, Initiative AsD. Predictive modelling using neuroimaging data in the presence of confounds. NeuroImage. 2017;150:23–49. Brown MR, Sidhu GS, Greiner R, Asgarian N, Bastani M, Silverstone PH, et al. ADHD-200 Global Competition: diagnosing ADHD using personal characteristic data can outperform resting state fMRI measurements. Front Syst Neurosci. 2012;6:69. Cabitza F, Rasoini R, Gensini GF. Unintended consequences of machine learning in medicine. Jama. 2017;318(6):517–8. Ting DSW, Cheung CY-L, Lim G, Tan GSW, Quang ND, Gan A, et al. Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes. Jama. 2017;318(22):2211–23. Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018;15(11):e1002683. Tables Table 1 summarizes values for Average Precision (AP) and mean Average Precision (mAP) in both the tuning and holdout dataset as well as results from a manual review by two expert readers in categories of fully or partially captured SIJs as well as missed predictions. Judgment of both reviewers was identical and is therefore not further differentiated in this summary. SIJ captured (Manual review) Calculated metrics Fully Partially Missed AP mAP Tuning data (n=106) 105 0 1 0.996 0.538 Holdout data (n=140) 132 1 7 0.981 0.515 Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-379664","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research article","associatedPublications":[],"authors":[{"id":19868401,"identity":"178d7e26-bd7c-408a-ac87-2238738f776e","order_by":0,"name":"Nils Friedrich Grauhan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABSElEQVRIie2RMWvCQBiG3xCoSyTrJ1b8CydCpCDmr5wE0sVOheLQISCci+LaQn+EU2unXgh0ujarIJS62KUFxzpYmtNSNEq7FpoH7jjeu4f3gwMyMv4iphkAdVgAhzSCVXYAHbFkl8n6CjcUnfi7ir9W5B4FawVawYYSrRTsUWo5ozMHfzq0cydTuRihcdNVzvNiFJdruYdpNG+jNNhWjjqGIPBTq9B7YWFfwbtSrVqlryaV294xk1KherldwyJDmMt3brGxj+QMj9ByKC8mnEkfMhRoDmVa0YNxy/1W7FensBSPnMWzRPlA825HCUgrjNZKg6jlFPNCct0rwyBpQVoRK4XUDGFfECeanRVLwqsMx0mLuqfqRaoljiI9mGt3fXO+EHWXbO+68CYaZRYnSfu8XhrsfMwW1Ezf04/vNe6vLzIyMjL+HZ/oYnvqO8vEoQAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-5682-5832","institution":"Charité Universitätsmedizin Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Nils","middleName":"Friedrich","lastName":"Grauhan","suffix":""},{"id":19868402,"identity":"bc3517fe-dd49-490d-99e3-151d0be9a015","order_by":1,"name":"Keno Kyrill Bressem","email":"","orcid":"","institution":"Charite University Hospital Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Keno","middleName":"Kyrill","lastName":"Bressem","suffix":""},{"id":19868403,"identity":"e752dfc6-25b4-4687-9cef-896852f6afe7","order_by":2,"name":"Yves Nicolas Manzoni","email":"","orcid":"","institution":"Charite University Hospital Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yves","middleName":"Nicolas","lastName":"Manzoni","suffix":""},{"id":19868404,"identity":"842951d4-bf23-442e-8480-c9507e51c5ec","order_by":3,"name":"Lisa Christine Adams","email":"","orcid":"","institution":"Charite University Hospital Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Lisa","middleName":"Christine","lastName":"Adams","suffix":""},{"id":19868405,"identity":"83f3f8ea-9f4c-4995-8ddb-0e0f97f252e3","order_by":4,"name":"Valeria Rios-Rodgriguez","email":"","orcid":"","institution":"Charité Universitätsmedizin Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Valeria","middleName":"","lastName":"Rios-Rodgriguez","suffix":""},{"id":19868406,"identity":"fe17439b-c3d8-434c-bca5-1ab9ff143d38","order_by":5,"name":"Fabian Nikolai Proft","email":"","orcid":"","institution":"Charité Universitätsmedizin Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Fabian","middleName":"Nikolai","lastName":"Proft","suffix":""},{"id":19868407,"identity":"5b1399e6-679a-4d96-8018-aae41b13eeb6","order_by":6,"name":"Hildrun Haibel","email":"","orcid":"","institution":"Charite University Hospital Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hildrun","middleName":"","lastName":"Haibel","suffix":""},{"id":19868408,"identity":"d8457f6f-438d-46b4-bdd6-bda770186757","order_by":7,"name":"Martin Rudwaleit","email":"","orcid":"","institution":"Charite University Hospital Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Martin","middleName":"","lastName":"Rudwaleit","suffix":""},{"id":19868409,"identity":"7f1804cf-51ca-425e-afc0-2648a236586c","order_by":8,"name":"Stefan Markus Niehues","email":"","orcid":"","institution":"Charite University Hospital Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Stefan","middleName":"Markus","lastName":"Niehues","suffix":""},{"id":19868410,"identity":"d7658264-26f6-4657-a4a4-9289b327a64b","order_by":9,"name":"Denis Poddubnyy","email":"","orcid":"","institution":"Charite University Hospital Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Denis","middleName":"","lastName":"Poddubnyy","suffix":""},{"id":19868411,"identity":"20e3378f-b683-46d4-a274-471a0a5b9317","order_by":10,"name":"Janis Lucas Vahldiek","email":"","orcid":"","institution":"Charite University Hospital Berlin: Charite Universitatsmedizin Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Janis","middleName":"Lucas","lastName":"Vahldiek","suffix":""}],"badges":[],"createdAt":"2021-03-31 13:01:58","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-379664/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-379664/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":7721808,"identity":"537790af-09bb-4f27-9885-61909c32a2b1","added_by":"auto","created_at":"2021-04-06 17:54:36","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":40996,"visible":true,"origin":"","legend":"demonstrates the allocation of all images used in this study. Training and tuning data were retrieved from the PROOF dataset and split randomly as shown above. The holdout data was exclusively taken from the GESPIC dataset.","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-379664/v1/f1827ea6ac8017758848bd7a.png"},{"id":7722125,"identity":"5290549d-0b8f-4d86-900a-8e2b86058981","added_by":"auto","created_at":"2021-04-06 18:00:36","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":25345,"visible":true,"origin":"","legend":"demonstrates the mathematical concept of the term Intersection over Union (IoU). The area of intersection (grey) between the predicted area (red) and the ground truth (green) is divided by the area enclosed (union) by both the predicted area (red) and the ground truth (green).","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-379664/v1/16639faaf997e8efb2d47494.png"},{"id":7722065,"identity":"44e74256-f08b-4bbe-a2e0-6e84c7a1d90a","added_by":"auto","created_at":"2021-04-06 17:57:36","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":556470,"visible":true,"origin":"","legend":"shows two examples from the holdout data. The top row images display ground truth (left) and prediction (right) of a case in which the left SIJ was missed. The bottom row images demonstrate an example where the left SIJ was only partially captured.","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-379664/v1/0045a55b4ac2fc324ff02e44.png"},{"id":13684182,"identity":"cf7fc23b-3f3a-4ecb-adf0-b5a738c93151","added_by":"auto","created_at":"2021-09-17 12:06:43","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":774211,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-379664/v1/91189cd5-9710-41b2-baff-5052e21f763d.pdf"}],"financialInterests":"","formattedTitle":"\u003cp\u003eTowards Accurate Detection of Axial Spondyloarthritis by Using Deep Learning to Capture Sacroiliac Joints on Plain Radiographs.\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003ePlain radiographs of the sacroiliac joints together with magnetic resonance imaging (MRI) play an important role in detecting axial spondyloarthritis (SpA) (1, 2). However, accuracy particularly in grading pathological findings in radiographs according to the modified New York criteria can differ significantly between observers (3, 4). Deep learning-based methods using convolutional neural networks (CNN) have recently demonstrated precision levels of human experts in reading medical images (5-7). They may therefore aid diagnosis of SpA as they have proven to excel when confronted with a specific task (8).\u003c/p\u003e\n\u003cp\u003eSuch an approach was recently pursued by training a CNN to detect definite radiographic sacroiliitis on plain radiographs (9). The results demonstrated a high accuracy of the developed CNN. While these are excellent results and comparable to the accuracy of domain experts, performance may still have been compromised: Firstly, computational capacity put a limit on the resolution of the radiographs included in the study, resulting in a significant downsizing. Although downsizing is a common practice in deep learning this entails the risk of losing crucial image information (10). Secondly, studies have shown that a CNN may not base its prediction solely on clinically relevant features (11). An excellent way to visually track a model\u0026rsquo;s decision is to use Gradient-weighted Class Activation Mappings (Grad-CAMs) (12). Grad-CAM highlights regions that have a strong influence on the model\u0026rsquo;s decision. In reviewing images that were incorrectly predicted in the holdout dataset of our previous study, several images were found where the model had based its decision on aspects unrelated to SpA features.\u003c/p\u003e\n\u003cp\u003eThe aim of this study was therefore to train and validate a CNN in the detection and segmentation of both sacroiliac joints on plain radiographs by drawing bounding boxes around the region of interest. If proven to be accurate, this CNN could be integrated in a future preprocessing pipeline augmenting the actual detection or classification of SpA no longer restrained by low image resolution or irrelevant information. To test the generalizability of our results, image data for training and testing were taken from two completely independent datasets.\u003c/p\u003e"},{"header":"Methods","content":"\u003ch2\u003eData\u003c/h2\u003e\n\u003cp\u003eIn this study, we used plain radiographs retrieved from two completely separate image collections: a) Patients With Axial Spondyloarthritis: Multic\u003cu\u003eo\u003c/u\u003euntry Registry of Clinical Characteristics (PROOF) is a continuous study involving medical facilities in 29 countries. Overall 2170 adult patients diagnosed with SpA (non-radiographic or radiographic) \u0026le;12 months before study enrolment and fulfilling the ASAS classification criteria for SpA are included (13). B) German Spondyloarthritis Inception Cohort (GESPIC) is a multicenter inception cohort study conducted in Germany that includes 646 patients with SpA (14).\u003c/p\u003e\n\u003ch2\u003ePre-processing and Segmentation\u003c/h2\u003e\n\u003cp\u003eBoth PROOF and GESPIC datasets contain radiographs of sacroiliac joints in DICOM format (Digital Imaging and Communications in Medicine). All grey-scale values were adjusted to a unified representation state using the Horos Project DICOM Viewer (version 4.0.0, www.horosproject.org). Afterward, all images were converted to the Tagged Image File Format (TIFF). Annotations for training, tuning and holdout datasets were done using the COCO Annotator (https://github.com/jsbroks/coco-annotator). Bounding boxes were placed by one expert radiologist (JLV) separately around each of the sacroiliac joints (SIJ) on every radiograph. A total of 529 annotated images from the PROOF dataset was available and then randomly split into a training dataset (80%, n=423) and a tuning dataset (20%, n=106). A total of 140 annotated images from GESPIC was available and taken to form the holdout dataset (see flowchart in Figure 1).\u003c/p\u003e\n\u003ch2\u003eModel Training\u003c/h2\u003e\n\u003cp\u003eWe used a RetinaNet with a ResNet-50 backbone, pre-trained in ImageNet, specifically suited for object detection (15). Model training was performed using IceVision (https://airctic.com, version 0.5.2), which is an object-detection framework built on top of Python (https://www.python.org, version 3.8), as well as the Fastai (https://fast.ai, version 2.2.5) and PyTorch (https://pytorch.org, version 1.7.0) libraries. Training was executed on a dedicated Ubuntu 18.04 Workstation with two Nvidia GeForce RTX 2080ti Graphic cards. All input images were resized and cropped to 512 x 512 px. Random transformations were used to augment training data. During model training, an early stopping approach was utilized to avoid overfitting. As we used a pre-trained ResNet50 as a backbone, we disabled weight updates for the backbone during the first 53 training epochs. Then we trained for 27 more epochs, updating all weights of the model. The model with the best performance on the tuning data was exported and used for inference on the holdout data.\u003c/p\u003e\n\u003ch2\u003eStatistical Analysis\u003c/h2\u003e\n\u003cp\u003eModel performance was evaluated using two separate approaches: first, in a swift manual review, two expert readers (NFG, YNM) independently checked the results of the model on the holdout data and counted all images in which the model either partially or completely failed to detect the SIJs (see Table 1). Second, average precision (AP) and mean average precision (mAP) were calculated based on Intersection over Union (IoU). This metric is commonly used in object detection competitions such as the PASCAL Visual Object Classes 2012 (VOC 2012) challenge (16) or in the Microsoft Common objects in context (COCO) challenge (17). IoU is an expression for the overlap of the predicted area and the actual image (ground truth) divided by the area enclosed by both (see Figure 2). A standard requirement to identify a prediction as a true positive is a ratio of 0.5 (17).\u003c/p\u003e"},{"header":"Results","content":" \u003cp\u003eIn a manual review of the inferred segmentations in the tuning data, both observers found no instances and only one instance in the holdout data in which the model had only partially captured the SIJ. In the tuning dataset, one instance was found in which the model had missed the SIJ, and seven instances in the holdout dataset. All remaining 105 (tuning data, 99.1%) and 132 (holdout data, 94.3%) instances were counted as fully captured by both observers. There was no disagreement between the two observers in assigning each image to the above categories.\u003c/p\u003e \u003cp\u003eObject detection metrics calculated from IoU were at 0.538 for mean average precision and 0.996 for average precision in the tuning dataset. In the holdout dataset, values were 0.515 for mean average precision and 0.981 for average precision. The results are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e "},{"header":"Discussion","content":"\u003cp\u003eOur model showed excellent results both in the manual review and displayed by high AP and mAP values. In addition, we would like to highlight that the manual review revealed only one instance of a partial capture of the SIJ. This underlines the fact that the model was very accurate and that a further increase in AP and mAP values would not necessarily reflect a more precise capture of the actual SIJ. Also, it should be noted that in the majority of cases where the model failed on holdout data radiographs included X-ray protective devices commonly absent in all the training and tuning data (see examples in Figure 3). All of this suggests that such a model could very well serve as a reliable pre-stage placed ahead of a second CNN charged with detecting specific features relevant to SpA.\u003c/p\u003e\n\u003cp\u003eOnly a few studies have tried to use deep learning approaches to either detect SpA in a simple binary task or to grade them according to the modified New York Criteria. The study by T\u0026uuml;rk et al. is the only one known to us that relied on conventional radiographs to grade SpA (18). In their work authors report that while their model performed well compared to human observers accuracy still dropped when asked to differentiate between early stages of SpA (grade 0 and grade 1). Notably, no segmentation of the included images was performed. Therefore, we are confident that isolating SIJs on radiographs before predicting grades of SpA will boost model performance, especially in detecting these early stages.\u003c/p\u003e\n\u003cp\u003eWe base this judgment on the possibility to reduce the need for downsizing images that contain rather subtle image information relevant to a correct prediction. By automatically cropping bounding boxes containing the SIJ and resizing the snippet to its original high resolution, we could achieve an efficient pipeline layout. Such an efficient pipeline is very important in grading SpA as the anatomy of SIJs is complex and features are often difficult to detect even for expert readers (19, 20). Also, lowering image resolution to facilitate model training can lead to unnecessary bias (10).\u003c/p\u003e\n\u003cp\u003eDeep learning models are known to be susceptible to various confounders in medical image data, such as the gender or age of the participants. (21, 22). Such confounders pose a serious problem as they can lead to overestimating the generalizability of the results (10). Furthermore, the frequently described \u0026ldquo;black box nature\u0026rdquo; of deep learning-based algorithms carries the risk that further systematic errors may be overlooked (23, 24). We are convinced that reducing image information to the clinically relevant core will help to decrease these errors.\u003c/p\u003e\n\u003cp\u003eAn important prerequisite to verify the generalizability of the results is to check the performance of a CNN on holdout data from external datasets (25). As with our previous study, we have chosen two completely independent datasets for training and testing. Detection metrics used to evaluate the tuning and the holdout data (AP and mAP) were very close and therefore imply a robust performance of our model.\u003c/p\u003e\n\u003cp\u003eBeyond the application of detecting SpA, our study promises to be valuable in the broader field of digital applications in Rheumatology. Radiographic diagnostics in Rheumatology often rely on very subtle changes in radiographs such as small bony erosions, soft tissue swelling, or joint space narrowing. A similar approach could, therefore, potentially enhance any automated detection of more delicate findings on plain radiographs.\u003c/p\u003e"},{"header":"Conclusions","content":" \u003cp\u003eConvolutional neural networks are highly accurate in detecting sacroiliac joints on plain radiographs and may therefore aid in a more focused image analysis of these regions.\u003c/p\u003e "},{"header":"Abbreviations","content":"\u003cp\u003eAP: Average precision; ASAS: Assessment of Spondylarthritis international Society; CNN: Convolutional neural network; COCO: Microsoft Common objects in context; DICOM: Digital Imaging and Communications in Medicine; GESPIC: German Spondyloarthritis Inception Cohort; IoU: Intersection over union; mAP: mean Average precision; MRI: Magnetic Resonance Imaging; PASCAL VOC 2012: Pascal Visual Object Classes Challenge 2012; PROOF: \u003cu\u003eP\u003c/u\u003eatients With Axial Spondyloarthritis (study); SIJ: Sacroiliac joint; SpA: axial Spondyloarthritis\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthical Approval and Consent to participate\u003c/h2\u003e\n\u003cp\u003eBoth PROOF and GESPIC cohorts were approved by the local ethics committees of each study center in accordance with the local laws and regulations and is being conducted in accordance with the Declaration of Helsinki and Good Clinical Practice. GESPIC was additionally approved by a central ethics committee of the coordinating center. Written informed consent to participate was obtained from all patients.\u003c/p\u003e\n\u003ch2\u003eConsent for publication\u003c/h2\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003ch2\u003eAvailability of supporting data\u003c/h2\u003e\n\u003cp\u003eThe data that support the findings of this study are available from the corresponding author, N.F.G., upon reasonable request.\u003c/p\u003e\n\u003ch2\u003eCompeting Interests:\u003c/h2\u003e\n\u003cp\u003eN.F.G, K.K.B., L.C.A., and Y.N.M. have nothing to disclose.\u003c/p\u003e\n\u003cp\u003eJ.L.V reports non-financial support from Bayer, non-financial support from Guerbet, non-financial support from Medtronic, personal fees and non-financial support from Merit Medical, outside the submitted work.\u003c/p\u003e\n\u003cp\u003eS.M.N. reports grants from the German Research Foundation (Deutsche Forschungsgemeinschaft), personal fees from Bracco Imaging, personal fees from Canon Medical Systems, personal fees from Guerbet, personal fees from Teleflex Medical GmbH, personal fees from Bayer Vital GmbH, outside the submitted work.\u003c/p\u003e\n\u003cp\u003eH.H. reports personal fees from Pfizer, personal fees from Janssen, personal fees from Novartis, personal fees from Roche, personal fees from MSD, outside the submitted work.\u003c/p\u003e\n\u003cp\u003eF.P. reports personal fees from AbbVie, personal fees from AMGEN, personal fees from BMS, personal fees from Celgene, from MSD, grants and personal fees from Novartis, personal fees from Pfizer, from Roche, personal fees from UCB, outside the submitted work.\u003c/p\u003e\n\u003cp\u003eM.R. received honoraria and/or consulting fees from AbbVie, BMS, Celgene, Janssen, Eli Lilly, MSD, Novartis, Pfizer, Roche, UCB Pharma.\u003c/p\u003e\n\u003cp\u003eD.P. reports grants and personal fees from AbbVie, during the conduct of the study; grants and personal fees from AbbVie, personal fees from BMS, personal fees from Celgene, grants and personal fees from Lilly, grants and personal fees from MSD, grants and personal fees from Novartis, grants and personal fees from Pfizer, personal fees from Roche, personal fees from UCB, outside the submitted work.\u003c/p\u003e\n\u003cp\u003eV.R.R. reports personal fees from Abbvie, personal fees from Novartis, outside the submitted work.\u003c/p\u003e\n\u003ch2\u003eFunding:\u003c/h2\u003e\n\u003cp\u003eGESPIC was initially supported by the German Federal Ministry of Education and Research (Bundesministerium f\u0026uuml;r Bildung und Forschung, BMBF). As scheduled, BMBF was reduced in 2005 and stopped in 2007; thereafter, complementary financial support was obtained also from Abbott, Amgen, Centocor, Schering\u0026ndash;Plough, and Wyeth. Starting in 2010, the core GESPIC cohort was supported by AbbVie. The PROOF study is funded by AbbVie. We thank AbbVie for allowing us to use the PROOF dataset for the aim of the current study.\u003c/p\u003e\n\u003ch2\u003eAuthor\u0026rsquo;s contributions:\u003c/h2\u003e\n\u003cp\u003eNFG and JLV: Data preparation, execution of model training and testing, statistical analysis, drafting and finalizing the manuscript.\u003c/p\u003e\n\u003cp\u003eKKB, YNM, DP, SMN and LCA: Essential contributions to the concept of the study, editing and correcting the manuscript.\u003c/p\u003e\n\u003cp\u003eVRR, FP, HH, and MR: Editing and correcting the manuscript.\u003c/p\u003e\n\u003ch2\u003eAuthors\u0026rsquo; information\u003c/h2\u003e\n\u003cp\u003e1 Charit\u0026eacute; \u0026ndash; Universit\u0026auml;tsmedizin Berlin, Department of Radiology, Berlin, Germany\u003c/p\u003e\n\u003cp\u003e2 Berlin Institute of Health (BIH), Berlin, Germany\u003c/p\u003e\n\u003cp\u003e3 Charit\u0026eacute; \u0026ndash; Universit\u0026auml;tsmedizin Berlin, Department of Gastroenterology, Infectiology and Rheumatology, Berlin, Germany\u003c/p\u003e\n\u003cp\u003e4 Department of Internal Medicine and Rheumatology, Klinikum Bielefeld Rosenhöhe, Bielefeld, Germany\u003c/p\u003e\n\u003ch2\u003eAcknowledgements:\u003c/h2\u003e\n\u003cp\u003eLCA is participant in the BIH-Charit\u0026eacute; Clinician Scientist Program funded by the Charit\u0026eacute; \u0026ndash;Universit\u0026auml;tsmedizin Berlin and the Berlin Institute of Health. KKB is participant in the BIH-Charit\u0026eacute; Digital Clinician Scientist Program funded by the Charit\u0026eacute; \u0026ndash;Universit\u0026auml;tsmedizin Berlin and the Berlin Institute of Health.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSlobodin G, Hussein H, Rosner I, Eshed I. Sacroiliitis\u0026ndash;early diagnosis is key. Journal of inflammation research. 2018;11:339.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeuft-Dorenbosch L, Landewe R, Weijers R, Wanders A, Houben H, Van Der Linden S, et al. Combining information obtained from magnetic resonance imaging and conventional radiographs to detect sacroiliitis in patients with recent onset inflammatory back pain. Ann Rheum Dis. 2006;65(6):804\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan den Berg R, Lenczner G, Feydy A, van der Heijde D, Reijnierse M, Saraux A, et al. Agreement between clinical practice and trained central reading in reading of sacroiliac joints on plain pelvic radiographs: results from the DESIR cohort. Arthritis rheumatology. 2014;66(9):2403\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeiss PF, Xiao R, Brandon TG, Biko DM, Maksymowych WP, Lambert RG, et al. Radiographs in screening for sacroiliitis in children: what is the value? Arthritis research therapy. 2018;20(1):141.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEsteva A, Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017;542(7639):115\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBressem KK, Adams LC, Erxleben C, Hamm B, Niehues SM, Vahldiek JL. Comparing different deep learning architectures for classification of chest radiographs. Scientific reports. 2020;10(1):1\u0026ndash;16.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUrakawa T, Tanaka Y, Goto S, Matsuzawa H, Watanabe K, Endo N. Detecting intertrochanteric hip fractures with orthopedist-level accuracy using a deep convolutional neural network. Skeletal Radiol. 2019;48(2):239\u0026ndash;44.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChea P, Mandell JC. Current applications and future directions of deep learning in musculoskeletal radiology. Skeletal radiology. 2020:1\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBressem KK, Vahldiek JL, Adams LC, Niehues SM, Haibel H, Rodriguez VR, et al. Detecting radiographic sacroiliitis using deep learning with expert-level accuracy in axial spondyloarthritis. medRxiv. 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Confounding variables can degrade generalization performance of radiological deep learning models. arXiv preprint arXiv:180700431. 2018.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFerrari E, Retico A, Bacciu D. Measuring the effects of confounders in medical supervised classification problems: the Confounding Index (CI). Artif Intell Med. 2020;103:101804.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSelvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D, editors. Grad-cam: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE international conference on computer vision; 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePoddubnyy D, Inman RD, Sieper J, Ganz F, Hojnik M. Region-Specific Differences in Clinical Presentation of Patients with Axial Spondyloarthritis \u0026ndash; Results from a Large Multinational Cohort Study [abstract]. Arthritis \u0026amp; Rheumatology. 2018;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRudwaleit M, Haibel H, Baraliakos X, Listing J, M\u0026auml;rker-Hermann E, Zeidler H, et al. The early disease stage in axial spondylarthritis: results from the German Spondyloarthritis Inception Cohort. Arthritis Rheumatism: Official Journal of the American College of Rheumatology. 2009;60(3):717\u0026ndash;27.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin T-Y, Goyal P, Girshick R, He K, Doll\u0026aacute;r P, editors. Focal loss for dense object detection. Proceedings of the IEEE international conference on computer vision; 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEveringham M, Van Gool L, Williams CK, Winn J, Zisserman A. The pascal visual object classes (voc) challenge. Int J Comput Vision. 2010;88(2):303\u0026ndash;38.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D et al, editors. Microsoft coco: Common objects in context. European conference on computer vision; 2014: Springer.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eT\u0026uuml;rk S, Cukur T, Hepdurgun C, Dar S, Karabulut A, Tamsel I et al, editors. Deep Learning Algorithm for Sacroiliac Joint Evaluation: Grading with Radiographs2020: European Congress of Radiology 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSpoorenberg A, De Vlam K, van der Linden S, Dougados M, Mielants H, van de Tempel H, et al. Radiological scoring methods in ankylosing spondylitis. Reliability and change over 1 and 2 years. The Journal of Rheumatology. 2004;31(1):125\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan Tubergen A, Heuft-Dorenbosch L, Schulpen G, Landewe R, Wijers R, Van Der Heijde D, et al. Radiographic assessment of sacroiliitis by radiologists and rheumatologists: does training improve quality? Annals of the rheumatic diseases. 2003;62(6):519\u0026ndash;25.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRao A, Monteiro JM, Mourao-Miranda J, Initiative AsD. Predictive modelling using neuroimaging data in the presence of confounds. NeuroImage. 2017;150:23\u0026ndash;49.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrown MR, Sidhu GS, Greiner R, Asgarian N, Bastani M, Silverstone PH, et al. ADHD-200 Global Competition: diagnosing ADHD using personal characteristic data can outperform resting state fMRI measurements. Front Syst Neurosci. 2012;6:69.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCabitza F, Rasoini R, Gensini GF. Unintended consequences of machine learning in medicine. Jama. 2017;318(6):517\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTing DSW, Cheung CY-L, Lim G, Tan GSW, Quang ND, Gan A, et al. Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes. Jama. 2017;318(22):2211\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018;15(11):e1002683.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003e\u003cstrong\u003eTable 1\u003c/strong\u003e summarizes values for Average Precision (AP) and mean Average Precision (mAP) in both the tuning and holdout dataset as well as results from a manual review by two expert readers in categories of fully or partially captured SIJs as well as missed predictions. Judgment of both reviewers was identical and is therefore not further differentiated in this summary.\u003c/p\u003e\n\u003ctable border=\"1\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" width=\"302\"\u003e\n\u003cp\u003eSIJ captured (Manual review)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"201\"\u003e\n\u003cp\u003eCalculated metrics\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003eFully\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003ePartially\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003eMissed\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003eAP\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003emAP\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003eTuning data (n=106)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e105\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e0.996\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e0.538\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003eHoldout data (n=140)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e132\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e0.981\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"101\"\u003e\n\u003cp\u003e0.515\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Axial spondyloarthritis, sacroiliac joint, radiograph, deep learning, convolutional neural network, object detection","lastPublishedDoi":"10.21203/rs.3.rs-379664/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-379664/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eWell-informed decisions about how to best treat patients with axial spondyloarthritis (SpA) regularly include an evaluation of the sacroiliac joints (SIJ) on plain radiographs. However, grading radiographic findings correctly has proven to be a considerable challenge to expert readers as well as to state-of-the-art convolutional neural networks (CNNs). A method to reduce image information to the clinically relevant core would undoubtedly lead to more accurate results. We, therefore, trained a CNN only to detect SIJs on radiographs and evaluated its potential as a preprocessing pipeline in the automated classification of SpA.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eMaterials and Methods\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eWe employed a CNN of the RetinaNet architecture, which was trained on a total of 423 plain radiographs of the sacroiliac joints (SIJs). Images were taken from two completely independent datasets. Training and tuning were performed on image data from the Patients With Axial Spondyloarthritis (PROOF) study and testing was executed using images from the German Spondyloarthritis Inception Cohort (GESPIC). Performance was evaluated by manual review and standard object detection metrics from PASCAL and Microsoft COCO.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eThe CNN produced excellent results in detecting SIJs on the tuning (n =106) and on the holdout dataset (n =140). Object detection metrics for the tuning data were AP = 0.996 and mAP = 0.538; values for the independent holdout data were AP = 0.981 and mAP = 0.515. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eThe developed CNN was highly accurate in detecting SIJs on radiographs. Such a model could increase the reliability of deep learning-based algorithms in detecting and grading SpA.\u003c/p\u003e","manuscriptTitle":"Towards Accurate Detection of Axial Spondyloarthritis by Using Deep Learning to Capture Sacroiliac Joints on Plain Radiographs.","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-04-06 17:54:34","doi":"10.21203/rs.3.rs-379664/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"de6b7296-cc11-4de8-82c5-8d4e46198576","owner":[],"postedDate":"April 6th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":3453922,"name":"Health Economics \u0026 Outcomes Research"}],"tags":[],"updatedAt":"2021-04-06T17:57:35+00:00","versionOfRecord":[],"versionCreatedAt":"2021-04-06 17:54:34","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-379664","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-379664","identity":"rs-379664","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.