Redefining Early Osteoarthritis Detection: Deep Learning Meets Historical Radiography | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Redefining Early Osteoarthritis Detection: Deep Learning Meets Historical Radiography Rifatul Islam Majumder This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5983981/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The Kellgren-Lawrence (KL) grading system, established by Lawrence and Aitken Swan [1] and Kellgren and Lawrence [2] in the early 1950s, remains the clinical standard for radiographic osteoarthritis (OA) classification despite well‐documented inter‐ and intra‐observer variability [3-6]. Early studies revealed that occupational factors, particularly in coal miners, are strongly associated with knee OA [1,2]. However, subsequent research has highlighted that, over reliance on osteophytes, ambiguous criterion especially for KL Grade 1 (doubtful OA) lead to diagnostic uncertainty [3-6]. Comparative evaluations of multiple radiographic scoring systems based on joint space narrowing (JSN) and osteophyte formation [7-10] have further underscored limited sensitivity of the KL system to early cartilage damage and inconsistent reproducibility [11-16]. Health sciences/Biomarkers/Diagnostic markers Health sciences/Health care/Medical imaging/Radiography Health sciences/Diseases/Rheumatic diseases/Osteoarthritis Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Figure 14 Figure 15 Figure 16 Full Text To address these issues, e‐learning module designed to standardize KL grading among novice clinicians, resulting in significantly improved interrater and intrarater reliability (weighted κ = 0.82 and 0.85, respectively) [17]. Complementing this educational approach, I employed deep learning techniques to evaluate the clinical relevance of KL Grade 1 and to identify potential biomarkers for early OA detection. Using 45 degree posteroanterior weight‐bearing knee radiographs from the OIA dataset [18,19], I trained multiple models including ResNet18 and DenseNet121 architectures, with and without segmentation [20-22] described in Section 2.3. These experiments consistently demonstrated that KL Grade 1 is a major source of misclassification, as models frequently conflate it with Grade 0, thereby blurring the diagnostic boundary for early OA detection [23]. Integration of Gradient weighted Class Activation Mapping (Grad CAM) [24] provided valuable visual insights into model decision making. The Grad CAM analyses (see Section 2.2, Figures 1-13) revealed that early OA is primarily associated with features in the lateral compartment and the lateral femoral condyle, with joint space narrowing being as significant as, if not more so than, osteophyte formation. In contrast, subtle abnormalities in the medial compartment were not consistently linked to early OA. Furthermore, the activation patterns for KL Grade 1 frequently overlapped with those for both normal (Grade 0) and advanced OA (Grade 2+), corroborating historical concerns regarding observer bias [25-27]. In addition, I developed a unique stacked input approach that combines the grayscale image, the segmentation mask, and the grayscale image. This method, integrated with deep learning model, outperforms current techniques for automating early OA detection [28-31]. To my knowledge, this segmentation training strategy is novel and has not been previously reported. Furthermore, no prior study has combined the KL grading system with expert human readings and model heatmap analyses to explore definitive early biomarkers of OA. These findings suggest that refining the KL grading criteria, potentially by redefining or eliminating the ambiguous Grade 1 could enhance early OA diagnosis. Moreover, relying solely on osteophyte formation is insufficient for early OA detection; a redefined KL grading scheme may not only improve OA classification but also help distinguish OA from other bone diseases, such as rheumatoid arthritis. The combination of standardized training and advanced deep learning with interpretability techniques holds promise for improving diagnostic accuracy and clinical decision making in OA management, with significant implications for early diagnosis and tailored treatment planning. 2.1) Descriptive Overview: This research builds upon the foundational work of Kellgren and Lawrence, particularly their studies from 1952 and 1957. In 1952, Lawrence and Aitken Swan highlighted the occupational influences on rheumatic complaints by revealing that coal miners exhibited significantly higher incidences of knee osteoarthritis (OA) compared to the general population [1]. That same year, Kellgren and Lawrence demonstrated a strong association between occupational strain and radiological signs of knee OA, establishing a statistically significant relationship (P < 0.01) between knee pain and OA severity [2]. Their grading system, which categorized osteoarthritic changes into five levels, later evolved into the widely used Kellgren Lawrence (KL) grading system formalized by the World Health Organization in 1961. A recent study by Braun and Gold (2011) emphasizes that OA is a disease of the entire joint and explores the role of imaging particularly for knee OA in its diagnosis. The KL grading system, which relies on radiographic features such as osteophytes and joint space narrowing (JSN), remains the standard for OA classification despite limitations, including poor sensitivity to early cartilage damage [3]. Felson et al. (1987), as part of the Framingham Osteoarthritis Study, assessed the prevalence of knee OA in a population based cohort using a 0-4 scale based on the Kellgren and Lawrence criteria [4],[5]. Additionally, Scott et al. developed and validated an atlas for evaluating individual radiographic features of knee OA within the Baltimore Longitudinal Study of Aging. In their study, four trained readers evaluated 30 standing anterior posterior knee radiographs for eight features including medial and lateral osteophytes, JSN, sclerosis, osteophytes of the tibial spines, and chondrocalcinosis along with the overall KL global scale, reporting inter reader reliability from 0.63 to 0.83 and intra reader reliability from 0.82 to 0.95 [6]. Despite KL grades being longstanding, studies since 1957 have revealed substantial inter and intra observer variability in the KL grading system. For example, the same observer demonstrated a high correlation (r = 0.88) when reading metacarpophalangeal joints twice, yet the lowest correlation was found for the dorso lumbar spine (r = 0.42). Knee OA grading showed relatively strong agreement (r = 0.83) both within and between observers, whereas wrist OA exhibited the lowest agreement (r = 0.10) [4]. This variability particularly in assessing features such as osteophytes, sclerosis, and JSN remains a major concern. The American Rheumatism Association (ARA) subcommittee, in its 1981 report and subsequent 1987 revision, described OA as a heterogeneous disease, categorizing it as either idiopathic (no known cause) or secondary (linked to conditions such as rheumatoid arthritis [RA]) [7]. Moreover, Altman et al. demonstrated that RA patients also exhibit joint space narrowing, with osteophytes present in 62% of cases , indicating that reliance solely on radiological features may be problematic [8]. Multiple standard systems including the International Knee Documentation Committee (IKDC) radiographic scale [9],[10], Fairbank [11],[12], Brandt et al. [13], Ahlbäck [14], and Jäger Wirth [15],[16] primarily depend on JSN and osteophytes. This dependence raises concerns regarding the validity of using JSN as a definitive feature due to its subjectivity and low reliability. In a prospective study of 632 patients undergoing revision ACL reconstruction (the Multicenter ACL Revision Study [MARS]), six radiographic classification systems were evaluated by comparing anteroposterior (AP) and 45° posteroanterior (Rosenberg) weight bearing radiographs with arthroscopic findings. Overall interobserver reliability was moderate for AP images (ICC = 0.55; 95% CI, 0.53-0.56) and good for Rosenberg images (ICC = 0.63; 95% CI, 0.61-0.65). Specifically, the KL system had an ICC of 0.38 (95% CI, 0.33-0.43) on AP views and 0.54 (95% CI, 0.48-0.59) on Rosenberg views, while the correlation between KL grading and arthroscopic cartilage degeneration was higher with Rosenberg radiographs (Spearman rho = 0.42; 95% CI, 0.33-0.49) than with AP radiographs (rho = 0.30; 95% CI, 0.23-0.38). Although the IKDC scale demonstrated the best overall performance (AP: ICC = 0.59, 95% CI, 0.55-0.63; Rosenberg: ICC = 0.66, 95% CI, 0.62-0.71), the comparatively lower reliability of the KL system underscores its limitations in accurately classifying early osteoarthritic changes based solely on radiographic features [17]. The literature further indicates that the KL grading system faces significant limitations in differentiating between features such as “possible” versus “definite” osteophytes and “doubtful” versus “definite” joint space narrowing. These variations, which can differ among observers, highlight a lack of standardization and underscore the need for comprehensive training. To address this issue, an e learning tool (Articulate Presenter) was developed. The tutorial refined over three iterative feedback rounds from clinical researchers, radiologists, and graduate students focused on key features such as osteophytes and JSN to distinguish KL Grades 0-4. Forty seven health sciences graduate students with no prior KL grading experience completed the training and subsequently graded 30 knee radiographs (including 15 duplicates for reliability testing). The results demonstrated almost perfect interrater reliability (weighted kappa, κw = 0.82) and intrarater reliability (κw = 0.85), significantly improving upon standard reliability levels reported in the literature . Students reported increased confidence and understanding of the KL scale although they noted difficulties distinguishing between Grades 3 and 4 suggesting that standardized online training can enhance the consistency of KL grading in both research and clinical practice [18] (see FIGURE A for a visualization of the differences between possible and definite osteophytes and Figure B we can visualize JSN and Figure C sclerosis). Although osteophytes are generally a reliable indicator, the lack of a clear definition for what constitutes an osteophyte hinders early OA detection, priority in current research. Bagge et al. conducted a comparative study on the prevalence of radiographic osteoarthritis (ROA) in two elderly European populations (Göteborg, Sweden, and Zoetermeer, The Netherlands). An interobserver reliability assessment on 150 radiographic films revealed that a binary “abnormal” versus “normal” classification yielded higher agreement percentages and kappa values than a five point scale, while an intraobserver analysis of 50 films demonstrated excellent reliability (kappa values well over 0.75). The study found that hand ROA was more prevalent in the Göteborg population, whereas knee ROA prevalence was similar between the two groups. Notably, neither population showed a sigwnificant age related increase in hand ROA; however, while knee ROA did not increase with age in the Swedish cohort, it exhibited a significant age related increase in both sexes in the Zoetermeer population [19]. From these observations, hypothesize that many of the validity issues associated with the KL grading system stem from its progressive nature, same concern also raised by Spector and Cooper reappraised the radiographic assessment of OA in population studies, focusing on the widely used KL grading system that assigns grades 0-4 based on specific radiographic features. They emphasize that OA, the most common joint disorder globally, is a complex interplay of clinical, radiographic, and pathological changes, and that an ideal radiographic system should be accurate, reproducible, noninvasive, convenient, and inexpensive. Spector and Cooper highlight two major criticisms: first, inconsistencies in definitions for example, knee Grade 2 has been variably defined as “definite osteophytes with minimal joint space narrowing” versus “definite osteophytes but unimpaired joint space” which contribute to poor interobserver and intercenter reproducibility; and second, an overemphasis on osteophytes that assumes joint space narrowing follows osteophyte formation, potentially excluding cases with significant joint space loss but no visible osteophytes . Their review, which included a preliminary analysis of 1,954 knee radiographs, demonstrated that while all measures (joint space narrowing, osteophyte formation, and sclerosis) showed good reproducibility, the presence of definite osteophytes best predicted knee pain, whereas for hip OA, joint space measurements were more reproducible and strongly associated with pain though incorporating osteophyte data further improved prediction. This reappraisal underscores that no single radiographic measure is universally applicable across different joints, and that standardizing definitions and evaluating individual joint features may enhance the reproducibility and clinical utility of the KL grading system [20]. Using the OIA dataset [21],[22] with 45° posteroanterior (Rosenberg) weight bearing radiographs, multiple deep learning models including ResNets [23], DenseNets [24], and EfficientNets [25], both with and without segmentation, and with fixed and adaptive learning rates across various training transformations were initially tested. From these, two deep learning models were selected for their effectiveness: five ResNet18 models without segmentation and five DenseNet121 models with segmentation. These models were trained on the Knee Osteoarthritis Severity Grading Dataset [5] to classify knee OA. The findings indicate that KL grade 1 is a major source of misclassification, causing overlap between KL grade 0 and KL grade 2+. In all instances, KL grade 1 was more closely related to KL grade 0, further blurring the diagnostic boundary for observers, clinicians, and radiologists. This observation aligns with Kellgren’s earlier findings, where more joints were graded as “1” (minimal or doubtful OA) in a second reading compared to the first, while fewer were graded as “0” (no OA) [3]. Performance was evaluated across multiple configurations using the best with regard to their own dataset splits with 10% test set as well as a standard set which consist of Grade 0 vs Grade 2 4 ,10% models brief summery on Table 1. Table 1 Configuration Task Model Accuracy (%) Predictions (Grade 0 / Grade 1 / Grade 2-4) Accuracy In Standard Set Common Predictions (Between Models) 1 Grade 0 vs. Grade (2-4) (Excluding Grade 1) ResNet18 (No Segmentation) 91.86 986 / / 527 91.86 753 (Grade 0), 333 (Grade 2-4) DenseNet121 (With Segmentation) 89.05 947 / / 548 89.05 2 Grade 0 vs. Grade (1-4) (Including Grade 1) ResNet18 81.09 645 / / 850 89.64 231 (Grade 0), 774 (Grade 1-4) DenseNet121 (With Segmentation) 79.39 307 / / 1,188 87.87 3 Grade 0 vs. Grade 1 ResNet18 75.53 745 / 750 / 85.65 638 (Grade 0), 237 (Grade 1) DenseNet121 (With Segmentation) 72.57 1,151 / 344 / 73.08 4 Grade 1 vs. Grade 2 ResNet18 74.26 / 746 / 749 86.98 652 (Grade 1), 205 (Grade 2) DenseNet121 (With Segmentation) 71.10 / 1,196 / 299 68.64 5 Grade (0-1) vs. Grade (2-4) ResNet18 88.61 1,320 / / 175 88.90 1,270 (Grade 0-1), 57 (Grade 2-4) DenseNet121 (With Segmentation) 86.91 1,388 / / 107 92.46 Overlap Analysis Config. 5 & 1: 277 (Grade 0), 57 (Positive) Config. 5 & 2: 630 (Grade 0), 55 (Positive) Config. 1, 2, & 5: 221 (Grade 0), 55 (Positive) Finally, in a multiclass setting, ResNet18 achieved an overall accuracy of 83.24%; however, its performance on KL Grade 1 was poor, with a precision of 52.38%, recall of 14.77%, and an F1 score of 23.04%, highlighting the inherent ambiguity of this grade. Notably, when KL Grade 0 and KL Grade 1 were combined into a single category, ResNet18’s accuracy improved to 88.61%, while DenseNet121 with segmentation reached 86.91%. In the standard evaluation set, DenseNet121 with segmentation when treating Grade 1 as part of the negative (Grade 0) class achieved the highest accuracy of 92.46%. This suggests that KL Grade 1 more closely resembles KL Grade 0, and that the combined image plus segmentation technique can enhance overall accuracy. Although ResNet18 generally exhibits slightly better performance in all related datasets, DenseNet121 appears to be more sensitive to subtle bone changes due to its segmentation component. In relative datasets, DenseNet121 achieved moderate accuracies (approximately 71-75%) on these classifications, indicating that while ResNet18 may learn overlapping features for Grade 0 as well as some Grade 2 like features from Grade 1, DenseNet121 might capture more distinct features to differentiate Grade 1 from Grade 2, albeit with a moderate accuracy and a notably low F1 score of 32.00% for Grade 2. This observation confirms that although some radiographs classified as Grade 1 do differ from Grade 2, the differences are so subtle with significant overlap between Grade 1 and Grade 0 that most models tend to learn nearly identical features for both Grade 1 and Grade 0. Furthermore, an evaluation of Grad CAM visualizations (Figures 1-13) across 10 deep learning models, using four standard class images, revealed key insights for OA detection. In clear cases, the activation maps demonstrated that joint space narrowing (JSN) plays a critical role often showing high intensity activation in the lateral compartment for true positive images. Although observers detected definite osteophytes and JSN, the models frequently exhibited dominant activation in regions corresponding to definite JSN, in cases of “possible” or “doubtful” JSN, models significant activation on osteophytes, suggesting that JSN may be as significant as, if not more valuable than osteophytes in diagnosing OA. Moreover, consistent lateral compartment activation in two standard images indicates that changes in this region could serve as early indicators of OA. Some models trained to differentiate between Grade 1 versus Grade 0 or Grade 2, displayed slight medial compartment abnormalities; however, these features do not appear to be strongly associated with OA and may represent early signs of other bone diseases. Overall, the nearly indistinguishable features between Grade 0 and Grade 1 as detailed in Section 2.2 underscore the challenge in accurately classifying KL Grade 1. These findings validate earlier results from Scott et al. [6] and demonstrate that the approach achieves higher accuracy than existing methods for early OA detection as well as slightly contradicting research like Altmen et al. on point of “JSN can be one of the dominant features of OA”, these results further support the clinical validity of the KL grading system when enhanced by Grad CAM analysis, despite historical concerns regarding observer bias. Moreover, the work opens avenues for further research into the identification of early OA biomarkers through improved segmentation techniques in collaboration with clinicians and radiologists. Comprehensive training methods, such as e learning module [18], have also proven to improve human diagnostic performance. Finally, historical insights from Kellgren’s research and related studies suggest that KL Grade 1 may reflect early rheumatoid arthritis (RA) rather than a distinct stage of OA [1,2,7,8]. Given that Kellgren and Lawrence originally developed their grading system within the context of rheumatic disease research with Lawrence and Aitken Swan (1952) correlating occupational history with joint degeneration and Kellgren (1956) reporting significant inconsistencies in grading for RA [26] this study suggests that redefining the KL criteria, potentially by eliminating or modifying Grade 1, could reduce ambiguity and improve early detection and classification for both OA and RA, ultimately aiding clinical decision making for interventions such as knee replacement surgery. 2.2) Relation between Grad Cam and Observer eye: Grad CAM (Gradient weighted Class Activation Mapping) is an interpretability technique that provides visual explanations for decisions made by deep learning model, particularly convolutional neural networks (CNNs) developed in 2016 [27]. This method operates by identifying the final convolutional layer of the network layer chosen for its retention of spatial information essential for localization and then computing the gradients of the target class score with respect to the feature maps of that layer. These gradients effectively reveal the significance of each feature map in contributing to the prediction. By globally averaging (or pooling) these gradients, the method assigns a weight to each feature map, which in turn is used to compute a weighted combination that results in a class activation map (CAM). This CAM highlights the regions within the input image that are most influential in the model's decision making process. Once generated, the activation map is resized to match the dimensions of the original image and overlaid as a heatmap, thereby visually emphasizing the areas that the network focused on during prediction. Beyond merely explaining the decisions of the model, Grad CAM offers practical benefits in validating predictions; for instance, in studies involving knee osteoarthritis classification, it is employed to verify that the regions highlighted in the heatmaps correspond to clinically relevant features. Additionally, Grad CAM has broader applications in various domains such as medical imaging, autonomous driving, and any field where understanding the "reasoning" behind model decisions is crucial. Its capacity to provide intuitive visual insights helps bridge the gap between complex model architectures and human interpretability, fostering trust and transparency in deep learning systems. This same method is going to help us more effectively understand KL Grading skim from inside. For effectively understand the and compare between all those models, my criteria were to select 4 standard image group for all 10 models: I) True Negative. II)False Negative. III)True Positive. IV) False Positive. In respect to that from the standard set I was not able to get any common false positive prediction from the models, It made me to just use 3 common images for all 10 models, 4 common images in 6 models, 4 common images in 2 model, 4 common images on 2 models (6+2+2=10). Except in case of false positive we can visualize the impact of Grade 1 on models. OBSERVATION FROM OBSERBER: 9778477L: This radiograph doesn’t show any potential sign of OA. No visible JSN, edge of both joint compartment looks smooth, no visible sign of osteophytes. Assigning Grade 0. (Figure 2,8,10) 9063955R: “Definite” JSN on lateral compartment, Multiple sharp edges can be seen in both lateral and median side of tibia multiple “Definite” osteophytes. White thickening on tibia possible sclerosis. No visible sign of trauma or cartilage degeneration. Assigned Grade 2. (Figure 2,8,10) 9442105L: “Possible” JSN in Median compartment, White thickening on cartilage of tibia possible sclerosis. Smooth edges of both compartments, no possible sign of osteophytes. No visible sign of trauma or cartilage degeneration. Assigned Grade 1. (Figure 2,8,10) 9803621R: “Slight Possibility” of JSN in median compartment. Multiple sharp edges in both median and lateral compartment side of tibia, multiple definite visible osteophytes, slight white thickening on the cartilage of tibia possible sclerosis. No visible sign of trauma or cartilage degeneration. Assigned Grade 2. (Figure 2) 9082901R: “Possible” JSN in median compartment, no sclerosis observed, slight possibility of sharp edge in lateral joint compartment of tibia but no visible osteophytes. No visible sign of trauma or cartilage degeneration. Assigned Grade 0. (Figure 8) 9460287R: “Possible” JSN in median compartment, no sclerosis or sharp edge observed in any of the joint compartment edge. No visible sign of trauma or cartilage degeneration. Assigned Grade 0. (Figure 10) Table 2 Configuration True Negative IMG:9778477L ALL Models True Positive IMG:9063955R ALL Models False Negative IMG:9442105L ALL Models False Positive IMG:9803621R (1 to 6) IMG:9082901R (7,8) IMG:9460287R (9,10) Model Grade 0 vs. Grade (2-4) (Excluding Grade 1) Excluding surrounding area of both joint compartments. Prediction is based partially on both compartments, mostly lateral joint compartment Partially excluding the surrounding area of both joint compartments, mainly lateral Prediction is totally based on lateral joint compartment (more to femur) 1)DenseNet121 (Figure 2) Excluding surrounding area of both joint compartments. Prediction is mainly based on lateral joint compartment Excluding surrounding area of both joint compartments. Prediction is totally based on lateral joint compartment (most to femur) 2)ResNet18 (Figure 3) Grade 0 vs. Grade (1-4) (Including Grade 1) Excluding partially both joint compartments, primarily lateral. Prediction is mainly based on lateral joint compartment Partially excluding lateral joint compartment Prediction is totally based on edge of lateral joint compartment 3)DenseNet121 (Figure 4) Excluding surrounding area of both joint compartments. Prediction is mainly based on lateral joint compartment Excluding almost everything except edge of lateral joint compartment mostly on tibia. Prediction is dominantly based on edge of lateral compartment but partially focusing on medial compartment edge 4)ResNet18 (Figure 5) Grade (0-1) vs. Grade (2-4) Just excluding both joint compartment area. Prediction is only based on lateral compartment Just excluding both joint compartment area. Prediction is only based on edge of lateral compartment dominantly on femur 5)DenseNet121 (Figure 6) Just excluding mostly median joint compartment and partially lateral Prediction is only based on lateral compartment Just excluding both joint compartment area. Prediction is only based on edge of lateral compartment mostly on femur 6)ResNet18 (Figure 7) Grade 0 vs. Grade 1 Excluding just medial joint compartment Prediction is only based on lateral compartment Excluding everything except edge of medial compartment on tibia Excluding almost everything except medial joint compartment dominantly on edge 7)ResNet18 (Figure 9) Almost random slight exclusion of both joint compartment Prediction is almost random if not possibly full joint space Almost random if not, slightly excluding edge of lateral Almost random, If not, slightly excluding medial joint compartment 8)DenseNet121 (Figure 10) Grade 1 vs. Grade 2 Excluding medial and slightly lateral joint compartment Prediction on lateral joint compartment and dominantly on femur Just excluding lateral compartment prediction dominantly based on median compartment edge Prediction just based on median compartment and edge of femur 9)ResNet18 (Figure 12) Excluding medial joint compartment prediction based on lateral side of tibia Prediction is based in total lateral joint compartment Almost random if not, slightly excluding medial joint compartment Prediction is based on edge of lateral joint compartment 10)DenseNet121 (Figure 13) The standard test datasets were constructed with Grade 0 images serving as negatives and Grade 2-4 images as positives. Notably, for models 7 and 8, the positive portion of the datasets, as well as all corresponding positive Grad CAM images remained undefined for the model. Conversely, for models 9 and 10, the negative portion of the dataset was not provided to the model. Based on the predictions from all 10 models, several key observations emerge: The integration of observer findings (Table 2) with Grad CAM visualizations reveals a complex and highly informative correlation that underscores both the strengths and limitations of current KL grading methods, particularly with respect to the ambiguous Grade 1. For instance, consider radiograph 9778477L : the expert observer identified no evidence of OA no joint space narrowing (JSN), smooth joint compartment edges, and no visible osteophytes assigning it a Grade 0. Correspondingly, Grad CAM outputs from models 1-6 consistently showed minimal or no activation in both the medial and lateral compartments, reinforcing the observer’s conclusion and validating the model’s ability to correctly identify a true negative case. In contrast, radiograph 9063955R was assigned a Grade 2 by the observer due to clear JSN in the lateral compartment, multiple sharp edges on both the lateral and medial sides of the tibia, and definite osteophytes with possible sclerosis. Grad CAM maps for this image predominantly highlighted the lateral joint compartment. This lateral activation pattern is clinically significant because it aligns well with the observer’s identification of clear pathological features, thus demonstrating that the models effectively focus on the regions that are most indicative of OA. Ambiguity arises in borderline cases such as radiograph 9442105L . Although the OIA dataset categorizes this image as Grade 2+, the observer noted only possible JSN and subtle changes in the medial compartment without distinct osteophyte formation, leading to an assignment of Grade 1. Grad CAM outputs for this image were less definitive some models configured to distinguish between Grade 0 and Grade 1 showed only minimal activation in the medial compartment or even near random patterns. This is particularly concerning, as the medial compartment in both 9778477L and 9442105L typically lacks robust edge specific activation. Such ambiguous activation may indicate that subtle medial changes could be early signs of pathology, or alternatively, that the inherent vagueness of Grade 1 is causing inconsistent model behavior. Furthermore, when comparing models trained to differentiate between Grade 1 and Grade 2, both ResNet and DenseNet architectures sometimes produced overlapping activation regions. Notably, some ResNet models lacked activation in the lateral compartment a region that might serve as a key differentiator between Grade 1 and Grade 2 thus suggesting that lateral compartment features may be critical for early OA detection. Additional insights emerge when examining training configurations. Excluding Grade 1 from the positive class leads to Grad CAM maps in both ResNet and DenseNet models that display clear, minimal activation in both joint compartments and their surrounding areas, supporting a more definitive classification of true negatives. In contrast, Grade 1 as positive often causes DenseNet models to exhibit slightly random activation patterns. Interestingly, when Grade 1 is instead included as part of the negative class, both model types are more precise in detecting features in both compartments. This observation raises the possibility that radiograph 9442105L might, in fact, be a true negative that was misclassified by human observers a finding that further underscores the problematic nature of Grade 1. False positives also provide valuable context. For example, image 9803621R graded as 2 by the observer due to multiple definite osteophytes was initially graded 0 “Normal” in OIA dataset exhibits lateral compartment activation predominantly along the femur when Grade 1 is excluded, closely mirroring true positive patterns. However, inclusion of Grade 1 shifts this activation towards the lateral compartment edge, reinforcing the importance of lateral features in confirming OA. In the case of image 9082901R , which was graded as 0 despite the observer noting a slight possibility of a sharp lateral edge and potential JSN, Grad CAM outputs from models trained on a Grade 0 versus Grade 1 configuration showed near random behavior with little lateral activation. Such variability further highlights the intrinsic ambiguity associated with Grade 1. Finally, image 9460287R , assigned as Grade 0 due to the absence of definitive OA features (aside from a slight possibility of medial JSN), produced contradictory Grad CAM results when models attempted to differentiate between Grade 1 and Grade 2, again underscoring the challenge posed by Grade 1’s subtle and overlapping features. 2.3) Explanation of Deep Learning Approach: After evaluating multiple models and training strategies with 75 % training, 5% validation and 10% testing set, was determined that training the ResNet18 model without segmentation yielded the best results when grayscale images were replicated across three channels. Specifically, using an adaptive learning rate starting at 0.0001 with a decay schedule in each 5 epoch that reduced the rate by 30% if no improvement was observed and applying rigorous data augmentation produced optimal performance for ResNet18. For training with segmentation, The study developed a novel approach designed to reduce dependency on segmentation quality and mitigate misclassification issues. In this method, I created a stacked input image where the first and third channels contain the original grayscale image, and the middle channel contains the corresponding segmentation mask. This approach proved particularly effective when applied to the DenseNet121 model. With a fixed learning rate of 0.0001 and the same rigorous transformation pipeline, the DenseNet121 model achieved superior performance. Notably, while ResNet18 reached its best accuracy within approximately 100 epochs, DenseNet121 required around 400 epochs to converge but ultimately provided more reliable outcomes. Both models were trained using a binary classification framework. The segmentation training method is unique to this study and, to my knowledge, has not been previously reported. Key Libraries and tools used: Hardware: Laptop: NVIDIA RTX 3060, AMD Ryzen 7 5800H RAM: 24 GB DDR4 3200 MHz Storage: 100 GB disk drive (SSD) Operating System: Windows 11 Software: IDE: Visual Studio Code with Jupyter Notebook extension Environment: Conda, Python 3.11 Key Libraries: PyTorch, NumPy, Pandas, Matplotlib, TorchVision, Scikit learn, Pillow Declarations Author Contribution R.I.M. wrote the manuscript, conducted the research, developed the deep learning models, analyzed the results, and prepared all figures and tables. The author reviewed and approved the final manuscript. References LAWRENCE JS, AITKEN SWAN J. Rheumatism in miners. Part I: Rheumatic complaints. Br J Ind Med . 1952;9(1):1 18. doi:10.1136/oem.9.1.1 Kellgren JH, Lawrence JS. Rheumatism in miners. Part II: X ray study. Occupational and Environmental Medicine. 1952;9(3):197 207. doi:10.1136/oem.9.3.197 Braun HJ, Gold GE. Diagnosis of osteoarthritis: imaging. Bone . 2012;51(2):278 288. doi:10.1016/j.bone.2011.11.019 Kellgren JH, Lawrence JS. Radiological assessment of osteo arthrosis. Ann Rheum Dis. 1957;16(4):494 502. doi:10.1136/ard.16.4.494 Felson DT, Naimark A, Anderson J, Kazis L, Castelli W, Meenan RF. The prevalence of knee osteoarthritis in the elderly. The Framingham Osteoarthritis Study. Arthritis Rheum . 1987;30(8):914 918. doi:10.1002/art.1780300811 Scott WW Jr, Lethbridge Cejku M, Reichle R, Wigley FM, Tobin JD, Hochberg MC. Reliability of grading scales for individual radiographic features of osteoarthritis of the knee. The Baltimore longitudinal study of aging atlas of knee osteoarthritis. Invest Radiol . 1993;28(6):497 501. Arnett FC, Edworthy SM, Bloch DA, et al. The American Rheumatism Association 1987 revised criteria for the classification of rheumatoid arthritis. Arthritis Rheum . 1988;31(3):315 324. doi:10.1002/art.1780310302 Altman R, Asch E, Bloch D, et al. Development of criteria for the classification and reporting of osteoarthritis. Classification of osteoarthritis of the knee. Diagnostic and Therapeutic Criteria Committee of the American Rheumatism Association. Arthritis Rheum . 1986;29(8):1039 1049. doi:10.1002/art.1780290816 Irrgang JJ, Anderson AF, Boland AL, et al. Responsiveness of the International Knee Documentation Committee Subjective Knee Form. Am J Sports Med . 2006;34(10):1567 1573. doi:10.1177/0363546506288855 Hefti F, Müller W, Jakob RP, Stäubli HU. Evaluation of knee ligament injuries with the IKDC form. Knee Surg Sports Traumatol Arthrosc . 1993;1(3 4):226 234. doi:10.1007/BF01560215 FAIRBANK TJ. Knee joint changes after meniscectomy. J Bone Joint Surg Br . 1948;30B(4):664 670. Tapper EM, Hoover NW. Late results after meniscectomy. J Bone Joint Surg Am . 1969;51(3):. Brandt KD, Fife RS, Braunstein EM, Katz B. Radiographic grading of the severity of knee osteoarthritis: relation of the Kellgren and Lawrence grade to a grade based on joint space narrowing, and correlation with arthroscopic evidence of articular cartilage degeneration. Arthritis Rheum . 1991;34(11):1381 1386. doi:10.1002/art.1780341106 Galli M, De Santis V, Tafuro L. Reliability of the Ahlbäck classification of knee osteoarthritis. Osteoarthritis Cartilage . 2003;11(8):580 584. doi:10.1016/s1063 4584(03)00095 5 Scheller G, Sobau C, Bülow JU. Arthroscopic partial lateral meniscectomy in an otherwise normal knee: Clinical, functional, and radiographic results of a long term follow up study. Arthroscopy . 2001;17(9):946 952. doi:10.1053/jars.2001.28952 Schroeder Boersch H, Töws P, Jani L. Reproduzierbarkeit von radiologischen Arthrosemerkmalen. Evaluierung eines Scores zur radiologischen Klassifikation von degenerativen Veränderungen bei Gonarthrose [Reproducibility of radiologic markers of osteoarthritis. Evaluating a score for radiologic classification of degenerative changes in osteoarthritis of the knee joint]. Z Orthop Ihre Grenzgeb . 1998;136(4):293 297. doi:10.1055/s 2008 1053740 Wright RW; MARS Group. Osteoarthritis Classification Scales: Interobserver Reliability and Arthroscopic Correlation. J Bone Joint Surg Am . 2014;96(14):1145 1151. doi:10.2106/JBJS.M.00929 Hayes B, Kittelson A, Loyd B, Wellsandt E, Flug J, Stevens Lapsley J. Assessing radiographic knee osteoarthritis: an online training tutorial for the Kellgren Lawrence grading scale. MedEdPORTAL. 2016;12:10503. doi:10.15766/mep_2374 8265.10503 Bagge E, Bjelle A, Valkenburg HA, Svanborg A. Prevalence of radiographic osteoarthritis in two elderly European populations. Rheumatol Int . 1992;12(1):33 38. doi:10.1007/BF00246874 Spector TD, Cooper C. Radiographic assessment of osteoarthritis in population studies: whither Kellgren and Lawrence?. Osteoarthritis Cartilage . 1993;1(4):203 206. doi:10.1016/s1063 4584(05)80325 5 Eckstein F, Kwoh CK, Link TM. Imaging research results from the Osteoarthritis Initiative (OAI): a review and lessons learned 10 years after start of enrolment. Ann Rheum Dis. 2014;73(7):1289 1300. doi:10.1136/annrheumdis 2014 205310 Chen P. Knee osteoarthritis severity grading dataset. Mendeley Data. 2018;V1. doi:10.17632/56rmx5bjcr.1 He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016:770 778. doi:10.1109/CVPR.2016.90 Huang G, Liu Z, Van Der Maaten L, Weinberger KQ. Densely connected convolutional networks. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2017:2261 2269. doi:10.1109/CVPR.2017.243 Tan M, Le QV. EfficientNet: rethinking model scaling for convolutional neural networks. Int Conf Mach Learn. 2019. doi:10.48550/arXiv.1905.11946 KELLGREN JH. Radiological signs of rheumatoid arthritis; a study of observer differences in the reading of hand films. Ann Rheum Dis . 1956;15(1):55 60. doi:10.1136/ard.15.1.55 Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad CAM: Visual explanations from deep networks via gradient based localization. Int J Comput Vis. 2019;128(2):336 359. doi:10.1007/s11263 019 01228 7 Tiulpin A, Thevenot J, Rahtu E, Lehenkari P, Saarakkala S. Automatic knee osteoarthritis diagnosis from plain radiographs: a deep learning based approach. Sci Rep. 2018;8(1):1727. doi:10.1038/s41598 018 20132 7 Shamir L, Ling SM, Scott W, Hochberg M, Ferrucci L, Goldberg IG. Early detection of radiographic knee osteoarthritis using computer aided analysis. Osteoarthritis Cartilage. 2009;17(10):1307 1312. doi:10.1016/j.joca.2009.04.010 Yeoh PSQ, Lai KW, Goh SL, Hasikin K, Hum YC, Tee YK, Dhanalakshmi S. Emergence of deep learning in knee osteoarthritis diagnosis. Computational Intelligence and Neuroscience. 2021;2021(1):4931437. doi:10.1155/2021/4931437 Kijowski R, Fritz J, Deniz CM. Deep learning applications in osteoarthritis imaging. Skeletal Radiol . 2023;52(11):2225 2238. doi:10.1007/s00256 023 04296 6 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5983981","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":415002674,"identity":"6efad9b7-79dd-421c-b71f-ef4d5d529807","order_by":0,"name":"Rifatul Islam Majumder","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABDklEQVRIie2PsWoCQRCG5xg4mwXbBeX2CQIjwsXiMG+R+iRw1xiwEssDwXuFwBV5BUPAesO2IfUGG1Ok3xAIKSyye5bRM+kE9yt2Yfg//hkAj+cUwWBj3wGDlv0M8cgN5aZRQQIJnAECBHeTpF8raXNNrYBTkJlsVLhhk3JZYvD5seBdMcf3JSOV31+rN9syjC6K/UpXIfKnBbfhMNac1O3DOiOr3PRjuV/h2JY7BSHW5JQqdYocrQ4qiN9OEfPWl07tYr0qN8eUsG4BxWItKUtFZ3y0JRw8v7hb2PS1oKS37IwnMqWGW9oK9WyaXImyXK23Wy5ElT8aMxtGh5TfUJ2kv8YdovhP2uPxeM6BH9WCXI3KoNgoAAAAAElFTkSuQmCC","orcid":"","institution":"","correspondingAuthor":true,"prefix":"","firstName":"Rifatul","middleName":"Islam","lastName":"Majumder","suffix":""}],"badges":[],"createdAt":"2025-02-07 21:53:06","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5983981/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5983981/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":76279130,"identity":"ffd3cefc-6f1a-4331-8d67-1b0bc9b0783e","added_by":"auto","created_at":"2025-02-14 10:22:30","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":299655,"visible":true,"origin":"","legend":"\u003cp\u003e4-images-6-models-3-common_all\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/cc4fd6d217c11a3f6a93eb38.png"},{"id":76279097,"identity":"cefa5c74-0a68-4050-b011-465596968f1b","added_by":"auto","created_at":"2025-02-14 10:22:28","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":233222,"visible":true,"origin":"","legend":"\u003cp\u003eDenseNet121-Grade 0 vs. Grade 2-4\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/fbac1621275af392edc7da1d.png"},{"id":76279090,"identity":"92ffcbef-b819-4d96-be86-01bc7c9415e7","added_by":"auto","created_at":"2025-02-14 10:22:28","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":227892,"visible":true,"origin":"","legend":"\u003cp\u003eResNet18-Grade 0 vs. Grade 2-4\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/0bffcf51044a4686cfb964dd.png"},{"id":76279088,"identity":"03bae8f0-3cd7-491c-a177-98d5f99955f8","added_by":"auto","created_at":"2025-02-14 10:22:28","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":228203,"visible":true,"origin":"","legend":"\u003cp\u003eDenseNet121-Grade 0 vs. Grade 1-4\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/76f9a66b130c2844f4691afc.png"},{"id":76279133,"identity":"e12de63c-af26-4c26-b07a-b3c739bae64d","added_by":"auto","created_at":"2025-02-14 10:22:30","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":228818,"visible":true,"origin":"","legend":"\u003cp\u003eResNet-Grade 0 vs. Grade 1-4\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/ff4949a8c5164ce8897003df.png"},{"id":76280446,"identity":"5919374b-e5da-4a11-82cb-deb1819e8113","added_by":"auto","created_at":"2025-02-14 10:30:28","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":228983,"visible":true,"origin":"","legend":"\u003cp\u003eDenseNet121-Grade 0-1 vs. Grade 2-4\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/adde691eeda0c0383ab6bf61.png"},{"id":76280442,"identity":"e32b3663-7676-4ed4-accc-e15885f8a328","added_by":"auto","created_at":"2025-02-14 10:30:27","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":222255,"visible":true,"origin":"","legend":"\u003cp\u003eResNet18-Grade 0-1 vs. Grade 2-4\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/32510201921eadcc2f04ac13.png"},{"id":76279083,"identity":"75f232fc-2b62-4664-b028-465f74b4a660","added_by":"auto","created_at":"2025-02-14 10:22:27","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":288272,"visible":true,"origin":"","legend":"\u003cp\u003e1-image-2-models (0 vs 1)-9082901R-added as-incorrect true pred\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/5fd80cc039c286127b26f0fa.png"},{"id":76280453,"identity":"c7e581c2-9157-4f8e-8926-ac31e3c55d09","added_by":"auto","created_at":"2025-02-14 10:30:29","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":224608,"visible":true,"origin":"","legend":"\u003cp\u003eResNet18-Grade 0 vs. Grade 1\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/1f5d03216d49485cd1e95bb5.png"},{"id":76279101,"identity":"b9966ee3-2806-4536-823c-5d2d4c00ec1b","added_by":"auto","created_at":"2025-02-14 10:22:28","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":244126,"visible":true,"origin":"","legend":"\u003cp\u003eDenseNet121-Grade 0 vs. Grade 1\u003c/p\u003e","description":"","filename":"10.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/71a24521fb2e63747a9b121e.png"},{"id":76279081,"identity":"9b1faa02-1ee6-4bf6-8943-c54ba0d19902","added_by":"auto","created_at":"2025-02-14 10:22:27","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":296110,"visible":true,"origin":"","legend":"\u003cp\u003e1-image-2-models (1 vs 2)-940287R-added as-incorrect true pred\u003c/p\u003e","description":"","filename":"11.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/93ccbbc1f2d6664fb7c83f0d.png"},{"id":76279114,"identity":"d8b80b5c-a89c-4ee9-8464-e172c7de1f40","added_by":"auto","created_at":"2025-02-14 10:22:29","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":235607,"visible":true,"origin":"","legend":"\u003cp\u003eResNet18-Grade 1 vs. Grade 2\u003c/p\u003e","description":"","filename":"12.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/18814a0de7ad4d2ade8c237b.png"},{"id":76279107,"identity":"5515fceb-69a4-4bc1-b101-4f5405ac9dbf","added_by":"auto","created_at":"2025-02-14 10:22:29","extension":"png","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":239747,"visible":true,"origin":"","legend":"\u003cp\u003eDenseNet121-Grade 1 vs. Grade 2\u003c/p\u003e","description":"","filename":"13.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/fe3491d60f7080a59aedc6fb.png"},{"id":76280454,"identity":"c5fa8324-61a3-449e-b3c1-6e0dd3d2d606","added_by":"auto","created_at":"2025-02-14 10:30:29","extension":"png","order_by":14,"title":"Figure 14","display":"","copyAsset":false,"role":"figure","size":138539,"visible":true,"origin":"","legend":"\u003cp\u003eFigure A (Osteophytes)\u003c/p\u003e","description":"","filename":"A.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/dff46b1210a61c4226ec4f83.png"},{"id":76279105,"identity":"7f572cce-b669-4e21-9bc2-ea6af3c54f2c","added_by":"auto","created_at":"2025-02-14 10:22:29","extension":"png","order_by":15,"title":"Figure 15","display":"","copyAsset":false,"role":"figure","size":80565,"visible":true,"origin":"","legend":"\u003cp\u003eFigure B (JSN)\u003c/p\u003e","description":"","filename":"B.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/20830f936c9a7308260acf77.png"},{"id":76281064,"identity":"98d01822-7777-4eab-bec4-daf2834db2d2","added_by":"auto","created_at":"2025-02-14 10:38:30","extension":"png","order_by":16,"title":"Figure 16","display":"","copyAsset":false,"role":"figure","size":63670,"visible":true,"origin":"","legend":"\u003cp\u003eFigure C (Sclerosis)\u003c/p\u003e","description":"","filename":"C.png","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/762b71fa1f566d259d492361.png"},{"id":76282624,"identity":"bee95e64-5d25-4e32-9702-44c73f3f2e08","added_by":"auto","created_at":"2025-02-14 10:46:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":6418950,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5983981/v1/ab7a8742-e2e3-4094-9ee8-d2afdba2328f.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Redefining Early Osteoarthritis Detection: Deep Learning Meets Historical Radiography","fulltext":[{"header":"Full Text","content":"\u003cp\u003eTo address these issues, e‐learning module designed to standardize KL grading among novice clinicians, resulting in significantly improved interrater and intrarater reliability (weighted \u0026kappa; = 0.82 and 0.85, respectively) [17]. Complementing this educational approach, I employed deep learning techniques to evaluate the clinical relevance of KL Grade 1 and to identify potential biomarkers for early OA detection. Using 45 degree posteroanterior weight‐bearing knee radiographs from the OIA dataset [18,19], I trained multiple models including ResNet18 and DenseNet121 architectures, with and without segmentation [20-22] described in Section 2.3. These experiments consistently demonstrated that KL Grade 1 is a major source of misclassification, as models frequently conflate it with Grade 0, thereby blurring the diagnostic boundary for early OA detection [23].\u003c/p\u003e\n\u003cp\u003eIntegration of Gradient weighted Class Activation Mapping (Grad CAM) [24] provided valuable visual insights into model decision making. The Grad CAM analyses (see Section 2.2, Figures 1-13) revealed that early OA is primarily associated with features in the lateral compartment and the lateral femoral condyle, with joint space narrowing being as significant as, if not more so than, osteophyte formation. In contrast, subtle abnormalities in the medial compartment were not consistently linked to early OA. Furthermore, the activation patterns for KL Grade 1 frequently overlapped with those for both normal (Grade 0) and advanced OA (Grade 2+), corroborating historical concerns regarding observer bias [25-27].\u003c/p\u003e\n\u003cp\u003eIn addition, I developed a unique stacked input approach that combines the grayscale image, the segmentation mask, and the grayscale image. This method, integrated with deep learning model, outperforms current techniques for automating early OA detection [28-31]. To my knowledge, this segmentation training strategy is novel and has not been previously reported. Furthermore, no prior study has combined the KL grading system with expert human readings and model heatmap analyses to explore definitive early biomarkers of OA.\u003c/p\u003e\n\u003cp\u003eThese findings suggest that refining the KL grading criteria, potentially by redefining or eliminating the ambiguous Grade 1 could enhance early OA diagnosis. Moreover, relying solely on osteophyte formation is insufficient for early OA detection; a redefined KL grading scheme may not only improve OA classification but also help distinguish OA from other bone diseases, such as rheumatoid arthritis. The combination of standardized training and advanced deep learning with interpretability techniques holds promise for improving diagnostic accuracy and clinical decision making in OA management, with significant implications for early diagnosis and tailored treatment planning.\u003c/p\u003e\n\u003cp\u003e2.1) Descriptive Overview:\u003c/p\u003e\n\u003cp\u003eThis research builds upon the foundational work of Kellgren and Lawrence, particularly their studies from 1952 and 1957. In 1952, Lawrence and Aitken Swan highlighted the occupational influences on rheumatic complaints by revealing that coal miners exhibited significantly higher incidences of knee osteoarthritis (OA) compared to the general population [1]. That same year, Kellgren and Lawrence demonstrated a strong association between occupational strain and radiological signs of knee OA, establishing a statistically significant relationship (P \u0026lt; 0.01) between knee pain and OA severity [2]. Their grading system, which categorized osteoarthritic changes into five levels, later evolved into the widely used Kellgren Lawrence (KL) grading system formalized by the World Health Organization in 1961.\u003c/p\u003e\n\u003cp\u003eA recent study by Braun and Gold (2011) emphasizes that OA is a disease of the entire joint and explores the role of imaging particularly for knee OA in its diagnosis. The KL grading system, which relies on radiographic features such as osteophytes and joint space narrowing (JSN), remains the standard for OA classification despite limitations, including poor sensitivity to \u003cstrong\u003eearly cartilage damage\u003c/strong\u003e [3]. Felson et al. (1987), as part of the Framingham Osteoarthritis Study, assessed the prevalence of knee OA in a population based cohort using a 0-4 scale based on the Kellgren and Lawrence criteria [4],[5]. Additionally, Scott et al. developed and validated an atlas for evaluating individual radiographic features of knee OA within the Baltimore Longitudinal Study of Aging. In their study, four trained readers evaluated 30 standing anterior posterior knee radiographs for eight features including \u003cstrong\u003emedial and lateral osteophytes, JSN, sclerosis, osteophytes of the tibial spines, and chondrocalcinosis\u0026nbsp;\u003c/strong\u003ealong with the overall KL global scale, reporting inter reader reliability from 0.63 to 0.83 and intra reader reliability from 0.82 to 0.95 [6].\u003c/p\u003e\n\u003cp\u003eDespite KL grades being longstanding, studies since 1957 have revealed substantial inter \u0026nbsp;and intra observer variability in the KL grading system. For example, the same observer demonstrated a high correlation (r = 0.88) when reading metacarpophalangeal joints twice, yet the lowest correlation was found for the dorso lumbar spine (r = 0.42). Knee OA grading showed relatively strong agreement (r = 0.83) both within and between observers, whereas wrist OA exhibited the lowest agreement (r = 0.10) [4]. This variability particularly in assessing features such as osteophytes, sclerosis, and JSN remains a major concern. The American Rheumatism Association (ARA) subcommittee, in its 1981 report and subsequent 1987 revision, described OA as a heterogeneous disease, categorizing it as either idiopathic (no known cause) or secondary (linked to conditions such as rheumatoid arthritis [RA]) [7]. Moreover, Altman et al. demonstrated that RA patients also exhibit joint space narrowing, with \u003cstrong\u003eosteophytes present in 62% of cases\u003c/strong\u003e, indicating that reliance solely on radiological features may be problematic [8].\u003c/p\u003e\n\u003cp\u003eMultiple standard systems including the International Knee Documentation Committee (IKDC) radiographic scale [9],[10], Fairbank [11],[12], Brandt et al. [13], Ahlb\u0026auml;ck [14], and J\u0026auml;ger Wirth [15],[16] primarily depend on JSN and osteophytes. This dependence raises concerns regarding the validity of using JSN as a definitive feature due to its subjectivity and low reliability. In a prospective study of 632 patients undergoing revision ACL reconstruction (the Multicenter ACL Revision Study [MARS]), six radiographic classification systems were evaluated by comparing anteroposterior (AP) and 45\u0026deg; posteroanterior (Rosenberg) weight bearing radiographs with arthroscopic findings. Overall interobserver reliability was moderate for AP images (ICC = 0.55; 95% CI, 0.53-0.56) and good for Rosenberg images (ICC = 0.63; 95% CI, 0.61-0.65). Specifically, the KL system had an ICC of 0.38 (95% CI, 0.33-0.43) on AP views and 0.54 (95% CI, 0.48-0.59) on Rosenberg views, while the correlation between KL grading and arthroscopic cartilage degeneration was higher with Rosenberg radiographs (Spearman rho = 0.42; 95% CI, 0.33-0.49) than with AP radiographs (rho = 0.30; 95% CI, 0.23-0.38). Although the IKDC scale demonstrated the best overall performance (AP: ICC = 0.59, 95% CI, 0.55-0.63; Rosenberg: ICC = 0.66, 95% CI, 0.62-0.71), the comparatively lower reliability of the KL system underscores its limitations in accurately classifying early osteoarthritic changes based solely on radiographic features [17].\u003c/p\u003e\n\u003cp\u003eThe literature further indicates that the KL grading system faces significant limitations in differentiating between features such as \u0026ldquo;possible\u0026rdquo; versus \u0026ldquo;definite\u0026rdquo; osteophytes and \u0026ldquo;doubtful\u0026rdquo; versus \u0026ldquo;definite\u0026rdquo; joint space narrowing. These variations, which can differ among observers, highlight a lack of standardization and underscore the need for comprehensive training. To address this issue, an e learning tool (Articulate Presenter) was developed. The tutorial refined over three iterative feedback rounds from clinical researchers, radiologists, and graduate students focused on key features such as osteophytes and JSN to distinguish KL Grades 0-4. Forty seven health sciences graduate students with no prior KL grading experience completed the training and subsequently graded 30 knee radiographs (including 15 duplicates for reliability testing). The results demonstrated almost perfect interrater reliability (weighted kappa, \u0026kappa;w = 0.82) and intrarater reliability (\u0026kappa;w = 0.85), \u003cstrong\u003esignificantly improving upon standard reliability levels reported in the literature\u003c/strong\u003e. Students reported increased confidence and understanding of the KL scale although they noted difficulties distinguishing between Grades 3 and 4 suggesting that standardized online training can enhance the consistency of KL grading in both research and clinical practice [18] (see FIGURE A for a visualization of the differences between possible and definite osteophytes and Figure B we can visualize JSN and Figure C sclerosis).\u003c/p\u003e\n\u003cp\u003eAlthough osteophytes are generally a reliable indicator, the lack of a clear definition for what constitutes an osteophyte hinders early OA detection, priority in current research. Bagge et al. conducted a comparative study on the prevalence of radiographic osteoarthritis (ROA) in two elderly European populations (G\u0026ouml;teborg, Sweden, and Zoetermeer, The Netherlands). An interobserver reliability assessment on 150 radiographic films revealed that a binary \u0026ldquo;abnormal\u0026rdquo; versus \u0026ldquo;normal\u0026rdquo; classification yielded higher agreement percentages and kappa values than a five point scale, while an intraobserver analysis of 50 films demonstrated excellent reliability (kappa values well over 0.75). The study found that hand ROA was more prevalent in the G\u0026ouml;teborg population, whereas knee ROA prevalence was similar between the two groups. Notably, neither population showed a sigwnificant age related increase in hand ROA; however, while knee ROA did not increase with age in the Swedish cohort, it exhibited a significant age related increase in both sexes in the Zoetermeer population [19].\u003c/p\u003e\n\u003cp\u003eFrom these observations, hypothesize that many of the validity issues associated with the KL grading system stem from its progressive nature, same concern also raised by Spector and Cooper reappraised the radiographic assessment of OA in population studies, focusing on the widely used KL grading system that assigns grades 0-4 based on specific radiographic features. They emphasize that OA, the most common joint disorder globally, is a complex interplay of clinical, radiographic, and pathological changes, and that an ideal radiographic system should be accurate, reproducible, noninvasive, convenient, and inexpensive. Spector and Cooper highlight two major criticisms: first, inconsistencies in definitions for example, knee Grade 2 has been variably defined as \u003cstrong\u003e\u0026ldquo;definite osteophytes with minimal joint space narrowing\u0026rdquo; versus \u0026ldquo;definite osteophytes but unimpaired joint space\u0026rdquo;\u003c/strong\u003e which contribute to poor interobserver and intercenter reproducibility; and second, an \u003cstrong\u003eoveremphasis on osteophytes\u003c/strong\u003e that assumes joint space narrowing follows osteophyte formation, potentially excluding cases with \u003cstrong\u003esignificant joint space loss but no visible osteophytes\u003c/strong\u003e. Their review, which included a preliminary analysis of 1,954 knee radiographs, demonstrated that while all measures (joint space narrowing, osteophyte formation, and sclerosis) showed good reproducibility, the presence of definite osteophytes best predicted knee pain, whereas for hip OA, joint space measurements were more reproducible and strongly associated with pain though incorporating osteophyte data further improved prediction. This reappraisal underscores that no single radiographic measure is universally applicable across different joints, and that standardizing definitions and evaluating individual joint features may enhance the reproducibility and clinical utility of the KL grading system [20].\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eUsing the OIA dataset [21],[22] with 45\u0026deg; posteroanterior (Rosenberg) weight bearing radiographs, multiple deep learning models including ResNets [23], DenseNets [24], and EfficientNets [25], both with and without segmentation, and with fixed and adaptive learning rates across various training transformations were initially tested. From these, two deep learning models were selected for their effectiveness: five ResNet18 models without segmentation and five DenseNet121 models with segmentation. These models were trained on the Knee Osteoarthritis Severity Grading Dataset [5] to classify knee OA.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe findings indicate that KL grade 1 is a major source of misclassification, causing overlap between KL grade 0 and KL grade 2+. In all instances, KL grade 1 was more closely related to KL grade 0, further blurring the diagnostic boundary for observers, clinicians, and radiologists. This observation aligns with Kellgren\u0026rsquo;s earlier findings, where more joints were graded as \u0026ldquo;1\u0026rdquo; (minimal or doubtful OA) in a second reading compared to the first, while fewer were graded as \u0026ldquo;0\u0026rdquo; (no OA) [3].\u003c/p\u003e\n\u003cp\u003ePerformance was evaluated across multiple configurations using the best with regard to their own dataset splits with 10% test set as well as a standard set which consist of Grade 0 vs Grade 2 4 ,10% \u0026nbsp; models brief summery on Table 1.\u003c/p\u003e\n\u003cp\u003eTable 1\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eConfiguration\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTask\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAccuracy (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ePredictions (Grade 0 / Grade 1 / Grade 2-4)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eAccuracy\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eIn\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eStandard\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eSet\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCommon Predictions (Between Models)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003eGrade 0 vs. Grade (2-4) (Excluding Grade 1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eResNet18 (No Segmentation)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e91.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e986 / \u0026nbsp; / 527\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e91.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e753 (Grade 0), 333 (Grade 2-4)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDenseNet121 (With Segmentation)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e89.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e947 / \u0026nbsp; / 548\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e89.05\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e\u003cstrong\u003e2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003eGrade 0 vs. Grade (1-4) (Including Grade 1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eResNet18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e81.09\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e645 / \u0026nbsp; / 850\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e89.64\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e231 (Grade 0), 774 (Grade 1-4)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDenseNet121 (With Segmentation)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e79.39\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e307 / \u0026nbsp; / 1,188\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e87.87\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003eGrade 0 vs. Grade 1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eResNet18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e75.53\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e745 / 750 / \u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e85.65\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e638 (Grade 0), 237 (Grade 1)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDenseNet121 (With Segmentation)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e72.57\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1,151 / 344 / \u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e73.08\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003eGrade 1 vs. Grade 2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eResNet18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e74.26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026nbsp; / 746 / 749\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e86.98\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e652 (Grade 1), 205 (Grade 2)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDenseNet121 (With Segmentation)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e71.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026nbsp; / 1,196 / 299\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e68.64\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e\u003cstrong\u003e5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003eGrade (0-1) vs. Grade (2-4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eResNet18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e88.61\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1,320 / \u0026nbsp; / 175\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e88.90\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e1,270 (Grade 0-1), 57 (Grade 2-4)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDenseNet121 (With Segmentation)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e86.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1,388 / \u0026nbsp; / 107\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e92.46\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003e\n \u003cp\u003e\u003cstrong\u003eOverlap Analysis\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"3\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd rowspan=\"3\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd rowspan=\"3\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd colspan=\"3\" valign=\"top\"\u003e\n \u003cp\u003eConfig. 5 \u0026amp; 1: 277 (Grade 0), 57 (Positive)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\"\u003e\n \u003cp\u003eConfig. 5 \u0026amp; 2: 630 (Grade 0), 55 (Positive)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\"\u003e\n \u003cp\u003eConfig. 1, 2, \u0026amp; 5: 221 (Grade 0), 55 (Positive)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eFinally, in a multiclass setting, ResNet18 achieved an overall accuracy of 83.24%; however, its performance on KL Grade 1 was poor, with a precision of 52.38%, recall of 14.77%, and an F1 score of 23.04%, highlighting the inherent ambiguity of this grade. Notably, when KL Grade 0 and KL Grade 1 were combined into a single category, ResNet18\u0026rsquo;s accuracy improved to 88.61%, while DenseNet121 with segmentation reached 86.91%. In the standard evaluation set, DenseNet121 with segmentation when treating Grade 1 as part of the negative (Grade 0) class achieved the highest accuracy of 92.46%. This suggests that KL Grade 1 more closely resembles KL Grade 0, and that the combined image plus segmentation technique can enhance overall accuracy. Although ResNet18 generally exhibits slightly better performance in all related datasets, DenseNet121 appears to be more sensitive to subtle bone changes due to its segmentation component. In relative datasets, DenseNet121 achieved moderate accuracies (approximately 71-75%) on these classifications, indicating that while ResNet18 may learn overlapping features for Grade 0 as well as some Grade 2 like features from Grade 1, DenseNet121 might capture more distinct features to differentiate Grade 1 from Grade 2, albeit with a moderate accuracy and a notably low F1 score of 32.00% for Grade 2. This observation confirms that although some radiographs classified as Grade 1 do differ from Grade 2, the differences are so subtle with significant overlap between Grade 1 and Grade 0 that most models tend to learn nearly identical features for both Grade 1 and Grade 0.\u003c/p\u003e\n\u003cp\u003eFurthermore, an evaluation of Grad CAM visualizations (Figures 1-13) across 10 deep learning models, using four standard class images, revealed key insights for OA detection. In clear cases, the activation maps demonstrated that joint space narrowing (JSN) plays a critical role often showing high intensity activation in the lateral compartment for true positive images. Although observers detected definite osteophytes and JSN, the models frequently exhibited dominant activation in regions corresponding to definite JSN, in cases of \u0026ldquo;possible\u0026rdquo; or \u0026ldquo;doubtful\u0026rdquo; JSN, models significant activation on osteophytes, suggesting that JSN may be as significant as, if not more valuable than osteophytes in diagnosing OA. Moreover, consistent lateral compartment activation in two standard images indicates that changes in this region could serve as early indicators of OA. Some models trained to differentiate between Grade 1 versus Grade 0 or Grade 2, displayed slight medial compartment abnormalities; however, these features do not appear to be strongly associated with OA and may represent early signs of other bone diseases. Overall, the nearly indistinguishable features between Grade 0 and Grade 1 as detailed in Section 2.2 underscore the challenge in accurately classifying KL Grade 1.\u003c/p\u003e\n\u003cp\u003eThese findings validate earlier results from Scott et al. [6] and demonstrate that the approach achieves higher accuracy than existing methods for early OA detection as well as slightly contradicting research like Altmen et al. on point of \u0026nbsp; \u0026ldquo;JSN can be one of the dominant features of OA\u0026rdquo;, these results further support the clinical validity of the KL grading system when enhanced by Grad CAM analysis, despite historical concerns regarding observer bias. Moreover, the work opens avenues for further research into the identification of early OA biomarkers through improved segmentation techniques in collaboration with clinicians and radiologists. Comprehensive training methods, such as e learning module [18], have also proven to improve human diagnostic performance. Finally, historical insights from Kellgren\u0026rsquo;s research and related studies suggest that KL Grade 1 may reflect early rheumatoid arthritis (RA) rather than a distinct stage of OA [1,2,7,8]. Given that Kellgren and Lawrence originally developed their grading system within the context of rheumatic disease research with Lawrence and Aitken Swan (1952) correlating occupational history with joint degeneration and Kellgren (1956) reporting significant inconsistencies in grading for RA [26] this study suggests that redefining the KL criteria, potentially by eliminating or modifying Grade 1, could reduce ambiguity and improve early detection and classification for both OA and RA, ultimately aiding clinical decision making for interventions such as knee replacement surgery.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.2) Relation between Grad Cam and Observer eye:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eGrad CAM (Gradient weighted Class Activation Mapping) is an interpretability technique that provides visual explanations for decisions made by deep learning model, particularly convolutional neural networks (CNNs) developed in 2016 [27]. This method operates by identifying the final convolutional layer of the network layer chosen for its retention of spatial information essential for localization and then computing the gradients of the target class score with respect to the feature maps of that layer. These gradients effectively reveal the significance of each feature map in contributing to the prediction. By globally averaging (or pooling) these gradients, the method assigns a weight to each feature map, which in turn is used to compute a weighted combination that results in a class activation map (CAM). This CAM highlights the regions within the input image that are most influential in the model\u0026apos;s decision making process. Once generated, the activation map is resized to match the dimensions of the original image and overlaid as a heatmap, thereby visually emphasizing the areas that the network focused on during prediction. Beyond merely explaining the decisions of the model, Grad CAM offers practical benefits in validating predictions; for instance, in studies involving knee osteoarthritis classification, it is employed to verify that the regions highlighted in the heatmaps correspond to clinically relevant features. Additionally, Grad CAM has broader applications in various domains such as medical imaging, autonomous driving, and any field where understanding the \u0026quot;reasoning\u0026quot; behind model decisions is crucial. Its capacity to provide intuitive visual insights helps bridge the gap between complex model architectures and human interpretability, fostering trust and transparency in deep learning systems. This same method is going to help us more effectively understand KL Grading skim from inside.\u003c/p\u003e\n\u003cp\u003eFor effectively understand the and compare between all those models, my criteria were to select 4 standard image group for all 10 models: I) True Negative. II)False Negative. III)True Positive. IV) False Positive. \u0026nbsp;In respect to that from the standard set I was not able to get any common false positive prediction from the models, It made me to just use 3 common images for all 10 models, 4 common images in 6 models, 4 common images in 2 model, 4 common images on 2 models (6+2+2=10). Except in case of false positive we can visualize the impact of Grade 1 on models.\u003c/p\u003e\n\u003cp\u003eOBSERVATION FROM OBSERBER:\u003c/p\u003e\n\u003col class=\"decimal_type\"\u003e\n \u003cli\u003e9778477L: This radiograph doesn\u0026rsquo;t show any potential sign of OA. No visible JSN, edge of both joint compartment looks smooth, no visible sign of osteophytes. Assigning Grade 0. (Figure 2,8,10)\u003c/li\u003e\n \u003cli\u003e9063955R: \u0026ldquo;Definite\u0026rdquo; JSN on lateral compartment, Multiple sharp edges can be seen in both lateral and median side of tibia multiple \u0026ldquo;Definite\u0026rdquo; osteophytes. White thickening on tibia possible sclerosis. No visible sign of trauma or cartilage degeneration. Assigned Grade 2. (Figure 2,8,10)\u003c/li\u003e\n \u003cli\u003e9442105L: \u0026ldquo;Possible\u0026rdquo; JSN in Median compartment, White thickening on cartilage of tibia possible sclerosis. Smooth edges of both compartments, no possible sign of osteophytes. No visible sign of trauma or cartilage degeneration. Assigned Grade 1. (Figure 2,8,10)\u003c/li\u003e\n \u003cli\u003e9803621R: \u0026ldquo;Slight Possibility\u0026rdquo; of JSN in median compartment. Multiple sharp edges in both median and lateral compartment side of tibia, multiple definite visible osteophytes, slight white thickening on the cartilage of tibia possible sclerosis. No visible sign of trauma or cartilage degeneration. Assigned Grade 2. (Figure 2)\u003c/li\u003e\n \u003cli\u003e9082901R: \u0026ldquo;Possible\u0026rdquo; JSN in median compartment, no sclerosis observed, slight possibility of sharp edge in lateral joint compartment of tibia but no visible osteophytes. No visible sign of trauma or cartilage degeneration. Assigned Grade 0. (Figure 8)\u003c/li\u003e\n \u003cli\u003e9460287R: \u0026ldquo;Possible\u0026rdquo; JSN in median compartment, no sclerosis or sharp edge observed in any of the joint compartment edge. No visible sign of trauma or cartilage degeneration. Assigned Grade 0. (Figure 10)\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eTable 2\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"726\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 80px;\"\u003e\n \u003cp\u003eConfiguration\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eTrue Negative\u003c/p\u003e\n \u003cp\u003eIMG:9778477L ALL Models\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eTrue Positive\u003c/p\u003e\n \u003cp\u003eIMG:9063955R ALL Models\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eFalse Negative\u003c/p\u003e\n \u003cp\u003eIMG:9442105L ALL Models\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eFalse Positive\u003c/p\u003e\n \u003cp\u003eIMG:9803621R (1 to 6)\u003c/p\u003e\n \u003cp\u003eIMG:9082901R (7,8)\u003c/p\u003e\n \u003cp\u003eIMG:9460287R (9,10)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003eModel\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 80px;\"\u003e\n \u003cp\u003eGrade 0 vs. Grade (2-4) (Excluding Grade 1)\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eExcluding surrounding area of both joint compartments.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction is based partially on both compartments, mostly lateral joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003ePartially excluding the surrounding area of both joint compartments, mainly lateral\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003ePrediction is totally based on lateral joint compartment (more to femur)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e1)DenseNet121\u003c/p\u003e\n \u003cp\u003e(Figure 2)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eExcluding surrounding area of both joint compartments.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction is mainly based on lateral joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eExcluding surrounding area of both joint compartments.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003ePrediction is totally based on lateral joint compartment (most to femur)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e2)ResNet18\u003c/p\u003e\n \u003cp\u003e(Figure 3)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 80px;\"\u003e\n \u003cp\u003eGrade 0 vs. Grade (1-4) (Including Grade 1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eExcluding partially both joint compartments, primarily lateral.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction is mainly based on lateral joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003ePartially excluding lateral joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003ePrediction is totally based on edge of lateral joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e3)DenseNet121\u003c/p\u003e\n \u003cp\u003e(Figure 4)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eExcluding surrounding area of both joint compartments.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction is mainly based on lateral joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eExcluding almost everything except edge of lateral joint compartment mostly on tibia.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003ePrediction is dominantly based on edge of lateral compartment but partially focusing on medial compartment edge\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e4)ResNet18\u003c/p\u003e\n \u003cp\u003e(Figure 5)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 80px;\"\u003e\n \u003cp\u003eGrade (0-1) vs.\u003c/p\u003e\n \u003cp\u003eGrade (2-4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eJust excluding both joint compartment area.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction is only based on lateral compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eJust excluding both joint compartment area.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003ePrediction is only based on edge of lateral compartment dominantly on femur\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e5)DenseNet121\u003c/p\u003e\n \u003cp\u003e(Figure 6)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eJust excluding mostly median joint compartment and partially lateral\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction is only based on lateral compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eJust excluding both joint compartment area.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003ePrediction is only based on edge of lateral compartment mostly on femur\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e6)ResNet18\u003c/p\u003e\n \u003cp\u003e(Figure 7)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 80px;\"\u003e\n \u003cp\u003eGrade 0\u003c/p\u003e\n \u003cp\u003evs.\u003c/p\u003e\n \u003cp\u003eGrade 1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eExcluding just medial joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction is only based on lateral compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eExcluding everything except edge of medial compartment on tibia\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eExcluding almost everything except medial joint compartment dominantly on edge\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e7)ResNet18\u003c/p\u003e\n \u003cp\u003e(Figure 9)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eAlmost random slight exclusion of both joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction is almost random if not possibly full joint space\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eAlmost random if not, slightly excluding edge of lateral\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eAlmost random, If not, slightly excluding medial joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e8)DenseNet121\u003c/p\u003e\n \u003cp\u003e(Figure 10)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 80px;\"\u003e\n \u003cp\u003eGrade 1 vs. Grade 2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eExcluding medial and slightly lateral joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction on lateral joint compartment and dominantly on femur\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eJust excluding lateral compartment prediction dominantly based on median compartment edge\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003ePrediction just based on median compartment and edge of femur\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e9)ResNet18\u003c/p\u003e\n \u003cp\u003e(Figure 12)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eExcluding medial joint compartment prediction based on lateral side of tibia\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003ePrediction is based in total lateral joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eAlmost random if not, slightly excluding medial joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003ePrediction is based on edge of lateral joint compartment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003e10)DenseNet121\u003c/p\u003e\n \u003cp\u003e(Figure 13)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eThe standard test datasets were constructed with Grade 0 images serving as negatives and Grade 2-4 images as positives. Notably, for models 7 and 8, the positive portion of the datasets, as well as all corresponding positive Grad CAM images remained undefined for the model. Conversely, for models 9 and 10, the negative portion of the dataset was not provided to the model. Based on the predictions from all 10 models, several key observations emerge:\u003c/p\u003e\n\u003cp\u003eThe integration of observer findings (Table \u0026nbsp;2) with Grad CAM visualizations reveals a complex and highly informative correlation that underscores both the strengths and limitations of current KL grading methods, particularly with respect to the ambiguous Grade 1. For instance, consider radiograph \u003cstrong\u003e9778477L\u003c/strong\u003e: the expert observer identified no evidence of OA no joint space narrowing (JSN), smooth joint compartment edges, and no visible osteophytes assigning it a Grade 0. Correspondingly, Grad CAM outputs from models 1-6 consistently showed minimal or no activation in both the medial and lateral compartments, reinforcing the observer\u0026rsquo;s conclusion and validating the model\u0026rsquo;s ability to correctly identify a true negative case.\u003c/p\u003e\n\u003cp\u003eIn contrast, radiograph \u003cstrong\u003e9063955R\u003c/strong\u003e was assigned a Grade 2 by the observer due to clear JSN in the lateral compartment, multiple sharp edges on both the lateral and medial sides of the tibia, and definite osteophytes with possible sclerosis. Grad CAM maps for this image predominantly highlighted the lateral joint compartment. This lateral activation pattern is clinically significant because it aligns well with the observer\u0026rsquo;s identification of clear pathological features, thus demonstrating that the models effectively focus on the regions that are most indicative of OA.\u003c/p\u003e\n\u003cp\u003eAmbiguity arises in borderline cases such as radiograph \u003cstrong\u003e9442105L\u003c/strong\u003e. Although the OIA dataset categorizes this image as Grade 2+, the observer noted only possible JSN and subtle changes in the medial compartment without distinct osteophyte formation, leading to an assignment of Grade 1. Grad CAM outputs for this image were less definitive \u0026nbsp;some models configured to distinguish between Grade 0 and Grade 1 showed only minimal activation in the medial compartment or even near random patterns. This is particularly concerning, as the medial compartment in both 9778477L and 9442105L typically lacks robust edge specific activation. Such ambiguous activation may indicate that subtle medial changes could be early signs of pathology, or alternatively, that the inherent vagueness of Grade 1 is causing inconsistent model behavior. Furthermore, when comparing models trained to differentiate between Grade 1 and Grade 2, both ResNet and DenseNet architectures sometimes produced overlapping activation regions. Notably, some ResNet models lacked activation in the lateral compartment a region that might serve as a key differentiator between Grade 1 and Grade 2 thus suggesting that lateral compartment features may be critical for early OA detection.\u003c/p\u003e\n\u003cp\u003eAdditional insights emerge when examining training configurations. Excluding Grade 1 from the positive class leads to Grad CAM maps in both ResNet and DenseNet models that display clear, minimal activation in both joint compartments and their surrounding areas, supporting a more definitive classification of true negatives. In contrast, Grade 1 as positive often causes DenseNet models to exhibit slightly random activation patterns. Interestingly, when Grade 1 is instead included as part of the negative class, both model types are more precise in detecting features in both compartments. This observation raises the possibility that radiograph \u003cstrong\u003e9442105L\u003c/strong\u003e might, in fact, be a true negative that was misclassified by human observers a finding that further underscores the problematic nature of Grade 1.\u003c/p\u003e\n\u003cp\u003eFalse positives also provide valuable context. For example, image \u003cstrong\u003e9803621R\u003c/strong\u003e graded as 2 by the observer due to multiple definite osteophytes was initially graded 0 \u0026ldquo;Normal\u0026rdquo; in OIA dataset exhibits lateral compartment activation predominantly along the femur when Grade 1 is excluded, closely mirroring true positive patterns. However, inclusion of Grade 1 shifts this activation towards the lateral compartment edge, reinforcing the importance of lateral features in confirming OA. In the case of image \u003cstrong\u003e9082901R\u003c/strong\u003e, which was graded as 0 despite the observer noting a slight possibility of a sharp lateral edge and potential JSN, Grad CAM outputs from models trained on a Grade 0 versus Grade 1 configuration showed near random behavior with little lateral activation. Such variability further highlights the intrinsic ambiguity associated with Grade 1. Finally, image \u003cstrong\u003e9460287R\u003c/strong\u003e, assigned as Grade 0 due to the absence of definitive OA features (aside from a slight possibility of medial JSN), produced contradictory Grad CAM results when models attempted to differentiate between Grade 1 and Grade 2, again underscoring the challenge posed by Grade 1\u0026rsquo;s subtle and overlapping features.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.3) Explanation of Deep Learning Approach:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAfter evaluating multiple models and training strategies with 75 % training, 5% validation and 10% testing set, was determined that training the ResNet18 model without segmentation yielded the best results when grayscale images were replicated across three channels. Specifically, using an adaptive learning rate starting at 0.0001 with a decay schedule in each 5 epoch that reduced the rate by 30% if no improvement was observed and applying rigorous data augmentation produced optimal performance for ResNet18.\u003c/p\u003e\n\u003cp\u003eFor training with segmentation, The study developed a novel approach designed to reduce dependency on segmentation quality and mitigate misclassification issues. In this method, I created a stacked input image where the first and third channels contain the original grayscale image, and the middle channel contains the corresponding segmentation mask. This approach proved particularly effective when applied to the DenseNet121 model. With a fixed learning rate of 0.0001 and the same rigorous transformation pipeline, the DenseNet121 model achieved superior performance. Notably, while ResNet18 reached its best accuracy within approximately 100 epochs, DenseNet121 required around 400 epochs to converge but ultimately provided more reliable outcomes.\u003c/p\u003e\n\u003cp\u003eBoth models were trained using a binary classification framework. The segmentation training method is unique to this study and, to my knowledge, has not been previously reported.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eKey Libraries and tools used:\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHardware:\u003c/strong\u003e\u003c/p\u003e\n\u003cul type=\"disc\"\u003e\n \u003cli\u003eLaptop: NVIDIA RTX 3060, AMD Ryzen 7 5800H\u003c/li\u003e\n \u003cli\u003eRAM: 24 GB DDR4 3200 MHz\u003c/li\u003e\n \u003cli\u003eStorage: 100 GB disk drive (SSD)\u003c/li\u003e\n \u003cli\u003eOperating System: Windows 11\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cstrong\u003eSoftware:\u003c/strong\u003e\u003c/p\u003e\n\u003cul type=\"disc\"\u003e\n \u003cli\u003eIDE: Visual Studio Code with Jupyter Notebook extension\u003c/li\u003e\n \u003cli\u003eEnvironment: Conda, Python 3.11\u003c/li\u003e\n \u003cli\u003eKey Libraries: PyTorch, NumPy, Pandas, Matplotlib, TorchVision, Scikit learn, Pillow\u003c/li\u003e\n\u003c/ul\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eR.I.M. wrote the manuscript, conducted the research, developed the deep learning models, analyzed the results, and prepared all figures and tables. The author reviewed and approved the final manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eLAWRENCE JS, AITKEN SWAN J. Rheumatism in miners. Part I: Rheumatic complaints. \u003cem\u003eBr J Ind Med\u003c/em\u003e. 1952;9(1):1 18. doi:10.1136/oem.9.1.1\u003c/li\u003e\n \u003cli\u003eKellgren JH, Lawrence JS. Rheumatism in miners. Part II: X ray study. \u003cem\u003eOccupational and Environmental Medicine.\u003c/em\u003e 1952;9(3):197 207. doi:10.1136/oem.9.3.197\u003c/li\u003e\n \u003cli\u003eBraun HJ, Gold GE. Diagnosis of osteoarthritis: imaging. \u003cem\u003eBone\u003c/em\u003e. 2012;51(2):278 288. doi:10.1016/j.bone.2011.11.019\u003c/li\u003e\n \u003cli\u003eKellgren JH, Lawrence JS. Radiological assessment of osteo arthrosis. \u003cem\u003eAnn Rheum Dis.\u003c/em\u003e 1957;16(4):494 502. doi:10.1136/ard.16.4.494\u003c/li\u003e\n \u003cli\u003eFelson DT, Naimark A, Anderson J, Kazis L, Castelli W, Meenan RF. The prevalence of knee osteoarthritis in the elderly. The Framingham Osteoarthritis Study. \u003cem\u003eArthritis Rheum\u003c/em\u003e. 1987;30(8):914 918. doi:10.1002/art.1780300811\u003c/li\u003e\n \u003cli\u003eScott WW Jr, Lethbridge Cejku M, Reichle R, Wigley FM, Tobin JD, Hochberg MC. Reliability of grading scales for individual radiographic features of osteoarthritis of the knee. The Baltimore longitudinal study of aging atlas of knee osteoarthritis. \u003cem\u003eInvest Radiol\u003c/em\u003e. 1993;28(6):497 501.\u003c/li\u003e\n \u003cli\u003eArnett FC, Edworthy SM, Bloch DA, et al. The American Rheumatism Association 1987 revised criteria for the classification of rheumatoid arthritis. \u003cem\u003eArthritis Rheum\u003c/em\u003e. 1988;31(3):315 324. doi:10.1002/art.1780310302\u003c/li\u003e\n \u003cli\u003eAltman R, Asch E, Bloch D, et al. Development of criteria for the classification and reporting of osteoarthritis. Classification of osteoarthritis of the knee. Diagnostic and Therapeutic Criteria Committee of the American Rheumatism Association. \u003cem\u003eArthritis Rheum\u003c/em\u003e. 1986;29(8):1039 1049. doi:10.1002/art.1780290816\u003c/li\u003e\n \u003cli\u003eIrrgang JJ, Anderson AF, Boland AL, et al. Responsiveness of the International Knee Documentation Committee Subjective Knee Form. \u003cem\u003eAm J Sports Med\u003c/em\u003e. 2006;34(10):1567 1573. doi:10.1177/0363546506288855\u003c/li\u003e\n \u003cli\u003eHefti F, M\u0026uuml;ller W, Jakob RP, St\u0026auml;ubli HU. Evaluation of knee ligament injuries with the IKDC form. \u003cem\u003eKnee Surg Sports Traumatol Arthrosc\u003c/em\u003e. 1993;1(3 4):226 234. doi:10.1007/BF01560215\u003c/li\u003e\n \u003cli\u003eFAIRBANK TJ. Knee joint changes after meniscectomy. \u003cem\u003eJ Bone Joint Surg Br\u003c/em\u003e. 1948;30B(4):664 670.\u003c/li\u003e\n \u003cli\u003eTapper EM, Hoover NW. Late results after meniscectomy. \u003cem\u003eJ Bone Joint Surg Am\u003c/em\u003e. 1969;51(3):.\u003c/li\u003e\n \u003cli\u003eBrandt KD, Fife RS, Braunstein EM, Katz B. Radiographic grading of the severity of knee osteoarthritis: relation of the Kellgren and Lawrence grade to a grade based on joint space narrowing, and correlation with arthroscopic evidence of articular cartilage degeneration. \u003cem\u003eArthritis Rheum\u003c/em\u003e. 1991;34(11):1381 1386. doi:10.1002/art.1780341106\u003c/li\u003e\n \u003cli\u003eGalli M, De Santis V, Tafuro L. Reliability of the Ahlb\u0026auml;ck classification of knee osteoarthritis. \u003cem\u003eOsteoarthritis Cartilage\u003c/em\u003e. 2003;11(8):580 584. doi:10.1016/s1063 4584(03)00095 5\u003c/li\u003e\n \u003cli\u003eScheller G, Sobau C, B\u0026uuml;low JU. Arthroscopic partial lateral meniscectomy in an otherwise normal knee: Clinical, functional, and radiographic results of a long term follow up study. \u003cem\u003eArthroscopy\u003c/em\u003e. 2001;17(9):946 952. doi:10.1053/jars.2001.28952\u003c/li\u003e\n \u003cli\u003eSchroeder Boersch H, T\u0026ouml;ws P, Jani L. Reproduzierbarkeit von radiologischen Arthrosemerkmalen. Evaluierung eines Scores zur radiologischen Klassifikation von degenerativen Ver\u0026auml;nderungen bei Gonarthrose [Reproducibility of radiologic markers of osteoarthritis. Evaluating a score for radiologic classification of degenerative changes in osteoarthritis of the knee joint]. \u003cem\u003eZ Orthop Ihre Grenzgeb\u003c/em\u003e. 1998;136(4):293 297. doi:10.1055/s 2008 1053740\u003c/li\u003e\n \u003cli\u003eWright RW; MARS Group. Osteoarthritis Classification Scales: Interobserver Reliability and Arthroscopic Correlation. \u003cem\u003eJ Bone Joint Surg Am\u003c/em\u003e. 2014;96(14):1145 1151. doi:10.2106/JBJS.M.00929\u003c/li\u003e\n \u003cli\u003eHayes B, Kittelson A, Loyd B, Wellsandt E, Flug J, Stevens Lapsley J. Assessing radiographic knee osteoarthritis: an online training tutorial for the Kellgren Lawrence grading scale. \u003cem\u003eMedEdPORTAL.\u003c/em\u003e 2016;12:10503. doi:10.15766/mep_2374 8265.10503\u003c/li\u003e\n \u003cli\u003eBagge E, Bjelle A, Valkenburg HA, Svanborg A. Prevalence of radiographic osteoarthritis in two elderly European populations. \u003cem\u003eRheumatol Int\u003c/em\u003e. 1992;12(1):33 38. doi:10.1007/BF00246874\u003c/li\u003e\n \u003cli\u003eSpector TD, Cooper C. Radiographic assessment of osteoarthritis in population studies: whither Kellgren and Lawrence?. \u003cem\u003eOsteoarthritis Cartilage\u003c/em\u003e. 1993;1(4):203 206. doi:10.1016/s1063 4584(05)80325 5\u003c/li\u003e\n \u003cli\u003eEckstein F, Kwoh CK, Link TM. Imaging research results from the Osteoarthritis Initiative (OAI): a review and lessons learned 10 years after start of enrolment. \u003cem\u003eAnn Rheum Dis.\u003c/em\u003e 2014;73(7):1289 1300. doi:10.1136/annrheumdis 2014 205310\u003c/li\u003e\n \u003cli\u003eChen P. Knee osteoarthritis severity grading dataset. \u003cem\u003eMendeley Data.\u003c/em\u003e 2018;V1. doi:10.17632/56rmx5bjcr.1\u003c/li\u003e\n \u003cli\u003eHe K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. \u003cem\u003e2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).\u003c/em\u003e 2016:770 778. doi:10.1109/CVPR.2016.90\u003c/li\u003e\n \u003cli\u003eHuang G, Liu Z, Van Der Maaten L, Weinberger KQ. Densely connected convolutional networks. \u003cem\u003e2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).\u003c/em\u003e 2017:2261 2269. doi:10.1109/CVPR.2017.243\u003c/li\u003e\n \u003cli\u003eTan M, Le QV. EfficientNet: rethinking model scaling for convolutional neural networks. \u003cem\u003eInt Conf Mach Learn.\u003c/em\u003e 2019. doi:10.48550/arXiv.1905.11946\u003c/li\u003e\n \u003cli\u003eKELLGREN JH. Radiological signs of rheumatoid arthritis; a study of observer differences in the reading of hand films. \u003cem\u003eAnn Rheum Dis\u003c/em\u003e. 1956;15(1):55 60. doi:10.1136/ard.15.1.55\u003c/li\u003e\n \u003cli\u003eSelvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad CAM: Visual explanations from deep networks via gradient based localization. \u003cem\u003eInt J Comput Vis.\u003c/em\u003e 2019;128(2):336 359. doi:10.1007/s11263 019 01228 7\u003c/li\u003e\n \u003cli\u003eTiulpin A, Thevenot J, Rahtu E, Lehenkari P, Saarakkala S. Automatic knee osteoarthritis diagnosis from plain radiographs: a deep learning based approach. \u003cem\u003eSci Rep.\u003c/em\u003e 2018;8(1):1727. doi:10.1038/s41598 018 20132 7\u003c/li\u003e\n \u003cli\u003eShamir L, Ling SM, Scott W, Hochberg M, Ferrucci L, Goldberg IG. Early detection of radiographic knee osteoarthritis using computer aided analysis. \u003cem\u003eOsteoarthritis Cartilage.\u003c/em\u003e 2009;17(10):1307 1312. doi:10.1016/j.joca.2009.04.010\u003c/li\u003e\n \u003cli\u003eYeoh PSQ, Lai KW, Goh SL, Hasikin K, Hum YC, Tee YK, Dhanalakshmi S. Emergence of deep learning in knee osteoarthritis diagnosis. \u003cem\u003eComputational Intelligence and Neuroscience.\u003c/em\u003e 2021;2021(1):4931437. doi:10.1155/2021/4931437\u003c/li\u003e\n \u003cli\u003eKijowski R, Fritz J, Deniz CM. Deep learning applications in osteoarthritis imaging. \u003cem\u003eSkeletal Radiol\u003c/em\u003e. 2023;52(11):2225 2238. doi:10.1007/s00256 023 04296 6\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-5983981/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5983981/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe Kellgren-Lawrence (KL) grading system, established by Lawrence and Aitken Swan [1] and Kellgren and Lawrence [2] in the early 1950s, remains the clinical standard for radiographic osteoarthritis (OA) classification despite well‐documented inter‐ and intra‐observer variability [3-6]. Early studies revealed that occupational factors, particularly in coal miners, are strongly associated with knee OA [1,2]. However, subsequent research has highlighted that, over reliance on osteophytes, ambiguous criterion especially for KL Grade 1 (doubtful OA) lead to diagnostic uncertainty [3-6]. Comparative evaluations of multiple radiographic scoring systems based on joint space narrowing (JSN) and osteophyte formation [7-10] have further underscored limited sensitivity of the KL system to early cartilage damage and inconsistent reproducibility [11-16].\u003c/p\u003e","manuscriptTitle":"Redefining Early Osteoarthritis Detection: Deep Learning Meets Historical Radiography","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-02-14 10:22:21","doi":"10.21203/rs.3.rs-5983981/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2fb63f2d-7ac2-467f-a3f2-7c037a55f7c6","owner":[],"postedDate":"February 14th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":44254478,"name":"Health sciences/Biomarkers/Diagnostic markers"},{"id":44254479,"name":"Health sciences/Health care/Medical imaging/Radiography"},{"id":44254480,"name":"Health sciences/Diseases/Rheumatic diseases/Osteoarthritis"}],"tags":[],"updatedAt":"2025-02-14T10:22:22+00:00","versionOfRecord":[],"versionCreatedAt":"2025-02-14 10:22:21","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5983981","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5983981","identity":"rs-5983981","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.