Evaluation of the Reproducibility and Consistency of the RUST Radiographic Scale in Tibial Fracture Healing: Comparison Between Orthopedic Residents and Artificial Intelligence

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-04 · read from full text

This retrospective observational study evaluated the reproducibility and consistency of the Radiographic Union Score for Tibial fractures (RUST) in 47 skeletally mature patients with tibial shaft fractures treated with intramedullary nailing, using radiographs assessed by four orthopedic residents (R1–R4) at two blinded time points and by a trained AI algorithm using RUST parameters. Residents showed excellent interobserver reproducibility (ICC = 0.93) and very good intraobserver consistency (global ICC = 0.72), but agreement with AI was limited, with only 17% absolute agreement and a tendency for residents to overestimate AI scores (45.2%) more often than underestimate them (37.8%). The paper reports that adding AI to the interobserver analysis reduced agreement (ICC = 0.61), and one resident (R2) differed significantly from AI (p = 0.009). This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Purpose To evaluate the reproducibility and consistency of the Radiographic Union Score for Tibial fractures (RUST) in diaphyseal fractures treated with intramedullary nailing among orthopedic residents with different levels of experience, and to compare their performance with artificial intelligence (AI) analysis. Methods Radiographs from 47 patients were assessed by four orthopedic residents (R1–R4) at two independent time points using the RUST scale. The same images were evaluated by AI. Intra- and interobserver agreement were calculated using the Intraclass Correlation Coefficient (ICC). Differences between human and AI evaluations were analyzed using Wilcoxon and Friedman tests with a 5% significance level. Results The sample consisted predominantly of men (71.7%) with a mean age of 32.9 years. AI classified 61.7% of fractures as consolidated (score ≥ 7). Absolute agreement between residents and AI was 17%, with residents overestimating scores in 45.2% and underestimating in 37.8% of cases. Observer R3 showed the best agreement with AI (ICC = 0.58), while R2 demonstrated a significant difference (p = 0.009). Interobserver reproducibility among residents was excellent (ICC = 0.93; 95% CI 0.84–0.96) but decreased to 0.61 when AI was included. Intraobserver consistency was very good (global ICC = 0.72; 95% CI 0.63–0.78), with R4 presenting the highest stability (ICC = 0.78). Conclusion The RUST scale showed high reproducibility among residents but limited agreement with AI, suggesting the need for additional training to improve consistency between human interpretation and algorithmic assessment.
Full text 86,997 characters · extracted from preprint-html · click to expand
Evaluation of the Reproducibility and Consistency of the RUST Radiographic Scale in Tibial Fracture Healing: Comparison Between Orthopedic Residents and Artificial Intelligence | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Evaluation of the Reproducibility and Consistency of the RUST Radiographic Scale in Tibial Fracture Healing: Comparison Between Orthopedic Residents and Artificial Intelligence Pedro José Labronici, Victoria Maria Silva Iatarola, Newton Luiz Lombardi Fonseca Silveira, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8886056/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 11 You are reading this latest preprint version Abstract Purpose To evaluate the reproducibility and consistency of the Radiographic Union Score for Tibial fractures (RUST) in diaphyseal fractures treated with intramedullary nailing among orthopedic residents with different levels of experience, and to compare their performance with artificial intelligence (AI) analysis. Methods Radiographs from 47 patients were assessed by four orthopedic residents (R1–R4) at two independent time points using the RUST scale. The same images were evaluated by AI. Intra- and interobserver agreement were calculated using the Intraclass Correlation Coefficient (ICC). Differences between human and AI evaluations were analyzed using Wilcoxon and Friedman tests with a 5% significance level. Results The sample consisted predominantly of men (71.7%) with a mean age of 32.9 years. AI classified 61.7% of fractures as consolidated (score ≥ 7). Absolute agreement between residents and AI was 17%, with residents overestimating scores in 45.2% and underestimating in 37.8% of cases. Observer R3 showed the best agreement with AI (ICC = 0.58), while R2 demonstrated a significant difference (p = 0.009). Interobserver reproducibility among residents was excellent (ICC = 0.93; 95% CI 0.84–0.96) but decreased to 0.61 when AI was included. Intraobserver consistency was very good (global ICC = 0.72; 95% CI 0.63–0.78), with R4 presenting the highest stability (ICC = 0.78). Conclusion The RUST scale showed high reproducibility among residents but limited agreement with AI, suggesting the need for additional training to improve consistency between human interpretation and algorithmic assessment. tibial fracture fracture healing radiographic union score reliability validity Figures Figure 1 Figure 2 INTRODUCTION Tibial shaft fracture is the most common among long bone fractures, showing a high incidence and predominantly affecting young males in their economically productive years. [ 1 – 6 ] Although locked intramedullary nailing is widely accepted as the standard treatment for tibial shaft fractures, there remains no consensus among orthopedic surgeons regarding the radiographic assessment of bone healing.[ 7 , 8 ] The presence of periosteal callus formation visible on radiographs, combined with the absence of a fracture line, constitutes the most reliable radiographic sign of bone healing among different observers.[ 5 , 9 ] Whelan et al.[ 10 ] developed the Radiographic Union Scale for Tibial fractures (RUST), based on specific parameters that assess the progression of bone healing after intramedullary fixation using a locked nail. Several studies have demonstrated substantial intra- and interobserver agreement for the RUST scale. [ 11 – 19 ] Recently, with the rise of Artificial Intelligence (AI), multiple studies have investigated its application in evaluating different long bone fractures.[ 20 , 21 ] However, to the best of our knowledge, no study has assessed the use of AI in determining radiographic evidence of bone union using the RUST score. Our hypothesis is that AI may enhance the ability of orthopedic surgeons to identify bone healing in tibial shaft fractures treated with intramedullary nailing using the RUST scale. The primary objective of this study is to evaluate the performance of an AI tool trained with the RUST scoring system in diaphyseal fractures treated with intramedullary nails and to compare its results with those of orthopedic residents. The secondary objective is to assess intra- and interobserver agreement among residents at different levels of training. MATERIALS AND METHODS This retrospective observational study included 47 skeletally mature patients with tibial shaft fractures treated with intramedullary nailing at a tertiary trauma center. Ethical approval was obtained from the institutional research ethics committee. Patients were included if postoperative follow-up radiographs were available. Exclusion criteria comprised pathological fractures, infection, and inadequate image quality. Four orthopedic residents (R1–R4) independently evaluated all radiographs using the RUST scale at two separate time points four weeks apart. Observers were blinded to clinical data and to each other’s assessments. Each cortex (anterior, posterior, medial, and lateral) received a score from 1 to 3 according to the RUST criteria: 1 — visible fracture line without callus 2 — callus with persistent fracture line 3 — bridging callus without fracture line The sum generated a total score ranging from 4 to 12, with scores ≥ 7 considered radiographic consolidation. The same radiographs were analyzed by an artificial intelligence algorithm trained using RUST parameters described in previous validation studies Statistical Analysis Intra- and interobserver reliability were assessed using the Intraclass Correlation Coefficient (ICC). Agreement between residents and AI was analyzed using Wilcoxon and Friedman tests. Significance was set at p < 0.05. RESULTS Forty-seven patients were included (71.7% male), with a mean age of 32.9 years. The median age was 29 years, the standard deviation was 11.2 years, and the coefficient of variation was 0.34, indicating high age variability. This heterogeneity was intentional, to ensure greater robustness in the concordance analysis. Table 1 presents the frequency distribution and statistical data for the total RUST scores of the 47 patients, obtained by both AI and the four evaluators. According to the AI classifications, most patients had total RUST scores between 7 and 9 (53.2%), with 61.7% of cases considered consolidated (scores ≥ 7). The frequency distributions of RUST total scores calculated by the evaluators showed different patterns. Table 1 presents the frequency distribution and statistical data for the total RUST scores of the 47 patients, obtained by both AI and the four evaluators. According to the AI. Evaluator RUST Total Score Range Frequency (n) % Median Mean Standard Deviation p-value of the Wilcoxon test comparing with AI IA 4 a 6 18 38,3% 8,0 7,4 1,7 7 a 9 25 53,2% - 10 a 12 4 8,5% R4 4 a 6 22 46,8% 7,0 6,8 2,3 0,076 7 a 9 18 38,3% 10 a 12 6 12,8% R3 4 a 6 10 21,3% 8,0 7,8 1,9 0,190 7 a 9 27 57,4% 10 a 12 10 21,3% R2 4 a 6 2 4,3% 8,0 8,2 1,6 0,009 7 a 9 36 76,6% 10 a 12 9 19,1% R1 4 a 6 12 25,5% 7,5 7,5 1,8 0,795 7 a 9 28 59,6% 10 a 12 7 14,9% Table 2 presents the agreement analysis between the scores obtained by AI and the mean scores assigned by each evaluator, as well as the overall agreement (independent of evaluator). Considering the overall results, regardless of the evaluator, only 17.0% of assessments matched the AI score, while 37.8% of evaluations underestimated and 45.2% overestimated the AI score. The evaluator who showed the highest concordance with AI was R3, whose scores coincided with AI in 21.3% of cases and achieved the highest ICC = 0.58. The lowest performances were observed for R4, who matched AI in only 10.6% of evaluations, and R2, who presented a statistically significant difference compared with AI scores (p = 0.009 in the Wilcoxon test; ICC = 0.44). Table 2 Agreement analysis between the scores obtained by AI and the mean scores of the evaluators, for each evaluator and overall. Comparison Result with AI Evaluator R4 R3 R2 R1 Global Absolute Agreement with AI 10,6% 21,3% 17,0% 19,1% 17,0% Evaluator Underestimated AI Score 59,6% 29,8% 23,4% 38,3% 37,8% Evaluator Overestimated AI Score 29,8% 48,9% 59,6% 42,6% 45,2% Wilcoxon Test p-value 0,076 0,190 0,009 0,795 0,355 ICC 0,47 0,58 0,44 0,45 0,49 95% CI of ICC 0,07 − 0,70 0,26 − 0,77 0,03 − 0,68 0,00–0,69 0,31 − 0,61 Agreement Classification Good Good Good Good Good Significance of Agreement Not confirmed Confirmed with reservation Not confirmed Not confirmed Confirmed with reservation Table 3 shows the interobserver reproducibility of the total RUST score. The agreement among the four evaluators (R1–R4) was excellent (ICC = 0.93; 95% CI: 0.84–0.96). When the AI was included in the interobserver analysis, the ICC decreased to 0.61, still classified as very good, indicating significant agreement, but with a greater dispersion related to the AI’s stricter and standardized criteria. Table 3 Interobserver reproducibility analysis (total RUST score). Evaluators Compared ICC 95% CI of ICC Agreement Classification Significance of Agreement Friedman Test p-value R4, R3, R2 and R1 0,93 (0,84;0,96) Excellent Confirmed < 0,001 IA, R4, R3,R2 and R1 0,61 (0,47; 0,73) Very Good Confirmed < 0,001 Table 4 displays the intraobserver reproducibility, evaluating each resident’s consistency across both assessment rounds and the learning stability of the proposed method. The global intraobserver ICC was 0.72 (95% CI: 0.63–0.78), demonstrating very good and statistically significant agreement. Evaluator R4 showed the highest intraobserver stability (ICC = 0.78), likely due to greater caution and rigor in applying scoring criteria, although this did not translate into higher concordance with AI. All evaluators exhibited “very good” intraobserver reliability, confirming that once the logic of the RUST scale is understood, consistent individual interpretation is achieved. Table 4 Intraobserver reproducibility analysis (learning stability of evaluators). Statistics Evaluator R4 R3 R2 R1 Overall ICC 0,78 0,68 0,65 0,67 0,72 95% CI of ICC (0,63; 0,87) (0,46; 0,82) (0,42; 0,81) (0,45; 0,81) (0,63; 0,78) Agreement Classification Very Good Very Good Very Good Very Good Very Good Significance of Agreement Confirmed Confirmed Confirmed Confirmed Confirmed DISCUSSION The radiographic assessment of tibial shaft fracture healing remains a challenge due to the variability of criteria used to define bone union. In this study, the application of the RUST scale allowed for a more objective and standardized analysis of bone healing progression in fractures treated with intramedullary nailing. The scoring method demonstrated high reproducibility, with agreement indices ranging from very good to excellent among residents from the first to fourth year of training. However, when compared with artificial intelligence (AI), the results showed low concordance, with only 17% absolute agreement, highlighting the need for further training to align human interpretation with algorithmic patterns. Several radiographic scoring systems have been proposed to evaluate fracture healing, but their use remains limited. The Hammer [ 9 ] scale, which classifies fractures into five groups according to callus formation and fracture line obliteration, showed low correlation with mechanical stability and only moderate interobserver agreement (k ≈ 0.60). The Tower et al. [ 12 ] scale, ranging from 0 to 10 based on callus presence and radiolucent lines, demonstrated good correlation with a combined clinical scale, although its validity as an isolated radiographic method could not be confirmed. Based on the reliable assessment of cortical bridging and the frequently used criterion of fracture line visibility, it was hypothesized that the RUST scale represents a more robust and reliable tool than conventional assessment methods. [ 5 , 13 ] Previous studies have demonstrated good to excellent interobserver reliability of the RUST scale (ICC ranging from 0.84 to 0.95). Reproducibility proved consistent among orthopaedic surgeons and radiologists, and even medical students achieved similar indices, albeit with greater intraobserver variability. Moreover, evaluator experience has been shown to increase reliability, being higher among trauma specialists compared to residents.[ 2 , 14 , 15 ] Azevedo Filho et al. reported an ICC of 0.87 (95% CI: 0.81–0.91). The authors noted that as evaluator experience increased, so did reliability: higher ICC values were observed among trauma surgeons compared to first, second, and third-year residents (ICC 0.94; 0.80; 0.92; and 0.90, respectively).[ 16 ] The heterogeneity of the criteria used to define bone healing across different studies compromises the direct comparability of results. The use of more objective and standardized parameters, such as those proposed by the RUST scale, has the potential to provide greater uniformity, reproducibility, and consistency in information related to bone healing time.[ 17 ] In this study, the sample consisted predominantly of male individuals (71.7%), aged between 17 and 60 years. This age diversity, characterized by high variability, was intentionally included to confer greater robustness to the agreement analyses. Our study confirms the findings reported by Azevedo et al. [ 16 ] when evaluating fracture healing among first- to fourth-year residents using the RUST scoring method in 47 cases. The overall agreement among the four evaluators, without discriminating specific measures, showed a value of 0.82 (95% CI: 0.75–0.87), confirming the significant reproducibility of the proposed classification method, with agreement classified as excellent when comparing the mean results of the two evaluation rounds. In the interobserver reproducibility analysis, the ICC among the four evaluators was 0.93 (95% CI: 0.84–0.96), indicating excellent agreement regardless of the residents’ training level (Fig. 1 ). This finding may be attributed to the robustness of the RUST score, whose simplicity and objectivity reduce subjectivity and promote uniformity of evaluations, even among observers with different levels of experience. However, agreement with artificial intelligence (AI) proved limited: when AI was included in the interobserver analysis, the ICC dropped from 0.93 to 0.61 (Fig. 2 ). This result suggests that AI applies stricter and more uniform criteria, exposing differences in human judgment. On the other hand, it highlights the potential of AI as a training support tool, contributing to evaluator calibration and promoting greater standardization in the application of the RUST score. Although all residents showed “good” agreement with AI, only for evaluator R3 was this agreement statistically significant, albeit with some limitations. These findings suggest that, when considering AI as a reference standard, residents require additional training to improve the application of the method and bring their interpretations closer to the algorithmic pattern. The frequency of overestimation (45.2%) and underestimation (37.8%) in comparison with AI suggests that residents have not yet fully mastered the criteria of each score (1 to 3 per cortex), leading to less accurate interpretations in intermediate cases. The tendency toward overestimation, in particular, may reflect an attempt to anticipate signs of clinical consolidation rather than strictly adhering to the established radiographic parameters. Our findings are consistent with those of Pettersson et al.[ 20 ], who reported that artificial intelligence, based systems show excellent internal performance (AUC/ICC > 0.8) and tend to yield comparable, though still divergent, results from human evaluations. This discrepancy reflects differences between subjective clinical judgment and the standardized criteria applied by AI. Like the Swedish study on elbow fractures, our findings reinforce the notion that AI should be viewed as a complementary tool, capable of enhancing human evaluator calibration, increasing standardization, and potentially reducing diagnostic errors. The convergence between both studies points to methodological consistency in the application of AI in orthopaedics, suggesting that algorithms trained with large volumes of labeled data following standardized criteria (such as AO/OTA or RUST) could, in the future, be safely integrated into clinical decision making, making and medical education workflows.[ 21 ] This study has limitations, including retrospective design, absence of clinical correlation with weight-bearing progression, and single-center population. Nevertheless, the controlled radiographic evaluation allowed consistent comparison between human observers and AI assessment. Conclusion The RUST score is highly reproducible among orthopedic residents but shows limited agreement with AI analysis. Additional training and integration of automated tools may improve standardization of fracture healing assessment. Declarations Ethical approval This study was approved by the institutional research ethics committee and conducted in accordance with the Declaration of Helsinki. Informed consent The requirement for informed consent was waived due to the retrospective design of the study. Conflict of interest The authors declare that they have no conflict of interest. Funding The authors received no financial support for the research, authorship, and/or publication of this article. Author Contribution Conceptualization: PJIL, VMSI, GWMethodology: PJIL, WDB, REPData collection: NLLFS, IMR, AF, MBMData analysis: GWWriting – original draft: PJIL, VMSI, WDB, REP, VGWriting – review & editing: All authorsSupervision: PJL References Court-Brown CM, Rimmer S, Prakash U, McQueen MM (1998) The epidemiology of open long bone fractures. Injury 29(7):529–342 Kooistra BW, Dijkman BG, Busse JW, Sprague S, Schemitsch EH, Bhandari M (2010) The radiographic union scale in tibial fractures: reliability and validity. J Orthop Trauma. ;24Suppl. 3:S81–6 Chua W, Murphy D, Siow W, Kagda F, Thambiah J (2012) Epidemiological analysis of outcomes in 323 open tibial diaphyseal fractures: a nine-year experience. Singap Med J 53(6):385–389 Kojima KE, Ferreira RV (2011) Fraturas da diáfise da tíbia. Rev Bras Ortop 46(2):130–155 Whelan DB, Bhandari M, McKee MD, Guyatt GH, Kreder HJ, Stephen D et al (2002) Interobserver and intraobserver variation in the assessment of the healing of tibial fractures after intramedullary fixation. J Bone Joint Surg Br 84:15–86 Zeckey C, Mommsen P, Andruszkow H, Macke C, Frink M, Stübig T et al (2011) The aseptic femoral and tibial shaft non-unionin healthy patients – an analysis of the health-related quality of life and the socioeconomic outcome. Open Orthop J 5:193–197 Bhandari M, Guyatt GH, Swiontkowski MF et al (2002) A lack of consensus in the assessment of fracture healing among orthopaedic surgeons. J Orthop Trauma 16:562–566 Kooistra BW, Dijkman BG, Busse JW et al (2010) The radiographic union scale in tibial fractures: reliability and validity. J Orthop Trauma 24:S81–S86 Hammer RR, Hammerby S, Lindholm B (1985) Accuracy of radiologic assessment of tibial shaft fracture union in humans. Clin Orthop Relat Res. ;(199):233–238s Whelan DB, Bhandari M, Stephen D et al (2010) Development of the radiographic union score for tibial fractures for the assessment of tibial fracture healing after intramedullary fixation. J Trauma 68:629–632 Leow JM, Clement ND, Tawonsawatruk T, Simpson CJ, Simpson AHRW (2016) The radiographic union scale in tibial (RUST) fractures: reliability of the outcome measure at an independent center. Bone joint Res 5(4):116–121 Tower SS, Beals RK, Duwelius PJ (1993) Resonant frequency analysis of the tibia as a measure of fracture healing. J Orthop Trauma 7:552–557 Morshed S, Corrales L, Genant H et al (2008) Outcome assessment in clinical trials of fracture-healing. J Bone Joint Surg Am 90:62–67 Ali S, Singh A, Agarwal A, Parihar A, Mahdi AA, Srivastava RN (2014) Reliability of the RUST score for the assessment of union in simple diaphyseal tibial fractures. Int J Biomed Res 5(5):333–335 Mane SS, Yamajala SNSMLV, Mitnala SRP (2024) Evaluating reliability of the RUST score in diaphyseal Tibia fractures: A collaborative assessment by orthopaedic surgeons and radiologists. J Orthop Rep 3(4):100325 Azevedo Filho FAS, Cotias RB, Azi ML, Teixeira AAA (2017) Reliability of the radiographic union scale in tibial fractures (RUST). Revista Brasileira de Ortop (English Edition) 52(1):35–3916 Chloros GD, Howard A, Giordano V, Giannoudis PV (2020) Radiographic Long Bone Fracture Healing Scores: Can they predict non-union? Injury 51(8):1693–1695 Wojahn RD, Bechtold D, Abraamyan T, Spraggs-Hughes A, Gardner MJ, Ricci WM, McAndrew CM (2022) Progression of tibia fracture healing using RUST: are early radiographs helpful? J Orthop Trauma 36(1):e6–e11 Misir A, Uzun E, Kizkapan TB, Yildiz KI, Onder M, Ozcamdalli M (2020) Reliability of RUST and Modified RUST Scores for the Evaluation of Union in Humeral Shaft Fractures Treated with Different Techniques. Indian J Orthop 54(Suppl 1):121–126 Pettersson A, Axenhus M, Stukan T et al (2025) Use of artificial intelligence for classification of fractures around the elbow in adults according to the 2018 AO/OTA classification system. BMC Musculoskelet Disord 26(1):848 Sevinç Hüseyin, Fatih et al (2025) Detection and classification of femoral neck fractures from plain pelvic X-rays using deep learning and machine learning methods. Ulus Travma Acil Cerrahi Derg 31(8):783–788 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 30 Apr, 2026 Reviews received at journal 29 Apr, 2026 Reviewers agreed at journal 28 Apr, 2026 Reviews received at journal 21 Apr, 2026 Reviewers agreed at journal 18 Apr, 2026 Reviews received at journal 01 Mar, 2026 Reviewers agreed at journal 01 Mar, 2026 Reviewers invited by journal 24 Feb, 2026 Editor assigned by journal 17 Feb, 2026 Submission checks completed at journal 17 Feb, 2026 First submitted to journal 15 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8886056","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":596042267,"identity":"1ebed22f-f4e4-48c3-b9a3-ce1fa79ea30a","order_by":0,"name":"Pedro José Labronici","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA8UlEQVRIiWNgGAWjYBACPgYeECXBwAYmGWyAFGPjAXxa2NC0pIG0NBCjBQKAWg6DGfi1sJ899uHjDgt5Punmg7crKs7brW0/DLSlxiYapxaevOSZM89IGLbJHEu2PHPmdvK2M4lALcfSchtwOizHmJm3TSKBTSLHTLKx7Xay2QGgFsaGw7i18L+Bacn/BtRyLtns/EMCWiQQtrABtRywM7tByBaJd8mMM9vAfjG2bDiTnGB2A2hLAh6/8PPnHmb42FYnLz+7+eHNhgo7e7Pz6Q8ffKixwakFASQgVCJYZQJB5Uha7IlSPApGwSgYBSMKAAC93FnMT7o3DwAAAABJRU5ErkJggg==","orcid":"","institution":"Fluminense Federal University","correspondingAuthor":true,"prefix":"","firstName":"Pedro","middleName":"José","lastName":"Labronici","suffix":""},{"id":596042268,"identity":"ddb29f4d-4b30-4aea-b695-957accfde2b3","order_by":1,"name":"Victoria Maria Silva Iatarola","email":"","orcid":"","institution":"Santa Teresa Hospital","correspondingAuthor":false,"prefix":"","firstName":"Victoria","middleName":"Maria Silva","lastName":"Iatarola","suffix":""},{"id":596042269,"identity":"4aa35b56-90e0-4de3-8254-ea5c1a9c8832","order_by":2,"name":"Newton Luiz Lombardi Fonseca Silveira","email":"","orcid":"","institution":"Santa Teresa Hospital","correspondingAuthor":false,"prefix":"","firstName":"Newton","middleName":"Luiz Lombardi Fonseca","lastName":"Silveira","suffix":""},{"id":596042270,"identity":"95f47fd2-669d-4e1b-8c8b-e1c26a5118ac","order_by":3,"name":"Isabela de Miranda Rosa","email":"","orcid":"","institution":"Fluminense Federal University","correspondingAuthor":false,"prefix":"","firstName":"Isabela","middleName":"de Miranda","lastName":"Rosa","suffix":""},{"id":596042271,"identity":"1bc3f957-f084-4c82-a417-bea95aae6e0e","order_by":4,"name":"William Dias Belangero","email":"","orcid":"","institution":"State University of Campinas","correspondingAuthor":false,"prefix":"","firstName":"William","middleName":"Dias","lastName":"Belangero","suffix":""},{"id":596042272,"identity":"25dead94-e59e-44fb-ad2a-3e5655a5cae4","order_by":5,"name":"Gustavo Waldolato","email":"","orcid":"","institution":"Faculty of Medical Sciences of Minas Gerais","correspondingAuthor":false,"prefix":"","firstName":"Gustavo","middleName":"","lastName":"Waldolato","suffix":""},{"id":596042273,"identity":"8e681277-0426-4ae8-b4e2-ffb57b110789","order_by":6,"name":"Robinson Esteves Pires","email":"","orcid":"","institution":"Federal University of Minas Gerais","correspondingAuthor":false,"prefix":"","firstName":"Robinson","middleName":"Esteves","lastName":"Pires","suffix":""},{"id":596042274,"identity":"cf022e30-7d26-4533-bfe1-ef7e31958429","order_by":7,"name":"Anderson Freitas","email":"","orcid":"","institution":"Orthopaedic and Specialized Medicine Hospital","correspondingAuthor":false,"prefix":"","firstName":"Anderson","middleName":"","lastName":"Freitas","suffix":""},{"id":596042275,"identity":"834feab2-4d23-4d93-81d7-b761ad2c8c86","order_by":8,"name":"Marcelo Bezerra Mathias","email":"","orcid":"","institution":"Fluminense Federal University","correspondingAuthor":false,"prefix":"","firstName":"Marcelo","middleName":"Bezerra","lastName":"Mathias","suffix":""},{"id":596042276,"identity":"2704b180-53c3-4404-af73-604dac64f95a","order_by":9,"name":"Vicenzo Giordano","email":"","orcid":"","institution":"Miguel Couto Municipal Hospital","correspondingAuthor":false,"prefix":"","firstName":"Vicenzo","middleName":"","lastName":"Giordano","suffix":""}],"badges":[],"createdAt":"2026-02-15 12:53:23","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8886056/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8886056/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":103577635,"identity":"5ed09a25-1ca7-46b1-90dd-abd210db838f","added_by":"auto","created_at":"2026-02-27 09:28:41","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":344940,"visible":true,"origin":"","legend":"\u003cp\u003eIntraobserver agreement measures (ICC) for each resident and for the overall analysis, with confidence intervals.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8886056/v1/30d0b13876c2c877f81850b0.jpeg"},{"id":104398496,"identity":"24e76b2e-512e-49f2-a034-9f4953516c51","added_by":"auto","created_at":"2026-03-11 12:02:40","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":182419,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of interobserver agreement: among residents (ICC = 0.93, Excellent) and among residents + AI (ICC = 0.61, Very Good).\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8886056/v1/e6c5beb51b86905d82efa39a.jpeg"},{"id":104410289,"identity":"5e4dc680-182d-49aa-8af7-614f914b6bfd","added_by":"auto","created_at":"2026-03-11 12:51:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1265853,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8886056/v1/35c52717-d108-4ff6-9ab6-ca613fcfb598.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Evaluation of the Reproducibility and Consistency of the RUST Radiographic Scale in Tibial Fracture Healing: Comparison Between Orthopedic Residents and Artificial Intelligence","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eTibial shaft fracture is the most common among long bone fractures, showing a high incidence and predominantly affecting young males in their economically productive years. [\u003cspan additionalcitationids=\"CR2 CR3 CR4 CR5\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] Although locked intramedullary nailing is widely accepted as the standard treatment for tibial shaft fractures, there remains no consensus among orthopedic surgeons regarding the radiographic assessment of bone healing.[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eThe presence of periosteal callus formation visible on radiographs, combined with the absence of a fracture line, constitutes the most reliable radiographic sign of bone healing among different observers.[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] Whelan et al.[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] developed the Radiographic Union Scale for Tibial fractures (RUST), based on specific parameters that assess the progression of bone healing after intramedullary fixation using a locked nail.\u003c/p\u003e \u003cp\u003eSeveral studies have demonstrated substantial intra- and interobserver agreement for the RUST scale. [\u003cspan additionalcitationids=\"CR12 CR13 CR14 CR15 CR16 CR17 CR18\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] Recently, with the rise of Artificial Intelligence (AI), multiple studies have investigated its application in evaluating different long bone fractures.[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] However, to the best of our knowledge, no study has assessed the use of AI in determining radiographic evidence of bone union using the RUST score.\u003c/p\u003e \u003cp\u003eOur hypothesis is that AI may enhance the ability of orthopedic surgeons to identify bone healing in tibial shaft fractures treated with intramedullary nailing using the RUST scale. The primary objective of this study is to evaluate the performance of an AI tool trained with the RUST scoring system in diaphyseal fractures treated with intramedullary nails and to compare its results with those of orthopedic residents. The secondary objective is to assess intra- and interobserver agreement among residents at different levels of training.\u003c/p\u003e"},{"header":"MATERIALS AND METHODS","content":"\u003cp\u003eThis retrospective observational study included 47 skeletally mature patients with tibial shaft fractures treated with intramedullary nailing at a tertiary trauma center. Ethical approval was obtained from the institutional research ethics committee.\u003c/p\u003e \u003cp\u003ePatients were included if postoperative follow-up radiographs were available. Exclusion criteria comprised pathological fractures, infection, and inadequate image quality.\u003c/p\u003e \u003cp\u003eFour orthopedic residents (R1\u0026ndash;R4) independently evaluated all radiographs using the RUST scale at two separate time points four weeks apart. Observers were blinded to clinical data and to each other\u0026rsquo;s assessments.\u003c/p\u003e \u003cp\u003eEach cortex (anterior, posterior, medial, and lateral) received a score from 1 to 3 according to the RUST criteria:\u003c/p\u003e \u003cp\u003e1 \u0026mdash; visible fracture line without callus\u003c/p\u003e \u003cp\u003e2 \u0026mdash; callus with persistent fracture line\u003c/p\u003e \u003cp\u003e3 \u0026mdash; bridging callus without fracture line\u003c/p\u003e \u003cp\u003eThe sum generated a total score ranging from 4 to 12, with scores\u0026thinsp;\u0026ge;\u0026thinsp;7 considered radiographic consolidation.\u003c/p\u003e \u003cp\u003eThe same radiographs were analyzed by an artificial intelligence algorithm trained using RUST parameters described in previous validation studies\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eIntra- and interobserver reliability were assessed using the Intraclass Correlation Coefficient (ICC). Agreement between residents and AI was analyzed using Wilcoxon and Friedman tests. Significance was set at p\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e \u003c/div\u003e"},{"header":"RESULTS","content":"\u003cp\u003eForty-seven patients were included (71.7% male), with a mean age of 32.9 years. The median age was 29 years, the standard deviation was 11.2 years, and the coefficient of variation was 0.34, indicating high age variability. This heterogeneity was intentional, to ensure greater robustness in the concordance analysis.\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e presents the frequency distribution and statistical data for the total RUST scores of the 47 patients, obtained by both AI and the four evaluators. According to the AI classifications, most patients had total RUST scores between 7 and 9 (53.2%), with 61.7% of cases considered consolidated (scores\u0026thinsp;\u0026ge;\u0026thinsp;7). The frequency distributions of RUST total scores calculated by the evaluators showed different patterns.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003epresents the frequency distribution and statistical data for the total RUST scores of the 47 patients, obtained by both AI and the four evaluators. According to the AI.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEvaluator\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRUST Total Score Range\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFrequency\u003c/p\u003e \u003cp\u003e(n)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e%\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedian\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eMean\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eStandard Deviation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003ep-value of the Wilcoxon test comparing with AI\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eIA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 a 6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e38,3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e8,0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e7,4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1,7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7 a 9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e53,2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10 a 12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e8,5%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eR4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 a 6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e46,8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7,0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e6,8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e2,3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0,076\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7 a 9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e38,3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10 a 12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e12,8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eR3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 a 6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e21,3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e8,0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e7,8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1,9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0,190\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7 a 9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e57,4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10 a 12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e21,3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eR2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 a 6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e4,3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e8,0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e8,2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1,6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0,009\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7 a 9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e76,6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10 a 12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e19,1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eR1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 a 6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e25,5%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7,5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e7,5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1,8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0,795\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7 a 9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e59,6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10 a 12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e14,9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e presents the agreement analysis between the scores obtained by AI and the mean scores assigned by each evaluator, as well as the overall agreement (independent of evaluator). Considering the overall results, regardless of the evaluator, only 17.0% of assessments matched the AI score, while 37.8% of evaluations underestimated and 45.2% overestimated the AI score. The evaluator who showed the highest concordance with AI was R3, whose scores coincided with AI in 21.3% of cases and achieved the highest ICC\u0026thinsp;=\u0026thinsp;0.58. The lowest performances were observed for R4, who matched AI in only 10.6% of evaluations, and R2, who presented a statistically significant difference compared with AI scores (p\u0026thinsp;=\u0026thinsp;0.009 in the Wilcoxon test; ICC\u0026thinsp;=\u0026thinsp;0.44).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAgreement analysis between the scores obtained by AI and the mean scores of the evaluators, for each evaluator and overall.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eComparison Result with AI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"5\" nameend=\"c6\" namest=\"c2\"\u003e \u003cp\u003eEvaluator\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eR4\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eR3\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eR2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eR1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eGlobal\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAbsolute Agreement with AI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10,6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e21,3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e17,0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e19,1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e17,0%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEvaluator Underestimated AI Score\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59,6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e29,8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e23,4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e38,3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e37,8%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEvaluator Overestimated AI Score\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e29,8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e48,9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e59,6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e42,6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e45,2%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWilcoxon Test p-value\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,076\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,190\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,009\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,795\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0,355\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eICC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0,49\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e95% CI of ICC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,07\u0026thinsp;\u0026minus;\u0026thinsp;0,70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,26\u0026thinsp;\u0026minus;\u0026thinsp;0,77\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,03\u0026thinsp;\u0026minus;\u0026thinsp;0,68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,00\u0026ndash;0,69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0,31\u0026thinsp;\u0026minus;\u0026thinsp;0,61\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAgreement Classification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGood\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGood\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGood\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eGood\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eGood\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSignificance of Agreement\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNot confirmed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eConfirmed with reservation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNot confirmed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNot confirmed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eConfirmed with reservation\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows the interobserver reproducibility of the total RUST score. The agreement among the four evaluators (R1\u0026ndash;R4) was excellent (ICC\u0026thinsp;=\u0026thinsp;0.93; 95% CI: 0.84\u0026ndash;0.96). When the AI was included in the interobserver analysis, the ICC decreased to 0.61, still classified as very good, indicating significant agreement, but with a greater dispersion related to the AI\u0026rsquo;s stricter and standardized criteria.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eInterobserver reproducibility analysis (total RUST score).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEvaluators Compared\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eICC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e95% CI of ICC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAgreement Classification\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eSignificance of Agreement\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eFriedman Test p-value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eR4, R3, R2 and R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0,93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e(0,84;0,96)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eConfirmed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0,001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIA, R4, R3,R2 and R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0,61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e(0,47; 0,73)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eVery Good\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eConfirmed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0,001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e displays the intraobserver reproducibility, evaluating each resident\u0026rsquo;s consistency across both assessment rounds and the learning stability of the proposed method. The global intraobserver ICC was 0.72 (95% CI: 0.63\u0026ndash;0.78), demonstrating very good and statistically significant agreement.\u003c/p\u003e \u003cp\u003eEvaluator R4 showed the highest intraobserver stability (ICC\u0026thinsp;=\u0026thinsp;0.78), likely due to greater caution and rigor in applying scoring criteria, although this did not translate into higher concordance with AI. All evaluators exhibited \u0026ldquo;very good\u0026rdquo; intraobserver reliability, confirming that once the logic of the RUST scale is understood, consistent individual interpretation is achieved.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eIntraobserver reproducibility analysis (learning stability of evaluators).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eStatistics\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"5\" nameend=\"c6\" namest=\"c2\"\u003e \u003cp\u003eEvaluator\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eR4\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eR3\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eR2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eR1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOverall\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eICC\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0,72\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e95% CI of ICC\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(0,63; 0,87)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e(0,46; 0,82)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e(0,42; 0,81)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e(0,45; 0,81)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e(0,63; 0,78)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAgreement Classification\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVery Good\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eVery Good\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eVery Good\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eVery Good\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eVery Good\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSignificance of Agreement\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eConfirmed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eConfirmed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eConfirmed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eConfirmed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eConfirmed\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eThe radiographic assessment of tibial shaft fracture healing remains a challenge due to the variability of criteria used to define bone union. In this study, the application of the RUST scale allowed for a more objective and standardized analysis of bone healing progression in fractures treated with intramedullary nailing. The scoring method demonstrated high reproducibility, with agreement indices ranging from very good to excellent among residents from the first to fourth year of training. However, when compared with artificial intelligence (AI), the results showed low concordance, with only 17% absolute agreement, highlighting the need for further training to align human interpretation with algorithmic patterns.\u003c/p\u003e \u003cp\u003eSeveral radiographic scoring systems have been proposed to evaluate fracture healing, but their use remains limited. The Hammer [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] scale, which classifies fractures into five groups according to callus formation and fracture line obliteration, showed low correlation with mechanical stability and only moderate interobserver agreement (k\u0026thinsp;\u0026asymp;\u0026thinsp;0.60). The Tower et al. [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] scale, ranging from 0 to 10 based on callus presence and radiolucent lines, demonstrated good correlation with a combined clinical scale, although its validity as an isolated radiographic method could not be confirmed.\u003c/p\u003e \u003cp\u003eBased on the reliable assessment of cortical bridging and the frequently used criterion of fracture line visibility, it was hypothesized that the RUST scale represents a more robust and reliable tool than conventional assessment methods. [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]\u003c/p\u003e \u003cp\u003ePrevious studies have demonstrated good to excellent interobserver reliability of the RUST scale (ICC ranging from 0.84 to 0.95). Reproducibility proved consistent among orthopaedic surgeons and radiologists, and even medical students achieved similar indices, albeit with greater intraobserver variability. Moreover, evaluator experience has been shown to increase reliability, being higher among trauma specialists compared to residents.[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] Azevedo Filho et al. reported an ICC of 0.87 (95% CI: 0.81\u0026ndash;0.91). The authors noted that as evaluator experience increased, so did reliability: higher ICC values were observed among trauma surgeons compared to first, second, and third-year residents (ICC 0.94; 0.80; 0.92; and 0.90, respectively).[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eThe heterogeneity of the criteria used to define bone healing across different studies compromises the direct comparability of results. The use of more objective and standardized parameters, such as those proposed by the RUST scale, has the potential to provide greater uniformity, reproducibility, and consistency in information related to bone healing time.[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eIn this study, the sample consisted predominantly of male individuals (71.7%), aged between 17 and 60 years. This age diversity, characterized by high variability, was intentionally included to confer greater robustness to the agreement analyses.\u003c/p\u003e \u003cp\u003eOur study confirms the findings reported by Azevedo et al. [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] when evaluating fracture healing among first- to fourth-year residents using the RUST scoring method in 47 cases. The overall agreement among the four evaluators, without discriminating specific measures, showed a value of 0.82 (95% CI: 0.75\u0026ndash;0.87), confirming the significant reproducibility of the proposed classification method, with agreement classified as excellent when comparing the mean results of the two evaluation rounds.\u003c/p\u003e \u003cp\u003eIn the interobserver reproducibility analysis, the ICC among the four evaluators was 0.93 (95% CI: 0.84\u0026ndash;0.96), indicating excellent agreement regardless of the residents\u0026rsquo; training level (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). This finding may be attributed to the robustness of the RUST score, whose simplicity and objectivity reduce subjectivity and promote uniformity of evaluations, even among observers with different levels of experience. However, agreement with artificial intelligence (AI) proved limited: when AI was included in the interobserver analysis, the ICC dropped from 0.93 to 0.61 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). This result suggests that AI applies stricter and more uniform criteria, exposing differences in human judgment. On the other hand, it highlights the potential of AI as a training support tool, contributing to evaluator calibration and promoting greater standardization in the application of the RUST score.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAlthough all residents showed \u0026ldquo;good\u0026rdquo; agreement with AI, only for evaluator R3 was this agreement statistically significant, albeit with some limitations. These findings suggest that, when considering AI as a reference standard, residents require additional training to improve the application of the method and bring their interpretations closer to the algorithmic pattern.\u003c/p\u003e \u003cp\u003eThe frequency of overestimation (45.2%) and underestimation (37.8%) in comparison with AI suggests that residents have not yet fully mastered the criteria of each score (1 to 3 per cortex), leading to less accurate interpretations in intermediate cases. The tendency toward overestimation, in particular, may reflect an attempt to anticipate signs of clinical consolidation rather than strictly adhering to the established radiographic parameters.\u003c/p\u003e \u003cp\u003eOur findings are consistent with those of Pettersson et al.[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], who reported that artificial intelligence, based systems show excellent internal performance (AUC/ICC\u0026thinsp;\u0026gt;\u0026thinsp;0.8) and tend to yield comparable, though still divergent, results from human evaluations. This discrepancy reflects differences between subjective clinical judgment and the standardized criteria applied by AI. Like the Swedish study on elbow fractures, our findings reinforce the notion that AI should be viewed as a complementary tool, capable of enhancing human evaluator calibration, increasing standardization, and potentially reducing diagnostic errors. The convergence between both studies points to methodological consistency in the application of AI in orthopaedics, suggesting that algorithms trained with large volumes of labeled data following standardized criteria (such as AO/OTA or RUST) could, in the future, be safely integrated into clinical decision making, making and medical education workflows.[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eThis study has limitations, including retrospective design, absence of clinical correlation with weight-bearing progression, and single-center population. Nevertheless, the controlled radiographic evaluation allowed consistent comparison between human observers and AI assessment.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003e The RUST score is highly reproducible among orthopedic residents but shows limited agreement with AI analysis. Additional training and integration of automated tools may improve standardization of fracture healing assessment.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthical approval\u003c/h2\u003e\n\u003cp\u003eThis study was approved by the institutional research ethics committee and conducted in accordance with the Declaration of Helsinki.\u003c/p\u003e\n\u003ch2\u003eInformed consent\u003c/h2\u003e\n\u003cp\u003eThe requirement for informed consent was waived due to the retrospective design of the study.\u003c/p\u003e\n\u003ch2\u003eConflict of interest\u003c/h2\u003e\n\u003cp\u003eThe authors declare that they have no conflict of interest.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThe authors received no financial support for the research, authorship, and/or publication of this article.\u003c/p\u003e\n\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\n\u003cp\u003eConceptualization: PJIL, VMSI, GWMethodology: PJIL, WDB, REPData collection: NLLFS, IMR, AF, MBMData analysis: GWWriting \u0026ndash; original draft: PJIL, VMSI, WDB, REP, VGWriting \u0026ndash; review \u0026amp; editing: All authorsSupervision: PJL\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eCourt-Brown CM, Rimmer S, Prakash U, McQueen MM (1998) The epidemiology of open long bone fractures. Injury 29(7):529\u0026ndash;342\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKooistra BW, Dijkman BG, Busse JW, Sprague S, Schemitsch EH, Bhandari M (2010) The radiographic union scale in tibial fractures: reliability and validity. J Orthop Trauma. ;24Suppl. 3:S81\u0026ndash;6\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChua W, Murphy D, Siow W, Kagda F, Thambiah J (2012) Epidemiological analysis of outcomes in 323 open tibial diaphyseal fractures: a nine-year experience. Singap Med J 53(6):385\u0026ndash;389\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKojima KE, Ferreira RV (2011) Fraturas da di\u0026aacute;fise da t\u0026iacute;bia. Rev Bras Ortop 46(2):130\u0026ndash;155\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWhelan DB, Bhandari M, McKee MD, Guyatt GH, Kreder HJ, Stephen D et al (2002) Interobserver and intraobserver variation in the assessment of the healing of tibial fractures after intramedullary fixation. J Bone Joint Surg Br 84:15\u0026ndash;86\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZeckey C, Mommsen P, Andruszkow H, Macke C, Frink M, St\u0026uuml;big T et al (2011) The aseptic femoral and tibial shaft non-unionin healthy patients \u0026ndash; an analysis of the health-related quality of life and the socioeconomic outcome. Open Orthop J 5:193\u0026ndash;197\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBhandari M, Guyatt GH, Swiontkowski MF et al (2002) A lack of consensus in the assessment of fracture healing among orthopaedic surgeons. J Orthop Trauma 16:562\u0026ndash;566\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKooistra BW, Dijkman BG, Busse JW et al (2010) The radiographic union scale in tibial fractures: reliability and validity. J Orthop Trauma 24:S81\u0026ndash;S86\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHammer RR, Hammerby S, Lindholm B (1985) Accuracy of radiologic assessment of tibial shaft fracture union in humans. Clin Orthop Relat Res. ;(199):233\u0026ndash;238s\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWhelan DB, Bhandari M, Stephen D et al (2010) Development of the radiographic union score for tibial fractures for the assessment of tibial fracture healing after intramedullary fixation. J Trauma 68:629\u0026ndash;632\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLeow JM, Clement ND, Tawonsawatruk T, Simpson CJ, Simpson AHRW (2016) The radiographic union scale in tibial (RUST) fractures: reliability of the outcome measure at an independent center. Bone joint Res 5(4):116\u0026ndash;121\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTower SS, Beals RK, Duwelius PJ (1993) Resonant frequency analysis of the tibia as a measure of fracture healing. J Orthop Trauma 7:552\u0026ndash;557\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMorshed S, Corrales L, Genant H et al (2008) Outcome assessment in clinical trials of fracture-healing. J Bone Joint Surg Am 90:62\u0026ndash;67\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAli S, Singh A, Agarwal A, Parihar A, Mahdi AA, Srivastava RN (2014) Reliability of the RUST score for the assessment of union in simple diaphyseal tibial fractures. Int J Biomed Res 5(5):333\u0026ndash;335\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMane SS, Yamajala SNSMLV, Mitnala SRP (2024) Evaluating reliability of the RUST score in diaphyseal Tibia fractures: A collaborative assessment by orthopaedic surgeons and radiologists. J Orthop Rep 3(4):100325\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAzevedo Filho FAS, Cotias RB, Azi ML, Teixeira AAA (2017) Reliability of the radiographic union scale in tibial fractures (RUST). Revista Brasileira de Ortop (English Edition) 52(1):35\u0026ndash;3916\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChloros GD, Howard A, Giordano V, Giannoudis PV (2020) Radiographic Long Bone Fracture Healing Scores: Can they predict non-union? Injury 51(8):1693\u0026ndash;1695\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWojahn RD, Bechtold D, Abraamyan T, Spraggs-Hughes A, Gardner MJ, Ricci WM, McAndrew CM (2022) Progression of tibia fracture healing using RUST: are early radiographs helpful? J Orthop Trauma 36(1):e6\u0026ndash;e11\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMisir A, Uzun E, Kizkapan TB, Yildiz KI, Onder M, Ozcamdalli M (2020) Reliability of RUST and Modified RUST Scores for the Evaluation of Union in Humeral Shaft Fractures Treated with Different Techniques. Indian J Orthop 54(Suppl 1):121\u0026ndash;126\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePettersson A, Axenhus M, Stukan T et al (2025) Use of artificial intelligence for classification of fractures around the elbow in adults according to the 2018 AO/OTA classification system. BMC Musculoskelet Disord 26(1):848\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSevin\u0026ccedil; H\u0026uuml;seyin, Fatih et al (2025) Detection and classification of femoral neck fractures from plain pelvic X-rays using deep learning and machine learning methods. Ulus Travma Acil Cerrahi Derg 31(8):783\u0026ndash;788\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"european-journal-of-orthopaedic-surgery-and-traumatology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ejos","sideBox":"Learn more about [European Journal of Orthopaedic Surgery \u0026 Traumatology](http://link.springer.com/journal/590)","snPcode":"590","submissionUrl":"https://submission.springernature.com/new-submission/590/3","title":"European Journal of Orthopaedic Surgery \u0026 Traumatology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"tibial fracture, fracture healing, radiographic union score, reliability, validity","lastPublishedDoi":"10.21203/rs.3.rs-8886056/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8886056/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003ePurpose\u003c/h2\u003e \u003cp\u003eTo evaluate the reproducibility and consistency of the Radiographic Union Score for Tibial fractures (RUST) in diaphyseal fractures treated with intramedullary nailing among orthopedic residents with different levels of experience, and to compare their performance with artificial intelligence (AI) analysis.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eRadiographs from 47 patients were assessed by four orthopedic residents (R1\u0026ndash;R4) at two independent time points using the RUST scale. The same images were evaluated by AI. Intra- and interobserver agreement were calculated using the Intraclass Correlation Coefficient (ICC). Differences between human and AI evaluations were analyzed using Wilcoxon and Friedman tests with a 5% significance level.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eThe sample consisted predominantly of men (71.7%) with a mean age of 32.9 years. AI classified 61.7% of fractures as consolidated (score\u0026thinsp;\u0026ge;\u0026thinsp;7). Absolute agreement between residents and AI was 17%, with residents overestimating scores in 45.2% and underestimating in 37.8% of cases. Observer R3 showed the best agreement with AI (ICC\u0026thinsp;=\u0026thinsp;0.58), while R2 demonstrated a significant difference (p\u0026thinsp;=\u0026thinsp;0.009). Interobserver reproducibility among residents was excellent (ICC\u0026thinsp;=\u0026thinsp;0.93; 95% CI 0.84\u0026ndash;0.96) but decreased to 0.61 when AI was included. Intraobserver consistency was very good (global ICC\u0026thinsp;=\u0026thinsp;0.72; 95% CI 0.63\u0026ndash;0.78), with R4 presenting the highest stability (ICC\u0026thinsp;=\u0026thinsp;0.78).\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003e The RUST scale showed high reproducibility among residents but limited agreement with AI, suggesting the need for additional training to improve consistency between human interpretation and algorithmic assessment.\u003c/p\u003e","manuscriptTitle":"Evaluation of the Reproducibility and Consistency of the RUST Radiographic Scale in Tibial Fracture Healing: Comparison Between Orthopedic Residents and Artificial Intelligence","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-27 09:28:33","doi":"10.21203/rs.3.rs-8886056/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-04-30T15:36:34+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-29T04:27:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"59673943067955407711248432278703481480","date":"2026-04-29T03:37:46+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-21T07:58:39+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"166342306601229441989146657336359351739","date":"2026-04-19T00:03:12+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-01T23:30:44+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"237869790448248756199932144464091854166","date":"2026-03-01T13:04:24+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-02-24T05:21:41+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-17T07:19:47+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-17T07:17:13+00:00","index":"","fulltext":""},{"type":"submitted","content":"European Journal of Orthopaedic Surgery \u0026 Traumatology","date":"2026-02-15T12:46:43+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"european-journal-of-orthopaedic-surgery-and-traumatology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ejos","sideBox":"Learn more about [European Journal of Orthopaedic Surgery \u0026 Traumatology](http://link.springer.com/journal/590)","snPcode":"590","submissionUrl":"https://submission.springernature.com/new-submission/590/3","title":"European Journal of Orthopaedic Surgery \u0026 Traumatology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"3de23f6b-1d72-4528-a74f-0035d1bfe4bd","owner":[],"postedDate":"February 27th, 2026","published":true,"recentEditorialEvents":[{"type":"decision","content":"Revision requested","date":"2026-04-30T15:36:34+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-29T04:27:35+00:00","index":21,"fulltext":""},{"type":"reviewerAgreed","content":"59673943067955407711248432278703481480","date":"2026-04-29T03:37:46+00:00","index":20,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-02T00:38:10+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-27 09:28:33","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8886056","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8886056","identity":"rs-8886056","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0