SHIELD: An AI Framework for Skeletal Health Intelligence and Early Lesion Detection to Improve Orthopedic Referrals

preprint OA: closed
Full text JSON View at publisher
AI-generated deep summary by claude@2026-06, 2026-06-24 · read from full text

The SHIELD study developed and validated an AI framework to speed and standardize triage of radiology reports for suspected metastatic bone disease, using RadBERT-RoBERTa fine-tuned on a decade of radiology reports (N=245 patients) from two academic medical centers. The model classified reports into three referral tiers (“No Referral,” “Referral,” and “Referral/High Risk”) and incorporated an LLM-based natural-language explanation component to improve transparency. On a hold-out test set, it reported 100% accuracy and AUC of 1.00 for distinguishing referral versus non-referral, and 89.52% overall accuracy for the three-class task, with a “fail-safe” constraint of never misclassifying high-risk cases as requiring no referral; the paper frames a key limitation as the preprint status (not peer reviewed) and the retrospective, cohort-specific data source. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Delays in the referral of patients with suspected metastatic bone disease (MBD) from radiology reports represent a critical challenge that can negatively impact patient outcomes. The conventional manual review process is often a significant bottleneck, leading to prolonged diagnostic timelines. We developed and validated SHIELD, an automated AI framework designed to accelerate and improve the accuracy of MBD referrals. We fine-tuned a RadBERT-RoBERTa model on a decade of radiology reports (N = 245 patients) from two academic medical centers to classify reports into three tiers: "No Referral," "Referral," and "Referral/High Risk." To ensure clinical utility and transparency, SHIELD incorporates a Large Language Model to generate natural-language explanations for its classifications. SHIELD demonstrated exceptional performance on a hold-out test set. It achieved 100% accuracy and an Area Under the Curve (AUC) of 1.00 in the primary binary task of distinguishing referral from non-referral cases. In the more granular three-class task, the model achieved an overall accuracy of 89.52%, with near-perfect performance in identifying "No Referral" reports (F1-score: 99.20%). Critically, the model operated in a clinically "fail-safe" manner, never misclassifying a high-risk case as requiring no referral. A retrospective timeline analysis revealed that SHIELD can reduce the referral period from a conventional average of 109.6 days to a computational time of 1–3 minutes. Proposed work provides high accuracy with a sophisticated explainability component using a large language model. Thus, SHIELD framework is a robust, explainable, and autonomous solution for triaging radiology reports. By drastically reducing administrative and diagnostic delays, it has the potential to significantly accelerate the clinical workflow, ensure timely specialist consultation, and ultimately improve the standard of care for patients with suspected MBD.
Full text 163,483 characters · extracted from preprint-html · click to expand
SHIELD: An AI Framework for Skeletal Health Intelligence and Early Lesion Detection to Improve Orthopedic Referrals | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article SHIELD: An AI Framework for Skeletal Health Intelligence and Early Lesion Detection to Improve Orthopedic Referrals Abbas Alili, Ava M. McKane, Fatih M. Demir, Cynthia L. Emory, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7926103/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 3 You are reading this latest preprint version Abstract Delays in the referral of patients with suspected metastatic bone disease (MBD) from radiology reports represent a critical challenge that can negatively impact patient outcomes. The conventional manual review process is often a significant bottleneck, leading to prolonged diagnostic timelines. We developed and validated SHIELD, an automated AI framework designed to accelerate and improve the accuracy of MBD referrals. We fine-tuned a RadBERT-RoBERTa model on a decade of radiology reports (N = 245 patients) from two academic medical centers to classify reports into three tiers: "No Referral," "Referral," and "Referral/High Risk." To ensure clinical utility and transparency, SHIELD incorporates a Large Language Model to generate natural-language explanations for its classifications. SHIELD demonstrated exceptional performance on a hold-out test set. It achieved 100% accuracy and an Area Under the Curve (AUC) of 1.00 in the primary binary task of distinguishing referral from non-referral cases. In the more granular three-class task, the model achieved an overall accuracy of 89.52%, with near-perfect performance in identifying "No Referral" reports (F1-score: 99.20%). Critically, the model operated in a clinically "fail-safe" manner, never misclassifying a high-risk case as requiring no referral. A retrospective timeline analysis revealed that SHIELD can reduce the referral period from a conventional average of 109.6 days to a computational time of 1–3 minutes. Proposed work provides high accuracy with a sophisticated explainability component using a large language model. Thus, SHIELD framework is a robust, explainable, and autonomous solution for triaging radiology reports. By drastically reducing administrative and diagnostic delays, it has the potential to significantly accelerate the clinical workflow, ensure timely specialist consultation, and ultimately improve the standard of care for patients with suspected MBD. Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction Metastatic bone disease represents a significant complication arising from advanced malignancies. In the United States, it is estimated that approximately 300,000 adults are living with MBD 1 . Although advancements in systemic therapies for various cancers have improved patient longevity, the prognosis after developing bone metastases is variable, with a wide range of survival rates depending on the primary type of cancer 2 . For example, one-year survival after a bone metastasis diagnosis can range from as high as 51% for breast cancer to as low as 10% for lung cancer 3 . The management of bone metastases also has significant implications for healthcare expenditures. Early intervention for patients with MBD has been demonstrated to reduce patient morbidity and overall healthcare costs, which are estimated to exceed $ 12.6 billion annually in the United States 4 . Despite the urgent nature of MBD, many patients are not referred to an orthopedic provider until after a pathologic fracture has occurred, leading to a missed opportunity for prophylactic stabilization 5 . A patient’s initial diagnosis of cancer or MBD can be overwhelming for patients and caregivers, with multiple competing priorities occurring at the same time. Knowledge gaps from referring providers on the acuity of impending fractures based upon location, size, or extent of a bone lesion can lead to disparate referral patterns. Access to musculoskeletal experts can also lead to delays in referrals to the treating surgeon. Other sites of metastasis can influence prioritization of referrals and subspecialty appointments, and unintended delays in orthopaedic referrals may occur as a result. This delay underscores the need for improved and standardized referral pathways to orthopedic oncology to enable timely intervention and optimize patient outcomes 6 , 7 . While Mirels' score 8 is a widely used tool within orthopedics to assess fracture risk based on radiographic and clinical features, its adoption outside of the specialty is limited. Furthermore, its low specificity (35%) creates a risk of overtreatment 9 . Referrals to orthopedic specialists are often based on subjective clinical judgment, non-standardized imaging terminology, or non-specific symptoms such as pain, leaving the interpretation and decision to refer largely at the discretion of the non-specialist provider. Recent advances in artificial intelligence (AI) offer promising solutions to expedite the referral process and reduce the workload of medical experts 10 , 11 . For instance, Vergara et al. 12 developed an AI model for gatekeeping referrals from primary to specialized care, which achieved moderate accuracy in distinguishing between authorized referrals, outperforming human gatekeepers by nearly 20% primarily due to higher specificity. Another study demonstrated that incorporating natural language processing (NLP)-extracted qualitative data from referral letters significantly increases the accuracy of machine learning (ML) models by up to 19.5% for triaging patients with low back pain to appropriate interventions, despite the overall model accuracy remaining low for clinical application 13 . The ARCHERY project also utilized NLP with a clinically based large language model (LLM) on free-text radiology reports to predict patient selection for total hip and knee arthroplasties, demonstrating promising potential for hip arthroplasty but not for knee arthroplasty. Their work highlighted the importance of further model testing and training for new clinical cohorts 14 . LLMs being at the forefront of AI have enabled the processing of vast data corpora for various downstream tasks such as classification, analysis, and prediction 15 – 17 . The promise of these transformer-based architectures has spurred interest in exploring even larger models, often containing billions of parameters. In the biomedical domain, researchers have developed specialized transformer models like BioBERT 18 and PubMedBERT 19 (each comprising 110 million parameters) by training them on biomedical literature from PubMed. Furthermore, clinically specific language models such as RadBERT have been adapted for the field of radiology 20 . These models have demonstrated strong performance in extracting relevant information from free-text radiology reports and have shown potential in predicting the need for surgical interventions based on the language used in imaging reports. The growing capabilities of AI in healthcare are accompanied by a corresponding scepticism regarding its adoption. A recent study investigated how explainable artificial intelligence (XAI) influences clinicians' trust in AI applications, addressing the challenge of fostering appropriate reliance on AI's often "black box" nature 21 . The majority of studies included suggest that XAI has the potential to enhance clinicians' trust in AI recommendations. However, complex or contradictory explanations can undermine this trust, whereas excessive trust in incorrect AI advice can adversely impact clinical accuracy 22 . This highlights a pressing need for XAI solutions that provide clear, understandable, and reliable explanations, especially due to ethical and regulatory concerns in healthcare 23 , 24 . Despite advancements in the application of AI in healthcare and the associated challenges of explainability, very limited research has specifically addressed the problem of delayed referrals for metastatic bone disease. The overall goal of this study was to introduce SHIELD (Skeletal Health Intelligence and Early Lesion Detection), an AI framework designed to prevent late referrals by providing highly accurate classifications and offering appropriate explainability to aid medical experts in their decision-making. Our study had three aims. The primary aim of this work was to 1) develop an AI framework to classify radiology reports into three distinct referral categories: no referral needed, referral recommended, and referral/high-risk. This was achieved by fine-tuning the RadBERT-RoBERTa-4m 20 model for MBD referral classification using a decade of data (January 2014 – May 2025) from two affiliated academic medical centers. RadBERT-RoBERTa-4m 20 is a transformer-based language model developed for radiology and clinical NLP. A secondary aim was to 2) provide explanations for the framework's classification decisions, thereby increasing trust and adoption among medical providers. We utilized the Llama-3.1-8B-Instruct, Meta’s instruction-tuned language model that was released in July 2024, to accomplish this task. Our final aim was to 3) validate the explainability of the proposed framework and assess the overlap between the terminology used by the AI and a predefined set of terms established by orthopedic experts. 2. Methods 2.1 Cohort Selection, Radiology Report Review, and Annotation by Clinical Experts Following institutional review board approval (IRB00127840), we conducted a retrospective review of all patients with a diagnosis of secondary neoplasm of bone treated at two affiliated academic medical centers over a 10-year period (January 2014 – May 2025) by analyzing 1404 Electronic Health Records (EHR). Patients were identified using ICD codes corresponding to secondary malignant neoplasm of bone (C79.51). Clinical documentation and radiology reports from both plain and advanced imaging modalities (PET, CT, MRI, and radiographs) were reviewed. Patients were included if they were ≥ 18 years of age at the time of MBD diagnosis and had evidence of osseous metastases to the appendicular long bones, defined as the femur , fibula , tibia , humerus , radius , or ulna . Patients were excluded if they were referred to orthopedic oncology prior to evidence of long-bone metastases (LBM) or were never referred to orthopedic oncology, had their initial imaging dated before January 2014, had only axial or non-LBM, were under 18 years of age at diagnosis, or had incomplete or inaccessible imaging and medical records. From an initial cohort of 514 patients with secondary neoplasm of bone, 245 met full inclusion criteria. A control cohort was generated by screening adult patients with any cancer diagnosis over the same 10-year period who had undergone at least one imaging study (PET, CT, MRI, and radiograph) indicating cancer at either institution. Patients were excluded from this control cohort if they had any mention of osseous metastases in the appendicular skeleton, a diagnosis of primary bone tumor, were under 18 years old, or lacked complete medical documentation. Of the 890 screened patients, 245 met the inclusion criteria for the control group. The cohort selection process is summarized in Fig. 1 . Data was de-identified and collected using REDCap, a secure, web-based platform. Patient demographics, such as age, sex, race/ethnicity, comorbidities, and primary cancer type, were recorded and listed in Table 1. For the MBD cohort, the timeline from first imaging mention of metastasis to referral to orthopedic oncology was documented, along with prior treatments including chemotherapy and radiation. Raw radiology reports were obtained for all relevant imaging studies from the date of the first long bone metastatic lesion to the date of referral. Reports were then analyzed for descriptive language using a list of predefined terms indicative of MBD or SREs. This list of terms was curated a priori by an orthopedic oncologist to avoid bias during data extraction. Specific lesion descriptors, signs of impending fracture, and evidence of progression were recorded. Table 1. Participant Demographic Information Based on 245 Patients with Metastatic Bone Disease (MBD) and Complete Trained Healthcare Professional Data. Characteristic Patients, No. (%) Age Mean = 62 years, (SD = 11.5) Sex a Female 138 (56.3%) Male 107 (43.7%) Race a White 167 (61.2%) Black or African American 43 (17.6%) Asian 3 (1.2%) Other 11 (4.5%) Ethnicity a Hispanic or Latino 8 (3.3%) Not Hispanic or Latino 237 (96.7%) Body Mass Index 30 92 (37.6%) Comorbidities b COPD 16 (6.5%) Diabetes 48 (19.6%) Hypercholesterolemia 19 (7.8%) Hyperlipidemia 53 (21.6%) Hypertension 121 (49.4%) Osteoporosis 10 (4.1%) Renal Insufficiency 12 (4.9%) None 81 (33.1%) a Data on patient demographic characteristics, including age, sex, race, and ethnicity, were collected from the electronic health records at two academic institutions and classified by standardized categories as defined by the investigators. b Indicates selected pertinent comorbidities present at time of first imaging demonstrating MBD. 2.2 Data Preprocessing We curated a dataset of 555 radiology reports annotated by clinical experts into three referral classes: 0 = No Referral, 1 = Referral, and 2 = Referral/High Risk. To ensure a robust evaluation, the dataset was split into a stratified hold-out test set (15%) and a train/validation pool (85%). The stratification preserved class balance across subsets. Randomization was fixed for reproducibility. We trained the 'No Referral' and 'Referral' classes using only radiology reports from the initial patient visit. In contrast, for the 'Referral/High Risk' class, due to the scarcity of the data, we augmented the dataset by treating reports from all of a patient's visits as distinct data samples. We applied a sliding-window tokenization strategy using the RadBERT tokenizer, as many reports exceeded the model’s maximum input length. Each report was split into overlapping segments with a maximum length of 512 tokens and a stride of 256 tokens. This approach enabled overlapping token spans and multi-window representation per report while maintaining compatibility with the pre-trained transformer backbone. Tokenization was performed after the data split to prevent leakage between sets. Each segment inherited the label of its original report, producing a window-level dataset for model training and evaluation. 2.3 SHIELD’s Model Architecture for Radiology Report Classification The proposed classification model is based on the RadBERT-RoBERTa-4m 20 transformer architecture fine-tuned for three-class classes (Fig. 1 .). The model takes tokenized windows of radiology reports as input. Each window is encoded by the pretrained RadBERT encoder to produce contextual embeddings. The [CLS] token representation is passed through a classification head consisting of a linear layer to generate logits for the three output classes. Cross-entropy loss is used for training. A 5-fold stratified cross-validation was performed on the train/validation set, maintaining label distribution in each fold. During each fold iteration, tokenized train and validation windows were generated independently. After model inference, window-level predictions were aggregated back to report-level predictions using two schemes: Majority vote of the predicted classes across all windows for a report. Mean probability aggregation, where class probabilities were averaged across windows and the class with the highest mean probability was selected. Training was conducted using the Hugging Face Trainer API with the following settings: Learning rate: 2e-5 Batch size: 4 for both training and evaluation Number of epochs: 8 Evaluation and model checkpoint saving at each epoch Best model selection based on validation macro F1-score Logging every 10 steps To evaluate generalization, each fold’s model was tested on the same unseen hold-out test set. This set’s tokenization and window-counts were precomputed once to ensure consistent evaluation across folds. Final ensemble predictions were derived by averaging predicted probabilities across all five folds. During inference, the model generates logits for each window, which are converted to probabilities using softmax. These probabilities are then aggregated per report using the two methods described earlier. To enhance performance and stability: We performed model ensembling by averaging predicted class probabilities from all five folds. Aggregated predictions were evaluated using both majority voting and mean-max probability , and both strategies were benchmarked on the hold-out test set. Receiver Operating Characteristic (ROC) curve analysis was performed for two key binary tasks: distinguishing “No Referral” from “Referral” and “Referral/High Risk” combined, and distinguishing “Referral” from “Referral/High Risk”, yielding interpretable AUC values to quantify discriminative performance. Final outputs included per-class performance metrics (mean ± std across folds), confusion matrices, and ROC curve plots. Proposed architecture leveraged domain-specific pretraining, sliding-window augmentation, cross-validated ensembling, and robust evaluation strategies to deliver reliable and interpretable triage predictions from unstructured radiology narratives. 2.4 Generating Interpretable Explanations with SHIELD's Generative AI To provide interpretable rationales for each referral prediction, we implemented an automated pipeline using the Llama-3.1-8B-Instruct 25 model hosted locally. Radiology reports and their predicted class labels (No Referral, Referral, Referral/High Risk) obtained using our fine-tuned UCSD-VA-health/RadBERT-RoBERTa-4m 20 model were stored in a structured CSV file and processed sequentially. For each case, the radiology report text and its predicted label were combined into a class-conditioned prompt. The prompts explicitly requested the model to identify and explain the linguistic or diagnostic features supporting the assigned decision. The following prompts were used: Class 0 (No Referral) : “Identify and explain the terminology or findings that would support NOT referring this patient to the Orthopedic Oncology department.” Class 1 (Referral) : “Identify and explain the key terminology or findings that would indicate the patient SHOULD be referred to the Orthopedic Oncology department.” Class 2 (Referral/High Risk) : “Identify and explain the specific terminology or findings that suggest the patient should be referred to the Orthopedic Oncology department due to risk of pathological fracture (emergency). ” The radiology report text was embedded within this prompt, and the model was queried to generate free-text explanations. Each generation was constrained to a maximum of 1024 tokens with a low sampling temperature (0.3) to encourage concise and reproducible outputs. The resulting explanations were appended to the original dataset as an additional column and exported as a new CSV file for subsequent analysis. This procedure is depicted in Fig. 2. and ensures that every report–prediction pair is accompanied by a standardized, model-generated explanation, enabling a systematic review of the linguistic cues underlying the automated decisions. 2.5 Comparative Analysis of Clinician and LLM-Generated Terminology To systematically compare terminology used by clinicians with that generated by a large language model (LLM) for orthopedic oncology referral, we implemented a three-phase computational analysis pipeline. The input consisted of two curated term lists: (i) unique medical terms provided by an expert orthopedist, which two medical students used to sort radiology reports into three classes and label them, and (ii) a text-based explanation of the decision generated by the LLM, which the medical students filtered out to keep the main terms only. Phase 1: Preprocessing : All terms were normalized to lowercase, stripped of punctuation, tokenized, and lemmatized using the Natural Language Toolkit (NLTK). This ensured consistent lexical forms across both datasets (e.g., “lesions” and “lesion” were reduced to a common root form). Phase 2: Lexical (Set-Based) Analysis : We applied set-theoretic comparisons to quantify overlap and exclusivity between clinician and LLM-derived terminology. Terms were categorized into three groups: shared terms, clinician-only terms (missed by the LLM), and LLM-only terms (absent from clinician usage). The lexical overlap was visualized using a Venn diagram, enabling rapid assessment of concordance and divergence. Phase 3: Semantic and Conceptual Analysis : To capture conceptual similarities beyond exact word matches, we employed a BioBERT (dmis-lab/biobert-v1.1) model to generate vector embeddings of clinician-only and LLM-only terms 26 , 27 . Cosine similarity was computed between all term pairs, producing a similarity matrix. This enabled identification of LLM terms that were semantically aligned with clinician terms despite lexical differences. Heatmaps were generated to display the 30 most descriptive term pairs. Furthermore, similarity categories were defined as High (likely synonyms, score > 0.9) , Moderate (related concepts, score 0.75–0.9) , or Low (likely unrelated, score < 0.75) . Results were exported to a CSV file for detailed inspection. This multi-phase pipeline provided both lexical and semantic perspectives on the relationship between clinician-reported and LLM-generated terminology, enabling systematic identification of overlaps, mismatches, and conceptually related terms. 3. Results 3.1 Classification Performance The performance of the proposed SHIELD model was evaluated on the hold-out test set, with performance metrics averaged across a 5-fold cross-validation. Although the majority vote and the mean probability aggregation methods showed very close results, we will only present the latter. The model achieved a high overall accuracy of 89.52% ± 2.84% as shown in Table 2 . For the "No Referral" category, the model demonstrated outstanding performance. It achieved a precision of 98.97% ± 2.29%, a recall of 99.46% ± 1.21%, and a resulting F1-score of 99.20% ± 1.18%. This indicates exceptional reliability in correctly identifying reports that do not require further clinical action, a crucial function for reducing unnecessary workload. The model was also highly effective in identifying reports requiring clinical follow-up. For the general "Referral" class, it yielded a recall of 96.55% ± 4.22%, signifying that very few referral cases were missed. The precision for this class was 78.82% ± 4.27%, leading to a well-balanced F1-score of 86.72% ± 3.27%. Distinguishing between standard “Referral” and urgent referrals (Referral/High Risk) proved the most challenging aspect of model performance.. In this task, while the model exhibited high precision at 94.05% ± 9.37%, indicating a low rate of false positives the recall for this class was considerably lower (57.78% ± 13.94%). This trade-off between high certainty and moderate sensitivity resulted in an F1-score of 70.53% ± 10.73% for this critical category. These results indicate that the model is more reliable at correctly identifying urgent referrals when it flags them as such, but misses over half of all actual high-risk cases. Table 2 Cross-Validated Classification Performance of the AI Model. This table summarizes the performance of the final model on the hold-out test dataset. The reported metrics of precision, recall, and F1-score averaged over five validation folds. Predictions for each report were aggregated by calculating the mean of the maximum probabilities across all text segments. All performance metrics are presented as the mean ± standard deviation. Classes Precision Recall F1-score Support No Referral 0.9897 ± 0.0229 0.9946 ± 0.0121 0.9920 ± 0.0118 37 Referral 0.7882 ± 0.0427 0.9655 ± 0.0422 0.8672 ± 0.0327 29 Referral/High Risk 0.9405 ± 0.0937 0.5778 ± 0.1394 0.7053 ± 0.1073 18 Accuracy 0.8952 ± 0.0284 84 To further visualize the classification performance, confusion matrices for both binary and multi-class scenarios were generated from the ensemble model's predictions on the hold-out test set. The first matrix (Fig. 3 (a)) illustrates the model's performance on the primary binary task of distinguishing reports that require a referral from those that do not. In this scenario, the "Referral" and "Referral/High Risk" classes were consolidated into a single "Referral" category. The model achieved perfect classification, correctly identifying all 37 "No Referral" cases and all 47 "Referral" cases. This result underscores the model's exceptional capability to reliably determine whether a clinical follow-up is necessary. The second matrix (Fig. 3 (b)) provides a more granular, three-class analysis. The model continued to demonstrate flawless performance for the "No Referral" and standard "Referral" categories, correctly classifying all 37 and 29 reports, respectively. The model's classification errors were isolated to the "Referral/High Risk" category. Of the 18 true "Referral/High Risk" cases, 10 were correctly identified. The remaining 8 cases were misclassified as a standard "Referral." Critically, no "Referral/High Risk" case was misclassified as "No Referral." This indicates that while the model struggles to distinguish the nuanced differences between a standard and a high-risk referral, its errors are "fail-safe" in a clinical sense—it successfully flags all high-risk patients as requiring a referral, even if it de-escalates the urgency level for a subset of them. The discriminative power of the model was further assessed using Receiver Operating Characteristic (ROC) curve analysis for two key classification scenarios, as shown in Fig. 3 (c). For the primary task of distinguishing "No Referral" reports from all referral cases ("Referral" and "Referral/High Risk" combined), the model achieved a perfect Area Under the Curve (AUC) of 1.00. The blue line, representing this task, aligns perfectly with the top-left corner of the plot, indicating a flawless separation between the two groups. This result confirms that the model can achieve a 100% true positive rate (sensitivity) with a 0% false positive rate (100% specificity), underscoring its reliability for initial screening. In the more challenging sub-classification task of differentiating between a standard "Referral" and a "Referral/High Risk" case, the model also demonstrated excellent performance. The red curve shows a strong ability to distinguish between these two classes, yielding a high AUC of 0.93. This indicates a robust capability to correctly rank and identify high-risk cases over standard referral cases. Collectively, the ROC analysis confirms the findings from the confusion matrices: the model is exceptionally accurate in the primary determination of referral need and maintains a high degree of discriminative ability when stratifying the urgency of that referral. 3.2 Comparison of Conventional vs. SHIELD-Based Referral Timelines To quantify the potential clinical impact of our AI framework, we compared the referral timelines of the proposed model against the conventional clinical pathway, with the results summarized in Table 3 . The analysis revealed that the conventional method is associated with substantial delays, requiring an average of 109.64 ± 286.09 days from the initial imaging report to a patient referral. Additional delays were observed within this period, including an average of 80.57 days from the first documented high-risk descriptor to referral and a diagnostic window spanning 97.50 days between the first and last imaging events prior to referral. In stark contrast, the AI model provides its classification recommendation within 1 to 3 minutes of processing the initial report. By generating an immediate recommendation from a single data point, the model renders the prolonged, multi-step diagnostic timelines of the conventional pathway obsolete, demonstrating its potential to drastically reduce administrative and clinical delays inherent in the current standard of care. Table 3 Comparison of Referral Timelines Between Conventional Methods and the Proposed AI Model. Metric Conventional Method (days, mean ± SD) Proposed AI Model (minutes) Time from Initial Imaging to Referral Decision 109.64 ± 286.09 ~ 1–3 Time from First High-Risk Finding to Referral 80.57 ± 256.87 Diagnostic Period (First to Last Imaging Pre-Referral) 97.50 ± 281.99 Of the 245 patients evaluated, 124 (50.61%) had undergone at least two imaging visits prior to referral. Additionally, 85 patients (34.69%) experienced delayed referrals, defined as a referral made more than 28 days after the first imaging report indicating potential bone malignancy. The proposed AI solution demonstrated promising results by substantially reducing the referral delay, successfully classifying referrals based solely on the first imaging report for "No Referral" and “Referral" classes. 3.3 Explainable AI Table 4 presents a side-by-side comparison of an original radiology report accurately classified as Referral/High-Risk and its corresponding AI-generated explanation. Table 4 An example of an explained radiology report provided by LLM. Original Radiology Report AI Explanation 6/21/22 - XR Femur Minimum 2 Vw Right Status: Final result Study Result Narrative & Impression DATE OF SERVICE: 6/21/2022 4:10 pm EXAM: Right femur two views Clinical history: Bone Mets, right femur pain. Comparison: CT pelvis 3/26/2022 Findings: Mixed sclerotic and lytic osseous lesions are again noted. There is a pathologic fracture without significant displacement at a lytic lesion of the lesser trochanter of the right proximal femur. 6/21/22 - XR Pelvis 1–2 Views Status: Edited Result - FINAL Addendum ADDENDUM REPORT: 6/21/2022 4:36 PM ADDENDUM: There is a pathologic fracture without significant displacement at a lytic lesion of the right proximal femur lesser trochanter. THIS IS AN ELECTRONICALLY VERIFIED FINAL REPORT 6/21/2022 4:36 PM DATE OF SERVICE: 6/21/2022 4:10 pm EXAM: Pelvis single view Comparison: CT pelvis 3/26/2022 Clinical history: Bone Mets, right femur pain. Findings: No acute fracture or dislocation. Multiple osseous sclerotic and a few lytic lesions are again noted, and there appears to be a couple new small foci of lytic lesions at the proximal right femur. Pelvic phleboliths are present. The report indicates that the patient has a pathologic fracture without significant displacement at a lytic lesion of the lesser trochanter of the right proximal femur. This suggests that the patient has a bone metastasis that has weakened the bone structure, leading to a fracture. The presence of lytic lesions, which are areas of bone destruction, and sclerotic lesions, which are areas of bone formation, indicates that the patient has a bone metastasis. The fact that the fracture is not significantly displaced suggests that it is a pathologic fracture, which is a fracture that occurs due to a underlying bone disease, rather than a traumatic fracture. The patient's clinical history of bone Mets (metastases) and right femur pain also supports the need for referral to the Orthopedic Oncology department. The presence of multiple osseous sclerotic and lytic lesions, as well as new small foci of lytic lesions at the proximal right femur, suggests that the patient's bone metastasis is progressive and requires further evaluation and management. Therefore, the specific terminology or findings that suggest the patient should be referred to the Orthopedic Oncology department due to risk of pathological fracture (emergency) are: * Pathologic fracture without significant displacement at a lytic lesion of the lesser trochanter of the right proximal femur * Presence of lytic lesions and sclerotic lesions indicating bone metastasis * Clinical history of bone Mets and right femur pain * Presence of multiple osseous sclerotic and lytic lesions, as well as new small foci of lytic lesions at the proximal right femur. These findings suggest that the patient requires urgent evaluation and management by the Orthopedic Oncology department to address the risk of pathological fracture and to manage the underlying bone metastasis. The complete list of reports explained is available as supplementary material. In our analysis, an expert orthopedic doctor provided 74 unique descriptive terms, while the LLM produced 199 unique terms after careful filtering. The Venn diagram in Fig. 4 illustrates the overlap between these term sets, showing 13 terms that were common to both the expert and the LLM: 'enhancing lesion', 'periosteal reaction', 'aggressive appearing', 'cortical disruption', 'cortical destruction', 'expansile', 'cortical erosion', 'aggressive', 'cortical thinning', 't1 hypointense t2 hyperintense', 'endosteal scalloping', 'lytic lesion', 'cortical breakthrough'. Some examples for clinician-only terms were 'predisposed to pathologic fracture', 'lytic change', 'increased risk of pathologic fracture', 'hypermetabolic osseous disease', 'cortical transgression', 'concerning for instability', 'hypointense', 'mixed lytic', and 'increased uptake in bone'. In contrast LLM specific terms were more lexically longer like 'localized pathologic fracture', 'sclerotic lesion', 'expansile lytic lesion involving the left anterior 5th rib', 'fracture', 'proximal femur', 'proximal right femur', 'presence of lytic lesion and sclerotic lesion indicating bone metastasis', 'possibility of a primary osseous neoplasm or lymphoma cannot be excluded', 'mildly displaced pathologic fracture', 'indeterminate lucent lesion'. To evaluate the semantic alignment between the LLM's generated explanations and standard clinical terminology, we performed a cosine similarity analysis on keywords and phrases that were unique to either the LLM's output ("LLM-Only") or the clinician's lexicon ("Clinician-Only") demonstrated in Fig. 5 . The results, presented in a heatmap, reveal a high degree of semantic congruence between the two vocabularies, even when the exact phrasing differed. Multiple term-pairs exhibited strong conceptual overlap with similarity scores exceeding 0.90. For instance, the LLM-generated phrase "permeation through the lateral cortex" was found to be almost identical in meaning to the clinician's term "transcortical permeation" (similarity = 0.947). Similarly, the LLM's more specific "femoral head" and "femoral neck" showed very high similarity (0.944 and 0.942, respectively) to the broader clinician term "femur." We also observed that the LLM's descriptive phrase "lytic metastatic bone lesion" was a strong semantic match to the concise clinical term "metastasis" (similarity = 0.900), confirming the model's ability to capture clinical meaning accurately. Another pair from LLM being "soft tissue mass" with the clinician's "marrow infiltration" at a score of 0.85 demonstrates a clear example of a lexically different but contextually similar pair. This heatmap shows that the LLM can generate lexically diverse but clinically relevant and semantically equivalent terms when compared to an expert clinician's vocabulary. 4. Discussion Prior work from our group has demonstrated the growing role of artificial intelligence in orthopedics and oncology. In orthopedics, we applied machine learning to model fracture recovery using gait analysis, offering predictive insights into patient outcomes following lower extremity fractures. Beyond orthopedics, we advanced interpretable AI frameworks for oncology, including 28 , 29 deep learning models for tumor classification and cancer recurrence prediction 30 , 31 . We also developed multimodal language–vision assistants in pathology, highlighting the feasibility of adapting LLMs for domain-specific interpretation 32 . Collectively, these contributions provide a foundation for the present work, which builds on this trajectory by addressing delayed referrals in metastatic bone disease with both accurate classification and explainable AI-driven explanations. In response to the critical issue of delayed referrals for metastatic bone disease, this study successfully developed and validated a dual-component AI framework of SHIELD designed to both classify and explain findings in radiology reports. Complementing its predictive power, the explainability component of Llama-3.1 model demonstrated its ability to generate clinically coherent rationales. SHIELD’s classification model demonstrates exceptional utility as a clinical decision support tool, primarily by automating the initial screening of radiology reports with a high degree of reliability. The model's near-perfect performance in identifying "No Referral" cases (F1-score of 99.20%) is a key finding, as it signifies a robust capability to reduce clinician workload by safely filtering out non-urgent reports. This strength is further underscored by the model's flawless binary classification, achieving a perfect AUC of 1.00 when distinguishing between reports that require a referral and those that do not. This near-perfect accuracy in the primary screening task suggests that the model can be trusted to ensure that no patient requiring follow-up is overlooked, a critical benchmark for implementation in a clinical workflow. The more nuanced challenge lies in differentiating the urgency between a standard "Referral" and a "Referral/High Risk" case. Here, the model exhibits a clinically meaningful trade-off: its high precision (94.05%) in the "Referral/High Risk" category ensures that when it flags a case as high-risk, the prediction is highly trustworthy. However, its moderate recall (57.78%) for this same category indicates a tendency to classify some high-risk cases into the standard referral pool. Critically, the confusion matrix reveals that these errors are "fail-safe"—no high-risk case was ever misclassified as "No Referral." This behaviour is paramount for patient safety; while the model may occasionally underestimate the level of urgency, it correctly identifies every at-risk patient as needing clinical attention. The high AUC of 0.93 for this sub-classification task further suggests that the model possesses strong discriminative power, and future work would focus on tuning the classification threshold by increasing number of training samples to optimize the balance between precision and recall based on specific clinical needs. The stark contrast in referral timelines highlights the most profound clinical implication of the SHIELD framework: its potential to dismantle the protracted and inefficient nature of the current diagnostic pathway. Clinical triaging with an average delay of nearly 110 days from initial imaging to referral is not merely slow but is characterized by systemic inefficiencies. Our findings show that over half of patients required multiple imaging visits and more than a third experienced referral delays exceeding a month, underscoring a period of clinical uncertainty and patient anxiety. In direct contrast, the AI model’s ability to generate a reliable classification in minutes from the very first imaging report represents a paradigm shift. It obviates the need for a prolonged "watch-and-wait" period and multiple follow-up scans, which are currently the primary drivers of these delays. By intervening at the earliest possible point, the framework transforms a reactive, multi-step process into a proactive and immediate decision point. This can potentially ensure the journey from initial suspicion to specialist consultation begins without delay. This investigation demonstrates the framework’s robust capability to process and reinterpret specialized medical language for clinical applications. The key finding is the significant disparity between the low lexical overlap and the high semantic similarity when comparing the LLM's output to an expert clinician's terminology. This suggests that the LLM is not merely extracting keywords but is generating a conceptually aligned, and often more descriptive, vocabulary to explain complex radiological findings. The model's ability to produce lexically diverse yet semantically equivalent terms underscores its potential as a sophisticated tool for clinical communication. By translating dense medical reports into structured and simplified explanations, as exemplified in the case study, the model can serve as a valuable aid in enhancing patient understanding and supporting clinical decision-making. Our work contributes to the growing body of research on AI-driven clinical triage, sharing the overarching goal of improving upon manual, often inefficient processes, as seen in both emergency 33 and specialist settings. While our framework addresses a specialized referral pathway similar to recent studies in gatekeeping and triage, its classification performance notably surpasses these prior efforts. For instance, the AI gatekeeping model by Vergara et al. 12 achieved an overall accuracy of 71.6% (AUC of 0.765) in a binary referral authorization task across multiple specialties. Similarly, the work by Maarseveen et al. 34 reported an AUC of 0.78 for prioritizing rheumatology referrals, and the system by Abdel-Hafez et al. 35 showed a 53.8% level of agreement for categorizing ENT referrals. A directly comparable study was recently introduced by Sangwon et al. 36 , who evaluated three NLP models for automating bone metastases referrals: a rule-based Regular Expression (RegEx) model, GPT-4, and a specialized BERT model (NYUTron). Their findings indicated that the RegEx model often outperformed the more complex LLMs for their specific clinical application, achieving an F1-score of 88.9% during validation. While this work represents a valuable contribution to AI-driven triage, it has notable limitations, such as the lack of a clinically validated timeline comparison and the use of a dataset covering a shorter time frame. In contrast to these studies, our SHIELD framework demonstrates a significant advance in performance, achieving 100% accuracy in the binary Referral vs. No Referral task and an excellent AUC of 0.93 in the more complex three-class problem. Furthermore, a key architectural advantage of our framework is its operational autonomy; following the initial data labeling phase, the model operates independently without requiring real-time clinician interrogation or rule-based adjustments for its predictions. This performance advantage likely stems from our model's unique design, which enables it to interpret the dense, unstructured narrative of a full radiology report for a specialized oncological pathway. However, the most significant departure remains our framework's integral dual-component architecture. Where other models focus primarily on the classification task, our work pairs its high accuracy with a sophisticated explainability component using a large language model. This distinction is crucial 37 , as our goal is not only to accelerate a workflow but also to build clinician trust and facilitate communication for a considered specialist referral—a focus on interpretability that moves beyond what is presented in related literature. This study has several limitations. Our dataset consisted of 1,404 EHRs collected over ten years. Furthermore, the scope was restricted to malignancies of the long bones (femur, fibula, tibia, humerus, radius, and ulna). The model's performance on the 'Referral/High Risk' class was impacted by a relatively small number of samples in this category, which presented a challenge for achieving optimal classification accuracy. Future work will aim to address these limitations. We plan to expand the dataset by including a larger number of EHRs and incorporating all sites of skeletal metastasis. Additionally, integrating multimodality by adding image analysis could further enhance the model's performance, making it an even more reliable tool for medical experts 38 . Collectively, these findings validate the framework's potential to significantly accelerate the patient referral pathway while fostering trust and adoption among clinicians through transparent and meaningful explanations. 5. Conclusion Given the aging population and increasing cancer incidence, there is a critical and growing need to identify patients at risk for metastatic bone disease and skeletal-related events earlier in their disease course. The promising results of this study suggest our proposed SHIELD framework can directly address this challenge by ameliorating the problem of late referrals. The framework facilitates the early recognition of high-risk radiographic findings, particularly when explicit descriptors, such as “lytic lesion” or “impending fracture,” are present. Its adoption can lead to significantly improved survival and functional outcomes by enabling healthcare professionals to treat patients prophylactically. Ultimately, leveraging AI to analyze language patterns within radiology reports provides a scalable and effective pathway to ensure more accurate and timely referrals to orthopedic oncology, along with reliable explanations. Declarations Participant Consent: The study used de-identified retrospective data and did not involve direct participant contact. The need for individual consent was waived by the Institutional Review Board (IRB) of Wake Forest School of Medicine (Approval ID: IRB00127840) in accordance with ethical guidelines and regulatory standards. Author Contribution: C.L.E. and M.N.G. conceptualized the study. C.L.E. identified the clinical need, proposed the original project idea, and provided clinical oversight. Under the supervision of C.L.E., A.M.M. was responsible for the acquisition, de-identification, and curation of the radiology report dataset from the participating medical centers. A.A. designed the computational methodology, developed and trained the SHIELD framework, conducted all computational experiments, and performed the statistical and retrospective timeline analyses under the direct supervision of M.N.G. F.M.D. provided medical consultation, aiding in the definition of the risk-stratification tiers and the clinical validity of the model’s outputs. A.A. and M.N.G. wrote the original draft of the manuscript. All authors participated in the critical review, editing, and final approval of the manuscript and approved the journal submission. Acknowledgements: The authors would like to thank Robert B. Lipsit, for his contributions to patient query, exclusion, and data collection. Funding: Ava M. McKane was part of the Medical Student Research Program sponsored by the Wake Forest University School of Medicine and Department of Orthopaedics. The project described was supported in part by R01 DC020715 (PIs: Gurcan, Moberly) from the National Institute on Deafness and Other Communication Disorders, R21 CA273665 (PI: Gurcan) from the National Cancer Institute, U01 TR003629 (PIs: Gurcan, Paige, Steube) from the National Center for Advancing Translational Sciences, and R01HL177046 (PIs: Gurcan, Michelson, Hachem) from the National Heart, Lung, and Blood Institute. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health, National Institute on Deafness and Other Communication Disorders, National Cancer Institute, National Heart, Lung, and Blood Institute, or National Center for Advancing Translational Sciences. Data Availability: The dataset from this study is held securely in coded form Wake Forest University School of Medicine. The full dataset creation plan are available from the authors upon request. The source code for implementing the methods is available at the following repository: https://github.com/CAIR-LAB-WFUSM/SHIELD.git. Human Ethics and Consent to Participate: This project was approved by the Wake Forest University School of Medicine Research Ethics Board Protocol review board approval (IRB00127840) in accordance with the Declaration of Helsinki. All the records in this study were de-identified. Clinical trial number: not applicable. Conflict of Interest: The authors have declared that no competing interests exist. Competing interests: The authors declare no competing interests. References DiCaprio MR, Murtaza H, Palmer B, Evangelist M. Narrative review of the epidemiology, economic burden, and societal impact of metastatic bone disease. Ann Jt . AME Publishing Company . 2022;7. doi:10.21037/aoj-20-97 Canavan ME, Wang X, Ascha MS, et al. Systemic Anticancer Therapy and Overall Survival in Patients with Very Advanced Solid Tumors. JAMA Oncol . 2024;10(7):887-895. doi:10.1001/jamaoncol.2024.1129 Svensson E, Christiansen CF, Ulrichsen SP, Rørth MR, Sørensen HT. Survival after bone metastasis by primary cancer type: A Danish population-based cohort study. BMJ Open . 2017;7(9). doi:10.1136/bmjopen-2017-016022 Mosher ZA, Patel H, Ewing MA, et al. Early Clinical and Economic Outcomes of Prophylactic and Acute Pathologic Fracture Treatment .; 2025. https://doi.org/10. Ardakani AHG, Faimali M, Nystrom L, et al. Metastatic bone disease: Early referral for multidisciplinary care. Cleve Clin J Med . 2022;89(7):393-399. doi:10.3949/ccjm.89a.21062 Levin A. Consequences of late referral on patient outcomes. Nephrology Dialysis Transplantation . 2000;15(suppl_3):8-13. doi:10.1093/oxfordjournals.ndt.a027977 Kotrych D, Ciechanowicz D, Pawlik J, Szyjkowska M, Kwapisz B, Mądry M. Delay in Diagnosis and Treatment of Primary Bone Tumors during COVID-19 Pandemic in Poland. Cancers (Basel) . 2022;14(24). doi:10.3390/cancers14246037 Jawad MU, Scully SP. In brief: Classifications in brief: Mirels’ classification: Metastatic disease in long bones and impending pathologic fracture. Clin Orthop Relat Res . Springer New York LLC . 2010;468(10):2825-2827. doi:10.1007/s11999-010-1326-4 Kimura T. Multidisciplinary approach for bone metastasis: A review. Cancers (Basel) . MDPI AG . 2018;10(6). doi:10.3390/cancers10060156 Breden S, Hinterwimmer F, Consalvo S, et al. Deep Learning-Based Detection of Bone Tumors around the Knee in X-rays of Children. J Clin Med . 2023;12(18). doi:10.3390/jcm12185960 Rizk PA, Gonzalez MR, Galoaa BM, et al. Machine Learning–Assisted Decision Making in Orthopaedic Oncology. JBJS Rev . 2024;12(7). doi:10.2106/JBJS.RVW.24.00057 Vergara PO, Oliveira JDC, Mattiello R, et al. Accuracy of Artificial Intelligence for Gatekeeping in Referrals to Specialized Care. JAMA Netw Open . Published online 2025. doi:10.1001/jamanetworkopen.2025.13285 Fudickar S, Bantel C, Spieker J, et al. Natural Language Processing of Referral Letters for Machine Learning–Based Triaging of Patients With Low Back Pain to the Most Appropriate Intervention: Retrospective Study. J Med Internet Res . 2024;26(1). doi:10.2196/46857 Farrow L, Zhong M, Anderson L. Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports. Bone Joint J . 2024;106(7):688-695. doi:10.1302/0301-620X.106B7 Hu Y, Xiang Y, Zhou YJ, et al. AI-based diagnosis of acute aortic syndrome from noncontrast CT. Nat Med . Published online 2025. doi:10.1038/s41591-025-03916-z Zhao Z, Zhang Y, Wu C, et al. Large-vocabulary segmentation for medical images with text prompts. NPJ Digit Med . 2025;8(1):566. doi:10.1038/s41746-025-01964-w Fogel AL, Kvedar JC. Artificial intelligence powers digital medicine. NPJ Digit Med . Nature Publishing Group . 2018;1(1). doi:10.1038/s41746-017-0012-2 Lee J, Yoon W, Kim S, et al. BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics . 2020;36(4):1234-1240. doi:10.1093/bioinformatics/btz682 Gu Y, Tinn R, Cheng H, et al. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Trans Comput Healthc . 2022;3(1). doi:10.1145/3458754 Yan A, McAuley J, Lu X, et al. RadBERT: Adapting Transformer-based Language Models to Radiology. Radiol Artif Intell . 2022;4(4). doi:10.1148/ryai.210258 Rosenbacke R, Melhus Å, McKee M, Stuckler D. How Explainable Artificial Intelligence Can Increase or Decrease Clinicians’ Trust in AI Applications in Health Care: Systematic Review. JMIR AI . JMIR Publications Inc. 2024;3. doi:10.2196/53207 Sadeghi Z, Alizadehsani R, CIFCI MA, et al. A review of Explainable Artificial Intelligence in healthcare. Computers and Electrical Engineering . 2024;118. doi:10.1016/j.compeleceng.2024.109370 Aziz NA, Manzoor A, Mazhar Qureshi MD, Qureshi MA, Rashwan W. Unveiling Explainable AI in Healthcare: Current Trends, Challenges, and Future Directions. Preprint posted online August 10, 2024. doi:10.1101/2024.08.10.24311735 Kibria MG, Kucirka L, Mostafa J. Assessing AI Explainability: A Usability Study Using a Novel Framework Involving Clinicians. In: Institute of Electrical and Electronics Engineers (IEEE); 2025:553-564. doi:10.1109/ichi64645.2025.00069 Grattafiori A, Dubey A, Jauhri A, et al. The Llama 3 Herd of Models. Published online November 23, 2024. http://arxiv.org/abs/2407.21783 Ren Y, Wu D, Khurana A, et al. Classification of Patient Portal Messages with BERT-based Language Models. In: Proceedings - 2023 IEEE 11th International Conference on Healthcare Informatics, ICHI 2023 . Institute of Electrical and Electronics Engineers Inc.; 2023:176-182. doi:10.1109/ICHI57859.2023.00033 Kroll H, Sackhoff P, Thang BM, Ksouri M, Balke WT. A Library Perspective on Supervised Text Processing in Digital Libraries: An Investigation in the Biomedical Domain. In: Proceedings of the ACM/IEEE Joint Conference on Digital Libraries . Institute of Electrical and Electronics Engineers Inc.; 2025. doi:10.1145/3677389.3702557 Rezapour M, Seymour RB, Medda S, et al. Analyzing Gait Dynamics and Recovery Trajectory in Lower Extremity Fractures Using Linear Mixed Models and Gait Analysis Variables. Bioengineering . 2025;12(1). doi:10.3390/bioengineering12010067 Rezapour M, Seymour RB, Sims SH, Karunakar MA, Habet N, Gurcan MN. Employing machine learning to enhance fracture recovery insights through gait analysis. Journal of Orthopaedic Research . 2024;42(8):1748-1761. doi:10.1002/jor.25837 Su Z, Guo Y, Wesolowski R, et al. Computational Pathology for Accurate Prediction of Breast Cancer Recurrence: Development and Validation of a Deep Learning-based Tool. Modern Pathology . Published online December 2025:100847. doi:10.1016/j.modpat.2025.100847 Tavolara TE, Su Z, Gurcan MN, Niazi MKK. One label is all you need: Interpretable AI-enhanced histopathology for oncology. Semin Cancer Biol . Academic Press . 2023;97:70-85. doi:10.1016/j.semcancer.2023.09.006 Afzaal U, Su Z, Sajjad U, et al. HistoChat: Instruction-tuning multimodal vision language assistant for colorectal histopathology on limited data. Patterns . Published online August 8, 2025. doi:10.1016/j.patter.2025.101284 Tyler S, Olis M, Aust N, et al. Use of Artificial Intelligence in Triage in Hospital Emergency Departments: A Scoping Review. Cureus . Published online May 8, 2024. doi:10.7759/cureus.59906 Maarseveen TD, Glas HK, Veris-van Dieren J, van den Akker E, Knevel R. Improving musculoskeletal care with AI enhanced triage through data driven screening of referral letters. NPJ Digit Med . 2025;8(1). doi:10.1038/s41746-025-01495-4 Abdel-Hafez A, Jones M, Ebrahimabadi M, et al. Artificial intelligence in medical referrals triage based on Clinical Prioritization Criteria. Front Digit Health . 2023;5. doi:10.3389/fdgth.2023.1192975 Sangwon KL, Han X, Becker A, et al. Automating the Referral of Bone Metastases Patients With and Without the Use of Large Language Models. Neurosurgery . Published online 2025. doi:10.1227/neu.0000000000003683 Bienefeld N, Boss JM, Lüthy R, et al. Solving the explainable AI conundrum by bridging clinicians’ needs and developers’ goals. NPJ Digit Med . 2023;6(1). doi:10.1038/s41746-023-00837-4 Huang J, Wittbrodt MT, Teague CN, et al. Efficiency and Quality of Generative AI-Assisted Radiograph Reporting. JAMA Netw Open . 2025;8(6). doi:10.1001/jamanetworkopen.2025.13921 Additional Declarations No competing interests reported. Supplementary Files Supplementarymaterial.docx Cite Share Download PDF Status: Under Review Version 1 posted Editor assigned by journal 06 Nov, 2025 Submission checks completed at journal 05 Nov, 2025 First submitted to journal 22 Oct, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7926103","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":540748600,"identity":"bee248a5-4c82-4514-ab80-35ba36ed3fd8","order_by":0,"name":"Abbas Alili","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABMklEQVRIie3QvWqDUBTA8SsXrsuxrrco5hUUoUtD+iqVC5kc2iUIKUQJXJekXQN9iTxCRWiW1K6GZmhwzRK62E/qx2hMOxZ6/4tw8Me5HIREor8YheqjE1kKnhGph7icIHqYAMHS2KwJrgjAT6T4jdBfEfV2Gu1y1IMjjLmXD2JDDR+j7MLrwZnmS5n31lyyTpgGiJUP46tJEtt0ybA9WzIA/Q5bD9fNNalrasVbKpIqvO/MESOawnHxsHNyHEwaopO69muORhW5/CyJmsnvyteolZipe1JcIK4IVnjXmVNGsOLHNfHzBrHWSf8UzEV1ZE1PujZNM1uD+wVA6oytwG8Q42kar3JvaHTCcLPbDqih3jibF7gaGvKMRRv/Y/+lkbl/LPlI4i3kQG1bRCKR6B/1DfaJXWNubD6yAAAAAElFTkSuQmCC","orcid":"","institution":"Wake Forest University School of Medicine","correspondingAuthor":true,"prefix":"","firstName":"Abbas","middleName":"","lastName":"Alili","suffix":""},{"id":540748601,"identity":"667cf0c7-515f-4505-b611-30945ee7dbaa","order_by":1,"name":"Ava M. McKane","email":"","orcid":"","institution":"Wake Forest University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Ava","middleName":"M.","lastName":"McKane","suffix":""},{"id":540748602,"identity":"86f88e3d-9d5f-4f13-8a0d-958bdcac5ca6","order_by":2,"name":"Fatih M. Demir","email":"","orcid":"","institution":"Wake Forest University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Fatih","middleName":"M.","lastName":"Demir","suffix":""},{"id":540748603,"identity":"de9b606e-ab1d-49b5-8626-86a479ec0962","order_by":3,"name":"Cynthia L. Emory","email":"","orcid":"","institution":"Wake Forest University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Cynthia","middleName":"L.","lastName":"Emory","suffix":""},{"id":540748604,"identity":"4bd12a5b-0587-4789-a423-173121405e15","order_by":4,"name":"Metin N. Gurcan","email":"","orcid":"","institution":"Wake Forest University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Metin","middleName":"N.","lastName":"Gurcan","suffix":""}],"badges":[],"createdAt":"2025-10-22 18:38:11","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7926103/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7926103/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":95662228,"identity":"fce037a7-df49-48cd-89dc-9228347a914b","added_by":"auto","created_at":"2025-11-11 16:37:17","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":13628776,"visible":true,"origin":"","legend":"","description":"","filename":"SHIELDAnAIFrameworkforSkeletalHealthIntelligenceandEarlyLesionDetectiontoImproveOrthopedicReferrals.docx","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/ca17584c102e2827d7564652.docx"},{"id":95662156,"identity":"e9069c5c-63ef-46ee-965b-6ed58c5beb52","added_by":"auto","created_at":"2025-11-11 16:37:13","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7834,"visible":true,"origin":"","legend":"","description":"","filename":"b10d27b8df4f47169d6240cbd4b5cd9f.json","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/87b0377aef13bce436c43cd7.json"},{"id":95661992,"identity":"a035559b-aceb-4d5b-afcf-c56b20640710","added_by":"auto","created_at":"2025-11-11 16:37:03","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":45738,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementarymaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/d8738d841757994767527836.docx"},{"id":95661969,"identity":"5c6be5b2-da7c-4b34-be3d-a086bae9586f","added_by":"auto","created_at":"2025-11-11 16:37:02","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":130094,"visible":true,"origin":"","legend":"","description":"","filename":"b10d27b8df4f47169d6240cbd4b5cd9f1enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/52af71cbf3091af9fc05451b.xml"},{"id":95662291,"identity":"868e7b51-8856-4fac-8b2b-d521e0783692","added_by":"auto","created_at":"2025-11-11 16:37:20","extension":"jpeg","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":697572,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/0a3800e97497daf2a494337d.jpeg"},{"id":95662343,"identity":"9a71f519-9650-4397-8e05-7931184dd2ae","added_by":"auto","created_at":"2025-11-11 16:37:25","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8251721,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/6dd3edae540f40fec513a2b1.png"},{"id":95662288,"identity":"756b2e0b-76c9-4e9d-b071-dbfb8930198d","added_by":"auto","created_at":"2025-11-11 16:37:20","extension":"jpeg","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":374285,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/c0b619728708d50b35586460.jpeg"},{"id":95662181,"identity":"a9475f02-474b-496f-b1b4-07807b23dddd","added_by":"auto","created_at":"2025-11-11 16:37:15","extension":"jpeg","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1161426,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/68852724bbd5d359729aacf6.jpeg"},{"id":95662132,"identity":"c56cdd3b-3afa-4b7d-af7b-536701b7b354","added_by":"auto","created_at":"2025-11-11 16:37:12","extension":"jpeg","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":736844,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/cbd13f9586813f345aef2715.jpeg"},{"id":95662009,"identity":"521f19ad-662b-40c8-bee2-84127f2024a9","added_by":"auto","created_at":"2025-11-11 16:37:03","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":255954,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/d9597a80a4bc0a905b0f9cca.png"},{"id":95662000,"identity":"dca2c9dc-12e4-4001-822c-2bf7ff5be27f","added_by":"auto","created_at":"2025-11-11 16:37:03","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":343910,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/ff4ae79a74cf6cdd86e6ce75.png"},{"id":95662062,"identity":"6753f98b-948a-4c70-a771-7befa3b97d4b","added_by":"auto","created_at":"2025-11-11 16:37:08","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":447874,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/b2f25cd5157488a84e8124c7.png"},{"id":95662020,"identity":"8cd02b08-cc05-4403-ad08-c7c7d4ce7485","added_by":"auto","created_at":"2025-11-11 16:37:04","extension":"jpeg","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1074,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage8.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/f6a44745feffd4f90a2e93a8.jpeg"},{"id":95662283,"identity":"323a1a22-3ad1-4d31-8d0b-7a0693a34203","added_by":"auto","created_at":"2025-11-11 16:37:19","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":246480,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/540ba9d84eabc585405e7cab.png"},{"id":95662354,"identity":"a8df64eb-b788-4cbf-8ee0-fde9c7af0704","added_by":"auto","created_at":"2025-11-11 16:37:26","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":100393,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/e576168757aec48d389c35fa.png"},{"id":95661849,"identity":"d64e8aed-aefe-450e-8191-c7ca7d05a213","added_by":"auto","created_at":"2025-11-11 16:36:55","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1810019,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/104062439c4ee798a9068978.png"},{"id":95662030,"identity":"116d658e-ff8a-4057-8086-91f7a9befb0d","added_by":"auto","created_at":"2025-11-11 16:37:05","extension":"png","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":73092,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/79e86bfed0c6b6dccab9ba2e.png"},{"id":95662320,"identity":"0832ae0d-13ed-4841-97e5-08f7bb3082f7","added_by":"auto","created_at":"2025-11-11 16:37:23","extension":"png","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":139919,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/bb6adc177d2be865e1346a01.png"},{"id":95662297,"identity":"c36f37b2-daf5-4caf-ba9d-4250c241487f","added_by":"auto","created_at":"2025-11-11 16:37:21","extension":"png","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":94843,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/2d172972955a48c2bc01d07c.png"},{"id":95661847,"identity":"fd31862b-565c-41ab-a028-b69062dbfeca","added_by":"auto","created_at":"2025-11-11 16:36:55","extension":"png","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":113858,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/e8528976efecb1926c70f502.png"},{"id":95662178,"identity":"2e05ce95-79b7-4038-877d-87c3e8315f95","added_by":"auto","created_at":"2025-11-11 16:37:15","extension":"png","order_by":21,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":144230,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/d1aa7a9cf207f0fb4211a4ee.png"},{"id":95662002,"identity":"f1858e12-e1bf-4e79-924e-c8f90c80fd6b","added_by":"auto","created_at":"2025-11-11 16:37:03","extension":"png","order_by":22,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":178976,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/16fa6ab6b5dc7f55074c855b.png"},{"id":95662172,"identity":"e166c5a5-e731-482c-8822-9974dfe141d2","added_by":"auto","created_at":"2025-11-11 16:37:15","extension":"png","order_by":23,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":935,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/32dde52b4c523c80d363351a.png"},{"id":95662315,"identity":"31379582-e7c6-4ec9-a185-2fb0486f0587","added_by":"auto","created_at":"2025-11-11 16:37:22","extension":"png","order_by":24,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":78764,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/d8958f281389f48995d636f1.png"},{"id":95662274,"identity":"918f5b41-ad5f-4d4d-af41-dffdf1ba54ce","added_by":"auto","created_at":"2025-11-11 16:37:19","extension":"xml","order_by":25,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":127980,"visible":true,"origin":"","legend":"","description":"","filename":"b10d27b8df4f47169d6240cbd4b5cd9f1structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/8f8180ec84a6d7f29b30df4f.xml"},{"id":95662298,"identity":"38dd0bb0-b9ab-4d5c-b6dc-88710892d95e","added_by":"auto","created_at":"2025-11-11 16:37:21","extension":"html","order_by":26,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":139740,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/71992d10cdbcc2eca43301d2.html"},{"id":95662324,"identity":"df9583d9-8300-4c2d-b09b-b7c426fe904a","added_by":"auto","created_at":"2025-11-11 16:37:23","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1462049,"visible":true,"origin":"","legend":"\u003cp\u003eFlow diagram of radiology report selection and preprocessing for SHIELD-based referral classification analysis.\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/af066a7b107301addb1f739b.jpg"},{"id":95662045,"identity":"b7c51712-7bee-4b33-b129-a73c184768b4","added_by":"auto","created_at":"2025-11-11 16:37:06","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1617496,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 1. Schematic of the SHIELD’s AI pipeline for automated classification of radiology reports. \u003c/strong\u003eThe workflow begins with the extraction of radiology reports, which are then reviewed and labeled by clinicians. Following data preprocessing, the text is tokenized, and a sliding window technique is applied to create manageable input sequences. These sequences are then used to fine-tune the RadBERT-RoBERTa-4m model. Finally, the model's classification layer categorizes each report into one of three classes: \"No Referral,\" \"Referral,\" or \"Referral/High Risk.\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/2d937bca5aac2c550c0407eb.jpg"},{"id":95662388,"identity":"020a24d8-5b72-454d-966f-a1b80c64ff16","added_by":"auto","created_at":"2025-11-11 16:37:27","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":866233,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 2. The explainable AI architecture of the SHIELD framework. \u003c/strong\u003eA new radiology report, along with its classification label (e.g., \"Referral/High Risk\") as determined by the classification model, is provided as input to the LLAMA 3.1 8B Instruct model. The LLM then processes this information to generate a natural-language explanation, outlining the key findings and clinical reasoning that led to the initial classification, thereby providing transparency for medical experts.\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/1f172262eb5e562faa8fab85.jpg"},{"id":95662137,"identity":"2db2de50-38ac-48d5-af8a-ba76c79829e0","added_by":"auto","created_at":"2025-11-11 16:37:13","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":83639,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 3.\u003c/strong\u003e \u003cstrong\u003eGranular Analysis of Classification Performance using Confusion Matrices and ROC curves.\u003c/strong\u003e The figure displays the prediction accuracy of the ensemble model on the hold-out test data. (a) The binary confusion matrix demonstrates perfect performance in the primary task of identifying the need for a referral. (b) The three-class confusion matrix provides a detailed view, showing that misclassifications are confined to the referral subtypes, where some \"Referral/High Risk\" cases were classified as standard \"Referral.\" (c) ROC curves for the ensemble model evaluated on the hold-out test set. The blue line represents the perfect separation (AUC = 1.00) between \"No Referral\" and all referral categories combined. The red line shows excellent discrimination (AUC = 0.93) when distinguishing between \"Referral\" and \"Referral/High Risk\" reports.\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/c6c1af495362ebe0d9a8b216.jpg"},{"id":95662006,"identity":"083b87f3-3701-4f7b-9f88-68db3f1dedad","added_by":"auto","created_at":"2025-11-11 16:37:03","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":246480,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 4:\u003c/strong\u003e Venn diagram illustrates the overlap between LLM and clinician selected term sets.\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/53c4eef59ddf6fbb5a07a4b2.png"},{"id":95661954,"identity":"79136e4e-f3d8-4421-afc5-86122393993c","added_by":"auto","created_at":"2025-11-11 16:37:02","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":529464,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 5.\u003c/strong\u003e Heatmap showing 30 demonstrative pairs for LLM and Clinician generated terms. By using BioBERT to analyze conceptual relationships, this visualization effectively identifies example pairs of terms that are semantically aligned despite differences in their exact wording. The colored squares indicate the cosine similarity score for a given pair, with warmer, red tones representing higher similarity and cooler, blue tones representing moderate similarity.\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/f187f06a4d2d2e61c136d6a9.png"},{"id":95663496,"identity":"4304762a-5246-4698-a17d-d03eb977bc35","added_by":"auto","created_at":"2025-11-11 16:39:00","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5788489,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/5168508d-924f-42c2-b664-671a53f29eba.pdf"},{"id":95661867,"identity":"9a243d2a-62aa-43b3-914b-1f9c9fda24b9","added_by":"auto","created_at":"2025-11-11 16:36:57","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":45738,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementarymaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-7926103/v1/6faf81af89df71f215b0680c.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"SHIELD: An AI Framework for Skeletal Health Intelligence and Early Lesion Detection to Improve Orthopedic Referrals","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eMetastatic bone disease represents a significant complication arising from advanced malignancies. In the United States, it is estimated that approximately 300,000 adults are living with MBD\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. Although advancements in systemic therapies for various cancers have improved patient longevity, the prognosis after developing bone metastases is variable, with a wide range of survival rates depending on the primary type of cancer\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. For example, one-year survival after a bone metastasis diagnosis can range from as high as 51% for breast cancer to as low as 10% for lung cancer\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. The management of bone metastases also has significant implications for healthcare expenditures. Early intervention for patients with MBD has been demonstrated to reduce patient morbidity and overall healthcare costs, which are estimated to exceed \u003cspan\u003e$\u003c/span\u003e12.6\u0026nbsp;billion annually in the United States\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eDespite the urgent nature of MBD, many patients are not referred to an orthopedic provider until after a pathologic fracture has occurred, leading to a missed opportunity for prophylactic stabilization\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. A patient\u0026rsquo;s initial diagnosis of cancer or MBD can be overwhelming for patients and caregivers, with multiple competing priorities occurring at the same time. Knowledge gaps from referring providers on the acuity of impending fractures based upon location, size, or extent of a bone lesion can lead to disparate referral patterns. Access to musculoskeletal experts can also lead to delays in referrals to the treating surgeon. Other sites of metastasis can influence prioritization of referrals and subspecialty appointments, and unintended delays in orthopaedic referrals may occur as a result. This delay underscores the need for improved and standardized referral pathways to orthopedic oncology to enable timely intervention and optimize patient outcomes\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. While Mirels' score\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e is a widely used tool within orthopedics to assess fracture risk based on radiographic and clinical features, its adoption outside of the specialty is limited. Furthermore, its low specificity (35%) creates a risk of overtreatment\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Referrals to orthopedic specialists are often based on subjective clinical judgment, non-standardized imaging terminology, or non-specific symptoms such as pain, leaving the interpretation and decision to refer largely at the discretion of the non-specialist provider.\u003c/p\u003e\u003cp\u003eRecent advances in artificial intelligence (AI) offer promising solutions to expedite the referral process and reduce the workload of medical experts\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e. For instance, Vergara et al.\u003csup\u003e12\u003c/sup\u003e developed an AI model for gatekeeping referrals from primary to specialized care, which achieved moderate accuracy in distinguishing between authorized referrals, outperforming human gatekeepers by nearly 20% primarily due to higher specificity. Another study demonstrated that incorporating natural language processing (NLP)-extracted qualitative data from referral letters significantly increases the accuracy of machine learning (ML) models by up to 19.5% for triaging patients with low back pain to appropriate interventions, despite the overall model accuracy remaining low for clinical application\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. The ARCHERY project also utilized NLP with a clinically based large language model (LLM) on free-text radiology reports to predict patient selection for total hip and knee arthroplasties, demonstrating promising potential for hip arthroplasty but not for knee arthroplasty. Their work highlighted the importance of further model testing and training for new clinical cohorts\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eLLMs being at the forefront of AI have enabled the processing of vast data corpora for various downstream tasks such as classification, analysis, and prediction\u003csup\u003e\u003cspan additionalcitationids=\"CR16\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. The promise of these transformer-based architectures has spurred interest in exploring even larger models, often containing billions of parameters. In the biomedical domain, researchers have developed specialized transformer models like BioBERT\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e and PubMedBERT\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e (each comprising 110\u0026nbsp;million parameters) by training them on biomedical literature from PubMed. Furthermore, clinically specific language models such as RadBERT have been adapted for the field of radiology\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. These models have demonstrated strong performance in extracting relevant information from free-text radiology reports and have shown potential in predicting the need for surgical interventions based on the language used in imaging reports.\u003c/p\u003e\u003cp\u003eThe growing capabilities of AI in healthcare are accompanied by a corresponding scepticism regarding its adoption. A recent study investigated how explainable artificial intelligence (XAI) influences clinicians' trust in AI applications, addressing the challenge of fostering appropriate reliance on AI's often \"black box\" nature\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. The majority of studies included suggest that XAI has the potential to enhance clinicians' trust in AI recommendations. However, complex or contradictory explanations can undermine this trust, whereas excessive trust in incorrect AI advice can adversely impact clinical accuracy\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. This highlights a pressing need for XAI solutions that provide clear, understandable, and reliable explanations, especially due to ethical and regulatory concerns in healthcare\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eDespite advancements in the application of AI in healthcare and the associated challenges of explainability, very limited research has specifically addressed the problem of delayed referrals for metastatic bone disease. The overall goal of this study was to introduce \u003cb\u003eSHIELD\u003c/b\u003e (Skeletal Health Intelligence and Early Lesion Detection), an AI framework designed to prevent late referrals by providing highly accurate classifications and offering appropriate explainability to aid medical experts in their decision-making. Our study had three aims. The \u003cem\u003eprimary aim\u003c/em\u003e of this work was to 1) develop an AI framework to classify radiology reports into three distinct referral categories: no referral needed, referral recommended, and referral/high-risk. This was achieved by fine-tuning the RadBERT-RoBERTa-4m\u003csup\u003e20\u003c/sup\u003e model for MBD referral classification using a decade of data (January 2014 \u0026ndash; May 2025) from two affiliated academic medical centers. RadBERT-RoBERTa-4m\u003csup\u003e20\u003c/sup\u003e is a transformer-based language model developed for radiology and clinical NLP. A \u003cem\u003esecondary aim\u003c/em\u003e was to 2) provide explanations for the framework's classification decisions, thereby increasing trust and adoption among medical providers. We utilized the Llama-3.1-8B-Instruct, Meta\u0026rsquo;s instruction-tuned language model that was released in July 2024, to accomplish this task. Our \u003cem\u003efinal aim\u003c/em\u003e was to 3) validate the explainability of the proposed framework and assess the overlap between the terminology used by the AI and a predefined set of terms established by orthopedic experts.\u003c/p\u003e"},{"header":"2. Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1 Cohort Selection, Radiology Report Review, and Annotation by Clinical Experts\u003c/h2\u003e\u003cp\u003e Following institutional review board approval (IRB00127840), we conducted a retrospective review of all patients with a diagnosis of secondary neoplasm of bone treated at two affiliated academic medical centers over a 10-year period (January 2014 \u0026ndash; May 2025) by analyzing 1404 Electronic Health Records (EHR). Patients were identified using ICD codes corresponding to secondary malignant neoplasm of bone (C79.51). Clinical documentation and radiology reports from both plain and advanced imaging modalities (PET, CT, MRI, and radiographs) were reviewed.\u003c/p\u003e\u003cp\u003ePatients were included if they were \u0026ge;\u0026thinsp;18 years of age at the time of MBD diagnosis and had evidence of osseous metastases to the appendicular long bones, defined as the \u003cem\u003efemur\u003c/em\u003e, \u003cem\u003efibula\u003c/em\u003e, \u003cem\u003etibia\u003c/em\u003e, \u003cem\u003ehumerus\u003c/em\u003e, \u003cem\u003eradius\u003c/em\u003e, or \u003cem\u003eulna\u003c/em\u003e. Patients were excluded if they were referred to orthopedic oncology prior to evidence of long-bone metastases (LBM) or were never referred to orthopedic oncology, had their initial imaging dated before January 2014, had only axial or non-LBM, were under 18 years of age at diagnosis, or had incomplete or inaccessible imaging and medical records. From an initial cohort of 514 patients with secondary neoplasm of bone, 245 met full inclusion criteria. A control cohort was generated by screening adult patients with any cancer diagnosis over the same 10-year period who had undergone at least one imaging study (PET, CT, MRI, and radiograph) indicating cancer at either institution. Patients were excluded from this control cohort if they had any mention of osseous metastases in the appendicular skeleton, a diagnosis of primary bone tumor, were under 18 years old, or lacked complete medical documentation. Of the 890 screened patients, 245 met the inclusion criteria for the control group. The cohort selection process is summarized in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eData was de-identified and collected using REDCap, a secure, web-based platform. Patient demographics, such as age, sex, race/ethnicity, comorbidities, and primary cancer type, were recorded and listed in Table\u0026nbsp;1. For the MBD cohort, the timeline from first imaging mention of metastasis to referral to orthopedic oncology was documented, along with prior treatments including chemotherapy and radiation. Raw radiology reports were obtained for all relevant imaging studies from the date of the first long bone metastatic lesion to the date of referral. Reports were then analyzed for descriptive language using a list of predefined terms indicative of MBD or SREs. This list of terms was curated a priori by an orthopedic oncologist to avoid bias during data extraction. Specific lesion descriptors, signs of impending fracture, and evidence of progression were recorded.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eTable 1.\u003c/strong\u003e Participant Demographic Information Based on 245 Patients with Metastatic Bone Disease (MBD) and Complete Trained Healthcare Professional Data.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Taba\" border=\"1\"\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCharacteristic\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003ePatients, No. (%)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMean\u0026thinsp;=\u0026thinsp;62 years, (SD\u0026thinsp;=\u0026thinsp;11.5)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eSex\u003c/b\u003e\u003csup\u003e\u003cb\u003ea\u003c/b\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFemale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e138 (56.3%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e107 (43.7%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eRace\u003c/b\u003e\u003csup\u003e\u003cb\u003ea\u003c/b\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWhite\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e167 (61.2%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBlack or African American\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e43 (17.6%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAsian\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e3 (1.2%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eOther\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e11 (4.5%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eEthnicity\u003c/b\u003e\u003csup\u003e\u003cb\u003ea\u003c/b\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHispanic or Latino\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e8 (3.3%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNot Hispanic or Latino\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e237 (96.7%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eBody Mass Index\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026lt;25\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e77 (31.4%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e25.0\u0026ndash;30.0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e169 (69.0%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026gt;30\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e92 (37.6%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eComorbidities\u003c/b\u003e\u003csup\u003e\u003cb\u003eb\u003c/b\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCOPD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e16 (6.5%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDiabetes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e48 (19.6%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHypercholesterolemia\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e19 (7.8%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHyperlipidemia\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e53 (21.6%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHypertension\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e121 (49.4%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eOsteoporosis\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e10 (4.1%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRenal Insufficiency\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e12 (4.9%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNone\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e81 (33.1%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cp\u003e\u003cstrong\u003e\u003csup\u003ea\u003c/sup\u003e\u003c/strong\u003e\u003csup\u003e\u0026nbsp;Data on patient demographic characteristics, including age, sex, race, and ethnicity, were collected from the electronic health records at two academic institutions and classified by standardized categories as defined by the investigators.\u003c/sup\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e\u003csup\u003eb\u003c/sup\u003e\u003c/strong\u003e\u003csup\u003e\u0026nbsp;Indicates selected pertinent comorbidities present at time of first imaging demonstrating MBD.\u003c/sup\u003e\u003c/p\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2 Data Preprocessing\u003c/h2\u003e\u003cp\u003eWe curated a dataset of 555 radiology reports annotated by clinical experts into three referral classes: 0\u0026thinsp;=\u0026thinsp;No Referral, 1\u0026thinsp;=\u0026thinsp;Referral, and 2\u0026thinsp;=\u0026thinsp;Referral/High Risk. To ensure a robust evaluation, the dataset was split into a stratified hold-out test set (15%) and a train/validation pool (85%). The stratification preserved class balance across subsets. Randomization was fixed for reproducibility. We trained the 'No Referral' and 'Referral' classes using only radiology reports from the initial patient visit. In contrast, for the 'Referral/High Risk' class, due to the scarcity of the data, we augmented the dataset by treating reports from all of a patient's visits as distinct data samples.\u003c/p\u003e\u003cp\u003eWe applied a sliding-window tokenization strategy using the RadBERT tokenizer, as many reports exceeded the model\u0026rsquo;s maximum input length. Each report was split into overlapping segments with a maximum length of 512 tokens and a stride of 256 tokens. This approach enabled overlapping token spans and multi-window representation per report while maintaining compatibility with the pre-trained transformer backbone. Tokenization was performed after the data split to prevent leakage between sets. Each segment inherited the label of its original report, producing a window-level dataset for model training and evaluation.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e2.3 SHIELD\u0026rsquo;s Model Architecture for Radiology Report Classification\u003c/h2\u003e\u003cp\u003eThe proposed classification model is based on the RadBERT-RoBERTa-4m\u003csup\u003e20\u003c/sup\u003e transformer architecture fine-tuned for three-class classes (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e1\u003c/span\u003e.). The model takes tokenized windows of radiology reports as input. Each window is encoded by the pretrained RadBERT encoder to produce contextual embeddings. The [CLS] token representation is passed through a classification head consisting of a linear layer to generate logits for the three output classes. Cross-entropy loss is used for training.\u003c/p\u003e\u003cp\u003eA 5-fold stratified cross-validation was performed on the train/validation set, maintaining label distribution in each fold. During each fold iteration, tokenized train and validation windows were generated independently. After model inference, window-level predictions were aggregated back to report-level predictions using two schemes:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eMajority vote of the predicted classes across all windows for a report.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eMean probability aggregation, where class probabilities were averaged across windows and the class with the highest mean probability was selected.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eTraining was conducted using the Hugging Face Trainer API with the following settings:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eLearning rate: 2e-5\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eBatch size: 4 for both training and evaluation\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eNumber of epochs: 8\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eEvaluation and model checkpoint saving at each epoch\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eBest model selection based on validation macro F1-score\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eLogging every 10 steps\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eTo evaluate generalization, each fold\u0026rsquo;s model was tested on the same unseen hold-out test set. This set\u0026rsquo;s tokenization and window-counts were precomputed once to ensure consistent evaluation across folds. Final ensemble predictions were derived by averaging predicted probabilities across all five folds. During inference, the model generates logits for each window, which are converted to probabilities using softmax. These probabilities are then aggregated per report using the two methods described earlier.\u003c/p\u003e\u003cp\u003eTo enhance performance and stability:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eWe performed model ensembling by averaging predicted class probabilities from all five folds.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eAggregated predictions were evaluated using both \u003cb\u003emajority voting\u003c/b\u003e and \u003cb\u003emean-max probability\u003c/b\u003e, and both strategies were benchmarked on the hold-out test set.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eReceiver Operating Characteristic (ROC) curve analysis was performed for two key binary tasks: distinguishing \u0026ldquo;No Referral\u0026rdquo; from \u0026ldquo;Referral\u0026rdquo; and \u0026ldquo;Referral/High Risk\u0026rdquo; combined, and distinguishing \u0026ldquo;Referral\u0026rdquo; from \u0026ldquo;Referral/High Risk\u0026rdquo;, yielding interpretable AUC values to quantify discriminative performance.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eFinal outputs included per-class performance metrics (mean\u0026thinsp;\u0026plusmn;\u0026thinsp;std across folds), confusion matrices, and ROC curve plots.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eProposed architecture leveraged domain-specific pretraining, sliding-window augmentation, cross-validated ensembling, and robust evaluation strategies to deliver reliable and interpretable triage predictions from unstructured radiology narratives.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e2.4 Generating Interpretable Explanations with SHIELD's Generative AI\u003c/h2\u003e\u003cp\u003eTo provide interpretable rationales for each referral prediction, we implemented an automated pipeline using the Llama-3.1-8B-Instruct\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e model hosted locally. Radiology reports and their predicted class labels (No Referral, Referral, Referral/High Risk) obtained using our fine-tuned UCSD-VA-health/RadBERT-RoBERTa-4m\u003csup\u003e20\u003c/sup\u003e model were stored in a structured CSV file and processed sequentially. For each case, the radiology report text and its predicted label were combined into a class-conditioned prompt. The prompts explicitly requested the model to identify and explain the linguistic or diagnostic features supporting the assigned decision. The following prompts were used:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eClass 0 (No Referral)\u003c/b\u003e:\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u0026ldquo;Identify and explain the terminology or findings that would support \u003cb\u003eNOT referring\u003c/b\u003e this patient to the Orthopedic Oncology department.\u0026rdquo;\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eClass 1 (Referral)\u003c/b\u003e:\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u0026ldquo;Identify and explain the key terminology or findings that would indicate the patient \u003cb\u003eSHOULD be referred\u003c/b\u003e to the Orthopedic Oncology department.\u0026rdquo;\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eClass 2 (Referral/High Risk)\u003c/b\u003e:\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u0026ldquo;Identify and explain the specific terminology or findings that suggest the patient should be referred to the Orthopedic Oncology department \u003cb\u003edue to risk of pathological fracture (emergency).\u003c/b\u003e\u0026rdquo;\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThe radiology report text was embedded within this prompt, and the model was queried to generate free-text explanations. Each generation was constrained to a maximum of 1024 tokens with a low sampling temperature (0.3) to encourage concise and reproducible outputs. The resulting explanations were appended to the original dataset as an additional column and exported as a new CSV file for subsequent analysis. This procedure is depicted in Fig. 2. and ensures that every report\u0026ndash;prediction pair is accompanied by a standardized, model-generated explanation, enabling a systematic review of the linguistic cues underlying the automated decisions.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003e2.5 Comparative Analysis of Clinician and LLM-Generated Terminology\u003c/h2\u003e\u003cp\u003eTo systematically compare terminology used by clinicians with that generated by a large language model (LLM) for orthopedic oncology referral, we implemented a three-phase computational analysis pipeline. The input consisted of two curated term lists: (i) unique medical terms provided by an expert orthopedist, which two medical students used to sort radiology reports into three classes and label them, and (ii) a text-based explanation of the decision generated by the LLM, which the medical students filtered out to keep the main terms only.\u003c/p\u003e\u003cp\u003e\u003cb\u003ePhase 1: Preprocessing\u003c/b\u003e: All terms were normalized to lowercase, stripped of punctuation, tokenized, and lemmatized using the Natural Language Toolkit (NLTK). This ensured consistent lexical forms across both datasets (e.g., \u0026ldquo;lesions\u0026rdquo; and \u0026ldquo;lesion\u0026rdquo; were reduced to a common root form).\u003c/p\u003e\u003cp\u003e\u003cb\u003ePhase 2: Lexical (Set-Based) Analysis\u003c/b\u003e: We applied set-theoretic comparisons to quantify overlap and exclusivity between clinician and LLM-derived terminology. Terms were categorized into three groups: shared terms, clinician-only terms (missed by the LLM), and LLM-only terms (absent from clinician usage). The lexical overlap was visualized using a Venn diagram, enabling rapid assessment of concordance and divergence.\u003c/p\u003e\u003cp\u003e\u003cb\u003ePhase 3: Semantic and Conceptual Analysis\u003c/b\u003e: To capture conceptual similarities beyond exact word matches, we employed a BioBERT (dmis-lab/biobert-v1.1) model to generate vector embeddings of clinician-only and LLM-only terms\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e,\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. Cosine similarity was computed between all term pairs, producing a similarity matrix. This enabled identification of LLM terms that were semantically aligned with clinician terms despite lexical differences. Heatmaps were generated to display the 30 most descriptive term pairs. Furthermore, similarity categories were defined as \u003cem\u003eHigh (likely synonyms, score\u0026thinsp;\u0026gt;\u0026thinsp;0.9)\u003c/em\u003e, \u003cem\u003eModerate (related concepts, score 0.75\u0026ndash;0.9)\u003c/em\u003e, or \u003cem\u003eLow (likely unrelated, score\u0026thinsp;\u0026lt;\u0026thinsp;0.75)\u003c/em\u003e. Results were exported to a CSV file for detailed inspection.\u003c/p\u003e\u003cp\u003e This multi-phase pipeline provided both lexical and semantic perspectives on the relationship between clinician-reported and LLM-generated terminology, enabling systematic identification of overlaps, mismatches, and conceptually related terms.\u003c/p\u003e\u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e3.1 Classification Performance\u003c/h2\u003e\u003cp\u003eThe performance of the proposed SHIELD model was evaluated on the hold-out test set, with performance metrics averaged across a 5-fold cross-validation. Although the majority vote and the mean probability aggregation methods showed very close results, we will only present the latter.\u003c/p\u003e\u003cp\u003eThe model achieved a high overall accuracy of 89.52% \u0026plusmn; 2.84% as shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e2\u003c/span\u003e. For the \"No Referral\" category, the model demonstrated outstanding performance. It achieved a precision of 98.97% \u0026plusmn; 2.29%, a recall of 99.46% \u0026plusmn; 1.21%, and a resulting F1-score of 99.20% \u0026plusmn; 1.18%. This indicates exceptional reliability in correctly identifying reports that do not require further clinical action, a crucial function for reducing unnecessary workload. The model was also highly effective in identifying reports requiring clinical follow-up. For the general \"Referral\" class, it yielded a recall of 96.55% \u0026plusmn; 4.22%, signifying that very few referral cases were missed. The precision for this class was 78.82% \u0026plusmn; 4.27%, leading to a well-balanced F1-score of 86.72% \u0026plusmn; 3.27%.\u003c/p\u003e\u003cp\u003eDistinguishing between standard \u0026ldquo;Referral\u0026rdquo; and urgent referrals (Referral/High Risk) proved the most challenging aspect of model performance.. In this task, while the model exhibited high precision at 94.05% \u0026plusmn; 9.37%, indicating a low rate of false positives the recall for this class was considerably lower (57.78% \u0026plusmn; 13.94%). This trade-off between high certainty and moderate sensitivity resulted in an F1-score of 70.53% \u0026plusmn; 10.73% for this critical category. These results indicate that the model is more reliable at correctly identifying urgent referrals when it flags them as such, but misses over half of all actual high-risk cases.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003e\u003cb\u003eCross-Validated Classification Performance of the AI Model.\u003c/b\u003e This table summarizes the performance of the final model on the hold-out test dataset. The reported metrics of precision, recall, and F1-score averaged over five validation folds. Predictions for each report were aggregated by calculating the mean of the maximum probabilities across all text segments. All performance metrics are presented as the mean\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eClasses\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003ePrecision\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRecall\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eF1-score\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSupport\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNo Referral\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e\u003cp\u003e0.9897\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0229\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e\u003cp\u003e0.9946\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0121\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e\u003cp\u003e0.9920\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0118\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e37\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eReferral\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e\u003cp\u003e0.7882\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0427\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e\u003cp\u003e0.9655\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0422\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e\u003cp\u003e0.8672\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0327\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e29\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eReferral/High Risk\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e\u003cp\u003e0.9405\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0937\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e\u003cp\u003e0.5778\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1394\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e\u003cp\u003e0.7053\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1073\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e18\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAccuracy\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e\u003cp\u003e0.8952\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0284\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e84\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eTo further visualize the classification performance, confusion matrices for both binary and multi-class scenarios were generated from the ensemble model's predictions on the hold-out test set. The first matrix (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e3\u003c/span\u003e(a)) illustrates the model's performance on the primary binary task of distinguishing reports that require a referral from those that do not. In this scenario, the \"Referral\" and \"Referral/High Risk\" classes were consolidated into a single \"Referral\" category. The model achieved perfect classification, correctly identifying all 37 \"No Referral\" cases and all 47 \"Referral\" cases. This result underscores the model's exceptional capability to reliably determine whether a clinical follow-up is necessary.\u003c/p\u003e\u003cp\u003eThe second matrix (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e3\u003c/span\u003e(b)) provides a more granular, three-class analysis. The model continued to demonstrate flawless performance for the \"No Referral\" and standard \"Referral\" categories, correctly classifying all 37 and 29 reports, respectively. The model's classification errors were isolated to the \"Referral/High Risk\" category. Of the 18 true \"Referral/High Risk\" cases, 10 were correctly identified. The remaining 8 cases were misclassified as a standard \"Referral.\" Critically, no \"Referral/High Risk\" case was misclassified as \"No Referral.\" This indicates that while the model struggles to distinguish the nuanced differences between a standard and a high-risk referral, its errors are \"fail-safe\" in a clinical sense\u0026mdash;it successfully flags all high-risk patients as requiring a referral, even if it de-escalates the urgency level for a subset of them.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eThe discriminative power of the model was further assessed using Receiver Operating Characteristic (ROC) curve analysis for two key classification scenarios, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e3\u003c/span\u003e(c). For the primary task of distinguishing \"No Referral\" reports from all referral cases (\"Referral\" and \"Referral/High Risk\" combined), the model achieved a perfect Area Under the Curve (AUC) of 1.00. The blue line, representing this task, aligns perfectly with the top-left corner of the plot, indicating a flawless separation between the two groups. This result confirms that the model can achieve a 100% true positive rate (sensitivity) with a 0% false positive rate (100% specificity), underscoring its reliability for initial screening. In the more challenging sub-classification task of differentiating between a standard \"Referral\" and a \"Referral/High Risk\" case, the model also demonstrated excellent performance. The red curve shows a strong ability to distinguish between these two classes, yielding a high AUC of 0.93. This indicates a robust capability to correctly rank and identify high-risk cases over standard referral cases. Collectively, the ROC analysis confirms the findings from the confusion matrices: the model is exceptionally accurate in the primary determination of referral need and maintains a high degree of discriminative ability when stratifying the urgency of that referral.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\u003ch2\u003e3.2 Comparison of Conventional vs. SHIELD-Based Referral Timelines\u003c/h2\u003e\u003cp\u003eTo quantify the potential clinical impact of our AI framework, we compared the referral timelines of the proposed model against the conventional clinical pathway, with the results summarized in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e3\u003c/span\u003e. The analysis revealed that the conventional method is associated with substantial delays, requiring an average of 109.64\u0026thinsp;\u0026plusmn;\u0026thinsp;286.09 days from the initial imaging report to a patient referral. Additional delays were observed within this period, including an average of 80.57 days from the first documented high-risk descriptor to referral and a diagnostic window spanning 97.50 days between the first and last imaging events prior to referral. In stark contrast, the AI model provides its classification recommendation within 1 to 3 minutes of processing the initial report. By generating an immediate recommendation from a single data point, the model renders the prolonged, multi-step diagnostic timelines of the conventional pathway obsolete, demonstrating its potential to drastically reduce administrative and clinical delays inherent in the current standard of care.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eComparison of Referral Timelines Between Conventional Methods and the Proposed AI Model.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMetric\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eConventional Method (days, mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eProposed AI Model (minutes)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTime from Initial Imaging to Referral Decision\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e\u003cp\u003e109.64\u0026thinsp;\u0026plusmn;\u0026thinsp;286.09\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003e~\u0026thinsp;1\u0026ndash;3\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTime from First High-Risk Finding to Referral\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e\u003cp\u003e80.57\u0026thinsp;\u0026plusmn;\u0026thinsp;256.87\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDiagnostic Period (First to Last Imaging Pre-Referral)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e\u003cp\u003e97.50\u0026thinsp;\u0026plusmn;\u0026thinsp;281.99\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eOf the 245 patients evaluated, 124 (50.61%) had undergone at least two imaging visits prior to referral. Additionally, 85 patients (34.69%) experienced delayed referrals, defined as a referral made more than 28 days after the first imaging report indicating potential bone malignancy. The proposed AI solution demonstrated promising results by substantially reducing the referral delay, successfully classifying referrals based solely on the first imaging report for \"No Referral\" and \u0026ldquo;Referral\" classes.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e3.3 Explainable AI\u003c/h2\u003e\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e4\u003c/span\u003e presents a side-by-side comparison of an original radiology report accurately classified as Referral/High-Risk and its corresponding AI-generated explanation.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eAn example of an explained radiology report provided by LLM.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eOriginal Radiology Report\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAI Explanation\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e6/21/22 - XR Femur Minimum 2 Vw Right Status: Final result Study Result Narrative \u0026amp; Impression\u003c/p\u003e\u003cp\u003eDATE OF SERVICE: 6/21/2022 4:10 pm\u003c/p\u003e\u003cp\u003eEXAM: Right femur two views\u003c/p\u003e\u003cp\u003eClinical history: Bone Mets, right femur pain.\u003c/p\u003e\u003cp\u003eComparison: CT pelvis 3/26/2022\u003c/p\u003e\u003cp\u003eFindings: Mixed sclerotic and lytic osseous lesions are again noted. There is a pathologic fracture without significant displacement at a lytic lesion of the lesser trochanter of the right proximal femur.\u003c/p\u003e\u003cp\u003e6/21/22 - XR Pelvis 1\u0026ndash;2 Views\u003c/p\u003e\u003cp\u003eStatus: Edited Result - FINAL Addendum\u003c/p\u003e\u003cp\u003eADDENDUM REPORT: 6/21/2022 4:36 PM\u003c/p\u003e\u003cp\u003eADDENDUM: There is a pathologic fracture without significant displacement at a lytic lesion of the right proximal femur lesser trochanter.\u003c/p\u003e\u003cp\u003eTHIS IS AN ELECTRONICALLY VERIFIED\u003c/p\u003e\u003cp\u003eFINAL REPORT 6/21/2022 4:36 PM\u003c/p\u003e\u003cp\u003eDATE OF SERVICE: 6/21/2022 4:10 pm\u003c/p\u003e\u003cp\u003eEXAM: Pelvis single view\u003c/p\u003e\u003cp\u003eComparison: CT pelvis 3/26/2022\u003c/p\u003e\u003cp\u003eClinical history: Bone Mets, right femur pain.\u003c/p\u003e\u003cp\u003eFindings: No acute fracture or dislocation. Multiple osseous sclerotic and a few lytic lesions are again noted, and there appears to be a couple new small foci of lytic lesions at the proximal right femur. Pelvic phleboliths are present.\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eThe report indicates that the patient has a pathologic fracture without significant displacement at a lytic lesion of the lesser trochanter of the right proximal femur. This suggests that the patient has a bone metastasis that has weakened the bone structure, leading to a fracture. The presence of lytic lesions, which are areas of bone destruction, and sclerotic lesions, which are areas of bone formation, indicates that the patient has a bone metastasis. The fact that the fracture is not significantly displaced suggests that it is a pathologic fracture, which is a fracture that occurs due to a underlying bone disease, rather than a traumatic fracture.\u003c/p\u003e\u003cp\u003eThe patient's clinical history of bone Mets (metastases) and right femur pain also supports the need for referral to the Orthopedic Oncology department. The presence of multiple osseous sclerotic and lytic lesions, as well as new small foci of lytic lesions at the proximal right femur, suggests that the patient's bone metastasis is progressive and requires further evaluation and management.\u003c/p\u003e\u003cp\u003eTherefore, the specific terminology or findings that suggest the patient should be referred to the Orthopedic Oncology department due to risk of pathological fracture (emergency) are:\u003c/p\u003e\u003cp\u003e* Pathologic fracture without significant displacement at a lytic lesion of the lesser trochanter of the right proximal femur\u003c/p\u003e\u003cp\u003e* Presence of lytic lesions and sclerotic lesions indicating bone metastasis\u003c/p\u003e\u003cp\u003e* Clinical history of bone Mets and right femur pain\u003c/p\u003e\u003cp\u003e* Presence of multiple osseous sclerotic and lytic lesions, as well as new small foci of lytic lesions at the proximal right femur.\u003c/p\u003e\u003cp\u003eThese findings suggest that the patient requires urgent evaluation and management by the Orthopedic Oncology department to address the risk of pathological fracture and to manage the underlying bone metastasis.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe complete list of reports explained is available as supplementary material. In our analysis, an expert orthopedic doctor provided 74 unique descriptive terms, while the LLM produced 199 unique terms after careful filtering. The Venn diagram in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003e illustrates the overlap between these term sets, showing 13 terms that were common to both the expert and the LLM: 'enhancing lesion', 'periosteal reaction', 'aggressive appearing', 'cortical disruption', 'cortical destruction', 'expansile', 'cortical erosion', 'aggressive', 'cortical thinning', 't1 hypointense t2 hyperintense', 'endosteal scalloping', 'lytic lesion', 'cortical breakthrough'. Some examples for clinician-only terms were 'predisposed to pathologic fracture', 'lytic change', 'increased risk of pathologic fracture', 'hypermetabolic osseous disease', 'cortical transgression', 'concerning for instability', 'hypointense', 'mixed lytic', and 'increased uptake in bone'. In contrast LLM specific terms were more lexically longer like 'localized pathologic fracture', 'sclerotic lesion', 'expansile lytic lesion involving the left anterior 5th rib', 'fracture', 'proximal femur', 'proximal right femur', 'presence of lytic lesion and sclerotic lesion indicating bone metastasis', 'possibility of a primary osseous neoplasm or lymphoma cannot be excluded', 'mildly displaced pathologic fracture', 'indeterminate lucent lesion'.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eTo evaluate the semantic alignment between the LLM's generated explanations and standard clinical terminology, we performed a cosine similarity analysis on keywords and phrases that were unique to either the LLM's output (\"LLM-Only\") or the clinician's lexicon (\"Clinician-Only\") demonstrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e5\u003c/span\u003e. The results, presented in a heatmap, reveal a high degree of semantic congruence between the two vocabularies, even when the exact phrasing differed. Multiple term-pairs exhibited strong conceptual overlap with similarity scores exceeding 0.90. For instance, the LLM-generated phrase \"permeation through the lateral cortex\" was found to be almost identical in meaning to the clinician's term \"transcortical permeation\" (similarity\u0026thinsp;=\u0026thinsp;0.947). Similarly, the LLM's more specific \"femoral head\" and \"femoral neck\" showed very high similarity (0.944 and 0.942, respectively) to the broader clinician term \"femur.\" We also observed that the LLM's descriptive phrase \"lytic metastatic bone lesion\" was a strong semantic match to the concise clinical term \"metastasis\" (similarity\u0026thinsp;=\u0026thinsp;0.900), confirming the model's ability to capture clinical meaning accurately. Another pair from LLM being \"soft tissue mass\" with the clinician's \"marrow infiltration\" at a score of 0.85 demonstrates a clear example of a lexically different but contextually similar pair. This heatmap shows that the LLM can generate lexically diverse but clinically relevant and semantically equivalent terms when compared to an expert clinician's vocabulary.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003ePrior work from our group has demonstrated the growing role of artificial intelligence in orthopedics and oncology. In orthopedics, we applied machine learning to model fracture recovery using gait analysis, offering predictive insights into patient outcomes following lower extremity fractures. Beyond orthopedics, we advanced interpretable AI frameworks for oncology, including\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e,\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e deep learning models for tumor classification and cancer recurrence prediction\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e,\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. We also developed multimodal language\u0026ndash;vision assistants in pathology, highlighting the feasibility of adapting LLMs for domain-specific interpretation\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e. Collectively, these contributions provide a foundation for the present work, which builds on this trajectory by addressing delayed referrals in metastatic bone disease with both accurate classification and explainable AI-driven explanations. In response to the critical issue of delayed referrals for metastatic bone disease, this study successfully developed and validated a dual-component AI framework of SHIELD designed to both classify and explain findings in radiology reports. Complementing its predictive power, the explainability component of Llama-3.1 model demonstrated its ability to generate clinically coherent rationales.\u003c/p\u003e\u003cp\u003eSHIELD\u0026rsquo;s classification model demonstrates exceptional utility as a clinical decision support tool, primarily by automating the initial screening of radiology reports with a high degree of reliability. The model's near-perfect performance in identifying \"No Referral\" cases (F1-score of 99.20%) is a key finding, as it signifies a robust capability to reduce clinician workload by safely filtering out non-urgent reports. This strength is further underscored by the model's flawless binary classification, achieving a perfect AUC of 1.00 when distinguishing between reports that require a referral and those that do not. This near-perfect accuracy in the primary screening task suggests that the model can be trusted to ensure that no patient requiring follow-up is overlooked, a critical benchmark for implementation in a clinical workflow.\u003c/p\u003e\u003cp\u003eThe more nuanced challenge lies in differentiating the urgency between a standard \"Referral\" and a \"Referral/High Risk\" case. Here, the model exhibits a clinically meaningful trade-off: its high precision (94.05%) in the \"Referral/High Risk\" category ensures that when it flags a case as high-risk, the prediction is highly trustworthy. However, its moderate recall (57.78%) for this same category indicates a tendency to classify some high-risk cases into the standard referral pool. Critically, the confusion matrix reveals that these errors are \"fail-safe\"\u0026mdash;no high-risk case was ever misclassified as \"No Referral.\" This behaviour is paramount for patient safety; while the model may occasionally underestimate the level of urgency, it correctly identifies every at-risk patient as needing clinical attention. The high AUC of 0.93 for this sub-classification task further suggests that the model possesses strong discriminative power, and future work would focus on tuning the classification threshold by increasing number of training samples to optimize the balance between precision and recall based on specific clinical needs.\u003c/p\u003e\u003cp\u003eThe stark contrast in referral timelines highlights the most profound clinical implication of the SHIELD framework: its potential to dismantle the protracted and inefficient nature of the current diagnostic pathway. Clinical triaging with an average delay of nearly 110 days from initial imaging to referral is not merely slow but is characterized by systemic inefficiencies. Our findings show that over half of patients required multiple imaging visits and more than a third experienced referral delays exceeding a month, underscoring a period of clinical uncertainty and patient anxiety. In direct contrast, the AI model\u0026rsquo;s ability to generate a reliable classification in minutes from the very first imaging report represents a paradigm shift. It obviates the need for a prolonged \"watch-and-wait\" period and multiple follow-up scans, which are currently the primary drivers of these delays. By intervening at the earliest possible point, the framework transforms a reactive, multi-step process into a proactive and immediate decision point. This can potentially ensure the journey from initial suspicion to specialist consultation begins without delay.\u003c/p\u003e\u003cp\u003eThis investigation demonstrates the framework\u0026rsquo;s robust capability to process and reinterpret specialized medical language for clinical applications. The key finding is the significant disparity between the low lexical overlap and the high semantic similarity when comparing the LLM's output to an expert clinician's terminology. This suggests that the LLM is not merely extracting keywords but is generating a conceptually aligned, and often more descriptive, vocabulary to explain complex radiological findings. The model's ability to produce lexically diverse yet semantically equivalent terms underscores its potential as a sophisticated tool for clinical communication. By translating dense medical reports into structured and simplified explanations, as exemplified in the case study, the model can serve as a valuable aid in enhancing patient understanding and supporting clinical decision-making.\u003c/p\u003e\u003cp\u003eOur work contributes to the growing body of research on AI-driven clinical triage, sharing the overarching goal of improving upon manual, often inefficient processes, as seen in both emergency\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e and specialist settings. While our framework addresses a specialized referral pathway similar to recent studies in gatekeeping and triage, its classification performance notably surpasses these prior efforts. For instance, the AI gatekeeping model by Vergara et al.\u003csup\u003e12\u003c/sup\u003e achieved an overall accuracy of 71.6% (AUC of 0.765) in a binary referral authorization task across multiple specialties. Similarly, the work by Maarseveen et al.\u003csup\u003e34\u003c/sup\u003e reported an AUC of 0.78 for prioritizing rheumatology referrals, and the system by Abdel-Hafez et al.\u003csup\u003e35\u003c/sup\u003e showed a 53.8% level of agreement for categorizing ENT referrals. A directly comparable study was recently introduced by Sangwon et al.\u003csup\u003e36\u003c/sup\u003e, who evaluated three NLP models for automating bone metastases referrals: a rule-based Regular Expression (RegEx) model, GPT-4, and a specialized BERT model (NYUTron). Their findings indicated that the RegEx model often outperformed the more complex LLMs for their specific clinical application, achieving an F1-score of 88.9% during validation. While this work represents a valuable contribution to AI-driven triage, it has notable limitations, such as the lack of a clinically validated timeline comparison and the use of a dataset covering a shorter time frame.\u003c/p\u003e\u003cp\u003eIn contrast to these studies, our SHIELD framework demonstrates a significant advance in performance, achieving 100% accuracy in the binary Referral vs. No Referral task and an excellent AUC of 0.93 in the more complex three-class problem. Furthermore, a key architectural advantage of our framework is its operational autonomy; following the initial data labeling phase, the model operates independently without requiring real-time clinician interrogation or rule-based adjustments for its predictions. This performance advantage likely stems from our model's unique design, which enables it to interpret the dense, unstructured narrative of a full radiology report for a specialized oncological pathway. However, the most significant departure remains our framework's integral dual-component architecture. Where other models focus primarily on the classification task, our work pairs its high accuracy with a sophisticated explainability component using a large language model. This distinction is crucial\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e, as our goal is not only to accelerate a workflow but also to build clinician trust and facilitate communication for a considered specialist referral\u0026mdash;a focus on interpretability that moves beyond what is presented in related literature.\u003c/p\u003e\u003cp\u003eThis study has several limitations. Our dataset consisted of 1,404 EHRs collected over ten years. Furthermore, the scope was restricted to malignancies of the long bones (femur, fibula, tibia, humerus, radius, and ulna). The model's performance on the 'Referral/High Risk' class was impacted by a relatively small number of samples in this category, which presented a challenge for achieving optimal classification accuracy. Future work will aim to address these limitations. We plan to expand the dataset by including a larger number of EHRs and incorporating all sites of skeletal metastasis. Additionally, integrating multimodality by adding image analysis could further enhance the model's performance, making it an even more reliable tool for medical experts\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e. Collectively, these findings validate the framework's potential to significantly accelerate the patient referral pathway while fostering trust and adoption among clinicians through transparent and meaningful explanations.\u003c/p\u003e"},{"header":"5. Conclusion","content":"\u003cp\u003eGiven the aging population and increasing cancer incidence, there is a critical and growing need to identify patients at risk for metastatic bone disease and skeletal-related events earlier in their disease course. The promising results of this study suggest our proposed SHIELD framework can directly address this challenge by ameliorating the problem of late referrals. The framework facilitates the early recognition of high-risk radiographic findings, particularly when explicit descriptors, such as \u0026ldquo;lytic lesion\u0026rdquo; or \u0026ldquo;impending fracture,\u0026rdquo; are present. Its adoption can lead to significantly improved survival and functional outcomes by enabling healthcare professionals to treat patients prophylactically. Ultimately, leveraging AI to analyze language patterns within radiology reports provides a scalable and effective pathway to ensure more accurate and timely referrals to orthopedic oncology, along with reliable explanations.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eParticipant Consent: The study used de-identified retrospective data and did not involve direct participant contact. The need for individual consent was waived by the Institutional Review Board (IRB) of Wake Forest School of Medicine (Approval ID: IRB00127840) in accordance with ethical guidelines and regulatory standards.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eAuthor Contribution:\u0026nbsp;\u003c/strong\u003eC.L.E. and M.N.G. conceptualized the study. C.L.E. identified the clinical need, proposed the original project idea, and provided clinical oversight. Under the supervision of C.L.E., A.M.M. was responsible for the acquisition, de-identification, and curation of the radiology report dataset from the participating medical centers. A.A. designed the computational methodology, developed and trained the SHIELD framework, conducted all computational experiments, and performed the statistical and retrospective timeline analyses under the direct supervision of M.N.G. F.M.D. provided medical consultation, aiding in the definition of the risk-stratification tiers and the clinical validity of the model\u0026rsquo;s outputs. A.A. and M.N.G. wrote the original draft of the manuscript. All authors participated in the critical review, editing, and final approval of the manuscript and approved the journal submission.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements:\u0026nbsp;\u003c/strong\u003eThe authors would like to thank Robert B. Lipsit, for his contributions to patient query, exclusion, and data collection.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u0026nbsp;\u003c/strong\u003eAva M. McKane was part of the Medical Student Research Program sponsored by the Wake Forest University School of Medicine and Department of Orthopaedics. The project described was supported in part by R01 DC020715 (PIs: Gurcan, Moberly) from the National Institute on Deafness and Other Communication Disorders, R21 CA273665 (PI: Gurcan) from the National Cancer Institute, U01 TR003629 (PIs: Gurcan, Paige, Steube) from the National Center for Advancing Translational Sciences, and R01HL177046 (PIs: Gurcan, Michelson, Hachem) from the National Heart, Lung, and Blood Institute. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health, National Institute on Deafness and Other Communication Disorders, National Cancer Institute, National Heart, Lung, and Blood Institute, or National Center for Advancing Translational Sciences.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability:\u0026nbsp;\u003c/strong\u003eThe dataset from this study is held securely in coded form Wake Forest University School of Medicine. The full dataset creation plan are available from the authors upon request. The source code for implementing the methods is available at the following repository: https://github.com/CAIR-LAB-WFUSM/SHIELD.git.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHuman Ethics and Consent to Participate:\u003c/strong\u003e This project was approved by the Wake Forest University School of Medicine Research Ethics Board Protocol review board approval (IRB00127840) in accordance with the Declaration of Helsinki. All the records in this study were de-identified.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClinical trial number:\u0026nbsp;\u003c/strong\u003enot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of Interest:\u003c/strong\u003e The authors have declared that no competing interests exist.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests:\u003c/strong\u003e The authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eDiCaprio MR, Murtaza H, Palmer B, Evangelist M. Narrative review of the epidemiology, economic burden, and societal impact of metastatic bone disease. \u003cem\u003eAnn Jt\u003c/em\u003e.\u003cem\u003eAME Publishing Company\u003c/em\u003e. 2022;7. doi:10.21037/aoj-20-97\u003c/li\u003e\n\u003cli\u003eCanavan ME, Wang X, Ascha MS, et al. Systemic Anticancer Therapy and Overall Survival in Patients with Very Advanced Solid Tumors. \u003cem\u003eJAMA Oncol\u003c/em\u003e. 2024;10(7):887-895. doi:10.1001/jamaoncol.2024.1129\u003c/li\u003e\n\u003cli\u003eSvensson E, Christiansen CF, Ulrichsen SP, R\u0026oslash;rth MR, S\u0026oslash;rensen HT. Survival after bone metastasis by primary cancer type: A Danish population-based cohort study. \u003cem\u003eBMJ Open\u003c/em\u003e. 2017;7(9). doi:10.1136/bmjopen-2017-016022\u003c/li\u003e\n\u003cli\u003eMosher ZA, Patel H, Ewing MA, et al. \u003cem\u003eEarly Clinical and Economic Outcomes of Prophylactic and Acute Pathologic Fracture Treatment\u003c/em\u003e.; 2025. https://doi.org/10.\u003c/li\u003e\n\u003cli\u003eArdakani AHG, Faimali M, Nystrom L, et al. Metastatic bone disease: Early referral for multidisciplinary care. \u003cem\u003eCleve Clin J Med\u003c/em\u003e. 2022;89(7):393-399. doi:10.3949/ccjm.89a.21062\u003c/li\u003e\n\u003cli\u003eLevin A. Consequences of late referral on patient outcomes. \u003cem\u003eNephrology Dialysis Transplantation\u003c/em\u003e. 2000;15(suppl_3):8-13. doi:10.1093/oxfordjournals.ndt.a027977\u003c/li\u003e\n\u003cli\u003eKotrych D, Ciechanowicz D, Pawlik J, Szyjkowska M, Kwapisz B, Mądry M. Delay in Diagnosis and Treatment of Primary Bone Tumors during COVID-19 Pandemic in Poland. \u003cem\u003eCancers (Basel)\u003c/em\u003e. 2022;14(24). doi:10.3390/cancers14246037\u003c/li\u003e\n\u003cli\u003eJawad MU, Scully SP. In brief: Classifications in brief: Mirels\u0026rsquo; classification: Metastatic disease in long bones and impending pathologic fracture. \u003cem\u003eClin Orthop Relat Res\u003c/em\u003e.\u003cem\u003eSpringer New York LLC\u003c/em\u003e. 2010;468(10):2825-2827. doi:10.1007/s11999-010-1326-4\u003c/li\u003e\n\u003cli\u003eKimura T. Multidisciplinary approach for bone metastasis: A review. \u003cem\u003eCancers (Basel)\u003c/em\u003e.\u003cem\u003eMDPI AG\u003c/em\u003e. 2018;10(6). doi:10.3390/cancers10060156\u003c/li\u003e\n\u003cli\u003eBreden S, Hinterwimmer F, Consalvo S, et al. Deep Learning-Based Detection of Bone Tumors around the Knee in X-rays of Children. \u003cem\u003eJ Clin Med\u003c/em\u003e. 2023;12(18). doi:10.3390/jcm12185960\u003c/li\u003e\n\u003cli\u003eRizk PA, Gonzalez MR, Galoaa BM, et al. Machine Learning\u0026ndash;Assisted Decision Making in Orthopaedic Oncology. \u003cem\u003eJBJS Rev\u003c/em\u003e. 2024;12(7). doi:10.2106/JBJS.RVW.24.00057\u003c/li\u003e\n\u003cli\u003eVergara PO, Oliveira JDC, Mattiello R, et al. Accuracy of Artificial Intelligence for Gatekeeping in Referrals to Specialized Care. \u003cem\u003eJAMA Netw Open\u003c/em\u003e. Published online 2025. doi:10.1001/jamanetworkopen.2025.13285\u003c/li\u003e\n\u003cli\u003eFudickar S, Bantel C, Spieker J, et al. Natural Language Processing of Referral Letters for Machine Learning\u0026ndash;Based Triaging of Patients With Low Back Pain to the Most Appropriate Intervention: Retrospective Study. \u003cem\u003eJ Med Internet Res\u003c/em\u003e. 2024;26(1). doi:10.2196/46857\u003c/li\u003e\n\u003cli\u003eFarrow L, Zhong M, Anderson L. Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports. \u003cem\u003eBone Joint J\u003c/em\u003e. 2024;106(7):688-695. doi:10.1302/0301-620X.106B7\u003c/li\u003e\n\u003cli\u003eHu Y, Xiang Y, Zhou YJ, et al. AI-based diagnosis of acute aortic syndrome from noncontrast CT. \u003cem\u003eNat Med\u003c/em\u003e. Published online 2025. doi:10.1038/s41591-025-03916-z\u003c/li\u003e\n\u003cli\u003eZhao Z, Zhang Y, Wu C, et al. Large-vocabulary segmentation for medical images with text prompts. \u003cem\u003eNPJ Digit Med\u003c/em\u003e. 2025;8(1):566. doi:10.1038/s41746-025-01964-w\u003c/li\u003e\n\u003cli\u003eFogel AL, Kvedar JC. Artificial intelligence powers digital medicine. \u003cem\u003eNPJ Digit Med\u003c/em\u003e.\u003cem\u003eNature Publishing Group\u003c/em\u003e. 2018;1(1). doi:10.1038/s41746-017-0012-2\u003c/li\u003e\n\u003cli\u003eLee J, Yoon W, Kim S, et al. BioBERT: A pre-trained biomedical language representation model for biomedical text mining. \u003cem\u003eBioinformatics\u003c/em\u003e. 2020;36(4):1234-1240. doi:10.1093/bioinformatics/btz682\u003c/li\u003e\n\u003cli\u003eGu Y, Tinn R, Cheng H, et al. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. \u003cem\u003eACM Trans Comput Healthc\u003c/em\u003e. 2022;3(1). doi:10.1145/3458754\u003c/li\u003e\n\u003cli\u003eYan A, McAuley J, Lu X, et al. RadBERT: Adapting Transformer-based Language Models to Radiology. \u003cem\u003eRadiol Artif Intell\u003c/em\u003e. 2022;4(4). doi:10.1148/ryai.210258\u003c/li\u003e\n\u003cli\u003eRosenbacke R, Melhus \u0026Aring;, McKee M, Stuckler D. How Explainable Artificial Intelligence Can Increase or Decrease Clinicians\u0026rsquo; Trust in AI Applications in Health Care: Systematic Review. \u003cem\u003eJMIR AI\u003c/em\u003e.\u003cem\u003eJMIR Publications Inc.\u003c/em\u003e 2024;3. doi:10.2196/53207\u003c/li\u003e\n\u003cli\u003eSadeghi Z, Alizadehsani R, CIFCI MA, et al. A review of Explainable Artificial Intelligence in healthcare. \u003cem\u003eComputers and Electrical Engineering\u003c/em\u003e. 2024;118. doi:10.1016/j.compeleceng.2024.109370\u003c/li\u003e\n\u003cli\u003eAziz NA, Manzoor A, Mazhar Qureshi MD, Qureshi MA, Rashwan W. Unveiling Explainable AI in Healthcare: Current Trends, Challenges, and Future Directions. Preprint posted online August 10, 2024. doi:10.1101/2024.08.10.24311735\u003c/li\u003e\n\u003cli\u003eKibria MG, Kucirka L, Mostafa J. Assessing AI Explainability: A Usability Study Using a Novel Framework Involving Clinicians. In: Institute of Electrical and Electronics Engineers (IEEE); 2025:553-564. doi:10.1109/ichi64645.2025.00069\u003c/li\u003e\n\u003cli\u003eGrattafiori A, Dubey A, Jauhri A, et al. The Llama 3 Herd of Models. Published online November 23, 2024. http://arxiv.org/abs/2407.21783\u003c/li\u003e\n\u003cli\u003eRen Y, Wu D, Khurana A, et al. Classification of Patient Portal Messages with BERT-based Language Models. In: \u003cem\u003eProceedings - 2023 IEEE 11th International Conference on Healthcare Informatics, ICHI 2023\u003c/em\u003e. Institute of Electrical and Electronics Engineers Inc.; 2023:176-182. doi:10.1109/ICHI57859.2023.00033\u003c/li\u003e\n\u003cli\u003eKroll H, Sackhoff P, Thang BM, Ksouri M, Balke WT. A Library Perspective on Supervised Text Processing in Digital Libraries: An Investigation in the Biomedical Domain. In: \u003cem\u003eProceedings of the ACM/IEEE Joint Conference on Digital Libraries\u003c/em\u003e. Institute of Electrical and Electronics Engineers Inc.; 2025. doi:10.1145/3677389.3702557\u003c/li\u003e\n\u003cli\u003eRezapour M, Seymour RB, Medda S, et al. Analyzing Gait Dynamics and Recovery Trajectory in Lower Extremity Fractures Using Linear Mixed Models and Gait Analysis Variables. \u003cem\u003eBioengineering\u003c/em\u003e. 2025;12(1). doi:10.3390/bioengineering12010067\u003c/li\u003e\n\u003cli\u003eRezapour M, Seymour RB, Sims SH, Karunakar MA, Habet N, Gurcan MN. Employing machine learning to enhance fracture recovery insights through gait analysis. \u003cem\u003eJournal of Orthopaedic Research\u003c/em\u003e. 2024;42(8):1748-1761. doi:10.1002/jor.25837\u003c/li\u003e\n\u003cli\u003eSu Z, Guo Y, Wesolowski R, et al. Computational Pathology for Accurate Prediction of Breast Cancer Recurrence: Development and Validation of a Deep Learning-based Tool. \u003cem\u003eModern Pathology\u003c/em\u003e. Published online December 2025:100847. doi:10.1016/j.modpat.2025.100847\u003c/li\u003e\n\u003cli\u003eTavolara TE, Su Z, Gurcan MN, Niazi MKK. One label is all you need: Interpretable AI-enhanced histopathology for oncology. \u003cem\u003eSemin Cancer Biol\u003c/em\u003e.\u003cem\u003eAcademic Press\u003c/em\u003e. 2023;97:70-85. doi:10.1016/j.semcancer.2023.09.006\u003c/li\u003e\n\u003cli\u003eAfzaal U, Su Z, Sajjad U, et al. HistoChat: Instruction-tuning multimodal vision language assistant for colorectal histopathology on limited data. \u003cem\u003ePatterns\u003c/em\u003e. Published online August 8, 2025. doi:10.1016/j.patter.2025.101284\u003c/li\u003e\n\u003cli\u003eTyler S, Olis M, Aust N, et al. Use of Artificial Intelligence in Triage in Hospital Emergency Departments: A Scoping Review. \u003cem\u003eCureus\u003c/em\u003e. Published online May 8, 2024. doi:10.7759/cureus.59906\u003c/li\u003e\n\u003cli\u003eMaarseveen TD, Glas HK, Veris-van Dieren J, van den Akker E, Knevel R. Improving musculoskeletal care with AI enhanced triage through data driven screening of referral letters. \u003cem\u003eNPJ Digit Med\u003c/em\u003e. 2025;8(1). doi:10.1038/s41746-025-01495-4\u003c/li\u003e\n\u003cli\u003eAbdel-Hafez A, Jones M, Ebrahimabadi M, et al. Artificial intelligence in medical referrals triage based on Clinical Prioritization Criteria. \u003cem\u003eFront Digit Health\u003c/em\u003e. 2023;5. doi:10.3389/fdgth.2023.1192975\u003c/li\u003e\n\u003cli\u003eSangwon KL, Han X, Becker A, et al. Automating the Referral of Bone Metastases Patients With and Without the Use of Large Language Models. \u003cem\u003eNeurosurgery\u003c/em\u003e. Published online 2025. doi:10.1227/neu.0000000000003683\u003c/li\u003e\n\u003cli\u003eBienefeld N, Boss JM, L\u0026uuml;thy R, et al. Solving the explainable AI conundrum by bridging clinicians\u0026rsquo; needs and developers\u0026rsquo; goals. \u003cem\u003eNPJ Digit Med\u003c/em\u003e. 2023;6(1). doi:10.1038/s41746-023-00837-4\u003c/li\u003e\n\u003cli\u003eHuang J, Wittbrodt MT, Teague CN, et al. Efficiency and Quality of Generative AI-Assisted Radiograph Reporting. \u003cem\u003eJAMA Netw Open\u003c/em\u003e. 2025;8(6). doi:10.1001/jamanetworkopen.2025.13921\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"journal-of-medical-systems","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Journal of Medical Systems](https://www.springer.com/journal/10916)","snPcode":"10916","submissionUrl":"https://submission.nature.com/new-submission/10916/3","title":"Journal of Medical Systems","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-7926103/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7926103/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDelays in the referral of patients with suspected metastatic bone disease (MBD) from radiology reports represent a critical challenge that can negatively impact patient outcomes. The conventional manual review process is often a significant bottleneck, leading to prolonged diagnostic timelines. We developed and validated SHIELD, an automated AI framework designed to accelerate and improve the accuracy of MBD referrals. We fine-tuned a RadBERT-RoBERTa model on a decade of radiology reports (N\u0026thinsp;=\u0026thinsp;245 patients) from two academic medical centers to classify reports into three tiers: \"No Referral,\" \"Referral,\" and \"Referral/High Risk.\" To ensure clinical utility and transparency, SHIELD incorporates a Large Language Model to generate natural-language explanations for its classifications. SHIELD demonstrated exceptional performance on a hold-out test set. It achieved 100% accuracy and an Area Under the Curve (AUC) of 1.00 in the primary binary task of distinguishing referral from non-referral cases. In the more granular three-class task, the model achieved an overall accuracy of 89.52%, with near-perfect performance in identifying \"No Referral\" reports (F1-score: 99.20%). Critically, the model operated in a clinically \"fail-safe\" manner, never misclassifying a high-risk case as requiring no referral. A retrospective timeline analysis revealed that SHIELD can reduce the referral period from a conventional average of 109.6 days to a computational time of 1\u0026ndash;3 minutes. Proposed work provides high accuracy with a sophisticated explainability component using a large language model. Thus, SHIELD framework is a robust, explainable, and autonomous solution for triaging radiology reports. By drastically reducing administrative and diagnostic delays, it has the potential to significantly accelerate the clinical workflow, ensure timely specialist consultation, and ultimately improve the standard of care for patients with suspected MBD.\u003c/p\u003e","manuscriptTitle":"SHIELD: An AI Framework for Skeletal Health Intelligence and Early Lesion Detection to Improve Orthopedic Referrals","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-11 16:20:45","doi":"10.21203/rs.3.rs-7926103/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorAssigned","content":"","date":"2025-11-06T05:10:55+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-11-05T23:15:58+00:00","index":"","fulltext":""},{"type":"submitted","content":"Journal of Medical Systems","date":"2025-10-22T18:23:38+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"journal-of-medical-systems","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Journal of Medical Systems](https://www.springer.com/journal/10916)","snPcode":"10916","submissionUrl":"https://submission.nature.com/new-submission/10916/3","title":"Journal of Medical Systems","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"738f914e-1ad6-452c-bd09-7cc12a684b29","owner":[],"postedDate":"November 11th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-03-27T04:39:29+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-11 16:20:45","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7926103","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7926103","identity":"rs-7926103","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00