Risk Stratification of Breast Cancer Metastasis: A Predictive Modelling Framework Using Clinical and Hormonal Receptor Data in a Ghanaian Cohort | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Risk Stratification of Breast Cancer Metastasis: A Predictive Modelling Framework Using Clinical and Hormonal Receptor Data in a Ghanaian Cohort Abdulzeid Yen Anafo, Senyefia Bosson-Amedenu, Emmanuel Ayitey, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6671298/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Breast cancer remains the most diagnosed cancer among women globally. It is a leading cause of cancer-related deaths, with disproportionately high mortality rates in sub-Saharan Africa due to late-stage presentation and limited access to diagnostic and treatment services. In Ghana, nearly 70% of breast cancer cases are diagnosed at advanced or metastatic stages, underscoring the urgent need for accurate and context-specific prognostic tools. Here, we assembled clinical, molecular and demographic data from 558 breast cancer patients at Korle-Bu Teaching Hospital in Ghana to develop and compare five supervised machine‐learning model logistic regression, random forest, support vector machine, XGBoost and naïve Bayes for metastasis prediction. Using standardized preprocessing and repeated 10-fold stratified cross-validation, logistic regression achieved the most balanced performance (96% cross-validated accuracy; 93% test‐set accuracy) and facilitated interpretability via SHAP analysis, which identified lymph node involvement, tumour size and stage as top predictors. We further refined prognostic utility by stratifying patients into low, intermediate and high metastatic risk groups, revealing distinct clinical and biological profiles high‐risk individuals exhibited aggressive subtypes (triple-negative, HER2-positive) and advanced staging, whereas low-risk patients showed favourable receptor status and earlier disease. These findings demonstrate that tailored, interpretable machine-learning tools can support accurate, resource-appropriate metastasis risk assessment in low-resource settings. Breast Cancer Metastasis Machine Learning Risk Stratification Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Breast cancer (BC) stands as the most diagnosed cancer and the top cause of cancer-related deaths among women worldwide, exhibiting significant disparities in outcomes across different regions [ 1 ]. High-income countries have made substantial progress in early detection and improving survival rates, achieving five-year survival rates that exceed 80% due to widespread screening and timely treatment interventions [ 2 ]. In contrast, sub-Saharan Africa experiences alarmingly high breast cancer mortality rates, primarily attributed to late-stage diagnosis resulting from limited access to screening programs, and significant delays in treatment [ 3 ]. In Ghana specifically, BC constitutes nearly one-third of all cancers affecting women, and a large proportion of cases (approximately 70%) are diagnosed at advanced or metastatic stages, leading to poor prognosis and high mortality rates [ 4 ], [ 5 ]. BC Metastasis also known as stage IV breast cancer, occurs when cancer cells spread beyond the breast and regional lymph nodes to distant organs such as the bones, liver, lungs, and brain [ 6 ]. The prognosis for patients with MBC remains challenging, with survival rates significantly lower than for early-stage disease. For instance, the five-year relative survival rate for metastatic breast cancer in high-income countries is approximately 29%, highlighting the severity of this stage [ 7 ]. In low-resource settings like Ghana, survival outcomes are further compromised due to barriers such as inadequate healthcare infrastructure, limited access to advanced treatments, high costs of medications, and delays in seeking medical care [ 5 ], [ 8 ], [ 9 ]. Additionally, socio-cultural factors and misconceptions surrounding breast cancer often discourage women from early healthcare engagement, exacerbating the severity of disease presentation at diagnosis and contributing to poor clinical outcomes [ 4 ]. Consequently, addressing the metastatic burden of breast cancer in regions such as Ghana requires multifaceted approaches, including enhanced public awareness, improved screening programs, timely referral systems, and equitable access to effective therapeutic interventions [ 9 ]. Accurate metastasis prediction is vital for optimizing treatment choices, resource distribution, and patient outcomes [ 10 ]. Conventional prognostic approaches heavily depend on clinical and pathological staging, hormone receptor status, and patient history to inform management strategies [ 11 ]. Nonetheless, these traditional methods might not fully reflect the complexity and diversity of breast cancer progression, particularly in resource-constrained environments where genomic profiling and advanced diagnostic tools are often unavailable [ 9 ], [ 12 ]. Machine learning (ML) techniques offer a promising alternative, using data-driven methods to reveal subtle and intricate patterns in clinical data that typical analysis might miss [ 13 ], [ 14 ]. These models can considerably enhance decision-making by effectively pinpointing high-risk individuals, allowing for personalized and timely management strategies adapted to individual risk profiles [ 15 ]. While there is increasing interest in machine learning (ML) tools for oncology, a notable shortage exists in predictive models tailored to African populations with breast cancer. Most current models derive from Western datasets, potentially misrepresenting the genetic, environmental, and healthcare characteristics unique to African communities. This study seeks to address this deficiency by creating and assessing machine learning models focused on predicting breast cancer metastasis, utilising clinical, molecular, and demographic data from a cohort in Ghana. Methods This research was a single-institution-based quantitative retrospective cohort study conducted among patients with breast cancer. We employed a predictive modelling framework featuring five supervised machine learning classifiers: Logistic Regression, Random Forest, Support Vector Machine, XGBoost, and Naive Bayes, specifically analysing clinical and hormonal receptor data from Ghanaian patients. Furthermore, the research intends to develop a practical, data-oriented tool for assessing metastasis risk, enhancing precision medicine efforts in low- and middle-income nations. Data Collection This study utilized secondary data from 558 breast cancer patients, aged 13 to 97 years, diagnosed at Korle-Bu Teaching Hospital in Ghana. The dataset included demographic information (such as age and menopausal status), clinical features (including tumor size, cancer stage, and lymph node involvement), and molecular characteristics (estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER) status). Data also captured tumor subtype classifications, recurrence events, survival time intervals, and mortality outcomes due to breast cancer or competing causes. Data Preprocessing Before model development, the dataset underwent several preprocessing steps to ensure compatibility with machine learning algorithms and optimize model performance. Continuous variables, such as tumor size and age at diagnosis, were standardized using the StandardScaler from scikit-learn [16], ensuring that all features contributed proportionally to the analysis without being influenced by differences in measurement scales. Categorical variables, including hormone receptor status (ER, PR, HER2), menopausal status, cancer stage, and tumor subtype, were encoded using one-hot encoding to transform them into a numerical format suitable for model training [17]. These preprocessing steps ensured that both clinical and molecular features were appropriately prepared for the subsequent machine learning analyses, allowing models to capture patterns effectively and improve prediction accuracy [15] Model Development and Selection This study assessed various models to analyse their effectiveness in predicting metastasis among breast cancer patients. The models examined include Logistic Regression (LR), Linear Support Vector Machine (SVM), Random Forest (RF), XGBoost, and Naïve Bayes classifiers, selected for their complementary abilities in conducting binary classification tasks typical in oncology research. Logistic Regression is favoured in cancer research because of its simplicity and clarity in calculating odds ratios and pinpointing crucial predictors [18]. Linear-kernel support Vector Machines are particularly adept at handling high-dimensional clinical data, providing strong outcomes in differentiating between metastatic and non-metastatic instances [19]. Meanwhile, ensemble models like Random Forest and XGBoost excel in predictive accuracy for cancer prognosis, thanks to their capacity to model intricate interactions and mitigate overfitting [20]. Although Naïve Bayes relies on the assumption of feature independence, it remains a useful tool in medical settings for its efficiency and surprisingly competitive accuracy in specific cancer prediction scenarios [15]. Model Evaluation and Cross-Validation Methodology We developed a thorough evaluation framework to establish the reliability and adaptability of our predictive models. We used a 10-fold stratified cross-validation method to ensure that the distribution of metastatic and non-metastatic cases remained consistent across all folds. This stratification is essential in medical datasets with imbalanced classes, as it guarantees fair representation of the minority group and mitigates performance bias [21]. To further bolster stability, we repeated the cross-validation process three times with different random seeds, a technique that reduces variance in model performance estimates and yields more dependable results [22]. In each fold, hyperparameters were fine-tuned to enhance model performance. The C and gamma, the penalty parameter and kernel coefficient, respectively, were optimized for the SVM, while n_estimators and max_depth were adjusted for the Random Forest and XGBoost classifiers [20]. We computed evaluation metrics such as accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (AUC) for each fold, averaging them to comprehensively assess model effectiveness. We performed a final assessment using a distinct test set following the training phase. We created confusion matrices for each model to provide a comprehensive analysis of classification performance, detailing rates of true positives, true negatives, false positives, and false negatives. Additionally, we plotted ROC curves to demonstrate each model’s ability to distinguish between metastatic and non-metastatic cases at different thresholds. We particularly focused on precision, recall, and F1-score due to their critical importance in clinical decision-making, where misclassifications especially false negatives can significantly impact patient outcomes. This methodology aligns with recent studies emphasising the importance of thorough evaluation metrics in medical machine learning [15]. Risk Stratification After evaluating the models, predictive algorithms categorised BC patients into significant risk categories based on their metastasis likelihood. This stratification process started by generating individual risk scores from each model's probabilistic outputs. These scores reflected each patient's predicted probability of metastatic progression, acting as an indicator of metastasis risk. To enhance understandability and aid clinical decisions, these continuous risk scores were grouped into three clear categories: low, intermediate, and high. The risk stratification method, utilising predicted probabilities, is commonly employed in cancer research to improve clinical decision-making and patient counseling. Risk scores are typically grouped into discrete categories- low, intermediate, and high risk- making them practical for clinical use [23]. This approach has been applied in breast cancer research to pinpoint patients at increased risk of metastasis or recurrence, enabling more tailored surveillance and treatment strategies [24]. Stratification thresholds are usually based on clinical relevance or the distribution characteristics of predicted probabilities, as illustrated by various machine learning models used for breast cancer prognosis [25]. Patients who scored below 0.33 were placed in the low-risk category, suggesting a lower chance of developing metastasis. Those with scores between 0.33 and 0.66 were labelled as intermediate risk, indicating a moderate likelihood of progression. Scores above 0.66 were categorised as high risk, signifying patients with a considerable probability of metastasis. SHAP Analysis To improve interpretability and quantify how individual features affect model predictions, we utilised SHAP (SHapley Additive exPlanations) analysis, a cohesive framework for understanding predictions from machine learning models [26]. SHAP values were calculated for the top-performing model among LR, RF, XGBoost, SVM, and Naïve Bayes using the TreeExplainer and LinearExplainer methods provided in the SHAP Python library based on model compatibility. This technique offers both global and local interpretability by attributing a Shapley value to each feature, indicating its average impact on the predicted metastasis risk for each patient. The SHAP beeswarm plot was employed to visualise impact, distribution and direction of each feature across the dataset. In contrast, the SHAP summary bar plot depicted the mean absolute importance of every feature. This analysis was crucial for validating the clinical significance of the model by pinpointing which variables had the strongest influence on metastasis predictions. Statistical Analysis Python (version 3.8), using libraries such as scikit-learn, NumPy, Pandas, Matplotlib, and Seaborn, was used for all data manipulations, analyses, and visualizations. This comprehensive suite of tools facilitated rigorous statistical analysis and ensured reproducibility of the results . Results Baseline Characteristics of Participants Table 1 below presents the baseline characteristics of participants. Among the 558 patients included in the study, 281 (50.4%) experienced metastases while 277 (49.6%) did not. Patients with metastasis were significantly older than those without spread, with a mean age of 52.3 ± 15.4 years compared to 47.6 ± 12.6 years (p < 0.001). A greater proportion of patients aged above 50 years was observed in the metastasis group (50.5%) compared to the no spread group (40.1%, p = 0.013). Recurrent disease was significantly more frequent among patients with metastasis (13.5% vs 3.6%, p < 0.001). Competing risk analysis revealed that 85.8% of metastasis patients experienced an event compared to none in the no spread group (p < 0.001). Hormonal receptor status differed markedly between groups: estrogen receptor (ER) positivity and progesterone receptor (PR) positivity were both significantly lower among metastasis patients (ER+: 43.1% vs 66.4%; PR+: 41.6% vs 64.3%; both p < 0.001). Human epidermal growth factor receptor 2 (HER2) status, however, did not differ significantly between groups (p = 0.230). With respect to molecular subtypes, triple-negative breast cancer was notably more frequent in the metastasis group (52.7% vs 33.2%), while Luminal-A subtype predominated among patients without metastasis (45.5% vs 18.9%; p < 0.001). Higher tumor grades (Grade 3) were significantly associated with metastasis (57.3% vs 35.0%, p < 0.001). Advanced tumor stages were also more prevalent in metastasis cases, with 69.4% at Stage 3 and 20.6% at Stage 4 compared to only 5.8% and 0.7% respectively among patients without spread (p < 0.001). Tumor size distribution showed that larger tumors (51-60 mm and above 60 mm) were predominantly found among metastasis cases (69.4% and 20.6%, respectively; p < 0.001). Lymph node involvement was significantly greater in the metastasis group, with 46.6% having two nodes involved and 33.1% having three nodes involved (p < 0.001). Menopausal status did not significantly differ between groups (p = 0.400). Additionally, hospitalization during the disease course was significantly more common among metastasis patients (33.5% vs none, p < 0.001), and the presence of Breast CAncer gene 1 ( BRCA1) genetic traces was more frequent among metastasis cases (52.7% vs 33.2%, p < 0.001). Table 1 : Baseline Characteristics of Participants Variable Total (N=558) No Metastasis(n=277) Metastasis (n=281) p-value Age, mean ± SD 50.0 ± 14.3 47.6 ± 12.6 52.3 ± 15.4 <0.001 1 Age, median (IQR) 49.0 (40.0, 58.0) 47.0 (38.0, 56.0) 51.0 (41.0, 62.0) <0.001 2 n (%) n (%) n (%) Age categories 0.013 3 13-50 305 (54.7) 166 (59.9) 139 (49.5) Above 50 253 (45.3) 111 (40.1) 142 (50.5) Ethnicity <0.001 3 Other 111 (19.9) 66 (23.8) 45 (16.0) Akan 217 (38.9) 88 (31.8) 129 (45.9) Ga/Adangbe 131 (23.5) 63 (22.7) 68 (24.2) Ewe 99 (17.7) 60 (21.7) 39 (13.9) Recurrent <0.001 3 No 510 (91.4) 267 (96.4) 243 (86.5) Yes 48 (8.6) 10 (3.6) 38 (13.5) Competing Risk <0.001 4 Censored 277 (49.6) 277 (100.0) 0 (0.0) Event 241 (43.2) 0 (0.0) 241 (85.8) Death due to other diseases 40 (7.2) 0 (0.0) 40 (14.2) Human ER2 0.230 3 Negative 429 (76.9) 219 (79.1) 210 (74.7) Positive 129 (23.1) 58 (20.9) 71 (25.3) ER <0.001 3 Negative 253 (45.3) 93 (33.6) 160 (56.9) Positive 305 (54.7) 184 (66.4) 121 (43.1) PR <0.001 3 Negative 263 (47.1) 99 (35.7) 164 (58.4) Positive 295 (52.9) 178 (64.3) 117 (41.6) MSubtype <0.001 3 Luminal-A 179 (32.1) 126 (45.5) 53 (18.9) Triple negative 240 (43.0) 92 (33.2) 148 (52.7) Luminal-B 121 (21.7) 54 (19.5) 67 (23.8) Her2 Negative 18 (3.2) 5 (1.8) 13 (4.6) Grade <0.001 3 Grade1 179 (32.1) 126 (45.5) 53 (18.9) Grade 2 121 (21.7) 54 (19.5) 67 (23.8) Grade 3 258 (46.2) 97 (35.0) 161 (57.3) Type of Breast Cancer <0.001 4 IDC 424 (76.0) 179 (64.6) 245 (87.2) LCIS 8 (1.4) 5 (1.8) 3 (1.1) IBC 40 (7.2) 30 (10.8) 10 (3.6) MBC 27 (4.8) 18 (6.5) 9 (3.2) TCB 2 (0.4) 2 (0.7) 0 (0.0) ILC 32 (5.7) 21 (7.6) 11 (3.9) DCIS 25 (4.5) 22 (7.9) 3 (1.1) Stage, mean ± SD 2.3 ± 1.1 1.5 ± 0.8 3.1 ± 0.6 <0.001 1 Stage <0.001 4 0 31 (5.6) 28 (10.1) 3 (1.1) 1 101 (18.1) 99 (35.7) 2 (0.7) 2 155 (27.8) 132 (47.7) 23 (8.2) 3 211 (37.8) 16 (5.8) 195 (69.4) 4 60 (10.8) 2 (0.7) 58 (20.6) Status <0.001 4 Survived 275 (49.3) 272 (98.2) 3 (1.1) Died 283 (50.7) 5 (1.8) 278 (98.9) Lymph Node <0.001 4 0 84 (15.1) 84 (30.3) 0 (0.0) 1 214 (38.4) 157 (56.7) 57 (20.3) 2 158 (28.3) 27 (9.7) 131 (46.6) 3 102 (18.3) 9 (3.2) 93 (33.1) Menopause 0.400 3 No 194 (34.8) 101 (36.5) 93 (33.1) Yes 364 (65.2) 176 (63.5) 188 (66.9) Hospitalized <0.001 4 No 464 (83.2) 277 (100.0) 187 (66.5) Yes 94 (16.8) 0 (0.0) 94 (33.5) Tumor Size <0.001 4 5-10 mm 31 (5.6) 28 (10.1) 3 (1.1) 11-20 mm 101 (18.1) 99 (35.7) 2 (0.7) 201-50 mm 155 (27.8) 132 (47.7) 23 (8.2) 51-60 mm 211 (37.8) 16 (5.8) 195 (69.4) Above 60 mm 60 (10.8) 2 (0.7) 58 (20.6) Genetics <0.001 3 Braca 1 traces 240 (43.0) 92 (33.2) 148 (52.7) No traces 318 (57.0) 185 (66.8) 133 (47.3) 1 t-test p-value; 2 Wilcoxon Rank-sum test p-value; 3 Chi-square test p-value; 4 Fisher’s exact test p-value Model Evaluation The Table 2 below presents the performance evaluation of five machine learning models (Logistic Regression, Random Forest, Support Vector Machine (SVM), XGBoost, and Naive Bayes). Logistic Regression and SVM achieved the highest accuracy rates, at 0.93 and 0.91, respectively, with SVM also recording the highest area under the curve (AUC = 0.98), indicating exceptional discriminatory ability. Random Forest exhibited strong precision (0.95) and specificity (0.94), which enhanced its effectiveness in accurately identifying non-metastatic conditions. Although XGBoost and Naive Bayes displayed notable recall values (0.96 and 0.95, respectively), they had lower specificity scores, indicating a propensity to over-predict metastatic cases. Nonetheless, their high recall demonstrates their usefulness in clinical scenarios where reducing false negatives is essential. In conclusion, Logistic Regression proved to be the most balanced model, excelling across all metrics, including a cross-validation accuracy of 0.96. Table 2 :Performance evaluation of five machine learning models; LR, RF, SVM, XGBoost, and Naive Bayes for predicting breast cancer metastasis in a Ghanaian cohort. The models were assessed using accuracy, specificity, precision, recall, F1-score, AUC, and 10-fo Models Accuracy Specificity Precision Recall F-Measure AUC 10-fold cross validation accuracy LR 0.93 0.92 0.92 0.94 0.93 0.94 0.96 RF 0.90 0.94 0.95 0.94 0.91 0.97 0.94 SVM 0.91 0.92 0.92 0.91 0.92 0.98 0.92 XGBoost 0.88 0.80 0.83 0.96 0.89 0.95 0.87 Naives Bayes 0.87 0.79 0.82 0.95 0.88 0.97 0.91 The confusion matrices presented in Figure 1 below illustrate the classification performance of five machine learning models for predicting metastasis in a Ghanaian breast cancer cohort. Logistic Regression and SVM models exhibited the most balanced and accurate predictions, each achieving 34 true positives and 49 true negatives, with minimal false classifications, highlighting their reliability in distinguishing metastatic from non-metastatic cases. The Random Forest model also performed well, achieving similar true positive rates but with a slight increase in false positives. On the other hand, Naive Bayes displayed the weakest performance, with a notable rise in false negatives (7), which is a significant concern in medical diagnostics, where missing a metastasis case could delay treatment. XGBoost, while maintaining a strong true positive count, suffered from the highest false positive rate, suggesting it might overpredict metastasis. Feature importance A SHAP analysis was conducted to determine and quantify the impact of individual logistic model features on predicting breast cancer metastasis within the Ghanaian cohort. Figure 2 displays the SHAP beeswarm plot alongside the corresponding bar plot of mean SHAP values. The Logistic Regression model reveals that Lymph Node Involvement (LymphNode) is the most influential feature, with a mean SHAP value of approximately 1.38, indicating a strong impact on metastasis prediction. This is followed closely by Tumor Size (TumorSize) at 1.26 and Cancer Stage (Stage) at 1.15, both of which are clinically established predictors of metastasis. Age at Cancer Treatment (AgeCAT) contributes moderately with a SHAP value of 0.61, along with Time Since Diagnosis (FlagTime) at 0.51 and Type of Breast Cancer (TypeofBC) at 0.46. Genetic Predisposition (Genetics) also plays a notable role (0.46), as do Menopausal Status (Menopause) at 0.29 and Estrogen Receptor Status (ER) at 0.27. Other features such as Molecular Subtype (MSubType), Ethnicity, Progesterone Receptor Status (PR), and Recurrent Tumour (Recurrent) have smaller contributions (SHAP values between 0.1 and 0.25). On the lowest end of influence are Patient Age (Age), Tumor Grade (Grade), and Human Epidermal Growth Factor Receptor 2 (HER2), each with mean SHAP values below 0.1, suggesting minimal predictive weight in this model. The SHAP beeswarm plot for the logistic regression model illustrates the global impact of each feature on metastasis prediction. Lymph node involvement emerged as the most influential predictor, where higher values (shown in red) strongly increased the predicted risk of metastasis, while lower values (blue) contributed negatively. Tumor size and cancer stage followed closely, with larger tumors and advanced stages (also represented in red) consistently shifting predictions toward AgeCAT and genetic factors showed moderate influence, indicating that demographic groupings and hereditary risk may modestly affect model output. TypeofBC also had a noticeable impact, especially certain subtypes associated with greater metastatic potential. In contrast, features such as follow-up time, ER status, PR status, menopausal status, and ethnicity exhibited limited influence, with SHAP values close to zero and mixed red-blue distributions, suggesting variability without strong directional impact. The least impactful features were age, tumor grade, and HER2 status, indicating minimal contribution to the logistic regression model’s predictions. Risk stratification The risk stratification outcomes from the Logistic Regression model highlight significant, biologically relevant patterns across clinical, hormonal, molecular, and demographic variables. Figures 3 and 4 depict variations in general clinical and demographic traits, presenting selected molecular and histopathological characteristics. Additional plots on risk stratification for other ML models are displayed in the supplementary file as Figure 5S-8S. Figure 3A shows that the Akan ethnic group exhibits an increasing prevalence from Low Risk (31%) to High Risk (46%), while the Ga-Adangbe and Ewe groups maintain relative consistency across categories. Other ethnic groups show a higher representation in the Low (24%) and Intermediate Risk (12%) categories than in the High-Risk category (17%). Regarding genetics, family history is notably more common among Low (69%) and Intermediate Risk (65%) individuals compared to those in the High-Risk group (55%). The occurrence of BRCA mutations rises steadily with increasing risk: 31% in Low Risk, 35% in Intermediate, and 55% in High Risk. Tumor grades demonstrate a clear trend, with Grade 1 being the most prevalent in Low Risk (47%) and decreasing in higher risk groups, while Grade 3 rises with risk levels: 33% in Low, 42% in Intermediate, and 59% in High Risk. Hospitalisation trends also mirror risk stratification: 100% of Low-Risk individuals were not hospitalised, compared to 96% in Intermediate and only 66% in the High-Risk group, with hospitalisation increasing to 34% in High-Risk individuals. Menopausal status reveals that post-menopausal patients are more frequent in the Intermediate (77%) and High Risk (66%) groups, while pre-menopausal individuals are more common in the High-Risk group (34%) compared to Low (36%) and Intermediate (23%) groups. Tumor staging indicates that Stage 3 is prevalent in High Risk (71%) and Intermediate (46%) categories, while early stages (Stage 0 and 1) are nearly absent in the High-Risk group. Stage 4 is primarily found in the High-Risk group (22%). Figure 3B reinforces the existence of biologically significant patterns across clinical, hormonal, molecular, and demographic features. As illustrated in Figures 3 and 4, high-risk patients significantly exhibit numerous aggressive disease markers. Within the high-risk group, 53% of patients experienced lymph node involvement, contrasting with 33% in the moderate-risk group and just 12% in the low-risk group. Likewise, tumors exceeding 5 cm were more common in high-risk patients (44%) compared to moderate-risk (27%) and low-risk patients (13%). The proportion of patients presenting with stage 3 or higher disease was 52% in the high-risk group, 35% in the moderate-risk group, and 17% in the low-risk category. In terms of recurrence and menopausal status, 16% of high-risk patients experienced recurrence, compared to 7% in the moderate-risk and 5% in the low-risk groups. Menopausal women represented 63% of the high-risk category, versus 50% in the moderate-risk group and 44% in the low-risk group. Notably, genetic predisposition was observed in 12% of high-risk patients, compared to only 4% and 6% in moderate and low-risk groups, respectively. The use of hormonal contraceptives showed a downward trend with increasing risk, being most prevalent in the low-risk group (64%), decreasing to 49% in the moderate-risk group, and just 42% in the high-risk category. Regarding clinical and molecular traits, 84% of low-risk patients were ER-positive, 82% were PR-positive, and 91% were HER2-negative, exhibiting a hormonal receptor-rich profile associated with better outcomes. Additionally, 73% of low-risk patients were categorized as Luminal A subtype, while only 1% were classified as triple-negative breast cancer (TNBC). In contrast, the high-risk group displayed more aggressive phenotypes: only 44% were ER-positive and 41% were PR-positive, while 29% were HER2-positive, and 25% had TNBC, marking a substantial increase compared to the low-risk group. The Luminal B subtype was distributed more evenly, with 13% in the low-risk, 22% in moderate-risk, and 21% in high-risk patients. Significantly, stage III and stage IV diseases were much more prevalent in the high-risk group 29% and 12%, respectively compared to just 6% and 1% in the low-risk category. Similarly, Grade 3 tumors occurred in 29% of high-risk patients, in contrast to only 6% of low-risk patients. Discussions Breast cancer in sub-Saharan Africa carries a disproportionately high mortality rate, yet most clinical decision-support tools for predicting metastasis have been developed and validated on Western cohorts. As a result, existing models may not account for the unique genetic, environmental, and health-system factors that influence disease progression in West African populations. Our study therefore fills a critical gap by presenting a comprehensive comparative evaluation of five machine-learning models Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), XGBoost, and Naive Bayes (NB trained and tested on clinical, molecular, and demographic data from a Ghanaian breast cancer cohort. By tailoring predictive algorithms to a local dataset, we aim to improve the accuracy of metastasis risk stratification in Ghana and ultimately support better patient management and resource allocation. Among all the models, Logistic Regression exhibited the most balanced and consistent performance across metrics, achieving high accuracy, recall, specificity, precision, and the highest 10-fold cross-validation score. These findings suggest that Logistic Regression, despite being a relatively simple and interpretable model, is highly effective in metastatic risk prediction when well-calibrated and trained on context-specific data [ 18 ]. The model's balance between sensitivity and specificity is particularly crucial in resource-constrained settings like Ghana, where diagnostic accuracy must be matched with practical decision-making efficiency to minimize both over-treatment and missed diagnoses. SVM showed the highest area under the ROC curve, indicating superior discriminatory ability between metastatic and non-metastatic cases. While its recall and precision were slightly lower than LR, its high AUC underscores its utility in settings where minimizing classification overlap is a priority. In agreement with prior studies [ 13 ], [ 15 ], SVM's strength lies in handling complex decision boundaries in high-dimensional clinical data. Random Forest, an ensemble learning method, excelled in specificity and precision, making it particularly suited for accurately identifying non-metastatic cases. This performance aligns with other research indicating RF’s robustness in minimizing false positives and enhancing model generalization [ 27 ]. However, its slightly lower recall compared to XGBoost and Naive Bayes implies that it may still miss a few metastatic cases. Interestingly, both XGBoost and Naive Bayes demonstrated exceptionally high recall, which is crucial in cancer prognostics where false negatives can lead to delayed intervention and worse outcomes. However, their lower specificity suggests a tendency to overpredict metastasis, resulting in a higher rate of false positives. In clinical settings, these models may serve effectively as initial screening tools that prioritize sensitivity, followed by confirmation with more balanced models like LR or RF [ 28 ]. The confusion matrices reinforce these patterns. Logistic Regression and SVM correctly identified positives and negatives, showing minimal misclassification. Random Forest also maintained a strong classification balance, though it exhibited a modest rise in false positives. Conversely, Naive Bayes produced a relatively higher number of false negatives, which raises clinical concerns in oncology where missing a metastasis can delay treatment. XGBoost demonstrated the highest number of false positives, suggesting its application would need to be complemented with confirmatory diagnostics in practice. These findings support the notion that model selection must be context-sensitive: while high recall models may be appropriate for screening, models with higher specificity are more suitable for confirmatory decision support in oncology [ 29 ]. To enhance the interpretability and clinical applicability of the logistic regression model for predicting breast cancer metastasis in a Ghanaian cohort, SHAP analysis was employed. This approach offers transparent, locally accurate explanations of model outputs and ranks features based on their individual contributions to prediction [ 26 ]. Such interpretability is crucial for building clinician trust and ensuring responsible deployment of machine learning in resource-constrained healthcare environments. The analysis revealed several clinically and contextually significant outcomes. Most notably, lymph node involvement was identified as the strongest driver of metastasis prediction. This finding aligns with global oncology literature, which consistently recognizes lymphatic spread as a principal route of breast cancer metastasis [ 29 ]. The SHAP beeswarm plot further confirmed that high values of lymph node involvement were consistently associated with increased predicted risk, validating the model’s biological coherence. Closely following were tumor size and cancer stage, both critical indicators in the TNM staging system [ 30 ]. These features were more influential than molecular markers such as ER, PR, and HER2, a trend that diverges from models trained on Western datasets. This distinction is particularly important in the Ghanaian context, where late-stage presentation is common due to limited screening and awareness [ 31 ]. In such settings, macroscopic clinical features (size and spread) may outweigh molecular subtypes in determining metastatic risk. Another notable outcome was that age categorized by treatment group proved significantly more predictive than raw age. This suggests that age-related risk is not linear, and stratified clinical groupings (under 40, 40–59, 60+) capture meaningful variation, particularly since younger African women often present with more aggressive disease [ 32 ]. Similarly, time since diagnosis was a meaningful factor, reflecting how disease progression may impact metastatic outcomes. Moderately influential features included type of breast cancer and genetic predisposition. Even in the absence of full genomic sequencing, self-reported family history and clinical categorization proved valuable, suggesting that structured clinical history is a low-cost but high-value addition to diagnostic modeling in resource-limited settings. Menopausal status and estrogen receptor status also played moderate roles, aligning with their known, but context-dependent, influence on recurrence and treatment response [ 33 ]. Interestingly, HER2 status and tumor grade, which are widely regarded as aggressive disease markers in Western literature [ 34 ], had minimal predictive value in this model. This may reflect either true epidemiological differences or systemic gaps in HER2 testing and data completeness in Ghanaian clinical practice [ 35 ]. It also raises questions about the generalizability of global biomarkers and highlights the need for locally validated tools. Similarly, ethnicity had minimal impact on model prediction. This suggests that while ethnic group classifications (Akan, Ewe, Ga-Adangbe and others) are socially and demographically relevant, they do not appear to independently predict metastasis in this cohort, reinforcing that biological and clinical characteristics carry more weight than sociocultural identifiers in risk modelling. The SHAP beeswarm plot provided valuable visual insights into the directionality and magnitude of feature impacts. Features with high SHAP values such as lymph node involvement, tumour size, and stage, strongly increased metastasis predictions. In contrast, features like ER, PR, menopause, and ethnicity exhibited near-zero SHAP values with mixed distributions, indicating limited or inconsistent contribution to model output. These patterns are not only statistically informative but also clinically actionable. The model emphasizes a clear set of high-priority variables nodal involvement, tumor size, stage, and categorized age that can be reliably assessed even in low-resource environments. These variables should be prioritized in screening, triage, and early intervention programs. Moreover, the findings highlight that simple, interpretable model such as logistic regression, when supported by SHAP, can replicate core clinical knowledge and offer actionable insights for oncological care in low- and middle-income countries (LMICs). The model’s alignment with established clinical pathways bolsters its validity and facilitates integration into decision-support systems. The risk stratification analysis via Logistic Regression revealed biologically and clinically relevant distinctions. on clinical and hormonal variables differ across breast cancer metastasis risk groups. These differences highlight established predictors of metastasis and may guide risk-tailored clinical decisions. High-risk individuals tend to be younger on average than those in the low-risk group. This finding aligns with prior literature suggesting that early-onset breast cancers may exhibit more aggressive biological behavior, particularly in triple-negative and HER2-positive subtypes, which are more common in younger patients [ 32 ]. Patients in the high-risk group exhibit significantly higher cancer stages, indicating later diagnosis or more advanced disease progression. Stage at diagnosis is a well-established prognostic factor, with higher stages correlating with greater tumor burden and metastatic potential [ 36 ]. This underscores the importance of early detection and screening in improving outcomes. Lymph node involvement is markedly higher among high-risk patients. Axillary lymph node status is one of the most reliable indicators of metastasis risk, as nodal metastasis often precedes distant spread [ 37 ]. This supports the high importance of SHAP, as seen in earlier model interpretations. Larger tumors (> 5 cm) are overrepresented in the high-risk group, while smaller tumors (< 2 cm) dominate the low-risk group. Tumor size is a core element of Tumor-Node-Metastasis TNM staging and has been directly correlated with poor prognosis due to its association with increased cellular proliferation and angiogenesis [ 33 ]. There is a higher proportion of PR-negative patients in the high-risk category. PR-negativity is a hallmark of more aggressive, hormone-independent tumors and is frequently observed in conjunction with HER2 positivity or triple-negative status, both linked to early recurrence and metastasis [ 38 ]. Pre-menopausal women are disproportionately represented in the high-risk group. Hormonal fluctuations and denser breast tissue in pre-menopausal women may contribute to diagnostic delays and more aggressive tumor subtypes [ 39 ]. This observation complements the age distribution and further emphasizes the need for age-specific screening strategies. The risk stratification analysis via Logistic Regression revealed biologically and clinically relevant distinctions on clinical and hormonal variables differ across breast cancer metastasis risk groups. Several key observations emerged from the stratified risk analysis that offer deeper insights into breast cancer metastasis patterns within the studied cohort. While the high-risk group exhibited the highest prevalence of genetic mutations, a notable and somewhat unexpected elevation in mutation frequency was also observed among intermediate-risk patients compared to low-risk patients. This suggests that some biologically predisposed individuals may be under-classified based on clinical features alone, highlighting the need to refine risk stratification by incorporating molecular markers [ 40 ], [ 41 ]. Additionally, the intermediate-risk group displayed considerable ethnic diversity, unlike the more demographically concentrated high-risk group, pointing toward potential disparities in healthcare access and diagnostic timing that may influence clinical risk categorization [ 42 ], [ 43 ]. Hospitalization was heavily skewed toward high-risk individuals, despite some intermediate-risk patients sharing similar genetic or hormonal profiles, indicating that hospitalization history may serve as a proxy for disease burden or healthcare engagement that traditional clinical indicators do not fully capture. Another unexpected finding was the dominance of HER2-negative status across all risk categories, including high-risk patients. This could suggest either a cohort-specific biological trend or the mitigating effects of HER2-targeted therapies, such as trastuzumab, in reducing risk classification [ 44 ]. In contrast, estrogen receptor (ER) positivity was most prevalent in the low- and intermediate-risk groups, underscoring its known protective role and reinforcing its value in prognostication and treatment planning [ 45 ]. Furthermore, aggressive subtypes such as Inflammatory Breast Cancer (IBC), Metaplastic Breast Cancer (MBC), and Triple-Negative Breast Cancer (TNBC) were found almost exclusively in the high-risk category, providing visual confirmation of their association with poor prognosis and highlighting their critical role in metastasis risk modeling [ 46 ]. These findings collectively emphasize the complex interplay between biological, clinical, and social determinants in breast cancer progression and advocate for more nuanced, personalized approaches to risk stratification and patient management. Limitations One significant limitation of our study is the absence of independent validation for the risk categories derived from our machine learning models. Without external validation using a separate dataset, our risk stratification system's generalizability and clinical utility remain uncertain. Specifically, this limitation raises concerns about how well the risk categories would perform across different patient populations or in varied clinical settings. Moreover, the calibration of the risk model how well the predicted probabilities match actual outcomes has not been empirically tested, which could lead to potential overestimation or underestimation of risk. Consequently, the reliability of these risk groups for making critical clinical decisions, such as determining the intensity of surveillance or pre-emptive treatments, is not established. Addressing this limitation in future studies would require gathering independent validation data, possibly through multicentre studies, and refining risk thresholds based on real-world clinical outcomes to enhance the model's applicability and trustworthiness in clinical practice. This approach will improve the model’s credibility and ensure that it can effectively aid in patient management and treatment planning. Conclusions and Recommendation These insights are highly relevant in Ghana, where healthcare delivery is challenged by late presentation, diagnostic delays, and inconsistent pathology services. They suggest that a predictive model trained on local data and guided by explainable ML methods can capture both global oncological principles and local disease presentation patterns. The stratified visualizations derived from these models revealed both expected and novel patterns such as the high prevalence of genetic mutations in intermediate-risk individuals, the concentration of aggressive subtypes in the high-risk group, and the unexpectedly dominant HER2-negative status across all risk tiers. These findings not only validated the models' predictive capacity but also uncovered clinically relevant nuances that may be overlooked in conventional staging. Additionally, the influence of ethnicity and hospitalization on risk profiles underscores the importance of incorporating socio-demographic variables into algorithmic models to enhance fairness and real-world applicability. We therefore recommend for the development of personalized, equitable, and explainable risk assessment frameworks to improve early identification and intervention for patients at risk of breast cancer metastasis. Abbreviations Abbreviation Full Meaning BC Breast Cancer MBC Metastatic Breast Cancer ER Estrogen Receptor PR Progesterone Receptor HER2 Human Epidermal Growth Factor Receptor 2 ML Machine Learning LR Logistic Regression RF Random Forest SVM Support Vector Machine AUC Area Under the Curve SHAP SHapley Additive exPlanations IDC Invasive Ductal Carcinoma ILC Invasive Lobular Carcinoma DCIS Ductal Carcinoma In Situ LCIS Lobular Carcinoma In Situ TCB Tubular Carcinoma of Breast IBC Inflammatory Breast Cancer TNBC Triple-Negative Breast Cancer WHO World Health Organization AJCC American Joint Committee on Cancer BRCA Breast Cancer gene (BRCA1/BRCA2) TNM Tumor-Node-Metastasis (Staging System) ROC Receiver Operating Characteristic NB Naive Bayes Declarations Ethics approval and consent to participate This retrospective study was reviewed and approved by the Institutional Review Board (IRB), Korle-Bu Teaching Hospital, Accra, Ghana (IRB KTHI5/1083703). The need for written informed consent was waived by the IRB because the study used de-identified, routinely collected clinical data, in accordance with national regulations. Also, since the research did not involve direct human interaction, no written consent process was required. Clinical trial number: not applicable Consent for publication Not applicable. Availability of data and materials The datasets used and/or analysed during the current study are available from the corresponding author on reasonable request. Competing interests The authors declare that they have no competing interests. Funding There was no specific funding for this study. Authors' contributions AYA conceptualized the study and led the study design, methodology development, data curation, formal analysis, machine learning modeling, manuscript drafting, visualization, and overall project supervision. SBA was responsible for data collection, validation of clinical findings, and critical manuscript revision. EA contributed to visualization, manuscript review, and editing. VUG conducted statistical analysis, assisted with machine learning modeling, validated results, and contributed to manuscript revision. JA was involved in manuscript editing and revision. SO contributed to manuscript revision and editing. SBB conducted the literature review, assisted with clinical data analysis, supported result validation, and edited the manuscript. AMB contributed to data interpretation, supervised data collection activities, and participated in manuscript review and editing. All authors read and approved the final manuscript. Acknowledgements Not applicable. Authors' information Not applicable References World Health Organization, “Breast cancer.” Accessed: May 14, 2025. [Online]. Available: https://www.who.int/news-room/fact-sheets/detail/breast-cancer F. Bray, J. Ferlay, I. Soerjomataram, R. L. Siegel, L. A. Torre, and A. Jemal, “Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries,” CA Cancer J Clin , vol. 68, no. 6, pp. 394–424, Nov. 2018, doi: 10.3322/CAAC.21492,. E. Jedy-Agba, V. McCormack, C. Adebamowo, and I. dos-Santos-Silva, “Stage at diagnosis of breast cancer in sub-Saharan Africa: a systematic review and meta-analysis,” Lancet Glob Health , vol. 4, no. 12, p. e923, Dec. 2016, doi: 10.1016/S2214-109X(16)30259-5. S. Y. Opoku, M. Benwell, and J. Yarney, “Knowledge, attitudes, beliefs, behaviour and breast cancer screening practices in Ghana, West Africa.,” 2012. Accessed: May 14, 2025. [Online]. Available: http://197.255.68.203/handle/123456789/3828 A. C. Mensah et al. , “Survival Outcomes of Breast Cancer in Ghana: An Analysis of Clinicopathological Features,” Open Access Library Journal , vol. 3, no. 1, pp. 1–11, Jan. 2016, doi: 10.4236/OALIB.1102145. American Cancer Society, “Treatment of Stage IV (Metastatic) Breast Cancer .” Accessed: May 14, 2025. [Online]. Available: https://www.cancer.org/cancer/types/breast-cancer/treatment/treatment-of-breast-cancer-by-stage/treatment-of-stage-iv-advanced-breast-cancer.html R. L. Siegel, K. D. Miller, N. S. Wagle, and A. Jemal, “Cancer statistics, 2023,” CA Cancer J Clin , vol. 73, no. 1, pp. 17–48, Jan. 2023, doi: 10.3322/CAAC.21763,. T. Reinert, A. B. A. de Souza, G. P. Sartori, F. M. Obst, and C. H. Barrios, “Highlights of the 17th St Gallen International Breast Cancer Conference 2021: Customising local and systemic therapies,” Ecancermedicalscience , vol. 15, 2021, doi: 10.3332/ECANCER.2021.1236,. E. M. Der et al. , “Triple-negative breast cancer in Ghanaian women: The Korle Bu Teaching Hospital experience,” Breast Journal , vol. 21, no. 6, pp. 627–633, Nov. 2015, doi: 10.1111/TBJ.12527,. N. Harbeck et al. , “Breast cancer,” Nat Rev Dis Primers , vol. 5, no. 1, Dec. 2019, doi: 10.1038/S41572-019-0111-2,. F. Cardoso et al. , “Early breast cancer: ESMO Clinical Practice Guidelines for diagnosis, treatment and follow-up,” Annals of Oncology , vol. 30, no. 8, pp. 1194–1220, Aug. 2019, doi: 10.1093/annonc/mdz173. E. Jiagge et al. , “Comparative Analysis of Breast Cancer Phenotypes in African American, White American, and West Versus East African patients: Correlation Between African Ancestry and Triple-Negative Breast Cancer,” Ann Surg Oncol , vol. 23, no. 12, pp. 3843–3849, Nov. 2016, doi: 10.1245/S10434-016-5420-Z,. J. A. Cruz and D. S. Wishart, “Applications of machine learning in cancer prediction and prognosis,” Cancer Inform , vol. 2, pp. 59–77, 2006, doi: 10.1177/117693510600200030. A. Esteva et al. , “A guide to deep learning in healthcare,” Nat Med , vol. 25, no. 1, pp. 24–29, Jan. 2019, doi: 10.1038/S41591-018-0316-Z;SUBJMETA=114,1305,1647,48,631,692,700;KWRD=BIOINFORMATICS,HEALTH+CARE,MACHINE+LEARNING. K. Kourou, T. P. Exarchos, K. P. Exarchos, M. V. Karamouzis, and D. I. Fotiadis, “Machine learning applications in cancer prognosis and prediction,” Comput Struct Biotechnol J , vol. 13, pp. 8–17, 2015, doi: 10.1016/j.csbj.2014.11.005. F. Pedregosa et al. , “Scikit-learn: Machine Learning in Python,” Jan. 2012, Accessed: May 14, 2025. [Online]. Available: http://arxiv.org/abs/1201.0490 S. B. Kotsiantis, I. D. Zaharakis, and P. E. Pintelas, “Machine learning: A review of classification and combining techniques,” Artif Intell Rev , vol. 26, no. 3, pp. 159–190, Nov. 2006, doi: 10.1007/S10462-007-9052-3/METRICS. D. W. Hosmer, S. Lemeshow, and R. X. Sturdivant, “Applied Logistic Regression: Third Edition,” Applied Logistic Regression: Third Edition , pp. 1–510, Aug. 2013, doi: 10.1002/9781118548387. I. Guyon, J. Weston, S. Barnhill, and V. Vapnik, “Gene selection for cancer classification using support vector machines,” Mach Learn , vol. 46, no. 1–3, pp. 389–422, 2002, doi: 10.1023/A:1012487302797/METRICS. T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , vol. 13-17-August-2016, pp. 785–794, Aug. 2016, doi: 10.1145/2939672.2939785. R. Kohavi, “A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection,” International Joint Conference on Artificial Intelligence , 1995. S. Varma and R. Simon, “Bias in error estimation when using cross-validation for model selection,” BMC Bioinformatics , vol. 7, no. 1, pp. 1–8, Feb. 2006, doi: 10.1186/1471-2105-7-91/FIGURES/4. E. W. Steyerberg et al. , “Assessing the performance of prediction models: A framework for traditional and novel measures,” Epidemiology , vol. 21, no. 1, pp. 128–138, Jan. 2010, doi: 10.1097/EDE.0B013E3181C30FB2,. L. Dihge, M. Ohlsson, P. Edén, P. O. Bendahl, and L. Rydén, “Artificial neural network models to predict nodal status in clinically node-negative breast cancer,” BMC Cancer , vol. 19, no. 1, Jun. 2019, doi: 10.1186/S12885-019-5827-6,. R. Ha et al. , “Predicting Breast Cancer Molecular Subtype with MRI Dataset Utilizing Convolutional Neural Network Algorithm,” J Digit Imaging , vol. 32, no. 2, p. 276, Apr. 2019, doi: 10.1007/S10278-019-00179-2. S. M. Lundberg and S. I. Lee, “A Unified Approach to Interpreting Model Predictions,” Adv Neural Inf Process Syst , vol. 2017-December, pp. 4766–4775, May 2017, Accessed: May 14, 2025. [Online]. Available: https://arxiv.org/pdf/1705.07874 L. Breiman, “Random forests,” Mach Learn , vol. 45, no. 1, pp. 5–32, Oct. 2001, doi: 10.1023/A:1010933404324/METRICS. P. Wang et al. , “Machine Learning Models for Diagnosing Glaucoma from Retinal Nerve Fiber Layer Thickness Maps,” Ophthalmol Glaucoma , vol. 2, no. 6, pp. 422–428, Nov. 2019, doi: 10.1016/j.ogla.2019.08.004. D. Delen, G. Walker, and A. Kadam, “Predicting breast cancer survivability: A comparison of three data mining methods,” Artif Intell Med , vol. 34, no. 2, pp. 113–127, Jun. 2005, doi: 10.1016/J.ARTMED.2004.07.002,. American Joint Committee on Cancer, “Cancer Staging Systems.” Accessed: May 14, 2025. [Online]. Available: https://www.facs.org/quality-programs/cancer-programs/american-joint-committee-on-cancer/cancer-staging-systems/ J. Clegg-Lamptey, J. Dakubo, and Y. N. Attobra, “Why Do Breast Cancer Patients Report Late or Abscond During Treatment in Ghana? A Pilot Study,” Ghana Med J , vol. 43, no. 3, p. 127, Sep. 2009, Accessed: May 14, 2025. [Online]. Available: https://pmc.ncbi.nlm.nih.gov/articles/PMC2810246/ C. K. Anders, R. Johnson, J. Litton, M. Phillips, and A. Bleyer, “Breast Cancer Before Age 40 Years,” Semin Oncol , vol. 36, no. 3, pp. 237–249, Jun. 2009, doi: 10.1053/j.seminoncol.2009.03.001. E. A. Rakha et al. , “Invasive lobular carcinoma of the breast: Response to hormonal therapy and outcomes,” Eur J Cancer , vol. 44, no. 1, pp. 73–83, Jan. 2008, doi: 10.1016/j.ejca.2007.10.009. C. M. Perou et al. , “Molecular portraits of human breast tumours,” Nature , vol. 406, no. 6797, pp. 747–752, Aug. 2000, doi: 10.1038/35021093,. L. Denny et al. , “Human papillomavirus prevalence and type distribution in invasive cervical cancer in sub-Saharan Africa,” Int J Cancer , vol. 134, no. 6, pp. 1389–1398, Mar. 2014, doi: 10.1002/IJC.28425,. S. B. Edge and C. C. Compton, “The american joint committee on cancer: The 7th edition of the AJCC cancer staging manual and the future of TNM,” Ann Surg Oncol , vol. 17, no. 6, pp. 1471–1474, Jun. 2010, doi: 10.1245/S10434-010-0985-4/TABLES/1. M. Cianfrocca and L. J. Goldstein, “Prognostic and Predictive Factors in Early-Stage Breast Cancer,” Oncologist , vol. 9, no. 6, pp. 606–616, Nov. 2004, doi: 10.1634/THEONCOLOGIST.9-6-606’)). A. Goldhirsch et al. , “Personalizing the treatment of women with early breast cancer: Highlights of the st gallen international expert consensus on the primary therapy of early breast Cancer 2013,” Annals of Oncology , vol. 24, no. 9, pp. 2206–2223, Sep. 2013, doi: 10.1093/annonc/mdt303. N. Hamajima et al. , “Menarche, menopause, and breast cancer risk: Individual participant meta-analysis, including 118 964 women with breast cancer from 117 epidemiological studies,” Lancet Oncol , vol. 13, no. 11, pp. 1141–1151, Nov. 2012, doi: 10.1016/S1470-2045(12)70425-4. N. Petrucelli, M. B. Daly, and G. L. Feldman, “Hereditary breast and ovarian cancer due to mutations in BRCA1 and BRCA2,” Genetics in Medicine , vol. 12, no. 5, pp. 245–259, May 2010, doi: 10.1097/GIM.0b013e3181d38f2f. W. D. Foulkes, I. E. Smith, and J. S. Reis-Filho, “Triple-negative breast cancer,” N Engl J Med , vol. 363, no. 20, pp. 1938–1948, Nov. 2010, doi: 10.1056/NEJMRA1001389. L. A. Carey et al. , “Race, breast cancer subtypes, and survival in the Carolina Breast Cancer Study,” J Am Med Assoc , vol. 295, no. 21, pp. 2492–2502, Jun. 2006, doi: 10.1001/JAMA.295.21.2492,. C. E. DeSantis et al. , “Breast cancer statistics, 2019,” CA Cancer J Clin , vol. 69, no. 6, pp. 438–451, Nov. 2019, doi: 10.3322/CAAC.21583,. D. J. Slamon et al. , “Use of Chemotherapy plus a Monoclonal Antibody against HER2 for Metastatic Breast Cancer That Overexpresses HER2,” New England Journal of Medicine , vol. 344, no. 11, pp. 783–792, Mar. 2001, doi: 10.1056/NEJM200103153441101,. E. A. Rakha, M. E. El-Sayed, A. R. Green, A. H. S. Lee, J. F. Robertson, and I. O. Ellis, “Prognostic markers in triple-negative breast cancer,” Cancer , vol. 109, no. 1, pp. 25–32, Jan. 2007, doi: 10.1002/CNCR.22381,. S. Dawood et al. , “International expert panel on inflammatory breast cancer: Consensus statement for standardized diagnosis and treatment,” Annals of Oncology , vol. 22, no. 3, pp. 515–523, 2011, doi: 10.1093/ANNONC/MDQ345,. Additional Declarations No competing interests reported. Supplementary Files Supplementaryfile.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6671298","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":473249756,"identity":"581d71fe-c0f7-4b97-bc94-1ab6c519eeec","order_by":0,"name":"Abdulzeid Yen Anafo","email":"","orcid":"","institution":"Department of Mathematical Sciences, University of Mines and Technology, P.O. Box 237, Tarkwa, Ghana","correspondingAuthor":false,"prefix":"","firstName":"Abdulzeid","middleName":"Yen","lastName":"Anafo","suffix":""},{"id":473249757,"identity":"a8d59a9f-5262-4476-a695-d2e7eed9d08e","order_by":1,"name":"Senyefia Bosson-Amedenu","email":"","orcid":"","institution":"Department of Mathematics, Statistics and Actuarial Science, Takoradi Technical University, Box 256 Takoradi, Ghana","correspondingAuthor":false,"prefix":"","firstName":"Senyefia","middleName":"","lastName":"Bosson-Amedenu","suffix":""},{"id":473249758,"identity":"033af6df-6838-4db7-bc2c-6cbc2a38051b","order_by":2,"name":"Emmanuel Ayitey","email":"","orcid":"","institution":"Department of Mathematics, Statistics and Actuarial Science, Takoradi Technical University, Box 256 Takoradi, Ghana","correspondingAuthor":false,"prefix":"","firstName":"Emmanuel","middleName":"","lastName":"Ayitey","suffix":""},{"id":473249759,"identity":"b18441e8-5c88-4064-951d-ad2dc6a9c017","order_by":3,"name":"Vincent Uwumboriyhie Gmayinaam","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABA0lEQVRIiWNgGAWjYPACCRDB+ICxAcwzAOIDRGlhNiBFCxiwSRClxeBG+tUNP9ss8vjbT6dV8+6ws2dgb94mwfDrDh4tOWU3e9skiiXO5G67zXsmObGB51iZBGPfM5xaJGfkpN3gbZNIbDgA0tLGnMAgkWMmwdhzGK+Wm3+BWuaff7utmLet3p5B/g1+LfwS6cdug2zZcCN3GzNv22HGBgkeMwmGH3i08Lxhuy1zTiJx4423myXnth1PbONJK7ZIbMCthY09/dnNN2V1ifPO52788Lat2p6f/fDGGx/+4NbCwMBjwMDIhmwIiEhsw6ODgf0BA8MfDFFMkVEwCkbBKBi5AAB95FyRu+EkbAAAAABJRU5ErkJggg==","orcid":"","institution":"Institute of Health Research, University of Health and Allied Sciences, Ho, Ghana","correspondingAuthor":true,"prefix":"","firstName":"Vincent","middleName":"Uwumboriyhie","lastName":"Gmayinaam","suffix":""},{"id":473249760,"identity":"52d5c66e-6d18-4728-b916-12bff12eb4b9","order_by":4,"name":"Joseph Acquah","email":"","orcid":"","institution":"Department of Mathematical Sciences, University of Mines and Technology, P.O. Box 237, Tarkwa, Ghana","correspondingAuthor":false,"prefix":"","firstName":"Joseph","middleName":"","lastName":"Acquah","suffix":""},{"id":473249761,"identity":"2b8cc79b-b55c-4f6e-9d76-319d4bf3ed93","order_by":5,"name":"Selasi Ocloo","email":"","orcid":"","institution":"Department of Engineering, Ashesi University, Brekuso, Ghana","correspondingAuthor":false,"prefix":"","firstName":"Selasi","middleName":"","lastName":"Ocloo","suffix":""},{"id":473249762,"identity":"4d24c809-1377-4eaa-a8e6-117f46026a94","order_by":6,"name":"Stephanies Brako Boateng","email":"","orcid":"","institution":"Medical Physics Unit, Sweden Ghana Medical Centre, East Legon, Ghana","correspondingAuthor":false,"prefix":"","firstName":"Stephanies","middleName":"Brako","lastName":"Boateng","suffix":""},{"id":473249763,"identity":"cadc8b75-6bb8-458e-bfa6-273d5a5de4d1","order_by":7,"name":"Alhassan Mohammed Baidoo","email":"","orcid":"","institution":"Medical Physics Unit, Sweden Ghana Medical Centre, East Legon, Ghana","correspondingAuthor":false,"prefix":"","firstName":"Alhassan","middleName":"Mohammed","lastName":"Baidoo","suffix":""}],"badges":[],"createdAt":"2025-05-15 09:53:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6671298/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6671298/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":85070507,"identity":"c1af140d-f865-44cd-ab81-2bc629cc054f","added_by":"auto","created_at":"2025-06-20 15:37:17","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":231691,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eConfusion matrices for five machine learning models; Logistic Regression, Random Forest, SVM, Naive Bayes, and XGBoost used in predicting breast cancer metastasis in a Ghanaian cohort.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6671298/v1/ff8bf1a0c50f6709359c957d.jpeg"},{"id":85069291,"identity":"65a5d03d-dda6-4500-bed1-4c665a7a04b1","added_by":"auto","created_at":"2025-06-20 15:29:17","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":248738,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eSHAP-based feature importance analysis of the Logistic model used for predicting breast cancer metastasis in a Ghanaian cohort. The left beeswarm plot illustrates the direction and magnitude of individual feature impacts on model predictions, with red indicating high feature values and blue indicating low values. The right bar chart presents the mean absolute SHAP values, reflecting each feature’s average contribution to the model's output. Lymph node involvement, tumor size, and cancer stage were the most influential predictors, while age, tumor grade, and HER2 status showed minimal effect.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6671298/v1/8549f4789911bb0ca5553b1c.jpeg"},{"id":85070508,"identity":"5f8ade13-b2dd-4eaa-95d8-fa680c30f9b9","added_by":"auto","created_at":"2025-06-20 15:37:17","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":313506,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eHeatmaps showing the distribution of clinical, demographic, molecular, and pathological features across Logistic Regression-predicted risk groups for breast cancer metastasis (Low Risk, Moderate Risk, and High Risk). Figuure 3A, illustrates variations in general clinical and demographic characteristics. High-risk groups showed elevated proportions of adverse indicators such as lymph node involvement (53%), advanced-stage disease (52%), and larger tumor sizes (44%), while low-risk groups were enriched with early-stage disease and negative nodal status. Figure 3B, displays selected molecular and histopathological features. High-risk patients had greater representation of HER2-positive (29%), triple-negative (25%), and Grade 3 tumors (29%), as well as higher recurrence rates (16%). In contrast, low-risk patients were predominantly ER/PR-positive (84%/82%) and of Luminal-A subtype (73%), highlighting the model’s ability to stratify patients in alignment with clinically validated metastatic risk patterns.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6671298/v1/711afb24b618d0390e74a337.jpeg"},{"id":85070509,"identity":"329fa165-fc83-4c99-ae20-7492a542ad41","added_by":"auto","created_at":"2025-06-20 15:37:17","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":847936,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eComprehensive visualization of clinical, hormonal, molecular, and demographic features across Logistic Regression-predicted risk groups for breast cancer metastasis (High Risk, Intermediate Risk, and Low Risk). The top row includes box plots showing differences in age, cancer stage, and lymph node involvement, along with bar plots for tumor size, PR status, and menopausal status. Notably, high-risk patients were more likely to be older, have advanced-stage tumors, lymph node involvement, and post-menopausal status. The bottom row presents labeled categorical distributions, illustrating that high-risk patients had higher proportions of HER2 positivity, ER/PR negativity, triple-negative and HER2-enriched subtypes, and genetic mutations. Low-risk patients, in contrast, were predominantly ER/PR-positive, Luminal-A subtype, and had lower hospitalization and recurrence rates. These patterns highlight the model’s capacity to stratify patients according to biologically and clinically validated metastatic risk factors.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6671298/v1/4842adbe573b420610f18423.jpeg"},{"id":88973813,"identity":"a0cce2bb-2deb-47ae-85c3-38e9d352b3f6","added_by":"auto","created_at":"2025-08-13 10:02:12","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2900038,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6671298/v1/23bdc9c8-f34a-4810-8ee0-f1bb819451e3.pdf"},{"id":85069302,"identity":"cb017c2f-0894-4ee0-8f1d-5ea0bf6adf04","added_by":"auto","created_at":"2025-06-20 15:29:17","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":1505095,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementaryfile.docx","url":"https://assets-eu.researchsquare.com/files/rs-6671298/v1/0d7de98643b8fcf9a5db5114.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Risk Stratification of Breast Cancer Metastasis: A Predictive Modelling Framework Using Clinical and Hormonal Receptor Data in a Ghanaian Cohort","fulltext":[{"header":"Introduction","content":"\u003cp\u003eBreast cancer (BC) stands as the most diagnosed cancer and the top cause of cancer-related deaths among women worldwide, exhibiting significant disparities in outcomes across different regions [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. High-income countries have made substantial progress in early detection and improving survival rates, achieving five-year survival rates that exceed 80% due to widespread screening and timely treatment interventions [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. In contrast, sub-Saharan Africa experiences alarmingly high breast cancer mortality rates, primarily attributed to late-stage diagnosis resulting from limited access to screening programs, and significant delays in treatment [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. In Ghana specifically, BC constitutes nearly one-third of all cancers affecting women, and a large proportion of cases (approximately 70%) are diagnosed at advanced or metastatic stages, leading to poor prognosis and high mortality rates [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eBC Metastasis also known as stage IV breast cancer, occurs when cancer cells spread beyond the breast and regional lymph nodes to distant organs such as the bones, liver, lungs, and brain [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. The prognosis for patients with MBC remains challenging, with survival rates significantly lower than for early-stage disease. For instance, the five-year relative survival rate for metastatic breast cancer in high-income countries is approximately 29%, highlighting the severity of this stage [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. In low-resource settings like Ghana, survival outcomes are further compromised due to barriers such as inadequate healthcare infrastructure, limited access to advanced treatments, high costs of medications, and delays in seeking medical care [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Additionally, socio-cultural factors and misconceptions surrounding breast cancer often discourage women from early healthcare engagement, exacerbating the severity of disease presentation at diagnosis and contributing to poor clinical outcomes [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Consequently, addressing the metastatic burden of breast cancer in regions such as Ghana requires multifaceted approaches, including enhanced public awareness, improved screening programs, timely referral systems, and equitable access to effective therapeutic interventions [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAccurate metastasis prediction is vital for optimizing treatment choices, resource distribution, and patient outcomes [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Conventional prognostic approaches heavily depend on clinical and pathological staging, hormone receptor status, and patient history to inform management strategies [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Nonetheless, these traditional methods might not fully reflect the complexity and diversity of breast cancer progression, particularly in resource-constrained environments where genomic profiling and advanced diagnostic tools are often unavailable [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Machine learning (ML) techniques offer a promising alternative, using data-driven methods to reveal subtle and intricate patterns in clinical data that typical analysis might miss [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. These models can considerably enhance decision-making by effectively pinpointing high-risk individuals, allowing for personalized and timely management strategies adapted to individual risk profiles [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWhile there is increasing interest in machine learning (ML) tools for oncology, a notable shortage exists in predictive models tailored to African populations with breast cancer. Most current models derive from Western datasets, potentially misrepresenting the genetic, environmental, and healthcare characteristics unique to African communities. This study seeks to address this deficiency by creating and assessing machine learning models focused on predicting breast cancer metastasis, utilising clinical, molecular, and demographic data from a cohort in Ghana.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eThis research was a single-institution-based quantitative retrospective cohort study conducted among patients with breast cancer. We employed a predictive modelling framework featuring five supervised machine learning classifiers: Logistic Regression, Random Forest, Support Vector Machine, XGBoost, and Naive Bayes, specifically analysing clinical and hormonal receptor data from Ghanaian patients. Furthermore, the research intends to develop a practical, data-oriented tool for assessing metastasis risk, enhancing precision medicine efforts in low- and middle-income nations.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Collection\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study utilized secondary data from 558 breast cancer patients, aged 13 to 97 years, diagnosed at Korle-Bu Teaching Hospital in Ghana. The dataset included demographic information (such as age and menopausal status), clinical features (including tumor size, cancer stage, and lymph node involvement), and molecular characteristics (estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER) status). Data also captured tumor subtype classifications, recurrence events, survival time intervals, and mortality outcomes due to breast cancer or competing causes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Preprocessing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBefore model development, the dataset underwent several preprocessing steps to ensure compatibility with machine learning algorithms and optimize model performance. Continuous variables, such as tumor size and age at diagnosis, were standardized using the StandardScaler from scikit-learn [16], ensuring that all features contributed proportionally to the analysis without being influenced by differences in measurement scales. Categorical variables, including hormone receptor status (ER, PR, HER2), menopausal status, cancer stage, and tumor subtype, were encoded using one-hot encoding to transform them into a numerical format suitable for model training [17]. These preprocessing steps ensured that both clinical and molecular features were appropriately prepared for the subsequent machine learning analyses, allowing models to capture patterns effectively and improve prediction accuracy [15]\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel Development and Selection\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study assessed various models to analyse their effectiveness in predicting metastasis among breast cancer patients. The models examined include Logistic Regression (LR), Linear Support Vector Machine (SVM), Random Forest (RF), XGBoost, and Na\u0026iuml;ve Bayes classifiers, selected for their complementary abilities in conducting binary classification tasks typical in oncology research. Logistic Regression is favoured in cancer research because of its simplicity and clarity in calculating odds ratios and pinpointing crucial predictors [18]. Linear-kernel support Vector Machines are particularly adept at handling high-dimensional clinical data, providing strong outcomes in differentiating between metastatic and non-metastatic instances [19]. Meanwhile, ensemble models like Random Forest and XGBoost excel in predictive accuracy for cancer prognosis, thanks to their capacity to model intricate interactions and mitigate overfitting [20]. Although Na\u0026iuml;ve Bayes relies on the assumption of feature independence, it remains a useful tool in medical settings for its efficiency and surprisingly competitive accuracy in specific cancer prediction scenarios [15].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel Evaluation and Cross-Validation Methodology\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe developed a thorough evaluation framework to establish the reliability and adaptability of our predictive models. We used a 10-fold stratified cross-validation method to ensure that the distribution of metastatic and non-metastatic cases remained consistent across all folds. This stratification is essential in medical datasets with imbalanced classes, as it guarantees fair representation of the minority group and mitigates performance bias [21]. To further bolster stability, we repeated the cross-validation process three times with different random seeds, a technique that reduces variance in model performance estimates and yields more dependable results [22].\u003c/p\u003e\n\u003cp\u003eIn each fold, hyperparameters were fine-tuned to enhance model performance. The C and gamma, the penalty parameter and kernel coefficient, respectively, were optimized for the SVM, while n_estimators and max_depth were adjusted for the Random Forest and XGBoost classifiers [20]. We computed evaluation metrics such as accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (AUC) for each fold, averaging them to comprehensively assess model effectiveness.\u003c/p\u003e\n\u003cp\u003eWe performed a final assessment using a distinct test set following the training phase. We created confusion matrices for each model to provide a comprehensive analysis of classification performance, detailing rates of true positives, true negatives, false positives, and false negatives. Additionally, we plotted ROC curves to demonstrate each model\u0026rsquo;s ability to distinguish between metastatic and non-metastatic cases at different thresholds. We particularly focused on precision, recall, and F1-score due to their critical importance in clinical decision-making, where misclassifications especially false negatives can significantly impact patient outcomes. This methodology aligns with recent studies emphasising the importance of thorough evaluation metrics in medical machine learning [15].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRisk Stratification\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAfter evaluating the models, predictive algorithms categorised BC patients into significant risk categories based on their metastasis likelihood. This stratification process started by generating individual risk scores from each model\u0026apos;s probabilistic outputs. These scores reflected each patient\u0026apos;s predicted probability of metastatic progression, acting as an indicator of metastasis risk. To enhance understandability and aid clinical decisions, these continuous risk scores were grouped into three clear categories: low, intermediate, and high.\u003c/p\u003e\n\u003cp\u003eThe risk stratification method, utilising predicted probabilities, is commonly employed in cancer research to improve clinical decision-making and patient counseling. Risk scores are typically grouped into discrete categories- low, intermediate, and high risk- making them practical for clinical use [23]. This approach has been applied in breast cancer research to pinpoint patients at increased risk of metastasis or recurrence, enabling more tailored surveillance and treatment strategies [24]. Stratification thresholds are usually based on clinical relevance or the distribution characteristics of predicted probabilities, as illustrated by various machine learning models used for breast cancer prognosis [25].\u003c/p\u003e\n\u003cp\u003ePatients who scored below 0.33 were placed in the low-risk category, suggesting a lower chance of developing metastasis. Those with scores between 0.33 and 0.66 were labelled as intermediate risk, indicating a moderate likelihood of progression. Scores above 0.66 were categorised as high risk, signifying patients with a considerable probability of metastasis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSHAP Analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo improve interpretability and quantify how individual features affect model predictions, we utilised SHAP (SHapley Additive exPlanations) analysis, a cohesive framework for understanding predictions from machine learning models [26]. SHAP values were calculated for the top-performing model among LR, RF, XGBoost, SVM, and Na\u0026iuml;ve Bayes using the TreeExplainer and LinearExplainer methods provided in the SHAP Python library based on model compatibility. This technique offers both global and local interpretability by attributing a Shapley value to each feature, indicating its average impact on the predicted metastasis risk for each patient. The SHAP beeswarm plot was employed to visualise impact, distribution and direction of each feature across the dataset. In contrast, the SHAP summary bar plot depicted the mean absolute importance of every feature. This analysis was crucial for validating the clinical significance of the model by pinpointing which variables had the strongest influence on metastasis predictions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatistical Analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePython (version 3.8), using libraries such as scikit-learn, NumPy, Pandas, Matplotlib, and Seaborn, was used for all data manipulations, analyses, and visualizations. This comprehensive suite of tools facilitated rigorous statistical analysis and ensured reproducibility of the results\u003cem\u003e.\u0026nbsp;\u003c/em\u003e\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eBaseline Characteristics of Participants\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTable 1 \u0026nbsp;below presents the baseline characteristics of participants. Among the 558 patients included in the study, 281 (50.4%) experienced metastases while 277 (49.6%) did not. Patients with metastasis were significantly older than those without spread, with a mean age of 52.3 \u0026plusmn; 15.4 years compared to 47.6 \u0026plusmn; 12.6 years (p \u0026lt; 0.001). A greater proportion of patients aged above 50 years was observed in the metastasis group (50.5%) compared to the no spread group (40.1%, p = 0.013). Recurrent disease was significantly more frequent among patients with metastasis (13.5% vs 3.6%, p \u0026lt; 0.001). Competing risk analysis revealed that 85.8% of metastasis patients experienced an event compared to none in the no spread group (p \u0026lt; 0.001). Hormonal receptor status differed markedly between groups: estrogen receptor (ER) positivity and progesterone receptor (PR) positivity were both significantly lower among metastasis patients (ER+: 43.1% vs 66.4%; PR+: 41.6% vs 64.3%; both p \u0026lt; 0.001). Human epidermal growth factor receptor 2 (HER2) status, however, did not differ significantly between groups (p = 0.230). With respect to molecular subtypes, triple-negative breast cancer was notably more frequent in the metastasis group (52.7% vs 33.2%), while Luminal-A subtype predominated among patients without metastasis (45.5% vs 18.9%; p \u0026lt; 0.001). Higher tumor grades (Grade 3) were significantly associated with metastasis (57.3% vs 35.0%, p \u0026lt; 0.001). Advanced tumor stages were also more prevalent in metastasis cases, with 69.4% at Stage 3 and 20.6% at Stage 4 compared to only 5.8% and 0.7% respectively among patients without spread (p \u0026lt; 0.001). Tumor size distribution showed that larger tumors (51-60 mm and above 60 mm) were predominantly found among metastasis cases (69.4% and 20.6%, respectively; p \u0026lt; 0.001). Lymph node involvement was significantly greater in the metastasis group, with 46.6% having two nodes involved and 33.1% having three nodes involved (p \u0026lt; 0.001). Menopausal status did not significantly differ between groups (p = 0.400). Additionally, hospitalization during the disease course was significantly more common among metastasis patients (33.5% vs none, p \u0026lt; 0.001), and the presence of Breast CAncer gene\u003cstrong\u003e\u0026nbsp;1 (\u003c/strong\u003eBRCA1) genetic traces was more frequent among metastasis cases (52.7% vs 33.2%, p \u0026lt; 0.001).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003cstrong\u003e: Baseline Characteristics of Participants\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eVariable\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal (N=558)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eNo Metastasis(n=277)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eMetastasis (n=281)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003ep-value\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAge, mean \u0026plusmn; SD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e50.0 \u0026plusmn; 14.3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e47.6 \u0026plusmn; 12.6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e52.3 \u0026plusmn; 15.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e1\u003c/sup\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAge, median (IQR)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e49.0 (40.0, 58.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e47.0 (38.0, 56.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e51.0 (41.0, 62.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003en (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003en (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003en (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAge categories\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e0.013 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;13-50\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e305 (54.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e166 (59.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e139 (49.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Above 50\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e253 (45.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e111 (40.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e142 (50.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eEthnicity\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Other\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e111 (19.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e66 (23.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e45 (16.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Akan\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e217 (38.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e88 (31.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e129 (45.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Ga/Adangbe\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e131 (23.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e63 (22.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e68 (24.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Ewe\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e99 (17.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e60 (21.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e39 (13.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eRecurrent\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;No\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e510 (91.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e267 (96.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e243 (86.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Yes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e48 (8.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e10 (3.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e38 (13.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCompeting Risk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Censored\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e277 (49.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e277 (100.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Event\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e241 (43.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e241 (85.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Death due to other diseases\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e40 (7.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e40 (14.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eHuman ER2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e0.230 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Negative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e429 (76.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e219 (79.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e210 (74.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Positive\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e129 (23.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e58 (20.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e71 (25.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eER\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Negative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e253 (45.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e93 (33.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e160 (56.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Positive\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e305 (54.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e184 (66.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e121 (43.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Negative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e263 (47.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e99 (35.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e164 (58.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Positive\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e295 (52.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e178 (64.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e117 (41.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMSubtype\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Luminal-A\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e179 (32.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e126 (45.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e53 (18.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Triple negative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e240 (43.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e92 (33.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e148 (52.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Luminal-B\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e121 (21.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e54 (19.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e67 (23.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Her2 Negative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e18 (3.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e5 (1.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e13 (4.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eGrade\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Grade1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e179 (32.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e126 (45.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e53 (18.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Grade 2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e121 (21.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e54 (19.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e67 (23.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Grade 3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e258 (46.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e97 (35.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e161 (57.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eType of Breast Cancer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;IDC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e424 (76.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e179 (64.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e245 (87.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;LCIS\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e8 (1.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e5 (1.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e3 (1.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;IBC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e40 (7.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e30 (10.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e10 (3.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;MBC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e27 (4.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e18 (6.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e9 (3.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;TCB\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e2 (0.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e2 (0.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;ILC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e32 (5.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e21 (7.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e11 (3.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;DCIS\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e25 (4.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e22 (7.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e3 (1.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStage, mean \u0026plusmn; SD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e2.3 \u0026plusmn; 1.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e1.5 \u0026plusmn; 0.8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e3.1 \u0026plusmn; 0.6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStage\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e31 (5.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e28 (10.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e3 (1.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e101 (18.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e99 (35.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e2 (0.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e155 (27.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e132 (47.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e23 (8.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e211 (37.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e16 (5.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e195 (69.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e60 (10.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e2 (0.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e58 (20.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStatus\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Survived\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e275 (49.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e272 (98.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e3 (1.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Died\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e283 (50.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e5 (1.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e278 (98.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eLymph Node\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e84 (15.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e84 (30.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e214 (38.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e157 (56.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e57 (20.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e158 (28.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e27 (9.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e131 (46.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e102 (18.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e9 (3.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e93 (33.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMenopause\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e0.400 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;No\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e194 (34.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e101 (36.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e93 (33.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Yes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e364 (65.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e176 (63.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e188 (66.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eHospitalized\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;No\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e464 (83.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e277 (100.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e187 (66.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Yes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e94 (16.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e94 (33.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eTumor Size\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;5-10 mm\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e31 (5.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e28 (10.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e3 (1.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;11-20 mm\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e101 (18.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e99 (35.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e2 (0.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;201-50 mm\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e155 (27.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e132 (47.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e23 (8.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;51-60 mm\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e211 (37.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e16 (5.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e195 (69.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Above 60 mm\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e60 (10.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e2 (0.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e58 (20.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eGenetics\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e\u0026lt;0.001 \u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Braca 1 traces\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e240 (43.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e92 (33.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e148 (52.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;No traces\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e318 (57.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e185 (66.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\n \u003cp\u003e133 (47.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003csup\u003e1\u0026nbsp;\u003c/sup\u003et-test p-value; \u003csup\u003e2\u003c/sup\u003e Wilcoxon Rank-sum test p-value; \u003csup\u003e3\u003c/sup\u003e Chi-square test p-value; \u003csup\u003e4\u003c/sup\u003e Fisher\u0026rsquo;s exact test p-value\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel Evaluation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Table 2 below presents the performance evaluation of five machine learning models (Logistic Regression, Random Forest, Support Vector Machine (SVM), XGBoost, and Naive Bayes). Logistic Regression and SVM achieved the highest accuracy rates, at 0.93 and 0.91, respectively, with SVM also recording the highest area under the curve (AUC = 0.98), indicating exceptional discriminatory ability. Random Forest exhibited strong precision (0.95) and specificity (0.94), which enhanced its effectiveness in accurately identifying non-metastatic conditions. Although XGBoost and Naive Bayes displayed notable recall values (0.96 and 0.95, respectively), they had lower specificity scores, indicating a propensity to over-predict metastatic cases. Nonetheless, their high recall demonstrates their usefulness in clinical scenarios where reducing false negatives is essential. In conclusion, Logistic Regression proved to be the most balanced model, excelling across all metrics, including a cross-validation accuracy of 0.96.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003e2\u003c/strong\u003e\u003cstrong\u003e:Performance evaluation of five machine learning models; LR, RF, SVM, XGBoost, and Naive Bayes for predicting breast cancer metastasis in a Ghanaian cohort. The models were assessed using accuracy, specificity, precision, recall, F1-score, AUC, and 10-fo\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModels\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 79px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAccuracy\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 86px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSpecificity\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 78px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePrecision\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRecall\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eF-Measure\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 71px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAUC\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 83px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e10-fold cross validation accuracy\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003eLR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 79px;\"\u003e\n \u003cp\u003e0.93\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 86px;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 78px;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e0.93\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 71px;\"\u003e\n \u003cp\u003e0.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 83px;\"\u003e\n \u003cp\u003e0.96\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003eRF\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 79px;\"\u003e\n \u003cp\u003e0.90\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 86px;\"\u003e\n \u003cp\u003e0.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 78px;\"\u003e\n \u003cp\u003e0.95\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e0.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 71px;\"\u003e\n \u003cp\u003e0.97\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 83px;\"\u003e\n \u003cp\u003e0.94\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003eSVM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 79px;\"\u003e\n \u003cp\u003e0.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 86px;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 78px;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 71px;\"\u003e\n \u003cp\u003e0.98\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 83px;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 79px;\"\u003e\n \u003cp\u003e0.88\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 86px;\"\u003e\n \u003cp\u003e0.80\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 78px;\"\u003e\n \u003cp\u003e0.83\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.96\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e0.89\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 71px;\"\u003e\n \u003cp\u003e0.95\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 83px;\"\u003e\n \u003cp\u003e0.87\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003eNaives Bayes\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 79px;\"\u003e\n \u003cp\u003e0.87\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 86px;\"\u003e\n \u003cp\u003e0.79\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 78px;\"\u003e\n \u003cp\u003e0.82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.95\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e0.88\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 71px;\"\u003e\n \u003cp\u003e0.97\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 83px;\"\u003e\n \u003cp\u003e0.91\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eThe confusion matrices presented in Figure 1 below illustrate the classification performance of five machine learning models for predicting metastasis in a Ghanaian breast cancer cohort. Logistic Regression and SVM models exhibited the most balanced and accurate predictions, each achieving 34 true positives and 49 true negatives, with minimal false classifications, highlighting their reliability in distinguishing metastatic from non-metastatic cases. The Random Forest model also performed well, achieving similar true positive rates but with a slight increase in false positives. On the other hand, Naive Bayes displayed the weakest performance, with a notable rise in false negatives (7), which is a significant concern in medical diagnostics, where missing a metastasis case could delay treatment. XGBoost, while maintaining a strong true positive count, suffered from the highest false positive rate, suggesting it might overpredict metastasis.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFeature importance\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA SHAP analysis was conducted to determine and quantify the impact of individual logistic model features on predicting breast cancer metastasis within the Ghanaian cohort. Figure 2 displays the SHAP beeswarm plot alongside the corresponding bar plot of mean SHAP values. The Logistic Regression model reveals that Lymph Node Involvement (LymphNode) is the most influential feature, with a mean SHAP value of approximately 1.38, indicating a strong impact on metastasis prediction. This is followed closely by Tumor Size (TumorSize) at 1.26 and Cancer Stage (Stage) at 1.15, both of which are clinically established predictors of metastasis. Age at Cancer Treatment (AgeCAT) contributes moderately with a SHAP value of 0.61, along with Time Since Diagnosis (FlagTime) at 0.51 and Type of Breast Cancer (TypeofBC) at 0.46. Genetic Predisposition (Genetics) also plays a notable role (0.46), as do Menopausal Status (Menopause) at 0.29 and Estrogen Receptor Status (ER) at 0.27. Other features such as Molecular Subtype (MSubType), Ethnicity, Progesterone Receptor Status (PR), and Recurrent Tumour (Recurrent) have smaller contributions (SHAP values between 0.1 and 0.25). On the lowest end of influence are Patient Age (Age), Tumor Grade (Grade), and Human Epidermal Growth Factor Receptor 2 (HER2), each with mean SHAP values below 0.1, suggesting minimal predictive weight in this model. \u0026nbsp;The SHAP beeswarm plot for the logistic regression model illustrates the global impact of each feature on metastasis prediction. Lymph node involvement emerged as the most influential predictor, where higher values (shown in red) strongly increased the predicted risk of metastasis, while lower values (blue) contributed negatively. Tumor size and cancer stage followed closely, with larger tumors and advanced stages (also represented in red) consistently shifting predictions toward AgeCAT and genetic factors showed moderate influence, indicating that demographic groupings and hereditary risk may modestly affect model output. TypeofBC also had a noticeable impact, especially certain subtypes associated with greater metastatic potential. In contrast, features such as follow-up time, ER status, PR status, menopausal status, and ethnicity exhibited limited influence, with SHAP values close to zero and mixed red-blue distributions, suggesting variability without strong directional impact. The least impactful features were age, tumor grade, and HER2 status, indicating minimal contribution to the logistic regression model\u0026rsquo;s predictions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRisk stratification\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe risk stratification outcomes from the Logistic Regression model highlight significant, biologically relevant patterns across clinical, hormonal, molecular, and demographic variables. Figures 3 and 4 depict variations in general clinical and demographic traits, presenting selected molecular and histopathological characteristics. Additional plots on risk stratification for other ML models are displayed in the supplementary file as Figure 5S-8S.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFigure 3A shows that the Akan ethnic group exhibits an increasing prevalence from Low Risk (31%) to High Risk (46%), while the Ga-Adangbe and Ewe groups maintain relative consistency across categories. Other ethnic groups show a higher representation in the Low (24%) and Intermediate Risk (12%) categories than in the High-Risk category (17%). Regarding genetics, family history is notably more common among Low (69%) and Intermediate Risk (65%) individuals compared to those in the High-Risk group (55%). The occurrence of BRCA mutations rises steadily with increasing risk: 31% in Low Risk, 35% in Intermediate, and 55% in High Risk. Tumor grades demonstrate a clear trend, with Grade 1 being the most prevalent in Low Risk (47%) and decreasing in higher risk groups, while Grade 3 rises with risk levels: 33% in Low, 42% in Intermediate, and 59% in High Risk. Hospitalisation trends also mirror risk stratification: 100% of Low-Risk individuals were not hospitalised, compared to 96% in Intermediate and only 66% in the High-Risk group, with hospitalisation increasing to 34% in High-Risk individuals. Menopausal status reveals that post-menopausal patients are more frequent in the Intermediate (77%) and High Risk (66%) groups, while pre-menopausal individuals are more common in the High-Risk group (34%) compared to Low (36%) and Intermediate (23%) groups. Tumor staging indicates that Stage 3 is prevalent in High Risk (71%) and Intermediate (46%) categories, while early stages (Stage 0 and 1) are nearly absent in the High-Risk group. Stage 4 is primarily found in the High-Risk group (22%). \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFigure 3B reinforces the existence of biologically significant patterns across clinical, hormonal, molecular, and demographic features. As illustrated in Figures 3 and 4, high-risk patients significantly exhibit numerous aggressive disease markers. Within the high-risk group, 53% of patients experienced lymph node involvement, contrasting with 33% in the moderate-risk group and just 12% in the low-risk group. Likewise, tumors exceeding 5 cm were more common in high-risk patients (44%) compared to moderate-risk (27%) and low-risk patients (13%). The proportion of patients presenting with stage 3 or higher disease was 52% in the high-risk group, 35% in the moderate-risk group, and 17% in the low-risk category.\u003c/p\u003e\n\u003cp\u003eIn terms of recurrence and menopausal status, 16% of high-risk patients experienced recurrence, compared to 7% in the moderate-risk and 5% in the low-risk groups. Menopausal women represented 63% of the high-risk category, versus 50% in the moderate-risk group and 44% in the low-risk group. Notably, genetic predisposition was observed in 12% of high-risk patients, compared to only 4% and 6% in moderate and low-risk groups, respectively. The use of hormonal contraceptives showed a downward trend with increasing risk, being most prevalent in the low-risk group (64%), decreasing to 49% in the moderate-risk group, and just 42% in the high-risk category. Regarding clinical and molecular traits, 84% of low-risk patients were ER-positive, 82% were PR-positive, and 91% were HER2-negative, exhibiting a hormonal receptor-rich profile associated with better outcomes. Additionally, 73% of low-risk patients were categorized as Luminal A subtype, while only 1% were classified as triple-negative breast cancer (TNBC). In contrast, the high-risk group displayed more aggressive phenotypes: only 44% were ER-positive and 41% were PR-positive, while 29% were HER2-positive, and 25% had TNBC, marking a substantial increase compared to the low-risk group. The Luminal B subtype was distributed more evenly, with 13% in the low-risk, 22% in moderate-risk, and 21% in high-risk patients. Significantly, stage III and stage IV diseases were much more prevalent in the high-risk group 29% and 12%, respectively compared to just 6% and 1% in the low-risk category. Similarly, Grade 3 tumors occurred in 29% of high-risk patients, in contrast to only 6% of low-risk patients. \u0026nbsp;\u003c/p\u003e"},{"header":"Discussions","content":"\u003cp\u003eBreast cancer in sub-Saharan Africa carries a disproportionately high mortality rate, yet most clinical decision-support tools for predicting metastasis have been developed and validated on Western cohorts. As a result, existing models may not account for the unique genetic, environmental, and health-system factors that influence disease progression in West African populations. Our study therefore fills a critical gap by presenting a comprehensive comparative evaluation of five machine-learning models Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), XGBoost, and Naive Bayes (NB trained and tested on clinical, molecular, and demographic data from a Ghanaian breast cancer cohort. By tailoring predictive algorithms to a local dataset, we aim to improve the accuracy of metastasis risk stratification in Ghana and ultimately support better patient management and resource allocation.\u003c/p\u003e \u003cp\u003eAmong all the models, Logistic Regression exhibited the most balanced and consistent performance across metrics, achieving high accuracy, recall, specificity, precision, and the highest 10-fold cross-validation score. These findings suggest that Logistic Regression, despite being a relatively simple and interpretable model, is highly effective in metastatic risk prediction when well-calibrated and trained on context-specific data [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. The model's balance between sensitivity and specificity is particularly crucial in resource-constrained settings like Ghana, where diagnostic accuracy must be matched with practical decision-making efficiency to minimize both over-treatment and missed diagnoses. SVM showed the highest area under the ROC curve, indicating superior discriminatory ability between metastatic and non-metastatic cases. While its recall and precision were slightly lower than LR, its high AUC underscores its utility in settings where minimizing classification overlap is a priority. In agreement with prior studies [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], SVM's strength lies in handling complex decision boundaries in high-dimensional clinical data.\u003c/p\u003e \u003cp\u003eRandom Forest, an ensemble learning method, excelled in specificity and precision, making it particularly suited for accurately identifying non-metastatic cases. This performance aligns with other research indicating RF’s robustness in minimizing false positives and enhancing model generalization [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. However, its slightly lower recall compared to XGBoost and Naive Bayes implies that it may still miss a few metastatic cases. Interestingly, both XGBoost and Naive Bayes demonstrated exceptionally high recall, which is crucial in cancer prognostics where false negatives can lead to delayed intervention and worse outcomes. However, their lower specificity suggests a tendency to overpredict metastasis, resulting in a higher rate of false positives. In clinical settings, these models may serve effectively as initial screening tools that prioritize sensitivity, followed by confirmation with more balanced models like LR or RF [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe confusion matrices reinforce these patterns. Logistic Regression and SVM correctly identified positives and negatives, showing minimal misclassification. Random Forest also maintained a strong classification balance, though it exhibited a modest rise in false positives. Conversely, Naive Bayes produced a relatively higher number of false negatives, which raises clinical concerns in oncology where missing a metastasis can delay treatment. XGBoost demonstrated the highest number of false positives, suggesting its application would need to be complemented with confirmatory diagnostics in practice. These findings support the notion that model selection must be context-sensitive: while high recall models may be appropriate for screening, models with higher specificity are more suitable for confirmatory decision support in oncology [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo enhance the interpretability and clinical applicability of the logistic regression model for predicting breast cancer metastasis in a Ghanaian cohort, SHAP analysis was employed. This approach offers transparent, locally accurate explanations of model outputs and ranks features based on their individual contributions to prediction [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. Such interpretability is crucial for building clinician trust and ensuring responsible deployment of machine learning in resource-constrained healthcare environments. The analysis revealed several clinically and contextually significant outcomes. Most notably, lymph node involvement was identified as the strongest driver of metastasis prediction. This finding aligns with global oncology literature, which consistently recognizes lymphatic spread as a principal route of breast cancer metastasis [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. The SHAP beeswarm plot further confirmed that high values of lymph node involvement were consistently associated with increased predicted risk, validating the model’s biological coherence. Closely following were tumor size and cancer stage, both critical indicators in the TNM staging system [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. These features were more influential than molecular markers such as ER, PR, and HER2, a trend that diverges from models trained on Western datasets. This distinction is particularly important in the Ghanaian context, where late-stage presentation is common due to limited screening and awareness [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn such settings, macroscopic clinical features (size and spread) may outweigh molecular subtypes in determining metastatic risk. Another notable outcome was that age categorized by treatment group proved significantly more predictive than raw age. This suggests that age-related risk is not linear, and stratified clinical groupings (under 40, 40–59, 60+) capture meaningful variation, particularly since younger African women often present with more aggressive disease [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. Similarly, time since diagnosis was a meaningful factor, reflecting how disease progression may impact metastatic outcomes. Moderately influential features included type of breast cancer and genetic predisposition. Even in the absence of full genomic sequencing, self-reported family history and clinical categorization proved valuable, suggesting that structured clinical history is a low-cost but high-value addition to diagnostic modeling in resource-limited settings. Menopausal status and estrogen receptor status also played moderate roles, aligning with their known, but context-dependent, influence on recurrence and treatment response [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Interestingly, HER2 status and tumor grade, which are widely regarded as aggressive disease markers in Western literature [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e], had minimal predictive value in this model. This may reflect either true epidemiological differences or systemic gaps in HER2 testing and data completeness in Ghanaian clinical practice [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. It also raises questions about the generalizability of global biomarkers and highlights the need for locally validated tools. Similarly, ethnicity had minimal impact on model prediction. This suggests that while ethnic group classifications (Akan, Ewe, Ga-Adangbe and others) are socially and demographically relevant, they do not appear to independently predict metastasis in this cohort, reinforcing that biological and clinical characteristics carry more weight than sociocultural identifiers in risk modelling.\u003c/p\u003e \u003cp\u003eThe SHAP beeswarm plot provided valuable visual insights into the directionality and magnitude of feature impacts. Features with high SHAP values such as lymph node involvement, tumour size, and stage, strongly increased metastasis predictions. In contrast, features like ER, PR, menopause, and ethnicity exhibited near-zero SHAP values with mixed distributions, indicating limited or inconsistent contribution to model output. These patterns are not only statistically informative but also clinically actionable. The model emphasizes a clear set of high-priority variables nodal involvement, tumor size, stage, and categorized age that can be reliably assessed even in low-resource environments. These variables should be prioritized in screening, triage, and early intervention programs. Moreover, the findings highlight that simple, interpretable model such as logistic regression, when supported by SHAP, can replicate core clinical knowledge and offer actionable insights for oncological care in low- and middle-income countries (LMICs). The model’s alignment with established clinical pathways bolsters its validity and facilitates integration into decision-support systems.\u003c/p\u003e \u003cp\u003eThe risk stratification analysis via Logistic Regression revealed biologically and clinically relevant distinctions. on clinical and hormonal variables differ across breast cancer metastasis risk groups. These differences highlight established predictors of metastasis and may guide risk-tailored clinical decisions. High-risk individuals tend to be younger on average than those in the low-risk group. This finding aligns with prior literature suggesting that early-onset breast cancers may exhibit more aggressive biological behavior, particularly in triple-negative and HER2-positive subtypes, which are more common in younger patients [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. Patients in the high-risk group exhibit significantly higher cancer stages, indicating later diagnosis or more advanced disease progression. Stage at diagnosis is a well-established prognostic factor, with higher stages correlating with greater tumor burden and metastatic potential [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. This underscores the importance of early detection and screening in improving outcomes.\u003c/p\u003e \u003cp\u003eLymph node involvement is markedly higher among high-risk patients. Axillary lymph node status is one of the most reliable indicators of metastasis risk, as nodal metastasis often precedes distant spread [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. This supports the high importance of SHAP, as seen in earlier model interpretations. Larger tumors (\u0026gt; 5 cm) are overrepresented in the high-risk group, while smaller tumors (\u0026lt; 2 cm) dominate the low-risk group. Tumor size is a core element of Tumor-Node-Metastasis TNM staging and has been directly correlated with poor prognosis due to its association with increased cellular proliferation and angiogenesis [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. There is a higher proportion of PR-negative patients in the high-risk category. PR-negativity is a hallmark of more aggressive, hormone-independent tumors and is frequently observed in conjunction with HER2 positivity or triple-negative status, both linked to early recurrence and metastasis [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. Pre-menopausal women are disproportionately represented in the high-risk group. Hormonal fluctuations and denser breast tissue in pre-menopausal women may contribute to diagnostic delays and more aggressive tumor subtypes [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. This observation complements the age distribution and further emphasizes the need for age-specific screening strategies.\u003c/p\u003e \u003cp\u003eThe risk stratification analysis via Logistic Regression revealed biologically and clinically relevant distinctions on clinical and hormonal variables differ across breast cancer metastasis risk groups. Several key observations emerged from the stratified risk analysis that offer deeper insights into breast cancer metastasis patterns within the studied cohort. While the high-risk group exhibited the highest prevalence of genetic mutations, a notable and somewhat unexpected elevation in mutation frequency was also observed among intermediate-risk patients compared to low-risk patients. This suggests that some biologically predisposed individuals may be under-classified based on clinical features alone, highlighting the need to refine risk stratification by incorporating molecular markers [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e], [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Additionally, the intermediate-risk group displayed considerable ethnic diversity, unlike the more demographically concentrated high-risk group, pointing toward potential disparities in healthcare access and diagnostic timing that may influence clinical risk categorization [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e], [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. Hospitalization was heavily skewed toward high-risk individuals, despite some intermediate-risk patients sharing similar genetic or hormonal profiles, indicating that hospitalization history may serve as a proxy for disease burden or healthcare engagement that traditional clinical indicators do not fully capture. Another unexpected finding was the dominance of HER2-negative status across all risk categories, including high-risk patients. This could suggest either a cohort-specific biological trend or the mitigating effects of HER2-targeted therapies, such as trastuzumab, in reducing risk classification [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. In contrast, estrogen receptor (ER) positivity was most prevalent in the low- and intermediate-risk groups, underscoring its known protective role and reinforcing its value in prognostication and treatment planning [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. Furthermore, aggressive subtypes such as Inflammatory Breast Cancer (IBC), Metaplastic Breast Cancer (MBC), and Triple-Negative Breast Cancer (TNBC) were found almost exclusively in the high-risk category, providing visual confirmation of their association with poor prognosis and highlighting their critical role in metastasis risk modeling [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThese findings collectively emphasize the complex interplay between biological, clinical, and social determinants in breast cancer progression and advocate for more nuanced, personalized approaches to risk stratification and patient management.\u003c/p\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eLimitations\u003c/h2\u003e \u003cp\u003eOne significant limitation of our study is the absence of independent validation for the risk categories derived from our machine learning models. Without external validation using a separate dataset, our risk stratification system's generalizability and clinical utility remain uncertain. Specifically, this limitation raises concerns about how well the risk categories would perform across different patient populations or in varied clinical settings. Moreover, the calibration of the risk model how well the predicted probabilities match actual outcomes has not been empirically tested, which could lead to potential overestimation or underestimation of risk. Consequently, the reliability of these risk groups for making critical clinical decisions, such as determining the intensity of surveillance or pre-emptive treatments, is not established. Addressing this limitation in future studies would require gathering independent validation data, possibly through multicentre studies, and refining risk thresholds based on real-world clinical outcomes to enhance the model's applicability and trustworthiness in clinical practice. This approach will improve the model’s credibility and ensure that it can effectively aid in patient management and treatment planning.\u003c/p\u003e \u003c/div\u003e "},{"header":"Conclusions and Recommendation","content":"\u003cp\u003eThese insights are highly relevant in Ghana, where healthcare delivery is challenged by late presentation, diagnostic delays, and inconsistent pathology services. They suggest that a predictive model trained on local data and guided by explainable ML methods can capture both global oncological principles and local disease presentation patterns. The stratified visualizations derived from these models revealed both expected and novel patterns such as the high prevalence of genetic mutations in intermediate-risk individuals, the concentration of aggressive subtypes in the high-risk group, and the unexpectedly dominant HER2-negative status across all risk tiers. These findings not only validated the models' predictive capacity but also uncovered clinically relevant nuances that may be overlooked in conventional staging. Additionally, the influence of ethnicity and hospitalization on risk profiles underscores the importance of incorporating socio-demographic variables into algorithmic models to enhance fairness and real-world applicability.\u003c/p\u003e\u003cp\u003eWe therefore recommend for the development of personalized, equitable, and explainable risk assessment frameworks to improve early identification and intervention for patients at risk of breast cancer metastasis.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003ctable border=\"0\" cellpadding=\"0\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAbbreviation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFull Meaning\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eBC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eBreast Cancer\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eMBC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eMetastatic Breast Cancer\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eER\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eEstrogen Receptor\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003ePR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eProgesterone Receptor\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eHER2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eHuman Epidermal Growth Factor Receptor 2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eML\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eMachine Learning\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eLR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eLogistic Regression\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eRF\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eRandom Forest\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eSVM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eSupport Vector Machine\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eAUC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eArea Under the Curve\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eSHAP\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eSHapley Additive exPlanations\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eIDC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eInvasive Ductal Carcinoma\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eILC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eInvasive Lobular Carcinoma\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eDCIS\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eDuctal Carcinoma In Situ\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eLCIS\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eLobular Carcinoma In Situ\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eTCB\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eTubular Carcinoma of Breast\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eIBC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eInflammatory Breast Cancer\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eTNBC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eTriple-Negative Breast Cancer\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eWHO\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eWorld Health Organization\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eAJCC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eAmerican Joint Committee on Cancer\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eBRCA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eBreast Cancer gene (BRCA1/BRCA2)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eTNM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eTumor-Node-Metastasis (Staging System)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eROC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eReceiver Operating Characteristic\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003eNB\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 225px;\"\u003e\n \u003cp\u003eNaive Bayes\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis retrospective study was reviewed and approved by the Institutional Review Board (IRB), Korle-Bu Teaching Hospital, Accra, Ghana (IRB KTHI5/1083703). The need for written informed consent was waived by the IRB because the study used de-identified, routinely collected clinical data, in accordance with national regulations. Also, since the research did not involve direct human interaction, no written consent process was required.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClinical trial number:\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003enot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets used and/or analysed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThere was no specific funding for this study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAYA conceptualized the study and led the study design, methodology development, data curation, formal analysis, machine learning modeling, manuscript drafting, visualization, and overall project supervision. SBA was responsible for data collection, validation of clinical findings, and critical manuscript revision. EA contributed to visualization, manuscript review, and editing. VUG conducted statistical analysis, assisted with machine learning modeling, validated results, and contributed to manuscript revision. JA was involved in manuscript editing and revision. SO contributed to manuscript revision and editing. SBB conducted the literature review, assisted with clinical data analysis, supported result validation, and edited the manuscript. AMB contributed to data interpretation, supervised data collection activities, and participated in manuscript review and editing. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eWorld Health Organization, \u0026ldquo;Breast cancer.\u0026rdquo; Accessed: May 14, 2025. [Online]. Available: https://www.who.int/news-room/fact-sheets/detail/breast-cancer\u003c/li\u003e\n\u003cli\u003eF. Bray, J. Ferlay, I. Soerjomataram, R. L. Siegel, L. A. Torre, and A. Jemal, \u0026ldquo;Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries,\u0026rdquo; \u003cem\u003eCA Cancer J Clin\u003c/em\u003e, vol. 68, no. 6, pp. 394\u0026ndash;424, Nov. 2018, doi: 10.3322/CAAC.21492,.\u003c/li\u003e\n\u003cli\u003eE. Jedy-Agba, V. McCormack, C. Adebamowo, and I. dos-Santos-Silva, \u0026ldquo;Stage at diagnosis of breast cancer in sub-Saharan Africa: a systematic review and meta-analysis,\u0026rdquo; \u003cem\u003eLancet Glob Health\u003c/em\u003e, vol. 4, no. 12, p. e923, Dec. 2016, doi: 10.1016/S2214-109X(16)30259-5.\u003c/li\u003e\n\u003cli\u003eS. Y. Opoku, M. Benwell, and J. Yarney, \u0026ldquo;Knowledge, attitudes, beliefs, behaviour and breast cancer screening practices in Ghana, West Africa.,\u0026rdquo; 2012. Accessed: May 14, 2025. [Online]. Available: http://197.255.68.203/handle/123456789/3828\u003c/li\u003e\n\u003cli\u003eA. C. Mensah \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Survival Outcomes of Breast Cancer in Ghana: An Analysis of Clinicopathological Features,\u0026rdquo; \u003cem\u003eOpen Access Library Journal\u003c/em\u003e, vol. 3, no. 1, pp. 1\u0026ndash;11, Jan. 2016, doi: 10.4236/OALIB.1102145.\u003c/li\u003e\n\u003cli\u003eAmerican Cancer Society, \u0026ldquo;Treatment of Stage IV (Metastatic) Breast Cancer .\u0026rdquo; Accessed: May 14, 2025. [Online]. Available: https://www.cancer.org/cancer/types/breast-cancer/treatment/treatment-of-breast-cancer-by-stage/treatment-of-stage-iv-advanced-breast-cancer.html\u003c/li\u003e\n\u003cli\u003eR. L. Siegel, K. D. Miller, N. S. Wagle, and A. Jemal, \u0026ldquo;Cancer statistics, 2023,\u0026rdquo; \u003cem\u003eCA Cancer J Clin\u003c/em\u003e, vol. 73, no. 1, pp. 17\u0026ndash;48, Jan. 2023, doi: 10.3322/CAAC.21763,.\u003c/li\u003e\n\u003cli\u003eT. Reinert, A. B. A. de Souza, G. P. Sartori, F. M. Obst, and C. H. Barrios, \u0026ldquo;Highlights of the 17th St Gallen International Breast Cancer Conference 2021: Customising local and systemic therapies,\u0026rdquo; \u003cem\u003eEcancermedicalscience\u003c/em\u003e, vol. 15, 2021, doi: 10.3332/ECANCER.2021.1236,.\u003c/li\u003e\n\u003cli\u003eE. M. Der \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Triple-negative breast cancer in Ghanaian women: The Korle Bu Teaching Hospital experience,\u0026rdquo; \u003cem\u003eBreast Journal\u003c/em\u003e, vol. 21, no. 6, pp. 627\u0026ndash;633, Nov. 2015, doi: 10.1111/TBJ.12527,.\u003c/li\u003e\n\u003cli\u003eN. Harbeck \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Breast cancer,\u0026rdquo; \u003cem\u003eNat Rev Dis Primers\u003c/em\u003e, vol. 5, no. 1, Dec. 2019, doi: 10.1038/S41572-019-0111-2,.\u003c/li\u003e\n\u003cli\u003eF. Cardoso \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Early breast cancer: ESMO Clinical Practice Guidelines for diagnosis, treatment and follow-up,\u0026rdquo; \u003cem\u003eAnnals of Oncology\u003c/em\u003e, vol. 30, no. 8, pp. 1194\u0026ndash;1220, Aug. 2019, doi: 10.1093/annonc/mdz173.\u003c/li\u003e\n\u003cli\u003eE. Jiagge \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Comparative Analysis of Breast Cancer Phenotypes in African American, White American, and West Versus East African patients: Correlation Between African Ancestry and Triple-Negative Breast Cancer,\u0026rdquo; \u003cem\u003eAnn Surg Oncol\u003c/em\u003e, vol. 23, no. 12, pp. 3843\u0026ndash;3849, Nov. 2016, doi: 10.1245/S10434-016-5420-Z,.\u003c/li\u003e\n\u003cli\u003eJ. A. Cruz and D. S. Wishart, \u0026ldquo;Applications of machine learning in cancer prediction and prognosis,\u0026rdquo; \u003cem\u003eCancer Inform\u003c/em\u003e, vol. 2, pp. 59\u0026ndash;77, 2006, doi: 10.1177/117693510600200030.\u003c/li\u003e\n\u003cli\u003eA. Esteva \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;A guide to deep learning in healthcare,\u0026rdquo; \u003cem\u003eNat Med\u003c/em\u003e, vol. 25, no. 1, pp. 24\u0026ndash;29, Jan. 2019, doi: 10.1038/S41591-018-0316-Z;SUBJMETA=114,1305,1647,48,631,692,700;KWRD=BIOINFORMATICS,HEALTH+CARE,MACHINE+LEARNING.\u003c/li\u003e\n\u003cli\u003eK. Kourou, T. P. Exarchos, K. P. Exarchos, M. V. Karamouzis, and D. I. Fotiadis, \u0026ldquo;Machine learning applications in cancer prognosis and prediction,\u0026rdquo; \u003cem\u003eComput Struct Biotechnol J\u003c/em\u003e, vol. 13, pp. 8\u0026ndash;17, 2015, doi: 10.1016/j.csbj.2014.11.005.\u003c/li\u003e\n\u003cli\u003eF. Pedregosa \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Scikit-learn: Machine Learning in Python,\u0026rdquo; Jan. 2012, Accessed: May 14, 2025. [Online]. Available: http://arxiv.org/abs/1201.0490\u003c/li\u003e\n\u003cli\u003eS. B. Kotsiantis, I. D. Zaharakis, and P. E. Pintelas, \u0026ldquo;Machine learning: A review of classification and combining techniques,\u0026rdquo; \u003cem\u003eArtif Intell Rev\u003c/em\u003e, vol. 26, no. 3, pp. 159\u0026ndash;190, Nov. 2006, doi: 10.1007/S10462-007-9052-3/METRICS.\u003c/li\u003e\n\u003cli\u003eD. W. Hosmer, S. Lemeshow, and R. X. Sturdivant, \u0026ldquo;Applied Logistic Regression: Third Edition,\u0026rdquo; \u003cem\u003eApplied Logistic Regression: Third Edition\u003c/em\u003e, pp. 1\u0026ndash;510, Aug. 2013, doi: 10.1002/9781118548387.\u003c/li\u003e\n\u003cli\u003eI. Guyon, J. Weston, S. Barnhill, and V. Vapnik, \u0026ldquo;Gene selection for cancer classification using support vector machines,\u0026rdquo; \u003cem\u003eMach Learn\u003c/em\u003e, vol. 46, no. 1\u0026ndash;3, pp. 389\u0026ndash;422, 2002, doi: 10.1023/A:1012487302797/METRICS.\u003c/li\u003e\n\u003cli\u003eT. Chen and C. Guestrin, \u0026ldquo;XGBoost: A scalable tree boosting system,\u0026rdquo; \u003cem\u003eProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining\u003c/em\u003e, vol. 13-17-August-2016, pp. 785\u0026ndash;794, Aug. 2016, doi: 10.1145/2939672.2939785.\u003c/li\u003e\n\u003cli\u003eR. Kohavi, \u0026ldquo;A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection,\u0026rdquo; \u003cem\u003eInternational Joint Conference on Artificial Intelligence\u003c/em\u003e, 1995.\u003c/li\u003e\n\u003cli\u003eS. Varma and R. Simon, \u0026ldquo;Bias in error estimation when using cross-validation for model selection,\u0026rdquo; \u003cem\u003eBMC Bioinformatics\u003c/em\u003e, vol. 7, no. 1, pp. 1\u0026ndash;8, Feb. 2006, doi: 10.1186/1471-2105-7-91/FIGURES/4.\u003c/li\u003e\n\u003cli\u003eE. W. Steyerberg \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Assessing the performance of prediction models: A framework for traditional and novel measures,\u0026rdquo; \u003cem\u003eEpidemiology\u003c/em\u003e, vol. 21, no. 1, pp. 128\u0026ndash;138, Jan. 2010, doi: 10.1097/EDE.0B013E3181C30FB2,.\u003c/li\u003e\n\u003cli\u003eL. Dihge, M. Ohlsson, P. Ed\u0026eacute;n, P. O. Bendahl, and L. Ryd\u0026eacute;n, \u0026ldquo;Artificial neural network models to predict nodal status in clinically node-negative breast cancer,\u0026rdquo; \u003cem\u003eBMC Cancer\u003c/em\u003e, vol. 19, no. 1, Jun. 2019, doi: 10.1186/S12885-019-5827-6,.\u003c/li\u003e\n\u003cli\u003eR. Ha \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Predicting Breast Cancer Molecular Subtype with MRI Dataset Utilizing Convolutional Neural Network Algorithm,\u0026rdquo; \u003cem\u003eJ Digit Imaging\u003c/em\u003e, vol. 32, no. 2, p. 276, Apr. 2019, doi: 10.1007/S10278-019-00179-2.\u003c/li\u003e\n\u003cli\u003eS. M. Lundberg and S. I. Lee, \u0026ldquo;A Unified Approach to Interpreting Model Predictions,\u0026rdquo; \u003cem\u003eAdv Neural Inf Process Syst\u003c/em\u003e, vol. 2017-December, pp. 4766\u0026ndash;4775, May 2017, Accessed: May 14, 2025. [Online]. Available: https://arxiv.org/pdf/1705.07874\u003c/li\u003e\n\u003cli\u003eL. Breiman, \u0026ldquo;Random forests,\u0026rdquo; \u003cem\u003eMach Learn\u003c/em\u003e, vol. 45, no. 1, pp. 5\u0026ndash;32, Oct. 2001, doi: 10.1023/A:1010933404324/METRICS.\u003c/li\u003e\n\u003cli\u003eP. Wang \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Machine Learning Models for Diagnosing Glaucoma from Retinal Nerve Fiber Layer Thickness Maps,\u0026rdquo; \u003cem\u003eOphthalmol Glaucoma\u003c/em\u003e, vol. 2, no. 6, pp. 422\u0026ndash;428, Nov. 2019, doi: 10.1016/j.ogla.2019.08.004.\u003c/li\u003e\n\u003cli\u003eD. Delen, G. Walker, and A. Kadam, \u0026ldquo;Predicting breast cancer survivability: A comparison of three data mining methods,\u0026rdquo; \u003cem\u003eArtif Intell Med\u003c/em\u003e, vol. 34, no. 2, pp. 113\u0026ndash;127, Jun. 2005, doi: 10.1016/J.ARTMED.2004.07.002,.\u003c/li\u003e\n\u003cli\u003eAmerican Joint Committee on Cancer, \u0026ldquo;Cancer Staging Systems.\u0026rdquo; Accessed: May 14, 2025. [Online]. Available: https://www.facs.org/quality-programs/cancer-programs/american-joint-committee-on-cancer/cancer-staging-systems/\u003c/li\u003e\n\u003cli\u003eJ. Clegg-Lamptey, J. Dakubo, and Y. N. Attobra, \u0026ldquo;Why Do Breast Cancer Patients Report Late or Abscond During Treatment in Ghana? A Pilot Study,\u0026rdquo; \u003cem\u003eGhana Med J\u003c/em\u003e, vol. 43, no. 3, p. 127, Sep. 2009, Accessed: May 14, 2025. [Online]. Available: https://pmc.ncbi.nlm.nih.gov/articles/PMC2810246/\u003c/li\u003e\n\u003cli\u003eC. K. Anders, R. Johnson, J. Litton, M. Phillips, and A. Bleyer, \u0026ldquo;Breast Cancer Before Age 40 Years,\u0026rdquo; \u003cem\u003eSemin Oncol\u003c/em\u003e, vol. 36, no. 3, pp. 237\u0026ndash;249, Jun. 2009, doi: 10.1053/j.seminoncol.2009.03.001.\u003c/li\u003e\n\u003cli\u003eE. A. Rakha \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Invasive lobular carcinoma of the breast: Response to hormonal therapy and outcomes,\u0026rdquo; \u003cem\u003eEur J Cancer\u003c/em\u003e, vol. 44, no. 1, pp. 73\u0026ndash;83, Jan. 2008, doi: 10.1016/j.ejca.2007.10.009.\u003c/li\u003e\n\u003cli\u003eC. M. Perou \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Molecular portraits of human breast tumours,\u0026rdquo; \u003cem\u003eNature\u003c/em\u003e, vol. 406, no. 6797, pp. 747\u0026ndash;752, Aug. 2000, doi: 10.1038/35021093,.\u003c/li\u003e\n\u003cli\u003eL. Denny \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Human papillomavirus prevalence and type distribution in invasive cervical cancer in sub-Saharan Africa,\u0026rdquo; \u003cem\u003eInt J Cancer\u003c/em\u003e, vol. 134, no. 6, pp. 1389\u0026ndash;1398, Mar. 2014, doi: 10.1002/IJC.28425,.\u003c/li\u003e\n\u003cli\u003eS. B. Edge and C. C. Compton, \u0026ldquo;The american joint committee on cancer: The 7th edition of the AJCC cancer staging manual and the future of TNM,\u0026rdquo; \u003cem\u003eAnn Surg Oncol\u003c/em\u003e, vol. 17, no. 6, pp. 1471\u0026ndash;1474, Jun. 2010, doi: 10.1245/S10434-010-0985-4/TABLES/1.\u003c/li\u003e\n\u003cli\u003eM. Cianfrocca and L. J. Goldstein, \u0026ldquo;Prognostic and Predictive Factors in Early-Stage Breast Cancer,\u0026rdquo; \u003cem\u003eOncologist\u003c/em\u003e, vol. 9, no. 6, pp. 606\u0026ndash;616, Nov. 2004, doi: 10.1634/THEONCOLOGIST.9-6-606\u0026rsquo;)).\u003c/li\u003e\n\u003cli\u003eA. Goldhirsch \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Personalizing the treatment of women with early breast cancer: Highlights of the st gallen international expert consensus on the primary therapy of early breast Cancer 2013,\u0026rdquo; \u003cem\u003eAnnals of Oncology\u003c/em\u003e, vol. 24, no. 9, pp. 2206\u0026ndash;2223, Sep. 2013, doi: 10.1093/annonc/mdt303.\u003c/li\u003e\n\u003cli\u003eN. Hamajima \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Menarche, menopause, and breast cancer risk: Individual participant meta-analysis, including 118 964 women with breast cancer from 117 epidemiological studies,\u0026rdquo; \u003cem\u003eLancet Oncol\u003c/em\u003e, vol. 13, no. 11, pp. 1141\u0026ndash;1151, Nov. 2012, doi: 10.1016/S1470-2045(12)70425-4.\u003c/li\u003e\n\u003cli\u003eN. Petrucelli, M. B. Daly, and G. L. Feldman, \u0026ldquo;Hereditary breast and ovarian cancer due to mutations in BRCA1 and BRCA2,\u0026rdquo; \u003cem\u003eGenetics in Medicine\u003c/em\u003e, vol. 12, no. 5, pp. 245\u0026ndash;259, May 2010, doi: 10.1097/GIM.0b013e3181d38f2f.\u003c/li\u003e\n\u003cli\u003eW. D. Foulkes, I. E. Smith, and J. S. Reis-Filho, \u0026ldquo;Triple-negative breast cancer,\u0026rdquo; \u003cem\u003eN Engl J Med\u003c/em\u003e, vol. 363, no. 20, pp. 1938\u0026ndash;1948, Nov. 2010, doi: 10.1056/NEJMRA1001389.\u003c/li\u003e\n\u003cli\u003eL. A. Carey \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Race, breast cancer subtypes, and survival in the Carolina Breast Cancer Study,\u0026rdquo; \u003cem\u003eJ Am Med Assoc\u003c/em\u003e, vol. 295, no. 21, pp. 2492\u0026ndash;2502, Jun. 2006, doi: 10.1001/JAMA.295.21.2492,.\u003c/li\u003e\n\u003cli\u003eC. E. DeSantis \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Breast cancer statistics, 2019,\u0026rdquo; \u003cem\u003eCA Cancer J Clin\u003c/em\u003e, vol. 69, no. 6, pp. 438\u0026ndash;451, Nov. 2019, doi: 10.3322/CAAC.21583,.\u003c/li\u003e\n\u003cli\u003eD. J. Slamon \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Use of Chemotherapy plus a Monoclonal Antibody against HER2 for Metastatic Breast Cancer That Overexpresses HER2,\u0026rdquo; \u003cem\u003eNew England Journal of Medicine\u003c/em\u003e, vol. 344, no. 11, pp. 783\u0026ndash;792, Mar. 2001, doi: 10.1056/NEJM200103153441101,.\u003c/li\u003e\n\u003cli\u003eE. A. Rakha, M. E. El-Sayed, A. R. Green, A. H. S. Lee, J. F. Robertson, and I. O. Ellis, \u0026ldquo;Prognostic markers in triple-negative breast cancer,\u0026rdquo; \u003cem\u003eCancer\u003c/em\u003e, vol. 109, no. 1, pp. 25\u0026ndash;32, Jan. 2007, doi: 10.1002/CNCR.22381,.\u003c/li\u003e\n\u003cli\u003eS. Dawood \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;International expert panel on inflammatory breast cancer: Consensus statement for standardized diagnosis and treatment,\u0026rdquo; \u003cem\u003eAnnals of Oncology\u003c/em\u003e, vol. 22, no. 3, pp. 515\u0026ndash;523, 2011, doi: 10.1093/ANNONC/MDQ345,.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Breast Cancer Metastasis, Machine Learning, Risk Stratification","lastPublishedDoi":"10.21203/rs.3.rs-6671298/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6671298/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBreast cancer remains the most diagnosed cancer among women globally. It is a leading cause of cancer-related deaths, with disproportionately high mortality rates in sub-Saharan Africa due to late-stage presentation and limited access to diagnostic and treatment services. In Ghana, nearly 70% of breast cancer cases are diagnosed at advanced or metastatic stages, underscoring the urgent need for accurate and context-specific prognostic tools. Here, we assembled clinical, molecular and demographic data from 558 breast cancer patients at Korle-Bu Teaching Hospital in Ghana to develop and compare five supervised machine‐learning model logistic regression, random forest, support vector machine, XGBoost and na\u0026iuml;ve Bayes for metastasis prediction. Using standardized preprocessing and repeated 10-fold stratified cross-validation, logistic regression achieved the most balanced performance (96% cross-validated accuracy; 93% test‐set accuracy) and facilitated interpretability via SHAP analysis, which identified lymph node involvement, tumour size and stage as top predictors. We further refined prognostic utility by stratifying patients into low, intermediate and high metastatic risk groups, revealing distinct clinical and biological profiles high‐risk individuals exhibited aggressive subtypes (triple-negative, HER2-positive) and advanced staging, whereas low-risk patients showed favourable receptor status and earlier disease. These findings demonstrate that tailored, interpretable machine-learning tools can support accurate, resource-appropriate metastasis risk assessment in low-resource settings.\u003c/p\u003e","manuscriptTitle":"Risk Stratification of Breast Cancer Metastasis: A Predictive Modelling Framework Using Clinical and Hormonal Receptor Data in a Ghanaian Cohort","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-20 15:29:12","doi":"10.21203/rs.3.rs-6671298/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c75dd96a-dcc2-4bd0-b92e-adbd151bdbfc","owner":[],"postedDate":"June 20th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-08-13T09:54:06+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-20 15:29:12","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6671298","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6671298","identity":"rs-6671298","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.