Full text
82,195 characters
· extracted from
preprint-html
· click to expand
Can interpretability and accuracy coexist in cancer survival analysis? | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Can interpretability and accuracy coexist in cancer survival analysis? View ORCID Profile Piyush Borole , Tongjie Wang , View ORCID Profile Antonio Vergari , View ORCID Profile Ajitha Rajan doi: https://doi.org/10.1101/2025.04.11.648380 Piyush Borole 1 School of Informatics, University of Edinburgh , 10 Crichton Street, Edinburgh, EH8 9AB, UK Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Piyush Borole For correspondence: p.borole{at}sms.ed.ac.uk arajan{at}ed.ac.uk Tongjie Wang 1 School of Informatics, University of Edinburgh , 10 Crichton Street, Edinburgh, EH8 9AB, UK Find this author on Google Scholar Find this author on PubMed Search for this author on this site Antonio Vergari 1 School of Informatics, University of Edinburgh , 10 Crichton Street, Edinburgh, EH8 9AB, UK Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Antonio Vergari Ajitha Rajan 1 School of Informatics, University of Edinburgh , 10 Crichton Street, Edinburgh, EH8 9AB, UK Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Ajitha Rajan For correspondence: p.borole{at}sms.ed.ac.uk arajan{at}ed.ac.uk Abstract Full Text Info/History Metrics Supplementary material Preview PDF Abstract Survival analysis refers to statistical procedures used to analyze data that focuses on the time until an event occurs, such as death in cancer patients. Traditionally, the linear Cox Proportional Hazards (CPH) model is widely used due to its inherent interpretability. CPH model help identify key disease-associated factors (through feature weights), providing insights into patient risk of death. However, their reliance on linear assumptions limits their ability to capture the complex, non-linear relationships present in real-world data. To overcome this, more advanced models, such as neural networks, have been introduced, offering significantly improved predictive accuracy. However, these gains come at the expense of interpretability, which is essential for clinical trust and practical application. To address the trade-off between predictive accuracy and interpretability in survival analysis, we propose ConSurv, a concept bottleneck model that maintains state-of-the-art performance while providing transparent and interpretable insights. Using gene expression and clinical data from breast cancer patients, ConSurv captures complex feature interactions and predicts patient risk. By offering clear, biologically meaningful explanations for each prediction, ConSurv attempts to build trust among clinicians and researchers in using the model for informed decision-making. 1 Introduction Survival analysis is essential for estimating the time until events such as death, relapse, or recovery occur. It lays the groundwork for assessing disease severity and understanding the factors that influence patient outcomes [ 1 , 2 ]. Traditional methods like the Cox Proportional Hazards (CPH) model, a linear approach, have been widely used due to their straightforward interpretability and robust ability to estimate hazard ratios (a type of relative risk) [ 2 , 3 ]. However, with the advent of high-dimensional data such as RNA sequencing (RNA-seq), these linear models face significant limitation in achieving high performance. RNA expression data measure gene activity levels and play a critical role in survival research by revealing dysregulated genes associated with disease progression. The complex and non-linear relationships inherent in high-dimensional RNA-seq data challenge the CPH model’s ability to accurately extract meaningful features. This limitation results in less precise predictions compared to more advanced machine learning models that can capture these intricate patterns. However, CPH, being a linear model, possesses inherent interpretability, allowing the importance of each input feature to be inferred from its feature weights. As illustrated in Figure 1 , such a linear model offers high interpretability (on the x-axis) but low accuracy (on the y-axis). Download figure Open in new tab Fig. 1 Accuracy and Interpretability trade-off. Linear and rule-based models are inherently interpretable due to their explicit rules and feature weights used for prediction. However, they often exhibit lower performance compared to more complex, black-box models such as neural networks. As complexity of models increase, they tend to improve in accuracy, but this often comes at the cost of interpretability, as their decision-making processes become more difficult to understand. Tree-based algorithms, while offering transparent decision pathways, can also become challenging to interpret due to their large number of decision nodes and depth of the tree. A concept bottleneck model, on the other hand, is a gray-box model that strikes a balance between accuracy and interpretability. It consists of two parts: a black-box component that predicts human-interpretable “concepts” and a second part that uses these concepts exclusively to make the final prediction, typically through a simple (often linear) model. At the other extreme, emerging non-linear neural network models such as DeepSurv [ 4 ], offer improved accuracy by capturing the complexity of the data but provide very low interpretability. This lack of interpretability makes it challenging for clinicians and researchers to derive meaningful causal relationships between input features and risk predictions. This is illustrated in Figure 1 , where the neural network ranks high on accuracy dimension but very low on interpretability dimension. The opacity of these models and the technical expertise required for their interpretation lead to a trust gap, preventing their adoption in clinical settings. Current research on interpretability of these neural networks for survival analysis focuses on post-hoc explanations which involve applying additional models to ascertain contribution of each feature (called explanations) to the predictions of these black-box models. For example, SurvLIME [ 5 ] uses the Local Interpretable Model-agnostic Explanations (LIME, [ 6 ]) framework, while SurvSHAP [ 7 ] and AUTOSurv [ 8 ] apply SHapley Additive exPlanations (SHAP, [ 9 ]). Both LIME and SHAP are perturbation-based models that require multiple evaluation passes over the model to explain a single prediction. Furthermore, empirical evidence demonstrates that these models could fail to assign importance to relevant features and produce explanations that are unfaithful [ 10 ]. These limitations make perturbation-based models not only computationally expensive but also, more importantly, unreliable [ 11 , 12 ]. Other approaches, utilize DeepLIFT [ 13 , 14 ] and backpropagation [ 15 ] to quantify the contribution of input features to the risk prediction. However, these gradient-based methods are highly sensitive to noise and non-linearities, leading to explanations that are often unreliable and unfaithful to the models they aim to interpret [ 16 ]. Relying on such explanations can lead to incorrect conclusions about the data, ultimately resulting in a trust deficit. The Cox-nnet model [ 17 ], a single-layer perceptron, takes an interesting approach to interpretability by capturing biologically relevant functions (such as the p53 pathway) in its hidden nodes, which serve as surrogate features for predicting patient survival. However, their approach of quantifying the contribution of each input to every node of the network quickly becomes infeasible as deeper networks are required to improve accuracy. Tree-based models, such as XGBoost [ 18 , 19 ], provide a partial solution by better balancing transparency and accuracy than deep learning models as seen in Figure 1 . They offer transparency by using features as nodes within the trees, which could be individually examined, while also capturing non-linear relationships more effectively than linear models such as CPH. However, when applied to high-dimensional data such as RNA-seq, the trained trees can have vast number of decision nodes making them overwhelmingly complex. This growth in complexity inherently hinders interpretability. As seen in the above examples of linear, tree-based, and neural network models, there is a trade-off between performance and interpretability, which is illustrated as accuracy-interpretability trade-off in Figure 1 . This trade-off suggests that as model complexity increases, accuracy improves, but interpretability declines. An ideal model should balance both aspects to ensure usability and trustworthiness. To balance the accuracy-interpretability trade-off in survival prediction, we propose a gray-box approach using concept bottleneck models [ 20 ]. These models consist of two parts: one predicts human-interpretable concepts, and the other—typically a linear model—uses these concepts for final predictions. We apply this approach to risk prediction on RNA-seq and clinical data for breast cancer and demonstrate performance comparable to the existing models. We also show that the key concepts can differentiate high- and low-risk patients demonstrating the usefulness of these concepts. Finally, we assess the stability of these concepts by evaluating their importance and robustness by examining the consistency in extracted concepts across multiple runs. Our model, ConSurv, inherently embeds interpretability as a core component by aggregating genetic and clinical features into concepts, providing humanly understandable and biologically meaningful insights. 2 Results In the results section, we evaluate the models based on two key dimensions: performance in predicting LogRisk and ability in explaining model decision-making (i.e. interpretability). The existing models assessed in this study are CPH, XGBoost, and DeepSurv that are popular and widely used in survival analysis and covers the spectrum on accuracy-interpretability tradeoff of Figure 1 . Performance is measured using the Concordance Index (or CI, see Section 4.4 ). The result section begins with description of the dataset ( Section 2.1 , TCGA breast cancer data) used in this study. Sections 2.2 and 2.3 present the performance evaluation of existing models and our proposed ConSurv framework. Finally, sections 2.4 to 2.6 focus on model interpretability. We begin by outlining the interpretability aspects of the linear CPH, the black-box DeepSurv (using SHAP), and the gray-box ConSurv model. This is followed by an exploration of the extracted concepts, assessing their alignment with biological insights, as well as their stability and robustness. Table 1 provides a summary of the models used in this study and their interpretability. View this table: View inline View popup Download powerpoint Table 1 Survival Analysis Models And Their Interpretability 2.1 Data In this study, we utilized breast cancer RNA-seq (RNA profiling table) data and corresponding clinical information obtained from The Cancer Genome Atlas (TCGA) through the Genomic Data Commons (GDC) portal. All data used in this study are de-identified and publicly available, eliminating the need for additional ethical approval. The study complies with the data usage policies of TCGA and GDC, which allow the use of the data for research purposes. RNA profiling table, contains more than 60,000 gene expression values per patient. Given the limitations of physical memory and training time constraints, we opt for the variance threshold method to filter out genes that contain less information. Therefore, we rank the index of dispersion for each gene and select the top 1000 highest-ranked genes. In addition to these 1000 genes, we added 8 clinical features associated with patients to the data set. These clinical features were one hot encoded. In total, there were 1209 patients with 1056 features in our final dataset. Each patient entry was accompanied with survival time ( t ) and event indicator ( e , which indicates if event, here death, occurred or not). 2.2 Non-linear models demonstrate significantly better performance than linear model The purpose of risk prediction is to stratify patients such that those at higher risk would have lower survival time or probabilities. To predict this risk (predicted as LogRisk), we used 1000 genes from RNA-seq data and eight associated clinical features for 1209 breast cancer patients. Model performance was evaluated using the Concordance Index (CI), as described in Section 4.4 . In our study, we utilized widely used models - CPH, XGBoost, and DeepSurv, from the linear, tree-based, and neural network categories, respectively, as shown in Table 1 (and Figure 1 ). CPH is a linear model that provides association between the survival time of patients and the features. XGBoost is tree-based modeling approach used for survival regression that predicts survival time [ 19 ]. DeepSurv is a neural network that predicts LogRisk using Cox-loss function (See Section 4.6 ) [ 4 ]. In addition to these, we trained a multi-layer perceptron (MLP) to establish baseline parameters (such as network depth) for our concept bottleneck model. This base MLP was subsequently transformed to develop our ConSurv model as described in Section 4.2 . Data was split ten times using different seeds into training, validation and test sets (70%, 15% and 15%). Training and validation sets were used for hyperparameter tuning (See Section 4.3 ). Performance (i.e. CI) of the best models was evaluated using the ten test sets. The linear CPH model exhibits the lowest median performance (median CI 0.65), while the black-box DeepSurv (median CI 0.75) achieves the highest median performance, with XGBoost (median CI 0.69) falling in between. The MLP demonstrates performance comparable to DeepSurv (median CI 0.73, p-value > 0.05, see Supplemental Table 2). Figures 2a,b plots the CI for the above described models. For XGBoost, hyperparameter tuning revealed that the highest performance was achieved with a tree depth of 10. During each run, 330–390 trained trees were generated. XGBoost also exhibits the highest variance in performance, with the lowest CI being 0.53 and the highest CI reaching 0.87 as seen in Figure 2a . We pruned the trained XGBoost trees to depths ranging from 1 to 10 and found that pruning upto a depth of 5 resulted in a marginal loss in performance (within 5% of the model with a depth of 10), as shown in Figures 2c,d . An illustration of tree pruning is in Supplemental Figure 1 which shows a tree with original depth 10 pruned upto depth 1. Across the runs, only 14 features (of 1056) were consistently used by XGBoost, as shown in Table 2 . This low overlap of common features could explain the high variability in its performance. Nevertheless, the most commonly used features (predominantly genes) are known to be associated with breast cancer survival, and their references are provided in Table 2 . Download figure Open in new tab Fig. 2 Performance of the survival models. a. Concordance Index or CI on test sets for Cox Proportional Hazards (CPH), XGBoost, DeepSurv, MLP and ConSurv models. Top 100 concepts were used for ConSurv-XGB (C-XGB Top100) and top 25 concepts for ConSurv-Rule (C-Rule Top25). b. Shows a upward CI trend as models progress from a fully interpretable to black-box. c. demonstrates the impact of pruning XGBoost trees to depths of 1–9 (with 10 being the original depth). d. Notably, the XGBoost performance with pruning depths of 5–9 remains within 5% of the original CI for depth 10 across all ten runs. e. Validation CI after training C-XGB using the top 5, 10, 50, 100, and 200 concepts ranked by weights in ConSurv-XGB using all concepts. f. Validation CI after training C-Rule model with top 10, 25, and 50 concepts ranked by weights from ConSurv-Rule using all concepts. (For all plots, errorbars: confidence interval at 95%, estimator: median) View this table: View inline View popup Download powerpoint Table 2 Common features across 10 different runs from XGBoost and RuleKit (Genes converted using https://www.biotools.fr/ ) Survival classification using Rule-based model As an additional means for pattern extraction used in our study, we trained a contrast set rule-based model (See Section 4.1 ), a form of association rule mining. The Contrast Set Rule-based model identifies relationships between features to cluster data points into classes using an ‘IF-THEN’ formulation. RuleKit is an implementation of this model which was used in our study. We transformed our problem into a classification task by categorizing patients into three groups based on survival times as standard practice: 10 years. The classification AUROC were 0.59 ( 10 years) for the three classes. Each run produced between 61–83 rules, with median rule length (or number of features per rule) of 21. About 75% of rules from each run belonged to class 10 survival years. Across runs, 44 features were consistently used by RuleKit, as shown in Table 2 . To determine whether these consistently relevant common genes of XGBoost and RuleKit play a central role in survival outcomes, we performed Gene Set Variation Analysis (GSVA) [ 21 ]. GSVA enables a non-parametric assessment of gene set enrichment. The dataset specifications for running GSVA are provided in Section 4.7 . Figures 3a,b illustrate the survival stratification based on common genes for both models. Download figure Open in new tab Fig. 3 Kaplan-Meier (KM) survival plots. a. The KM plot for common genes selected across runs (listed in Table 2 ) for XGBoost shows a high level of stratification between high- and low-risk groups ( p = 0.00093). b. The KM plot for common genes selected across runs (listed in Table 2 ) for RuleKit also demonstrates significant stratification between high- and low-risk groups ( p = 0.0029). c. The KM plots for the top 5 concepts produced by the ConSurv-XGB model show significant separation between low- and high-risk groups for all but one concept (concept 2). The features forming the concept are listed under respective KM plot. d. Similarly, the KM plots for the top 5 concepts produced by the ConSurv-Rule model demonstrate significant separation between low- and high-risk groups for all but one concept (concept 4). The processes obtained through ORA for each concept are listed below the concept. For XGBoost and RuleKit, a significant stratification between high-risk and low-risk groups is observed ( p = 0.00093 and p = 0.0029, respectively). GSVA further reveals that the common genes identified by XGBoost are associated with the Gene Ontology (GO) molecular function GO:0005198 (i.e., structural molecule activity), which is relatively nonspecific. In contrast, the common genes identified by RuleKit are linked to a more specific biological process—programmed cell death (GO biological process GO:0012501). The ability of both models to leverage common genes for survival stratification underscores their effectiveness in identifying core biological features that contribute to predictive performance. 2.3 ConSurv enhances interpretability while maintaining competitive performance Our proposed model ConSurv, is a concept bottleneck model that predicts risk as described in Equation (5) with human interpretable concepts as an intermediate prediction. The architecture is based on the bilayered MLP model with the second layer modified to capture predetermined concepts as illustrated in Figure 4a . In our framework, concepts are defined as sets of grouped features extracted either from trained trees generated by XGBoost or from rules learned using the Contrast Set Rule-based model (RuleKit), as described in Section 4.2 . The two ConSurv models are named as ConSurv-XGB-all and ConSurv-Rule-all respectively. The ‘-all’ indicates use of all concepts extracted by XGBoost or RuleKit. Download figure Open in new tab Fig. 4 ConSurv illustration. a. The network on the left depicts a typical multi-layered perceptron model with two hidden layers. This densely connected architecture provides no explanation for its decision-making process. The network on the right illustrates our concept bottleneck model, ConSurv. ConSurv consists of one hidden layer followed by a concept layer, where each concept receives input only from the features associated with that concept. The final LogRisk prediction is computed as a linear combination of these concepts. b. Concepts are extracted using either a trained XGBoost model or rules from the RuleKit model. For RuleKit, each rule’s unique features are grouped as a concept. For XGBoost, concepts are extracted at each tree depth. Each run of XGBoost (pruned to depth 5) generated between 540 and 770 concepts. However, having such a large number of concepts can hinder the model’s interpretability. To address this, we investigated reducing the number of concepts while maintaining performance. We first ranked the concepts from the trained ConSurv-XGB model based on their absolute weights. Then, we trained the model using only the top 5, 10, 50, 100, and 200 concepts. Supplemental Figure 2 illustrates the ranked concepts in descending order of absolute weight across different runs. As shown in Figure 2e , the validation set CI improves with increasing number of concepts upto 100 but plateaus (and even slightly declines) beyond that, when using all concepts. Based on this, we consider ConSurv-XGB with the top 100 concepts (hereafter referred to simply as ConSurv-XGB or C-XGB Top100) to be the best model in our analysis, achieving a median test set CI of 0.71. This approach reduces the number of concepts by approximately 80–88%, significantly simplifying the model while maintaining performance. Each run of RuleKit generated between 61 and 83 concepts, with an average of 25–30 features per concept. Similar to ConSurv-XGB, we evaluated models using the top 5, 10, 25, 50, and all available concepts. Supplemental Figure 3 illustrates the ranked concepts in descending order of absolute weight across different runs. As shown in Figure 2f , validation set performance does not substantially improve beyond the top 25 concepts (median CI 0.69). We use the top 25 concept model’s test set performance for comparison with existing models in Figure 2a (hereafter referred to simply as ConSurv-Rule or C-Rule Top25). In terms of median test performance, ConSurv models’ performance falls between XGBoost and DeepSurv. A Mann–Whitney U test confirms that ConSurv models, MLP, and DeepSurv perform significantly better than CPH. However, differences in performance among XGBoost, ConSurv models, MLP, and DeepSurv are not statistically significant (p-value > 0.05). The p-values for these comparisons are provided in Supplemental Table 2. We identified three key advantages of ConSurv. First, it achieves performance comparable to existing models. Second, it demonstrates stable performance across runs, as evidenced by a CI standard deviation of 0.06 for ConSurv-XGB, lower than the 0.1 observed for XGBoost, indicating consistency comparable to black-box models. Finally, as a gray-box model, ConSurv prioritizes interpretability through concepts as a core feature, ensuring that predictions rely exclusively on these concepts, thereby enhancing fidelity (or faithfulness). In the following sections, we explore the interpretability offered through concepts by ConSurv. 2.4 Concepts demonstrate significant stratification between high- and low-risk groups Interpretability of CPH Among the existing models, the CPH model has the lowest performance but provides inherent interpretability, with feature coefficients directly reflecting their importance. In Supplemental Figure 4, we present the top 40 features ranked by their absolute value of CPH coefficients. With these coefficients one can recognize the contribution made by every feature to the final prediction. These coefficients are faithful to the model as they are directly used for risk prediction. Interpretability of DeepSurv (SHAP limitations) In contrast, DeepSurv achieves the highest median performance among the evaluated models; however, interpreting its predictions requires post-hoc explanation techniques such as SHAP, which provides local explanations for individual patients. SHAP summary plots (Supplemental Figure 5), which aggregate feature importance across all patients, indicate that days to birth (age) is the most important feature driving Deep-Surv’s predictions. However, when we examined SHAP explanations for individual patients (Supplemental Figure 5), we observed substantial inconsistencies. For one patient, days to birth reduced the predicted risk (Supplemental Figure 6a), while for another, it increased the risk (Supplemental Figure 6b) and for a third patient (Supplemental Figure 6c ), the feature had no effect at all (SHAP score of 0.0). Notably, for the third patient, where SHAP suggested age was irrelevant, changing only the age resulted in changes in the predicted risk (ΔLogRisk = 22 and 15 for ages 25 and 80, respectively). This example highlights a key limitation of SHAP: even for the most important features, the explanations provided may not reliably reflect how the model truly uses them. Similar inconsistencies were observed for other top feature ajcc pathology t T 2, which showed conflicting effects on risk across different patients. These findings underscore broader concerns regarding the faithfulness of post-hoc XAI techniques in accurately representing the inner workings of the models they are intended to explain, and there is growing research highlighting these limitations [ 11 , 12 , 66 ]. Our framework addresses these limitations through interpretable concepts that group features together, offering insight into potential interactions. Furthermore, because the final prediction relies exclusively on these concepts, they present faith-fulness in the predictions. Additionally, it provides concept importance through the learned weights assigned to each concept. Interpretability of ConSurv Original concept bottleneck models [ 20 ] use predefined, human-interpretable concepts with ground truth for supervised concept training. However, availability of such concepts in many cases is rare. This is especially true in biomedical applications where AI is employed to gather insights from the data. In our case, clinicians not only want predictions but also features important for those predictions. In such scenarios, unsupervised extraction of concepts is a possible strategy where the extracted concepts are later ratified with domain knowledge based on the features they capture. The work by [ 67 ] demonstrates a time-series concept bottleneck model using this strategy. We use a similar approach where we first extract concepts by grouping features using either XGBoost or RuleKit (See Figure 4b , details in Section 4.2 ). Supplemental Tables 4–9 provides the top five concepts for ConSurv-XGB and ConSurv-Rule. We evaluate the effectiveness of each of the top five concepts from both models in distinguishing high- and low-risk groups using Kaplan-Meier survival plots. We obtained the output for all patients for each of the top five concepts using the ConSurv-XGB and ConSurv-Rule models. These outputs were then used to generate the KM plots presented in Figure 3 (c for ConSurv-XGB and d for ConSurv-Rule). For the ConSurv-XGB model, all concepts except concept 2 showed significant stratification (i.e., p < 0.05) between high- and low-risk patients. Similarly, for the ConSurv-Rule model, all concepts except concept 4 demonstrated significant differentiation (i.e., p < 0.05) between patient groups. This suggests that the top five most influential concepts in risk prediction from both models can effectively distinguish patients. In the next section, we analyze and interpret these top five concepts in the context of their biological relevance. 2.5 Concepts capture biological processes The concepts used in ConSurv are inherently interpretable and are extracted through two distinct approaches: ConSurv-XGB, where concepts are extracted using XGBoost, and ConSurv-Rule, where RuleKit is employed to define the concepts. First we look at the interpretation of top 5 concepts from ConSurv-XGB. Following which we look at interpretation of top 5 concepts extracted from ConSurv-Rule. Interpretation of top 5 ConSurv-XGB concepts Concept 1: The top-ranked concept in ConSurv-XGB, which exhibits the most significant stratification between high- and low-risk groups. This concept comprises the features: age, lymph, IGLV2-23, and MIR1244-4. Age is a key factor in survival outcomes, with women under 40 and over 80 showing poorer prognosis [ 35 ]. Lymph node pathology is also crucial, as one pathway for breast cancer metastasis is through the lymphatic system and lymph nodes [ 68 ]. This is especially significant for women under 40 with axillary lymph node-negative breast cancer shown to have poor prognosis [ 35 ]. Older patients with a high lymph node ratio are shown to have a threefold increased risk of breast cancer death. [ 69 , 70 ]. The IGLV family (here, IGLV2-23) is linked to older age at diagnosis and a distinct stromal microenvironment in breast cancer [ 71 ]. This concept connects age, lymph node pathology, and the IGLV gene. While we couldn’t find details of mechanism for MIR1244-2, there is increasing evidence of microRNAs’ role in breast cancer malignancy [ 72 ]. Concept 2: Comprises of NEAT1, TFF1, S100A6, RPS4Y1 and SETSIP genes. NEAT1 [ 73 ] promotes breast cancer, while TFF1 [ 74 ] and S100A6 [ 75 ] suppress it. All three interact with or affect estrogen expression [ 74 , 76 , 77 ]. RPS4Y1 [ 33 ], a male breast cancer marker, may serve as a proxy for estrogen status. SETSIP, while not directly linked, is associated with angiogenesis [ 78 ], potentially supporting tumor growth. This concept reflects estrogen response in breast cancer. Concept 3: Includes four genes—AQP8, PCSK1, HSPB1, and LGALS4—along with one clinical feature, lymph node pathology. AQP8 is highly expressed in basal and luminal B breast cancer types, which have low estrogen receptor (ER) levels [ 79 ]. PCSK1 is upregulated in breast cancer, promoting tumor progression, estrogen dependency, and anti-estrogen resistance in cell lines [ 80 ]. HSPB1 is linked to metastasis via epithelial-to-mesenchymal transition and is correlated with lymph node status and estrogen receptors [ 81 ]. LGALS4 high expression is a good prognostic factor for LN-negative patients [ 82 ]. This concept highlights estrogen dependency (via AQP8, PCSK1, and HSPB1) and lymph node involvement (via lymph node pathology, HSPB1, and LGALS4). Concept 4: Includes features COL11A2, TAGLN2, LPL, CPB1 and lymph node pathology. COL11A2 [ 83 ] and TAGLN2 [ 84 , 85 ] are linked to lymph node pathology. LPL promotes tumor growth by altering the microenvironment through lipid hydrolysis [ 86 ], while all three are potential therapeutic targets. CPB1 down-regulates tumor suppressors like SFRP1 and OS9 [ 87 ] and is up-regulated in lymph node-positive patients [ 88 ]. This concept highlights features associated with lymph node involvement and potential therapeutic targets. Concept 5: Comprises of HBD, CD24, IGHV5-51, RPS4Y1 and age. The loss of the immunomodulatory HBD [ 89 ] and overexpression of CD24 [ 27 , 90 ] promote metastasis by enabling escape from cell death. The IGHV family (e.g., IGHV5-51) is linked to vacuolation and degeneration in mouse breast tumorigenesis [ 91 ]. Together, these features suggest immune dysregulation, promoting cell death avoidance and tumor proliferation. However, the role of age and RPS4Y1 in conjunction with these genes requires further exploration. We find ConSurv-XGB concepts are smaller in size (only 1–5 features per concept), with each capturing a biologically meaningful property. Further, understanding relationships between features within each concept often requires domain knowledge as seen for the five concepts discussed above. Interpretation of top 5 ConSurv-Rule concepts The concepts extracted from RuleKit tend to be larger (25–30 features per concept), leading us to hypothesize that they capture broader biological processes. To investigate this, we performed Overrepresentation Analysis (ORA) (See Section 4.8 ) to identify associated processes, with results listed in Table 3 . We used the Hallmark [ 92 ] and C6 gene set databases in MSigD to determine relevant processes. Hallmark gene sets represent well-defined biological states with coherent expression patterns and were our primary reference. When no significant process were identified in Hallmark, we turned to C6. We found that all top 5 concepts in RuleKit can be associated with broader biological processes. Concept 1 primarily captures signaling pathways but also identifies hypoxia, a key feature of the tumor microenvironment that promotes angiogenesis in breast cancer [ 93 ]. Concept 2 is associated with immune-related processes, as indicated by the Hallmark allograft rejection and interferon-gamma response gene sets, both critical in breast cancer progression [ 94 , 95 ]. Concept 4, which did not significantly differentiate between high- and low-risk patients ( Figure 3d ), lacks Hallmark features but captures the C6 oncogenic signature CTIP DN.V1 UP, linked to early-onset breast cancer [ 96 ]. View this table: View inline View popup Download powerpoint Table 3 Biological processes of top 5 ConSurv-Rule concepts RuleKit concepts exhibit two key characteristics: 1. Certain gene sets, such as estrogen response and angiogenesis, appears repeatedly across concepts (both processes are crucial for breast cancer survival [ 93 , 97 ]); and 2. they often encompass multiple biological processes suggesting that RuleKit identifies high-level biological pattern. Both ConSurv models provide biologically meaningful concepts. While ConSurv-XGB requires domain knowledge to interpret feature associations within each concept, ConSurv-Rule concepts can be more easily linked to well-established biological processes. Unlike conventional methods that rely on individual feature weights or costly and often unreliable post-hoc explanations, ConSurv models offer direct biological insights through their inherently interpretable design. In the next section, we analyze quantitative properties of the concepts from both models. 2.6 Quantitative analysis of the interpretable concepts In this section, we explore the quantitative aspects of the interpretable concepts in our framework. First, we conduct an exploratory analysis of the concepts ( Figure 5 ) and then look at the stability of the concepts and robustness of the concept extraction. Download figure Open in new tab Fig. 5 Exploratory analysis of XGBoost, RuleKit, and extracted concepts across ten runs. a. Trained XGBoost trees covered 35–50% of the 1,056 features in the dataset and generated 330–390 trees ( b ). c. The number of unique concepts extracted from trained trees drops significantly from 2,900–3,600 at depth d = 10 to 4–8 at d = 1. d. For trees with pruning depth d = 5 the number of unique concepts ranges from 540–770. e. RuleKit-generated rules cover 43–53% of the 1,056 available features. f. RuleKit produced 61–83 rules, with each rule containing an average of 25–30 conditions ( g , Each box shows the 1st and 3rd quartiles, with the middle line indicating the median. Whiskers show the minimum and maximum, and outliers are shown as circles). h. Between 61–83 unique concepts were extracted, with one concept per rule. Figure 5a shows that for each run, XGBoost utilizes between 35–50% of the total 1,056 features in the dataset. Each XGBoost run produces between 330–390 trees ( Figure 5b ). The depth of the tree dictates the number of unique concepts that can be extracted. As seen in Figure 5c , the number of concepts drops drastically from depth 10 to depth 1. Given that pruning depth 5 achieves performance comparable to depth 10, we extract concepts at this pruning depth, yielding approximately 540–770 unique concepts ( Figure 5d ). Since the top 100 concepts contribute most to the predictive capability of ConSurv-XGB, we analyzed their origin. Specifically, we examined whether these concepts emerge from early or late iterations of XGBoost training. Figure 6a demonstrates that most highly ranked concepts emerge during the initial and final iterations, a trend consistent across runs. Figure 6b shows that larger-sized concepts (based on number of features within each concept) are more frequent in the top 100 concepts, with concepts of size 5 being predominant. Download figure Open in new tab Fig. 6 Quantitative analysis of extracted concepts across ten runs. a. The distribution of the top 100 concepts from ConSurv-XGB shows that most originate from XGBoost trees generated in early or late training iterations. b. Among the top 100 concepts from ConSurv-XGB, the majority comprise 5 features, with gradually fewer concepts having 4, 3, or 2 features. c. For ConSurv-Rule, the top 25 concepts typically contain 20–40 features per concept. d. To assess the stability of concept rankings, the rankings of concepts in ConSurv-XGB and ConSurv-Rule were compared to their original rankings in the ConSurv-XGB-all and ConSurv-Rule-all models, respectively, using Kendall’s Tau. Kendall’s Tau values for ConSurv-XGB ranged from -0.8 to 1 0 7 .32, indicating minimal similarity, while ConSurv-Rule concepts showed higher similarity with Kendall’s Tau values between 0.13 and 0.57. Each box shows the 1st and 3rd quartiles, with the middle line indicating the median. Whiskers show the minimum and maximum, and outliers are shown as circles. e. Kendall’s Tau values across ten runs showing change in the rankings of the top 100 ConSurv-XGB concepts between the original ranking and the new ranking. f. Kendall’s Tau values across ten runs reflecting changes in the rankings of the top 25 ConSurv-Rule concepts between the original ranking and the new ranking. g. The distribution of Jaccard Similarity Index between concepts captured across run pairs shows very limited overlap for both ConSurv-XGB and ConSurv-Rule concepts. h. Similarly, the distribution of Cosine Similarity Index indicates limited overlap between concepts across runs for both methods. However, ConSurv-Rule concepts exhibit greater overlap than ConSurv-XGB concepts on both similarity indices. In each run of RuleKit, between 43–53% of the 1,056 dataset features were utilized that produced 61–83 rules, with each rule containing an average of 25–30 conditions ( Figure 5e–g ). Unlike XGBoost, each unique rule roughly translates to one unique concept, resulting in 61–83 unique concepts per run ( Figure 5h ). Stability Next, we examined whether training the model using only the top-ranking concepts impacts their rankings. A model that preserves most of the concept rankings would suggest that these concepts and their importance remain stable. We compared the ranking of 100 concepts in ConSurv-XGB model with original rankings from ConSurv-XGB with all concepts. We used Kendall’s Tau to assess the similarity of rankings and found that it ranged between 0.08 to 0.31 (plotted in Figure 6d,e ) which indicates there is little to no correlations between the two rankings. Similarly, we compared the ranking of the 25 concepts in ConSurv-Rule model with original rankings from ConSurv-Rule with all concepts. We found that the Kendall’s Tau value ranged between 0.02 to 0.56 (plotted in Figure 6d , f) which indicates there is little correlations (better than ConSurv-XGB) between the two rankings. These results, which show a high mismatch between the original and retrained rankings for both ConSurv-XGB and ConSurv-Rule concepts, indicate that the concepts and their importance are not stable (i.e. the concepts are not stable). Robustness The concept extraction process can be considered robust if the concepts extracted across multiple runs exhibit a high degree of similarity. We evaluated the similarity of concepts across runs in both ConSurv models. Concepts represent sets of features, with sizes ranging from 1–5 features for ConSurv-XGB and averaging 25–30 features for ConSurv-Rule. To measure similarity, we used the commonly applied Jaccard Similarity Index (SI). However, as Jaccard SI is influenced by set size, we also used the Cosine SI, which focuses on overall similarity between sets. Figures 6g,h shows the distributions of Jaccard and Cosine SI across the 45 run pairs (10 runs yields 45 pairs), comparing every concept from one run to every concept in another run. For each pair, comparing 100 concepts in ConSurv-XGB results in 4,950 comparisons and 300 comparisons for the 25 concepts in ConSurv-Rule. ConSurv-Rule concepts were noticeably more consistent across runs compared to those from XGBoost, as indicated by higher Jaccard and Cosine similarity indices. This consistency is further supported by two observations: First, a trained RuleKit model utilizes more features per run than a trained XGBoost model (43%–53% vs. 35%–50%, respectively), allowing for more potential overlap across runs. Second, 44 features were consistently present across RuleKit runs, compared to only 14 features in XGBoost runs. However, despite this relative difference, the overall similarity indices for both remained quite low, indicating limited overlap in the concepts across runs. In summary, ConSurv-XGB concepts are noticeably smaller (1–5 features per concept) than ConSurv-Rule concepts (25–30 features per concept). Similarity in concepts across runs is limited for both models. However, in contrast to ConSurv-XGB, ConSurv-Rule concepts are more consistent across runs. Additionally, concepts from both model appear not stable, as their rankings change after retraining. 3 Discussion and Conclusion Our study aims to improve interpretability in survival analysis while balancing the trade-off with performance. Through an evaluation of existing models and our proposed ConSurv framework, we demonstrated that incorporating interpretable concepts enhances model transparency without compromising performance. Trade-off between interpretability and accuracy The CPH, a fully white-box approach, offers inherent interpretability via feature coefficients but suffers from the lowest median performance. Conversely, black-box models such as DeepSurv achieve the highest median performance but require unreliable posthoc explanations. ConSurv, a gray-box model, successfully balances these concerns. By leveraging interpretable concepts from ConSurv-XGB and ConSurv-Rule, we achieve performance similar to high performing models while providing human-interpretable insights. Importantly, ConSurv-XGB demonstrates lower CI variance than XGBoost, indicating increased stability of the performance. Concepts enable significant stratification among patients A key advantage of our framework lies in its ability to extract and leverage meaningful concepts in survival prediction. Unlike post-hoc methods that provide individual feature importance without capturing interactions, ConSurv integrates concepts directly into the model. This allows a more interpretable and faithful predictions with biologically relevant representation of survival risk factors. This is evident in our analysis of concepts and their biological significance. Additionally, Kaplan-Meier survival plots ( Figure 3 ) show that top ConSurv-XGB and ConSurv-Rule provides significant survival stratification between high- and low-risk patients. Biological relevance of concepts To assess the interpretability and biological relevance of our learned concepts, we performed Over-Representation Analysis (ORA) to visualize gene interaction networks. The extracted concepts align with known biological processes relevant to cancer survival, such as hypoxia, angiogenesis, etc. We apply very stringent condition that only Hallmark and C6 databases were used (with preference for Hallmark) to ensure only most relevant processes are obtained. However, this can be left at the discretion of the end user to set the criteria. Stability and reproducibility of extracted concepts An important consideration in modeling approach is the consistency of learned concepts across different runs. Our analysis of concept robustness ( Figure 6 ) reveals that ConSurv-Rule concepts exhibit greater consistency across runs compared to ConSurv-XGB concepts. However, despite this advantage, overall concept similarity remains low, indicating variability in the specific features grouped within each concept. Additionally, the stability of concepts for both models was low. More research is needed to develop strategies for extracting concepts more effectively, ensuring greater consistency across different runs. Reducing variability in concept formation across random data splits will be crucial for enhancing both the stability and robustness of learned concepts, thereby allowing the ConSurv framework to be more readily used. Limitations Despite the advantages of ConSurv, some limitations warrant further investigation. First, the variability in concept extraction across runs suggests a need for more robust feature selection or concept aggregation strategies. One potential approach is to integrate ensemble learning techniques that uniformize concept extraction across multiple runs. Additionally, while ConSurv-Rule concepts are relatively robust, they include many features per concept which capture multiple processes, and exhibit redundancy across concepts. Future work could explore hybrid approaches that optimize both concept size and redundancy. With only 1–5 features per concept, ConSurv-XGB concepts are small, but interpreting their biological meaning requires extensive domain knowledge. Conclusion In conclusion, our study introduces ConSurv, a gray-box survival analysis framework that balances interpretability and accuracy by leveraging concept-based learning. Through comparisons with existing models, we demonstrate that ConSurv achieves high predictive performance while maintaining transparency in its decision-making with biologically relevant concepts. This unique ability positions ConSurv as a valuable tool for clinical decision support, enabling both accurate risk prediction and actionable insights. Future research will focus on robust concept extraction and exploring broader applications in precision medicine. 4 Methods 4.1 Contrast Set Mining Association rule mining is used in problems with multi-dimensional datasets to identify relations between features [ 98 ]. These associations takes a form of “if-then” rules. Contrast set mining [ 99 – 101 ] is a special type of association rule mining that identifies key differences between groups within a dataset. It specifically aims to find attribute-value (i.e. feature and its value) combinations, known as contrast sets, that vary significantly in frequency or magnitude across these groups. For a dataset D containing two classes—positive ( P ) and negative ( N )—a contrast set S is defined as a combination of attribute-value pairs (e.g., A 1 = v 1 ∧ A 2 = v 2, where v 1 and v 2 are specific values of the attributes). The support of S within positive or negative class, represented as support( S, P ) or support( S, N ), is the proportion of instances in class that satisfy S . A contrast set S is deemed significant if there is at least one pair of classes, P and N , for which the support difference support( S, P ) support( S, N ) exceeds a given threshold δ and is statistically significant. For multiclass problems, it employs a one-vs-all strategy to generate rules specific to each class. In this work, we leverage the RuleKit [ 101 , 102 ] package to extract contrast sets. 4.2 ConSurv ConSurv is an interpretable concept bottleneck model that relies solely on concepts for its final predictions. Its interpretability stems from the use of interpretable concepts, while maintaining complete faithfulness by basing predictions exclusively on a linear combination of these concepts, as shown below: where C 1 .. C m are m interpretable concepts and w 1 .. w m their respective weights. ConSurv has two layers: one hidden layer and one m -sized concept layer. The architecture of ConSurv is designed such that each concept receives input only from the features grouped for that concept (see Figure 4a ). The concept layer output is normalized to values between 0–1 to ensure that the impact of each concept is influenced only by the static weight w for that concept. To extract concepts, we rely on trained XGBoost trees and RuleKit rules, as illustrated in Figure 4b . Contrast set rule are a form of association rule learning which find patterns among the features that aim to distinguish between different groups. As illustrated in Figure 4 , for each learned rule, the features are collected and grouped into concepts. On the other hands, tree based algorithms such as XGBoost are good at identifying features which are non-linearly related to each other and output. For XGBoost, concepts are extracted from each depth level of trained tree. In the illustration of Figure 4 , the tree has two levels nodes, Node Pathology feature at root and Beta hemoglobin feature at level 1. There the extracted concept from root level is just the Node Pathology feature (Concept 1) and from level 1 is Node Pathology and Beta − hemoglobin feature (Concept 2). 4.3 Hyperparameter tuning and model training To ensure optimal parameters for model training, we conducted hyperparameter tuning for CPH, XGBoost, RuleKit, DeepSurv, MLP, and ConSurv. Ten runs were performed, with data splits generated using different random seeds (run IDs and their corresponding seeds are listed in Supplemental Table 1). For each run, the data was split into 70% training, 15% validation, and 15% test sets. Hyperparameter tuning was performed on the training and validation sets across all ten runs. The test sets were used to evaluate performance of the tuned models. The categorical features were one hot encoded and continuous features were min-max normalized. The parameter combinations evaluated for each model are detailed in Supplemental Table 3, with the best-performing values highlighted in the last column. For ConSurv, the optimal parameters from MLP were adopted, with only the learning rate tuned separately. For XGBoost, DeepSurv, MLP, and ConSurv models, early stopping was applied with a patience of 50 rounds. Adaptive Moment Estimation (Adam) [ 103 ] was used for the gradient descent algorithm. The same training set used for training the ConSurv model was also used to train the corresponding XGBoost and RuleKit models for concept extraction. 4.4 Concordance Index Concordance Index or CI [ 104 , 105 ] is a used widely metric for evaluating performance of survival models. It is calculated as: Where, N c is number of concordant pairs such that for patients i and j , NN ( x i ) T j and e j = 1. N d is number of discordant pairs such that for patients i and j , NN ( x i ) > NN ( x j ), T i > T j and e j = 1. NN (.) is predicted LogRisk, T is time-to-event and e is event variable. 4.5 Set similarity indices Jaccard similarity index for set A and B is given by Equation (3) . It ranges from 0–1 with 0 indicating no overlap or similarity and 1 indicating perfect overlap or similarity. Cosine similarity index for set A and B is given by Equation (4) . First the features from the two concepts being compared are one hot encoded into vectors a and b and then the Cosine similarity is calculated. The range of Cosine similarity is from -1–1, however, in our case of binary vectors, it ranges only from 0–1. 4.6 Objective function: Cox-loss Deep learning models for survival analysis such as AutoSurv and DeepSurv often employ Cox-loss for training the models. We used the same loss function in training ConSurv and MLP models. The loss function is as follows (from [ 4 , 8 ]): where θ are model parameters, NN ( xi ) (LogRisk) is model output for input xi , λ is l 2 regularization parameter, e is the event variable, Ne =1 are number of patients with an observable event and R ( Ti ) are patients still at risk of failure at time t . 4.7 Gene Set Variation Analysis Gene Set Variation Analysis (GSVA) is a non-parametric, unsupervised method that estimates the variation of pathway activity (or gene set enrichment) across all samples in a transcriptomic dataset [ 21 ]. Rather than examining individual gene expression, GSVA computes an enrichment score per sample for each predefined gene set, effectively transforming the gene-expression matrix into a sample-by-gene-set score matrix. In our work, we used GSVA to aggregate the genes associated with each concept (rule) into a single “concept activity score” per patient. The aggregated GSVA scores enable stratification of patients into survival groups, which we visualize using Kaplan-Meier plots. 4.8 Over-Representation Analysis We performed Over-Representation Analysis (ORA) to identify biological pathways or gene sets enriched among top genes from each ConSurv rule. Using the Molecular Signatures Database (MSigDB) collections [ 92 ] (H - Hallmark gene sets and C6 - oncogenic signatures) and all measured genes as a background set, we applied Fisher’s exact test to evaluate whether a given genes in a concept appeared more frequently in any particular gene set than expected by chance. We corrected for multiple testing using the Benjamini–Hochberg method (FDR < 0.05). These enriched pathways provide functional context for each rule, underscoring potential mechanisms underlying risk stratification. Data Availability The breast cancer dataset used in this study is available at https://anonymous.4open.science/r/ConSurv-13BE . Code availability All models were implemented using Python-3.9. Following packages were used: xgboost (for XGBoost model), RuleKit (for Contrast Set Rule mining), lifelines (for CPH and concordance index), TorchSurv [ 106 ] (for loss function), PyTorch (For deep learning models - DeepSurv, MLP and ConSurv). The code and complete list of packages used is available at: https://anonymous.4open.science/r/ConSurv-13BE Author contributions P.B ., A.V., and A.R., conceptualized and designed experiments. P.B. and T.W. conducted the experiments and generated results. A.R. supervised the project. P.B., T.W., A.V., A.R. contributed to writing the manuscript. Acknowledgment This work and authors P.B., A.R., are supported by the European Union’s Horizon 2020 research and innovation programme under grant agreement no. 101017453. A.V. is supported by the “UNREAL: Unified Reasoning Layer for Trustworthy ML” project (EP/Y023838/1) selected by the ERC and funded by UKRI EPSRC. This work has made use of the resources provided by the Edinburgh Compute and Data Facility (ECDF) ( http://www.ecdf.ed.ac.uk/ ). Funding European Union https://ror.org/019w4f821 101017453 UK Research and Innovation https://ror.org/001aqnf71 EP/Y023838/1 Footnotes Contributing authors: tongjie.wang{at}ed.ac.uk ; avergari{at}ed.ac.uk ; References [1]. ↵ Clark , T. G. , Bradburn , M. J. , Love , S. B. & Altman , D. G . Survival analysis part i: basic concepts and first analyses . British journal of cancer 89 , 232 – 238 ( 2003 ). OpenUrl CrossRef PubMed Web of Science [2]. ↵ Bradburn , M. J. , Clark , T. G. , Love , S. B. & Altman , D. G . Survival analysis part ii: multivariate data analysis–an introduction to concepts and methods . British journal of cancer 89 , 431 – 436 ( 2003 ). OpenUrl CrossRef PubMed Web of Science [3]. ↵ Cox , D. R . Regression models and life-tables. breakthroughs in statistics . Stat. Soc 372 , 527 – 541 ( 1992 ). OpenUrl [4]. ↵ Katzman , J. L. et al. Deepsurv: personalized treatment recommender system using a cox proportional hazards deep neural network . BMC medical research methodology 18 , 1 – 12 ( 2018 ). OpenUrl CrossRef PubMed [5]. ↵ Kovalev , M. S. , Utkin , L. V. & Kasimov , E. M . Survlime: A method for explaining machine learning survival models . Knowledge-Based Systems 203 , 106164 ( 2020 ). [6]. ↵ Ribeiro , M. T. , Singh , S. & Guestrin , C . “why should i trust you” explaining the predictions of any classifier , 1135 – 1144 ( 2016 ). [7]. ↵ Krzyziński , M. , Spytek , M. , Baniecki , H. & Biecek , P . Survshap (t): timedependent explanations of machine learning survival models . Knowledge-Based Systems 262 , 110234 ( 2023 ). [8]. ↵ Jiang , L. et al. Autosurv: interpretable deep learning framework for cancer survival analysis incorporating clinical and multi-omics data . NPJ precision oncology 8 , 4 ( 2024 ). 9. ↵ Lundberg , S. M. & Lee , S.-I . Guyon , I. , et al. (eds) A unified approach to interpreting model predictions . (eds Guyon , I. et al. ) Advances in Neural Information Processing Systems , Vol. 30 ( Curran Associates, Inc ., 2017 ). URL https://proceedings.neurips.cc/paperfiles/paper/2017/file/8a20a8621978632d76c43dfd28b67767-Paper.pdf . [10]. ↵ Huang , X. & Marques-Silva , J . On the failings of shapley values for explainability . International Journal of Approximate Reasoning 109112 ( 2024 ). [11]. ↵ Rudin , C . Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead . Nature machine intelligence 1 , 206 – 215 ( 2019 ). OpenUrl CrossRef PubMed [12]. ↵ Bilodeau , B. , Jaques , N. , Koh , P. W. & Kim , B . Impossibility theorems for feature attribution . Proceedings of the National Academy of Sciences 121 , e2304406120 ( 2024 ). OpenUrl CrossRef PubMed [13]. ↵ Cho , H. J. , Shu , M. , Bekiranov , S. , Zang , C. & Zhang , A . Interpretable metalearning of multi-omics data for survival analysis and pathway enrichment . Bioinformatics 39 , btad113 ( 2023 ). 14. ↵ Shrikumar , A. , Greenside , P . & Kundaje , A . Learning important features through propagating activation differences , 3145 – 3153 (PMlR, 2017 ). [15]. ↵ Yousefi , S. et al. Predicting clinical outcomes from large scale cancer genomic profiles with deep survival models . Scientific reports 7 , 1 – 11 ( 2017 ). OpenUrl CrossRef PubMed [16]. ↵ Ancona , M. , Ceolini , E. , Öztireli , C. & Gross , M . Gradient-based attribution methods . Explainable AI: Interpreting, explaining and visualizing deep learning 169 – 191 ( 2019 ). [17]. ↵ Ching , T. , Zhu , X. & Garmire , L. X . Cox-nnet: an artificial neural network method for prognosis prediction of high-throughput omics data . PLoS computational biology 14 , e1006076 ( 2018 ). OpenUrl CrossRef [18]. ↵ Chen , T. & Guestrin , C . Xgboost: A scalable tree boosting system , 785 – 794 ( 2016 ). [19]. ↵ Barnwal , A. , Cho , H. & Hocking , T . Survival regression with accelerated failure time model in xgboost . Journal of Computational and Graphical Statistics 31 , 1292 – 1302 ( 2022 ). OpenUrl CrossRef 20. ↵ Koh , P. W. et al. Concept bottleneck models , 5338 – 5348 (PMLR, 2020). [21]. ↵ Hänzelmann , S. , Castelo , R. & Guinney , J. Gsva: gene set variation analysis for microarray and rna-seq data . BMC bioinformatics 14 , 1 – 15 ( 2013 ). OpenUrl CrossRef PubMed [22]. Woods , N. T. et al. Charting the landscape of tandem brct domain–mediated protein interactions . Science signaling 5 , rs6–rs6 ( 2012 ). [23]. Chan , S.-H. , Tseng , H.-J. & Wang , L.-H . Cd24a knockout transforms the tumor microenvironment from cold to hot by promoting tumor-killing immune cell infiltration in a murine triple-negative breast cancer model . bioRxiv 2024 – 07 ( 2024 ). [24]. Lee , S. et al. Carbonic anhydrases reduce the acidity of the tumor microenvironment, promote immune infiltration, decelerate tumor growth, and improve survival in erbb2/her2-enriched breast cancer . Breast Cancer Research 25 , 46 ( 2023 ). [25]. Adnane , J. et al. Bek and flg, two receptors to members of the fgf family, are amplified in subsets of human breast cancers . Oncogene 6 , 659 – 663 ( 1991 ). OpenUrl PubMed Web of Science [26]. Bechmann , M. B. , Brydholm , A. V. , Codony , V. L. , Kim , J. & Villadsen , R . Heterogeneity of ceacam5 in breast cancer . Oncotarget 11 , 3886 ( 2020 ). OpenUrl CrossRef PubMed [27]. ↵ Kristiansen , G. et al. Cd24 expression is a new prognostic marker in breast cancer . Clinical cancer research 9 , 4906 – 4913 ( 2003 ). OpenUrl Abstract / FREE Full Text [28]. Bera , A. et al. Functional role of vitronectin in breast cancer . PLoS One 15 , e0242141 ( 2020 ). OpenUrl CrossRef PubMed [29]. Booy , E. P. , McRae , E. K. , Koul , A. , Lin , F. & McKenna , S. A . The long noncoding rna bc200 (bcyrn1) is critical for cancer cell survival and proliferation . Molecular cancer 16 , 1 – 15 ( 2017 ). OpenUrl PubMed [30]. Cui , Y. , Jiao , Y. , Wang , K. , He , M. & Yang , Z . A new prognostic factor of breast cancer: High carboxyl ester lipase expression related to poor survival . Cancer genetics 239 , 54 – 61 ( 2019 ). OpenUrl CrossRef PubMed [31]. Li , J. et al. Dusp1 promoter methylation in peripheral blood leukocyte is associated with triple-negative breast cancer risk . Scientific reports 7 , 43011 ( 2017 ). [32]. Candas , D. et al. Mitochondrial mkp1 is a target for therapy-resistant her2-positive breast cancer cells . Cancer research 74 , 7498 – 7509 ( 2014 ). OpenUrl Abstract / FREE Full Text [33]. ↵ Li , Y. et al. Male breast cancer differs from female breast cancer in molecular features that affect prognoses and drug responses . Translational Oncology 45 , 101980 ( 2024 ). [34]. Fernandes , J. O. et al. Differences in breast cancer survival and stage by age in off-target screening groups: a population-based retrospective study . AJOG Global Reports 3 , 100208 ( 2023 ). [35]. ↵ Brandt , J. , Garne , J. P. , Tengrup , I. & Manjer , J . Age at diagnosis in relation to survival following breast cancer: a cohort study . World journal of surgical oncology 13 , 1 – 11 ( 2015 ). OpenUrl CrossRef PubMed [36]. Rakha , E. A. , Tse , G. M. & Quinn , C. M . An update on the pathological classification of breast cancer . Histopathology 82 , 5 – 16 ( 2023 ). OpenUrl CrossRef PubMed [37]. Nabet , B. Y. et al. Exosome rna unshielding couples stromal activation to pattern recognition receptor signaling in cancer . Cell 170 , 352 – 366 ( 2017 ). OpenUrl CrossRef PubMed [38]. Lv , X. et al. Identification of potential key genes and pathways predicting pathogenesis and prognosis for triple-negative breast cancer . Cancer cell international 19 , 1 – 12 ( 2019 ). OpenUrl CrossRef PubMed [39]. Sauer , N. et al. Prognostic role of prolactin-induced protein (pip) in breast cancer . Cells 12 , 2252 ( 2023 ). OpenUrl CrossRef [40]. Morgat , C. et al. Expression of neurotensin receptor-1 (nts 1) in primary breast tumors, cellular distribution, and association with clinical and biological factors . Breast Cancer Research and Treatment 190 , 403 – 413 ( 2021 ). OpenUrl CrossRef PubMed [41]. Kimura , N. , Yoshida , R. , Shiraishi , S.-i. , Pilichowska , M. & Ohuchi , N . Chromogranin a and chromogranin b in noninvasive and invasive breast carcinoma . Endocrine pathology 13 , 117 – 122 ( 2002 ). OpenUrl CrossRef PubMed Web of Science [42]. Emberley , E. D. , Murphy , L. C. & Watson , P. H . S100a7 and the progression of breast cancer . Breast Cancer Research 6 , 1 – 7 ( 2004 ). OpenUrl CrossRef PubMed Web of Science [43]. Wang , D. et al. Clinical significance of elevated s100a8 expression in breast cancer patients . Frontiers in oncology 8 , 496 ( 2018 ). [44]. Sharma , S. et al. Putative interactions between transthyretin and endosulfan ii and its relevance in breast cancer . International Journal of Biological Macromolecules 235 , 123670 ( 2023 ). [45]. Li , Q. et al. The oncoprotein muc1 facilitates breast cancer progression by promoting pink1-dependent mitophagy via atad3a destabilization . Cell Death & Disease 13 , 899 ( 2022 ). [46]. Vikramdeo , K. S. et al. Profiling mitochondrial dna mutations in tumors and circulating extracellular vesicles of triple-negative breast cancer patients for potential biomarker development . FASEB BioAdvances 5 , 412 ( 2023 ). [47]. Zhang , X. et al. The emerging role of snornas in human disease . Genes & Diseases 10 , 2064 – 2081 ( 2023 ). OpenUrl CrossRef PubMed [48]. Gu , X.-L. et al. Expression of cxcl14 and its anticancer role in breast cancer . Breast cancer research and treatment 135 , 725 – 735 ( 2012 ). OpenUrl CrossRef PubMed [49]. Cheriyath , V. et al. G1p3 (ifi6), a mitochondrial localised antiapoptotic protein, promotes metastatic potential of breast cancer cells through mtros . British journal of cancer 119 , 52 – 64 ( 2018 ). OpenUrl CrossRef PubMed [50]. Tong , W. L. , Tu , Y. N. , Samy , M. D. , Sexton , W. J. & Blanck , G . Identification of immunoglobulin v (d) j recombinations in solid tumor specimen exome files: Evidence for high level b-cell infiltrates in breast cancer . Human Vaccines & Immunotherapeutics 13 , 501 – 506 ( 2017 ). OpenUrl CrossRef PubMed [51]. Leung , T. et al. Cytochrome p450 2e1 (cyp2e1) regulates the response to oxidative stress and migration of breast cancer cells . Breast Cancer Research 15 , 1 – 12 ( 2013 ). OpenUrl CrossRef Web of Science [52]. Kannan , A. et al. Cancer testis antigen promotes triple negative breast cancer metastasis and is traceable in the circulating extracellular vesicles . Scientific reports 9 , 11632 ( 2019 ). [53]. Naderi , A. et al. Bex2 is overexpressed in a subset of primary breast cancers and mediates nerve growth factor/nuclear factor- κ b inhibition of apoptosis in breast cancer cell lines . Cancer research 67 , 6725 – 6736 ( 2007 ). OpenUrl Abstract / FREE Full Text [54]. Morita , T. & Hayashi , K . Tumor progression is mediated by thymosin- β 4 through a tgf β /mrtf signaling axis . Molecular Cancer Research 16 , 880 – 893 ( 2018 ). OpenUrl Abstract / FREE Full Text [55]. Wang , R. et al. Btg2 as a tumor target for the treatment of luminal a breast cancer . Experimental and Therapeutic Medicine 23 , 1 – 11 ( 2022 ). OpenUrl [56]. Xie , T. et al. Peg10 as an oncogene: expression regulatory mechanisms and role in tumor progression . Cancer cell international 18 , 1 – 10 ( 2018 ). OpenUrl CrossRef PubMed [57]. Sanchez-Lopez , J. M. et al. Integrative analysis of transcriptional profile reveals linc00052 as a suppressor of breast cancer cell migration . Cancer Biomarkers 30 , 365 – 379 ( 2021 ). OpenUrl CrossRef PubMed [58]. Link , T. et al. Exploratory investigation of psca-protein expression in primary breast cancer patients reveals a link to her2/neu overexpression . Oncotarget 8 , 54592 ( 2017 ). [59]. Ganguly , D. et al. Pleiotrophin drives a prometastatic immune niche in breast cancer . Journal of Experimental Medicine 220 , e20220610 ( 2023 ). OpenUrl CrossRef PubMed [60]. Wang , Y. et al. Low expression of crisp3 predicts a favorable prognosis in patients with mammary carcinoma . Journal of Cellular Physiology 234 , 13629 – 13638 ( 2019 ). OpenUrl CrossRef PubMed [61]. Zou , J. et al. Identification of c4bpa as biomarker associated with immune infiltration and prognosis in breast cancer . Translational Cancer Research 13 , 25 ( 2024 ). [62]. Kim , J. et al. Long noncoding rna malat1 suppresses breast cancer metastasis . Nature genetics 50 , 1705 – 1715 ( 2018 ). OpenUrl CrossRef PubMed [63]. Arun , G. & Spector , D. L . Malat1 long non-coding rna and breast cancer . RNA biology 16 , 860 – 863 ( 2019 ). OpenUrl CrossRef PubMed [64]. Yu , W. et al. Increased expression of cyp4z1 promotes tumor angiogenesis and growth in human breast cancer . Toxicology and applied pharmacology 264 , 73 – 83 ( 2012 ). OpenUrl CrossRef PubMed [65]. Brauer , H. A. et al. Dermcidin expression is associated with disease progression and survival among breast cancer patients . Breast cancer research and treatment 144 , 299 – 306 ( 2014 ). OpenUrl CrossRef PubMed [66]. ↵ Slack , D. , Hilgard , S. , Jia , E. , Singh , S. & Lakkaraju , H . Fooling lime and shap: Adversarial attacks on post hoc explanation methods , 180 – 186 ( 2020 ). 67. ↵ Wu , C. , Parbhoo , S. , Havasi , M. & Doshi-Velez , F . Learning optimal summaries of clinical time-series with concept bottleneck models , Vol. 182 of Proceedings of Machine Learning Research , 648 – 672 ( PMLR , 2022 ). URL https://proceedings.mlr.press/v182/wu22a.html . OpenUrl [68]. ↵ Nathanson , S. D. et al. Mechanisms of breast cancer metastasis . Clinical & experimental metastasis 39 , 117 – 137 ( 2022 ). OpenUrl CrossRef PubMed [69]. ↵ Wildiers , H. et al. Relationship between age and axillary lymph node involvement in women with breast cancer . Journal of clinical oncology 27 , 2931 – 2937 ( 2009 ). OpenUrl Abstract / FREE Full Text [70]. ↵ Vinh-Hung , V. et al. Age and axillary lymph node ratio in postmenopausal women with t1-t2 node positive breast cancer . The oncologist 15 , 1050 – 1062 ( 2010 ). OpenUrl Abstract / FREE Full Text [71]. ↵ Brouwers , B. et al. The footprint of the ageing stroma in older patients with breast cancer . Breast Cancer Research 19 , 1 – 14 ( 2017 ). OpenUrl CrossRef PubMed [72]. ↵ Muñoz , J. P. , Pérez-Moreno , P. , Pérez , Y. & Calaf , G. M . The role of micrornas in breast cancer and the challenges of their clinical application . Diagnostics 13 , 3072 ( 2023 ). OpenUrl CrossRef PubMed [73]. ↵ Shin , V. Y. et al. Long non-coding rna neat1 confers oncogenic role in triple-negative breast cancer through modulating chemoresistance and cancer stemness . Cell death & disease 10 , 270 ( 2019 ). [74]. ↵ Yi , J. et al. Trefoil factor 1 (tff1) is a potential prognostic biomarker with functional significance in breast cancers . Biomedicine & Pharmacotherapy 124 , 109827 ( 2020 ). [75]. ↵ Qi , M. et al. S100a6 inhibits mdm2 to suppress breast cancer growth and enhance sensitivity to chemotherapy . Breast Cancer Research 25 , 55 ( 2023 ). [76]. ↵ Knutsen , E. , Harris , A. L. & Perander , M . Expression and functions of long non-coding rna neat1 and isoforms in breast cancer . British journal of cancer 126 , 551 – 561 ( 2022 ). OpenUrl CrossRef PubMed [77]. ↵ Desai , K. V. , et al. S100a6 as a biomarker in human breast cancer , Vol. 46 , 448 ( 2005 ). [78]. ↵ Margariti , A. et al. Direct reprogramming of fibroblasts into endothelial cells capable of angiogenesis and reendothelialization in tissue-engineered vessels . Proceedings of the National Academy of Sciences 109 , 13793 – 13798 ( 2012 ). URL https://api.semanticscholar.org/CorpusID:21888753 . OpenUrl Abstract / FREE Full Text [79]. ↵ Zhu , L. et al. Significant prognostic values of aquaporin mrna expression in breast cancer . Cancer management and research 1503 – 1515 ( 2019 ). [80]. ↵ Jaaks , P. & Bernasconi , M . The proprotein convertase furin in tumour progression . International journal of cancer 141 , 654 – 663 ( 2017 ). OpenUrl CrossRef PubMed [81]. ↵ Huo , Q. , Wang , J. & Xie , N . High hspb1 expression predicts poor clinical outcomes and correlates with breast cancer metastasis . BMC cancer 23 , 501 ( 2023 ). [82]. ↵ Grosset , A.-A. et al. Galectin signatuares contribute to the heterogeneity of breast cancer and provide new prognostic information and therapeutic targets . Oncotarget 7 , 18183 ( 2016 ). [83]. ↵ Luo , Q. et al. Col11a1 serves as a biomarker for poor prognosis and correlates with immune infiltration in breast cancer . Frontiers in Genetics 13 , 935860 ( 2022 ). [84]. ↵ Xu , S.-G. , Yan , P.-J. & Shao , Z.-M . Differential proteomic analysis of a highly metastatic variant of human breast cancer cells using two-dimensional differential gel electrophoresis . Journal of cancer research and clinical oncology 136 , 1545 – 1556 ( 2010 ). OpenUrl CrossRef PubMed [85]. ↵ Meng , T. , Liu , L. , Hao , R. , Chen , S. & Dong , Y . Transgelin-2: a potential oncogenic factor . Tumor Biology 39 , 1010428317702650 ( 2017 ). [86]. ↵ Bavis , M. M. , Nicholas , A. M. , Tobin , A. J. , Christian , S. L. & Brown , R. J . The breast cancer microenvironment and lipoprotein lipase: Another negative notch for a beneficial enzyme? FEBS Open Bio 13 , 586 – 596 ( 2023 ). OpenUrl CrossRef PubMed [87]. ↵ Kothari , C. et al. Is carboxypeptidase b1 a prognostic marker for ductal carcinoma in situ? Cancers 13 , 1726 ( 2021 ). OpenUrl CrossRef PubMed [88]. ↵ Bouchal , P. et al. Combined proteomics and transcriptomics identifies carboxypeptidase b1 and nuclear factor κ b (nf- κ b) associated proteins as putative biomarkers of metastasis in low grade breast cancer . Molecular & Cellular Proteomics 14 , 1814 – 1830 ( 2015 ). OpenUrl CrossRef PubMed [89]. ↵ Pandurangi , R. , Sekar , T. & Paulmurugan , R . Restoration of the lost human beta defensin-1 protein in cancer as a strategy to improve the efficacy of chemotherapy . Journal of Medicinal Chemistry 67 , 14200 – 14209 ( 2024 ). URL doi: 10.1021/acs.jmedchem.4c01040 pmid: . PMID: 39137365 . OpenUrl CrossRef PubMed 90. ↵ Huang , S. , Zhang , X. , Wei , Y. & Xiao , Y . Checkpoint cd24 function on tumor and immunotherapy . Frontiers in Immunology 15 ( 2024 ). URL https://api.semanticscholar.org/CorpusID:268212342 . [91]. ↵ Ganaie , I. A. , Malik , M. Z. , Naqvi , S. H. , Jain , S. K. & Wajid , S . Differential levels of alpha-1-inhibitor iii, immunoglobulin heavy chain variable region, and hypertrophied skeletal muscle protein gtf3 in rat mammary tumorigenesis . Biochimie 174 , 57 – 68 ( 2020 ). OpenUrl CrossRef PubMed [92]. ↵ Liberzon , A. et al. The molecular signatures database hallmark gene set collection . Cell systems 1 , 417 – 425 ( 2015 ). OpenUrl CrossRef PubMed [93]. ↵ Elayat , G. & Selim , A . Angiogenesis in breast cancer: insights and innovations . Clinical and Experimental Medicine 24 , 178 ( 2024 ). 94. ↵ Oshi , M. , et al. Association of allograft rejection response score with biological cancer aggressiveness and with better survival in triple-negative breast cancer (tnbc) . ( 2021 ). [95]. ↵ Oshi , M. et al. Enhanced immune response outperform aggressive cancer biology and is associated with better survival in triple-negative breast cancer . NPJ Breast Cancer 8 , 92 ( 2022 ). [96]. ↵ Zarrizi , R. et al. Germline rbbp8 variants associated with early-onset breast cancer compromise replication fork stability . The Journal of clinical investigation 130 , 4069 – 4080 ( 2020 ). OpenUrl CrossRef PubMed [97]. ↵ Oshi , M. et al. Degree of early estrogen response predict survival after endocrine therapy in primary and metastatic er-positive breast cancer . Cancers 12 , 3557 ( 2020 ). OpenUrl CrossRef PubMed [98]. ↵ Agrawal , R. , Imieliński , T. & Swami , A . Mining association rules between sets of items in large databases , 207 – 216 ( 1993 ). [99]. ↵ Bay , S. D. & Pazzani , M. J . Detecting group differences: Mining contrast sets . Data mining and knowledge discovery 5 , 213 – 246 ( 2001 ). OpenUrl CrossRef [100]. Novak , P. K. , Lavrăc , N. , Gamberger , D. & Krstăcić , A . Csm-sd: Methodology for contrast set mining through subgroup discovery . Journal of Biomedical Informatics 42 , 113 – 122 ( 2009 ). OpenUrl CrossRef PubMed [101]. ↵ Gudyś , A. , Sikora , M. & Wróbel , Ł . Separate and conquer heuristic allows robust mining of contrast sets in classification, regression, and survival data . Expert Systems with Applications 123376 ( 2024 ). [102]. ↵ Gudyś , A. , Sikora , M. & Wrobel , Ł . Rulekit: A comprehensive suite for rule-based learning . Knowledge-Based Systems 194 , 105480 ( 2020 ). [103]. ↵ Kingma , D. P. & Ba , J . Bengio , Y. & LeCun , Y . (eds) Adam: A method for stochastic optimization . (eds Bengio , Y. & LeCun , Y .) 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings ( 2015 ). URL http://arxiv.org/abs/1412.6980 . [104]. ↵ Uno , H. , Cai , T. , Pencina , M. J. , D’Agostino , R. B. & Wei , L.-J . On the c-statistics for evaluating overall adequacy of risk prediction procedures with censored survival data . Statistics in medicine 30 , 1105 – 1117 ( 2011 ). OpenUrl CrossRef PubMed [105]. ↵ Harrell Jr, F. E. , Lee , K. L. & Mark , D. B . Multivariable prognostic models: issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors . Statistics in medicine 15 , 361 – 387 ( 1996 ). OpenUrl CrossRef PubMed Web of Science [106]. ↵ Monod , M. et al. Torchsurv: A lightweight package for deep survival analysis . Journal of Open Source Software 9 , 7341 ( 2024 ). URL doi: 10.21105/joss.07341 . OpenUrl CrossRef View the discussion thread. Back to top Previous Next Posted April 17, 2025. Download PDF Supplementary Material Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Can interpretability and accuracy coexist in cancer survival analysis? Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Can interpretability and accuracy coexist in cancer survival analysis? Piyush Borole , Tongjie Wang , Antonio Vergari , Ajitha Rajan bioRxiv 2025.04.11.648380; doi: https://doi.org/10.1101/2025.04.11.648380 Share This Article: Copy Citation Tools Can interpretability and accuracy coexist in cancer survival analysis? Piyush Borole , Tongjie Wang , Antonio Vergari , Ajitha Rajan bioRxiv 2025.04.11.648380; doi: https://doi.org/10.1101/2025.04.11.648380 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Bioinformatics Subject Areas All Articles Animal Behavior and Cognition (7642) Biochemistry (17715) Bioengineering (13907) Bioinformatics (42005) Biophysics (21472) Cancer Biology (18624) Cell Biology (25534) Clinical Trials (138) Developmental Biology (13390) Ecology (19935) Epidemiology (2067) Evolutionary Biology (24356) Genetics (15617) Genomics (22529) Immunology (17753) Microbiology (40437) Molecular Biology (17200) Neuroscience (88697) Paleontology (667) Pathology (2840) Pharmacology and Toxicology (4829) Physiology (7653) Plant Biology (15171) Scientific Communication and Education (2046) Synthetic Biology (4304) Systems Biology (9827) Zoology (2272)
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.