Multidimensional proteomics and explainable AI feature selection identify cross-platform lung cancer molecular signature in blood plasma

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Lung cancer is the leading cause of cancer mortality worldwide with 70% diagnosed late stage despite low-dose computed tomography (LDCT) screening availability. We combined data-independent acquisition mass-spectrometry (DIA-MS) and proximity extension assay (PEA) with explainable artificial intelligence (XAI)-led machine learning (ML) for plasma-based biomarker discovery. From a 490 lung cancer and 124 matched-control cohort, ML models were trained to predict lung cancer achieving an AUROC of 0.91 [95% CI: 0.88–0.93] and 0.97 [95% CI: 0.92–0.98] in DIA-MS and PEA, respectively. XAI characterised networks of model-consistent features primarily related to infection and inflammatory responses. We then introduced a DNA-aptamer proteomics method and identified a cross-platform concordance panel, with performances of 0.88 [95% CI: 0.80–0.90] and 0.88 [95% CI: 0.81–0.95] in DIA-MS and PEA, respectively. This study demonstrates that combining multi-dimensional proteomics with XAI-ML can characterise robust biomarker signatures.
Full text 239,147 characters · extracted from preprint-html · click to expand
Multidimensional proteomics and explainable AI feature selection identify cross-platform lung cancer molecular signature in blood plasma | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Multidimensional proteomics and explainable AI feature selection identify cross-platform lung cancer molecular signature in blood plasma Peter Jianrui Liu, Harriet Ferguson, Nikola Gushterov, Benedikt Kessler, and 12 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7660411/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Lung cancer is the leading cause of cancer mortality worldwide with 70% diagnosed late stage despite low-dose computed tomography (LDCT) screening availability. We combined data-independent acquisition mass-spectrometry (DIA-MS) and proximity extension assay (PEA) with explainable artificial intelligence (XAI)-led machine learning (ML) for plasma-based biomarker discovery. From a 490 lung cancer and 124 matched-control cohort, ML models were trained to predict lung cancer achieving an AUROC of 0.91 [95% CI: 0.88–0.93] and 0.97 [95% CI: 0.92–0.98] in DIA-MS and PEA, respectively. XAI characterised networks of model-consistent features primarily related to infection and inflammatory responses. We then introduced a DNA-aptamer proteomics method and identified a cross-platform concordance panel, with performances of 0.88 [95% CI: 0.80–0.90] and 0.88 [95% CI: 0.81–0.95] in DIA-MS and PEA, respectively. This study demonstrates that combining multi-dimensional proteomics with XAI-ML can characterise robust biomarker signatures. Biological sciences/Cancer Health sciences/Biomarkers Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Lung cancer is the leading cause of cancer mortality worldwide accounting for one fifth of total global cancer mortality and two million deaths per annum globally 1 . Over 70% of lung cancer diagnoses are made at late stages, when regional-to-distant disease have a five-year survivorship of 9–37% for non-small cell lung cancer (NSCLC) and 3–18% for small cell lung cancer (SCLC), in the US, respectively 2 . In contrast, early stage, localised lung cancer five-year survivorship is 65% for NSCLC and 30% for SCLC. Therefore, enabling earlier lung cancer diagnosis has been a priority for improving patient survivorship and outcome. Low dose computed tomography (LDCT) has been widely studied for lung cancer screening in high-risk populations. A number of studies including the NELSON study in the EU and the National Lung Screening Trial (NLSC) in the United States showed a 20% reduction in mortality in the LDCT screening group 3 , 4 . However, uptake and adoption have widely been reported to be poor ranging from as low as < 5% in the US to up to ~ 50% in the UK. Limiting factors including radiation exposure, cost, and throughput have been cited to prevent LDCT accessibility and uptake, demonstrating challenges for scaling up 5,6 . Further, as smoking rates fall in Western and Asian countries, a greater fraction of lung cancer are occurring in individuals with a light or never-smoking history 7 ; these individuals are ineligible for LDCT screening according to current criteria. Therefore, a minimally invasive and scalable approach for identifying high-risk lung cancer individuals with the potential for population level deployment may help improve current lung cancer screening paradigms. The blood proteome harbours promising information for the identification of a molecular signature for lung cancer, which can be applied as a liquid biopsy blood test. Prior efforts using a cell-free DNA based approach have shown limited sensitivity and specificity especially in earlier stage and low tumour burden disease, likely due to a low level of circulating tumour DNA 8 – 10 . In contrast, the blood proteome provides important insight including tumour-specific inflammatory and host response even during early and low-tumour burden disease, serving as a potent repertoire for novel biomarker discovery 11 , 12 . Quantification of such circulating proteins may create a molecular signature for lung cancer. To this end, past studies have focused on targeted proteomics approaches such as proximity extension assay (PEA) and parallel/multiple reaction monitoring, liquid-chromatography tandem-mass spectrometry (LC-MS) or untargeted approaches including data-dependent/ data-independent LC-MS. However, different platforms may exhibit biased preferences for different proteins and pathways and show incongruence between quantification platforms, limiting biomarker signatures to a specific platform 13 . To enable the identification of a robust blood biomarker signature, the use of reliable machine learning (ML) models that generalise to unseen data and can model the interactions between the features is critical 14 . In particular it is crucial to aggregate the information and importance of biomarkers across many different models to identify a robust panel that functions regardless of underlying models or model parameters. In addition, the black-box nature of more advanced ML models has been an ongoing concern, especially when clinical guidelines and regulatory approval processes benefit from understanding the rationale behind results. In this present study, we deployed multidimensional-proteomics combined with an explainable artificial intelligence (XAI)-guided ML approach on a large lung cancer cohort to identify a robust plasma-based molecular signature across diverse platforms. This combined data-independent acquisition mass spectrometry (DIA-MS) as an unbiased proteomics approach with a targeted PEA (Olink) for biomarker discovery from a cohort of 614 individuals (490 lung cancer cases across all stages, histologies, smoking exposures; and 124 age-, sex-matched controls from a research Canadian screening cohort). Using an 80/20 train/holdout strategy with cross-validated Shapley values for interpretability, we identified high-importance model-consistent features and uncovered key feature interactions relevant to lung cancer prediction. A third proteomics approach, targeted DNA-aptamer (SomaScan), was deployed on an expanded subset of patients to characterise a focused cross-platform biomarker signature. This multidimensional proteomics and XAI-driven methodology addresses core challenges in liquid biopsy development, including black-box model transparency, feature generalisability, and platform-specific constraints. We report the first cross-platform plasma-based lung cancer molecular signature identified through integrated multi-platform proteomics and explainable AI. Results Synergy between DIA-MS and PEA proteomics approaches enhances proteome coverage Single-shot DIA-MS and PEA were used for protein quantification of plasma samples from 490 post-diagnosis, treatment-naive lung cancer cases covering all stages and major histologies, and 124 age-/sex-matched cancer-free controls (n = 124) recruited through a lung screening programme, as described previously 15 , 16 (Fig. 1a, Extended Data Fig. 1a). Clinical and demographic characteristics are summarised in Table 1. Table 1 Clinical and demographic characteristics of healthy control vs lung cancer cohort Control (n = 124) NSCLC (n = 433) SCLC (n = 55) Rare (n = 2) Number of females, n (%) 63 (50.8) 219 (50.6) 21 (38.2) 1 (50) Mean age in years, (sd) 69 (10) 68 (11) 68 (10) 57 (13) Age range (10 year intervals), n (%) : < 50 0 (0) 23 (5.3) 2 (3.6) 1 (50) 50–60 27 (21.8) 76 (17.6) 12 (21.8) 0 (0) 60–70 30 (24.2) 130 (30) 16 (29.1) 1 (50) 70–80 54 (43.5) 148 (34.2) 15 (27.3) 0 (0) 80–90 13 (10.5) 49 (11.3) 10 (18.2) 0 (0) > 90 0 (0) 7 (1.6) 0 (0) 0 (0) Ethnicity, n (%) : Asian/Pacific Islander 5 (4) 66 (15.2) 3 (5.5) 0 (0) Black/African-Canadian 0 (0) 11 (2.5) 0 (0) 0 (0) First Nations 0 (0) 2 (0.5) 0 (0) 0 (0) Latino/Hispanics 0 (0) 2 (0.5) 1 (1.8) 0 (0) Mixed 0 (0) 4 (0.9) 0 (0) 0 (0) White/Caucasian 118 (95.2) 268 (61.9) 45 (81.8) 2 (100) Unknown 1 (0.8) 80 (18.5) 6 (10.9) 0 (0) Smoking history: Mean pack-years (sd) 40 (21) 32 (31) 50 (29) 27 (11) Heavy (20 py) 107 (86.3) 252 (58.2) 49 (89.1) 1 (50) Non-heavy (< 20 py) 17 (13.7) 175 (40.4) 5 (9.1) 1 (50) Unknown 0 (0) 6 (1.4) 1 (1.8) 0 (0) Smoking status at diagnosis, n (%) : Current smoker 43 (34.7) 128 (29.6) 32 (58.2) 1 (50) Ex-smoker 81 (65.3) 150 (34.6) 16 (29.1) 1 (50) Ever smoker, NOS 0 (0) 1 (0.2) 0 (0) 0 (0) Light ex-smoker 0 (0) 11 (2.5) 0 (0) 0 (0) Never smoker 0 (0) 119 (27.5) 2 (3.6) 0 (0) Stage of cancer, n (%) : Stage 0 NA 1 (0.2) 0 (0) 0 (0) Stage 1 NA 139 (32.1) 4 (7.3) 1 (50) Stage 2 NA 44 (10.2) 2 (3.6) 0 (0) Stage 3 NA 86 (19.9) 17 (30.9) 1 (50) Stage 4 NA 161 (37.2) 32 (58.2) 0 (0) Unknown NA 2 (0.5) 0 (0) 0 (0) Clinical morphology, n (%) : LUAD NA 251 (58) NA 2 (100) LUSC NA 107 (24.7) NA 0 (0) Large cell carcinoma NA 32 (7.4) NA 0 (0) Other NA 43 (9.9) 55 (100) 0 (0) Molecular alteration, n (%) : EGFR- NA 4 (0.9) 0 (0) 0 (0) EGFR+ NA 118 (27.3) 0 (0) 0 (0) EGFR Unknown NA 3 (0.7) 0 (0) 0 (0) EGFR NR NA 308 (71.1) 55 (100) 2 (100) ALK+ NA 1 (0.2) 0 (0) 0 (0) ALK- NA 114 (26.3) 0 (0) 0 (0) ALK NR NA 308 (71.1) 55 (100) 2 (100) ALK Unknown NA 10 (2.3) 0 (0) 0 (0) KRAS+ NA 1 (0.2) 0 (0) 0 (0) TP53+ NA 1 (0.2) 0 (0) 0 (0) NSCLC, non-small cell lung cancer; SCLC, small cell lung cancer; n, number of; sd, standard deviation; py, pack-years; NOS, not otherwise specified; LUAD, lung adenocarcinoma; LUSC, lung squamous cell carcinoma +, positive for mutation; -, negative for mutation; NR, not reported; NA, not applicable. DIA-MS quantified a wider-range of proteins (n = 3,655) compared with PEA (n = 2,884), with 1,136 quantified by both (Fig. 1b, Supplementary Fig. 1a). Given the high dynamic range of plasma protein concentrations, we assessed available concentrations as reported in the Human Protein Atlas (HPA) 17 . DIA-MS quantified 1,736 mid-low abundance proteins (< 0.1mg/L), while PEA quantified 1,486 (Fig. 1c). Data missingness was higher in DIA-MS (37.07% ± 34.82%) compared with PEA (13.09% ± 24.33%). Both datasets showed a small but statistically significant (FDR adjusted P < 0.05) negative Spearman correlation between protein concentration and missingness, suggesting poorer quantification at lower abundance (Fig. 1d). Assessment of secretome location in the HPA showed most proteins were not known to be secreted, with secreted proteins predominantly secreted to the blood across the abundance range (Extended Data Fig. 1b). Matrix-matched technical replicates were acquired by PEA alongside clinical samples (n = 14), while for DIA-MS a pool of clinical samples was evaluated (n = 42). PEA had lower CVs than DIA-MS with a median CV of 21.4% versus 34.1%, although high CV proteins were not consistent across platforms (Fig. 1e). Analysis of Spearman correlation across protein abundance range indicated higher cross-platform variability for low-abundance proteins (Fig. 1f). Implementing dual protein quantification platforms expanded information for biomarker discovery (Fig. 1b). Over-representation analysis (ORA) showed over-representation of proteins associated with cell migration, signalling, haemostasis and immune response in PEA. Also enriched by DIA-MS were cytoskeleton- and chromatin organisation-associated proteins (Fig. 1g). Together, these results show that combining DIA-MS and PEA increases the pool of potential biomarker candidates. Univariate analysis identified lung cancer-associated plasma proteins and sub-cohort dependent variation Principal component analysis (PCA) showed a gradual stage-dependent separation of lung cancer cases from cancer-free controls (Fig. 2a,b, Supplementary Fig. 2). In both datasets, early-stage lung cancer cases (Stage I-II), were closer to cancer-free controls compared with late stage (Stage III-IV). PCA loadings showed that separation was distributed across many features (Fig. 2c,d). SERPINA3, PI16 and GSN contributed to separation in both datasets. Analysis of differentially abundant proteins between cancer-free controls and lung cancer cases showed a larger level of significant downregulation (FDR-adjusted P 1) (Fig. 2e,f). Differential abundance was also evaluated on sub-cohorts including early and late lung cancer, lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC) and male-/female-only comparisons. Minimal difference in differential abundance was observed across sub-cohorts, with the exception of LUAD/LUSC in both datasets, as well as early lung cancer in PEA (Extended Data Fig. 2a,b). Gene set enrichment analysis (GSEA) showed the direction of regulation was similar across platforms, although little overlap was observed in significantly regulated pathways (FDR-adjusted P < 0.05) (Fig. 2g, Extended Data Fig. 2c), suggesting low concordance/overlap between approaches could generate different interpretations. Regulation of actin cytoskeleton organisation, protein complex assembly, and hydrogen peroxide metabolism associated proteins was observed in DIA-MS, irrespective of sub-cohort. In PEA, pathways associated with coagulation and GTPase/RAS signalling, were significantly downregulated across most sub-cohorts, while RNA metabolism, immune response and chromatin organisation pathways were significantly upregulated in SCLC. Overlapping pathways were related to acute-phase response and actin cytoskeleton organisation, with differences among subtypes largely derived from smaller populations (SCLC and LUSC). Overall, univariate analysis of plasma proteins revealed limited insight into biological processes, particularly in the more heterogeneous sub-populations. XAI-informed multivariate feature selection to build panels of lung cancer plasma biomarker candidates We next trained 75 tree-based ML models using combinations of different model architectures, feature de-selection, data augmentation, feature selection, multi-objective optimisation and XAI techniques to identify biomarkers that were consistently predictive of lung cancer across multiple models. All data underwent identical sample stratification, train/holdout splitting with separate pre-processing, minimising data-leakage to unseen holdout data to promote generalisation. All models were evaluated across nine lung cancer sub-cohorts for both datasets, giving rise to 1,350 ML models (Fig. 3a, Extended Data Fig. 3a, Supplementary Table 1). PEA models had better performance than DIA-MS models; all and late lung cancer, NSCLC, LUSC and LUAD and male-only PEA models achieved better performance than other sub-cohorts, as assessed by area under the receiver operator characteristic curve (AUROC) and Matthews Correlation Coefficient (MCC) (Fig. 3b-e). In DIA-MS models NSCLC performance was higher than SCLC. Male-only models outperformed female-only models in both datasets. Female-only lung cancer cases had a lower incidence of LUSC and SCLC compared with males and higher incidence of LUAD, with comparable age distribution (Supplementary Fig. 3a, b). Interestingly, assessment of an all-lung cancer trained-model predicting sub-cohorts showed similar decreased performance in female-only versus male-only cohort (Extended Data Fig. 3b). To identify robust feature sets, we next assessed Shapley value feature importance across all lung cancer models 18 . Some high-importance features were consistently selected for a model (WFDC2, CILP, NCF2, MMP7, GLRX) while some high-importance features showed strong model-dependency (SH3BP2, VWA8, KIF20B, CLEC4G and BCAT1) (Fig. 3f). Upon assessment of individual LC models, MB, a feature selected by DIA-MS and PEA, showed an early contribution during recursive feature addition (RFA) for both approaches (Fig. 3g), also showcasing how using Shapley values during RFA promoted incremental performance increases. Smaller performance increases suggested incorporation of redundancy during feature selection, while minor AUROC changes accompanied by larger changes in MCC highlighted insights from considering two performance metrics. Univariate analysis is often used to select features for ML 19 . During feature de-selection, features that had low information about the target (mutual information (MI) 3.5%, however a high number of proteins with high MI were not differentially abundant (Extended Data Fig. 3c,d). Moreover, a large number of high-importance features did not pass the threshold for differential abundance, particularly for DIA-MS models, while some high importance features did not have MI > 3.5%, emphasising the utility in using multiple feature selection-configurations. A direct comparison of univariate with the XAI-informed feature selection approach used here showed that despite containing non-differentially abundant proteins, overall ML-driven feature selection achieved higher performance even with a simple logistic regression model (Extended Data Fig. 3e). Characteristics of features selected The Rashomon effect describes the phenomena in ML where similar performances can be achieved with different internal workings of modell 20 We addressed this by comparing feature selection across many different models. To evaluate feature selection, we considered models that generalised, determined by comparing holdout performance to 95% CI from training data (Extended Data Fig. 3a). Five features were selected in generalisable PEA and DIA-MS models (LRG1, SDC1, MB, LTA4H and FGL1) (Fig. 4a). LTA4H was the only common feature regulated differently across approaches, with minor upregulation (significant in SCLC) in DIA-MS and significant downregulation across PEA sub-cohorts ( P 1). Other proteins quantified by both methodologies were only selected as features for DIA-MS (n = 25) or PEA (n = 13), albeit mostly regulated similarly across approaches. Like LTA4H, some showed divergent abundance (THBS2, THOP1 and ADAMTSL4) while others showed strong regulation in one dataset with minimal regulation by the alternative method (MUC16, CEACAM5, CILP, HAGH and LXN). Across generalisable models, high importance features according to Shapley values were similar for all/late lung cancer, NSCLC and LUAD. Variation across sub-cohorts was generally lower for PEA. SCLC and sex-split sub-cohorts diverged most from other models. In female-only DIA-MS models, high importance features for many cohorts (MB, WFDC2, HBA2, TAGLN2 and PPBP) were less informative while other candidates (POSTN, VWA8, HAGH, NAGLU and MUC1) were never selected. Similarly for male-only models in DIA-MS, while the overall order of importance was preserved, many features were never selected. In early lung cancer models, SDC1, POSTN, PPBP and SERPINA3 were less important/absent compared with all/late lung cancer models, suggesting their presence arises from influence of later stage lung cancer. Higher importance features had low missingness in DIA-MS, whereas in PEA two candidates (LTA4H and NEXN) were largely quantified below the LLOD (Extended Data Fig. 4a-b). High-importance lung cancer features were lower in abundance in PEA models (Extended Data Fig. 4c). Biological pathways enriched in lung cancer features showed that proteins involved in chemotaxis, cell adhesion, wound healing and immune response (acute-phase, complement activation and humoral immune response) were present at varying levels across sub-cohorts in both datasets (Fig. 4c). This challenged previous GSEA showing immune response pathways enrichment in PEA only (Fig. 2g). Both datasets relied on features associated with cell migration, while in PEA features were enriched for receptor-mediated endocytosis, integrin and cytokine signalling. DIA-MS was also enriched for oxidative stress features (Fig. 4c). In both datasets, if known, secretion to blood was more prevalent (Extended Data 4d). XAI-driven feature selection captures strong interactions between features. To further understand how models used features, we investigated how Shapley values individually and through interactions explained differences between cancer-free controls and sub-cohorts. Feature quantification from a lung cancer model separated cancer-free control samples from early and late lung cancer samples by t-SNE (Extended Data Fig. 4e). Shapley values better separated cancer-free controls from lung cancer samples when compared to protein quantification, with early/late differences being less obvious (Extended Data Fig. 4f). This suggested high-importance lung cancer features could be used regardless of stage. This is further supported by the high performance of lung cancer models predicting early lung cancer (AUROC = 0.89 [95% CI: 0.79–0.98], MCC = 0.65 [95% CI: 0.46–0.82]) (Extended Data Fig. 3b) compared with the performance across early lung cancer models (AUROC = 0.82 [95% CI: 0.77–0.88], MCC = 0.48 [95% CI: 0.34–0.64]) in DIA-MS (Fig. 3b,d). Analysis of Shapley value model-decision paths showed individual feature contribution towards a prediction (Extended Data Fig. 6a,b), while interactions demonstrated how features are mediated by each other. Shapley value interactions from a lung cancer model identified complex networks of feature interactions (Fig. 4d,e, Extended Data Fig. 6c,d). For DIA-MS, this network contained a higher level of interactions compared with PEA. Shapley value importance scores were comparatively lower and more evenly distributed within the PEA network. For DIA-MS, highly important features were involved in the strongest interactions (MB, WFDC2, HBA2, CILP and SDC1). Increased levels of CILP, a universal feature for all DIA-MS models, interacted with decreased levels of SDC1, SELENBP1 and/or WFDC2 to be predictive of cancer-free controls (Extended Data Fig. 6e). Similarly, decreased levels of SDC1 interacted with increased levels of MB and decreased levels of CEACAM5 and/or MMP7 to be predictive of cancer-free controls in the analysis of top PEA interactions. (Extended Data Fig. 6f). Several interactions (TAGLN2-MUC1, DDTL-ACTB, CASP8-CLIP2 and NCF2-CLIP) identified two groups of controls. These interactions emphasise the importance of using models capable of capturing feature interactions for disease identification. The identification of a platform-agnostic lung cancer plasma protein signature. We next introduced a third DNA-aptamer-based proteomic approach (DNA-APT) to further evaluate concordance across platforms. A panel of 100 SomaScan quantified proteins selected from previously identified features (Fig. 5a, Extended Data Fig. 7a,b) was applied to a cohort of 140 patients selected from the previously described cohort (Table 1), 72 cancer-free controls and 68 lung cancer cases. Variability was low in DNA-APT assay and dimensionality reduction showed separability between cancer-free controls, early and late lung cancer samples (Extended Data Fig. 7c). The majority of proteins quantified in all three methods (n = 27, 57.1%) had a significant Spearman correlation (FDR-adjusted P < 0.05) (Fig. 5b, Extended Data Fig. 7d). Six features showed no concordance with the DNA-APT approach. A group of predominantly PEA features (n = 16) had a strong, significant correlation between PEA/DNA-APT only. Similarly, a group of DIA-MS features (n = 16) had strong correlation between DIA-MS/DNA-APT only. High-importance feature SERPINA3 had opposite correlation in DIA-MS/DNA-APT vs PEA/DNA-APT. Overall, concordance between DNA-APT and PEA was higher than concordance with DIA-MS, with median Spearman R of 0.567 versus 0.363 (Fig. 5c). A panel of proteins with Spearman R > 0.4 across all three panels were evaluated (n = 18) and a second panel that contained proteins with Spearman R > 0.4 (n = 25) between DIA-MS/DNA-APT or PEA/DNA-APT data were evaluated using DIA-MS and PEA (see methods). Similar performances were seen between DIA-MS and PEA for the three-platform panel, with AUROC of 0.88 [95% CI: 0.80–0.95] and 0.88 [95% CI: 0.81-0. 95], and MCC of 0.49 [95% CI: 0.31–0.67] and 0.60 [95% CI: 0.44–0.77] for DIA-MS and PEA, respectively (Fig. 5d). The PEA/DNA-APT two-platform panel achieved the highest performance, with AUROC of 0.97 [95% CI: 0.94–0.99] and MCC of 0.71 [95% CI: 0.54–0.86]. The DIA-MS/DNA-APT model outperformed the three-platform panel with AUROC of 0.93 [95% CI: 0.86–0.97] and MCC of 0.60 [95% CI: 0.44–0.78]. To evaluate the DNA-APT assay, a partial least squares (PLS) regression was used to show differences between cancer-free controls and lung cancer, which were similarly distinguishable across all platforms (Fig. 5e). Analysis of loadings showed that similar features were responsible for separation on all platforms, separating cancer-free controls (MB, PF4, PPBP, CNTN3 and KIT) from lung cancer (HAGH, ALDH1A1, FGL1, MMP9, C9, CFHR5 and LBP). The same was carried out for the two-platform panels which achieved similar separation by PLS regression (Extended Data Fig. 7e). We show the identification of feature panels compatible across multiple proteomics platforms, as demonstrated by a similar degree and driver(s) of separation between the cross-platform panels. Discussion This study deployed unbiased DIA-MS and targeted PEA proteomics as discovery approaches in combination with XAI-guided ML feature selection. On a cohort of 614 patients including 490 lung cancer cases (stages I-IV) covering all major histologies and 124 age-, sex-matched controls we evaluated plasma molecular signatures for lung cancer across 1,350 ML models. Further characterisation using a DNA-aptamer approach identified an eighteen-feature lung cancer plasma signature with concordance across all platforms, achieving AUROC of 88.2% and 88.3% in DIA-MS and PEA holdout validation, respectively. Signatures with high concordance across two of the three panels achieved even higher AUROC between 92.4% to 97%. PLS regression on the DNA-APT validation dataset clearly distinguished lung cancer from controls. The proteomics approaches deployed are among the most widely used, DIA-MS offering unbiased discovery, while PEA and DNA-APT are often favoured in clinical research for straightforward data processing 21 . LC-MS based proteomics is increasingly considered for clinical purposes, particularly with regards to cancer detection and disease monitoring due to high-throughput and multiplexing 22 . This study leveraged both the overlap and divergence between platforms to expand coverage and to assess feature robustness. In-line with previous studies, we observed platform-specific discrepancies 21 , 23 , 24 , potentially arising from either DIA-MS pre-processing or immunoassay epitope-specific biases, reinforcing concerns about the reliability of signatures developed from a single method. Dynamic range is often considered a limitation of DIA-MS. Low abundance proteins showed greater variability and reduced concordance (Extended Data Fig. 4c), however models from all platforms relied on such features to varying degrees and yielded predictive signatures. By focusing on high-importance features with strong cross-platform concordance, we identified a robust plasma signature for lung cancer, supporting the development of transferable, platform-agnostic biomarker panels, facilitating scalable and flexible clinical implementation. Despite differences across platforms, cytokine/chemokine-mediated signalling were associated with both DIA-MS and PEA biomarker panels, with several proteins previously considered for lung cancer detection, monitoring or prognosis. Many candidates have also been previously associated with cancer pathology but would be considered novel biomarkers of lung cancer prediction (Extended Data Fig. 5a, b). Five proteins were selected across both proteomics platforms (MB, SDC1, LRG1, FGL1 and LTA4H). MB had high importance/concordance across both platforms and had previously been associated with skeletal-muscle wasting in cancer-related cachexia 25 . We observed that low plasma MB coupled with regulation of other candidates (SDC1 and COTL1) were important in lung cancer prediction. SDC1 had high importance across both datasets, with a known role in lung cancer pathology and as a prognostic marker in multiple solid-tumours 26 . We observed significant upregulation of SDC1 in all but early lung cancer sub-cohorts in both datasets, indicating increased circulating SDC1 with stage. To the best of our knowledge, MB and SDC1 have not previously been reported as being predictive of lung cancer. We also characterised high importance lung cancer features associated with immune/inflammatory responses. SERPINA3 and CRP, associated with acute-phase response, have previously been proposed as biomarkers for early NSCLC. We, however, observed SERPINA3 upregulation to be associated with later-stage NSCLC, LUSC/male-only lung cancer. There were however early-stage associated proteins characterised by both approaches. In PEA, the transferrin receptor (TFRC) was important for early lung cancer among other sub-cohorts (Fig. 4b). TFRC upregulation is known in NSCLC and other solid-tumours, also expressed on the surface of activated immune cells 27 , potentially with a dual-role in mediating lung cancer progression 28 . Similarly, HAGH showed upregulation in early-stage lung cancer ( P < 0.05 in DIA-MS only) (Fig. 4a, Fig. 5e), the only candidate to show an early-dominant response across both platforms, suggesting HAGH upregulation in plasma may occur early and reduce over disease progression. This is in line with HPA data showing HAGH has favourable prognostic value in other cancers including liver, renal and pancreatic cancer 29 . Pathway analysis showed that proteins related to oxidative stress and wound healing were informative, even in early-stage disease. Together, this underscores the potential of protein-based approaches to detect early-stage and low-tumour burden diseases by identifying amplified inflammatory signals, addressing concerns about limited ctDNA signal in such cases 8 – 10 . Finally, we showed that despite having used cohort-specific models for biomarker selection, it appears that models trained on the complete cohort were similar/superior at understanding the subclassification task, possibly due to additional context from other samples. For example, predictions of early cancer patients in unseen holdout data were improved when late cancer patients were included in training. Future work focused on early detection of cancer may consider the inclusion of late-stage samples to improve separability. Model performances and feature selection varied between male and female-only models indicative of sex-based variation in plasma molecular signatures (Fig. 3b,c). Male-only candidates had a high association with cell migration and apoptotic pathways in line with other cohorts while some more universal features were not selected in female-only models (LRG1, SELENBP1, KIF20B, SERPINA3, MUC1, SAA1, MUC16, KRT19, CIT and TFRC) (Fig. 4a-c). Moreover, many biological pathways associated with lung cancer were not major candidates in female-only models, particularly immune response-related pathways, aligning with reports at the transcript level 30 . This is consistent with previous studies showing similar differences between sex-separated molecular signatures, proposing the use of different features and models for prediction of sex-based sub-cohorts 30 , 31 . Although we found lung cancer trained-models predicted female-only sub-cohorts to similar levels as female-only models, further assessment of sex-based differences in plasma molecular signatures and other confounding data in a larger cohort could further characterise the differences. This highlights the potential utility of evaluating sex-based differences in prediction using plasma molecular signatures, particularly in early detection where changes in protein abundance can be small in magnitude. The strengths in this study lie in the application of a multi-dimensional proteomics approach to a large cohort of lung cancer cases (n = 490) covering all major histologies with a good representation of early-stage cases. Additionally, we applied rigorous XAI and ML methodologies, with separate train/holdout processing where possible, using two complementary metrics, AUROC and MCC and exploring both boosting (XGBoost) and bagging (BRF) algorithms. XGBoost was prioritised to achieve a more concise set of informative features, giving rise to smaller panels of ~ 20 proteins. We also overcame poor model interpretability by using TreeSHAP approximation to efficiently compute cross-validated Shapley values across a large feature set. The resulting stable feature set is expected to maintain strong performance even in simple models. Finally, combining two distinct proteomics approaches for lung cancer circulating biomarker discovery broadened potential candidate coverage and combining three independent protein detection platforms allowed the assessment of a concordant lung cancer molecular signature across platforms. A limitation of this study is that, with highly correlated features, Shapley values can distribute across redundant features, diluting the contribution of individual biomarkers. This is a common limitation in multivariate feature selection methods. Similarly, we focused on XGBoost models which can overfit when not carefully designed, and we mitigated these potential limitations by applying unseen-holdout generalisation constraints, regularisation and evaluating multiple models. Another limitation was the over-representation of lung cancer and rarer subtypes. Although this enabled sufficient power to provide reliable insight into subtype specific biomarkers, a more-representative cohort can assess the utility of this ML-approach and feature panels in populations more representative of epidemiological distributions. Similarly, assessment of real-world efficacy and health economic feasibility of plasma molecular signatures were not within the scope of this study. Future work seeks to further evaluate and validate our approach with prospective population level-based detection studies. In conclusion, we deployed a multidimensional proteomics and XAI-guided ML biomarker discovery approach in a lung cancer cohort composed of early stage, late stage, and all common histologies with age-, sex-matched controls from a screening program cohort. The combined synergy between different proteomics approaches and interpretable ML models identified a cross-platform lung cancer plasma protein signature with explainable feature ranking, contribution, and interaction mechanisms for decision making. We anticipate this signature to be highly versatile, transferable, and scalable across different protein quantification platforms and look forward to future assessment of its cost effectiveness and real-world impact on patient outcomes. Given promising results of sub-cohort analysis demonstrating the ability to detect specific lung cancer histologies and stages, customisable panels may be deployed in diverse clinical use cases including at-risk population risk stratification, screening, and disease classification. Materials and methods Cohort selection and collection Lung cancer cases and healthy controls were recruited from Princess Margaret Cancer Centre’s Lung-CALIBRE program (Lung cancer Clinical And Liquid biopsy, Implementation, Breathomics, Radiomics for Early detection). Ethics approval for the CALIBRE program was obtained from The University Health Network Research Ethics Board (REB 06-639). Lung cancer cases were selected using a stage-stratified random process from a biospecimen database of lung cancer patients to sufficiently cover all major histologies and stages. Case participants included in this analysis were recruited sequentially from 2008–2019 which met the conditions of age over 18 years; histological confirmation of lung cancer, adequate plasma specimen (1.8ml vial) taken after initial diagnosis prior to commencing first treatment. Controls were distribution-matched to age and gender of lung cancer cases from the research lung cancer screening program. This program used a broad eligibility of 10-pack year minimum cigarette smoking history and a minimum age of 45 years old. Healthy control samples were collected at Princess Margaret Cancer Centre and lung cancer cases collected from Toronto General Hospital and Princess Margaret Hospital. All blood samples used in this study were collected in EDTA tubes and processed by the same operator to plasma using centrifugation within two hours of collection. All processing was carried out at ≤ 4°C and samples were stored at -80°C at Princess Margaret Cancer Centre for further use. To achieve > 80% power for showing a minimum lung cancer detection sensitivity and specificity of 70% and 90% respectively (compared to a null hypothesis of 50% each, with alpha = 0.05), an overall sample size of 490 cancers and 124 controls, with a 80/20% train/holdout split was selected to satisfy the requirement for a minimum number of holdout sample size of 49 cancers and 12 controls 32 . Sample preparation and processing for DIA-MS Sample randomisation The order of the 614 clinical samples to be processed were randomised based on demographic and clinical characteristics. The randomised samples were processed in seven individual 96-well plates. Each plate contained a patient's plasma samples and six “pool plasma” which were later used for quality control. The pool sample was formed by pooling 3 µL of clinical plasma samples of mixed types. Top 14 most abundant protein depletion Plasma samples stored at -80°C were thawed at RT. Five µL of neat plasma samples were used for each depletion. Top 14 depletion was performed for each sample using high-selection top 14 depletion resin (A36372, Thermo Scientific) following manufacturer’s instructions. Depleting the plasma high-abundance proteins improves the dynamic range in LC-MS proteomics analysis. Approximately 100 µL of depleted plasma per well were collected for each sample; 25 µL of depleted plasma from each sample were transferred to a fresh 96-well PCR plate (LoBind, Eppendorf) for digestion. Haemolysis was evaluated as described previously 33 on the unprocessed plasma sample. SP3 in solution digestion 5 µL of 60 mM TCEP (XH347846, Thermo Scientific), 240 mM CAA (CO267-100G, Merck) mix were added to each sample in the well of a PCR 96-well plate. The plate was sealed with a heat sealer (E0030127838, Eppendorf) on a heat seal (E5391EN100431, Eppendorf). Reduction and alkylation were performed by incubating the samples at 95°C for 15 min on a Thermomixer. Single pot, solid-phase enhanced sample preparation protocol (SP3) was used to prepare peptide samples. Briefly, Sera-Mag™ Carboxylate-Modified Magnetic Beads (45152105050250 and 65152105050250, Cytiva) were washed twice in HPLC grade water and 2 µL added to each sample along with 80 µL 99.8% EtOH (E/0650/17, Fisher chemical) and incubated at RT for 5 mins. The beads were subsequently washed with 80% EtOH three times. 1 µg MS grade Pierce™ trypsin protease (90058, Thermo Scientific) in 100 mM ammonium bicarbonate (09831-500G, Merck) was added to each sample and incubated overnight at 37°C. Digested samples were transferred to a new plate, lyophilised and stored at -20°C ready for data acquisition. Sample loading Dried samples were resuspended in 100 µL 0.1% FA (85170, Thermo Scientific). For analysis on the Evosep One LC system, the sample was loaded onto Evotips (EV2011, Evosep) according to Evotip sample loading protocol. Briefly, the Evotips were rinsed with 20 µL solvent B (85174, Thermo Scientific), conditioned with isopropanol (99.5%, 383920025, Thermo Scientific), then equilibrated with 20 µL of solvent A. 5 µL of digests were loaded onto each Evotip. The Evotips were washed once with 20 µL of solvent A and kept wet with 100 µL of solvent A before LC-MS analysis. DIA-MS data acquisition Peptides samples were loaded onto Evosep One (EV-100 S00394, Evosep) coupled to a Orbitrap Exploris 480 mass spectrometry (MA10103C, Thermo Scientific) interfaced with a FAIMS PRO device. A 150um x 15cm PepMap RSLC C18 easy spray column (ES906 Thermo Scientific) was used as the analytical column on Evosep One system. An optimised data independent acquisition (DIA) method was used for MS data acquisition to maximise the depth of plasma proteome. The MS1 scan instrument setting: orbitrap resolution 120000; scan range (m/z): 345–1055; FAIMS CV (V): three sequential CVs − 40, -55, -70; Normalised AGC target (%): 300; Maximum injection time mode: Auto. The DIA MS2 scan the instrument setting: Orbitrap resolution: 30000; DIA window and m/z range: 38 variable DIA windows covering from 350 to 1050; FAIMS CV (V): three sequential CVs − 40, -55, -70 the same as MS1 FAIMS setting; RF length (%): 40; Normalised AGC target (%): 1000, Customed Maximum injection time (ms): 54; Data type: Profile; Loop control N = 13. Samples were run using Evosep One 15 sample per day method (gradient of 88min). Master pool samples were loaded as a QC after each plate row of samples. Data pre-processing In total, 714 DIA MS files were acquired including 614 lung cancer samples, 42 pooled quality control plasma samples and 58 master pool digests. Data was first converted to .htrms format using HTRMS converter software available with Spectronaut V18.5 (Biognosys). All 714 HTRMS files were loaded into Spectronaut V18.5 (Biognosys) for library-free directDIA analysis. Unless specified, default settings parameters were used. Default BSG factory settings were used for peptide/protein identification and quantification. Qvalue cutoff was 0.01 for both peptides and protein identification. LFQ method for protein quantitation at MS1 level and cross run global normalisation were enabled. Pulsar search workflow: directDIA+(Deep) was used with peptide fixed modification cysteine carbamidomethylation, and variable modification: acetyl (protein N-term) and oxidation (M). Two miss cleavages were allowed for trypsin cleavage. Human uniport FASTA file (UP000005640_9606_Hsapiens, Feb 2023) was used for in silico spectra library generation and protein inference. Proximity extension assay (PEA) data acquisition and processing Sample selection For PEA analysis, 600 samples were selected prioritising samples with low haemolysis score. The 600 selected samples were processed in two stages; an initial study of 88 samples (one plate), PEA 01, consisting of 29 healthy controls and 59 lung cancer cases, was followed by a larger cohort of 528 samples across six plates. PEA 02 consisted of 100 healthy controls and 428 lung cancer, with 16 bridging samples (6 healthy control, 10 LC) used across the two studies. Olink Explore 3072 Panel The Olink® Explore 3072/384 assay (referred to as PEA) consists of eight panels: Cardiometabolic, cardiometabolic II, inflammation, inflammation II, neurology, neurology II, oncology and oncology II. Each panel targets a maximum of 384 proteins, which quantify protein abundance using pairs of antibodies with complimentary oligonucleotide tags targeting the same protein. When both antibodies are bound to the protein of interest (i.e. in close proximity) the oligonucleotide tags hybridise, allowing DNA polymerase-dependent extension. The extended dsDNA is then tagged with barcodes and quantified using Next Generation Sequencing. Samples were randomised across the seven plates. Each assay is spiked with internal controls that act as incubation controls, extension controls and detection controls. Alongside internal controls, each plate contains sample controls, negative samples and inter-plate controls, used for plate normalisation, limit of detection (LOD) estimation and variation assessment respectively. The Olink® Explore 3072/384 assay was carried out by Randox Laboratories Ltd. Six µL of plasma from each sample was mixed with the antibodies from the eight panels and the resulting dsDNA specific for target protein extended, labelled with barcoded sequences and sequenced using NovaSeq 6000 Sequencing System (Illumina) for relative quantification. Data pre-processing Raw data was processed using NGS2COUNTS and imported into Olink® NPX software. Raw NGS reads were converted into normalised protein expression (NPX) values through a series of normalisation steps. Firstly, samples are normalised to extension controls and converted to Log 2 scale. Secondly, plate-control normalisation was performed by normalising the median of each sample to the median of plate controls. Intensity normalisation was subsequently performed, centering the median of each sample to the median of all samples (excluding controls) of the plate. These NPX values were used for all downstream analysis. DNA-aptamer data acquisition and pre-processing Sample selection A sub-cohort of the Princess Margaret patient cohort was selected for protein quantification on a third platform. This cohort consisted of 72 healthy control patients that were matched to 32 early lung cancer and 32 late lung cancer according to age, sex and smoking status (Extended Data Fig. 7a). DNA-aptamer quantification of feature subset A DNA-aptamer approach, SomaScan 11K V5.0 (SomaLogic), was used to assess feature concordance on a third proteomics technology. From the 11K SomaScan panel, 100 proteins from PEA and DIA-MS feature selection were chosen (Extended Data Fig. 7b). Samples were randomised across three plates according to cancer status (cancer-free, early lung cancer and late lung cancer). The SomaScan assays consist of SOMAmer probes composed of fluorophores, biotin and a photocleavable linker, which bind specific proteins. Once bound, biotinylated protein complexes are captured on Streptavidin beads and washed to remove unbound/weakly bound proteins. Photocleavage dissociates SOMAmers from protein complexes, non-specific complexes are dissociated and prevented from reforming using polyanionic competitors. Remaining SOMAmer-protein complexes are recaptured on streptavidin, eluted and binding of complementary DNA sequences to a microarray allows measurement of fluorescence. Data pre-processing Data processing of the SomaScan 11K V5.0 assay was performed using SomaScan in-house software. Results are reported in relative fluorescence units (RFU) and processed as follows: Firstly, using the 12 control SOMAmers spiked into each sample prior to microarray hybridisation normalisation was carried out. This was followed by intraplate median signal normalisation using plasma calibrator controls run alongside patient samples to account for dilution differences. Plate scaling was also carried out using calibrator samples, normalising the medians of calibrator reference ratios. Calibration was performed to normalise calibrator controls to reference values for each SOMAmer. Finally, Adaptive Normalisation by Maximum Likelihood (ANML) was used to normalise signal across samples. Briefly, SOMAmer reagents within a sample that are within two population standard deviations from reference are used to generate scaling factors; this was iterated over until convergence. The 11K panel was used for pre-processing while the data for the 100 chosen SOMAmers plus the 12 control SOMAmers were reported. Bioinformatic analysis Sample and protein filtering DIA-MS and PEA sample filtering are described in Extended Data Fig. 1a. Data filtering and quality assessment were performed in the R environment (version 4.2.2). For DIA-MS, any raw PG.Quantity values < 100 were replaced with NA. Samples with high haemolysis (≥ 2) were removed, along with samples not quantified by PEA. One clinical sample (cancer) that was later identified to be post-treatment was also removed. A minimum of 70% coverage across all remaining clinical samples was applied. PEA data from both studies (PEA 01 and PEA 02) were loaded and bridge-normalisation performed using the OlinkAnalyse package (version 3.8.2) available in R. Assays with QC warning were removed for all patients, as well as assays with incomplete data across the two studies. For bridging samples, data for PEA 02 was discarded and for assays with repeat measurements across panels, measurements with a warning were removed and the mean assay calculated across the panels. Patients with high haemolysis were removed to give the same population for comparison with DIA-MS. DNA-APT ADAT file was loaded using the SomaDataIO package (version 6.1.0) in R; all 140 samples passed Somalogics quality assessment and 20% dilution values were taken forward for further analysis. Data imputation and normalisation For non-ML related analysis, DIA-MS data for the 563 clinical samples were imputed using the MissForest algorithm with default parameters from the missingpy package (version 0.2.0, https://github.com/epsilon-machine/missingpy ) available in the Python environment and adapted to be compatible with the ML workflow. MissForest is a python implementation of the missForest package in R 34 that imputes data using random forests typically used for imputation of data that is considered missing at random (MAR) 34 . Assessment of missingness often relies on grouping of samples according to clinical characteristics. In this case where introduction of bias from imputation should be avoided and not be dependent on ML target, a target-independent model that can be used to impute unseen (unlabeled) data was needed. Given missForest has previously been shown to be a high performing imputation algorithm for missing data imputation in LC-MS based proteomics 35 and does not depend on knowing clinical characteristics of each instance, this method of imputation was selected for DIA-MS imputation. Following imputation, data was Log 2 transformed and normalised using median centering to account for differences in sample loading. Imputation was not required for PEA or DNA-APT data. The NPX values for PEA are already on the Log 2 scale, while DNA-APT data was Log 2 transformed. Post-analysis quality, missingness and dynamic range assessment For DIA-MS, CVs were calculated using pooled plasma samples run across seven plates without normalisation/Log2 transformation using standard CV formula (standard deviation/mean*100). For PEA, NPX data are both normalised and on Log 2 scale making the previous CV formula unsuitable 36 . For PEA and DNA-APT Log 2 transformed ANML values, the following formula was used: $$\:{CV}_{i}=100\bullet\:\sqrt{{e}^{{{Sln}_{i}}^{2}}-1}\:,\:\:\:where\:\:{Sln}_{i}=ln2\bullet\:{SD}_{isk}\:.\:$$ Missingness in DIA-MS data was assessed using all pre-treatment clinical samples with an NA for PG.Quantity. For PEA, missing data is not reported. Instead, values below the lower limit of detection (LLOD) were replaced with NA and classed as missing. Average missingness for each protein was reported as mean ± standard deviation. To assess dynamic range, known concentrations of proteins according to the Human Protein Atlas (HPA) (v24.0) 17 were mapped to UniProt IDs in PEA and DIA-MS data. For ProteinGroups in DIA-MS data the first protein was taken. Where blood concentration from mass spectrometry was not available, immunoassay known concentrations were used. “Secretome location” was used to identify secretion location of proteins predicted to be secreted. Prediction of secreted protein in HPA was performed using amino acid sequence and location derived from literature/available data 17 . Univariate and exploratory analysis Principal component analysis (PCA) was performed using prcomp() available in the stats package available in R (version 4.4.1). Default parameters were used except for scale = TRUE to equalise variance contribution of features and prevent bias towards high-variance features. Partial least squares regression (PLS) was used to distinguish lung cancer cases from healthy controls using the plsr() function from the pls package in R, with cross-validation (validation = “CV”) and scale = TRUE. T-distributed stochastic neighbourhood embedding (t-SNE) was performed using Rtsne from Rtsne (version 0.17) package in R. Uniform Manifold Approximation and Projection (UMAP) was performed using umap() from umap package (version 0.2.10.0) in R. For univariate analysis a linear model was fit to each protein using lmFit() from limma, followed by empirical Bayes adjustment using eBayes() from limma to improve estimation of variance 37 . Differentially expressed proteins were extracted using decideTests() with p.value = 0.05, lfc = 1 and adjustment.method = “fdr”. Summary statistics were extracted using topTable(). Proteins were considered differentially expressed if Log2 fold change > 1 and FDR-adjusted p-value < 0.05. Results are visualised as volcano plots (Fig. 2e,f and Extended Data Fig. 3c,d) and heatmaps (Supp Fig. 2B, 2C and Fig. 6A). Over-representation analysis (ORA) and gene set enrichment analysis (GSEA) 38 of Gene Ontology (GO) biological processes 39 were carried out using enirchGO() and gseGO() from clusterProfiler package (version 4.12.2) available in R 40 . The analysis was performed across sub-cohorts: For protein groups the first ID was used, pvalueCutoff = 0.05 and pAdjustMethod = “fdr”, human database was used (org.Hs.eg.db) and minGSsize = 5. Semantic similarity analysis was used to remove redundant terms, applied using the mgoSim() and clusterSim() functions available from clusterProfiler. Data was visualised as heatmaps. Spearman correlation across platforms was chosen due to it not being dependent on scale and calculated using cor() function in R for the 152 patient cohort quantified using all three methods. For assessing significance of correlation p -values were FDR-adjusted. For DNA-APT data a second method of normalisation was used (median centering) as an alternative to ANML. Spearman correlation was calculated on both median-centered and ANML normalised datasets, features with Spearman correlation R < 0.4 in the appropriate comparison were removed and only the intersection between the two normalisation methods were considered. This would account for bias potentially introduced following ANML normalisation 41 . Unsupervised, hierarchical clustering of differentially abundant proteins, ORA/GSEA results, Shapley value interactions and Spearman correlation co-efficient were performed using dist() and hclust() available in stats package in R and clusters generated using cutree() function in R. Data visualisation All data visualisations were generated in the R environment with the exception of networks. For heatmaps the heatmap.2() function from the gplots package was used with the exception of Fig. 4C visualised using ggplot2. Venn diagrams were generated using ggvenn. All other plots were generated using ggplot2. ML analysis for feature selection All ML data generation was performed using Python (version 3.10). All models undergo the same sample stratification, train/holdout split as well as normalisation and imputation. Following this, different combinations of model selection, feature de-selection, augmentation, feature selection and model optimisation are trained on training data only to give rise to the 75 different models produced for each sub-cohort of each dataset, as described in Extended Data Fig. 3a. Each step in the ML process is detailed below. Sample stratification and data split The 563 clinical samples that passed pre-processing were stratified using sex > age-range (10 year intervals) > stage (control vs early vs late LC) > histology (NSCLC vs SCLC) > smoking history (non-heavy vs heavy-smoker). An 80/20% split was chosen for generating train/holdout data; Training data was used for training models and analysis, while the 20% holdout dataset was used to assess model generalisability and evaluate feature importances. This single train/holdout split for healthy control vs lung cancer samples was used to generate sub-cohorts for all controls vs early only lung cancer patients (Stage 0–2), late lung cancer patients (Stage 3–4), NSCLC only, SCLC only, LUAD only and LUSC only. For male only and female only models, only male and female healthy controls and lung cancer patients were included, respectively. Using the same split for different cohorts allowed comparison of sub-cohort analyses as well as maintaining the unseen nature of the holdout dataset. Normalisation, imputation and scaling for ML analysis Data was median centered and Log2 transformed. missForest() from missingpy was used to impute training data as described previously. This model was fit on the training data only, and the subsequent model used to impute the training data and unseen holdout data. Fitting the imputation model on the training data only and applying to holdout avoids data-leakage between the two datasets and subsequent over-inflated ML performances. Similarly, scaling using StandardScaler() from scikit-learn (version 1.3.0, used for all scikit-learn functions) whereby the fit was applied to training only and used to transform train and holdout. Model selection Tree-based models are suitable for handling high-dimension datasets, maintaining feature interactions and capturing non-linear relationships, unlike linear models often used to build ML models on omics datasets. Extreme gradient boosting (XGBoost), a tree-based model, was chosen to classify healthy controls vs lung cancer (and sub-cohorts). XGBClassifier() function from the XGBoost Python model (version 1.7.6) was used. Balanced random forest (BRF) were chosen as an alternative tree-based model to XGBoost. BRF handles class imbalances by randomly under-sampling the majority class and maintaining balanced class representation during training, unlike random forests (RF) which favour the majority class. The use of bootstrapping in RF/BRF also prevents overfitting. BalancedRandomForestClassifier() from the Python package imblearn (version 0.12.2, https://github.com/scikit-learn-contrib/imbalanced-learn ) was used to build BRF models. Hyperparameters were determined during model-optimisation. Feature de-selection Feature de-selection is the process by which features with low information content about the target (control vs lung cancer) are removed, with the aim of improving ML performance and reducing overfitting. Two methods of feature de-selection were used along with no feature de-selection. The first method uses mutual information (MI) that assesses the dependency between two variables (cancer status and protein abundance) that is capable of capturing non-linear relationships. The mutual_info_classif() was used to compute MI and was cross-validated using RepeatedStratifiedKFold(), stratified according to target (healthy control vs LC) with n_splits = 2 and n_repeats = 25. Any feature with cross-validated MI less than 3.5% was considered to not contain information related to classification, in this case lung cancer case, and was subsequently removed from the feature panel. Alternatively, BorutaShap a two-step algorithm that combines the Boruta Algorithm 42 with Shapley values 43 was used to perform feature de-selection. Boruta works by creating shadow features by random permutation using actual features. By comparing feature importances of shadow and actual features, features where shadow features out-perform can be removed. In the case of BorutaShap, feature importances are calculated using Shapley values for each sample rather than model importances 18 . Shapley values were derived from cooperative game theory and applied to ML models to assess feature importance. Shapley values fairly attribute a "reward" among features based on their individual contributions to the prediction. This approach was implemented using the BorutaShap function from the BorutaShap package (version 1.0.17, ( https://github.com/Ekeany/Boruta-Shap ), using an XGBoost model with importance_measure = “shap”. Augmentation In some models augmentation was applied to balance classes between healthy controls and lung cancer cases by up-sampling the minority class. A two-step process was implemented to augment the minority class. Firstly, synthetic minority over-sampling technique (SMOTE) was used to generate synthetic samples from the minority class. SMOTE selects a random point from the minority class and interpolates a new datapoint between it and a neighbour actual sample, maintaining correlations/relationships between features in the synthetic data 44 . Tomek links are then used to remove a portion of the majority class with overlapping decision boundaries with the augmented minority class. SMOTE-Tomek augmentation was applied to training data following feature de-selection using the SMOTETomek() function available in imbalanced-learn (version 0.12.2). Recursive feature addition (RFA) with Shapley values Proteomics models that require large feature sets have limited clinical utility if needing to be translated onto alternative platforms. For example, immunoassay approaches require small panels for single-plex assays (such as ELISA) up to tens of analytes for multiplexed assays. Alternatively targeted mass-spectrometry methods such as single-reaction, multiple-reaction and parallel reaction monitoring (SRM, MRM and PRM, respectively) can be used, however such methods depending on instrumentation can also be limited to as little as 15 peptides 45 . Recursive feature addition (RFA) was used to build models that rely on a small subset (< 30) of proteins. RFA first selected the highest importance feature according to cross validated Shapley values. Using mean absolute Shapley value, high importance features were sequentially added and model performance evaluated to select the smallest subset of features that achieves the highest performance. RFA was coupled with repeated stratified cross validation (RSCV), as described for BorutaShap, to give more robust performance estimates across different feature sets. Traditionally, model feature importances (such as those that come from XGBoost or BRF) are used to rank features and determine order of feature addition in RFA. The use of Shapley values instead of model feature importances captures both global (overall) and local (sample-specific) attributions, as opposed to only global. This ensures a fair and consistent distribution of feature importance and fairly distributes importance among correlated features. Another key advantage is their ability to account for feature interactions, revealing how features work together to give a prediction rather than evaluating them in isolation. Shapley values were calculated using the TreeExplainer() function from the Shap Python package, a fast implementation of tree-based Shapley value (TreeSHAP) generation using interventional perturbations to maintain dependencies between features. As an alternative Shapley value method, the package EjectSHAP was also used 46 . EjectSHAP works similarly to TreeSHAP with the exception of how it attributes importances to features that never contributed to a prediction; if a branch or single node in a tree is not used to make a prediction, the importance of features within that tree are zero (i.e. “ejected”). Two metrics are used to evaluate model performance, area under the receiver operator characteristic curve (AUROC) and Matthews correlation coefficient (MCC). These metrics were evaluated on the top 30 features in XGBoost models and top 50 features in BRF, which often required higher numbers of features to converge during RFA. ROC curves plot the true positive rate (TPR/sensitivity) vs false positive rate (FPR) at different thresholds. AUC measures the area under the ROC curve; higher AUROC indicates higher model performance. MCC on the other hand utilises true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN), giving a more meaningful performance in cases where data is imbalanced. MCC can be interpreted as the correlation between the true labels and the predicted labels for binary classification problems 47 , 48 . MCC was calculated using the matthews_corrcoef() function in scikit-learn, using the following formula: \(\:MCC\:=\) \(\:\frac{TP\bullet\:TN-FP\bullet\:FN}{\sqrt{(TP+FP)\bullet\:(TP+FN)\bullet\:(TN+FP)\bullet\:(TN+FN)}}\:.\:\) MCC and AUROC performance gives rise to five different models. The first maximising mean AUROC across folds in RFA, the second maximising mean MCC across folds, the third approach maximises the stability of MCC across folds. MCC stability was calculated by dividing the mean of MCC across folds by the standard deviation across folds. This will give a subset of features for which MCC were minimally impacted by the fold used in RSCV. Finally, a ttest MCC and ttest AUROC feature set are generated to give a smaller subset of features, using ttest_ind_from_stats() available in scipy (version 1.11.1). In some cases, mean AUROC, MCC or MCC stability, or ttest MCC and AUROC might converge on the same subset of features and therefore produce less than five models, while in some cases all five RFA performance metrics give rise to models with different feature sets. Model optimisation XGBoost models require fine-tuning of model hyperparameters to avoid overfitting/underfitting on training data. For this, NSGAIISampler() from Optuna (version 3.6.1) Python module was used for multi-objective optimisation of model hyperparameters. RFA using baseline XGBoost or BRF hyperparameters were used to select top 30 features. Baseline hyperparameters were chosen to mitigate the overfitting potential of the models. For the XGBoost models we limited the max_depth to 2, the n_estimators to 200, the learning rate to 0.05, lambda to 0, and setting the scale_pos_weight as the square root of the ratio of number of controls to cancer cases. For the BRF models we limited the max_depth to 10, the n_estimators to 500 and used sampling without replacement. A model using these features was then optimised in Optuna to maximise three combinations of model performance metrics: MCC only, MCC + sensitivity at 99% specificity (Sens@90%Spec) and MCC + AUROC. If followed by optimisation, RFA was repeated using the optimised model hyperparameters to give a selection of features as described previously. Optuna was selected over other optimisation tools (such as GridSearchCV from Scikit-learn) as a result of its efficiency in searching higher-dimensional parameters spaces due to its Bayesian optimisation underpinning. Optimised model-hyperparameters were chosen from a local neighbourhood of similarly performing models, improving the stability of model sensitivity to minor hyperparameter perturbations. Model evaluation Following training of the models, each model was evaluated using the unseen holdout test set. Assessment of model generalisability, rather than performance, promoted the refinement of features based on similar performance across train/holdout data with the aim of improving generalisability in future unseen data. To do this in a high throughput manner (1,350 models to evaluate), 95% confidence intervals (CI) were generated for the training data during model training. These 95% CI evaluated model performance during training using RSCV with n_splits = 4 (equivalent in size to holdout data) and n_repeats = 25. 95% CI were calculated for the following performance metrics: MCC, AUROC, Sens@90%Spec, Sens@95%Spec and Sens@99%Spec. These specificities were chosen to ensure reliable prediction in a low-prevalence setting, where high specificity reduces FPRs while maintaining high TPR. If four out of five of the performance metrics in holdout were within the 95% CI of training data a model was considered to have generalised, and was taken forward for further interpretation. Model interpretability (XAI) Shapley values for each model were calculated as described during RFA. By calculating the mean of absolute Shapley values the contribution of each feature to all predictions can be interpreted. Within each generalisable model, features were ranked from largest mean absolute Shapley value to lowest. The average rank of a feature across generalisable models was used to understand the most important features for each sub-cohort. All data used in plotting Shapley value-related data comes from the SHAP module in Python. To assess the interactions between features, rather than the importances of a single feature, Shapley value interactions were used. Shapley value interactions capture pairwise interactions by assessing how the presence of one feature influences the contribution of a different feature to model prediction. Shapley value interactions aim to improve model-interpretability while identifying synergies and redundancies between features 39 . For a single representative healthy control vs lung cancer model, Shapley value interactions were calculated using the shap_interaction_values method from the TreeExplainer class from the SHAP module in Python. High values for Shapley value interaction equate to a higher proportion of Shapley value explained by the interaction between the two features. Shapley value interactions were normalised to the sum of all Shapley value interactions for both DIA-MS and PEA data to account for variation in scale of Shapley values between the two datasets. Normalised Shapley value interactions are visualised using a heatmap (Extended Data Fig. 6c,d), while a network of interactions between features accounting for more than 5% of actual Shapley value was visualised using Cytoscape (version 3.10.3) 49 (Fig. 4d,e). ML analysis for cross-platform concordance ML models were used to assess the ability of features with concordance across the three platforms to predict lung cancer using DIA-MS and PEA data. Two panels were fit to both datasets; the first panel contained 18 features that were concordant (definition of concordance used: Spearman R 2 > 0.4) across DIA-MS, PEA and DNA-APT. The second panel for DIA-MS were features concordant between DIA-MS and DNA-APT that if present were also concordant in PEA, and for PEA those that were concordant between PEA and DNA-APT that if present were also concordant in DIA-MS. This accounted for non-overlapping features between DIA-MS and PEA. A logistic regression was chosen as the model for assessing concordance due to its interpretability and low number of parameters to fit, allowing easier comparison of performances across datasets and/or features. The LogisticRegressionCV class from Scikit-learn was used to fit a balanced logistic regression on training data from PEA/DIA-MS data using relevant features. Performances were evaluated on holdout data not used in training. Comparing univariate feature selection with ML-based feature selection To compare ML-dependent feature selection with a univariate approach, as has been used in other studies 19 , 50 , an analysis of variance (ANOVA) was used to select features had the smallest FDR-adjusted p-values when comparing lung cancer cases with healthy controls in all data. Twenty features were selected for each cohort that had the smallest adjusted p-values according to an f_test from Scikit-learn and compared with the twenty features that had the highest Shapley value rank in that cohort. A balanced logistic regression was used to assess the ability of these features to predict lung cancer as previously described. MCC and AUROC from the two logistic regression models were compared with their respective metric from an XGBoost model (Extended Data Fig. 3e). Declarations Data and code availability All data analysed and code used in this study is available upon request. Acknowledgements We thank all participants in the CALIBRE study. We also thank members of the Princess Margaret Cancer Centre, Toronto General Hospital and Princess Margaret Hospital for their contribution to the study. We would like to thank members of the Target Discovery Institute, in particular Dr Iolanda Vendrell for their technical advice in this study. We thank team members at Randox for their support collecting and processing the PEA (Olink) data. We would like to thank all former and current employees of Oxford Cancer Analytics Ltd for their guidance in the preparation of this manuscript, especially Dr Honglei Huang and Dr Heinrich Roder. B.M.K. and R.F. were supported by the Chinese Academy of Medical Sciences (CAMS) Innovation Fund for Medical Science (CIFMS), China (grant number: 2024-I2M-2-001-1). Author contributions H.R.F and N.G performed data analysis and developed the ML workflow, produced all figures and tables and wrote the text of the manuscript. L.H developed the ML workflow and contributed to interpretation of the data and data analysis. I.K developed and optimised the DIA-MS workflow and experimental design, carried out sample preparation, data acquisition and pre-processing of DIA-MS dataset. Ella.M and Emma.M contributed to data analysis, interpretation including cohort selection and stratification. D.A.S and J.S contributed to conceptualisation of the study, cohort selection and experimental design. G.L, L.J.Z and D.P provided insights into clinical cohorts, sourced the clinical samples and performed clinical data abstraction in this study. M.D contributed to proteomic sample selection and preparation. B.K and R.F provided insights into proteomic methodologies and analysis. A.H contributed to conceptualisation and methodology of the study. P.J.L conceptualised, supervised and guided the design, analysis, interpretation of the study, and wrote the text of the manuscript. All authors contributed to the drafting of the manuscript. Competing interests H.R.F, N.G, L.H, I.K, M.D, J.S, Emma.M, Ella.M, D.A.S, A.H, and P.J.L are current/former employees, shareholders, and/or share option holders of Oxford Cancer Analytics Ltd. R.F and B.M.K are share option holders of Oxford Cancer Analytics Ltd. Oxford Cancer Analytics Ltd sponsored this study. References World Health Organisation. Lung cancer. World Health Organisation https://www.who.int/news-room/fact-sheets/detail/lung-cancer (2023). American Cancer Society. Lung Cancer Survival Rates. https://www.cancer.org/cancer/types/lung-cancer/detection-diagnosis-staging/survival-rates.html https://www.cancer.org/cancer/types/lung-cancer/detection-diagnosis-staging/survival-rates.html (2024). Reduced Lung-Cancer Mortality with Low-Dose Computed Tomographic Screening. New England Journal of Medicine 365, (2011). de Koning, H. J. et al. Reduced Lung-Cancer Mortality with Volume CT Screening in a Randomized Trial. New England Journal of Medicine 382, (2020). Jemal, A. & Fedewa, S. A. Lung cancer screening with low-dose computed tomography in the United States – 2010 to 2015. JAMA Oncol 3, (2017). Dickson, J. L. et al. Uptake of invitations to a lung health check offering low-dose CT lung cancer screening among an ethnically and socioeconomically diverse population at risk of lung cancer in the UK (SUMMIT): a prospective, longitudinal cohort study. Lancet Public Health 8, (2023). LoPiccolo, J., Gusev, A., Christiani, D. C. & Jänne, P. A. Lung cancer in patients who have never smoked — an emerging disease. Nature Reviews Clinical Oncology vol. 21 Preprint at https://doi.org/10.1038/s41571-023-00844-0 (2024). Chin, R. I. et al. Detection of Solid Tumor Molecular Residual Disease (MRD) Using Circulating Tumor DNA (ctDNA). Molecular Diagnosis and Therapy vol. 23 Preprint at https://doi.org/10.1007/s40291-019-00390-5 (2019). Chaudhuri, A. A. et al. Early detection of molecular residual disease in localized lung cancer by circulating tumor DNA profiling. Cancer Discov 7, (2017). Chabon, J. J. et al. Integrating genomic features for non-invasive early lung cancer detection. Nature 580, (2020). Kelly-Spratt, K. S. et al. Plasma proteome profiles associated with inflammation, angiogenesis, and cancer. PLoS One 6, (2011). Pitteri, S. J. et al. Tumor microenvironment-derived proteins dominate the plasma proteome response during breast cancer induction and progression. Cancer Res 71, (2011). Bhardwaj, M., Terzer, T., Schrotz-King, P. & Brenner, H. Comparison of proteomic technologies for blood-based detection of colorectal cancer. Int J Mol Sci 22, (2021). Ng, S., Masarone, S., Watson, D. & Barnes, M. R. The benefits and pitfalls of machine learning for biomarker discovery. Cell and Tissue Research vol. 394 Preprint at https://doi.org/10.1007/s00441-023-03816-z (2023). Shen, S. Y. et al. Sensitive tumour detection and classification using plasma cell-free DNA methylomes. Nature 563, (2018). Khodayari Moez, E. et al. Circulating proteome for pulmonary nodule malignancy. JNCI: Journal of the National Cancer Institute 115, (2023). Uhlén, M. et al. The human secretome. Sci Signal 12, (2019). Lundberg, S. M. et al. From local explanations to global understanding with explainable AI for trees. Nat Mach Intell 2, (2020). Liang, H. et al. LcProt: Proteomics-based identification of plasma biomarkers for lung cancer multievent, a multicentre study. Clin Transl Med 15, e70160 (2025). Müller, S. et al. An Empirical Evaluation of the Rashomon Effect in Explainable Machine Learning. ArXiv (2023) doi: https://doi.org/10.48550/arXiv.2306.15786 . Katz, D. H. et al. Proteomic profiling platforms head to head: Leveraging genetics and clinical traits to compare aptamer- And antibody-based methods. Sci Adv 8, (2022). Birhanu, A. G. Mass spectrometry-based proteomics as an emerging tool in clinical laboratories. Clinical Proteomics vol. 20 Preprint at https://doi.org/10.1186/s12014-023-09424-x (2023). Eldjarn, G. H. et al. Large-scale plasma proteomics comparisons through genetics and disease associations. Nature 622, (2023). Petrera, A. et al. Multiplatform Approach for Plasma Proteomics: Complementarity of Olink Proximity Extension Assay Technology to Mass Spectrometry-Based Protein Profiling. J Proteome Res 20, (2021). Weber, M. A. et al. Myoglobin plasma level related to muscle mass and fiber composition - A clinical marker of muscle wasting? J Mol Med 85, (2007). Czarnowski, D. Syndecans in cancer: A review of function, expression, prognostic value, and therapeutic significance. Cancer Treatment and Research Communications vol. 27 Preprint at https://doi.org/10.1016/j.ctarc.2021.100312 (2021). Dinh, H. Q. et al. Coexpression of CD71 and CD117 Identifies an Early Unipotent Neutrophil Progenitor Population in Human Bone Marrow. Immunity 53, (2020). Daniels, T. R., Delgado, T., Rodriguez, J. A., Helguera, G. & Penichet, M. L. The transferrin receptor part I: Biology and targeting with cytotoxic antibodies for the treatment of cancer. Clinical Immunology vol. 121 Preprint at https://doi.org/10.1016/j.clim.2006.06.010 (2006). Human Protein Atlas. HAGH: Cancer. HAGH https://www.proteinatlas.org/ENSG00000063854-HAGH/cancer (2024). Yang, W. & Rubin, J. B. Treating sex and gender differences as a continuous variable can improve precision cancer treatments. Biol Sex Differ 15, 35 (2024). Budnik, B., Amirkhani, H., Forouzanfar, M. H. & Afshin, A. Novel proteomics-based plasma test for early detection of multiple cancers in the general population. BMJ Oncology 3, (2024). Bujang, M. A. & Adnan, T. H. Requirements for Minimum Sample Size for Sensitivity and Specificity Analysis. JOURNAL OF CLINICAL AND DIAGNOSTIC RESEARCH 10, YE01–YE06 (2016). Searfoss, R. et al. Impact of hemolysis on multi-OMIC pancreatic biomarker discovery to derisk biomarker development in precision medicine studies. Sci Rep 12, (2022). Stekhoven, D. J. & Bühlmann, P. Missforest-Non-parametric missing value imputation for mixed-type data. Bioinformatics 28, (2012). Jin, L. et al. A comparative study of evaluating missing value imputation methods in label-free proteomics. Sci Rep 11, (2021). Canchola, J. A. Correct Use of Percent Coefficient of Variation (%CV) Formula for Log-Transformed Data. MOJ Proteom Bioinform 6, (2017). Ritchie, M. E. et al. Limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res 43, (2015). Subramanian, A. et al. Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci U S A 102, (2005). Aleksander, S. A. et al. The Gene Ontology knowledgebase in 2023. Genetics 224, (2023). Wu, T. et al. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innovation 2, (2021). Pietzner, M. et al. Synergistic insights into human health from aptamer- and antibody-based proteomic profiling. Nat Commun 12, (2021). Kursa, M. B. & Rudnicki, W. R. Feature selection with the boruta package. J Stat Softw 36, (2010). Lundberg, S. M. & Lee, S. I. Lundberg, S.M., Lee, S.I.: A unified approach to interpreting model predictions. In: Advances in Neural Information Processing Systems. pp. 4765–4774 (2017). NIPS-2017 Advances in Neural Information Processing Systems 32, (2017). Chawla, N. V., Bowyer, K. W., Hall, L. O. & Kegelmeyer, W. P. SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research 16, (2002). Van Bentum, M. & Selbach, M. An introduction to advanced targeted acquisition methods. Molecular and Cellular Proteomics vol. 20 Preprint at https://doi.org/10.1016/J.MCPRO.2021.100167 (2021). Campbell, T. W., Roder, H., Georgantas III, R. W. & Roder, J. Exact Shapley values for local and model-true explanations of decision tree ensembles. Machine Learning with Applications 9, (2022). Chicco, D. & Jurman, G. The Matthews correlation coefficient (MCC) should replace the ROC AUC as the standard metric for assessing binary classification. BioData Min 16, (2023). Chicco, D. & Jurman, G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics 21, (2020). Shannon, P. et al. Cytoscape: A software Environment for integrated models of biomolecular interaction networks. Genome Res 13, (2003). Davies, M. P. A. et al. Plasma protein biomarkers for early prediction of lung cancer. EBioMedicine 93, (2023). Additional Declarations Yes there is potential Competing Interest. H.R.F, N.G, L.H, I.K, M.D, J.S, Emma.M, Ella.M, D.A.S, A.H, and P.J.L are current/former employees, shareholders, and/or share option holders of Oxford Cancer Analytics Ltd. R.F and B.M.K are share option holders of Oxford Cancer Analytics Ltd. Oxford Cancer Analytics Ltd sponsored this study. Supplementary Files SupplementaryTable1.xlsx Data Set 1 SupplementaryTable2.xlsx Data Set 2 SupplementaryTable3.xlsx Data Set 3 SupplementaryTable4.xlsx Data Set 4 ExtendedData.docx Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7660411","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":523799476,"identity":"c3a35ca4-493f-48b9-892a-ed523b996589","order_by":0,"name":"Peter Jianrui Liu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5UlEQVRIiWNgGAWjYNCCChsZNgbmA0CWhAwRypmB+EwaDxsDWwJICw9xWhjbDgNV8hiAuIS1mLP3H3xc2cbMwyeR8/nVjRoLHgb2w0c34NNi2XOY2fDMOTYeNoncbdY5x4AO40lLu4FPi8GNZDbJhjIesBbjHDagFgkeM/xa7j9m/9kAVMkmkfPMOOcfMVpuMLMxNrQZgLQwP85tI0KLZU+ysWTDmQQeNp5nZsy5fUDrCPnFnP3gw48NFf/l5NuTH3/O+VYnx89++Bh+hyGx2STAJD7l6FqYPxBSPQpGwSgYBSMTAAD3WD+y+Pkb1QAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-3608-750X","institution":"Oxford Cancer Analytics","correspondingAuthor":true,"prefix":"","firstName":"Peter","middleName":"Jianrui","lastName":"Liu","suffix":""},{"id":523799477,"identity":"13213f87-2930-4d3c-af9f-65d65bd32229","order_by":1,"name":"Harriet Ferguson","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Harriet","middleName":"","lastName":"Ferguson","suffix":""},{"id":523799478,"identity":"e4bfad95-655d-4853-9a21-e82b0a167046","order_by":2,"name":"Nikola Gushterov","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Nikola","middleName":"","lastName":"Gushterov","suffix":""},{"id":523799479,"identity":"df9e79fb-3970-4d5b-a765-a2c0932599b8","order_by":3,"name":"Benedikt Kessler","email":"","orcid":"https://orcid.org/0000-0002-8160-2446","institution":"University of Oxford","correspondingAuthor":false,"prefix":"","firstName":"Benedikt","middleName":"","lastName":"Kessler","suffix":""},{"id":523799480,"identity":"2c445413-7d5a-4340-b3e5-04b3a6ef7572","order_by":4,"name":"Geoffrey Liu","email":"","orcid":"","institution":"University Health Network / Princess Margaret Cancer Centre","correspondingAuthor":false,"prefix":"","firstName":"Geoffrey","middleName":"","lastName":"Liu","suffix":""},{"id":523799481,"identity":"b7a9ba46-bebd-4cd1-b4ea-27c8edb1aa44","order_by":5,"name":"Andreas Halner","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Andreas","middleName":"","lastName":"Halner","suffix":""},{"id":523799482,"identity":"975e4863-033f-418c-842b-f8a474457f25","order_by":6,"name":"Luke Hankey","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Luke","middleName":"","lastName":"Hankey","suffix":""},{"id":523799484,"identity":"47065512-beb0-4ed3-9b7c-28f3b6df89a1","order_by":7,"name":"Iliyana Kaneva","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Iliyana","middleName":"","lastName":"Kaneva","suffix":""},{"id":523799485,"identity":"21215565-7a57-471a-93fd-30b329d8be97","order_by":8,"name":"Mrunmayee Dupalliwar","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Mrunmayee","middleName":"","lastName":"Dupalliwar","suffix":""},{"id":523799486,"identity":"cd077a83-3a6e-4f22-957e-bbedf722cf6e","order_by":9,"name":"Junetha Syed","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Junetha","middleName":"","lastName":"Syed","suffix":""},{"id":523799487,"identity":"51966f71-d669-4d4d-803f-b53f4d086f10","order_by":10,"name":"Emma Mi","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Emma","middleName":"","lastName":"Mi","suffix":""},{"id":523799488,"identity":"51eb7311-a4b8-417d-9826-9ce8109354ec","order_by":11,"name":"Ella Mi","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Ella","middleName":"","lastName":"Mi","suffix":""},{"id":523799489,"identity":"85d126f2-6933-473c-b655-75a93666223b","order_by":12,"name":"Daniel Szulc","email":"","orcid":"","institution":"Oxford Cancer Analytics","correspondingAuthor":false,"prefix":"","firstName":"Daniel","middleName":"","lastName":"Szulc","suffix":""},{"id":523799490,"identity":"3522d270-61e2-4839-a8aa-0b0e95606e7e","order_by":13,"name":"Devalben Patel","email":"","orcid":"","institution":"Princess Margaret Cancer Centre","correspondingAuthor":false,"prefix":"","firstName":"Devalben","middleName":"","lastName":"Patel","suffix":""},{"id":523799491,"identity":"71e4356f-9260-4a03-95f4-abdb2a2ae08c","order_by":14,"name":"Luna Zhan","email":"","orcid":"","institution":"University Health Network / Princess Margaret Cancer Centre / University of Toronto","correspondingAuthor":false,"prefix":"","firstName":"Luna","middleName":"","lastName":"Zhan","suffix":""},{"id":523799492,"identity":"56bd5ed2-8229-47f9-9b3f-659156e4668b","order_by":15,"name":"Roman Fischer","email":"","orcid":"https://orcid.org/0000-0002-9715-5951","institution":"University of Oxford","correspondingAuthor":false,"prefix":"","firstName":"Roman","middleName":"","lastName":"Fischer","suffix":""}],"badges":[],"createdAt":"2025-09-19 16:45:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7660411/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7660411/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":94847683,"identity":"923d8bd0-210f-4d0c-900a-ba7e225bf5fc","added_by":"auto","created_at":"2025-10-31 10:31:18","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":415257,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCharacterisation of the blood plasma proteome using DIA-MS and PEA. a\u003c/strong\u003e, Proteomics and clinical workflow. \u003cstrong\u003eb\u003c/strong\u003e, Overlap in proteins identified by DIA-MS (blue) and PEA (yellow). \u003cstrong\u003ec\u003c/strong\u003e, Distribution of proteins with known concentrations according to Human Protein Atlas (HPA), proteins with concentration below 0.1 mg/L (shaded blue) are considered in the mid-low abundance range. \u003cstrong\u003ed\u003c/strong\u003e, Concentration vs missingness (%) of proteins with known concentration in HPA showing Spearman correlation with FDR-adjusted p-value.\u003cstrong\u003e e\u003c/strong\u003e, Inter-plate CV (%) of proteins quantified by both DIA-MS and PEA across pooled plasma samples. Lines between boxplots link matching protein IDs. \u003cstrong\u003ef\u003c/strong\u003e, Boxplot showing relationship between median protein quantification by DIA-MS and strength of Spearman correlation of proteins quantified by DIA-MS and PEA. Spearman correlation is shown as R\u003csup\u003e2\u003c/sup\u003e, and divided into four groups. * FDR-adjusted \u003cem\u003ep-\u003c/em\u003evalue \u0026lt; 0.05. \u003cstrong\u003eg\u003c/strong\u003e, Heatmap of hierarchical clustering (Manhattan average) of the top 10 over-represented gene ontology biological pathways in proteins quantified by DIA-MS only (DIA-MS), PEA only (PEA) or quantified by both DIA-MS and PEA (BOTH). Top 10 ORA colour key indicates which analysis identified pathway as in top 10 overrepresented pathways. Purple = top 10 ORA pathway in DIA-MS only proteins, orange Top 10 pathways only in proteins quantified by both DIA-Ms and PEA, blue = top 10 pathway in proteins quantified by PEA only; pink = top 10 pathway in proteins quantified by PEA only and quantified by both PEA and DIA-MS. * indicates significant over-representation (FDR adjusted p-value \u0026lt; 0.05).\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/3ec2087c19b052af1356b809.png"},{"id":94985874,"identity":"eeaffcb7-b93a-4f74-8ec8-6381acda265e","added_by":"auto","created_at":"2025-11-03 06:59:09","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":417741,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCharacterisation of blood plasma proteins associated with lung cancer by stage and common histological subtypes. \u003c/strong\u003ePCA of patient samples using QC qualified DIA-MS quantification (\u003cstrong\u003ea\u003c/strong\u003e) or PEA quantification (\u003cstrong\u003eb\u003c/strong\u003e) of proteins across patients. Blue = control, dark pink = early lung cancer (LC) and light pink = late LC. Ellipsis drawn assuming multivariate t-distribution, confidence level = 0.95. Loadings for PCA of DIA-MS (\u003cstrong\u003ec\u003c/strong\u003e) and PEA (\u003cstrong\u003ed\u003c/strong\u003e); highly separated genes are labelled. Volcano plot of DIA-MS (\u003cstrong\u003ee\u003c/strong\u003e) and PEA (\u003cstrong\u003ef\u003c/strong\u003e) protein differential expression in lung cancer patients compared with control patients. Blue = decreased expression in lung cancer patients compared with controls, red = increased expression in lung cancer patients compared with controls, light grey = not differentially expressed in lung cancer patients compared with controls. Criteria for differential expression marked by dashed line; proteins with FDR-adjusted p-value \u0026lt; 0.05 and Log2 fold change (LogFC) \u0026gt; 1 are considered differentially expressed. \u003cstrong\u003ee\u003c/strong\u003e, Heatmap of biological processes differentially regulated according to Gene set enrichment analysis (GSEA) of proteins within different patient sub-cohorts. Rows and columns ordered according to hierarchical clustering (Manhattan average) of Normalised enrichment score in sub-cohort compared with controls, identifying six clusters of pathways. Pathways GSEA clustering. SCLC = Small cell lung cancer, LUSC = Lung Squamous Cell Carcinoma, LC = All lung cancer, ELC = Early (Stage I-II) lung cancer patients only, LLC = Late (stage III-IV) lung cancer patients only, LUAD = Lung Adenocarcinoma patients only.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/9ab1c47f945186319057fb6a.png"},{"id":94985450,"identity":"a22ccd38-01f7-4fa8-b26b-f127547a1328","added_by":"auto","created_at":"2025-11-03 06:58:12","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":567972,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExplainable machine learning identifies focused plasma molecular signature for lung cancer.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e, ML workflow used to classify lung cancer and subtypes using proteomic plasma signatures. Models were trained on 80% of stratified samples with a 20% unseen holdout test set. Sub-cohorts with the same sample stratification and train/holdout split applied to both DIA-MS and PEA datasets. Feature engineering was used to reduce feature lists, followed by feature selection using recursive feature addition (RFA) with repeated stratified cross validation (RSCV) to produce a refined set of features. The model (a reduced set of features with model hyperparameters) was trained on the training data and applied to holdout data. Model performance consistency between training data and holdout data was used to assess generalisability of models, using explainable AI (XAI), artifacts and performances. T = Training data only, H = Holdout data only. LC = all lung cancer, NSCLC = non-small cell lung cancer, SCLC = small cell lung cancer, LUAD = lung adenocarcinoma, LUSC = lung squamous cell carcinoma, ELC = early lung cancer, LLC = late lung cancer. Area under the receiver operator characteristic curve (AUROC) of all DIA-MS models (\u003cstrong\u003eb\u003c/strong\u003e) and PEA models (\u003cstrong\u003ec\u003c/strong\u003e), with 95% confidence intervals (CI) across all models (shaded area) and mean performance (solid line) for all sub-cohort comparisons as indicated by colour key. Mean MCC with 95% CI for all sub-cohort models for DIA-MS (\u003cstrong\u003ed\u003c/strong\u003e) and PEA (\u003cstrong\u003ee\u003c/strong\u003e). \u003cstrong\u003ef\u003c/strong\u003e, Dot plot showing frequency of top 20 most consistent features in the 75 control vs lung cancer models evaluated for DIA-MS (left) and PEA (right) datasets. Genes are ordered from high to low (top to bottom) global mean of rank based on MAD (mean absolute deviation) of Shapley value. Nodes are coloured according to mean rank of MAD of Shapley value within five sub-models (mean MCC, mean AUROC, stability of MCC, ttest AUROC and ttest MCC) assessed with each feature engineering, selection and model combination (grid below plot). + indicates use in model training; - indicates absence in model training; X: XGBoost; B: Balanced random forest; MI: Mutual information; B: BorutaSHAP; T: TreeSHAP; E: EjectSHAP; Sens@99: Sensitivity at 99% specificity; AUROC, Sens@99 and MCC refer to multi-objective optimisation. \u003cstrong\u003eg\u003c/strong\u003e, Plot showing change in the representative DIA-MS (left) and PEA (right) model mean ROC AUC (blue) and mean MCC (purple) during RFA+RSCV as additional features are considered in the model.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/8506139e5ae1d240d9f06466.png"},{"id":94847685,"identity":"21999250-6d08-4216-abd4-04549275a17c","added_by":"auto","created_at":"2025-10-31 10:31:18","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":524896,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExplainable AI integrating Shapley values identifies key features and pathways associated with lung cancer detection using plasma proteins. a\u003c/strong\u003e, Heatmap of differentially expressed (*; LogFC \u0026gt; 0.5 and FDR adjusted p-value \u0026lt; 0.05) features quantified by both DIA-MS (upper panel) and PEA (lower panel). Each row of heatmap is a different sub-cohort comparison with controls. Colour indicates Log2 fold change (LogFC) of sub-cohort compared with appropriate controls. Column colour shows if a lung cancer feature was used by DIA-MS lung cancer models only (pale turquoise), PEA lung cancer models only (pale yellow) or used by both datasets (pale purple) for lung cancer models. \u003cstrong\u003eb\u003c/strong\u003e, Average rank of absolute Shapley value within individual models for DIA-MS (upper) and PEA (lower) for each sub-cohort comparison. High average rank = on average higher Shapley value within models. Missing points indicate that a feature was never selected in generalisable models. Features are ordered on x-axis, from left to right, according to average rank in all lung cancer models, followed by sub-cohorts. Only the top 20 features according to rank of absolute Shapley value from each sub-cohort are shown.\u003cstrong\u003e c\u003c/strong\u003e,\u003cstrong\u003e \u003c/strong\u003eHeatmap of biological pathways over-represented in all DIA-MS and PEA selected features. Top 20 features were taken from all models; top 20 pathways are shown. ORA was performed on all features. Each cell represents the number of features assigned to each pathway. Order of pathways is based on hierarchical clustering of number of features within each pathway to group similar pathways/patterns together. Shapley value interaction network built using Shapley values from a DIA-MS (\u003cstrong\u003ed\u003c/strong\u003e) and PEA (\u003cstrong\u003ee\u003c/strong\u003e) control vs lung cancer model. Shapley value interaction networks were built using only interactions between two proteins that explain 5% of the average absolute Shapley value within a model. Dark red edge = high relative Shapley interaction, pale pink = low relative Shapley value interaction. Nodes are coloured according to each protein’s normalised absolute Shapley value. Dark purple node = high absolute Shapley value, light blue node = low absolute Shapley value.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/f2c1c5c9c6fdd548b9c22bff.png"},{"id":94847688,"identity":"0b337e52-b36d-4bf9-9ebc-37efda32a21b","added_by":"auto","created_at":"2025-10-31 10:31:18","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":374891,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCharacterisation of a platform agnostic lung cancer molecular signature from blood plasma across DIA-MS, PEA, and DNA-aptamer platforms.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e, Workflow for DNA-APT characterisation of features selected by DIA-MS and PEA. Briefly, a cohort from the 563 individuals was selected to assess platform concordance. Twenty highly ranked features were then selected according to Shapley values across generalisable models were selected for each sub-cohort/platform. Additional explanation of DNA-aptamer candidate selection available in Extended Data Fig. 7b. SomaScan 11K panel was used to quantify these 100 proteins and processed in line with SomaLogic recommendations. Following QC, concordance and ML performance was evaluated on complete DIA-MS and PEA cohorts using a logistic regression, and using partial least squares (PLS) regression across all three platforms. \u003cstrong\u003eb\u003c/strong\u003e, Hierarchical clustering (Manhattan average) of Spearman correlation co-efficient of protein quantification across the three platforms (DIA-MS, PEA and DNA-APT) for 42 proteins quantified by all three methods. Columns correspond to comparison (purple = PEA vs DNA-APT, green = DIA-MS vs PEA and blue = DIA-MS vs DNA-APT); rows correspond to ML features (lilac = feature in PEA and DIA-MS models, turquoise = feature in DIA-MS models only, yellow = feature in PEA models only); * = FDR adjusted p-value \u0026lt; 0.05 for significance of Spearman correlation. \u003cstrong\u003ec\u003c/strong\u003e, Histogram of Spearman correlation coefficients across the three comparisons. \u003cstrong\u003ed\u003c/strong\u003e, ROC curves of holdout performance for logistic regression models trained on DIA-MS (turquoise) or PEA (yellow) training data using either features with \u0026gt; 0.4 spearman correlation coefficient of protein quantification between DIA-MS and DNA-APT or PEA and DNA-APT for DIA-MS (n = 25) and PEA (n = 25) models respectively (dashed line), or \u0026gt; 0.4 spearman correlation coefficient across all three platforms (n = 18). AUROC for each model is shown in plot. \u003cstrong\u003ee\u003c/strong\u003e, PLS plots (upper) and corresponding PLS-loadings (lower) trained with features with concordance \u0026gt; 0.4 spearman correlation coefficient across all three platforms, using data from DNA-APT (left), DIA-MS (middle) or PEA \u0026gt; 0.4 (right). PLS plots show separability between control samples (blue) (n = 70) and lung cancer samples (orange) (n = 69). Loadings are labelled with gene names.\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/ebcf44f394b42f4e0a7f8b8c.png"},{"id":94990356,"identity":"22e18aed-54f5-4532-845e-e8acd3da42dc","added_by":"auto","created_at":"2025-11-03 07:16:34","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3699749,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/3c113f18-0072-4d01-8b5a-941383c43eb1.pdf"},{"id":94847692,"identity":"bf5569bd-fa86-46fa-b29e-dadea814fc72","added_by":"auto","created_at":"2025-10-31 10:31:19","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":21149309,"visible":true,"origin":"","legend":"\u003cp\u003eData Set 1\u003c/p\u003e","description":"","filename":"SupplementaryTable1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/4304b31e4e93c38d95dbbd08.xlsx"},{"id":94985886,"identity":"e545764e-d96a-4d32-8c44-7beb14e9915c","added_by":"auto","created_at":"2025-11-03 06:59:11","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":16034871,"visible":true,"origin":"","legend":"\u003cp\u003eData Set 2\u003c/p\u003e","description":"","filename":"SupplementaryTable2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/0275dac1e382e3215271c63f.xlsx"},{"id":94847689,"identity":"c2730648-ff9d-4404-aada-0dec4087e183","added_by":"auto","created_at":"2025-10-31 10:31:18","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":206587,"visible":true,"origin":"","legend":"\u003cp\u003eData Set 3\u003c/p\u003e","description":"","filename":"SupplementaryTable3.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/b976933f2ef85e64ba84cc2d.xlsx"},{"id":94984848,"identity":"0da3d3b3-24a2-444f-9bf8-8f3a0afd95cf","added_by":"auto","created_at":"2025-11-03 06:56:36","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":27130,"visible":true,"origin":"","legend":"\u003cp\u003eData Set 4\u003c/p\u003e","description":"","filename":"SupplementaryTable4.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/12a1ee133d5db6c99588812b.xlsx"},{"id":94847691,"identity":"ec852de1-9292-4fda-b20f-5a8c1f96a570","added_by":"auto","created_at":"2025-10-31 10:31:18","extension":"docx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":5315242,"visible":true,"origin":"","legend":"","description":"","filename":"ExtendedData.docx","url":"https://assets-eu.researchsquare.com/files/rs-7660411/v1/17c878b19824b8fc7cde90a4.docx"}],"financialInterests":"\u003cb\u003eYes\u003c/b\u003e there is potential Competing Interest.\nH.R.F, N.G, L.H, I.K, M.D, J.S, Emma.M, Ella.M, D.A.S, A.H, and P.J.L are current/former employees, shareholders, and/or share option holders of Oxford Cancer Analytics Ltd.\r\nR.F and B.M.K are share option holders of Oxford Cancer Analytics Ltd.\r\nOxford Cancer Analytics Ltd sponsored this study.","formattedTitle":"\u003cp\u003eMultidimensional proteomics and explainable AI feature selection identify cross-platform lung cancer molecular signature in blood plasma\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eLung cancer is the leading cause of cancer mortality worldwide accounting for one fifth of total global cancer mortality and two million deaths per annum globally\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. Over 70% of lung cancer diagnoses are made at late stages, when regional-to-distant disease have a five-year survivorship of 9\u0026ndash;37% for non-small cell lung cancer (NSCLC) and 3\u0026ndash;18% for small cell lung cancer (SCLC), in the US, respectively\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. In contrast, early stage, localised lung cancer five-year survivorship is 65% for NSCLC and 30% for SCLC. Therefore, enabling earlier lung cancer diagnosis has been a priority for improving patient survivorship and outcome.\u003c/p\u003e\u003cp\u003eLow dose computed tomography (LDCT) has been widely studied for lung cancer screening in high-risk populations. A number of studies including the NELSON study in the EU and the National Lung Screening Trial (NLSC) in the United States showed a 20% reduction in mortality in the LDCT screening group\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e,\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. However, uptake and adoption have widely been reported to be poor ranging from as low as \u0026lt;\u0026thinsp;5% in the US to up to ~\u0026thinsp;50% in the UK. Limiting factors including radiation exposure, cost, and throughput have been cited to prevent LDCT accessibility and uptake, demonstrating challenges for scaling up\u003csup\u003e5,6\u003c/sup\u003e. Further, as smoking rates fall in Western and Asian countries, a greater fraction of lung cancer are occurring in individuals with a light or never-smoking history\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e; these individuals are ineligible for LDCT screening according to current criteria. Therefore, a minimally invasive and scalable approach for identifying high-risk lung cancer individuals with the potential for population level deployment may help improve current lung cancer screening paradigms.\u003c/p\u003e\u003cp\u003eThe blood proteome harbours promising information for the identification of a molecular signature for lung cancer, which can be applied as a liquid biopsy blood test. Prior efforts using a cell-free DNA based approach have shown limited sensitivity and specificity especially in earlier stage and low tumour burden disease, likely due to a low level of circulating tumour DNA\u003csup\u003e\u003cspan additionalcitationids=\"CR9\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. In contrast, the blood proteome provides important insight including tumour-specific inflammatory and host response even during early and low-tumour burden disease, serving as a potent repertoire for novel biomarker discovery\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e,\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Quantification of such circulating proteins may create a molecular signature for lung cancer. To this end, past studies have focused on targeted proteomics approaches such as proximity extension assay (PEA) and parallel/multiple reaction monitoring, liquid-chromatography tandem-mass spectrometry (LC-MS) or untargeted approaches including data-dependent/ data-independent LC-MS. However, different platforms may exhibit biased preferences for different proteins and pathways and show incongruence between quantification platforms, limiting biomarker signatures to a specific platform\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eTo enable the identification of a robust blood biomarker signature, the use of reliable machine learning (ML) models that generalise to unseen data and can model the interactions between the features is critical\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. In particular it is crucial to aggregate the information and importance of biomarkers across many different models to identify a robust panel that functions regardless of underlying models or model parameters. In addition, the black-box nature of more advanced ML models has been an ongoing concern, especially when clinical guidelines and regulatory approval processes benefit from understanding the rationale behind results.\u003c/p\u003e\u003cp\u003eIn this present study, we deployed multidimensional-proteomics combined with an explainable artificial intelligence (XAI)-guided ML approach on a large lung cancer cohort to identify a robust plasma-based molecular signature across diverse platforms. This combined data-independent acquisition mass spectrometry (DIA-MS) as an unbiased proteomics approach with a targeted PEA (Olink) for biomarker discovery from a cohort of 614 individuals (490 lung cancer cases across all stages, histologies, smoking exposures; and 124 age-, sex-matched controls from a research Canadian screening cohort). Using an 80/20 train/holdout strategy with cross-validated Shapley values for interpretability, we identified high-importance model-consistent features and uncovered key feature interactions relevant to lung cancer prediction. A third proteomics approach, targeted DNA-aptamer (SomaScan), was deployed on an expanded subset of patients to characterise a focused cross-platform biomarker signature. This multidimensional proteomics and XAI-driven methodology addresses core challenges in liquid biopsy development, including black-box model transparency, feature generalisability, and platform-specific constraints. We report the first cross-platform plasma-based lung cancer molecular signature identified through integrated multi-platform proteomics and explainable AI.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eSynergy between DIA-MS and PEA proteomics approaches enhances proteome coverage\u003c/h2\u003e\u003cp\u003eSingle-shot DIA-MS and PEA were used for protein quantification of plasma samples from 490 post-diagnosis, treatment-naive lung cancer cases covering all stages and major histologies, and 124 age-/sex-matched cancer-free controls (n\u0026thinsp;=\u0026thinsp;124) recruited through a lung screening programme, as described previously\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e (Fig.\u0026nbsp;1a, Extended Data Fig.\u0026nbsp;1a). Clinical and demographic characteristics are summarised in Table\u0026nbsp;1.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eClinical and demographic characteristics of healthy control vs lung cancer cohort\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eControl (n\u0026thinsp;=\u0026thinsp;124)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNSCLC (n\u0026thinsp;=\u0026thinsp;433)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSCLC (n\u0026thinsp;=\u0026thinsp;55)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eRare (n\u0026thinsp;=\u0026thinsp;2)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNumber of females, \u003cem\u003en (%)\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e63 (50.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e219 (50.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e21 (38.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1 (50)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMean age in years, \u003cem\u003e(sd)\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e69 (10)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e68 (11)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e68 (10)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e57 (13)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e\u003cp\u003eAge range (10 year intervals), \u003cem\u003en (%)\u003c/em\u003e:\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;50\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e23 (5.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2 (3.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1 (50)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e50\u0026ndash;60\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e27 (21.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e76 (17.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e12 (21.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e60\u0026ndash;70\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e30 (24.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e130 (30)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e16 (29.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1 (50)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e70\u0026ndash;80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e54 (43.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e148 (34.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e15 (27.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e80\u0026ndash;90\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e13 (10.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e49 (11.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e10 (18.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026gt;\u0026thinsp;90\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e7 (1.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e\u003cp\u003eEthnicity, \u003cem\u003en (%)\u003c/em\u003e:\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAsian/Pacific Islander\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e5 (4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e66 (15.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e3 (5.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBlack/African-Canadian\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e11 (2.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFirst Nations\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2 (0.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLatino/Hispanics\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2 (0.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1 (1.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMixed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e4 (0.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWhite/Caucasian\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e118 (95.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e268 (61.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e45 (81.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2 (100)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eUnknown\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e1 (0.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e80 (18.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e6 (10.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e\u003cp\u003eSmoking history:\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMean pack-years \u003cem\u003e(sd)\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e40 (21)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e32 (31)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e50 (29)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e27 (11)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHeavy (20 py)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e107 (86.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e252 (58.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e49 (89.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1 (50)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNon-heavy (\u0026lt;\u0026thinsp;20 py)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e17 (13.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e175 (40.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e5 (9.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1 (50)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eUnknown\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e6 (1.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1 (1.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e\u003cp\u003eSmoking status at diagnosis, \u003cem\u003en (%)\u003c/em\u003e:\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCurrent smoker\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e43 (34.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e128 (29.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e32 (58.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1 (50)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEx-smoker\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e81 (65.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e150 (34.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e16 (29.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1 (50)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEver smoker, NOS\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1 (0.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLight ex-smoker\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e11 (2.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNever smoker\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e119 (27.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2 (3.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e\u003cp\u003eStage of cancer, \u003cem\u003en (%)\u003c/em\u003e:\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStage 0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1 (0.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStage 1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e139 (32.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e4 (7.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1 (50)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStage 2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e44 (10.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2 (3.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStage 3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e86 (19.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e17 (30.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1 (50)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStage 4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e161 (37.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e32 (58.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eUnknown\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2 (0.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e\u003cp\u003eClinical morphology, \u003cem\u003en (%)\u003c/em\u003e:\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLUAD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e251 (58)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2 (100)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLUSC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e107 (24.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLarge cell carcinoma\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e32 (7.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eOther\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e43 (9.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e55 (100)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e\u003cp\u003eMolecular alteration, \u003cem\u003en (%)\u003c/em\u003e:\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEGFR-\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e4 (0.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEGFR+\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e118 (27.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEGFR Unknown\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3 (0.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEGFR NR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e308 (71.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e55 (100)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2 (100)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eALK+\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1 (0.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eALK-\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e114 (26.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eALK NR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e308 (71.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e55 (100)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2 (100)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eALK Unknown\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e10 (2.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eKRAS+\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1 (0.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTP53+\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1 (0.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eNSCLC, non-small cell lung cancer; SCLC, small cell lung cancer; n, number of; sd, standard deviation; py, pack-years; NOS, not otherwise specified; LUAD, lung adenocarcinoma; LUSC, lung squamous cell carcinoma +, positive for mutation; -, negative for mutation; NR, not reported; NA, not applicable.\u003c/p\u003e\u003cp\u003eDIA-MS quantified a wider-range of proteins (n\u0026thinsp;=\u0026thinsp;3,655) compared with PEA (n\u0026thinsp;=\u0026thinsp;2,884), with 1,136 quantified by both (Fig.\u0026nbsp;1b, Supplementary Fig.\u0026nbsp;1a). Given the high dynamic range of plasma protein concentrations, we assessed available concentrations as reported in the Human Protein Atlas (HPA)\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. DIA-MS quantified 1,736 mid-low abundance proteins (\u0026lt;\u0026thinsp;0.1mg/L), while PEA quantified 1,486 (Fig.\u0026nbsp;1c). Data missingness was higher in DIA-MS (37.07% \u0026plusmn; 34.82%) compared with PEA (13.09% \u0026plusmn; 24.33%). Both datasets showed a small but statistically significant (FDR adjusted \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05) negative Spearman correlation between protein concentration and missingness, suggesting poorer quantification at lower abundance (Fig.\u0026nbsp;1d). Assessment of secretome location in the HPA showed most proteins were not known to be secreted, with secreted proteins predominantly secreted to the blood across the abundance range (Extended Data Fig.\u0026nbsp;1b).\u003c/p\u003e\u003cp\u003eMatrix-matched technical replicates were acquired by PEA alongside clinical samples (n\u0026thinsp;=\u0026thinsp;14), while for DIA-MS a pool of clinical samples was evaluated (n\u0026thinsp;=\u0026thinsp;42). PEA had lower CVs than DIA-MS with a median CV of 21.4% versus 34.1%, although high CV proteins were not consistent across platforms (Fig.\u0026nbsp;1e). Analysis of Spearman correlation across protein abundance range indicated higher cross-platform variability for low-abundance proteins (Fig.\u0026nbsp;1f).\u003c/p\u003e\u003cp\u003eImplementing dual protein quantification platforms expanded information for biomarker discovery (Fig.\u0026nbsp;1b). Over-representation analysis (ORA) showed over-representation of proteins associated with cell migration, signalling, haemostasis and immune response in PEA. Also enriched by DIA-MS were cytoskeleton- and chromatin organisation-associated proteins (Fig.\u0026nbsp;1g). Together, these results show that combining DIA-MS and PEA increases the pool of potential biomarker candidates.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eUnivariate analysis identified lung cancer-associated plasma proteins and sub-cohort dependent variation\u003c/h3\u003e\n\u003cp\u003ePrincipal component analysis (PCA) showed a gradual stage-dependent separation of lung cancer cases from cancer-free controls (Fig.\u0026nbsp;2a,b, Supplementary Fig.\u0026nbsp;2). In both datasets, early-stage lung cancer cases (Stage I-II), were closer to cancer-free controls compared with late stage (Stage III-IV). PCA loadings showed that separation was distributed across many features (Fig.\u0026nbsp;2c,d). SERPINA3, PI16 and GSN contributed to separation in both datasets.\u003c/p\u003e\u003cp\u003eAnalysis of differentially abundant proteins between cancer-free controls and lung cancer cases showed a larger level of significant downregulation (FDR-adjusted \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05, One-way ANOVA, with Log\u003csub\u003e2\u003c/sub\u003e foldchange (Log\u003csub\u003e2\u003c/sub\u003eFC)\u0026thinsp;\u0026gt;\u0026thinsp;1) (Fig.\u0026nbsp;2e,f). Differential abundance was also evaluated on sub-cohorts including early and late lung cancer, lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC) and male-/female-only comparisons. Minimal difference in differential abundance was observed across sub-cohorts, with the exception of LUAD/LUSC in both datasets, as well as early lung cancer in PEA (Extended Data Fig.\u0026nbsp;2a,b).\u003c/p\u003e\u003cp\u003eGene set enrichment analysis (GSEA) showed the direction of regulation was similar across platforms, although little overlap was observed in significantly regulated pathways (FDR-adjusted \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05) (Fig.\u0026nbsp;2g, Extended Data Fig.\u0026nbsp;2c), suggesting low concordance/overlap between approaches could generate different interpretations. Regulation of actin cytoskeleton organisation, protein complex assembly, and hydrogen peroxide metabolism associated proteins was observed in DIA-MS, irrespective of sub-cohort. In PEA, pathways associated with coagulation and GTPase/RAS signalling, were significantly downregulated across most sub-cohorts, while RNA metabolism, immune response and chromatin organisation pathways were significantly upregulated in SCLC. Overlapping pathways were related to acute-phase response and actin cytoskeleton organisation, with differences among subtypes largely derived from smaller populations (SCLC and LUSC). Overall, univariate analysis of plasma proteins revealed limited insight into biological processes, particularly in the more heterogeneous sub-populations.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eXAI-informed multivariate feature selection to build panels of lung cancer plasma biomarker candidates\u003c/h3\u003e\n\u003cp\u003eWe next trained 75 tree-based ML models using combinations of different model architectures, feature de-selection, data augmentation, feature selection, multi-objective optimisation and XAI techniques to identify biomarkers that were consistently predictive of lung cancer across multiple models. All data underwent identical sample stratification, train/holdout splitting with separate pre-processing, minimising data-leakage to unseen holdout data to promote generalisation. All models were evaluated across nine lung cancer sub-cohorts for both datasets, giving rise to 1,350 ML models (Fig.\u0026nbsp;3a, Extended Data Fig.\u0026nbsp;3a, Supplementary Table\u0026nbsp;1).\u003c/p\u003e\u003cp\u003ePEA models had better performance than DIA-MS models; all and late lung cancer, NSCLC, LUSC and LUAD and male-only PEA models achieved better performance than other sub-cohorts, as assessed by area under the receiver operator characteristic curve (AUROC) and Matthews Correlation Coefficient (MCC) (Fig.\u0026nbsp;3b-e). In DIA-MS models NSCLC performance was higher than SCLC. Male-only models outperformed female-only models in both datasets. Female-only lung cancer cases had a lower incidence of LUSC and SCLC compared with males and higher incidence of LUAD, with comparable age distribution (Supplementary Fig.\u0026nbsp;3a, b). Interestingly, assessment of an all-lung cancer trained-model predicting sub-cohorts showed similar decreased performance in female-only versus male-only cohort (Extended Data Fig.\u0026nbsp;3b).\u003c/p\u003e\u003cp\u003eTo identify robust feature sets, we next assessed Shapley value feature importance across all lung cancer models\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. Some high-importance features were consistently selected for a model (WFDC2, CILP, NCF2, MMP7, GLRX) while some high-importance features showed strong model-dependency (SH3BP2, VWA8, KIF20B, CLEC4G and BCAT1) (Fig.\u0026nbsp;3f). Upon assessment of individual LC models, MB, a feature selected by DIA-MS and PEA, showed an early contribution during recursive feature addition (RFA) for both approaches (Fig.\u0026nbsp;3g), also showcasing how using Shapley values during RFA promoted incremental performance increases. Smaller performance increases suggested incorporation of redundancy during feature selection, while minor AUROC changes accompanied by larger changes in MCC highlighted insights from considering two performance metrics.\u003c/p\u003e\u003cp\u003eUnivariate analysis is often used to select features for ML\u003csup\u003e19\u003c/sup\u003e. During feature de-selection, features that had low information about the target (mutual information (MI)\u0026thinsp;\u0026lt;\u0026thinsp;3.5%) were removed. The majority of differentially abundant proteins had MI\u0026thinsp;\u0026gt;\u0026thinsp;3.5%, however a high number of proteins with high MI were not differentially abundant (Extended Data Fig.\u0026nbsp;3c,d). Moreover, a large number of high-importance features did not pass the threshold for differential abundance, particularly for DIA-MS models, while some high importance features did not have MI\u0026thinsp;\u0026gt;\u0026thinsp;3.5%, emphasising the utility in using multiple feature selection-configurations. A direct comparison of univariate with the XAI-informed feature selection approach used here showed that despite containing non-differentially abundant proteins, overall ML-driven feature selection achieved higher performance even with a simple logistic regression model (Extended Data Fig.\u0026nbsp;3e).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eCharacteristics of features selected\u003c/h3\u003e\n\u003cp\u003eThe Rashomon effect describes the phenomena in ML where similar performances can be achieved with different internal workings of modell\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e We addressed this by comparing feature selection across many different models.\u003c/p\u003e\u003cp\u003eTo evaluate feature selection, we considered models that generalised, determined by comparing holdout performance to 95% CI from training data (Extended Data Fig.\u0026nbsp;3a). Five features were selected in generalisable PEA and DIA-MS models (LRG1, SDC1, MB, LTA4H and FGL1) (Fig.\u0026nbsp;4a). LTA4H was the only common feature regulated differently across approaches, with minor upregulation (significant in SCLC) in DIA-MS and significant downregulation across PEA sub-cohorts (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05, Log2FC\u0026thinsp;\u0026gt;\u0026thinsp;1). Other proteins quantified by both methodologies were only selected as features for DIA-MS (n\u0026thinsp;=\u0026thinsp;25) or PEA (n\u0026thinsp;=\u0026thinsp;13), albeit mostly regulated similarly across approaches. Like LTA4H, some showed divergent abundance (THBS2, THOP1 and ADAMTSL4) while others showed strong regulation in one dataset with minimal regulation by the alternative method (MUC16, CEACAM5, CILP, HAGH and LXN).\u003c/p\u003e\u003cp\u003eAcross generalisable models, high importance features according to Shapley values were similar for all/late lung cancer, NSCLC and LUAD. Variation across sub-cohorts was generally lower for PEA. SCLC and sex-split sub-cohorts diverged most from other models. In female-only DIA-MS models, high importance features for many cohorts (MB, WFDC2, HBA2, TAGLN2 and PPBP) were less informative while other candidates (POSTN, VWA8, HAGH, NAGLU and MUC1) were never selected. Similarly for male-only models in DIA-MS, while the overall order of importance was preserved, many features were never selected. In early lung cancer models, SDC1, POSTN, PPBP and SERPINA3 were less important/absent compared with all/late lung cancer models, suggesting their presence arises from influence of later stage lung cancer. Higher importance features had low missingness in DIA-MS, whereas in PEA two candidates (LTA4H and NEXN) were largely quantified below the LLOD (Extended Data Fig.\u0026nbsp;4a-b). High-importance lung cancer features were lower in abundance in PEA models (Extended Data Fig.\u0026nbsp;4c).\u003c/p\u003e\u003cp\u003eBiological pathways enriched in lung cancer features showed that proteins involved in chemotaxis, cell adhesion, wound healing and immune response (acute-phase, complement activation and humoral immune response) were present at varying levels across sub-cohorts in both datasets (Fig.\u0026nbsp;4c). This challenged previous GSEA showing immune response pathways enrichment in PEA only (Fig.\u0026nbsp;2g). Both datasets relied on features associated with cell migration, while in PEA features were enriched for receptor-mediated endocytosis, integrin and cytokine signalling. DIA-MS was also enriched for oxidative stress features (Fig.\u0026nbsp;4c). In both datasets, if known, secretion to blood was more prevalent (Extended Data 4d).\u003c/p\u003e\u003cp\u003e\u003cb\u003eXAI-driven feature selection captures strong interactions between features.\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo further understand how models used features, we investigated how Shapley values individually and through interactions explained differences between cancer-free controls and sub-cohorts.\u003c/p\u003e\u003cp\u003eFeature quantification from a lung cancer model separated cancer-free control samples from early and late lung cancer samples by t-SNE (Extended Data Fig.\u0026nbsp;4e). Shapley values better separated cancer-free controls from lung cancer samples when compared to protein quantification, with early/late differences being less obvious (Extended Data Fig.\u0026nbsp;4f). This suggested high-importance lung cancer features could be used regardless of stage. This is further supported by the high performance of lung cancer models predicting early lung cancer (AUROC\u0026thinsp;=\u0026thinsp;0.89 [95% CI: 0.79\u0026ndash;0.98], MCC\u0026thinsp;=\u0026thinsp;0.65 [95% CI: 0.46\u0026ndash;0.82]) (Extended Data Fig.\u0026nbsp;3b) compared with the performance across early lung cancer models (AUROC\u0026thinsp;=\u0026thinsp;0.82 [95% CI: 0.77\u0026ndash;0.88], MCC\u0026thinsp;=\u0026thinsp;0.48 [95% CI: 0.34\u0026ndash;0.64]) in DIA-MS (Fig.\u0026nbsp;3b,d).\u003c/p\u003e\u003cp\u003eAnalysis of Shapley value model-decision paths showed individual feature contribution towards a prediction (Extended Data Fig.\u0026nbsp;6a,b), while interactions demonstrated how features are mediated by each other. Shapley value interactions from a lung cancer model identified complex networks of feature interactions (Fig.\u0026nbsp;4d,e, Extended Data Fig.\u0026nbsp;6c,d). For DIA-MS, this network contained a higher level of interactions compared with PEA. Shapley value importance scores were comparatively lower and more evenly distributed within the PEA network. For DIA-MS, highly important features were involved in the strongest interactions (MB, WFDC2, HBA2, CILP and SDC1). Increased levels of CILP, a universal feature for all DIA-MS models, interacted with decreased levels of SDC1, SELENBP1 and/or WFDC2 to be predictive of cancer-free controls (Extended Data Fig.\u0026nbsp;6e). Similarly, decreased levels of SDC1 interacted with increased levels of MB and decreased levels of CEACAM5 and/or MMP7 to be predictive of cancer-free controls in the analysis of top PEA interactions. (Extended Data Fig.\u0026nbsp;6f). Several interactions (TAGLN2-MUC1, DDTL-ACTB, CASP8-CLIP2 and NCF2-CLIP) identified two groups of controls. These interactions emphasise the importance of using models capable of capturing feature interactions for disease identification.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eThe identification of a platform-agnostic lung cancer plasma protein signature.\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWe next introduced a third DNA-aptamer-based proteomic approach (DNA-APT) to further evaluate concordance across platforms. A panel of 100 SomaScan quantified proteins selected from previously identified features (Fig.\u0026nbsp;5a, Extended Data Fig.\u0026nbsp;7a,b) was applied to a cohort of 140 patients selected from the previously described cohort (Table\u0026nbsp;1), 72 cancer-free controls and 68 lung cancer cases.\u003c/p\u003e\u003cp\u003eVariability was low in DNA-APT assay and dimensionality reduction showed separability between cancer-free controls, early and late lung cancer samples (Extended Data Fig.\u0026nbsp;7c).\u003c/p\u003e\u003cp\u003eThe majority of proteins quantified in all three methods (n\u0026thinsp;=\u0026thinsp;27, 57.1%) had a significant Spearman correlation (FDR-adjusted \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05) (Fig.\u0026nbsp;5b, Extended Data Fig.\u0026nbsp;7d). Six features showed no concordance with the DNA-APT approach. A group of predominantly PEA features (n\u0026thinsp;=\u0026thinsp;16) had a strong, significant correlation between PEA/DNA-APT only. Similarly, a group of DIA-MS features (n\u0026thinsp;=\u0026thinsp;16) had strong correlation between DIA-MS/DNA-APT only. High-importance feature SERPINA3 had opposite correlation in DIA-MS/DNA-APT vs PEA/DNA-APT. Overall, concordance between DNA-APT and PEA was higher than concordance with DIA-MS, with median Spearman \u003cem\u003eR\u003c/em\u003e of 0.567 versus 0.363 (Fig.\u0026nbsp;5c).\u003c/p\u003e\u003cp\u003eA panel of proteins with Spearman \u003cem\u003eR\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.4 across all three panels were evaluated (n\u0026thinsp;=\u0026thinsp;18) and a second panel that contained proteins with Spearman \u003cem\u003eR\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.4 (n\u0026thinsp;=\u0026thinsp;25) between DIA-MS/DNA-APT or PEA/DNA-APT data were evaluated using DIA-MS and PEA (see methods). Similar performances were seen between DIA-MS and PEA for the three-platform panel, with AUROC of 0.88 [95% CI: 0.80\u0026ndash;0.95] and 0.88 [95% CI: 0.81-0. 95], and MCC of 0.49 [95% CI: 0.31\u0026ndash;0.67] and 0.60 [95% CI: 0.44\u0026ndash;0.77] for DIA-MS and PEA, respectively (Fig.\u0026nbsp;5d). The PEA/DNA-APT two-platform panel achieved the highest performance, with AUROC of 0.97 [95% CI: 0.94\u0026ndash;0.99] and MCC of 0.71 [95% CI: 0.54\u0026ndash;0.86]. The DIA-MS/DNA-APT model outperformed the three-platform panel with AUROC of 0.93 [95% CI: 0.86\u0026ndash;0.97] and MCC of 0.60 [95% CI: 0.44\u0026ndash;0.78].\u003c/p\u003e\u003cp\u003eTo evaluate the DNA-APT assay, a partial least squares (PLS) regression was used to show differences between cancer-free controls and lung cancer, which were similarly distinguishable across all platforms (Fig.\u0026nbsp;5e). Analysis of loadings showed that similar features were responsible for separation on all platforms, separating cancer-free controls (MB, PF4, PPBP, CNTN3 and KIT) from lung cancer (HAGH, ALDH1A1, FGL1, MMP9, C9, CFHR5 and LBP). The same was carried out for the two-platform panels which achieved similar separation by PLS regression (Extended Data Fig.\u0026nbsp;7e). We show the identification of feature panels compatible across multiple proteomics platforms, as demonstrated by a similar degree and driver(s) of separation between the cross-platform panels.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study deployed unbiased DIA-MS and targeted PEA proteomics as discovery approaches in combination with XAI-guided ML feature selection. On a cohort of 614 patients including 490 lung cancer cases (stages I-IV) covering all major histologies and 124 age-, sex-matched controls we evaluated plasma molecular signatures for lung cancer across 1,350 ML models. Further characterisation using a DNA-aptamer approach identified an eighteen-feature lung cancer plasma signature with concordance across all platforms, achieving AUROC of 88.2% and 88.3% in DIA-MS and PEA holdout validation, respectively. Signatures with high concordance across two of the three panels achieved even higher AUROC between 92.4% to 97%. PLS regression on the DNA-APT validation dataset clearly distinguished lung cancer from controls.\u003c/p\u003e\u003cp\u003eThe proteomics approaches deployed are among the most widely used, DIA-MS offering unbiased discovery, while PEA and DNA-APT are often favoured in clinical research for straightforward data processing\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. LC-MS based proteomics is increasingly considered for clinical purposes, particularly with regards to cancer detection and disease monitoring due to high-throughput and multiplexing\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. This study leveraged both the overlap and divergence between platforms to expand coverage and to assess feature robustness. In-line with previous studies, we observed platform-specific discrepancies\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e, potentially arising from either DIA-MS pre-processing or immunoassay epitope-specific biases, reinforcing concerns about the reliability of signatures developed from a single method. Dynamic range is often considered a limitation of DIA-MS. Low abundance proteins showed greater variability and reduced concordance (Extended Data Fig.\u0026nbsp;4c), however models from all platforms relied on such features to varying degrees and yielded predictive signatures. By focusing on high-importance features with strong cross-platform concordance, we identified a robust plasma signature for lung cancer, supporting the development of transferable, platform-agnostic biomarker panels, facilitating scalable and flexible clinical implementation.\u003c/p\u003e\u003cp\u003eDespite differences across platforms, cytokine/chemokine-mediated signalling were associated with both DIA-MS and PEA biomarker panels, with several proteins previously considered for lung cancer detection, monitoring or prognosis. Many candidates have also been previously associated with cancer pathology but would be considered novel biomarkers of lung cancer prediction (Extended Data Fig.\u0026nbsp;5a, b). Five proteins were selected across both proteomics platforms (MB, SDC1, LRG1, FGL1 and LTA4H). MB had high importance/concordance across both platforms and had previously been associated with skeletal-muscle wasting in cancer-related cachexia\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. We observed that low plasma MB coupled with regulation of other candidates (SDC1 and COTL1) were important in lung cancer prediction. SDC1 had high importance across both datasets, with a known role in lung cancer pathology and as a prognostic marker in multiple solid-tumours\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. We observed significant upregulation of SDC1 in all but early lung cancer sub-cohorts in both datasets, indicating increased circulating SDC1 with stage. To the best of our knowledge, MB and SDC1 have not previously been reported as being predictive of lung cancer. We also characterised high importance lung cancer features associated with immune/inflammatory responses. SERPINA3 and CRP, associated with acute-phase response, have previously been proposed as biomarkers for early NSCLC. We, however, observed SERPINA3 upregulation to be associated with later-stage NSCLC, LUSC/male-only lung cancer. There were however early-stage associated proteins characterised by both approaches. In PEA, the transferrin receptor (TFRC) was important for early lung cancer among other sub-cohorts (Fig.\u0026nbsp;4b). TFRC upregulation is known in NSCLC and other solid-tumours, also expressed on the surface of activated immune cells\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e, potentially with a dual-role in mediating lung cancer progression\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. Similarly, HAGH showed upregulation in early-stage lung cancer (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05 in DIA-MS only) (Fig.\u0026nbsp;4a, Fig.\u0026nbsp;5e), the only candidate to show an early-dominant response across both platforms, suggesting HAGH upregulation in plasma may occur early and reduce over disease progression. This is in line with HPA data showing HAGH has favourable prognostic value in other cancers including liver, renal and pancreatic cancer\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. Pathway analysis showed that proteins related to oxidative stress and wound healing were informative, even in early-stage disease. Together, this underscores the potential of protein-based approaches to detect early-stage and low-tumour burden diseases by identifying amplified inflammatory signals, addressing concerns about limited ctDNA signal in such cases\u003csup\u003e\u003cspan additionalcitationids=\"CR9\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. Finally, we showed that despite having used cohort-specific models for biomarker selection, it appears that models trained on the complete cohort were similar/superior at understanding the subclassification task, possibly due to additional context from other samples. For example, predictions of early cancer patients in unseen holdout data were improved when late cancer patients were included in training. Future work focused on early detection of cancer may consider the inclusion of late-stage samples to improve separability.\u003c/p\u003e\u003cp\u003eModel performances and feature selection varied between male and female-only models indicative of sex-based variation in plasma molecular signatures (Fig.\u0026nbsp;3b,c). Male-only candidates had a high association with cell migration and apoptotic pathways in line with other cohorts while some more universal features were not selected in female-only models (LRG1, SELENBP1, KIF20B, SERPINA3, MUC1, SAA1, MUC16, KRT19, CIT and TFRC) (Fig.\u0026nbsp;4a-c). Moreover, many biological pathways associated with lung cancer were not major candidates in female-only models, particularly immune response-related pathways, aligning with reports at the transcript level\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. This is consistent with previous studies showing similar differences between sex-separated molecular signatures, proposing the use of different features and models for prediction of sex-based sub-cohorts\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e,\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. Although we found lung cancer trained-models predicted female-only sub-cohorts to similar levels as female-only models, further assessment of sex-based differences in plasma molecular signatures and other confounding data in a larger cohort could further characterise the differences. This highlights the potential utility of evaluating sex-based differences in prediction using plasma molecular signatures, particularly in early detection where changes in protein abundance can be small in magnitude.\u003c/p\u003e\u003cp\u003eThe strengths in this study lie in the application of a multi-dimensional proteomics approach to a large cohort of lung cancer cases (n\u0026thinsp;=\u0026thinsp;490) covering all major histologies with a good representation of early-stage cases. Additionally, we applied rigorous XAI and ML methodologies, with separate train/holdout processing where possible, using two complementary metrics, AUROC and MCC and exploring both boosting (XGBoost) and bagging (BRF) algorithms. XGBoost was prioritised to achieve a more concise set of informative features, giving rise to smaller panels of ~\u0026thinsp;20 proteins. We also overcame poor model interpretability by using TreeSHAP approximation to efficiently compute cross-validated Shapley values across a large feature set. The resulting stable feature set is expected to maintain strong performance even in simple models. Finally, combining two distinct proteomics approaches for lung cancer circulating biomarker discovery broadened potential candidate coverage and combining three independent protein detection platforms allowed the assessment of a concordant lung cancer molecular signature across platforms.\u003c/p\u003e\u003cp\u003eA limitation of this study is that, with highly correlated features, Shapley values can distribute across redundant features, diluting the contribution of individual biomarkers. This is a common limitation in multivariate feature selection methods. Similarly, we focused on XGBoost models which can overfit when not carefully designed, and we mitigated these potential limitations by applying unseen-holdout generalisation constraints, regularisation and evaluating multiple models. Another limitation was the over-representation of lung cancer and rarer subtypes. Although this enabled sufficient power to provide reliable insight into subtype specific biomarkers, a more-representative cohort can assess the utility of this ML-approach and feature panels in populations more representative of epidemiological distributions. Similarly, assessment of real-world efficacy and health economic feasibility of plasma molecular signatures were not within the scope of this study. Future work seeks to further evaluate and validate our approach with prospective population level-based detection studies.\u003c/p\u003e\u003cp\u003eIn conclusion, we deployed a multidimensional proteomics and XAI-guided ML biomarker discovery approach in a lung cancer cohort composed of early stage, late stage, and all common histologies with age-, sex-matched controls from a screening program cohort. The combined synergy between different proteomics approaches and interpretable ML models identified a cross-platform lung cancer plasma protein signature with explainable feature ranking, contribution, and interaction mechanisms for decision making. We anticipate this signature to be highly versatile, transferable, and scalable across different protein quantification platforms and look forward to future assessment of its cost effectiveness and real-world impact on patient outcomes. Given promising results of sub-cohort analysis demonstrating the ability to detect specific lung cancer histologies and stages, customisable panels may be deployed in diverse clinical use cases including at-risk population risk stratification, screening, and disease classification.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003eCohort selection and collection\u003c/h2\u003e\u003cp\u003eLung cancer cases and healthy controls were recruited from Princess Margaret Cancer Centre\u0026rsquo;s Lung-CALIBRE program (Lung cancer Clinical And Liquid biopsy, Implementation, Breathomics, Radiomics for Early detection). Ethics approval for the CALIBRE program was obtained from The University Health Network Research Ethics Board (REB 06-639). Lung cancer cases were selected using a stage-stratified random process from a biospecimen database of lung cancer patients to sufficiently cover all major histologies and stages. Case participants included in this analysis were recruited sequentially from 2008\u0026ndash;2019 which met the conditions of age over 18 years; histological confirmation of lung cancer, adequate plasma specimen (1.8ml vial) taken after initial diagnosis prior to commencing first treatment.\u003c/p\u003e\u003cp\u003eControls were distribution-matched to age and gender of lung cancer cases from the research lung cancer screening program. This program used a broad eligibility of 10-pack year minimum cigarette smoking history and a minimum age of 45 years old.\u003c/p\u003e\u003cp\u003eHealthy control samples were collected at Princess Margaret Cancer Centre and lung cancer cases collected from Toronto General Hospital and Princess Margaret Hospital. All blood samples used in this study were collected in EDTA tubes and processed by the same operator to plasma using centrifugation within two hours of collection. All processing was carried out at \u0026le;\u0026thinsp;4\u0026deg;C and samples were stored at -80\u0026deg;C at Princess Margaret Cancer Centre for further use.\u003c/p\u003e\u003cp\u003eTo achieve\u0026thinsp;\u0026gt;\u0026thinsp;80% power for showing a minimum lung cancer detection sensitivity and specificity of 70% and 90% respectively (compared to a null hypothesis of 50% each, with alpha\u0026thinsp;=\u0026thinsp;0.05), an overall sample size of 490 cancers and 124 controls, with a 80/20% train/holdout split was selected to satisfy the requirement for a minimum number of holdout sample size of 49 cancers and 12 controls\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eSample preparation and processing for DIA-MS\u003c/h3\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003eSample randomisation\u003c/h2\u003e\u003cp\u003eThe order of the 614 clinical samples to be processed were randomised based on demographic and clinical characteristics. The randomised samples were processed in seven individual 96-well plates. Each plate contained a patient's plasma samples and six \u0026ldquo;pool plasma\u0026rdquo; which were later used for quality control. The pool sample was formed by pooling 3 \u0026micro;L of clinical plasma samples of mixed types.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003eTop 14 most abundant protein depletion\u003c/h2\u003e\u003cp\u003ePlasma samples stored at -80\u0026deg;C were thawed at RT. Five \u0026micro;L of neat plasma samples were used for each depletion. Top 14 depletion was performed for each sample using high-selection top 14 depletion resin (A36372, Thermo Scientific) following manufacturer\u0026rsquo;s instructions. Depleting the plasma high-abundance proteins improves the dynamic range in LC-MS proteomics analysis. Approximately 100 \u0026micro;L of depleted plasma per well were collected for each sample; 25 \u0026micro;L of depleted plasma from each sample were transferred to a fresh 96-well PCR plate (LoBind, Eppendorf) for digestion.\u003c/p\u003e\u003cp\u003eHaemolysis was evaluated as described previously\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e on the unprocessed plasma sample.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003eSP3 in solution digestion\u003c/h2\u003e\u003cp\u003e5 \u0026micro;L of 60 mM TCEP (XH347846, Thermo Scientific), 240 mM CAA (CO267-100G, Merck) mix were added to each sample in the well of a PCR 96-well plate. The plate was sealed with a heat sealer (E0030127838, Eppendorf) on a heat seal (E5391EN100431, Eppendorf). Reduction and alkylation were performed by incubating the samples at 95\u0026deg;C for 15 min on a Thermomixer.\u003c/p\u003e\u003cp\u003eSingle pot, solid-phase enhanced sample preparation protocol (SP3) was used to prepare peptide samples. Briefly, Sera-Mag\u0026trade; Carboxylate-Modified Magnetic Beads (45152105050250 and 65152105050250, Cytiva) were washed twice in HPLC grade water and 2 \u0026micro;L added to each sample along with 80 \u0026micro;L 99.8% EtOH (E/0650/17, Fisher chemical) and incubated at RT for 5 mins. The beads were subsequently washed with 80% EtOH three times. 1 \u0026micro;g MS grade Pierce\u0026trade; trypsin protease (90058, Thermo Scientific) in 100 mM ammonium bicarbonate (09831-500G, Merck) was added to each sample and incubated overnight at 37\u0026deg;C. Digested samples were transferred to a new plate, lyophilised and stored at -20\u0026deg;C ready for data acquisition.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\u003ch2\u003eSample loading\u003c/h2\u003e\u003cp\u003eDried samples were resuspended in 100 \u0026micro;L 0.1% FA (85170, Thermo Scientific). For analysis on the Evosep One LC system, the sample was loaded onto Evotips (EV2011, Evosep) according to Evotip sample loading protocol. Briefly, the Evotips were rinsed with 20 \u0026micro;L solvent B (85174, Thermo Scientific), conditioned with isopropanol (99.5%, 383920025, Thermo Scientific), then equilibrated with 20 \u0026micro;L of solvent A. 5 \u0026micro;L of digests were loaded onto each Evotip. The Evotips were washed once with 20 \u0026micro;L of solvent A and kept wet with 100 \u0026micro;L of solvent A before LC-MS analysis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003eDIA-MS data acquisition\u003c/h2\u003e\u003cp\u003ePeptides samples were loaded onto Evosep One (EV-100 S00394, Evosep) coupled to a Orbitrap Exploris 480 mass spectrometry (MA10103C, Thermo Scientific) interfaced with a FAIMS PRO device. A 150um x 15cm PepMap RSLC C18 easy spray column (ES906 Thermo Scientific) was used as the analytical column on Evosep One system.\u003c/p\u003e\u003cp\u003eAn optimised data independent acquisition (DIA) method was used for MS data acquisition to maximise the depth of plasma proteome. The MS1 scan instrument setting: orbitrap resolution 120000; scan range (m/z): 345\u0026ndash;1055; FAIMS CV (V): three sequential CVs \u0026minus;\u0026thinsp;40, -55, -70; Normalised AGC target (%): 300; Maximum injection time mode: Auto. The DIA MS2 scan the instrument setting: Orbitrap resolution: 30000; DIA window and m/z range: 38 variable DIA windows covering from 350 to 1050; FAIMS CV (V): three sequential CVs \u0026minus;\u0026thinsp;40, -55, -70 the same as MS1 FAIMS setting; RF length (%): 40; Normalised AGC target (%): 1000, Customed Maximum injection time (ms): 54; Data type: Profile; Loop control N\u0026thinsp;=\u0026thinsp;13.\u003c/p\u003e\u003cp\u003eSamples were run using Evosep One 15 sample per day method (gradient of 88min). Master pool samples were loaded as a QC after each plate row of samples.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\u003ch2\u003eData pre-processing\u003c/h2\u003e\u003cp\u003eIn total, 714 DIA MS files were acquired including 614 lung cancer samples, 42 pooled quality control plasma samples and 58 master pool digests. Data was first converted to .htrms format using HTRMS converter software available with Spectronaut V18.5 (Biognosys). All 714 HTRMS files were loaded into Spectronaut V18.5 (Biognosys) for library-free directDIA analysis. Unless specified, default settings parameters were used. Default BSG factory settings were used for peptide/protein identification and quantification. Qvalue cutoff was 0.01 for both peptides and protein identification. LFQ method for protein quantitation at MS1 level and cross run global normalisation were enabled. Pulsar search workflow: directDIA+(Deep) was used with peptide fixed modification cysteine carbamidomethylation, and variable modification: acetyl (protein N-term) and oxidation (M). Two miss cleavages were allowed for trypsin cleavage. Human uniport FASTA file (UP000005640_9606_Hsapiens, Feb 2023) was used for \u003cem\u003ein silico\u003c/em\u003e spectra library generation and protein inference.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\u003ch2\u003eProximity extension assay (PEA) data acquisition and processing\u003c/h2\u003e\u003cdiv id=\"Sec18\" class=\"Section3\"\u003e\u003ch2\u003eSample selection\u003c/h2\u003e\u003cp\u003eFor PEA analysis, 600 samples were selected prioritising samples with low haemolysis score. The 600 selected samples were processed in two stages; an initial study of 88 samples (one plate), PEA 01, consisting of 29 healthy controls and 59 lung cancer cases, was followed by a larger cohort of 528 samples across six plates. PEA 02 consisted of 100 healthy controls and 428 lung cancer, with 16 bridging samples (6 healthy control, 10 LC) used across the two studies.\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e\u003ch2\u003eOlink Explore 3072 Panel\u003c/h2\u003e\u003cp\u003eThe Olink\u0026reg; Explore 3072/384 assay (referred to as PEA) consists of eight panels: Cardiometabolic, cardiometabolic II, inflammation, inflammation II, neurology, neurology II, oncology and oncology II. Each panel targets a maximum of 384 proteins, which quantify protein abundance using pairs of antibodies with complimentary oligonucleotide tags targeting the same protein. When both antibodies are bound to the protein of interest (i.e. in close proximity) the oligonucleotide tags hybridise, allowing DNA polymerase-dependent extension. The extended dsDNA is then tagged with barcodes and quantified using Next Generation Sequencing. Samples were randomised across the seven plates. Each assay is spiked with internal controls that act as incubation controls, extension controls and detection controls. Alongside internal controls, each plate contains sample controls, negative samples and inter-plate controls, used for plate normalisation, limit of detection (LOD) estimation and variation assessment respectively. The Olink\u0026reg; Explore 3072/384 assay was carried out by Randox Laboratories Ltd. Six \u0026micro;L of plasma from each sample was mixed with the antibodies from the eight panels and the resulting dsDNA specific for target protein extended, labelled with barcoded sequences and sequenced using NovaSeq 6000 Sequencing System (Illumina) for relative quantification.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec20\" class=\"Section2\"\u003e\u003ch2\u003eData pre-processing\u003c/h2\u003e\u003cp\u003eRaw data was processed using NGS2COUNTS and imported into Olink\u0026reg; NPX software. Raw NGS reads were converted into normalised protein expression (NPX) values through a series of normalisation steps. Firstly, samples are normalised to extension controls and converted to Log\u003csub\u003e2\u003c/sub\u003e scale. Secondly, plate-control normalisation was performed by normalising the median of each sample to the median of plate controls. Intensity normalisation was subsequently performed, centering the median of each sample to the median of all samples (excluding controls) of the plate. These NPX values were used for all downstream analysis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec21\" class=\"Section2\"\u003e\u003ch2\u003eDNA-aptamer data acquisition and pre-processing\u003c/h2\u003e\u003cdiv id=\"Sec22\" class=\"Section3\"\u003e\u003ch2\u003eSample selection\u003c/h2\u003e\u003cp\u003eA sub-cohort of the Princess Margaret patient cohort was selected for protein quantification on a third platform. This cohort consisted of 72 healthy control patients that were matched to 32 early lung cancer and 32 late lung cancer according to age, sex and smoking status (Extended Data Fig.\u0026nbsp;7a).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec23\" class=\"Section3\"\u003e\u003ch2\u003eDNA-aptamer quantification of feature subset\u003c/h2\u003e\u003cp\u003eA DNA-aptamer approach, SomaScan 11K V5.0 (SomaLogic), was used to assess feature concordance on a third proteomics technology. From the 11K SomaScan panel, 100 proteins from PEA and DIA-MS feature selection were chosen (Extended Data Fig.\u0026nbsp;7b). Samples were randomised across three plates according to cancer status (cancer-free, early lung cancer and late lung cancer). The SomaScan assays consist of SOMAmer probes composed of fluorophores, biotin and a photocleavable linker, which bind specific proteins. Once bound, biotinylated protein complexes are captured on Streptavidin beads and washed to remove unbound/weakly bound proteins. Photocleavage dissociates SOMAmers from protein complexes, non-specific complexes are dissociated and prevented from reforming using polyanionic competitors. Remaining SOMAmer-protein complexes are recaptured on streptavidin, eluted and binding of complementary DNA sequences to a microarray allows measurement of fluorescence.\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec24\" class=\"Section2\"\u003e\u003ch2\u003eData pre-processing\u003c/h2\u003e\u003cp\u003eData processing of the SomaScan 11K V5.0 assay was performed using SomaScan in-house software. Results are reported in relative fluorescence units (RFU) and processed as follows: Firstly, using the 12 control SOMAmers spiked into each sample prior to microarray hybridisation normalisation was carried out. This was followed by intraplate median signal normalisation using plasma calibrator controls run alongside patient samples to account for dilution differences. Plate scaling was also carried out using calibrator samples, normalising the medians of calibrator reference ratios. Calibration was performed to normalise calibrator controls to reference values for each SOMAmer. Finally, Adaptive Normalisation by Maximum Likelihood (ANML) was used to normalise signal across samples. Briefly, SOMAmer reagents within a sample that are within two population standard deviations from reference are used to generate scaling factors; this was iterated over until convergence. The 11K panel was used for pre-processing while the data for the 100 chosen SOMAmers plus the 12 control SOMAmers were reported.\u003c/p\u003e\u003cdiv id=\"Sec25\" class=\"Section3\"\u003e\u003ch2\u003eBioinformatic analysis\u003c/h2\u003e\u003cdiv id=\"Sec26\" class=\"Section4\"\u003e\u003ch2\u003eSample and protein filtering\u003c/h2\u003e\u003cp\u003eDIA-MS and PEA sample filtering are described in Extended Data Fig.\u0026nbsp;1a. Data filtering and quality assessment were performed in the R environment (version 4.2.2).\u003c/p\u003e\u003cp\u003eFor DIA-MS, any raw PG.Quantity values\u0026thinsp;\u0026lt;\u0026thinsp;100 were replaced with NA. Samples with high haemolysis (\u0026ge;\u0026thinsp;2) were removed, along with samples not quantified by PEA. One clinical sample (cancer) that was later identified to be post-treatment was also removed. A minimum of 70% coverage across all remaining clinical samples was applied.\u003c/p\u003e\u003cp\u003ePEA data from both studies (PEA 01 and PEA 02) were loaded and bridge-normalisation performed using the OlinkAnalyse package (version 3.8.2) available in R. Assays with QC warning were removed for all patients, as well as assays with incomplete data across the two studies. For bridging samples, data for PEA 02 was discarded and for assays with repeat measurements across panels, measurements with a warning were removed and the mean assay calculated across the panels. Patients with high haemolysis were removed to give the same population for comparison with DIA-MS.\u003c/p\u003e\u003cp\u003eDNA-APT ADAT file was loaded using the SomaDataIO package (version 6.1.0) in R; all 140 samples passed Somalogics quality assessment and 20% dilution values were taken forward for further analysis.\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec27\" class=\"Section3\"\u003e\u003ch2\u003eData imputation and normalisation\u003c/h2\u003e\u003cp\u003eFor non-ML related analysis, DIA-MS data for the 563 clinical samples were imputed using the MissForest algorithm with default parameters from the missingpy package (version 0.2.0, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/epsilon-machine/missingpy\u003c/span\u003e\u003cspan address=\"https://github.com/epsilon-machine/missingpy\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) available in the Python environment and adapted to be compatible with the ML workflow. MissForest is a python implementation of the missForest package in R\u003csup\u003e34\u003c/sup\u003e that imputes data using random forests typically used for imputation of data that is considered missing at random (MAR)\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e. Assessment of missingness often relies on grouping of samples according to clinical characteristics. In this case where introduction of bias from imputation should be avoided and not be dependent on ML target, a target-independent model that can be used to impute unseen (unlabeled) data was needed. Given missForest has previously been shown to be a high performing imputation algorithm for missing data imputation in LC-MS based proteomics\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e and does not depend on knowing clinical characteristics of each instance, this method of imputation was selected for DIA-MS imputation. Following imputation, data was Log\u003csub\u003e2\u003c/sub\u003e transformed and normalised using median centering to account for differences in sample loading.\u003c/p\u003e\u003cp\u003eImputation was not required for PEA or DNA-APT data. The NPX values for PEA are already on the Log\u003csub\u003e2\u003c/sub\u003e scale, while DNA-APT data was Log\u003csub\u003e2\u003c/sub\u003e transformed.\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec28\" class=\"Section2\"\u003e\u003ch2\u003ePost-analysis quality, missingness and dynamic range assessment\u003c/h2\u003e\u003cp\u003eFor DIA-MS, CVs were calculated using pooled plasma samples run across seven plates without normalisation/Log2 transformation using standard CV formula (standard deviation/mean*100). For PEA, NPX data are both normalised and on Log\u003csub\u003e2\u003c/sub\u003e scale making the previous CV formula unsuitable\u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e. For PEA and DNA-APT Log\u003csub\u003e2\u003c/sub\u003e transformed ANML values, the following formula was used:\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:{CV}_{i}=100\\bullet\\:\\sqrt{{e}^{{{Sln}_{i}}^{2}}-1}\\:,\\:\\:\\:where\\:\\:{Sln}_{i}=ln2\\bullet\\:{SD}_{isk}\\:.\\:$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eMissingness in DIA-MS data was assessed using all pre-treatment clinical samples with an NA for PG.Quantity. For PEA, missing data is not reported. Instead, values below the lower limit of detection (LLOD) were replaced with NA and classed as missing. Average missingness for each protein was reported as mean\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation.\u003c/p\u003e\u003cp\u003eTo assess dynamic range, known concentrations of proteins according to the Human Protein Atlas (HPA) (v24.0)\u003csup\u003e17\u003c/sup\u003e were mapped to UniProt IDs in PEA and DIA-MS data. For ProteinGroups in DIA-MS data the first protein was taken. Where blood concentration from mass spectrometry was not available, immunoassay known concentrations were used. \u0026ldquo;Secretome location\u0026rdquo; was used to identify secretion location of proteins predicted to be secreted. Prediction of secreted protein in HPA was performed using amino acid sequence and location derived from literature/available data\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec29\" class=\"Section2\"\u003e\u003ch2\u003eUnivariate and exploratory analysis\u003c/h2\u003e\u003cp\u003ePrincipal component analysis (PCA) was performed using prcomp() available in the stats package available in R (version 4.4.1). Default parameters were used except for scale\u0026thinsp;=\u0026thinsp;TRUE to equalise variance contribution of features and prevent bias towards high-variance features.\u003c/p\u003e\u003cp\u003ePartial least squares regression (PLS) was used to distinguish lung cancer cases from healthy controls using the plsr() function from the pls package in R, with cross-validation (validation = \u0026ldquo;CV\u0026rdquo;) and scale\u0026thinsp;=\u0026thinsp;TRUE.\u003c/p\u003e\u003cp\u003eT-distributed stochastic neighbourhood embedding (t-SNE) was performed using Rtsne from Rtsne (version 0.17) package in R. Uniform Manifold Approximation and Projection (UMAP) was performed using umap() from umap package (version 0.2.10.0) in R.\u003c/p\u003e\u003cp\u003eFor univariate analysis a linear model was fit to each protein using lmFit() from limma, followed by empirical Bayes adjustment using eBayes() from limma to improve estimation of variance\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. Differentially expressed proteins were extracted using decideTests() with p.value\u0026thinsp;=\u0026thinsp;0.05, lfc\u0026thinsp;=\u0026thinsp;1 and adjustment.method = \u0026ldquo;fdr\u0026rdquo;. Summary statistics were extracted using topTable(). Proteins were considered differentially expressed if Log2 fold change\u0026thinsp;\u0026gt;\u0026thinsp;1 and FDR-adjusted p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05. Results are visualised as volcano plots (Fig.\u0026nbsp;2e,f and Extended Data Fig.\u0026nbsp;3c,d) and heatmaps (Supp Fig.\u0026nbsp;2B, 2C and Fig.\u0026nbsp;6A).\u003c/p\u003e\u003cp\u003eOver-representation analysis (ORA) and gene set enrichment analysis (GSEA)\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e of Gene Ontology (GO) biological processes\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e were carried out using enirchGO() and gseGO() from clusterProfiler package (version 4.12.2) available in R\u003csup\u003e40\u003c/sup\u003e. The analysis was performed across sub-cohorts: For protein groups the first ID was used, pvalueCutoff\u0026thinsp;=\u0026thinsp;0.05 and pAdjustMethod = \u0026ldquo;fdr\u0026rdquo;, human database was used (org.Hs.eg.db) and minGSsize\u0026thinsp;=\u0026thinsp;5. Semantic similarity analysis was used to remove redundant terms, applied using the mgoSim() and clusterSim() functions available from clusterProfiler. Data was visualised as heatmaps.\u003c/p\u003e\u003cp\u003eSpearman correlation across platforms was chosen due to it not being dependent on scale and calculated using cor() function in R for the 152 patient cohort quantified using all three methods. For assessing significance of correlation \u003cem\u003ep\u003c/em\u003e-values were FDR-adjusted. For DNA-APT data a second method of normalisation was used (median centering) as an alternative to ANML. Spearman correlation was calculated on both median-centered and ANML normalised datasets, features with Spearman correlation \u003cem\u003eR\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.4 in the appropriate comparison were removed and only the intersection between the two normalisation methods were considered. This would account for bias potentially introduced following ANML normalisation\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eUnsupervised, hierarchical clustering of differentially abundant proteins, ORA/GSEA results, Shapley value interactions and Spearman correlation co-efficient were performed using dist() and hclust() available in stats package in R and clusters generated using cutree() function in R.\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eData visualisation\u003c/h3\u003e\n\u003cp\u003eAll data visualisations were generated in the R environment with the exception of networks. For heatmaps the heatmap.2() function from the gplots package was used with the exception of Fig.\u0026nbsp;4C visualised using ggplot2. Venn diagrams were generated using ggvenn. All other plots were generated using ggplot2.\u003c/p\u003e\u003cdiv id=\"Sec31\" class=\"Section2\"\u003e\u003ch2\u003eML analysis for feature selection\u003c/h2\u003e\u003cp\u003eAll ML data generation was performed using Python (version 3.10). All models undergo the same sample stratification, train/holdout split as well as normalisation and imputation. Following this, different combinations of model selection, feature de-selection, augmentation, feature selection and model optimisation are trained on training data only to give rise to the 75 different models produced for each sub-cohort of each dataset, as described in Extended Data Fig.\u0026nbsp;3a. Each step in the ML process is detailed below.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec32\" class=\"Section2\"\u003e\u003ch2\u003eSample stratification and data split\u003c/h2\u003e\u003cp\u003eThe 563 clinical samples that passed pre-processing were stratified using sex\u0026thinsp;\u0026gt;\u0026thinsp;age-range (10 year intervals)\u0026thinsp;\u0026gt;\u0026thinsp;stage (control vs early vs late LC)\u0026thinsp;\u0026gt;\u0026thinsp;histology (NSCLC vs SCLC)\u0026thinsp;\u0026gt;\u0026thinsp;smoking history (non-heavy vs heavy-smoker). An 80/20% split was chosen for generating train/holdout data; Training data was used for training models and analysis, while the 20% holdout dataset was used to assess model generalisability and evaluate feature importances.\u003c/p\u003e\u003cp\u003eThis single train/holdout split for healthy control vs lung cancer samples was used to generate sub-cohorts for all controls vs early only lung cancer patients (Stage 0\u0026ndash;2), late lung cancer patients (Stage 3\u0026ndash;4), NSCLC only, SCLC only, LUAD only and LUSC only. For male only and female only models, only male and female healthy controls and lung cancer patients were included, respectively. Using the same split for different cohorts allowed comparison of sub-cohort analyses as well as maintaining the unseen nature of the holdout dataset.\u003c/p\u003e\u003cdiv id=\"Sec33\" class=\"Section3\"\u003e\u003ch2\u003eNormalisation, imputation and scaling for ML analysis\u003c/h2\u003e\u003cp\u003eData was median centered and Log2 transformed. missForest() from missingpy was used to impute training data as described previously. This model was fit on the training data only, and the subsequent model used to impute the training data and unseen holdout data. Fitting the imputation model on the training data only and applying to holdout avoids data-leakage between the two datasets and subsequent over-inflated ML performances.\u003c/p\u003e\u003cp\u003eSimilarly, scaling using StandardScaler() from scikit-learn (version 1.3.0, used for all scikit-learn functions) whereby the fit was applied to training only and used to transform train and holdout.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec34\" class=\"Section3\"\u003e\u003ch2\u003eModel selection\u003c/h2\u003e\u003cp\u003eTree-based models are suitable for handling high-dimension datasets, maintaining feature interactions and capturing non-linear relationships, unlike linear models often used to build ML models on omics datasets. Extreme gradient boosting (XGBoost), a tree-based model, was chosen to classify healthy controls vs lung cancer (and sub-cohorts). XGBClassifier() function from the XGBoost Python model (version 1.7.6) was used.\u003c/p\u003e\u003cp\u003eBalanced random forest (BRF) were chosen as an alternative tree-based model to XGBoost. BRF handles class imbalances by randomly under-sampling the majority class and maintaining balanced class representation during training, unlike random forests (RF) which favour the majority class. The use of bootstrapping in RF/BRF also prevents overfitting. BalancedRandomForestClassifier() from the Python package imblearn (version 0.12.2, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/scikit-learn-contrib/imbalanced-learn\u003c/span\u003e\u003cspan address=\"https://github.com/scikit-learn-contrib/imbalanced-learn\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was used to build BRF models.\u003c/p\u003e\u003cp\u003eHyperparameters were determined during model-optimisation.\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\n\u003ch3\u003eFeature de-selection\u003c/h3\u003e\n\u003cp\u003eFeature de-selection is the process by which features with low information content about the target (control vs lung cancer) are removed, with the aim of improving ML performance and reducing overfitting. Two methods of feature de-selection were used along with no feature de-selection.\u003c/p\u003e\u003cp\u003eThe first method uses mutual information (MI) that assesses the dependency between two variables (cancer status and protein abundance) that is capable of capturing non-linear relationships. The mutual_info_classif() was used to compute MI and was cross-validated using RepeatedStratifiedKFold(), stratified according to target (healthy control vs LC) with n_splits\u0026thinsp;=\u0026thinsp;2 and n_repeats\u0026thinsp;=\u0026thinsp;25. Any feature with cross-validated MI less than 3.5% was considered to not contain information related to classification, in this case lung cancer case, and was subsequently removed from the feature panel.\u003c/p\u003e\u003cp\u003eAlternatively, BorutaShap a two-step algorithm that combines the Boruta Algorithm\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e with Shapley values\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e was used to perform feature de-selection. Boruta works by creating shadow features by random permutation using actual features. By comparing feature importances of shadow and actual features, features where shadow features out-perform can be removed. In the case of BorutaShap, feature importances are calculated using Shapley values for each sample rather than model importances\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. Shapley values were derived from cooperative game theory and applied to ML models to assess feature importance. Shapley values fairly attribute a \"reward\" among features based on their individual contributions to the prediction. This approach was implemented using the BorutaShap function from the BorutaShap package (version 1.0.17, (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/Ekeany/Boruta-Shap\u003c/span\u003e\u003cspan address=\"https://github.com/Ekeany/Boruta-Shap\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), using an XGBoost model with importance_measure = \u0026ldquo;shap\u0026rdquo;.\u003c/p\u003e\n\u003ch3\u003eAugmentation\u003c/h3\u003e\n\u003cp\u003eIn some models augmentation was applied to balance classes between healthy controls and lung cancer cases by up-sampling the minority class. A two-step process was implemented to augment the minority class. Firstly, synthetic minority over-sampling technique (SMOTE) was used to generate synthetic samples from the minority class. SMOTE selects a random point from the minority class and interpolates a new datapoint between it and a neighbour actual sample, maintaining correlations/relationships between features in the synthetic data\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. Tomek links are then used to remove a portion of the majority class with overlapping decision boundaries with the augmented minority class. SMOTE-Tomek augmentation was applied to training data following feature de-selection using the SMOTETomek() function available in imbalanced-learn (version 0.12.2).\u003c/p\u003e\u003cdiv id=\"Sec37\" class=\"Section2\"\u003e\u003ch2\u003eRecursive feature addition (RFA) with Shapley values\u003c/h2\u003e\u003cp\u003eProteomics models that require large feature sets have limited clinical utility if needing to be translated onto alternative platforms. For example, immunoassay approaches require small panels for single-plex assays (such as ELISA) up to tens of analytes for multiplexed assays. Alternatively targeted mass-spectrometry methods such as single-reaction, multiple-reaction and parallel reaction monitoring (SRM, MRM and PRM, respectively) can be used, however such methods depending on instrumentation can also be limited to as little as 15 peptides\u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eRecursive feature addition (RFA) was used to build models that rely on a small subset (\u0026lt;\u0026thinsp;30) of proteins. RFA first selected the highest importance feature according to cross validated Shapley values. Using mean absolute Shapley value, high importance features were sequentially added and model performance evaluated to select the smallest subset of features that achieves the highest performance. RFA was coupled with repeated stratified cross validation (RSCV), as described for BorutaShap, to give more robust performance estimates across different feature sets.\u003c/p\u003e\u003cp\u003eTraditionally, model feature importances (such as those that come from XGBoost or BRF) are used to rank features and determine order of feature addition in RFA. The use of Shapley values instead of model feature importances captures both global (overall) and local (sample-specific) attributions, as opposed to only global. This ensures a fair and consistent distribution of feature importance and fairly distributes importance among correlated features. Another key advantage is their ability to account for feature interactions, revealing how features work together to give a prediction rather than evaluating them in isolation. Shapley values were calculated using the TreeExplainer() function from the Shap Python package, a fast implementation of tree-based Shapley value (TreeSHAP) generation using interventional perturbations to maintain dependencies between features. As an alternative Shapley value method, the package EjectSHAP was also used\u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e. EjectSHAP works similarly to TreeSHAP with the exception of how it attributes importances to features that never contributed to a prediction; if a branch or single node in a tree is not used to make a prediction, the importance of features within that tree are zero (i.e. \u0026ldquo;ejected\u0026rdquo;).\u003c/p\u003e\u003cp\u003eTwo metrics are used to evaluate model performance, area under the receiver operator characteristic curve (AUROC) and Matthews correlation coefficient (MCC). These metrics were evaluated on the top 30 features in XGBoost models and top 50 features in BRF, which often required higher numbers of features to converge during RFA. ROC curves plot the true positive rate (TPR/sensitivity) vs false positive rate (FPR) at different thresholds. AUC measures the area under the ROC curve; higher AUROC indicates higher model performance. MCC on the other hand utilises true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN), giving a more meaningful performance in cases where data is imbalanced. MCC can be interpreted as the correlation between the true labels and the predicted labels for binary classification problems\u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e,\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e. MCC was calculated using the matthews_corrcoef() function in scikit-learn, using the following formula:\u003c/p\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:MCC\\:=\\)\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\frac{TP\\bullet\\:TN-FP\\bullet\\:FN}{\\sqrt{(TP+FP)\\bullet\\:(TP+FN)\\bullet\\:(TN+FP)\\bullet\\:(TN+FN)}}\\:.\\:\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003cp\u003eMCC and AUROC performance gives rise to five different models. The first maximising mean AUROC across folds in RFA, the second maximising mean MCC across folds, the third approach maximises the stability of MCC across folds. MCC stability was calculated by dividing the mean of MCC across folds by the standard deviation across folds. This will give a subset of features for which MCC were minimally impacted by the fold used in RSCV. Finally, a ttest MCC and ttest AUROC feature set are generated to give a smaller subset of features, using ttest_ind_from_stats() available in scipy (version 1.11.1).\u003c/p\u003e\u003cp\u003eIn some cases, mean AUROC, MCC or MCC stability, or ttest MCC and AUROC might converge on the same subset of features and therefore produce less than five models, while in some cases all five RFA performance metrics give rise to models with different feature sets.\u003c/p\u003e\u003cdiv id=\"Sec38\" class=\"Section3\"\u003e\u003ch2\u003eModel optimisation\u003c/h2\u003e\u003cp\u003eXGBoost models require fine-tuning of model hyperparameters to avoid overfitting/underfitting on training data. For this, NSGAIISampler() from Optuna (version 3.6.1) Python module was used for multi-objective optimisation of model hyperparameters.\u003c/p\u003e\u003cp\u003eRFA using baseline XGBoost or BRF hyperparameters were used to select top 30 features. Baseline hyperparameters were chosen to mitigate the overfitting potential of the models. For the XGBoost models we limited the max_depth to 2, the n_estimators to 200, the learning rate to 0.05, lambda to 0, and setting the scale_pos_weight as the square root of the ratio of number of controls to cancer cases. For the BRF models we limited the max_depth to 10, the n_estimators to 500 and used sampling without replacement.\u003c/p\u003e\u003cp\u003eA model using these features was then optimised in Optuna to maximise three combinations of model performance metrics: MCC only, MCC\u0026thinsp;+\u0026thinsp;sensitivity at 99% specificity (Sens@90%Spec) and MCC\u0026thinsp;+\u0026thinsp;AUROC. If followed by optimisation, RFA was repeated using the optimised model hyperparameters to give a selection of features as described previously. Optuna was selected over other optimisation tools (such as GridSearchCV from Scikit-learn) as a result of its efficiency in searching higher-dimensional parameters spaces due to its Bayesian optimisation underpinning. Optimised model-hyperparameters were chosen from a local neighbourhood of similarly performing models, improving the stability of model sensitivity to minor hyperparameter perturbations.\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec39\" class=\"Section2\"\u003e\u003ch2\u003eModel evaluation\u003c/h2\u003e\u003cp\u003eFollowing training of the models, each model was evaluated using the unseen holdout test set. Assessment of model generalisability, rather than performance, promoted the refinement of features based on similar performance across train/holdout data with the aim of improving generalisability in future unseen data. To do this in a high throughput manner (1,350 models to evaluate), 95% confidence intervals (CI) were generated for the training data during model training. These 95% CI evaluated model performance during training using RSCV with n_splits\u0026thinsp;=\u0026thinsp;4 (equivalent in size to holdout data) and n_repeats\u0026thinsp;=\u0026thinsp;25. 95% CI were calculated for the following performance metrics: MCC, AUROC, Sens@90%Spec, Sens@95%Spec and Sens@99%Spec. These specificities were chosen to ensure reliable prediction in a low-prevalence setting, where high specificity reduces FPRs while maintaining high TPR. If four out of five of the performance metrics in holdout were within the 95% CI of training data a model was considered to have generalised, and was taken forward for further interpretation.\u003c/p\u003e\u003cdiv id=\"Sec40\" class=\"Section3\"\u003e\u003ch2\u003eModel interpretability (XAI)\u003c/h2\u003e\u003cp\u003eShapley values for each model were calculated as described during RFA. By calculating the mean of absolute Shapley values the contribution of each feature to all predictions can be interpreted. Within each generalisable model, features were ranked from largest mean absolute Shapley value to lowest. The average rank of a feature across generalisable models was used to understand the most important features for each sub-cohort. All data used in plotting Shapley value-related data comes from the SHAP module in Python.\u003c/p\u003e\u003cp\u003eTo assess the interactions between features, rather than the importances of a single feature, Shapley value interactions were used. Shapley value interactions capture pairwise interactions by assessing how the presence of one feature influences the contribution of a different feature to model prediction. Shapley value interactions aim to improve model-interpretability while identifying synergies and redundancies between features\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e. For a single representative healthy control vs lung cancer model, Shapley value interactions were calculated using the shap_interaction_values method from the TreeExplainer class from the SHAP module in Python. High values for Shapley value interaction equate to a higher proportion of Shapley value explained by the interaction between the two features. Shapley value interactions were normalised to the sum of all Shapley value interactions for both DIA-MS and PEA data to account for variation in scale of Shapley values between the two datasets. Normalised Shapley value interactions are visualised using a heatmap (Extended Data Fig.\u0026nbsp;6c,d), while a network of interactions between features accounting for more than 5% of actual Shapley value was visualised using Cytoscape (version 3.10.3)\u003csup\u003e49\u003c/sup\u003e (Fig.\u0026nbsp;4d,e).\u003c/p\u003e\u003cp\u003e\u003cb\u003eML analysis for cross-platform concordance\u003c/b\u003e\u003c/p\u003e\u003cp\u003eML models were used to assess the ability of features with concordance across the three platforms to predict lung cancer using DIA-MS and PEA data. Two panels were fit to both datasets; the first panel contained 18 features that were concordant (definition of concordance used: Spearman \u003cem\u003eR\u003c/em\u003e\u003csup\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sup\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.4) across DIA-MS, PEA and DNA-APT. The second panel for DIA-MS were features concordant between DIA-MS and DNA-APT that if present were also concordant in PEA, and for PEA those that were concordant between PEA and DNA-APT that if present were also concordant in DIA-MS. This accounted for non-overlapping features between DIA-MS and PEA.\u003c/p\u003e\u003cp\u003eA logistic regression was chosen as the model for assessing concordance due to its interpretability and low number of parameters to fit, allowing easier comparison of performances across datasets and/or features. The LogisticRegressionCV class from Scikit-learn was used to fit a balanced logistic regression on training data from PEA/DIA-MS data using relevant features. Performances were evaluated on holdout data not used in training.\u003c/p\u003e\u003cp\u003e\u003cb\u003eComparing univariate feature selection with ML-based feature selection\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo compare ML-dependent feature selection with a univariate approach, as has been used in other studies\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e,\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e, an analysis of variance (ANOVA) was used to select features had the smallest FDR-adjusted p-values when comparing lung cancer cases with healthy controls in all data. Twenty features were selected for each cohort that had the smallest adjusted p-values according to an f_test from Scikit-learn and compared with the twenty features that had the highest Shapley value rank in that cohort. A balanced logistic regression was used to assess the ability of these features to predict lung cancer as previously described. MCC and AUROC from the two logistic regression models were compared with their respective metric from an XGBoost model (Extended Data Fig.\u0026nbsp;3e).\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003eData and code availability\u003c/p\u003e\n\u003cp\u003eAll data analysed and code used in this study is available upon request.\u003c/p\u003e\n\u003cp\u003eAcknowledgements\u003c/p\u003e\n\u003cp\u003eWe thank all participants in the CALIBRE study. We also thank members of the Princess Margaret Cancer Centre, Toronto General Hospital and Princess Margaret Hospital for their contribution to the study. We would like to thank members of the Target Discovery Institute, in particular Dr Iolanda Vendrell for their technical advice in this study. We thank team members at Randox for their support collecting and processing the PEA (Olink) data. We would like to thank all former and current employees of Oxford Cancer Analytics Ltd for their guidance in the preparation of this manuscript, especially Dr Honglei Huang and Dr Heinrich Roder.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eB.M.K. and R.F. were supported by the Chinese Academy of Medical Sciences (CAMS) Innovation Fund for Medical Science (CIFMS), China (grant number: 2024-I2M-2-001-1).\u003c/p\u003e\n\u003cp\u003eAuthor contributions\u003c/p\u003e\n\u003cp\u003eH.R.F and N.G performed data analysis and developed the ML workflow, produced all figures and tables and wrote the text of the manuscript. L.H developed the ML workflow and contributed to interpretation of the data and data analysis. I.K developed and optimised the DIA-MS workflow and experimental design, carried out sample preparation, data acquisition and pre-processing of DIA-MS dataset. Ella.M and Emma.M contributed to data analysis, interpretation including cohort selection and stratification. D.A.S and J.S contributed to conceptualisation of the study, cohort selection and experimental design. G.L, L.J.Z and D.P provided insights into clinical cohorts, sourced the clinical samples and performed clinical data abstraction in this study. M.D contributed to proteomic sample selection and preparation. B.K and R.F provided insights into proteomic methodologies and analysis. A.H contributed to conceptualisation and methodology of the study. P.J.L conceptualised, supervised and guided the design, analysis, interpretation of the study, and wrote the text of the manuscript. All authors contributed to the drafting of the manuscript.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCompeting interests\u003c/p\u003e\n\u003cp\u003eH.R.F, N.G, L.H, I.K, M.D, J.S, Emma.M, Ella.M, D.A.S, A.H, and P.J.L are current/former employees, shareholders, and/or share option holders of Oxford Cancer Analytics Ltd.\u003c/p\u003e\n\u003cp\u003eR.F and B.M.K are share option holders of Oxford Cancer Analytics Ltd.\u003c/p\u003e\n\u003cp\u003eOxford Cancer Analytics Ltd sponsored this study.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eWorld Health Organisation. Lung cancer. \u003cem\u003eWorld Health Organisation\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.who.int/news-room/fact-sheets/detail/lung-cancer\u003c/span\u003e\u003cspan address=\"https://www.who.int/news-room/fact-sheets/detail/lung-cancer\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAmerican Cancer Society. Lung Cancer Survival Rates. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.cancer.org/cancer/types/lung-cancer/detection-diagnosis-staging/survival-rates.html https://www.cancer.org/cancer/types/lung-cancer/detection-diagnosis-staging/survival-rates.html\u003c/span\u003e\u003cspan address=\"https://www.cancer.org/cancer/types/lung-cancer/detection-diagnosis-staging/survival-rates.html https://www.cancer.org/cancer/types/lung-cancer/detection-diagnosis-staging/survival-rates.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eReduced Lung-Cancer Mortality with Low-Dose Computed Tomographic Screening. \u003cem\u003eNew England Journal of Medicine\u003c/em\u003e 365, (2011).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ede Koning, H. J. \u003cem\u003eet al.\u003c/em\u003e Reduced Lung-Cancer Mortality with Volume CT Screening in a Randomized Trial. \u003cem\u003eNew England Journal of Medicine\u003c/em\u003e 382, (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJemal, A. \u0026amp; Fedewa, S. A. Lung cancer screening with low-dose computed tomography in the United States \u0026ndash;\u0026thinsp;2010 to 2015. \u003cem\u003eJAMA Oncol\u003c/em\u003e 3, (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDickson, J. L. \u003cem\u003eet al.\u003c/em\u003e Uptake of invitations to a lung health check offering low-dose CT lung cancer screening among an ethnically and socioeconomically diverse population at risk of lung cancer in the UK (SUMMIT): a prospective, longitudinal cohort study. \u003cem\u003eLancet Public Health\u003c/em\u003e 8, (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLoPiccolo, J., Gusev, A., Christiani, D. C. \u0026amp; J\u0026auml;nne, P. A. Lung cancer in patients who have never smoked \u0026mdash; an emerging disease. \u003cem\u003eNature Reviews Clinical Oncology\u003c/em\u003e vol. 21 Preprint at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41571-023-00844-0\u003c/span\u003e\u003cspan address=\"10.1038/s41571-023-00844-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChin, R. I. \u003cem\u003eet al.\u003c/em\u003e Detection of Solid Tumor Molecular Residual Disease (MRD) Using Circulating Tumor DNA (ctDNA). \u003cem\u003eMolecular Diagnosis and Therapy\u003c/em\u003e vol. 23 Preprint at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s40291-019-00390-5\u003c/span\u003e\u003cspan address=\"10.1007/s40291-019-00390-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChaudhuri, A. A. \u003cem\u003eet al.\u003c/em\u003e Early detection of molecular residual disease in localized lung cancer by circulating tumor DNA profiling. \u003cem\u003eCancer Discov\u003c/em\u003e 7, (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChabon, J. J. \u003cem\u003eet al.\u003c/em\u003e Integrating genomic features for non-invasive early lung cancer detection. \u003cem\u003eNature\u003c/em\u003e 580, (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKelly-Spratt, K. S. \u003cem\u003eet al.\u003c/em\u003e Plasma proteome profiles associated with inflammation, angiogenesis, and cancer. \u003cem\u003ePLoS One\u003c/em\u003e 6, (2011).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePitteri, S. J. \u003cem\u003eet al.\u003c/em\u003e Tumor microenvironment-derived proteins dominate the plasma proteome response during breast cancer induction and progression. \u003cem\u003eCancer Res\u003c/em\u003e 71, (2011).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBhardwaj, M., Terzer, T., Schrotz-King, P. \u0026amp; Brenner, H. Comparison of proteomic technologies for blood-based detection of colorectal cancer. \u003cem\u003eInt J Mol Sci\u003c/em\u003e 22, (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNg, S., Masarone, S., Watson, D. \u0026amp; Barnes, M. R. The benefits and pitfalls of machine learning for biomarker discovery. \u003cem\u003eCell and Tissue Research\u003c/em\u003e vol. 394 Preprint at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s00441-023-03816-z\u003c/span\u003e\u003cspan address=\"10.1007/s00441-023-03816-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShen, S. Y. \u003cem\u003eet al.\u003c/em\u003e Sensitive tumour detection and classification using plasma cell-free DNA methylomes. \u003cem\u003eNature\u003c/em\u003e 563, (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKhodayari Moez, E. \u003cem\u003eet al.\u003c/em\u003e Circulating proteome for pulmonary nodule malignancy. \u003cem\u003eJNCI: Journal of the National Cancer Institute\u003c/em\u003e 115, (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eUhl\u0026eacute;n, M. \u003cem\u003eet al.\u003c/em\u003e The human secretome. \u003cem\u003eSci Signal\u003c/em\u003e 12, (2019).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLundberg, S. M. \u003cem\u003eet al.\u003c/em\u003e From local explanations to global understanding with explainable AI for trees. \u003cem\u003eNat Mach Intell\u003c/em\u003e 2, (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiang, H. \u003cem\u003eet al.\u003c/em\u003e LcProt: Proteomics-based identification of plasma biomarkers for lung cancer multievent, a multicentre study. \u003cem\u003eClin Transl Med\u003c/em\u003e 15, e70160 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eM\u0026uuml;ller, S. \u003cem\u003eet al.\u003c/em\u003e An Empirical Evaluation of the Rashomon Effect in Explainable Machine Learning. \u003cem\u003eArXiv\u003c/em\u003e (2023) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.2306.15786\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2306.15786\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKatz, D. H. \u003cem\u003eet al.\u003c/em\u003e Proteomic profiling platforms head to head: Leveraging genetics and clinical traits to compare aptamer- And antibody-based methods. \u003cem\u003eSci Adv\u003c/em\u003e 8, (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBirhanu, A. G. Mass spectrometry-based proteomics as an emerging tool in clinical laboratories. \u003cem\u003eClinical Proteomics\u003c/em\u003e vol. 20 Preprint at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12014-023-09424-x\u003c/span\u003e\u003cspan address=\"10.1186/s12014-023-09424-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEldjarn, G. H. \u003cem\u003eet al.\u003c/em\u003e Large-scale plasma proteomics comparisons through genetics and disease associations. \u003cem\u003eNature\u003c/em\u003e 622, (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePetrera, A. \u003cem\u003eet al.\u003c/em\u003e Multiplatform Approach for Plasma Proteomics: Complementarity of Olink Proximity Extension Assay Technology to Mass Spectrometry-Based Protein Profiling. \u003cem\u003eJ Proteome Res\u003c/em\u003e 20, (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWeber, M. A. \u003cem\u003eet al.\u003c/em\u003e Myoglobin plasma level related to muscle mass and fiber composition - A clinical marker of muscle wasting? \u003cem\u003eJ Mol Med\u003c/em\u003e 85, (2007).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCzarnowski, D. Syndecans in cancer: A review of function, expression, prognostic value, and therapeutic significance. \u003cem\u003eCancer Treatment and Research Communications\u003c/em\u003e vol. 27 Preprint at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ctarc.2021.100312\u003c/span\u003e\u003cspan address=\"10.1016/j.ctarc.2021.100312\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDinh, H. Q. \u003cem\u003eet al.\u003c/em\u003e Coexpression of CD71 and CD117 Identifies an Early Unipotent Neutrophil Progenitor Population in Human Bone Marrow. \u003cem\u003eImmunity\u003c/em\u003e 53, (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDaniels, T. R., Delgado, T., Rodriguez, J. A., Helguera, G. \u0026amp; Penichet, M. L. The transferrin receptor part I: Biology and targeting with cytotoxic antibodies for the treatment of cancer. \u003cem\u003eClinical Immunology\u003c/em\u003e vol. 121 Preprint at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.clim.2006.06.010\u003c/span\u003e\u003cspan address=\"10.1016/j.clim.2006.06.010\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2006).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHuman Protein Atlas. HAGH: Cancer. \u003cem\u003eHAGH\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.proteinatlas.org/ENSG00000063854-HAGH/cancer\u003c/span\u003e\u003cspan address=\"https://www.proteinatlas.org/ENSG00000063854-HAGH/cancer\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang, W. \u0026amp; Rubin, J. B. Treating sex and gender differences as a continuous variable can improve precision cancer treatments. \u003cem\u003eBiol Sex Differ\u003c/em\u003e 15, 35 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBudnik, B., Amirkhani, H., Forouzanfar, M. H. \u0026amp; Afshin, A. Novel proteomics-based plasma test for early detection of multiple cancers in the general population. \u003cem\u003eBMJ Oncology\u003c/em\u003e 3, (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBujang, M. A. \u0026amp; Adnan, T. H. Requirements for Minimum Sample Size for Sensitivity and Specificity Analysis. \u003cem\u003eJOURNAL OF CLINICAL AND DIAGNOSTIC RESEARCH\u003c/em\u003e 10, YE01\u0026ndash;YE06 (2016).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSearfoss, R. \u003cem\u003eet al.\u003c/em\u003e Impact of hemolysis on multi-OMIC pancreatic biomarker discovery to derisk biomarker development in precision medicine studies. \u003cem\u003eSci Rep\u003c/em\u003e 12, (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eStekhoven, D. J. \u0026amp; B\u0026uuml;hlmann, P. Missforest-Non-parametric missing value imputation for mixed-type data. \u003cem\u003eBioinformatics\u003c/em\u003e 28, (2012).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJin, L. \u003cem\u003eet al.\u003c/em\u003e A comparative study of evaluating missing value imputation methods in label-free proteomics. \u003cem\u003eSci Rep\u003c/em\u003e 11, (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCanchola, J. A. Correct Use of Percent Coefficient of Variation (%CV) Formula for Log-Transformed Data. \u003cem\u003eMOJ Proteom Bioinform\u003c/em\u003e 6, (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRitchie, M. E. \u003cem\u003eet al.\u003c/em\u003e Limma powers differential expression analyses for RNA-sequencing and microarray studies. \u003cem\u003eNucleic Acids Res\u003c/em\u003e 43, (2015).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSubramanian, A. \u003cem\u003eet al.\u003c/em\u003e Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles. \u003cem\u003eProc Natl Acad Sci U S A\u003c/em\u003e 102, (2005).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAleksander, S. A. \u003cem\u003eet al.\u003c/em\u003e The Gene Ontology knowledgebase in 2023. \u003cem\u003eGenetics\u003c/em\u003e 224, (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWu, T. \u003cem\u003eet al.\u003c/em\u003e clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. \u003cem\u003eInnovation\u003c/em\u003e 2, (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePietzner, M. \u003cem\u003eet al.\u003c/em\u003e Synergistic insights into human health from aptamer- and antibody-based proteomic profiling. \u003cem\u003eNat Commun\u003c/em\u003e 12, (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKursa, M. B. \u0026amp; Rudnicki, W. R. Feature selection with the boruta package. \u003cem\u003eJ Stat Softw\u003c/em\u003e 36, (2010).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLundberg, S. M. \u0026amp; Lee, S. I. Lundberg, S.M., Lee, S.I.: A unified approach to interpreting model predictions. In: Advances in Neural Information Processing Systems. pp. 4765\u0026ndash;4774 (2017). \u003cem\u003eNIPS-2017 Advances in Neural Information Processing Systems\u003c/em\u003e 32, (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChawla, N. V., Bowyer, K. W., Hall, L. O. \u0026amp; Kegelmeyer, W. P. SMOTE: Synthetic minority over-sampling technique. \u003cem\u003eJournal of Artificial Intelligence Research\u003c/em\u003e 16, (2002).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eVan Bentum, M. \u0026amp; Selbach, M. An introduction to advanced targeted acquisition methods. \u003cem\u003eMolecular and Cellular Proteomics\u003c/em\u003e vol. 20 Preprint at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/J.MCPRO.2021.100167\u003c/span\u003e\u003cspan address=\"10.1016/J.MCPRO.2021.100167\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCampbell, T. W., Roder, H., Georgantas III, R. W. \u0026amp; Roder, J. Exact Shapley values for local and model-true explanations of decision tree ensembles. \u003cem\u003eMachine Learning with Applications\u003c/em\u003e 9, (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChicco, D. \u0026amp; Jurman, G. The Matthews correlation coefficient (MCC) should replace the ROC AUC as the standard metric for assessing binary classification. \u003cem\u003eBioData Min\u003c/em\u003e 16, (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChicco, D. \u0026amp; Jurman, G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. \u003cem\u003eBMC Genomics\u003c/em\u003e 21, (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShannon, P. \u003cem\u003eet al.\u003c/em\u003e Cytoscape: A software Environment for integrated models of biomolecular interaction networks. \u003cem\u003eGenome Res\u003c/em\u003e 13, (2003).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDavies, M. P. A. \u003cem\u003eet al.\u003c/em\u003e Plasma protein biomarkers for early prediction of lung cancer. \u003cem\u003eEBioMedicine\u003c/em\u003e 93, (2023).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-7660411/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7660411/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eLung cancer is the leading cause of cancer mortality worldwide with 70% diagnosed late stage despite low-dose computed tomography (LDCT) screening availability. We combined data-independent acquisition mass-spectrometry (DIA-MS) and proximity extension assay (PEA) with explainable artificial intelligence (XAI)-led machine learning (ML) for plasma-based biomarker discovery. From a 490 lung cancer and 124 matched-control cohort, ML models were trained to predict lung cancer achieving an AUROC of 0.91 [95% CI: 0.88\u0026ndash;0.93] and 0.97 [95% CI: 0.92\u0026ndash;0.98] in DIA-MS and PEA, respectively. XAI characterised networks of model-consistent features primarily related to infection and inflammatory responses. We then introduced a DNA-aptamer proteomics method and identified a cross-platform concordance panel, with performances of 0.88 [95% CI: 0.80\u0026ndash;0.90] and 0.88 [95% CI: 0.81\u0026ndash;0.95] in DIA-MS and PEA, respectively. This study demonstrates that combining multi-dimensional proteomics with XAI-ML can characterise robust biomarker signatures.\u003c/p\u003e","manuscriptTitle":"Multidimensional proteomics and explainable AI feature selection identify cross-platform lung cancer molecular signature in blood plasma","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-31 10:31:14","doi":"10.21203/rs.3.rs-7660411/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"communications-medicine","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"commsmed","sideBox":"Learn more about [Communications Medicine](http://www.nature.com/commsmed)","snPcode":"43856","submissionUrl":"https://mts-commsmed.nature.com/cgi-bin/main.plex","title":"Communications Medicine","twitterHandle":"@commsmedicine","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Communications Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0aa36e4b-df0c-41b4-b23f-36e56559d073","owner":[],"postedDate":"October 31st, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":57217602,"name":"Biological sciences/Cancer"},{"id":57217603,"name":"Health sciences/Biomarkers"}],"tags":[],"updatedAt":"2025-10-31T10:31:14+00:00","versionOfRecord":[],"versionCreatedAt":"2025-10-31 10:31:14","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7660411","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7660411","identity":"rs-7660411","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00