Predicting High Confidence ctDNA Somatic Variants with Ensemble Machine Learning Models

preprint OA: gold CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Circulating tumour DNA (ctDNA) is a minimally invasive cancer biomarker that can be used to inform treatment of cancer patients. The utility of ctDNA as a cancer biomarker depends on the ability to accurately detect somatic variants associated with cancer. Accurate somatic variant detection in circulating cell free DNA (cfDNA) NGS data requires filtering strategies to remove germline variants, and NGS artifacts. Rule-based variant filtering methods either remove a substantial number of true positive ctDNA variants along with false variant calls or retain an implausibly large number of total variants. Machine Learning (ML) enables identification of complex patterns which may improve ability to distinguish between real somatic ctDNA variants and false positive calls. We built two Random Forest (RF) models for predicting high confidence somatic ctDNA variants in low and high depth cfDNA NGS data. Low depth models were fitted and evaluated on whole exome sequencing (WES) cfDNA data at depths of approximately 10X while the high depth data was sequenced at approximately 500X. Both models utilise a set of 15 features from variants detected by bcftools, FreeBayes, LoFreq and Mutect2. High confidence ground truth sets were obtained from matched tissue biopsy samples. We benchmarked our models against rule-based filtering with a set of hard, medium, and soft thresholds. Precision-recall curves showed the high depth model outperformed rule-based filtering at all thresholds in validation data (PR-AUC 0.71). Partial dependence plots showed membership in the COSMIC database, absence from the dbSNP common variants database, and increasing read depth increased mean probability of high confidence somatic variant prediction in both models. Our results demonstrate the utility of supervised ML models for filtering variants in cfDNA data.
Full text 147,850 characters · extracted from preprint-html · click to expand
Predicting High Confidence ctDNA Somatic Variants with Ensemble Machine Learning Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Predicting High Confidence ctDNA Somatic Variants with Ensemble Machine Learning Models Rugare Maruzani, Anna Fowler, Liam Brierley, Andrea Jorgensen This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5735554/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 26 May, 2025 Read the published version in Scientific Reports → Version 1 posted 12 You are reading this latest preprint version Abstract Circulating tumour DNA (ctDNA) is a minimally invasive cancer biomarker that can be used to inform treatment of cancer patients. The utility of ctDNA as a cancer biomarker depends on the ability to accurately detect somatic variants associated with cancer. Accurate somatic variant detection in circulating cell free DNA (cfDNA) NGS data requires filtering strategies to remove germline variants, and NGS artifacts. Rule-based variant filtering methods either remove a substantial number of true positive ctDNA variants along with false variant calls or retain an implausibly large number of total variants. Machine Learning (ML) enables identification of complex patterns which may improve ability to distinguish between real somatic ctDNA variants and false positive calls. We built two Random Forest (RF) models for predicting high confidence somatic ctDNA variants in low and high depth cfDNA NGS data. Low depth models were fitted and evaluated on whole exome sequencing (WES) cfDNA data at depths of approximately 10X while the high depth data was sequenced at approximately 500X. Both models utilise a set of 15 features from variants detected by bcftools, FreeBayes, LoFreq and Mutect2. High confidence ground truth sets were obtained from matched tissue biopsy samples. We benchmarked our models against rule-based filtering with a set of hard, medium, and soft thresholds. Precision-recall curves showed the high depth model outperformed rule-based filtering at all thresholds in validation data (PR-AUC 0.71). Partial dependence plots showed membership in the COSMIC database, absence from the dbSNP common variants database, and increasing read depth increased mean probability of high confidence somatic variant prediction in both models. Our results demonstrate the utility of supervised ML models for filtering variants in cfDNA data. Biological sciences/Cancer Biological sciences/Computational biology and bioinformatics Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Background Circulating tumour DNA (ctDNA) is fragmented DNA released into the bloodstream by tumour cells. ctDNA has been shown to be a predictor of disease, relapse and prognosis for cancer patients [ 1 – 3 ]. ctDNA is minimally invasive to sample and has a short half-life, enabling near real-time monitoring for patients at risk of relapse [ 4 ]. Tumours can display intratumoral and spatial heterogeneity [ 5 , 6 ] however as all tumour tissues release ctDNA, the entire genetic profile of a patient’s tumours can be analysed. This is in contrast to tissue biopsy where only genetic information at the biopsy site can be analysed [ 7 ]. The utility of ctDNA as a cancer biomarker depends on the ability to accurately detect somatic variants associated with cancer. ctDNA can be in low abundance in the pool of total cell free DNA (cfDNA), particularly in the early stages of disease or when monitoring for recurrence. Additionally, the abundance of ctDNA fragments released from subclones will be even lower. Ultimately, it is difficult to distinguish between real low frequency ctDNA variants and NGS artifacts [ 8 ]. Accurate somatic variant detection in cfDNA NGS data requires variant filtering methods. Accuracy of called variants is influenced by features including base quality and read depth. Base quality scores represent confidence in the called base by sequencing platforms. Read depth describes the total number of reads covering the variant site. Higher read depths and a higher number of reads supporting an alternate allele increase confidence a called variant is real. Rule-based filtering involves defining a set of thresholds for measures including read depth and base quality, and then discarding variants which fall below those thresholds. It also often involves discarding variants reported in population databases including the 1000 Genomes Project [ 9 ] and dbSNP [ 10 ]. Another approach to the variant filtering problem is consensus, or ensemble variant calling. Variant callers utilise a wide range of statistical methods and rely on differing assumptions about the input data to call variants. The result is that overlapping but different call sets are returned by different tools from the same input data. Variants commonly called by more than one caller have been shown to be more likely to be real somatic variants [ 11 ]. As such, ensemble variant calling aims to reduce the number of false positive calls by combining the output of multiple variant callers [ 12 ]. Naively, this may involve filtering out variants detected by a single caller [ 12 ]. Both these filtering strategies have limitations. Depending on the thresholds specified, rule-based filtering can discard a large fraction of true positive variants or retain an implausibly large number of variants post-filtering. Benchmarking studies have shown variant callers have low concordance in variants detected from the same input data [ 13 ]. Naturally, ensemble variant call sets tend to have high precision, with a trade-off in low recall [ 14 ]. Machine Learning (ML) algorithms can enable identification of complex, non-linear patterns which may improve ability to distinguish between real somatic ctDNA variants and false positive calls. Recently, ML tools for predicting somatic variants in cancer samples have been published [ 14 – 17 ]. However, few studies have focused on developing models for predicting somatic variants specifically in cfDNA NGS data. Jongbloed et al., (2023) [ 19 ] present an ML model for detecting somatic variants in cfDNA data from breast cancer patients. The model, however, was fitted on, and developed for, targeted sequencing data. Depending on the size of the panel, targeted sequencing can return orders of magnitude fewer variants compared to Whole Exome Sequencing (WES), reducing the resources required to validate calls, and consequently, the need for specialist filtering tools. Fitting supervised ML models requires a truth set to label positive examples. However, obtaining a truth set for predicting somatic variants in NGS data can be a challenge. This process may involve resource-intensive orthogonal sequencing or manual review of candidate variants. The use of multiple sequencing platforms and ultra-high sequencing depths often utilised for orthogonal validation [ 20 ] can prohibitively increase the costs associated with obtaining a truth set. Given the labour-intensive nature of manually reviewing variants [ 15 ], typically only a subset of candidate true positive variants are validated [ 16 ], potentially impacting the performance of the fitted model. One approach to obtaining a truth set is using common variants between matched tissue biopsy and the cfDNA sample. Tumour variants are much easier to detect in tissue than in ctDNA, and therefore the variants identified at low frequencies in ctDNA which are also confidently identified in the tissue sample are highly likely to be true positives. This method is less labour-intensive compared to manual curation and does not necessarily require independent sequencing technology. It therefore allows for a larger number of variants to be included in the truth set. Here, we present two ML models for predicting high confidence somatic variants in low and high depth cfDNA WES datasets. The high confidence truth set was defined as variants detected in a matched tissue sample after stringent filtering. We extracted features from the output of 4 variant callers and predicted high confidence somatic variants with Random Forest models. Our models do not rely on matched normal samples; we utilise allele frequency and population databases to filter germline variants. The models were trained to predict single nucleotide variants (SNVs) only as they are the most common variant type in cancer. Figure 1 illustrates the workflow for model training and evaluation. Our code is available at https://github.com/rugare-m/somVar . Methods Datasets We used a set of publicly available cfDNA WES samples to fit low and high depth ML models. The low depth model was trained and evaluated using samples from Dietz et al (2016) [ 21 ] ’s publication. This dataset included 6 cfDNA samples; each of the samples was sequenced with a matched tissue sample.The high depth model was trained and evaluated using samples from Li et al. (2022) [ 4 ] and one sample from Butler et al. (2015) [ 22 ], all sequenced with matched tissue samples. Table 1 shows the samples used to fit both models, including the accession numbers. Data pre-processing. To generate BAM files, FASTQ files were mapped to GRCh38 using BWA-MEM2 v2.2.1 [ 23 ] at default parameters. Output SAM files were converted to BAM file format using Samtools v1.2 [ 24 ]. Duplicate reads in BAM files were marked with GATK [ 25 ] MarkDuplicates v4.3.0.0, before recalibrating base quality scores with GATK BaseRecalibrator v4.3.0.0 and GATK ApplyBQSR v4.3.0.0. The sequencing depth profiles of all processed BAM files was visualised by plotting a histogram of depth at every position in the GRCh38 exome. Variant calling pipeline. Four variant callers were used to detect variants; bcftools v1.16 [ 26 ], FreeBayes v1.3.6 [ 27 ], LoFreq v2.1.5 [ 28 ], Mutect2 v4.3.0.0 [ 29 ]. All callers were run at default parameters. Output VCF files were decomposed to split multiallelic sites into individual records, and indels were removed. We discarded variants with allele frequencies between 0.4 and 0.6, and above 0.9 as these were assumed to be likely germline variants [ 18 ]. VCF files were annotated with strand bias, mapping quality, base quality, fragment length, read position, allele frequency and number of reads supporting alleles using GATK v4.3.0.0 VariantAnnotator. SnpSift v4.3t [ 30 ] was used to annotate coding variants reported in COSMIC v98 and common variants in dbSNP v151. For each sample, VCF files from all callers were merged into a single VCF file using GATK MergeVcfs. Obtaining truth sets. We used a set of high confidence variants as true positive calls. These variants were defined as merged tissue variants with read depth greater than 70 for low depth data (or greater than 500 for high depth data), with Phred quality scores > = 50, not reported in the common variants dbSNP v151 VCF file and reported in the COSMIC v98 coding mutations VCF. To minimise the number of germline variants in the truth set, variants with previously described allele frequencies were discarded. Machine learning features. A set of 15 features per variant in the merged cfDNA VCF files were extracted into tabular data from data in the BAM, reference FASTA and VCF files. BAM level features included read depth, strand bias, median fragment length of reads supporting alternate allele, mapping quality, allele frequency, and median distance of variant from end of reads supporting alternate allele. The read depth feature was standardised by removing mean and scaling to unit variance using scikit-learn v1.4.1 [ 31 ]. Reference FASTA features included GC percentage of reference sequence in a 20-nucleotide region flanking the variant (41 bases total), and a weighted homopolymer score calculated from the same region [ 32 ]. VCF features included a Boolean of presence of the variant in each of dbSNP and COSMIC databases, and bcftools, FreeBayes, LoFreq, and Mutect2 VCF files. Table 2 describes features used in the ML models. For the target, variants were labelled as high confidence somatic variants (HCSVs) if they were present in the truth VCF file and artefact variants (AVs) otherwise. Train, Test and Validation Data. One validation sample from each of the low and high depth datasets was randomly chosen. This was kept independent of the training and testing pipeline. Variants from the remaining samples were allocated to training and testing sets, with a 70:30 training: testing split. The models were optimised using the training and testing datasets and then used to predict high confidence somatic variants in the validation samples. We trained our models using the scikit-learn implementation of the Random Forest algorithm. The Random Forest algorithm fits decision trees on subsamples of input data and averages tree predictions to classify samples. Class balancing and hyperparameter optimisation loop. ML models fitted on imbalanced datasets perform poorly when predicting the minority class. These models tend to classify all new data as majority class, which is often the class of least interest [ 33 ]. Here, the ratio of true positive somatic variants to total variants per low depth sample was in the order of 1:1000. To mitigate against this, we treated the undersampling ratio as a parameter to be optimised in a nested loop that undersampled Artifact Variants (AVs) in Training Data to one of 6 ratios and optimised hyperparameters for models fitted on the undersampled Training Data (Fig. 1 ). In the outer loop, Training Data was undersampled to each of 1:10, 1:5, 2:5, 3:5, 4:5 and 1:1 HCSV to AV ratios using the RandomUnderSampler (RUS) algorithm in the Python library imbalanced-learn v0.12.0 [ 34 ]. RUS randomly downsamples AVs until the user-defined HCSVs to AV ratio is reached. The inner loop optimised hyperparameters for RF models fitted on the undersampled Training Data using RandomizedSearchCV in scikit-learn. The random search was set to evaluate 100 different hyperparameter combinations using 3-fold cross validation, and to score recall when outputting the best model. The best model from each of the 6 undersampling ratios was used to predict classes in Test Data, and precision and recall scores were calculated. Test Data was not undersampled to reflect real data where the ground truth is unknown. Prioritising recall, the model with the highest precision and recall scores on Test Data predictions was used to evaluate performance on the previously unseen Validation Data. Class balancing ratios and hyperparameters were separately optimised for low and high depth models. High Confidence Somatic Variant Probability Thresholds. The Random Forest algorithm predicts the probability of a variant being a HCSV. A default threshold of 0.5 is required to label variants as HCSV. Increasing the threshold increases the model precision score at the expense of recall. To reduce the false positive calls associated with high recall selected for in hyperparameter optimisation, we set the HCSV prediction threshold default to 0.75. Predictions in Validation Data were with the threshold value set to 0.75. Evaluating models on Validation Data. We produced confusion matrices on Validation Data for the exported models to show the raw numbers of true positive (TP), true negative (TN), false positive (FP) and false negative (FN) predictions by the models. Due to the large imbalance in classes, we opted to use precision-recall (PR) curves over receiver operator characteristics (ROC) to evaluate the models’ performance at different thresholds [ 14 ]. Feature importance analysis. We assessed features’ contribution to model performance using permutation feature importance in scikit-learn. This method shuffles a feature in Test Data n times and calculates the mean decrease of the model’s score on the permuted Test Data. Here, each feature was shuffled 30 times and we used Average Precision (AP) as the scoring metric - AP summarises the precision-recall curve (Table 3). Next, we calculated partial dependence of the top three features from permutation feature importance analysis in the Test Data. Partial dependence plots (PDPs) are calculated by varying a feature and observing resulting change in model-predicted marginal probability of HCSV, i.e. averaging out all other features. Benchmarking against rule-based filtering. We aimed to benchmark the RF models against rule-based filtering strategies using the SRR3401408 and SRR12083529 samples which were held out for validation. We used the RF models to predict high confidence variants on corresponding Validation Data and plotted PR curves. Next, we performed rule-based filtering of the merged validation sample cfDNA VCF files at three different thresholds; hard, medium, and soft, and plotted precision and recall on PR curves. Rule-based filtering at all thresholds included the previously described allele frequency filters to discard germline variants. Table 4 illustrates the thresholds applied at each filtering threshold. Results Sequencing depth profiles of matched tissue and Validation Data samples. We applied a read depth filter in the matched tissue samples to obtain a HCSV truth set and for benchmarking models against rule-based filtering. The thresholds for read depth are dependent on the sequencing depths of the data. To inform selection of the appropriate read depth thresholds, we plotted histograms of read depths at every position in the human exome for all samples. The low depth validation sample showed a mode read depth of approximately 10X (Figure 2 A). 10X read depth was selected as the medium depth threshold, with 5X and 20X as the soft and hard thresholds respectively. In the high depth validation sample, the mode was approximately 750X, which was selected as the medium depth threshold, with 500X and 1000X selected as soft and hard depth filter thresholds respectively (Figure 2B). Matched tissue samples showed read depth of approximately 50X and 200X in low and high depth data respectively. We defined high confidence variants as variants passing a stringent depth filter in matched tissue data. We therefore selected thresholds of 70X and 500X for low and high depth data truth sets respectively. Hyperparameter and class balancing optimisation. We fitted two Random Forest models for predicting high confidence somatic ctDNA variants. We optimised the hyperparameters of each model using a random search and 3-fold cross validation. Low depth models fitted on data undersampled to ratios of 2:5, 3:5, and 4:5 all returned the greatest performance, with 1.00 and 0.07 recall and precision scores. To select the best model from this set of 3, we compared the PR areas under curve (PR-AUCs) on Test Data predictions. The PR-AUC metric ranges from 0 to 1, with a perfect classifier returning a PR-AUC 1 while a model with no predictive power returns a PR-AUC of 0. Here, the model fitted on Training Data undersampled to 2:5 returned the highest score of 0.43. The 3:5 and 4:5 models returned PR-AUCs of 0.31 and 0.28 respectively. The model fitted on Training Data undersampled to 2:5 was selected as the final low depth model. Undersampling the low depth Training Data to a 2:5 ratio reduced AVs from 1,151,933 to 4,457. The high depth model fitted on Training Data undersampled to 1:5 ratio returned the highest precision and recall scores; 0.15 and 1.00 respectively (Table 6).Undersampling high depth Training Data to 1:5 ratio reduced AVs variants from 1,803,351 to 14,975. High Depth Random Forest model outperformed rule-based filtering. We evaluated the performance of both models using PR-AUC, and confusion matrices on Validation Data predictions. Low and high depth model PR-AUCs applied to these unseen data were 0.45 and 0.71 respectively (Figure 3 A-B). The high depth model outperformed soft, medium, and hard rule-based filtering thresholds at predicting HCSV in respective Validation Data (Figure 3B). At a 0.75 HCSV probability threshold, the low depth model returned a recall of 0.97 and precision of 0.11; and the high depth model returned a recall of 0.97 and precision of 0.26 in respective Validation Data predictions. Figures 3 C-D show the confusion matrices and PR-AUCs for low and high depth model predictions. Read Depth, COSMIC and dbSNP membership are key features for ML models. We used permutation feature importance to assess the contribution of features in predicting HCSV (Figure 4). Presence in COSMIC and dbSNP, and Read Depth were the top 3 features contributing to both models’ performance. Random permutation of COSMIC, dbSNP and Read Depth features resulted in a mean decrease in AP of 0.15, 0.13 and 0.07 respectively in low depth data predictions. In high depth data, the mean decrease in AP after permuting COSMIC, Read Depth and dbSNP were 0.63, 0.53 and 0.45 respectively. Permuting Median Alternate Fragment Length notably had a relatively low impact on AP in both model predictions. The PDPs showed variant presence in the COSMIC database results in the largest increase in mean probability of HCSV predictions for both models (Figure 5). Absence in dbSNP also increased mean probability of HCSV predictions in both models. Increase in read depth increased mean probability of HCSV predictions, however the increase is small and plateaus, particularly in the low depth model. Discussion Variant filtering is an important step in detecting somatic mutations in WES cfDNA NGS data. The aim is to filter out false positive calls without discarding too many real variants. Variant filtering in cfDNA data is complicated by the fact that ctDNA is typically in low abundance relative to healthy cfDNA. Ultimately true positive ctDNA variants occur at low allele frequencies that are difficult to distinguish from sequencing and PCR artifacts. We developed two ML models for predicting high confidence somatic ctDNA variants in low and high depth data. Our models were trained on WES samples with high confidence labels acquired from matched tissue samples. At a probability threshold of 0.75, both models correctly classified over 97% of HCSVs and outperformed all rule-based filtering thresholds. Hyperparameter optimisation for both models used recall as the scoring metric. While a balance of precision and recall is important, we determined recall is the more important metric in variant filtering of liquid biopsy NGS data. False positive variants can be removed in downstream orthogonal validation [ 15 , 35 ]; however, there is no way to rectify potentially actionable variants being discarded. Variants detected in a clinical context by any tool will likely require confirmation with Sanger sequencing for the foreseeable future [ 36 ]. The goal is to reduce the number of variants requiring confirmation, in turn reducing the resources required for validating variants, while retaining as many true positive variants as possible. The balance between precision and recall, however, can often be dependent on application of the tool. We therefore implemented a HCSV probability threshold parameter in our models to allow users to set the trade-off between precision and recall. We generated partial dependence plots for the top-ranking features for each of our models. PDPs provide insight into how ML models make predictions. In both models, presence in COSMIC increased the mean probability of HCSV predictions more than any other feature. This was as expected given that the COSMIC database is a well-curated catalogue of cancer-associated variants, and the ground truth contains COSMIC variants only. Absence from dbSNP also increased probability of HCSV predictions in both models. This was expected as variants described in dbSNP are unlikely to be cancer associated somatic variants. Read depth, mapping quality and base quality are routinely used features in ML models for predicting variants in NGS data [ 14 , 17 , 18 , 37 ]. In addition to these features, we included median fragment length of reads supporting the alternate allele in our models for predicting high confidence ctDNA somatic variants. Other researchers have previously demonstrated ctDNA fragments tend to be shorter than healthy cell free DNA [ 38 , 39 ]. True somatic ctDNA variants, therefore, are more likely to originate from shorter fragments compared to germline variants, and artefacts in normal cfDNA fragments. Permutation feature importance however demonstrated this feature had minimal effect on AP when permuted suggesting low predictive power in both models. The substantial overlap in the distribution of fragment lengths of ctDNA and healthy cell free DNA may explain this low predictive power [ 40 ]. One of the known sources of sequencing errors in Illumina platforms is base calling in homopolymer regions. After a run of the same base, Illumina platforms may erroneously substitute the first base after the homopolymer with the homopolymer base. Stoler and Nekrutenko, (2021) [ 41 ] found up to 5.3% of sequencing errors can be attributed to this effect. Repetitive regions are also associated with low mapping quality and base qualities [ 42 ]. We included the WHR feature in our models to filter out errors associated with repetitive sequences in the genome. Feature importance analysis, however, suggests WHR also has little predictive power. While both models showed good performance, our model training and evaluation pipeline had some limitations. Filtering out germline variants in somatic variant calling pipelines can be a challenge without a matched normal sample [ 43 ]. Here, we attempted to remove germline variants by discarding variants reported in dbSNP and variants with an allele frequency of between 0.4 and 0.6, and above 0.9. Using population frequency databases to filter germline variants is a well-established approach [ 44 , 45 ]. However, it is possible we discard some real HCSVs with the allele frequency filter. Another limitation in our approach was assembly of the truth set. Commonly detected variants in cfDNA and matched tissue samples after stringent filtering were labelled true positive high confidence variants. The matched tissue biopsy for each sample only represents a fraction of the tumour tissues releasing cfDNA, therefore the ground truth sets were necessarily incomplete. A possible consequence of this is that the false positive rate in our models is inflated if the models are predicting real HCSVs that are absent from the tissue biopsy sample, and so the truth set. Some approaches may be investigated to mitigate this limitation of our approach. The truth set may be derived from multiple tissue biopsies from the same patient. This approach would increase the probability of capturing all somatic variants compared to a single biopsy approach. The challenge of this approach is the increased sequencing costs and disadvantages associated with collecting biopsies from patients. Orthogonal validation of predicted variants may also be employed to more accurately assess model performance. However, the large number of high confidence variants returned from WES experiments may render orthogonal validation of all variants unfeasible as previously discussed. Declarations Ethics approval and consent to participate. Not applicable Consent for publication Not applicable Availability of data and materials Sequence data that support the findings of this study have been deposited in the European Nucleotide Archive with the primary accession code PRJNA318450, SRR12083523, SRR12083524, SRR12083526, SRR12083527, ERR855950, ERR855951 SRR12083529, SRR12083530. Competing interests The authors declare that they have no competing interests. Funding This work was funded by a Medical Research Council DiMeN Doctoral Training Partnership iCASE studentship, with industrial partners Nonacus. MRC Grant Reference Number: MR/R015902/1 References Asante D-B, Calapre L, Ziman M, Meniawy TM, Gray ES. Liquid biopsy in ovarian cancer using circulating tumor DNA and cells: Ready for prime time? Cancer Letters. 2020;468:59–71. Bortolini Silveira A, Bidard F-C, Tanguy M-L, Girard E, Trédan O, Dubot C, et al. Multimodal liquid biopsy for early monitoring and outcome prediction of chemotherapy in metastatic breast cancer. NPJ Breast Cancer. 2021;7:115. Wu X, Zhang Y, Hu T, He X, Zou Y, Deng Q, et al. A novel cell-free DNA methylation-based model improves the early detection of colorectal cancer. Molecular Oncology. 2021;15:2702–14. Li S, Zeng W, Ni X, Zhou Y, Stackpole ML, Noor ZS, et al. cfTrack: A Method of Exome-Wide Mutation Analysis of Cell-free DNA to Simultaneously Monitor the Full Spectrum of Cancer Treatment Outcomes Including MRD, Recurrence, and Evolution. Clin Cancer Res. 2022;28:1841–53. Chicard M, Colmet-Daage L, Clement N, Danzon A, Bohec M, Bernard V, et al. Whole-Exome Sequencing of Cell-Free DNA Reveals Temporo-spatial Heterogeneity and Identifies Treatment-Resistant Clones in Neuroblastoma. Clinical Cancer Research. 2018;24:939–49. Herzog H, Dogan S, Aktas B, Nel I. Targeted Sequencing of Plasma-Derived vs. Urinary cfDNA from Patients with Triple-Negative Breast Cancer. Cancers (Basel). 2022;14:4101. Bos MK, Angus L, Nasserinejad K, Jager A, Jansen MPHM, Martens JWM, et al. Whole exome sequencing of cell-free DNA – A systematic review and Bayesian individual patient data meta-analysis. Cancer Treatment Reviews. 2020;83:101951. Nix DA, Hellwig S, Conley C, Thomas A, Fuertes CL, Hamil CL, et al. The stochastic nature of errors in next-generation sequencing of circulating cell-free DNA. PLoS One. 2020;15:e0229063. Auton A, Abecasis GR, Altshuler DM, Durbin RM, Abecasis GR, Bentley DR, et al. A global reference for human genetic variation. Nature. 2015;526:68–74. Sayers EW, Barrett T, Benson DA, Bolton E, Bryant SH, Canese K, et al. Database resources of the National Center for Biotechnology Information. Nucleic Acids Res. 2011;39 Database issue:D38-51. Trevarton AJ, Chang JT, Symmans WF. Simple combination of multiple somatic variant callers to increase accuracy. Sci Rep. 2023;13:8463. Wang M, Luo W, Jones K, Bian X, Williams R, Higson H, et al. SomaticCombiner: improving the performance of somatic variant calling based on evaluation tests and a consensus approach. Sci Rep. 2020;10:12898. Hwang S, Kim E, Lee I, Marcotte EM. Systematic comparison of variant calling pipelines using gold standard personal exome variants. Scientific Reports. 2015;5:17875. Anzar I, Sverchkova A, Stratford R, Clancy T. NeoMutate: an ensemble machine learning framework for the prediction of somatic mutations in cancer. BMC Medical Genomics. 2019;12:63. Ainscough BJ, Barnell EK, Ronning P, Campbell KM, Wagner AH, Fehniger TA, et al. A deep learning approach to automate refinement of somatic variant calling from cancer sequencing data. Nat Genet. 2018;50:1735–43. Díaz-Navarro A, Bousquets-Muñoz P, Nadeu F, López-Tamargo S, Beà S, Campo E, et al. RFcaller: a machine learning approach combined with read-level features to detect somatic mutations. NAR Genomics and Bioinformatics. 2023;5:lqad056. McLaughlin RT, Asthana M, Di Meo M, Ceccarelli M, Jacob HJ, Masica DL. Fast, accurate, and racially unbiased pan-cancer tumor-only variant calling with tabular machine learning. npj Precis Onc. 2023;7:1–12. Jongbloed EM, Jansen MPHM, de Weerd V, Helmijr JA, Beaufort CM, Reinders MJT, et al. Machine learning-based somatic variant calling in cell-free DNA of metastatic breast cancer patients using large NGS panels. Sci Rep. 2023;13:10424. Jongbloed EM, Jansen MPHM, de Weerd V, Helmijr JA, Beaufort CM, Reinders MJT, et al. Machine learning-based somatic variant calling in cell-free DNA of metastatic breast cancer patients using large NGS panels. Sci Rep. 2023;13:10424. Spinella J-F, Mehanna P, Vidal R, Saillour V, Cassart P, Richer C, et al. SNooPer: a machine learning-based method for somatic variant identification from low-pass next-generation sequencing. BMC Genomics. 2016;17:912. Dietz S, Schirmer U, Mercé C, Bubnoff N von, Dahl E, Meister M, et al. Low Input Whole-Exome Sequencing to Determine the Representation of the Tumor Exome in Circulating DNA of Non-Small Cell Lung Cancer Patients. PLOS ONE. 2016;11:e0161012. Butler TM, Johnson-Camacho K, Peto M, Wang NJ, Macey TA, Korkola JE, et al. Exome Sequencing of Cell-Free DNA from Metastatic Cancer Patients Identifies Clinically Actionable Mutations Distinct from Primary Disease. PLOS ONE. 2015;10:e0136407. Li H, Durbin R. Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics. 2009;25:1754–60. Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25:2078–9. McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, et al. The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 2010;20:1297–303. Li H. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data. Bioinformatics. 2011;27:2987–93. Garrison E, Marth G. Haplotype-based variant detection from short-read sequencing. arXiv:12073907 [q-bio]. 2012. Wilm A, Aw PPK, Bertrand D, Yeo GHT, Ong SH, Wong CH, et al. LoFreq: a sequence-quality aware, ultra-sensitive variant caller for uncovering cell-population heterogeneity from high-throughput sequencing datasets. Nucleic Acids Res. 2012;40:11189–201. DePristo MA, Banks E, Poplin R, Garimella KV, Maguire JR, Hartl C, et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat Genet. 2011;43:491–8. Cingolani P, Platts A, Wang LL, Coon M, Nguyen T, Wang L, et al. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly (Austin). 2012;6:80–92. Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research. 2011;12:2825–30. Seo H, Park Y, Min BJ, Seo ME, Kim JH. Evaluation of exome variants using the Ion Proton Platform to sequence error-prone regions. PLoS One. 2017;12:e0181304. Spelmen VS, Porkodi R. A Review on Handling Imbalanced Data. In: 2018 International Conference on Current Trends towards Converging Technologies (ICCTCT). 2018. p. 1–11. Lemaître G, Nogueira F, Aridas CK. Imbalanced-learn: a python toolbox to tackle the curse of imbalanced datasets in machine learning. J Mach Learn Res. 2017;18:559–63. Holt JM, Kelly M, Sundlof B, Nakouzi G, Bick D, Lyon E. Reducing Sanger confirmation testing through false positive prediction algorithms. Genet Med. 2021;23:1255–62. Rehm HL, Bale SJ, Bayrak-Toydemir P, Berg JS, Brown KK, Deignan JL, et al. ACMG clinical laboratory standards for next-generation sequencing. Genetics in Medicine. 2013;15:733–47. Fang LT, Afshar PT, Chhibber A, Mohiyuddin M, Fan Y, Mu JC, et al. An ensemble approach to accurately detect somatic mutations using SomaticSeq. Genome Biology. 2015;16:197. Liu X, Lang J, Li S, Wang Y, Peng L, Wang W, et al. Fragment Enrichment of Circulating Tumor DNA With Low-Frequency Mutations. Front Genet. 2020;11. van ’t Erve I, Medina JE, Leal A, Papp E, Phallen J, Adleff V, et al. Metastatic Colorectal Cancer Treatment Response Evaluation by Ultra-Deep Sequencing of Cell-Free DNA and Matched White Blood Cells. Clinical Cancer Research. 2023;29:899–909. Underhill HR. Leveraging the Fragment Length of Circulating Tumour DNA to Improve Molecular Profiling of Solid Tumour Malignancies with Next-Generation Sequencing: A Pathway to Advanced Non-invasive Diagnostics in Precision Oncology? Mol Diagn Ther. 2021. https://doi.org/10.1007/s40291-021-00534-6. Stoler N, Nekrutenko A. Sequencing error profiles of Illumina sequencing instruments. NAR Genomics and Bioinformatics. 2021;3:lqab019. Cliften P. Chapter 7 - Base Calling, Read Mapping, and Coverage Analysis. In: Kulkarni S, Pfeifer J, editors. Clinical Genomics. Boston: Academic Press; 2015. p. 91–107. McNulty SN, Parikh BA, Duncavage EJ, Heusel JW, Pfeifer JD. Optimization of Population Frequency Cutoffs for Filtering Common Germline Polymorphisms from Tumor-Only Next-Generation Sequencing Data. The Journal of Molecular Diagnostics. 2019;21:903–12. Hiltemann S, Jenster G, Trapman J, van der Spek P, Stubbs A. Discriminating somatic and germline mutations in tumor DNA samples without matching normals. Genome Res. 2015;25:1382–90. Parikh K, Huether R, White K, Hoskinson D, Beaubier N, Dong H, et al. Tumor Mutational Burden From Tumor-Only Sequencing Compared With Germline Subtraction From Paired Tumor and Normal Specimens. JAMA Network Open. 2020;3:e200202. Tables Table 1: Datasets used to train and validate models. All samples are publicly available on the Sequence Read Archive (SRA) with corresponding accession numbers. cfDNA Accession Matched Tissue Reference Cancer Type Depth Usage SRR3401415 SRR3401405 Dietz et al., (2016) Squamous cell carcinoma Low Train/Test SRR3401416 SRR3401406 Dietz et al., (2016) Squamous cell carcinoma Low Train/Test SRR3401417 SRR3401411 Dietz et al., (2016) Lung adenocarcinoma Low Train/Test SRR3401418 SRR3401412 Dietz et al., (2016) lung adenocarcinoma Low Train/Test SRR3401407 SRR3401413 Dietz et al., (2016) Lung adenocarcinoma Low Train/Test SRR3401408 SRR3401414 Dietz et al., (2016) Squamous cell carcinoma Low Validation SRR12083523 SRR12083524 (Li et al., 2022) Glioma High Train/Test SRR12083526 SRR12083527 (Li et al., 2022) Glioma High Train/Test ERR855950 ERR855951 (Butler et al., 2015) Sarcoma High Train/Test SRR12083529 SRR12083530 (Li et al., 2022) Glioma High Validation Table 2: Description of features used for training ML models. Feature Description Read Depth Total number of reads at variant locus calculated as reads supporting reference + reads supporting alternate dbSNP Is variant described in dbSNP v151 common variants database COSMIC Is variant described in the COSMIC v98 database Strand Bias Strand bias estimated using Fisher's exact test, Phred-scaled p-value Median Alt Fragment Length Median fragment length of reads supporting alternate allele Weighted Homopolymer Rate (WHR) The sum of squares of the homopolymer lengths divided by the number of homopolymers in a 20-nucleotide region flanking the variant, including the variant. Homopolymers were defined as 4-mers and above. GC Percentage GC percentage in a 20-nucleotide region flanking the variant, including the variant Mapping Quality Median mapping quality of reads supporting alternate allele Allele Frequency Variant allele frequency calculated as (reads supporting alternate)/(reads supporting reference + reads supporting alternate) bcftools Is variant called by bcftools at default parameters FreeBayes Is variant called by FreeBayes at default parameters LoFreq Is variant called by LoFreq at default parameters Mutect2 Is variant called by Mutect2 at default parameters Distance from End of Read Median distance of variant starts from ends of reads supporting alternate allele Median alt base quality Median base quality of bases supporting alternate allele Table 3: Description of evaluation metrics used to assess model performance. Table 4: Rule-based filtering thresholds used to benchmark against RF models. Variants at or above thresholds were retained. Thresholds Read Depth Low Read Depth High QUAL dbSNP v151 COSMIC v98 Hard >20 >1000 >=50 0 1 Medium >10 >750 >=40 0 1 Soft >5 >500 >=20 0 1 Table 5: Total number and high confidence variants returned by from low and high depth datasets. cfDNA Accession Total Number of Variants High Confidence Variants Depth SRR3401415 298,225 376 Low SRR3401416 318,772 446 Low SRR3401417 261,362 286 Low SRR3401418 411,688 602 Low SRR3401407 358,119 838 Low SRR3401408 222,430 369 Low SRR12083523 257,374 1,446 High SRR12083526 309,852 1,592 High ERR855950 2,013,269 1,183 High SRR12083529 261,256 1,817 High Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 26 May, 2025 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Revision requested 24 Feb, 2025 Reviews received at journal 22 Feb, 2025 Reviews received at journal 17 Feb, 2025 Reviews received at journal 09 Feb, 2025 Reviewers agreed at journal 08 Feb, 2025 Reviewers agreed at journal 07 Feb, 2025 Reviewers agreed at journal 07 Feb, 2025 Reviewers invited by journal 05 Feb, 2025 Editor assigned by journal 05 Feb, 2025 Editor invited by journal 22 Jan, 2025 Submission checks completed at journal 22 Jan, 2025 First submitted to journal 30 Dec, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5735554","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":405638046,"identity":"bc177f24-5d75-419c-bc11-6da3a6c75ab2","order_by":0,"name":"Rugare Maruzani","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABB0lEQVRIiWNgGAWjYAgACTkJBuYGCJNYLcYSDIykaWFInEFIi/mM9IufKyoYovlnnzH+8HOHRfrM9oMNDD9qGBJnNmDXInMjp1jyzBmG3BnncgwMe89I5M7mSWxg7DnGkDgbl9slchIkG9sYchvO8Bgk8LZJ5M6TADqMt4EhcR5uLck/G/8x5M4Hajn4t00iXQ6ohfEvXi3pxyQbGxhyN5zhMWwG2pIgDdTCDLIFp8N43rBZNhyTyN14hq2YWbZNwnBmT2LDYZljEsa4vC/Bnv74ZkONTe68M8ybP75tq5OXOH744MM3NTayMw7gsIaBx4ABIw4O4I9I9gd4JEfBKBgFo2AUAAEAgj1WjUIN1zQAAAAASUVORK5CYII=","orcid":"","institution":"University of Liverpool","correspondingAuthor":true,"prefix":"","firstName":"Rugare","middleName":"","lastName":"Maruzani","suffix":""},{"id":405638047,"identity":"416886b4-d1ec-4e3a-9732-319bba16823b","order_by":1,"name":"Anna Fowler","email":"","orcid":"","institution":"University of Liverpool","correspondingAuthor":false,"prefix":"","firstName":"Anna","middleName":"","lastName":"Fowler","suffix":""},{"id":405638048,"identity":"442f1616-8ae3-4445-a0e8-9aa9f4965d2b","order_by":2,"name":"Liam Brierley","email":"","orcid":"","institution":"University of Glasgow","correspondingAuthor":false,"prefix":"","firstName":"Liam","middleName":"","lastName":"Brierley","suffix":""},{"id":405638049,"identity":"4adb19a5-abd2-4014-8d72-23174d7c1749","order_by":3,"name":"Andrea Jorgensen","email":"","orcid":"","institution":"University of Liverpool","correspondingAuthor":false,"prefix":"","firstName":"Andrea","middleName":"","lastName":"Jorgensen","suffix":""}],"badges":[],"createdAt":"2024-12-30 12:38:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5735554/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5735554/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-025-01326-2","type":"published","date":"2025-05-26T15:57:49+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":74691781,"identity":"f91b343f-6d83-49fd-9b01-7f9579ada8b2","added_by":"auto","created_at":"2025-01-24 18:45:39","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":83825,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eWorkflow for training and evaluating ML models.\u003c/strong\u003e In the outer loop, Training Data was randomly undersampled to each of six ratios. In the inner loop, hyperparameters were optimised for models fitted on undersampled data. The output best model was evaluated on Test Data. Test Data was not undersampled to reflect real data where ground truth is not known. \u0026nbsp;The best performing model on Test Data predictions was evaluated on previously unseen Validation Data. Double brakes indicate loop break points at the end of iterations.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-5735554/v1/9c066f1c724403261e1ef927.png"},{"id":74691783,"identity":"42fbfae2-bd0d-410f-875d-91a55e2cfe4f","added_by":"auto","created_at":"2025-01-24 18:45:39","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":56785,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eA-B: \u003c/strong\u003eSequencing depth histograms of low (A) high (B) depth validation samples.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-5735554/v1/29a6b0bfdc3d71d9a1757053.png"},{"id":74691782,"identity":"8c96d429-777c-4a78-9a87-d52455be9a83","added_by":"auto","created_at":"2025-01-24 18:45:39","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":55926,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eA-B: \u003c/strong\u003ePrecision Recall curves of low (A) and high (B) depth models Validation Data. X indicate rule-based filtering. \u003cstrong\u003eFigure 3 C-D: \u003c/strong\u003eConfusion matrices for low (C) and high (D) depth models on Validation Data with probability threshold of 0.75.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-5735554/v1/20d81d813022a365b1e8afc4.png"},{"id":74691787,"identity":"3960c239-a07e-4d12-be23-52e68ca2c463","added_by":"auto","created_at":"2025-01-24 18:45:39","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":49445,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eA)\u003c/strong\u003e Permutation feature importance rankings for low (A) and high (B) depth models calculated in Test Data. Error bars indicate the standard deviation across 30 permutations per feature.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-5735554/v1/d0e011c3aedb0cdd3408fa75.png"},{"id":74692370,"identity":"93177aa0-082b-45c2-b375-a0bc63a55a18","added_by":"auto","created_at":"2025-01-24 18:53:40","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":21616,"visible":true,"origin":"","legend":"\u003cp\u003ePDPs of the top 3 features in low depth (A) and high depth (B) models.\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-5735554/v1/dfdb74a1ea6976ef5612f14c.png"},{"id":83782955,"identity":"1215364d-f151-4113-b698-4c80a434cce0","added_by":"auto","created_at":"2025-06-02 16:09:11","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1187823,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5735554/v1/7f0f6e9f-578f-4903-ba33-ff92f2f6e4a5.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Predicting High Confidence ctDNA Somatic Variants with Ensemble Machine Learning Models","fulltext":[{"header":"Background","content":"\u003cp\u003eCirculating tumour DNA (ctDNA) is fragmented DNA released into the bloodstream by tumour cells. ctDNA has been shown to be a predictor of disease, relapse and prognosis for cancer patients [\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. ctDNA is minimally invasive to sample and has a short half-life, enabling near real-time monitoring for patients at risk of relapse [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Tumours can display intratumoral and spatial heterogeneity [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] however as all tumour tissues release ctDNA, the entire genetic profile of a patient\u0026rsquo;s tumours can be analysed. This is in contrast to tissue biopsy where only genetic information at the biopsy site can be analysed [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe utility of ctDNA as a cancer biomarker depends on the ability to accurately detect somatic variants associated with cancer. ctDNA can be in low abundance in the pool of total cell free DNA (cfDNA), particularly in the early stages of disease or when monitoring for recurrence. Additionally, the abundance of ctDNA fragments released from subclones will be even lower. Ultimately, it is difficult to distinguish between real low frequency ctDNA variants and NGS artifacts [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAccurate somatic variant detection in cfDNA NGS data requires variant filtering methods. Accuracy of called variants is influenced by features including base quality and read depth. Base quality scores represent confidence in the called base by sequencing platforms. Read depth describes the total number of reads covering the variant site. Higher read depths and a higher number of reads supporting an alternate allele increase confidence a called variant is real. Rule-based filtering involves defining a set of thresholds for measures including read depth and base quality, and then discarding variants which fall below those thresholds. It also often involves discarding variants reported in population databases including the 1000 Genomes Project [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] and dbSNP [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAnother approach to the variant filtering problem is consensus, or ensemble variant calling.\u003c/p\u003e \u003cp\u003eVariant callers utilise a wide range of statistical methods and rely on differing assumptions about the input data to call variants. The result is that overlapping but different call sets are returned by different tools from the same input data. Variants commonly called by more than one caller have been shown to be more likely to be real somatic variants [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. As such, ensemble variant calling aims to reduce the number of false positive calls by combining the output of multiple variant callers [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Naively, this may involve filtering out variants detected by a single caller [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eBoth these filtering strategies have limitations. Depending on the thresholds specified, rule-based filtering can discard a large fraction of true positive variants or retain an implausibly large number of variants post-filtering. Benchmarking studies have shown variant callers have low concordance in variants detected from the same input data [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Naturally, ensemble variant call sets tend to have high precision, with a trade-off in low recall [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Machine Learning (ML) algorithms can enable identification of complex, non-linear patterns which may improve ability to distinguish between real somatic ctDNA variants and false positive calls.\u003c/p\u003e \u003cp\u003eRecently, ML tools for predicting somatic variants in cancer samples have been published [\u003cspan additionalcitationids=\"CR15 CR16\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. However, few studies have focused on developing models for predicting somatic variants specifically in cfDNA NGS data. Jongbloed et al., (2023) [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] present an ML model for detecting somatic variants in cfDNA data from breast cancer patients. The model, however, was fitted on, and developed for, targeted sequencing data. Depending on the size of the panel, targeted sequencing can return orders of magnitude fewer variants compared to Whole Exome Sequencing (WES), reducing the resources required to validate calls, and consequently, the need for specialist filtering tools.\u003c/p\u003e \u003cp\u003eFitting supervised ML models requires a truth set to label positive examples. However, obtaining a truth set for predicting somatic variants in NGS data can be a challenge. This process may involve resource-intensive orthogonal sequencing or manual review of candidate variants. The use of multiple sequencing platforms and ultra-high sequencing depths often utilised for orthogonal validation [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] can prohibitively increase the costs associated with obtaining a truth set. Given the labour-intensive nature of manually reviewing variants [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], typically only a subset of candidate true positive variants are validated [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], potentially impacting the performance of the fitted model.\u003c/p\u003e \u003cp\u003eOne approach to obtaining a truth set is using common variants between matched tissue biopsy and the cfDNA sample. Tumour variants are much easier to detect in tissue than in ctDNA, and therefore the variants identified at low frequencies in ctDNA which are also confidently identified in the tissue sample are highly likely to be true positives. This method is less labour-intensive compared to manual curation and does not necessarily require independent sequencing technology. It therefore allows for a larger number of variants to be included in the truth set.\u003c/p\u003e \u003cp\u003eHere, we present two ML models for predicting high confidence somatic variants in low and high depth cfDNA WES datasets. The high confidence truth set was defined as variants detected in a matched tissue sample after stringent filtering. We extracted features from the output of 4 variant callers and predicted high confidence somatic variants with Random Forest models. Our models do not rely on matched normal samples; we utilise allele frequency and population databases to filter germline variants. The models were trained to predict single nucleotide variants (SNVs) only as they are the most common variant type in cancer. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e illustrates the workflow for model training and evaluation. Our code is available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/rugare-m/somVar\u003c/span\u003e\u003cspan address=\"https://github.com/rugare-m/somVar\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eDatasets\u003c/h2\u003e \u003cp\u003eWe used a set of publicly available cfDNA WES samples to fit low and high depth ML models. The low depth model was trained and evaluated using samples from Dietz et al (2016) [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] \u0026rsquo;s publication. This dataset included 6 cfDNA samples; each of the samples was sequenced with a matched tissue sample.The high depth model was trained and evaluated using samples from Li et al. (2022) [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] and one sample from Butler et al. (2015) [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], all sequenced with matched tissue samples. Table\u0026nbsp;1 shows the samples used to fit both models, including the accession numbers.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eData pre-processing.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo generate BAM files, FASTQ files were mapped to GRCh38 using BWA-MEM2 v2.2.1 [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] at default parameters. Output SAM files were converted to BAM file format using Samtools v1.2 [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Duplicate reads in BAM files were marked with GATK [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] MarkDuplicates v4.3.0.0, before recalibrating base quality scores with GATK BaseRecalibrator v4.3.0.0 and GATK ApplyBQSR v4.3.0.0. The sequencing depth profiles of all processed BAM files was visualised by plotting a histogram of depth at every position in the GRCh38 exome.\u003c/p\u003e \u003cp\u003e \u003cb\u003eVariant calling pipeline.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eFour variant callers were used to detect variants; bcftools v1.16 [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], FreeBayes v1.3.6 [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], LoFreq v2.1.5 [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], Mutect2 v4.3.0.0 [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. All callers were run at default parameters. Output VCF files were decomposed to split multiallelic sites into individual records, and indels were removed. We discarded variants with allele frequencies between 0.4 and 0.6, and above 0.9 as these were assumed to be likely germline variants [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. VCF files were annotated with strand bias, mapping quality, base quality, fragment length, read position, allele frequency and number of reads supporting alleles using GATK v4.3.0.0 VariantAnnotator. SnpSift v4.3t [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] was used to annotate coding variants reported in COSMIC v98 and common variants in dbSNP v151. For each sample, VCF files from all callers were merged into a single VCF file using GATK MergeVcfs.\u003c/p\u003e \u003cp\u003e \u003cb\u003eObtaining truth sets.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eWe used a set of high confidence variants as true positive calls. These variants were defined as merged tissue variants with read depth greater than 70 for low depth data (or greater than 500 for high depth data), with Phred quality scores\u0026thinsp;\u0026gt;\u0026thinsp;=\u0026thinsp;50, not reported in the common variants dbSNP v151 VCF file and reported in the COSMIC v98 coding mutations VCF. To minimise the number of germline variants in the truth set, variants with previously described allele frequencies were discarded.\u003c/p\u003e \u003cp\u003e \u003cb\u003eMachine learning features.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eA set of 15 features per variant in the merged cfDNA VCF files were extracted into tabular data from data in the BAM, reference FASTA and VCF files. BAM level features included read depth, strand bias, median fragment length of reads supporting alternate allele, mapping quality, allele frequency, and median distance of variant from end of reads supporting alternate allele. The read depth feature was standardised by removing mean and scaling to unit variance using scikit-learn v1.4.1 [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. Reference FASTA features included GC percentage of reference sequence in a 20-nucleotide region flanking the variant (41 bases total), and a weighted homopolymer score calculated from the same region [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. VCF features included a Boolean of presence of the variant in each of dbSNP and COSMIC databases, and bcftools, FreeBayes, LoFreq, and Mutect2 VCF files. Table\u0026nbsp;2 describes features used in the ML models. For the target, variants were labelled as high confidence somatic variants (HCSVs) if they were present in the truth VCF file and artefact variants (AVs) otherwise.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eTrain, Test and Validation Data.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eOne validation sample from each of the low and high depth datasets was randomly chosen. This was kept independent of the training and testing pipeline. Variants from the remaining samples were allocated to training and testing sets, with a 70:30 training: testing split. The models were optimised using the training and testing datasets and then used to predict high confidence somatic variants in the validation samples.\u003c/p\u003e \u003cp\u003eWe trained our models using the scikit-learn implementation of the Random Forest algorithm. The Random Forest algorithm fits decision trees on subsamples of input data and averages tree predictions to classify samples.\u003c/p\u003e \u003cp\u003e \u003cb\u003eClass balancing and hyperparameter optimisation loop.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eML models fitted on imbalanced datasets perform poorly when predicting the minority class. These models tend to classify all new data as majority class, which is often the class of least interest [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Here, the ratio of true positive somatic variants to total variants per low depth sample was in the order of 1:1000. To mitigate against this, we treated the undersampling ratio as a parameter to be optimised in a nested loop that undersampled Artifact Variants (AVs) in Training Data to one of 6 ratios and optimised hyperparameters for models fitted on the undersampled Training Data (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn the outer loop, Training Data was undersampled to each of 1:10, 1:5, 2:5, 3:5, 4:5 and 1:1 HCSV to AV ratios using the RandomUnderSampler (RUS) algorithm in the Python library imbalanced-learn v0.12.0 [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. RUS randomly downsamples AVs until the user-defined HCSVs to AV ratio is reached.\u003c/p\u003e \u003cp\u003eThe inner loop optimised hyperparameters for RF models fitted on the undersampled Training Data using RandomizedSearchCV in scikit-learn. The random search was set to evaluate 100 different hyperparameter combinations using 3-fold cross validation, and to score recall when outputting the best model. The best model from each of the 6 undersampling ratios was used to predict classes in Test Data, and precision and recall scores were calculated. Test Data was not undersampled to reflect real data where the ground truth is unknown.\u003c/p\u003e \u003cp\u003ePrioritising recall, the model with the highest precision and recall scores on Test Data predictions was used to evaluate performance on the previously unseen Validation Data. Class balancing ratios and hyperparameters were separately optimised for low and high depth models.\u003c/p\u003e \u003cp\u003e \u003cb\u003eHigh Confidence Somatic Variant Probability Thresholds.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe Random Forest algorithm predicts the probability of a variant being a HCSV. A default threshold of 0.5 is required to label variants as HCSV. Increasing the threshold increases the model precision score at the expense of recall. To reduce the false positive calls associated with high recall selected for in hyperparameter optimisation, we set the HCSV prediction threshold default to 0.75. Predictions in Validation Data were with the threshold value set to 0.75.\u003c/p\u003e \u003cp\u003e \u003cb\u003eEvaluating models on Validation Data.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eWe produced confusion matrices on Validation Data for the exported models to show the raw numbers of true positive (TP), true negative (TN), false positive (FP) and false negative (FN) predictions by the models. Due to the large imbalance in classes, we opted to use precision-recall (PR) curves over receiver operator characteristics (ROC) to evaluate the models\u0026rsquo; performance at different thresholds [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cb\u003eFeature importance analysis.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eWe assessed features\u0026rsquo; contribution to model performance using permutation feature importance in scikit-learn. This method shuffles a feature in Test Data \u003cem\u003en\u003c/em\u003e times and calculates the mean decrease of the model\u0026rsquo;s score on the permuted Test Data. Here, each feature was shuffled 30 times and we used Average Precision (AP) as the scoring metric - AP summarises the precision-recall curve (Table\u0026nbsp;3). Next, we calculated partial dependence of the top three features from permutation feature importance analysis in the Test Data. Partial dependence plots (PDPs) are calculated by varying a feature and observing resulting change in model-predicted marginal probability of HCSV, i.e. averaging out all other features.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eBenchmarking against rule-based filtering.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eWe aimed to benchmark the RF models against rule-based filtering strategies using the SRR3401408 and SRR12083529 samples which were held out for validation. We used the RF models to predict high confidence variants on corresponding Validation Data and plotted PR curves. Next, we performed rule-based filtering of the merged validation sample cfDNA VCF files at three different thresholds; hard, medium, and soft, and plotted precision and recall on PR curves. Rule-based filtering at all thresholds included the previously described allele frequency filters to discard germline variants. Table\u0026nbsp;4 illustrates the thresholds applied at each filtering threshold.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eSequencing depth profiles of matched tissue and Validation Data samples.\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe applied a read depth filter in the matched tissue samples to obtain a HCSV truth set and for benchmarking models against rule-based filtering. The thresholds for read depth are dependent on the sequencing depths of the data. To inform selection of the appropriate read depth thresholds, we plotted histograms of read depths at every position in the human exome for all samples.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe low depth validation sample showed a mode read depth of approximately 10X (Figure 2 A). 10X read depth was selected as the medium depth threshold, with 5X and 20X as the soft and hard thresholds respectively. In the high depth validation sample, the mode was approximately 750X, which was selected as the medium depth threshold, with 500X and 1000X selected as soft and hard depth filter thresholds respectively (Figure 2B).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMatched tissue samples showed read depth of approximately 50X and 200X in low and high depth data respectively. We defined high confidence variants as variants passing a stringent depth filter in matched tissue data. We therefore selected thresholds of 70X and 500X for low and high depth data truth sets respectively.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHyperparameter and class balancing optimisation.\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe fitted two Random Forest models for predicting high confidence somatic ctDNA variants. We optimised the hyperparameters of each model using a random search and 3-fold cross validation.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLow depth models fitted on data undersampled to ratios of 2:5, 3:5, and 4:5 all returned the greatest performance, with 1.00 and 0.07 recall and precision scores. To select the best model from this set of 3, we compared the PR areas under curve (PR-AUCs) on Test Data predictions. The PR-AUC metric ranges from 0 to 1, with a perfect classifier returning a PR-AUC 1 while a model with no predictive power returns a PR-AUC of 0. Here, the model fitted on Training Data undersampled to 2:5 returned the highest score of 0.43. The 3:5 and 4:5 models returned PR-AUCs of 0.31 and 0.28 respectively. The model fitted on Training Data undersampled to 2:5 was selected as the final low depth model. Undersampling the low depth Training Data to a 2:5 ratio reduced AVs from 1,151,933 to 4,457.\u003c/p\u003e\n\u003cp\u003eThe high depth model fitted on Training Data undersampled to 1:5 ratio returned the highest precision and recall scores; 0.15 and 1.00 respectively (Table 6).Undersampling high depth Training Data to 1:5 ratio reduced AVs variants from 1,803,351 to 14,975.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHigh Depth Random Forest model outperformed rule-based filtering.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe evaluated the performance of both models using PR-AUC, and confusion matrices on Validation Data predictions. Low and high depth model PR-AUCs applied to these unseen data were 0.45 and 0.71 respectively (Figure 3 A-B).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe high depth model outperformed soft, medium, and hard rule-based filtering thresholds at predicting HCSV in respective Validation Data (Figure 3B). At a 0.75 HCSV probability threshold, the low depth model returned a recall of 0.97 and precision of 0.11; and the high depth model returned a recall of 0.97 and precision of 0.26 in respective Validation Data predictions. Figures 3 C-D show the confusion matrices and PR-AUCs for low and high depth model predictions. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRead Depth, COSMIC and dbSNP membership are key features for ML models.\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe used permutation feature importance to assess the contribution of features in predicting HCSV (Figure 4). Presence in COSMIC and dbSNP, and Read Depth were the top 3 features contributing to both models’ performance. Random permutation of COSMIC, dbSNP and Read Depth features resulted in a mean decrease in AP of 0.15, 0.13 and 0.07 respectively in low depth data predictions. In high depth data, the mean decrease in AP after permuting COSMIC, Read Depth and dbSNP were 0.63, 0.53 and 0.45 respectively. Permuting Median Alternate Fragment Length notably had a relatively low impact on AP in both model predictions.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe PDPs showed variant presence in the COSMIC database results in the largest increase in mean probability of HCSV predictions for both models (Figure 5). Absence in dbSNP also increased mean probability of HCSV predictions in both models. Increase in read depth increased mean probability of HCSV predictions, however the increase is small and plateaus, particularly in the low depth model.\u0026nbsp;\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eVariant filtering is an important step in detecting somatic mutations in WES cfDNA NGS data. The aim is to filter out false positive calls without discarding too many real variants. Variant filtering in cfDNA data is complicated by the fact that ctDNA is typically in low abundance relative to healthy cfDNA. Ultimately true positive ctDNA variants occur at low allele frequencies that are difficult to distinguish from sequencing and PCR artifacts. We developed two ML models for predicting high confidence somatic ctDNA variants in low and high depth data. Our models were trained on WES samples with high confidence labels acquired from matched tissue samples. At a probability threshold of 0.75, both models correctly classified over 97% of HCSVs and outperformed all rule-based filtering thresholds.\u003c/p\u003e \u003cp\u003eHyperparameter optimisation for both models used recall as the scoring metric. While a balance of precision and recall is important, we determined recall is the more important metric in variant filtering of liquid biopsy NGS data. False positive variants can be removed in downstream orthogonal validation [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]; however, there is no way to rectify potentially actionable variants being discarded. Variants detected in a clinical context by any tool will likely require confirmation with Sanger sequencing for the foreseeable future [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. The goal is to reduce the number of variants requiring confirmation, in turn reducing the resources required for validating variants, while retaining as many true positive variants as possible. The balance between precision and recall, however, can often be dependent on application of the tool. We therefore implemented a HCSV probability threshold parameter in our models to allow users to set the trade-off between precision and recall.\u003c/p\u003e \u003cp\u003eWe generated partial dependence plots for the top-ranking features for each of our models. PDPs provide insight into how ML models make predictions. In both models, presence in COSMIC increased the mean probability of HCSV predictions more than any other feature. This was as expected given that the COSMIC database is a well-curated catalogue of cancer-associated variants, and the ground truth contains COSMIC variants only. Absence from dbSNP also increased probability of HCSV predictions in both models. This was expected as variants described in dbSNP are unlikely to be cancer associated somatic variants.\u003c/p\u003e \u003cp\u003eRead depth, mapping quality and base quality are routinely used features in ML models for predicting variants in NGS data [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. In addition to these features, we included median fragment length of reads supporting the alternate allele in our models for predicting high confidence ctDNA somatic variants. Other researchers have previously demonstrated ctDNA fragments tend to be shorter than healthy cell free DNA [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. True somatic ctDNA variants, therefore, are more likely to originate from shorter fragments compared to germline variants, and artefacts in normal cfDNA fragments. Permutation feature importance however demonstrated this feature had minimal effect on AP when permuted suggesting low predictive power in both models. The substantial overlap in the distribution of fragment lengths of ctDNA and healthy cell free DNA may explain this low predictive power [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eOne of the known sources of sequencing errors in Illumina platforms is base calling in homopolymer regions. After a run of the same base, Illumina platforms may erroneously substitute the first base after the homopolymer with the homopolymer base. Stoler and Nekrutenko, (2021) [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e] found up to 5.3% of sequencing errors can be attributed to this effect. Repetitive regions are also associated with low mapping quality and base qualities [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. We included the WHR feature in our models to filter out errors associated with repetitive sequences in the genome. Feature importance analysis, however, suggests WHR also has little predictive power.\u003c/p\u003e \u003cp\u003eWhile both models showed good performance, our model training and evaluation pipeline had some limitations. Filtering out germline variants in somatic variant calling pipelines can be a challenge without a matched normal sample [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. Here, we attempted to remove germline variants by discarding variants reported in dbSNP and variants with an allele frequency of between 0.4 and 0.6, and above 0.9. Using population frequency databases to filter germline variants is a well-established approach [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. However, it is possible we discard some real HCSVs with the allele frequency filter.\u003c/p\u003e \u003cp\u003eAnother limitation in our approach was assembly of the truth set. Commonly detected variants in cfDNA and matched tissue samples after stringent filtering were labelled true positive high confidence variants. The matched tissue biopsy for each sample only represents a fraction of the tumour tissues releasing cfDNA, therefore the ground truth sets were necessarily incomplete. A possible consequence of this is that the false positive rate in our models is inflated if the models are predicting real HCSVs that are absent from the tissue biopsy sample, and so the truth set.\u003c/p\u003e \u003cp\u003eSome approaches may be investigated to mitigate this limitation of our approach. The truth set may be derived from multiple tissue biopsies from the same patient. This approach would increase the probability of capturing all somatic variants compared to a single biopsy approach. The challenge of this approach is the increased sequencing costs and disadvantages associated with collecting biopsies from patients. Orthogonal validation of predicted variants may also be employed to more accurately assess model performance. However, the large number of high confidence variants returned from WES experiments may render orthogonal validation of all variants unfeasible as previously discussed.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate. \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eNot applicable\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSequence data that support the findings of this study have been deposited in the European Nucleotide Archive with the primary accession code PRJNA318450, SRR12083523, SRR12083524, SRR12083526, SRR12083527, ERR855950, ERR855951 SRR12083529, SRR12083530.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was funded by a Medical Research Council DiMeN Doctoral Training Partnership iCASE studentship, with industrial partners Nonacus. MRC Grant Reference Number: MR/R015902/1\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAsante D-B, Calapre L, Ziman M, Meniawy TM, Gray ES. Liquid biopsy in ovarian cancer using circulating tumor DNA and cells: Ready for prime time? Cancer Letters. 2020;468:59\u0026ndash;71.\u003c/li\u003e\n\u003cli\u003eBortolini Silveira A, Bidard F-C, Tanguy M-L, Girard E, Tr\u0026eacute;dan O, Dubot C, et al. Multimodal liquid biopsy for early monitoring and outcome prediction of chemotherapy in metastatic breast cancer. NPJ Breast Cancer. 2021;7:115.\u003c/li\u003e\n\u003cli\u003eWu X, Zhang Y, Hu T, He X, Zou Y, Deng Q, et al. A novel cell-free DNA methylation-based model improves the early detection of colorectal cancer. Molecular Oncology. 2021;15:2702\u0026ndash;14.\u003c/li\u003e\n\u003cli\u003eLi S, Zeng W, Ni X, Zhou Y, Stackpole ML, Noor ZS, et al. cfTrack: A Method of Exome-Wide Mutation Analysis of Cell-free DNA to Simultaneously Monitor the Full Spectrum of Cancer Treatment Outcomes Including MRD, Recurrence, and Evolution. Clin Cancer Res. 2022;28:1841\u0026ndash;53.\u003c/li\u003e\n\u003cli\u003eChicard M, Colmet-Daage L, Clement N, Danzon A, Bohec M, Bernard V, et al. Whole-Exome Sequencing of Cell-Free DNA Reveals Temporo-spatial Heterogeneity and Identifies Treatment-Resistant Clones in Neuroblastoma. Clinical Cancer Research. 2018;24:939\u0026ndash;49.\u003c/li\u003e\n\u003cli\u003eHerzog H, Dogan S, Aktas B, Nel I. Targeted Sequencing of Plasma-Derived vs. Urinary cfDNA from Patients with Triple-Negative Breast Cancer. Cancers (Basel). 2022;14:4101.\u003c/li\u003e\n\u003cli\u003eBos MK, Angus L, Nasserinejad K, Jager A, Jansen MPHM, Martens JWM, et al. Whole exome sequencing of cell-free DNA \u0026ndash; A systematic review and Bayesian individual patient data meta-analysis. Cancer Treatment Reviews. 2020;83:101951.\u003c/li\u003e\n\u003cli\u003eNix DA, Hellwig S, Conley C, Thomas A, Fuertes CL, Hamil CL, et al. The stochastic nature of errors in next-generation sequencing of circulating cell-free DNA. PLoS One. 2020;15:e0229063.\u003c/li\u003e\n\u003cli\u003eAuton A, Abecasis GR, Altshuler DM, Durbin RM, Abecasis GR, Bentley DR, et al. A global reference for human genetic variation. Nature. 2015;526:68\u0026ndash;74.\u003c/li\u003e\n\u003cli\u003eSayers EW, Barrett T, Benson DA, Bolton E, Bryant SH, Canese K, et al. Database resources of the National Center for Biotechnology Information. Nucleic Acids Res. 2011;39 Database issue:D38-51.\u003c/li\u003e\n\u003cli\u003eTrevarton AJ, Chang JT, Symmans WF. Simple combination of multiple somatic variant callers to increase accuracy. Sci Rep. 2023;13:8463.\u003c/li\u003e\n\u003cli\u003eWang M, Luo W, Jones K, Bian X, Williams R, Higson H, et al. SomaticCombiner: improving the performance of somatic variant calling based on evaluation tests and a consensus approach. Sci Rep. 2020;10:12898.\u003c/li\u003e\n\u003cli\u003eHwang S, Kim E, Lee I, Marcotte EM. Systematic comparison of variant calling pipelines using gold standard personal exome variants. Scientific Reports. 2015;5:17875.\u003c/li\u003e\n\u003cli\u003eAnzar I, Sverchkova A, Stratford R, Clancy T. NeoMutate: an ensemble machine learning framework for the prediction of somatic mutations in cancer. BMC Medical Genomics. 2019;12:63.\u003c/li\u003e\n\u003cli\u003eAinscough BJ, Barnell EK, Ronning P, Campbell KM, Wagner AH, Fehniger TA, et al. A deep learning approach to automate refinement of somatic variant calling from cancer sequencing data. Nat Genet. 2018;50:1735\u0026ndash;43.\u003c/li\u003e\n\u003cli\u003eD\u0026iacute;az-Navarro A, Bousquets-Mu\u0026ntilde;oz P, Nadeu F, L\u0026oacute;pez-Tamargo S, Be\u0026agrave; S, Campo E, et al. RFcaller: a machine learning approach combined with read-level features to detect somatic mutations. NAR Genomics and Bioinformatics. 2023;5:lqad056.\u003c/li\u003e\n\u003cli\u003eMcLaughlin RT, Asthana M, Di Meo M, Ceccarelli M, Jacob HJ, Masica DL. Fast, accurate, and racially unbiased pan-cancer tumor-only variant calling with tabular machine learning. npj Precis Onc. 2023;7:1\u0026ndash;12.\u003c/li\u003e\n\u003cli\u003eJongbloed EM, Jansen MPHM, de Weerd V, Helmijr JA, Beaufort CM, Reinders MJT, et al. Machine learning-based somatic variant calling in cell-free DNA of metastatic breast cancer patients using large NGS panels. Sci Rep. 2023;13:10424.\u003c/li\u003e\n\u003cli\u003eJongbloed EM, Jansen MPHM, de Weerd V, Helmijr JA, Beaufort CM, Reinders MJT, et al. Machine learning-based somatic variant calling in cell-free DNA of metastatic breast cancer patients using large NGS panels. Sci Rep. 2023;13:10424.\u003c/li\u003e\n\u003cli\u003eSpinella J-F, Mehanna P, Vidal R, Saillour V, Cassart P, Richer C, et al. SNooPer: a machine learning-based method for somatic variant identification from low-pass next-generation sequencing. BMC Genomics. 2016;17:912.\u003c/li\u003e\n\u003cli\u003eDietz S, Schirmer U, Merc\u0026eacute; C, Bubnoff N von, Dahl E, Meister M, et al. Low Input Whole-Exome Sequencing to Determine the Representation of the Tumor Exome in Circulating DNA of Non-Small Cell Lung Cancer Patients. PLOS ONE. 2016;11:e0161012.\u003c/li\u003e\n\u003cli\u003eButler TM, Johnson-Camacho K, Peto M, Wang NJ, Macey TA, Korkola JE, et al. Exome Sequencing of Cell-Free DNA from Metastatic Cancer Patients Identifies Clinically Actionable Mutations Distinct from Primary Disease. PLOS ONE. 2015;10:e0136407.\u003c/li\u003e\n\u003cli\u003eLi H, Durbin R. Fast and accurate short read alignment with Burrows\u0026ndash;Wheeler transform. Bioinformatics. 2009;25:1754\u0026ndash;60.\u003c/li\u003e\n\u003cli\u003eLi H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25:2078\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eMcKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, et al. The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 2010;20:1297\u0026ndash;303.\u003c/li\u003e\n\u003cli\u003eLi H. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data. Bioinformatics. 2011;27:2987\u0026ndash;93.\u003c/li\u003e\n\u003cli\u003eGarrison E, Marth G. Haplotype-based variant detection from short-read sequencing. arXiv:12073907 [q-bio]. 2012.\u003c/li\u003e\n\u003cli\u003eWilm A, Aw PPK, Bertrand D, Yeo GHT, Ong SH, Wong CH, et al. LoFreq: a sequence-quality aware, ultra-sensitive variant caller for uncovering cell-population heterogeneity from high-throughput sequencing datasets. Nucleic Acids Res. 2012;40:11189\u0026ndash;201.\u003c/li\u003e\n\u003cli\u003eDePristo MA, Banks E, Poplin R, Garimella KV, Maguire JR, Hartl C, et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat Genet. 2011;43:491\u0026ndash;8.\u003c/li\u003e\n\u003cli\u003eCingolani P, Platts A, Wang LL, Coon M, Nguyen T, Wang L, et al. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly (Austin). 2012;6:80\u0026ndash;92.\u003c/li\u003e\n\u003cli\u003ePedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research. 2011;12:2825\u0026ndash;30.\u003c/li\u003e\n\u003cli\u003eSeo H, Park Y, Min BJ, Seo ME, Kim JH. Evaluation of exome variants using the Ion Proton Platform to sequence error-prone regions. PLoS One. 2017;12:e0181304.\u003c/li\u003e\n\u003cli\u003eSpelmen VS, Porkodi R. A Review on Handling Imbalanced Data. In: 2018 International Conference on Current Trends towards Converging Technologies (ICCTCT). 2018. p. 1\u0026ndash;11.\u003c/li\u003e\n\u003cli\u003eLema\u0026icirc;tre G, Nogueira F, Aridas CK. Imbalanced-learn: a python toolbox to tackle the curse of imbalanced datasets in machine learning. J Mach Learn Res. 2017;18:559\u0026ndash;63.\u003c/li\u003e\n\u003cli\u003eHolt JM, Kelly M, Sundlof B, Nakouzi G, Bick D, Lyon E. Reducing Sanger confirmation testing through false positive prediction algorithms. Genet Med. 2021;23:1255\u0026ndash;62.\u003c/li\u003e\n\u003cli\u003eRehm HL, Bale SJ, Bayrak-Toydemir P, Berg JS, Brown KK, Deignan JL, et al. ACMG clinical laboratory standards for next-generation sequencing. Genetics in Medicine. 2013;15:733\u0026ndash;47.\u003c/li\u003e\n\u003cli\u003eFang LT, Afshar PT, Chhibber A, Mohiyuddin M, Fan Y, Mu JC, et al. An ensemble approach to accurately detect somatic mutations using SomaticSeq. Genome Biology. 2015;16:197.\u003c/li\u003e\n\u003cli\u003eLiu X, Lang J, Li S, Wang Y, Peng L, Wang W, et al. Fragment Enrichment of Circulating Tumor DNA With Low-Frequency Mutations. Front Genet. 2020;11.\u003c/li\u003e\n\u003cli\u003evan \u0026rsquo;t Erve I, Medina JE, Leal A, Papp E, Phallen J, Adleff V, et al. Metastatic Colorectal Cancer Treatment Response Evaluation by Ultra-Deep Sequencing of Cell-Free DNA and Matched White Blood Cells. Clinical Cancer Research. 2023;29:899\u0026ndash;909.\u003c/li\u003e\n\u003cli\u003eUnderhill HR. Leveraging the Fragment Length of Circulating Tumour DNA to Improve Molecular Profiling of Solid Tumour Malignancies with Next-Generation Sequencing: A Pathway to Advanced Non-invasive Diagnostics in Precision Oncology? Mol Diagn Ther. 2021. https://doi.org/10.1007/s40291-021-00534-6.\u003c/li\u003e\n\u003cli\u003eStoler N, Nekrutenko A. Sequencing error profiles of Illumina sequencing instruments. NAR Genomics and Bioinformatics. 2021;3:lqab019.\u003c/li\u003e\n\u003cli\u003eCliften P. Chapter 7 - Base Calling, Read Mapping, and Coverage Analysis. In: Kulkarni S, Pfeifer J, editors. Clinical Genomics. Boston: Academic Press; 2015. p. 91\u0026ndash;107.\u003c/li\u003e\n\u003cli\u003eMcNulty SN, Parikh BA, Duncavage EJ, Heusel JW, Pfeifer JD. Optimization of Population Frequency Cutoffs for Filtering Common Germline Polymorphisms from Tumor-Only Next-Generation Sequencing Data. The Journal of Molecular Diagnostics. 2019;21:903\u0026ndash;12.\u003c/li\u003e\n\u003cli\u003eHiltemann S, Jenster G, Trapman J, van der Spek P, Stubbs A. Discriminating somatic and germline mutations in tumor DNA samples without matching normals. Genome Res. 2015;25:1382\u0026ndash;90.\u003c/li\u003e\n\u003cli\u003eParikh K, Huether R, White K, Hoskinson D, Beaubier N, Dong H, et al. Tumor Mutational Burden From Tumor-Only Sequencing Compared With Germline Subtraction From Paired Tumor and Normal Specimens. JAMA Network Open. 2020;3:e200202.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003e\u003cstrong\u003eTable 1:\u003c/strong\u003e Datasets used to train and validate models. All samples are publicly available on the Sequence Read Archive (SRA) with corresponding accession numbers.\u0026nbsp;\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"720\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003e\u003cstrong\u003ecfDNA Accession\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMatched Tissue\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eReference\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCancer Type\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDepth\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eUsage\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eSRR3401415\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eSRR3401405\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eDietz et al., (2016)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003eSquamous cell carcinoma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eTrain/Test\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eSRR3401416\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eSRR3401406\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eDietz et al., (2016)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003eSquamous cell carcinoma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eTrain/Test\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eSRR3401417\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eSRR3401411\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eDietz et al., (2016)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003eLung adenocarcinoma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eTrain/Test\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eSRR3401418\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eSRR3401412\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eDietz et al., (2016)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003elung adenocarcinoma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eTrain/Test\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eSRR3401407\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eSRR3401413\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eDietz et al., (2016)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003eLung adenocarcinoma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eTrain/Test\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eSRR3401408\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eSRR3401414\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eDietz et al., (2016)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003eSquamous cell carcinoma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eValidation\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eSRR12083523\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eSRR12083524\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003e(Li et al., 2022)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003eGlioma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eTrain/Test\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eSRR12083526\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eSRR12083527\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003e(Li et al., 2022)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003eGlioma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eTrain/Test\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eERR855950\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eERR855951\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003e(Butler et al., 2015)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003eSarcoma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eTrain/Test\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.6144%;\"\u003e\n \u003cp\u003eSRR12083529\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eSRR12083530\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003e(Li et al., 2022)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.8627%;\"\u003e\n \u003cp\u003eGlioma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.0083%;\"\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.5049%;\"\u003e\n \u003cp\u003eValidation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2:\u0026nbsp;\u003c/strong\u003eDescription of features used for training ML models. \u0026nbsp; \u0026nbsp;\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"709\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFeature\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDescription\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eRead Depth\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eTotal number of reads at variant locus calculated as reads supporting reference + reads supporting alternate\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003edbSNP\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eIs variant described in dbSNP v151 common variants database\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eCOSMIC\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eIs variant described in the COSMIC v98 database\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eStrand Bias\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eStrand bias estimated using Fisher\u0026apos;s exact test, Phred-scaled p-value\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eMedian Alt Fragment Length\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eMedian fragment length of reads supporting alternate allele\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eWeighted Homopolymer Rate (WHR)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eThe sum of squares of the homopolymer lengths divided by the number of homopolymers in a 20-nucleotide region flanking the variant, including the variant. Homopolymers were defined as 4-mers and above.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eGC Percentage\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eGC percentage in a 20-nucleotide region flanking the variant, including the variant\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eMapping Quality\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eMedian mapping quality of reads supporting alternate allele\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eAllele Frequency\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eVariant allele frequency calculated as (reads supporting alternate)/(reads supporting reference + reads supporting alternate)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003ebcftools\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eIs variant called by bcftools at default parameters\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eFreeBayes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eIs variant called by FreeBayes at default parameters\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eLoFreq\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eIs variant called by LoFreq at default parameters\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eMutect2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eIs variant called by Mutect2 at default parameters\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eDistance from End of Read\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eMedian distance of variant starts from ends of reads supporting alternate allele\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 17.9126%;\"\u003e\n \u003cp\u003eMedian alt base quality\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 82.0874%;\"\u003e\n \u003cp\u003eMedian base quality of bases supporting alternate allele\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3:\u0026nbsp;\u003c/strong\u003eDescription of evaluation metrics used to assess model performance.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cimg src=\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAlgAAADoCAYAAAAzM5qAAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAADsMAAA7DAcdvqGQAAEivSURBVHhe7Z1rsBXFub+H/2e8bPWLRFMKRo+AJcJWTBBSYCFG/GBKECOWNygVQ6HGOxglKh5R462iYsrb2RUviBbUCShIpayAGEBQPAE88XgrQ8EXFBS+899P73mXvXvPzJq11qy111r791T1WtP37pme6bff7pkedKibSAghhBBCFMb/i/+FEEIIIURBSMASQgghhCgYCVhCCCGEEAUjAUsIIYQQomAkYAkhhBBCFIwELCGEEEKIgpGAJYQQQghRMBKwhBBCCCEKRgKWEEIIIUTBSMASQgghhCgYCVh1ZNCgQammP/nyyy+jo446KjrppJNiFyFan6T77Mwzz4yee+65OER+vv322+jSSy8tpQPcL9w33D95yHOf6V4Uon2RgNUAOjs7oylTpvQyjWTBggXR+eefH73++uuxixDti91vHR0d0ZYtW6Lrr7/eCVqV8Nhjj0VLly51aS1evDh2rR3di0IMINjsWdQHTi9m/fr1sUv/0N3ZuHJ0dxSxixDtR9L9tmrVqkPdglbF7b9e94zuRSEGDtJg9TOffPKJG10zDcE0AdMZHDPKhffff7/XNAU8/PDDvcKQhj+dQTo2Qsa+Zs0ad3zHHXeU0rGwBlMiN9xwg5uuwJ0yvf3227Hvj+FJl/Q5Jn/iCdGsXHDBBdGdd97pjh966CH3D9berV377R17nnvG7Fn3hB+Hf92LQgwgYkFL1AFOL6azs9ONXM3Mnz8/DnGoNLrGfc6cOaU42IHRuLkZjH79MMOGDXN20sXP0vziiy+c3fwJbyNn7BiDMmLn3y+HaQPMTtqkY3ksWbLE+QvR34Rt1ti2bVvJj3sCrL2H9wxx894zZs+6J/w4uheFGFj8eFeLwrEHYWh4KALTF9h56Bo8dP0weQQs/u0YOMbfHshmtwc6+GlaHn45eFjjNmPGDGe38HRWYP5+vkL0J9ZGrd37+H7W3hFgzI6ghRsCDZS7Z8DsWfdEGEf3ohADB00RNoDuhyZPz5JZvXq1c9++fbv7R81v/OIXv4iP8vPMM89EY8aMcdMEmM2bN8c++fjggw/cv1+OESNGuH8W+vqcfvrp7t/8hWh2mEI3hgwZUmrvLIAfP368Mw8++KBzy/uGoE+R94TuRSHaBwlYLQ4dQvdoN3r22WedkHX33Xc7uxCiB1v31NHREQ0dOtQdw4wZM3oNfPzBjxBC1IoErH5k5MiR7v/zzz93/2AjWOOwww6Lj37ko48+io+i6M0333T/c+bMiRYtWhSdc845zu5z5JFHxkfJmNbML8euXbvc/5QGf1JCiCLhJRFb3G6L3a29o8Ey7OUSFpfXE92LQgwcJGA1gJtvvrk0fWcGxo4d60bVX3zxhXsLEMPbRT5MAxAGeJsI8+677zo7/PSnP3X/aLB4u5C0/Y4DJk6c6P7paCxvH4Syzs5OVw78+VaPdTTz5s1z/0K0Cna/8RYe03/79u1z7fv22293/mF757656667nN+FF17o/uuF7kUhBhCHRN3g9KYZg4WqLGjFjUW3SQtWX3vttdKbQix0tcXxFsYW6BKGY8JgZ8Es7N27t+SGAf8YCMMCX8uHspCPEYa3xbhaWCuaBWujvrF7KiRs79yD3GcG7Rr3PIvcjaR7Igyje1GIgcMgfrpvTtEkMKXBqLv7Yan1IEIIIUSLoilCIYQQQoiCkYAlhBBCCFEwmiIUQgghhCiYXALWwoUL4yMhhBBCCBESykrSYNUAJ/OWW26JbUIIIYQYqITfrdQaLCGEEEKIgpGAJYQQQghRMBKwhBBCtB3/+Mc/4qO+ZPkJURQ1C1hs58AeXr5hy5dqdqWvBPYO83fJTyNvOCFE63L44YcnmieeeKLk3wr0Vznfeustt+1WUv74ff3117GtNfjnP//pPtbMP/h1CP0awe9+9ztXBmuXofHbqW/YKs22RkMo/PWvf+2Om4GwXVBe0ZtCNFhsNLx+/fqSYR+t8847L/atD+wdtnHjxtiWTt5wQojW5YcffnBmzZo1vew33XSTs4ts/vKXv0SPP/64O2ch+O3Zsye2tQannXaaqwv/4Nch9Ks3CCF/+9vfoosvvrjULrPaKX7mfuONN0bTpk1zafz85z93/s2ifWvFdtFoChGwTjjhBLdJqZmlS5c6Ievtt9+OQxTPd999F1133XWxLZ284YQQ7Y2vpbnvvvti1x6NBpoC0xikaTbwT0qDDi8cvWO3jjCMh7YCrYTliXbDp5py+nmkdcBhfNOMYEcAuPrqq92xj/mh8TE//i2v559/PjGOlSHvuQU/XY6z6u5veE844uCOhsfyIKz9J9UBwrQsH//Y8vTLnpZnEl1dXU64qoZrrrnG/Zsgc/nll0fLli1zxyFhea1e/jHkrSNuaW0Kv/CcQjVtt52pyxqsoUOHuv/t27e7f6YNEbaYTrQd5L/99ls3lehPK+JmcMwu8mnTjrixbx+Q9kknneTc2EGfHegNP1xSmv70IW6kZeUirddffz32FUK0Mhs2bIi++eYb9//oo486rQADsKlTp0ZLlixxGgP+saeRlEYe/Hj33HOP22f0vffec/HffPPNXh2OhcWPDguTp5zEW7duXUnT4WPxyZv4aB9MM4L93HPPLWlOfNL8LK8RI0bELn2p9NyCX3eEN7/uSWWn4yfMtm3b3LP8l7/8ZfTYY4/FqfWQVgdA8PH3fF2+fLlzyyp7njx9EETOPvvs2FYZL774ovs/9thj3f+YMWNcewnJOkfV1NFIa1Np59SuH//V3mPtRF0ELBOERo4c6f5h3rx50VVXXRUtXrzY2WfOnOkEmL179zrDMW4Gxx9++KHThOEPM2bMcP8+CE1crKeeeort5aMtW7Y4DRprr0LCNMmThuYLdpTzoosucmkhaCGQCSFaH+sEbWoIrQDPixNPPLHUgfDf0dGRqgVKSiMPYbzp06e75w+GTvPgwYPOHSysPZ927NiRq5zEYzYhCeITnvSAcnC8YsUKZ6+UrLyMSs8t+HWfPXt2qe5pZUfw+Oqrr5yWiM78iiuu6KMRzILwvsCC4IRGKqvsleb58ccfRyeffHJsK49phTBPPvmkK5+da/7379/fRwOUdY6qqaOR5zr72PXz749q2kG7UIiARSNDS2QGQWjYsGHR2LFj4xCRE4AQWFAPIoAh+d5///3R0Ucf7cytt97q3PAzf9YEoA3Df+7cue5C+Vos+P777+OjHgj/wAMP9BlZWZqMiixN8kfYWrt2bRzqx3LCZZddFu3bt88dCyHaj507d7rOkg7HTDPd83SydNC1lpP4NrNg8CyuJ7WW2a97Wtnp/Hmuo0maMGFCNHHiROeeF+LT+TOFZkILwkFW2avJsxIhhbTR9GDII2k9sy+QQ7lzVGkdi6RR+TQjhQhYzz77bDR+/PiSQbjiYiLEGNwsxu7du93/McccU5quIw7gZ/6s5zI4RqsUNiLs8+fPd1ospgnROB133HG94oKl6T9UKF9nZ6dTaRp+OQ2bYhRCtBfDhw93GiSmTczwPLDRdn9DJ3vGGWfUXE7iM1XjQ+ddT2ots1/3tLIzuB88eHD00ksvubRZR8Y6pUq48sor3RQa5TMNUFbZi8izaMpd30rrWCSNyqcZKUTAYtoP4ccM65ZCQSgJP46ZUDDKw6JFi5wmCs0VjQwhL+/aKV8IFEIMLBhgbd26tTSy5/nBIt1wCiYL6yhsysPWzVSLvbJPWdC4o42vtZzER4tgi52Jx1QRyyHy4Gv5fbLqXk2ZbaotrHta2Ukf4YbwQGeeRlodJk2a5KbQbOoMsspeSZ4GQlmRnHLKKfFRD+Wub6V1zEvaOfUpIp9WpS5rsMoxZMgQ9+9L2KyDYgE8bkn+TPGh6fLdADvrrRDomNpDsEKj9fLLL8chekjLE3UsF1sIMfBgrc+qVavcEgS016NGjXKdvK0hyQtvTLF2hmfJEUccEbtWx/HHH+/SYWqHdUhoHGotp8W/7bbbXHwEBH9tTxZ8hifpbUEjre7VlHncuHGpdU8qO/4Y0maxOYNtFniHZNXBptDAypZV9rx5GiwG/+yzz2JbbSDIUlbK51Pu+lZaxzyUaxdGrfm0MjVv9oxQhHR8++23xy59QTDi+1i+dop4CDgsSEc4YmoP4chGBb4/N+3vf/97J51//vnnzt/SZHNFLpilTxwWs5PmM8880yvvtDQ3bdrkNFlhOZkaRBsWlt3QZs9CCFEMdL5MC7YbaPV4o94WgNcCae3atcu9LSiaj6bZ7PmVV16JzjzzTLf2CsEGDRWvLRu+P2u1ELxQb4awpgrpmDcUSYewRx55pFvAHpKWpqYJhRBC1AOm6fhUgykPaoG3CnkrULQGNWuwBjLSYAkhhCgHmidevqplhxOmB5lt0e4EzUuowZKAVQMSsIQQQggBTTNFKIQQQgjRrkjAEkIIIYQoGAlYou1gMSlvJKUZg1esfXdemLDvyAghhBC1IAFLtB3/+te/3Ddf+Lgf34LhrVJe/+abZ/YtGLBNwQmHP9+24ds7QgghRK1IwBJtx6effuq+E8MH7jZu3Og+9Ad8dXr06NHuGA4cOOC24rCP9k2ePNltpCqEqD9Zm/1m+QnRKkjAEm3HNddc47RRwGaxfCzWYP8wY8OGDSXhC9ivEm2XEGn4U8yNhO/18YXzpPzxK3orlnrDNil8/d22S/HrEPr1Fwh59bje4fXqrzZVKc3a9tP8fOp1LcshAUu0NXzg79RTT41tvWF/LPY6A27Am2++2W3hIESzwVYsjz/+eOKXzvHbs2dPbGsN2CaFuth2KX4dQr92oxWvV39Sru2n+TUDErBE22LTDEm7trMQHu0WO+Ezsrn++uudcKWP+LUuXEd/RMseeZA0esVu7SOMx2bLvOzASw/YQ6E7KQ9A42Jx+Pc1MH4elm9IGN9euMDOQMHaqo/5ofExP/4tr6S94rBbGbLKHOKny3FW3f2XRQhHHNzZaNjyIKz9J9UBwrQsH//Y8vTLnpZnEsRPq5dRrg3lzY84YV2hmjblg39SGuXKHcZrxbbv+5EP8c34bccn7Xpl1aMaJGCJtoW1WP4UoM+WLVvcPyMfDJuAS7hqfZj2/eabb9z/o48+mnvqzI/H+r3Vq1e7rbuIz4sS/oPWwuLHAx2DwD516tRoyZIlrj3xj92HeOvWrUsV+AlP3sRnZD5t2jSXB3baMS9pcOyT5md5mYY2iTxlDvHrjvDm1z2p7HSohNm2bZu7x3hzN9yTL60OwFQ/18JYvny5c8sqe548Q5LqlZdK8kurq+XPv7XbSq9PUhp58OO1ettnFoL4XIdHHnnEbX4dkna9Kj3feZCAJdoWNunm5kmCxe+2Tku0D9ax2fRS3qmYMN706dPdyw+YMWPGRAcPHnTuYGHxow3t2LHDCey8oWodCP8dHR29RuzEO+GEE2Jbb4hPeGuTlIPjFStWOHulZOVl5ClziF933ri1uqeV/dhjj42++uqrqKury3WY7KNXyTQ84enkDTpGNA5ZZa8mz6R65aXWOoLl77fbSq9PUhp5COO1cttfunRpqR58VZ3rEpJ2vSo933mQgCXaEkYjqI+TwI8HNTeTELXAVAJTzTt37nQPbTp/M/v27YtDlYf4Q4cOjW09MEVRT2ots1/3tLLTqaJhQFMwYcKEaOLEic49L8TnPmWqxzQpdKBZZa81T6tXXmrNL41ar0+9aca2v3btWheXMixbtix27U3a9arH+ZaAJdqSWbNmuU8uoHZmlOLDyIUbCTV6LaMTIZhK4FMfw4cPd6N9prDMMJVio+FyEB/B34cOoJ7UWma/7mll594bPHiwe3uXtFkvc/nllzu/vFx55ZVu2orymZYjq+y15mn1yksRdUyi1utTb5qt7duUJdOblAFNXBJp16se51sClmhLuDl4AGAYsfiw1sr8muVhJeqHXWMTpl988UX3Xy0sBAY6BdbrsM6ps7PTvZVqWhb8WESbd5Es8RH6bVEu8egsLrroImcvByP3JLLqXk2ZbeorrHta2Umfzss6UDqxNNLqMGnSJDdNSJpoFSCr7JXkaSTVi6kkc8s6j9Xkl1ZXn1rbFAyEtu/DZ3aYvgRfg+Vfy7TrVcT5DpGAJYRoe3hriDe3eGAeccQRsWt1HH/88S4dBHfW66BV4aG+atWqaO7cuW7qZNSoUa7TtvUg5bD4LMolPh0AQkU4OEhizpw5rrMjXhJpda+mzOPGjUute1LZ8ceQNushFy1a5BYhh2TVgXRsOt/KllX2vHn6JNWL41tvvdX9I6CkncdK8yt3vYxa25TRzm3fh6k+0iINpgnPP//82KenDdm1HDJkSOL1Kup8+ww61E18LCpk4cKF0S233BLbhBCifaHTQevbbrRrvUTjYWG9jzRYQgghhBAFIwFLCCFEWdpVyyPtlagXErCEEEIIIQpGa7BqQGuwiof1EO2GRshCCNH+hGuwJGDVgAQsIYQQQoAWuQshhBBC1BkJWEIIkYB9nDGJLD8hhAAJWEIIEcDXm/k4o33FmS9L25ZLoV8j4IOHlAFYpxga+8J56M4HF+0L2QiFFk4IUX8kYImmgu0Jwk7CNwZf4PXd/Y6kaMK8zPhajEaWR9Qfvt7Mywn2FWe+9Lxnzx53HPrVGwQ7Ni7n69MGm9VSBjNsDWX4fjfeeGM0bdo0l0a4bYoQor5IwBJNxb/+9S+3LQYdAtslsLcUHQWdhm2XAQsWLHD/hMOfzoetG/JQ6Uje8rJOC0NZrMOCWsojigHBFi0PW3lwzBYhBtomhN4k4ZdwxMGddmGaKcLaPwIOWivfDcK0LB//2PL0NV5peSbR1dXVS7iqhGuuucb9m3DINiT+Hm1CiPohAUs0FZ9++ml0zz33uH2hNm7cGJ177rnOHWFm9OjR7hgOHDjgdnInHEyePDnav3+/Oy4a8rJyGOFu740sj0hnw4YNbgd8BF32e0PgQis6depU164QftFGmVYHYZsw27Ztc9cUTeRjjz0Wp9YDcbj+phnyQfBZvXp1bOvZZBw3y3PJkiUuDv/YIU+ePgh3Z599dmyrDNvc1za7HTNmjBu4CCHqjwQs0VQw4rbR+scffxyNHz/eHcNLL70UH/V0pL7Qs3v3bqftqgfkRSeYRSPLI9IxQQVBFw3ijh07oi1btkQdHR2ldsXUHscrVqxwggc7+aMlQuC64oor3HqnvBDeF1gQnNBIkaev5eSfMiBcVZon98HJJ58c23owbZoZf9rP93vyySd7bZzLP4J/lsZMCFEMErBE08LI/dRTT41tvdm6dWs0YsQId0zncvPNN1fUMVYCeaH9sE6LTjSkkeUR+eBaIZzs3LkzGjp0aOzag03jIXCgmUKTNGHCBLcjfyUQH0GKaUITWhDgyBMhCmHLzL59+5x/NXmagGT466ww/nS170ce5513XuzzIwcPHoyPhBD1QgKWaEpsRO53HAbTL3ScV199tetEr7/+eifM3HTTTXGIvphwhGGEj/Dmuz3xxBNxyN5YXv/zP//jOiy0VGGHWE15RP3hejFtO3z4cHeNfBA8AA3S4MGDnXaUqUWuIeuUKuHKK69004Q2PQjkyXQcbmZIn/ZcRJ5CiOZHApZoSliLFa57Mph+AX+UXk6YsbAYRvik7bulxbepHtMg0FHaOiujmvKI+mBaQwQq1mChVezs7HTaJFuMjqYJLeRFF13kNI8INyaAIRilsXbt2vioN5MmTXLTcDY9CORJ2qbVIn0WtWOvJE8DoaxITjnllPhICFEvJGCJpmT9+vWp655Y/G6agnrjL7RPo5HlEdmMGzfOCTIIxKzB4rogEK9atSq67bbbnIYR4cbWJeGPGTVqlGtvixYtcovgQ+bMmeMENuKHkI694WqfbrA8586d6+KQPsKfrf/Kk6dB+/vss89iW22gGaas4SBBCFE82ouwBrQXYX1gZG8dUqgJwo8pOjoo1kVVA53Mww8/7LRRWeTJq4jyiGJAkEGD2G7wJuD27dsz3zTMC2nt2rVLbVWIOqC9CEXTM2vWLPemE51AODXS1dXlpnseffTR0jqtepEnr0aWRwxMmMpkzaBNKdYCbxXy1qIQov5Ig1UD0mAJIRoBmqfjjjsu8Y3AvDAA2LRpk9YHClEnQg2WBKwakIAlhBBCCNAUoRBCCCFEnZGAJYQQQghRMBKwhBBCCCEKRgKWEEIIIUTBSMASQgghhCgYvUVYA7xFiBFCCCGE8JEGSwghhBCiYCRgCSGEEEIUjAQsIYQQbcf7778fH/Uly0+IomgKAev888+PBg0a1MuceeaZ0dtvvx2HqD/k6d90oV0I0bz4zw7fsKm3+bcC/VXO119/PTrqqKMS88fvyy+/jG2twSeffBKNHz/e/YNfh9CvEdxwww2uDNYuQ+O3U9+cdNJJpX6Q/oi+slkI2wXlFb1pGg3WnDlzovXr15fMsGHDoqlTpzb0JhBCtCa8q4Ph2eHbb7/9dmcX2bz88svRM888485ZCH67d++Oba3B6aef7urCP/h1CP3qDULIu+++G1166aWldpnVTvEz99tuu831g6RxzjnnOP9mGfi3YrtoNE0jYJ1wwgmuAZl5+umnnfuaNWvcvxBC1IKvpVmwYEHs2qPRQFNgGoO0QR3+SWnQ4YWjd+zWEYbx0FaglbA80W74VFNOP4+0DjiMb5oR7Dxnf/Ob37hjH/ND42N+/FteCGVJcawMec8t+OlynFV3f3aDcMTBHQ2P5UFY+0+qA4RpWT7+seXplz0tzyReeOGFaMaMGbGtMq677jr3b4LMVVddFb366qvuOCQsr9XLP4a8dcQtrU3hF55TqKbttjXdUnK/M2XKlEOLFy+ObT3s3buXodSh1157rWTvbqTODcMxbgbHc+bM6eX/xRdfxL6HDm3btq1XfPL04+PWPXKIbX3tSdx7773xkRCiGeCe5d4NwY3nA/AswM7zgWdAR0dH6V7nH3sSaWkk5Ynd0kyKh528LX/cwQ+L37Bhw9wzsFw5LZ7/zPOx+PY89csPPA8t7ZDQz88Ld+w+2HEvV+YQSxf88maVnTQ5R4QB+hGe8+CXK6kOMH/+/FKegB2TVfasPJPo7Ow8tGrVqtjWA2n45TNw88u5ZMmSUl2B/6RzmHWOqqkjED+rTaW1C/DzL5dPO9P3CvcDXCguDCfeDA2WhmmN2MJgx3CMm8Ex4e2CWnwDf9wsvtkNGoM1AAjtSUjAEqK54J7l3g0J3ez+puPznxNA55l076elkZSn+dmxj+8HPIvSwlpnWK6cYbwQ4hPeh+efDWz9MoSEfn5euKfVr5JzC1l1Tys7z3vi2THPdl9YNdLqQBy/sycf4meVPSvPJAhLOJ+k8wa4+YY8KYsP7mF+5c5RpXUE8skiq10AdvzL5dPONM0U4bPPPuvUjWa6G0W0evXq6Oijj3bzz6gj77//fmfH3Hrrrc4NP/N//PHHo6FDhzr/uXPnRlu2bHF+xl133VWKP2nSpGj//v2xjxBiILJ9+3b3rGGax8x3330X+/Y/RxxxRLR169aay0l8mxYyRo8eHR/Vh1rL7Nc9rew877s76uijjz6KujvxaOzYsc49L8Tv7uzdFJpNW7E2K6vs1eRJnLyQdnff7Mznn38eXXDBBbHPjxw4cCA+6qHcOaq0jkXSqHyakaYRsLol7VKj6h4ROLeZM2e6f5t/PuaYY9wcLoYGA/iZvy0CBI5Jyxr2K6+8En366aduXpgL/Oc//9m5CyEGLiNHjozOOussN5gzw8Pff5b0J99//300ZsyYmstJ/G+//Ta29YCAUE9qLbNf97SyM4A+7LDD3Nof0r722mujiy++2PnlZfbs2dHKlSujN954o7RWKqvsReRZNOWub6V1LJJG5dOMNI2A5YOG6Q9/+ENJQ2WYAOabPBeJhscoY8WKFdG4ceOcMMdNIYQY2PBc2Lx5c2lkz7OCRbqVLMK1Z5AtBH7uuefcf7XYK/uUBc3+aaedVnM5iY8WwRY7E2/p0qXRtGnTnL0c77zzTnzUm6y6V1NmW/Af1j2t7KSPcEN4oDNPI60OkydPdsISaV5yySXOLavsleRp+P1YEZx66qnxUQ/lrm+ldcxL2jn1KSKflqVbSOl3mMtlrjiE4rHIr7vh9Jl3tnVUuCX5+27M9YZVJT/iG/j7c8KhPQmtwRKiuUi61yF08+9vnhGsEcGNtSpJzyLISoM4Fp+Fxr5fVjzw17LgR3zS4Zg1SEZWOcM8kiA+a18Iy7+/tidcT+NDOCsPhHll1T3vuQXCZNU9reyEIw75YAgLhDXK1YF4pOuTVfa0PJPg3PrlBc5PWAbALe06AH5hOY2scwSV1jGpfD7lzin2atpBO5F9BhsEDTDphONubyVwzAVCcALcuVCG74/whb81Ji4uF9ZuAruJiWOEFz1PI5CAJYQQxRB20O0CSgLrx2qFtHzBUzQ3TTlFaDD/zgfagDVUfN29W2hya7BQub733nvOD3x/1moxx/vWW285Pxb0dTfKaNSoUS7uunXr3PdbfLA/9NBDbn1Wkl0IIYSoFKbp6MdsSrEWHnnkkWjWrFmxTTQ7g5Cy4mNRIQsXLnRGCCGESIO1accff3ziG4F5YZ3bBx98oN0JWggJWDUgAUsIIYQQSTT1FKEQQgghRCsiAUsIIYQQomAkYIm2g8Wk9kHaJGPwUoTvzpeQ7TsyQgghRC1IwBJtB1/s521SdgRYtWpV1NHRwfvfbgsK3A0+ZguEw58vHF9++eXOTQghhKgFCVii7dixY0f0wAMPuB0BNmzYEJ133nnOna9Od3Z2umP44YcfnJ1w8Ktf/Srat2+fOxZC1Bf7+nsSWX5CtAoSsETbcd1110WXXnqpO2az2IkTJ7pjYLsIg++hmfAFu3btctouIdLwp5gbCe2W7UWS8sev6K1Y6g3bpLCpv22X4tch9OsvEPLqcb3D69VfbapSmrXtp/n51OtalkMClmhr2M9yxIgRsa03H374odvrDLgB2QftzjvvdHYhmomXX37Zffw46as6+NmG960CH3+mLvyDX4fQr91oxevVn5Rr+2l+TUF3wUSVaKuc5iZtvy9gOyX8zLCt0kDZH6td4Tom7WWX1A6w2z5pYTzaAfus0Saw+9uchGH9bUvCveBsay7w41m+IWF820sOu298kvz4t7yefvrpkruB3cqQVeYQP12Os+ru74NHOOLgzvZklgdh7d835gZhWpaPf0xY/v2yp+WZBPGT6uW3m3JtKG9+xPGNuVXTpnzS0ihX7jBeq7d98iG+GUsnPA9p1yurHtXQu8SiIiRgNTfs2+XvN+nDjeffcKL14Xpah8CDETt7k+bpZMJ42BHCMTyI7UHrh8WPhzAPdQtnafKP3bB4tpdqiMUnLfDLD7RjSzsk9PPzwh27D3bcy5U5xNIFv7xZZSdNzhFhgA58xowZ7tgvV1IdgI7Q8gTsmKyyZ+WZRFq9SMfK4R8b2HGvNL+06wX+ucuqY0haGsTh2Ae7pZkUDzt5W/64gx8Wv2Zs+355TVAEwpCmHSddr3L1qIbeZ15UhASs5oabhpsnCR7SWQ9B0XrYA9TAzkPSf7ga5mfHPr4f+A/xMKx1+DzM2Wzeh4d4WrwQvzMw/PZbSSfj54V7Wv3KlTkkq+5pZaeTJJ4d04n5HbaRVgfi+J0c+RA/q+xZeSaRVi//3PnHBvZq8su6XmDpVnJ90tLApPnZsY/vB63W9sNjy9s/TrtelZzvvGgNlmhL+BaWbRQegt/SpUt7fbJBiGo44ogj3IsU27dvj7of1m5zeDNsOJ8X4vMdNp/Ro0fHR/Wh1jL7dU8r+9ChQ93nUT766CP3xu7YsWOde16Iz33K9+ls0Ttrs7LKXmueVq+81JpfGrVen3rTjG3/nXfecWlRhldffTV27U3a9arH+ZaAJdqSmTNnuk8u3HHHHX3esHrhhRfcjfTggw/qdXBRE99//300ZsyYaOTIkdFZZ50VrV69umR4OPNpkDwQH8Hfhw6gntRaZr/uaWXn3jvssMPc216kfe2110YXX3yx88vL7Nmzo5UrV0ZvvPGG+1YdZJW91jytXnkpoo5J1Hp96k2ztX3OPwPnTZs2uTJcdtllsU9v0q5XPc63BCzRlnBzHDrkpsDdiMWH3ejNr1keVqJ+2DU2Yfq5555z/9Xy8MMPu386hWeffda9icooePPmzSUtC368Pp73UwPER+i3nQSIR2cxbdo0Zy8HI/cksupeTZl50xbCuqeVnfTpvKwDpRNLI60OkydPLnWel1xyiXPLKnsleRpJ9RoyZEjJLes8VpNfWl19am1TMBDavk9HR0fpu4a+Bsu/lmnXq4jz3YfuTkZUidZgCdE8hI8z7LZ+gvUW2FnPw2Ja3y8rHoTrUIhPOhyzBsVgHQdrOHDH39aQQJhHEsRnzQdh+WdNiJG1DoVwVh4I88qqe1aZQwiTVfe0shOOOOSDISwQ1ihXB+KRrk9W2dPyTIL4afXiGDfOV9Z5rCS/cnX1082qo09WGlnlzooHzd72fT/WUmH30/Dz9q9l2vXKe77zMoif7sREFSxcuNAZIYRod/hQYzt2F+1aL9H/aIpQCCGEEKJgJGAJIYQoS7tqeaS9EvVCApYQQgghRMFIwBJNA2sh2tEIIYQYeEjAEk0Dqvp2NEIIIQYeErCEEEIIIQpGApYQQtSIdgQQQoRIwBJCiArhy+K2BRNfeh4/fnxtX3yuAL46Tv6QtOaPPdTMz75Q7oO/fZEbwdDCCyGKRQKWaBrYmiDsLHxjnHnmmb3c2dzTtlkomjAvM77Ggvxxs04LOMZNnVdrgZB06aWXxrZ0Xn755Wj37t3umM2HWWvHf71BqGMTc7+M69ev77Xmj22iDLYzybo3wq1UhBDFIQFLNA2ffvqp2zl/79690apVq9y+UnQYdCC4G3/4wx/cP+HwZwPYyy+/3LmVo9IRu+Xld2CUxTomeOqpp6IpU6ZEy5Yti1169juExYsXu39RfxBo0eywfxjHCxYsiH16BCcThPn3tU2Eww3DHmVz58517qRl7hgTVEhjzZo1TmvFsbmBHw4sX/84rQyUG3fap+/nw0bltuFxHubPn+/uDdO2JXHVVVf12rdNCFEMErBE07Bjx47ogQcecJt1btiwITrvvPOcO8JMZ2enO4YffvjB2W1Tz1/96lfRvn373HHRkBfCk8/nn38eH/VAmLvvvjvasmVLr46MMjZCqyF+ZN26dW4H/G3btkUPPvigux5oRidOnOi0TgjI/GMHhCg0QlxTDEL9rl27nB/Ta2+99ZZzR4ieN2+ecycN2oRpjnwQflauXBnbouiNN95wblllQOhnc9v/+7//c3lNmjQp+s///E/nF0JZx40bF9vKw72BtitpqtA466yzSlOOQojikIAlmobrrruuNPWxdevWUgcEfgdAJ2rCF9Ah0jHWA/Kiw8uCMOzWTke6du1a50anyfSiaCzPPPOM+zfBlmm8TZs29dI68o+2qNy02H//93+X0jn88MPdjv/lmDVrVq+2iuB0ySWXZJaBtkPaaKcQCEnjrrvucuFCEOL/4z/+I7b1YJo0M2G97r//fie4Pffcc7FLb4YOHeoGKGlaMyFEdUjAEk0JUzAjRoyIbb358MMPo9NOO80d05kwOr/zzjudvWjI64477ih1XkkjfcLQSV100UXR8uXLndsHH3wQTZgwwR2L/mX79u1OgGHqzQxaLkCgR7hgag6D4G5C/jvvvFOasss7hUY7QJBimtAEFoS0rDIQB23YRx995LSeY8eOde5pEN4nXINlQpyBpheN2fXXX58pRB04cCA+EkIUQvcNKark3nvvjY9EkXR3GMy7xLbe7N271/mZ6e7MDi1evDj2TcYPn2TS4lte3R2js0+ZMsW5+WCfMWOGOyYc4c1t27Ztzl00Bs69D3ba0qpVq9y1S2LJkiWH5syZE9t+5LXXXnNty6532CZJDzfD97M058+f7wxklYF247cV2iN5J5FWxyRCP9LtFuCcCdt8VjpCiOqQBks0HazFYo1LEky1QHfbdYapD1tQnoaFxXR3Ii5t3y0tvk3rmMaAt7Ns3ZdBmNGjR7tjwqGBYJqQtTJaf9UcoBHavHlzSXvDeiim57BPmzbNaSVNQ4nGyqbSmHa2652kwULDlcTkyZNdmjY9CFllwJ3F9bjByJEj3X8aWQvWs7B2zjRjEqeeeqr7R4NXr7dyhRhISMASTcd7772Xuu6Jxe+sdWoE/kL7NAjzi1/8IrZF0fTp06MVK1a4hcOiOUBIok3Nnj3bCVE/+9nP3JQyAvBjjz3mjk3YZlE7U2kIScQzoevCCy+MU+uBBe98AgH/EJsmBBOys8qAQEObxo11e/fee68rRxIMDv73f/83tlXO888/32e9ItPslJcyIrwxOPj3v/8d+wohqqb7oSKqRFOExcOUTHcHkDhthx9TJzbtUg1Mg6RN1fjkycvCMP1jMNXDbZU27SiaC6ZymdIzuH60v2YlbUqzFkizlntKCJGMNFiiqZg5c6ZbdMzC8nAqhLesWCjM6/fl3gCrlTx5WRhe4TfQSKANKDfNI5oD3tbjxQS0VGiP0DDx9mCzwpQmGiabTiyCRx55xL25KIQolkFIWfGxqJCFCxc6I4QQjYI1Yscff3x0wQUXxC7Vw+CBN17LrWMUQlSOBKwakIAlhBBCiCQ0RSiEEEIIUTASsIQQQgghCkYClhBCCCFEwUjAEkIIIYQoGAlYQgghhBAFIwFLCCGEEKJgJGAJIYQQQhSMBCwhhBBCiIKRgCWEEEIIUTASsIQQQgghCkYClhBCCCFEwUjAEkIIIYQoGAlYQgghhBAFIwFLCCGEEKJgJGAJIYQQQhTMoEPdxMeiQhYuXBjdcsstsU0IIYQQA5XDDjssPupBGiwhhBBCiIKRgCWEEEIIUTASsIQQokH84x//iI9Eq5LnGuo692UgnpO6C1jvv/9+NGjQoOiGG26IXYQQolgOP/zwRPPEE0+U/Pubf/7zn9GUKVPcfxZZZX3rrbeir7/+OrZVxq9//eu26OQ4Bz/96U/75ZrmuYZ5r3M9+N3vfufOD9g9YOb000+P3n33XedHO6A9FE3aNWn0Oam2rddyfyVRdwHr1VdfjYYNGxa9/vrrsYsQQhTLDz/84MyaNWt62W+66SZnbwZOO+00Vyb+Q+h4rr766tiWzl/+8pdoz549sW1gwjl4/PHH3blsNFnX0MgTph4gGPztb3+LLr744tglcvcDZcHceOON0bRp01y4n//8586/UQJ3f52TSin6/qq7gIVg9cADD0T79u2L3n777di1OViwYIHTrgkh2h9f83HffffFrj3CDaN7G+UnjbJ/+ctfljQDwAgZNwPNgNmz0vNH+JQBf8zll18eXXvttbFPclk5pgNFE2DpZOVFGubn1zcJwvh5ovmjThYfzYiRJ08zpjGBMI+sMoV5WDrYOQcIoxyHZOXh+yFYZNWDY64nfoQ3TSjgZpA+/sSnTVgafpi0ugBuaeVNSzuNrq6uXsJVyDXXXOP+TYCgzS1btswdh2Rdx6xzA1nn3wjPidUtKS/c/OMwDlh5wzyTSKsbccP7q1bqKmCZ1urSSy+NZsyYEa1cudLZ4aSTTnICjg/2M8880x1/++23Lh4CEIZj3AzcENjOP/98Z+CTTz7pFQd3Pw7lIV/8+N+6dWv0xRdfOL9y+QkhWpsNGzZE33zzjft/9NFH3Uj+u+++i6ZOnRotWbLEjbD5xx5CB0c8g2fHV199VZpO2LhxY3TuuefmTo+HPA9znlmYI488Mtq9e3fsm1xW0iMP00pk5YXfzTffHD3yyCPOb8SIES6/LPw877nnnmj16tXRe++95/J+8803XYdWrn7kiRaAOpH3bbfdFvv0kFSvEMuDMpAH6ZnmBbt/DpKwPAj//PPPu3Nt4Ldu3brolFNOyTx3HM+bN8/5rVq1ypXF79ABIY20t23b5uqL0PHYY4/Fvj1k1cVIOid50g7h+p599tmxrS8vvvii+z/22GPd/5gxY9x1TSLtOuY5N0n18clqQwiItDtj+fLlzi0rDn6VtPW0uhG3XNuqlLoKWCtWrHCCClx00UVOwDGhBYHLl1Rh6dKl0ezZs93xzJkzo6OOOirau3evMxzj5sNFvuqqq6LFixc7+x133OH+LQ789re/df9ffvll9Jvf/MY1Wj799dRTT0WbN2+Ohg4d6vzz5CeEaF2sg7JpCkbyW7ZsiU488cTSlAn/HR0dfaZOePBaZ8RzCzsdJYIW8EBHCMubXjmSyhqSlZf5nXfeec6PTooyZxHmOX36dPccxNAZHzx4sGz9eIZbfL4JhBDqk7depGnaGMJyTH+SB8uDctOf7Nixw9kBvxNOOKHsuQvzp8O1MhsIKtSvq6vLCRFXXHFFL00f5KlL0jnJk3bIxx9/HJ188smxrQfTxmCefPJJ14apP/C/f//+PoIjpF3HPOem3DXOOvfU0xf66K/L3Vfml7etl2ujRVI3AQtBiopceOGFzj558mQ3Tbh27VpnnzVrljsxCD6ANIk2iXC4IUXef//90dFHH+3Mrbfe6twsPCAkIcCh5jPuuuuuUpxJkya5BgSMDlkLZmEvuOACVx4W4efNTwjRXuzcudM9YHmIm+G5EMIDmQc6nRHaKp5r48aNc1p5OkDSIEze9OgEeDbZNAUdgnVaecnKCz8EjKIpVz+e79QH97Tpp3KQhw18Df8ZXwkIFggeIeXOXZh/Eggo9BH0XRMmTIgmTpwY+/xItXXJk3YSJjwZpIEAhCEtE0J8EJxD0q5j3nOTRda5p/wISwxiTPArd1/hV0lbL6KN5qVuApYJUjRwhJhPP/006uzsLEnuXCTsFo6GgLSNu6nKjznmmNKUHcIR+Gr0cJ70lVdecfkw1cj04J///OfYJ4qGDBniBDgaGaBN44F5zjnn5M5PCNFeDB8+3GlnmIoww/SGjZR9EILQVDGqJg6dHna0WGizIG96TNeQHs8jDNMslZKVF35MnRRNVp6cFwzTirijAauGpLLbc7tSECzOOOOM2PYj5c5dnoE1gvXgwYOjl156ycVlXRjrmnyqrUuetOtF1nXMe26yyDr3cOWVV7ppQtxt0FHueuVt60W10bzUTcD64x//6P7Hjx9fMmis0GrZNCHqW+bIAUmS6T4fpvJCg0CUBGmOHTvWCXCMLJk29BeNIrghwPFQRIC6++673Un2qSQ/IUTrwyAPAclGyzyoWaCbNG3CoI0Oj/VSjLQZNTPaZk2HrQPNmx5LJpgKsekbRtS2RqYcNijNyoupIjQ3NnVHp5K1LiUv5erHuTFtQrXaAfJAW2FLSEib8nPO8mBTaZSN/oU1OSFZ9cAP7Qh5AsJO0jUkPkIPcYGOPqTauuRJOwnKWimsRwtJu455z00WWecemHni3iAPtEyQFafStl6ujdr9VQR1EbCQcBGmWAAXCixojawCTAcSDg0X/9gBbRP4kj4CFA+xNOkfzRUaKjRTTP+FalgWxH/++efuwlAOji1MNfkJIVofHrQ8p+bOnesEnVGjRrkO2tZo+DC9Qufir++gA2Cqjw4A8qb3pz/9ybnb9A2LbvN8UmLOnDlOaCDtrLwQABEGWeuKH+uQstal5CUrTwav+OPOs9WEzkqxPFh8TFoIGv7aoXIwwKbzJTyD+KSp16x6mB9LUPBjii7pGpIuhrgsQl+0aJG7jj7V1iVP2iFc388++yy2lQeBhAECZfTJuo55z00WlkbaPcK5oVxgbllxKmnr5dqof38VQV02e37uuefcWiiTvn344OiHH37oDPDWIA8tHlD21iFQcVvHhfaJePhbmmih1q9fX9IwIQhx0nnrghNHWOKcddZZTt1oAhOCnEGenEw70Vn5JaHNnoUQ1cCUD1p9e3WekThvRTHtIaqHjhGBdSCCBnT79u2lReblIPyuXbuqmp4WyTRks2defbS3B0NYHIqQY/O4CChonkKVKeupEL5YC4UwRfhwSs8HIWn+/PlOyCI8r+I+88wzsW8Ubdq0yQly5IVMyT/p2+im0vyEEKJaGH3/9a9/dc8tNBSMzBncCVEt9KFMjWUpBXx4q5C39kT9qIsGqxlBQ8Xc7u233x679Gzjwyiy2lMgDZYQQohmAa3Ucccdl/i2oA/Tgygdmmmng3agIRqsZoQ3EFjQxjQg8M86COZthRBCiFaHKedywhXw9p2Eq/ozYAQsVPJMAf7sZz9zU4D8s9jt6aefjkMIIYQQQhTDgJkirAeaIhRCCCEEDNgpQiGEEEKIRiEBSwghhBCiYCRgCSGE6HfsS9zNSDOXrRrarT7NigQsIUTbw/YZ/jYieb/UXG28Wqg2DzrNcnH58nxa58qLQLYFSjlIw7YxKQI+tMpWZrYVSjNdr7BsrU6z1idsv+Xas99eCecbvi/3xBNPOL+i22ol1Cxg8X0p3srzDR8ZrXVDyHLwtfg829jkDSeEaF/YZmTPnj2xLT/Vxms1EEr4SKV9eLkctjFvmrBWKWx5whfYbWuUZrpeYdlanXaoT1J7XbNmjasXZsmSJe4L9bTPottqJRSiwWL/HratMcNX0vN8i6MW2Ipn48aNsS2dvOGEEK0NI1dGtOxDx/F9991XcudhzKidYyMprE818XAzfx7oaAkYTePOv681IC7hcGeE7ftl5RGmaRsJJ0E6FjapjkZXV1evzoqvgfOFeQwQN0yDffXKbehMfMpgUE9LEyi72Unf/ht1vbKuj4+fHsdpaRKf+uCOv2lRwI9Xrm0QDjcz4TUmT9LBz287WWn64G/45eI46dyCH65c+TlOOw/l6paHsL2GIFSxHyF7FEOetloPChGw2GyRPQHNsOUDQhYbLNcLHgDXXXddbEsnbzghROuzYcMGt58f/48++qgb6TKi5WFrI1wjKaxPtfFwZ6uuU045xe0vyGjaRtXYgQ6Kjoa9U9Gw0xn5e8il5cHzjDQYnZMmGptp06b1KQMQ9uabb3ZblxF2xIgRTgBJAvezzz47tvV0YMuXL48+/vhj1+HSibOhLnaDjzezcXEWxKMOxtatW6OvvvqqVF4Gv+HGvNWe90rj2blMuj7lSErT0ps3b55Lj82JuU6+4EH4cm0DuG5cW9oG14/Noo20tlN0fZLIU/5y5yGrbnkJ22sIX7SnrbF7C+Rpq/WgLmuw2CwZ2HgSmDZE2GI6EQN8SZ2pRH9a0b6yDhyz4XLatCNubHUDpH3SSSc5Nz4eumDBAucOfrikNP3pQ9xIy8pFWv4G1EKI5saEFJv+yJouqiSsT7l4+DPoZM/VE088sTRFwX9HR4frII899lgnaHR1CzJ0ZuwJx5oSIy0P0iQNG73jz/GKFSuc3cfyt9kEwoXCjIHgdPLJJ8e2yH3lm84RJk+e7PL5/vvvoyOPPNK5AXXcv39/LwEihPysY0NTgR2BkM4P6CgRwvJQ9PXKuj7lSEoz6dpwDi0M5GkbgJLC4vFtJdqKkdZ2iq5PEnnKX+48ZNUtL2F7BdNcYvhGJflQVsjTVutBXQQsE4RGjhzp/gFp9qqrrooWL17s7DNnznQCzN69e53hGDeD4w8//NBpwvCHpG1tEJqQlp966im3pyAXlxPL2quQME3ypBH4gh3lZNNM0kLQQiATQohK2blzp+s8ECDMsOE88MBH08IAb8KECdHEiROdezlI0wawBtMsSRCWZ1xerDMyEILOOOOMUifKgJm9W0MOHjwYH/WFjpTOlo4NbRWb/Y8bNy5auXKlEw44P9bZNpqs61MNSdcmjXJ5r1271l1X3MOprbS2U3R9ssjKq9x5yKpbJYTt1TSXmNmzZ0fXX3997PMjWW21HhQiYHGjoCUygyA0bNiwaOzYsXGIyAlACCycWAQwTsb9998fHX300c7ceuutzg0/83/88cfdhcKf3eYRnnwtFjCq8iH8Aw884NThPpbm888/X0qT/BG2uOCGlRMuu+yyujVQIUR7M3z4cDc1wVSbGaZhEFh4Zg4ePNhNu+F29dVXu3Ui5SBNpmB8fC28T1LYStixY0cvjRfaJrYbqxTSIC7TWpwPBALsCHBos/qLrOtTDaQX9k9pZOXNecK89957zn369OlxrB7S2k7R9ckiK6+s81CubkWBvIAA2GiNVUghAtazzz7rRjZmEK5QByPEGKjtjN27d7v/Y445xk3FYYgD+Jk/67kMjtEqhZIx9vnz5zstFtOEaJzYTdyPC5amP9qjfJ2dna5hGH45DZtibDdYeGgq1WqNEK2CP5CqhGrj8WxBiLCHPMIOC36x406naAIQnVIeSJOOwxYGkxYdFlp3po7A0mQKhakUmyIiHIJNGnTcPoT1B6rk+5Of/MStb/FhPQ550tEndWgsC0EYYHoRrQNaNaaXWItjS0aSqPf1yro+1UB6DMg5z8D5TEuvXN6cK9M+hlqetLZTdH2yyMqr3HlIq1vYfkN7SNhefUjfBHsf2mojKUTAYtoP4ccM65ZCQSgJP46ZUDDKw6JFi5wmCs0VFwMhL+/aKV8IHGgwd8+DDtAgmnq1nKFhM3UgRKvAm85orysdFFQbD3jIs8CX0TTxR40a5dbKMCXG0gQMbixS5hnGwt9yWJosDCZNOlrWOCG4YLiP+Ueo4h/BhhkFwoYaKR/cP/vss9jWQ7jOhY6RaR2EOSAPnh+UiecCnVmSAMcaMDpcP2/SYU0MnXESjbheWdenGiw9ZkFIj+m7tPSy8kbDhz/uKARCITSt7RRdnyyy8so6D1l1S2q/vt0nqb2GcG7+/ve/u2O/rTaSmjd75gSxUv/222+PXfqChorPN5jwhPoQjRVvQZhGiXVQrJFCWGPhW+gfxrE0Ccvcvv+mIIvcka5Xr15dCjdkyJDEPNGivfbaa6WF7X450VwhrPluPu2w2TOjCtZEAA/qvJ/XYASNep8HqxCitUErxRorW+ycB+Ls2rXLvSEGaLCuvfbaukxJCeFTaXsN22q9aIrNntFuseKfhWg2V/v73/8+2rx5sxN+Qn8EIV4dRUAy4ciHxWw2jUdYhKukqcQwTfIkTd6SGagwqmCUC5ybvOrken/nTAjRONBKoX1Km45J4sknn3RacEBDwAJ2CVeiEVTaXv222kj6RcCCV155xS2YRMBBc4TQw8I3w/dHy8SJtDldHwQu1JG8oUg6hEWVzQL2kLQ0B/I0IaBuRrhCZY/KN2+jlfZKiPaAqZMbb7zRvUiUBwQqNFZM3wCCFc8RIRpBJe01bKuNpOYpwoFMO0wRGghVrItg3QXCViVTBUIIIcRApymmCEXzwYjgv/7rv5z2j0WiSdpCIYQQQuRDApYogQoV4QrSXrkWQgghRHkkYIlesHjdNvv0v7VSCXzGgTT8TTx5ZZZXc82wdq4/tGTNUo4iSKqLf86FEEL0HxKwRB/Yh4zvjPBhwUrXmKH1uvLKK90bG/6bhrY/JMIXi+PRkPGhwUbTLOUogrAu9rKCEEKI/kcClkjkhRdecOux0O6EX27Ogu+MPPTQQ33e2Dhw4ID7OKl96I3XbHlrMfyAXL1plnIUQVgXPjdCXYQoRyu2dzFwadX2KgFLJEKnzecvAI1WnvVY3AR8biPpWzgbNmxwe1cZ9hVewjINyXQXBpheZMrLpiqLpFnKUQTUxf86NttBIRQPVBgMsCUH16yRkC9aRKPR+VcK9zLfBGzkGstGn6Nq0uf5lRWPt6z7u6OvdxmasS3nba9ZZQ3r1SgkYIlU+Agp+xXmZdOmTanfwuHjr/bFeB4QbO9ggktXV5fb+JNPRODGQ4SPn2IvmmYpRxFQF9srjrow1cmWFAMVtgthg/hGf5+NfPfs2RPbmh/ua84R/42i1c7RQKUZr1NWe0XoYplHOfqrXhKwRCpodPh0A/tB5XkYs+9T0iiCdBBSuBHwZ2+0efPmOc0Y8G+dItNc5PX999+namNII8ukCYVFl6MIWJSe5wEREtaF3QwQrqwuAw3OAV92tvNhbqbVQgDlYcyLALiHLwT4YTmmDeFv4dMEV/zIlxE2x4aflq8BDcuQNirHPy0N3y+pXpYmx34dLRwQ1gjjWxzS9sMBdtyBMlEG4jAYyapLJecIe576QVYZ8l4D/xz5EN/C+fHTqPZ8ZdWvkjJYWDN+vbLOk0EeRbdlZgOIb5C3zRAAZTR7Vnp+eSgD/hhexGJ7JiOprBwn1asRSMASqbB5NtS6f5N9bRfhhWkt1gn503SANob1RDa9yD5T7AOZBOlkmTQho+hy1AoPAPZzrGbdlF8XzCeffDJghSvgHDBdumbNGndscJ3XrVvndtGfOnWqa8v4M6Ll3PvTBoT95ptv3D/h2MuU3SUIwz6dSR1IVr6WFtt8kQZCMWVYsmSJC8s/9jQsDeKG36bDz69XUppok6mDgXY21DBbmbLOSxIIDZSHvV1pe3SSaR8nruQcGXnqV64MSennrS/h0Ag/8sgjLhyaYjrpakkra1abqLQMhKU+pE8cNgSHvNeKPIpuywhUxDN4vvLylJ1v9hEmz7zpUQ/OAfXAMPhlaYSRVNa0ejUCCVgiERoynQoarFrhJrIHO1ohGvuKFSuc3Qh3+ucmYlujImmWcgAaErRk/oiwEvy6iHToSHjhAoG0o6Oj1/Xn2L/+1ungB9OnT3drETEI4gcPHnTueQjTYnqCMrCjvwnv/FMm03CEWBrkz9uhtE3Dr1damrzJyz1scE/T4fnkOS9JHHvssa6j7Orqcp0YeVU6PZ10jow89StXhrRrkKe+lq+9CU0Y/7lQKWllzapfpWVYunRpqa58UZz8oNZrlXYe08rtQ3mtDaKtwo5Ai6AFPF9pk3nTK0dWm+oPJGCJPjBSZzTEDRu+DZgFIyNGDz6MTHiwc/MYhGM60YcbzdYTAQ+En/zkJxW9wZhFs5TDQNtkD5NKSapLCGGoHwYQ5Hy1+UBj586dfTaAZ4qhkVAG2hMdipl9+/bFvtlw7ZLWAmalyb1LG6FjM+2bdTxGteeFtNEIoEWYMGFCNHHixNinWMrVr9Iy5K0v4ezt3CJIK2tW/Sotw9q1a11dSGPZsmWxa32uVVa5fWhvCEq0PwaFbAjOGtiVK1c6YY80CJM3PYRMNP7UE4PAZsJyMyIBS/SCjpkNnxnhVCoAjB071gkoPoyauHFQ19poZPjw4S6cv0aAzuPkk0+ObZFT/XKT8RmFImiWchRBUl1CCNNKC/brDdeatu1Dh9NIKAOaMK6LGQYkee4zpjaYug4plybfpGOaEPekjqja80LnOHjwYNemyI+1b6yHKZqs+lVThrz1TQpXC2llzapfJWVgwIVhSps00L4a9bhWWeUOQQjiOUv5iIOAhx0tFtosyJseA13S45phal2+Um8kYIlesO6KUVM163m4GRBI/E6fdOgcMHazoPLG7n+IFLs/uubmQrtU1CiyWcqRBx5ECLhpD9ekuoRYGGjUgv1mprOz0wmlJkwzouY8FyU4oz0oB2WgUzFtEteXBblmD7FpHMKxBsvXrBrl0pw0aZKboqGuCNkhWefF2pbdz74WlzzppK2N0kGWI885CsmqXzVlyKovU2lAegyyGIxY3QkTDh5DqjlfWfWrtAzc2/ac8jVYlZ6notvy+eef74Q7yoc2jTKiWWWWBD/Imx7XifaMRheDFivv7EI17a9WJGCJEtzANF4+Mlotf/rTn9wbbYyaRHWgaWCUOWvWrNJDsRp4YDVqwX6zw0Od77qx8JcHMx0ObZ0Hfq3MmTPHCUCkm4WVAQ0xYUeNGuWEKF+g92EqhU6GMrIGK0kDVS5N4tpUclI+5c4L2k/evqIcRxxxhHMDyoIhP6ah+dwJC6zTyHuOQrLqV2kZIKu+GN6Y5p+1OwgFvGlMuHBtZhqVnq+s+lGOvGVAK0RahEPoMMEFKjlP9WjLDGCZ7vPLjrDPVB+CFeRNj/4FdxtgUo88yoBq21+tDDrUTXwsKmThwoUVbyXTrDBS4IHO2xdJN0klIFx1dXVFZ599di/tkOgLi93RkKESD2FkhlBkCzcrxdZbmRqdBy8PpFqvr6g/dASmgRRC9MDAk0HiNddc4+z0W7xtyExDM8DLBT7SYAmnJWEkR2dfSefLKMSfDjQYedGpS7gqT1YnijqcUVe1WiymExqxYF8IIRoB2qu//vWvbrCINg6NFy9jNSsSsITTwo0ePbo0KshLuTUJIh2mY3lAsFCd88gxbj6ozVGrs3C1Glphwb5IRtorIfqCAgBtPwvc0fxj0tahNgMSsAY4aDM++uij6I9//GPsko+0r6UPNJjKyTJp54k1ETwcbC0Bx7gl8e9//zs+qgzS9TWS/bFgXwghBipag1UDrb4Gy9Zd1YJG2vUFjRParTwLOYUQQvQfWoMlSjT7N0SEEEKIVkUC1gCGuWyboqrWiPpz/PHHx0dCCCFaBQlYQjQpvD3IAni+cCyEEKK1kIAlRJPS1dXlPjDJZy+EEEK0FhKwhGhCeAGBN/7uvvvu2EUIIUQrIQFLiCaD72GxRQZbFumTCkII0ZpIwBJ9QHvCXlp8w4mtCfiek22OKuoP38NiexwJV0II0bpIwBJ9YKNTtmhBi8LWBOxpd+DAgdhXCCGEEOWQgCX6wB6CGzdudN/J4kvgbEugN9mEEEKI/EjAEon4nwfgGA3W119/7exCCCGEyEYClugD31/av39/r88DMF2ozwUIIYQQ+ZCAJfrA4mqmBQ02CWbRtRBCCCHyIQFLCCGEEKJgJGAJIYQQQhSMBCwhhBBCiIKRgCWEEEIIUTASsIQQQgghCkYClhBCCCFEwUjAEkIIIYQoGAlYQgghhBAFIwFLCCGEEKJgJGAJIYQQQhSMBCwhhBBCiIKRgCWEEEIIUTASsIQQQgghCkYClhBCCCFEwUjAEkIIIYQolCj6/znaMWM2rCn5AAAAAElFTkSuQmCC\" width=\"600\" height=\"232\"\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 4:\u003c/strong\u003e Rule-based filtering thresholds used to benchmark against RF models. Variants at or above thresholds were retained.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.141%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eThresholds\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 17.0686%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRead Depth Low\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 17.4397%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRead Depth High\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6883%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eQUAL\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.7792%;\"\u003e\n \u003cp\u003e\u003cstrong\u003edbSNP\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003ev151\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.8831%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCOSMIC v98\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.141%;\"\u003e\n \u003cp\u003eHard\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 17.0686%;\"\u003e\n \u003cp\u003e\u0026gt;20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 17.4397%;\"\u003e\n \u003cp\u003e\u0026gt;1000\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6883%;\"\u003e\n \u003cp\u003e\u0026gt;=50\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.7792%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.8831%;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.141%;\"\u003e\n \u003cp\u003eMedium\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 17.0686%;\"\u003e\n \u003cp\u003e\u0026gt;10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 17.4397%;\"\u003e\n \u003cp\u003e\u0026gt;750\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6883%;\"\u003e\n \u003cp\u003e\u0026gt;=40\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.7792%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.8831%;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.141%;\"\u003e\n \u003cp\u003eSoft\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 17.0686%;\"\u003e\n \u003cp\u003e\u0026gt;5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 17.4397%;\"\u003e\n \u003cp\u003e\u0026gt;500\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6883%;\"\u003e\n \u003cp\u003e\u0026gt;=20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.7792%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16.8831%;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eTable 5:\u003c/strong\u003e Total number and high confidence variants returned by from low and high depth datasets.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"473\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003e\u003cstrong\u003ecfDNA Accession\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal Number of Variants\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eHigh Confidence Variants\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDepth\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eSRR3401415\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e298,225\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e376\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eSRR3401416\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e318,772\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e446\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eSRR3401417\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e261,362\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e286\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eSRR3401418\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e411,688\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e602\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eSRR3401407\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e358,119\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e838\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eSRR3401408\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e222,430\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e369\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eSRR12083523\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e257,374\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e1,446\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eSRR12083526\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e309,852\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e1,592\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eERR855950\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e2,013,269\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e1,183\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 30.0847%;\"\u003e\n \u003cp\u003eSRR12083529\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e261,256\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 27.9661%;\"\u003e\n \u003cp\u003e1,817\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.9831%;\"\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-5735554/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5735554/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eCirculating tumour DNA (ctDNA) is a minimally invasive cancer biomarker that can be used to inform treatment of cancer patients. The utility of ctDNA as a cancer biomarker depends on the ability to accurately detect somatic variants associated with cancer. Accurate somatic variant detection in circulating cell free DNA (cfDNA) NGS data requires filtering strategies to remove germline variants, and NGS artifacts. Rule-based variant filtering methods either remove a substantial number of true positive ctDNA variants along with false variant calls or retain an implausibly large number of total variants. Machine Learning (ML) enables identification of complex patterns which may improve ability to distinguish between real somatic ctDNA variants and false positive calls.\u003c/p\u003e \u003cp\u003eWe built two Random Forest (RF) models for predicting high confidence somatic ctDNA variants in low and high depth cfDNA NGS data. Low depth models were fitted and evaluated on whole exome sequencing (WES) cfDNA data at depths of approximately 10X while the high depth data was sequenced at approximately 500X. Both models utilise a set of 15 features from variants detected by bcftools, FreeBayes, LoFreq and Mutect2. High confidence ground truth sets were obtained from matched tissue biopsy samples. We benchmarked our models against rule-based filtering with a set of hard, medium, and soft thresholds.\u003c/p\u003e \u003cp\u003ePrecision-recall curves showed the high depth model outperformed rule-based filtering at all thresholds in validation data (PR-AUC 0.71). Partial dependence plots showed membership in the COSMIC database, absence from the dbSNP common variants database, and increasing read depth increased mean probability of high confidence somatic variant prediction in both models. Our results demonstrate the utility of supervised ML models for filtering variants in cfDNA data.\u003c/p\u003e","manuscriptTitle":"Predicting High Confidence ctDNA Somatic Variants with Ensemble Machine Learning Models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-01-24 18:45:35","doi":"10.21203/rs.3.rs-5735554/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-02-24T05:13:18+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-02-22T18:06:00+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-02-17T08:37:42+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-02-09T13:00:33+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"131677874851583547665470439368509396882","date":"2025-02-08T17:16:55+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"294207432559441561788798341510947324653","date":"2025-02-07T20:01:38+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"4706038865681080070287746144981588622","date":"2025-02-07T10:24:25+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-02-05T08:04:48+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-02-05T08:00:38+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-01-22T13:18:49+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-01-22T13:15:43+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2024-12-30T12:28:14+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9bf32a48-2f80-4ef4-86eb-4bc9751a99de","owner":[],"postedDate":"January 24th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":43251748,"name":"Biological sciences/Cancer"},{"id":43251749,"name":"Biological sciences/Computational biology and bioinformatics"}],"tags":[],"updatedAt":"2025-06-02T16:03:26+00:00","versionOfRecord":{"articleIdentity":"rs-5735554","link":"https://doi.org/10.1038/s41598-025-01326-2","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2025-05-26 15:57:49","publishedOnDateReadable":"May 26th, 2025"},"versionCreatedAt":"2025-01-24 18:45:35","video":"","vorDoi":"10.1038/s41598-025-01326-2","vorDoiUrl":"https://doi.org/10.1038/s41598-025-01326-2","workflowStages":[]},"version":"v1","identity":"rs-5735554","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5735554","identity":"rs-5735554","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-4.0