{"paper_id":"03ecd11d-a2da-4f07-b235-553ce16ca547","body_text":"1 \nTitle: PhyloSOLID: Robust phylogeny reconstruction from single -cell data despite inherent 1 \nerror and sparsity 2 \nAuthors: 3 \nQing Yang 1,2,3,5, Yuheng Liu 4,5, Jijian Yang 4, Xing Wu 2,3, Zhirui Yang 2,3, Yonghe Xia 2,3, 4 \nYunchao Zheng 2,3, Jinhong Lu 2,3, Mengdie Yao 2,3, Yiheng Du 2,3, Huan Liu 2,3, Nan Li 4*, 5 \nYanmei Dou2,3* 6 \n1. College of Life Sciences, Zhejiang University, Hangzhou, Zhejiang, China 7 \n2. School of Life Sciences, Westlake University, Hangzhou, Zhejiang, 310024, China  8 \n3. Westlake Laboratory of Life Sciences and Biomedicine, Hangzhou, Zhejiang, 310024, 9 \nChina 10 \n4. Westlake High-Performance Computing Center, Westlake University, Hangzhou, 11 \nZhejiang, 310024, China 12 \n5. These authors contributed equally. 13 \n*Correspondence: linan@westlake.edu.cn (N.L.), douyanmei@westlake.edu.cn (Y.Dou). 14 \nAbstract: 15 \nWhile lineage tracing based on somatic mutations in single -cell sequencing data offers a 16 \npowerful approach to reconstructing cellular histories in vivo, its reliability is fundamentally 17 \nlimited by pervasive technical artifacts—specifically, high error rates and data sparsity. These 18 \nissues introduce false phylogenetic signals that corrupt tree topology and lead to spurious 19 \nevolutionary conclusions. To overcome these limitations, we present PhyloSOLID, a 20 \nphylogenetic algorithm designed to be inherently robust to these data imperfections. 21 \nPhyloSOLID employs a progressive scaffolding strategy that begins with graph -based 22 \nconstruction of a low -resolution, high-confidence backbone tree from reliable and uniformly 23 \ncovered mutations. This scaffold is then refined through the iterative integration of remaining 24 \ndata, guided by a Bayesian statistical model that penalizes phylogenetic inconsistencies to 25 \neffectively separate the true evolutionary signal from technical artifacts. Benchmarking on both 26 \nsimulated datasets and multiple ground-truth datasets demonstrates that PhyloSOLID achieves 27 \nsuperior lineage reconstruction accuracy over existing methods, for both single -cell RNA-seq 28 \nand DNA -seq data. Additionally, a user -friendly web interface enables customized quality 29 \nassessment, artifact removal, and interpretation of lineage structures. PhyloSOLID provides a 30 \npowerful solution for decoding cellular evolution in developmental and disease contexts. 31 \nMain: 32 \nSomatic mutations serve as natural recorders of cellular lineage, providing a powerful means 33 \nto reconstruct developmental and evolutionary trajectories in vivo1–3. The advent of single-cell 34 \nsequencing has brought this promise into sharper focus by enabling the resolution of mutational 35 \nhistories at the cellular level4,5. However, the technical challenges inherent to these methods—36 \nincluding significant noise and extreme data sparsity—often obscure true biological signals. In 37 \nsingle-cell whole-genome sequencing (scWGS), whole -genome amplification can introduce 38 \npervasive amplification artifacts6. In single-cell RNA sequencing (scRNA-seq), technical noise 39 \nfrom library preparation and sequencing is compounded by biological processes such as RNA 40 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 2 \nediting7. A universal challenge is allele dropout (ADO), where the signal from one allele is lost 41 \ndue to stochastic sampling, uneven coverage, or amplification bias. In scRNA -seq, this is 42 \nfurther confounded by transcriptional dropout 8, which occurs when a gene or an allele is not 43 \nexpressed in a cell, preventing the detection of its underlying genotype. These technical 44 \nchallenges lead to high error rates in somatic mutation calling and lineage tracing, especially 45 \nwhen detecting mosaic mutations in non -cancerous cells, even with the most advanced 46 \ndetection pipelines9. 47 \nDespite the high error rates in somatic mutation detection, most existing phylogenetic tree 48 \nconstruction methods for single -cell data incorporate these noisy sites without proper 49 \nidentification, filtering, or consideration during tree construction 10–13. Additionally, most 50 \ncurrent methods fail to account for critical issues including allele dropout, extreme data sparsity, 51 \nand cell -type-specific gene expression biases 14–17. Consequently, the resulting phylogenetic 52 \ntrees are often inaccurate, distorting lineage relationships and complicating the interpretation 53 \nof cellular evolutionary history (Fig. 1a -b). Given the importance of precise phylogeny 54 \nreconstruction for understanding developmental processes and disease progression, there is a 55 \npressing need for error -tolerant methods that can effectively address the noise and technical 56 \nartifacts inherent in single-cell sequencing data. 57 \nTo address this critical methodological gap, we developed PhyloSOLID, a robust phylogenetic 58 \nreconstruction algorithm designed for single -cell somatic mutation data. Unlike conventional 59 \nmethods, PhyloSOLID incorporates several key innovations  (Fig. 1c, Supplementary Fig. 1, 60 \nMethods): (1) It learns the properties of artifact sites, germline variants, and true somatic 61 \nmutations from multiple cross-validated, real-world data, excluding most germline and artifact 62 \nsites prior to phylogeny construction (Supplementary Fig. 1, Supplementary Table 1)18–23. (2) 63 \nIt calculates genotyping posterior probabilities while accounting for allele imbalance issues. (3) 64 \nIt identifies correlated clones and mutations using a graph -based method, by explicitly 65 \nmodeling the inherent noise in the data  (Supplementary Fig.  1b-e. (4) It first constructs a 66 \nbackbone tree of high-confidence clones supported by mutations  from uniformly covered 67 \nregions and then places remaining mutations onto this scaffold using a Bayesian penalty 68 \napproach (Methods)24, preventing tree structure distortion from cell -type-specific gene 69 \nexpressions. (5) It refines the final topology by correcting misplaced variants, removing likely 70 \nfalse-positive mutations and filtering out suspected doublet cells. (6) It provides an interactive 71 \nwebsite for user -customized quality control of the tree structure. Together, these features 72 \nenable PhyloSOLID to  overcome the inherent challenges of single -cell data, such as noise, 73 \nsparsity and allele dropouts. 74 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 3 \n 75 \nFig. 1: Development and benchmarking of PhyloSOLID. a-b, Single-cell phylogenetic tree 76 \nstructure can be distorted by several factors: germline variants misclassified as somatic 77 \nmutations, erroneous mutation calls, allelic/depth dropout, and the presence of doublet cells. c, 78 \nPhyloSOLID workflow.  The PhyloSOLID framework takes a list of candidate somatic 79 \nmutations and raw sequencing data as input. It then reconstructs an accurate phylogenetic tree 80 \nby explicitly accounting for the major sources of error that distort tree structure. See Methods 81 \nfor details.  d, Benchmarking phylogenetic inference with varying false -positive levels. 82 \nPhyloSOLID and other tools were evaluated on a scDNA -seq dataset containing a controlled 83 \ngradient of artifactual and germline variants (false positives). Phylogenetic accuracy was 84 \nquantified using the normalized Robinson-Foulds (RF) distance. PhISCS execution failed at a 85 \n90% false positive rate. e, Tanglegram comparison at 50% false-positive rate. Compared with 86 \nother tools, t he tree inferred by PhyloSOLID remains highly congruent with the reference 87 \ntopology, as shown by the high node correspondence. f, Resolving distinct clones from sparse 88 \ndata. A sparse simulated cell-by-mutation matrix was generated, comprising two major clones 89 \n(red and blue cells, rows). PhyloSOLID was benchmarked against other tools and correctly 90 \npartitioned the cells into two independent clones.  g, across six single-cell datasets with 91 \nestablished ground -truth tree structures, PhyloSOLID consistently achieved superior 92 \nphylogenetic accuracy.  h, Accurate phylogenies from scRNA -seq data.  Phylogenetic trees 93 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 4 \ninferred by PhyloSOLID from somatic mutations in scRNA-seq data are highly congruent with 94 \nthe ground-truth phylogeny defined by CRISPR barcode mutations. In the LUAD-2 sample, no 95 \nclone-specific mosaic mutations of C1-3 were detected. Instead, we identified a founding clone 96 \n(Cm) that gave rise to the three descendent clones. “ncells” represents the number of cells. 97 \n 98 \nTo evaluate performance under noise and data sparsity, we first used a scDNA-seq dataset18 (22 99 \ncells; 78 somatic mutations) and simulated varying noise levels by introducing artificial false-100 \npositive sites (Methods, Supplementary Table 2). PhyloSOLID consistently achieved superior 101 \naccuracy, as measured by the normalized Robinson-Foulds (RF) distance25 to the ground-truth 102 \ntree (Methods, Fig. 1d). Notably, even at a 90% false-positive rate, PhyloSOLID maintained a 103 \nnormalized RF distance below 0.1, with a representative example illustrated in Fig. 1e.  To 104 \nfurther test robustness under realistic data sparsity , we used a second simulated dataset 105 \ncomprising 50 mutations across 1,761 cells with a high dropout rate of ~90% (Supplementary 106 \nTable 3). Here, PhyloSOLID accurately reconstructed the ground-truth monoclonal phylogeny 107 \n(Fig. 1f, Supplementary Fig. 2), whereas other widely used tools (e.g., HUNTREE, CellPhy, 108 \nBCITE, ScisTree) erroneously inferred polyclonality. 109 \nAfter compiling a benchmark of eight ground truth datasets —comprising four scDNAseq 110 \nbenchmarks post whole -genome amplification 18,19, two single -cell colony benchmarks  with 111 \npublished phylogenies20,21, and two scRNAseq benchmarks with knock -in CRISPR barcode 112 \nmutations26—we further evaluated the performance of PhyloSOLID against eight state-of-the-113 \nart methods ( e.g., HUNTREE, CellPhy, ScisTree).  Across all the scDNA -seq benchmarks, 114 \nPhyloSOLID consistently and significantly outperformed other methods (Methods, Fig. 1f, 115 \nSupplementary Fig. 3).  116 \nNotably, on two scRNA-seq benchmarks comprising over a thousand cells each, PhyloSOLID 117 \ncorrectly assigned 94.03% and 96.24% of mutant cells to their original clones, as validated by 118 \nthe CRISPR barcode trees, achieving high phylogenetic accuracy despite the inherent noise 119 \nand sparsity. In contrast, other tools exhibited high clone misassignment rates, ranging from 120 \n~30% to 80% (Fig. 1g, Supplementary Fig. 4-6). Remarkably, using only somatic mutations, 121 \nPhyloSOLID achieved phylogenetic resolution comparable to CRISPR-based barcoding across 122 \nthe scRNA-seq benchmarks tested. This result demonstrates that our method successfully 123 \nunlocks the potential of somatic mutations for in vivo lineage tracing.  Furthermore, the 124 \nprojection of cell lineages onto the gene expression UMAP revealed distinct transcriptional 125 \nprofiles for each clone ( Supplementary Fig. 6), demonstrating that lineage history is coupled 126 \nwith changes in cell state and fate. 127 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 5 \n 128 \nFig. 2: Advanced visualization and user -friendly interface . a-c, A suite of tailored tree 129 \nlayouts designed to intuitively integrate and visualize multiple layers of data —such as per -130 \nclade mutation burden, cell -type annotations, mutation profiles, and highlighting specific 131 \nmutations or cells —in a clear and publication -ready format. d, A standard rectangular tree 132 \nlayout is also provided, offering flexibility to accommodate diverse user preferences. e, 133 \nInteractive quality control and data exploration. A circular phylogenetic layout sorts and 134 \ndisplays shared mutations across concentric rings, enabling direct visual comparison of input 135 \nand output genotypes for each cell. This is integrated into an interactive web interface that 136 \nallows users to perform point -and-click operations for in -depth analysis, such as inspecting 137 \nread-level evidence, flagging potential false -positive mutations, and annotating doublet cells. 138 \nAn example of an \"orphaned mutation\" (black triangle) is illustrated here. This variant shares 139 \nonly one mutant cell with its assigned mother mutation yet causes many false -negative allele 140 \nflips. This pattern suggests the lack of a true clonal mother mutation, likely resulting from a 141 \nsequencing error, a doublet cell, or an independent mutation misassigned to this lineage. 142 \n 143 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 6 \nPhylogenetic trees reconstructed from single-cell data often involve large numbers of cells and 144 \nmutations, along with additional layers of information such as genotype calls and cell -type 145 \nannotations. These trees can be highly complex, making well -designed visualization schemes 146 \nessential for meaningful interpretation. However, existing phylogenetic tools typically offer 147 \nlimited or inflexible visualization options —ill-suited to diverse scenarios such as trees with 148 \nfew mutations across many cells, or vice versa —and often require advanced coding skills for 149 \ncustomization27. Furthermore, even robust methods like PhyloSOLID are not immune to 150 \nlimitations; residual false -positive mutations or doublet cells may still be incorporated and 151 \ndistort tree topology. Manual inspection and quality assessment are therefore crucial for 152 \naccurate biological interpretation. Yet most current approaches deliver only a basic tree 153 \nstructure15,28, leaving users without adequate support for systematic quality evaluation. 154 \nTo meet the demand for advanced visualization, we developed a suite of tailored tree 155 \nvisualization schemes (Fig. 2a –c). Each is designed to present multiple data layers —such as 156 \nmutation number, cell -type annotations, and mutational profiles—in a clear, efficient, and 157 \npublication-ready format. A rectangular tree layout is also provided to accommodate different 158 \nuser preferences (Fig. 2d). Beyond these general layouts, we developed specialized 159 \nvisualizations to suit different data structures and biological questions. These include trees that 160 \nintegrate single -cell mutation burdens, shared variants, and cell -type annotations (Fig. 2a); 161 \nrepresentations optimized for large numbers of mutations across few cells, with optional 162 \nmutational signature overlay (Fig. 2b); and layouts tailored to large cell sets that emphasize a 163 \nlimited set of key driver mutations (Fig. 2c). Together, these options allow users to adapt tree 164 \nvisualization to their specific dataset and analytical goals. 165 \nFor streamlined phylogeny quality control, we introduced a circular layout in which shared 166 \nmutations are sorted and displayed across concentric rings, enabling direct comparison of input 167 \nand output genotypes per cell (Fig. 2 e). We provide an interactive web interface that enables 168 \nusers to perform a range of exploratory and quality-control operations. For example, users can 169 \nvisually inspect read-level evidence for mutations, flag potential false-positive sites, and mark 170 \nquestionable mutations  and likely doublet cells —all through an intuitive, point -and-click 171 \nenvironment. This interactive system greatly lowers the barrier to rigorous phylogenetic quality 172 \nassessment, helping ensure robust interpretation of lineage relationships. 173 \nThe PhyloSOLID source code is being prepared for public release. It will be available on 174 \nGitHub ( https://github.com/douymLab/PhyloSOLID) by March 1, 2026, or upon the 175 \nmanuscript's formal acceptance, whichever comes first. An interactive website is accessible 176 \nonline at https://phylosolid.westlake.edu.cn. 177 \nDiscussion: 178 \nThe rapid expansion of single-cell sequencing has established somatic mutation and barcode -179 \nbased lineage tracing as a central technique for decoding cellular fate, development, and disease. 180 \nHowever, the intrinsic noisiness and sparsity of single -cell data—including pervasive allele 181 \ndropout and low mutation -calling accuracy —pose a fundamental challenge to reliable 182 \nphylogenetic inference. When unaddressed, these technical artifacts can be misinterpreted as 183 \nbiological signal, leading to incorrect reconstructions of cellular relationships and, 184 \nconsequently, erroneous biological conclusions. Here, we introduced PhyloSOLID, a 185 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 7 \ncomputational framework designed to be inherently robust to these pervasive data 186 \nimperfections. 187 \nThe core strength of PhyloSOLID lies in its progressive scaffolding strategy, which prioritizes 188 \nhigh-confidence phylogenetic signals before integrating more ambiguous data. By first 189 \nconstructing a robust backbone from reliable mutations and then iteratively refining it with a 190 \nBayesian model that penalizes phylogenetic inconsistencies, our method effectively 191 \ndisentangles true evolutionary history from technical noise. Extensive benchmarking on 192 \ndatasets with known ground truth demonstrated that PhyloSOLID consistently outperforms 193 \nexisting methods in reconstruction accuracy across both scRNA -seq and scDNA -seq 194 \nmodalities. Its performance on large -scale scRNA-seq datasets was particularly noteworthy, 195 \nwhere phylogenetic trees built solely from somatic mutations recapitulated the clonal structure 196 \ndefined by independent CRISPR barcoding with remarkably high accuracy . This result 197 \nunderscores that somatic mutations, when analyzed with a sufficiently robust phylogenetic 198 \nframework, can provide resolution for lineage tracing that rivals or even surpasses engineered 199 \nbarcoding approaches.  Beyond constructing phylogenies from somatic mutations, 200 \nPhyloSOLID is a flexible toolkit that can also leverage barcode mutations for phylogenetic 201 \ninference. 202 \nLooking forward, the need for reliable single  cell phylogenetics will become ever more 203 \npressing. PhyloSOLID meets this need by providing a statistically rigorous, error -aware 204 \nframework for lineage reconstruction. To maximize its accessibility and utility for the broader 205 \nresearch community, we have complemented the core algorithm with a user -friendly web 206 \ninterface for quality assessment, artifact removal, and lineage interpretation. We anticipate that 207 \nPhyloSOLID will become an essential tool, empowering researchers to extract trustworthy 208 \nevolutionary insights from the complex and noisy landscape of single -cell data and thereby 209 \nensuring that the burgeoning field of cellular lineage tracing is built upon a foundation of 210 \ncomputational rigor. 211 \nData availability 212 \nAll scWGS and scRNA-seq data used in this study are publicly available under the following 213 \naccession codes: dbGaP: phs001485.v3.p1, SRA: SRA053195, NDA: NDAR#2330, EGA: 214 \nEGAD00001007032, and GEO: GSE144239, GSE234814 and GSE161363. 215 \nCode availability: 216 \nPhyloSOLID is implemented in Python and R and is licensed under the MIT License. The 217 \nsource code, documentation and examples are available on GitHub at 218 \nhttps://github.com/douymLab/PhyloSOLID. The PhyloSOLID source code is being prepared 219 \nfor public release. It will be available on GitHub by March 1, 2026, or upon the manuscript's 220 \nformal acceptance, whichever comes first. 221 \nAcknowledgements: This work was supported by the Westlake Laboratory of Life Sciences 222 \nand Biomedicine (Hangzhou 310024, Zhejiang, China), under the grant \"Key R&D Program 223 \nof Zhejiang\" (2024SSYS0032) as well as the National Natural Science Foundation of China 224 \n(32270682) to Y.D. We thank the High-Performance Computing Center for technical support. 225 \nWe gratefully acknowledge the creators and submitters of the datasets used in this study. Data 226 \nused in this study under the dbGaP accession phs001485.v3.p1 was generated by Drs. 227 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 8 \nChristopher A. Walsh and Peter J. Park with funding from NINDS grant R01NS032457. The 228 \npublic datasets were obtained from dbGaP at http://www.ncbi.nlm.nih.gov/gap, EGA at 229 \nhttps://ega-archive.org/, NDA at https://nda.nih.gov/ and GEO at 230 \nhttps://www.ncbi.nlm.nih.gov/geo/.  231 \nAuthor contribution: Y.Dou conceived and supervised the project and secured the funding. 232 \nN.L. supervised the website construction. Q.Y. and Y.L. both made significant contributions 233 \nto the project, with Q.Y. developed the software  and accomplished bioinformatics analysis, 234 \nand Y.L. constructed the website with help from J. Y. X.W., Y.X, Z.Y., J.L, Y.Z., J.L., M.Y., 235 \nY.Du and H.L.  assisted in single  cell genotyping , variant detection, and Y. X.  assisted in 236 \nvisualization of phylogenies . Y.Dou and Q.Y. wrote the manuscript. All authors carefully 237 \nreviewed and approved the final manuscript. 238 \nInclusion & Ethics statement : This study was conducted in accordance with ethical 239 \nguidelines and principles. It did not involve human participant or animal subjects. There are no 240 \nsequencing data generated in this study. The authors ensure that the study adheres to the highest 241 \nethical standards in research, data generation, and usage. 242 \nCompeting interests: The authors declare no competing interests. 243 \n  244 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 9 \nReferences 245 \n1. Bizzotto, S., and Walsh, C.A. (2022). Genetic mosaicism in the human brain: from lineage 246 \ntracing to neuropsychiatric disorders. Nat Rev Neurosci, 1–12. https://doi.org/10.1038/s41583-247 \n022-00572-x. 248 \n2. Fasching, L., Jang, Y., Tomasi, S., Schreiner, J., Tomasini, L., Brady, M.V., Bae, T., Sarangi, 249 \nV., Vasmatzis, N., Wang, Y., et al. (2021). Early developmental asymmetries in cell lineage 250 \ntrees in living individuals. Science 371, 1245–1248. https://doi.org/10.1126/science.abe0981. 251 \n3. Tian, C., Zhang, Y., Tong, Y., Kock, K.H., Sim, D.Y., Liu, F., Dong, J., Jing, Z., Wang, W., 252 \nGao, J., et al. (2024). Single-cell RNA sequencing of peripheral blood links cell-type-specific 253 \nregulation of splicing to autoimmune and inflammatory diseases. Nat. Genet., 1 –14. 254 \nhttps://doi.org/10.1038/s41588-024-02019-8. 255 \n4. Gonzalez-Pena, V., Natarajan, S., Xia, Y., Klein, D., Carter, R., Pang, Y., Shaner, B., Annu, 256 \nK., Putnam, D., Chen, W., et al. (2021). Accurate genomic variant detection in single cells with 257 \nprimary template -directed amplification. Proc National Acad Sci 118, e2024176118. 258 \nhttps://doi.org/10.1073/pnas.2024176118. 259 \n5. Tanay, A., and Regev, A. (2017). Scaling single -cell genomics from phenomenology to 260 \nmechanism. Nature 541, 331–338. https://doi.org/10.1038/nature21350. 261 \n6. Estévez-Gómez, N., Prieto, T., Tomás, L., Alvariño, P., Guillaumet -Adkins, A., Heyn, H., 262 \nPrado-López, S., and Posada, D. (2025). Differential performance of strategies for single -cell 263 \nwhole-genome amplification. Cell Rep. methods 5, 101025. 264 \nhttps://doi.org/10.1016/j.crmeth.2025.101025. 265 \n7. Peng, Z., Cheng, Y., Tan, B.C. -M., Kang, L., Tian, Z., Zhu, Y., Zhang, W., Liang, Y., Hu, 266 \nX., Tan, X., et al. (2012). Comprehensive analysis of RNA -Seq data reveals extensive RNA 267 \nediting in a human transcriptome. Nat. Biotechnol. 30, 253 –260. 268 \nhttps://doi.org/10.1038/nbt.2122. 269 \n8. Lähnemann, D., Köster, J., Szczurek, E., McCarthy, D.J., Hicks, S.C., Robinson, M.D., 270 \nVallejos, C.A., Campbell, K.R., Beerenwinkel, N., Mahfouz, A., et al. (2020). Eleven grand 271 \nchallenges in single -cell data science. Genome Biol. 21, 31. https://doi.org/10.1186/s13059-272 \n020-1926-6. 273 \n9. Dou, Y., Gold, H.D., Luquette, L.J., and Park, P.J. (2018). Detecting Somatic Mutations in 274 \nNormal Cells. Trends Genet 34, 545–557. https://doi.org/10.1016/j.tig.2018.04.003. 275 \n10. Kızılkale, C., Mehrabadi, F.R., Azer, E.S., Pérez-Guijarro, E., Marie, K.L., Lee, M.P., Day, 276 \nC.-P., Merlino, G., Ergün, F., Buluç, A., et al. (2022). Fast intratumor heterogeneity inference 277 \nfrom single-cell sequencing data. Nat Comput Sci, 1 –7. https://doi.org/10.1038/s43588-022-278 \n00298-x. 279 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 10 \n11. Wu, Y. (2019). Accurate and efficient cell lineage tree inference from noisy single cell data: 280 \nthe maximum likelihood perfect phylogeny approach. Bioinformatics. 281 \nhttps://doi.org/10.1093/bioinformatics/btz676. 282 \n12. Zhang, H., Zhang, Y., Gao, T., and Wu, Y. (2025). ScisTree2 enables large-scale inference 283 \nof cell lineage trees and genotype calling using efficient local search. Genome Res. 35, 2781–284 \n2791. https://doi.org/10.1101/gr.280542.125. 285 \n13. Hoang, D.T., Vinh, L.S., Flouri, T., Stamatakis, A., Haeseler, A. von, and Minh, B.Q. 286 \n(2018). MPBoot: fast phylogenetic maximum parsimony tree inference and bootstrap 287 \napproximation. Bmc Evol Biol 18, 11. https://doi.org/10.1186/s12862-018-1131-3. 288 \n14. Malikic, S., Jahn, K., Kuipers, J., Sahinalp, S.C., and Beerenwinkel, N. (2019). Integrative 289 \ninference of subclonal tumour evolution from single -cell and bulk sequencing data. Nat 290 \nCommun 10, 2750. https://doi.org/10.1038/s41467-019-10737-5. 291 \n15. Kozlov, A., Alves, J.M., Stamatakis, A., and Posada, D. (2022). CellPhy: accurate and fast 292 \nprobabilistic inference of single-cell phylogenies from scDNA-seq data. Genome Biol 23, 37. 293 \nhttps://doi.org/10.1186/s13059-021-02583-w. 294 \n16. Zafar, H., Tzen, A., Navin, N., Chen, K., and Nakhleh, L. (2017). SiFit: inferring tumor 295 \ntrees from single -cell sequencing data under finite -sites models. Genome Biol 18, 178. 296 \nhttps://doi.org/10.1186/s13059-017-1311-2. 297 \n17. Malikic, S., Mehrabadi, F.R., Ciccolella, S., Rahman, Md.K., Ricketts, C., Haghshenas, E., 298 \nSeidman, D., Hach, F., Hajirasouliha, I., and Sahinalp, S.C. (2019). PhISCS: a combinatorial 299 \napproach for subperfect tumor phylogeny reconstruction via integrative use of single -cell and 300 \nbulk sequencing data. Genome Res 29, 1860–1877. https://doi.org/10.1101/gr.234435.118. 301 \n18. Lodato, M.A., Rodin, R.E., Bohrson, C.L., Coulter, M.E., Barton, A.R., Kwon, M., 302 \nSherman, M.A., Vitzthum, C.M., Luquette, L.J., Yandava, C.N., et al. (2018). Aging and 303 \nneurodegeneration are associated with increased mutations in single human neurons. Science 304 \n359, 555–559. https://doi.org/10.1126/science.aao4426. 305 \n19. Wang, Y., Waters, J., Leung, M.L., Unruh, A., Roh, W., Shi, X., Chen, K., Scheet, P., 306 \nVattathil, S., Liang, H., et al. (2014). Clonal evolution in breast cancer revealed by single 307 \nnucleus genome sequencing. Nature 512, 155–160. https://doi.org/10.1038/nature13600. 308 \n20. Bae, T., Tomasini, L., Mariani, J., Zhou, B., Roychowdhury, T., Franjic, D., Pletikos, M., 309 \nPattni, R., Chen, B.-J., Venturini, E., et al. (2018). Different mutational rates and mechanisms 310 \nin human cells at pregastrulation and neurogenesis. Science 359, 550 –555. 311 \nhttps://doi.org/10.1126/science.aan8690. 312 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 11 \n21. Park, S., Mali, N.M., Kim, R., Choi, J.-W., Lee, J., Lim, J., Park, J.M., Park, J.W., Kim, D., 313 \nKim, T., et al. (2021). Clonal dynamics in early human embryogenesis inferred from somatic 314 \nmutation. Nature 597, 393–397. https://doi.org/10.1038/s41586-021-03786-8. 315 \n22. Ji, A.L., Rubin, A.J., Thrane, K., Jiang, S., Reynolds, D.L., Meyers, R.M., Guo, M.G., 316 \nGeorge, B.M., Mollbrink, A., Bergenstråhle, J., et al. (2020). Multimodal Analysis of 317 \nComposition and Spatial Architecture in Human Squamous Cell Carcinoma. Cell 182, 497-318 \n514.e22. https://doi.org/10.1016/j.cell.2020.05.039. 319 \n23. Kumar, T., Nee, K., Wei, R., He, S., Nguyen, Q.H., Bai, S., Blake, K., Pein, M., Gong, Y., 320 \nSei, E., et al. (2023). A spatially resolved single -cell genomic atlas of the adult human breast. 321 \nNature, 1–11. https://doi.org/10.1038/s41586-023-06252-9. 322 \n24. Wu, X. MosaiSC: https://github.com/douymLab/mosaicSC. 323 \n25. Robinson, D.F., and Foulds, L.R. (1981). Comparison of phylogenetic trees. Math. Biosci. 324 \n53, 131–147. https://doi.org/10.1016/0025-5564(81)90043-2. 325 \n26. Quinn, J.J., Jones, M.G., Okimoto, R.A., Nanjo, S., Chan, M.M., Yosef, N., Bivona, T.G., 326 \nand Weissman, J.S. (2021). Single -cell lineages reveal the rates, routes, and drivers of 327 \nmetastasis in cancer xenografts. Science 371. https://doi.org/10.1126/science.abc1944. 328 \n27. Huerta-Cepas, J., Serra, F., and Bork, P. (2016). ETE 3: Reconstruction, Analysis, and 329 \nVisualization of Phylogenomic Data. Mol. Biol. Evol. 33, 1635 –1638. 330 \nhttps://doi.org/10.1093/molbev/msw046. 331 \n28. Ross, E.M., and Markowetz, F. (2016). OncoNEM: inferring tumor evolution from single-332 \ncell sequencing data. Genome Biol 17, 69. https://doi.org/10.1186/s13059-016-0929-9. 333 \n29. Dou, J., Tan, Y., Kock, K.H., Wang, J., Cheng, X., Tan, L.M., Han, K.Y., Hon, C.-C., Park, 334 \nW.-Y., Shin, J.W., et al. (2023). Single-nucleotide variant calling in single-cell sequencing data 335 \nwith Monopogen. Nat. Biotechnol., 1–10. https://doi.org/10.1038/s41587-023-01873-x. 336 \n30. Muyas, F., Sauer, C.M., Valle-Inclán, J.E., Li, R., Rahbari, R., Mitchell, T.J., Hormoz, S., 337 \nand Cortés -Ciriano, I. (2023). De novo detection of somatic mutations in high -throughput 338 \nsingle-cell profiling data sets. Nat. Biotechnol., 1 –10. https://doi.org/10.1038/s41587-023-339 \n01863-z. 340 \n31. STAligner enables the integration and alignment of multiple spatial transcriptomics 341 \ndatasets (2023). Nat. Comput. Sci. 3, 831–832. https://doi.org/10.1038/s43588-023-00543-x. 342 \n32. Yang, Z., Yao, M., Yang, Q., Du, Y., Lu, J., and Dou, Y. Detection of Mosaic Point 343 \nMutations Directly from Spatial Transcriptomics Enables in -vivo Spatiotemporal Lineage 344 \nTracing in Tumor and Normal Tissues. 345 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 12 \n33. Galili, T. (2015). dendextend: an R package for visualizing, adjusting and comparing trees 346 \nof hierarchical clustering. Bioinformatics 31, 3718 –3720. 347 \nhttps://doi.org/10.1093/bioinformatics/btv428. 348 \n34. Schliep, K.P. (2010). phangorn: phylogenetic analysis in R. Bioinformatics 27, 592–593. 349 \nhttps://doi.org/10.1093/bioinformatics/btq706. 350 \n  351 \n  352 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 13 \nMethods 353 \nThe PhyloSOLID framework 354 \n1. Pre-phylogeny genotyping and variant filtering 355 \nWe collected eight single-cell benchmarks, comprising six scDNA-seq benchmarks with 356 \northogonally validated ground-truth mutations (Supplementary Table 4) and two scRNA-seq 357 \nbenchmarks with ground-truth CRISPR barcoding trees (Supplementary Table 5-6), from 358 \npublic sources. Mutations were genotyped using MosaicSC24, a Bayesian approach that our lab 359 \nhave developed, accounting for single -cell-specific challenges such as allelic imbalance, 360 \nsequencing errors, and population allele frequencies. Following genotyping, we applied a 361 \nlogistic regression classifier trained with an \" leave-one-donor-out\" approach to filter artifacts 362 \nand germline variants. This model distinguished true mosaic mutations based on 10 (for 363 \nscDNA-seq) or 14 (for scRNA -seq) features derived from raw sequencing reads, including 364 \nmismatch frequency, base quality, read alignment position, etc. (Supplementary Table 1). We 365 \nfurther excluded mutations with a low mutant-allele fraction (VAF < 1%) or a population allele 366 \nfrequency >0.1% in gnomAD. For scRNA-seq-derived mutations, we further excluded thos e 367 \nmapping to known sites in three classical RNA editing databases29–31 or residing within allele-368 \nspecific expressed (ASE) genes 32, those with significantly imbalanced base‑substitution 369 \nprofiles, and those with low UMI consistency.  Of note, when BAM fil es are unavailable, 370 \nPhyloSOLID offers a practical alternative by constructing phylogenies directly from a cell-by-371 \nmutation matrix. 372 \n2. Leveraging clonal architecture to further filter germline variants 373 \nTo distinguish true somatic mutations from germline variants with allele dropout in single-cell 374 \nDNA sequencing data, we leverage the fundamental principle that genuine somatic mutations 375 \nexhibit structured patterns of co -occurrence and mutual exclusivity reflective of underlying 376 \nclonal architecture. Specifically, mutations belonging to the same subclone tend to co -occur 377 \nwithin the same group of cells, while those originating from distinct, mutually exclusive 378 \nsubclones are distributed across different cell populations. In contrast, germline variants 379 \naffected by stochastic allele dropout demonstrate no systematic association with clonal groups; 380 \ninstead, they appear randomly distributed across cells and show no consistent correlation with 381 \nother mutations, resulting in a nonspecific and biologically incoherent pattern. 382 \n2.1) Identification of mutations with co-occurrence patterns 383 \nLet 𝑛 and 𝑚 be the number of cells and mutations retained after initial filtering, respectively. 384 \nThe final matrix used for analysis is 𝐼 = {0, 1, 𝑁𝐴}!×#, where 𝑁𝐴 values are excluded from 385 \ncalculations. This matrix is derived from three input matrices: 386 \n• 𝑃$,& ∈ [0,1]!×#: the matrix of somatic posterior probabilities, where 𝑃$,& represents the 387 \nprobability that the somatic mutation  𝑗  is present in cell 𝑖 . Somatic mutation 388 \nlikelihoods were calculated based on a beta-binomial sampling model. 389 \n• 𝑀$,& ∈ [0,1]!×#: the mutant allele frequency matrix, where 𝑀$,& represents the mutant 390 \nallele frequency of somatic mutation 𝑗 in cell 𝑖. 391 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 14 \n• 𝐶$,& ∈ ℤ'(\n!×# : the sequencing co verage matrix, where 𝐶$,&  represents the number of 392 \nreads (coverage depth) at the locus of mutation 𝑗 in cell 𝑖. 393 \nFor each mutation 𝑗 in each cell 𝑖, an entry in the binary matrix is defined as: 394 \n𝐼$,& = 5\n1\t\t\t\t\t\t\t𝑖𝑓\t𝑀$,& > 0\t ⋀ 𝑃$,& > 0.5\t\t\t\t\t\t\t\t\t\t\n0\t\t\t\t\t\t𝑖𝑓\t𝑀$,& = 0\t ⋀ 𝑃$,& ≤ 0.5\t ⋀ 𝐶$,& ≥ 1\t\n𝑁𝐴\t\t\t\t\t𝑖𝑓\t𝐶$,& = 0\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\n(1)  395 \nFor each mutation 𝑗, the set of cells in which it was detected was defined as:  396 \n𝒮(𝑗) = A𝑖 ∈ {1,2, … , 𝑛}D\t𝐼$,& = 1E (2) 397 \nMutation processing order: To ensure robust phylogenetic reconstruction, mutations are 398 \nprocessed in descending order of their cellular prevalence |  𝒮(𝑗)|, with higher -confidence 399 \nmutations processed first. This sorting strategy prioritizes mutations with stronger phylogenetic 400 \nsignals for initial tree construction. 401 \nFor any pair of mutations 𝑗)  and 𝑗*, the following contingency counts were computed: 402 \n𝑁)) (𝑗) , 𝑗*) = F 𝟏(𝐼$,&! = 1\t⋀\t𝐼$,&\" = 1)\n!\n$+)\n 403 \n𝑁)( (𝑗) , 𝑗*) = F 𝟏(𝐼$,&! = 1\t⋀\t𝐼$,&\" = 0)\n!\n$+)\n 404 \n𝑁() (𝑗) , 𝑗*) = F 𝟏(𝐼$,&! = 0\t⋀\t𝐼$,&\" = 1)\n!\n$+)\n 405 \n𝑁(( (𝑗) , 𝑗*) = F 𝟏H𝐼$,&! = 0\t⋀\t𝐼$,&\" = 0I\n!\n$+)\n(3) 406 \nwhere 𝟏(∙) was the indicator function, returning 1 if its argument is true and 0 otherwise. 407 \nWe applied the Jaccard index to measure the similarity between the sets of cells containing 408 \neach mutation: 409 \n𝐽(𝑗) , 𝑗*) = 𝑁)) (𝑗) , 𝑗*)\n𝑁)) (𝑗) , 𝑗*) + 𝑁)( (𝑗) , 𝑗*) + 𝑁() (𝑗) , 𝑗*) (4) 410 \nJaccard index was selected for scRNAseq data because this metric ignores double -negative 411 \ncells ( 𝑁(( ), which is suitable for sparse scRNAseq data. However, this measure alone is 412 \ninsufficient for identifying parent -child clonal relationships, as th e Jaccard index between a 413 \nparent and its subclone may not necessarily be high. To specifically capture such hierarchical 414 \nrelationships, we introduced an additional metric defined as follows: 415 \n𝑓(𝑗) , 𝑗*) = 𝑁)) (𝑗) , 𝑗*)\n|𝒮(𝑗) )| (5) 416 \n𝑓(𝑗*, 𝑗) ) = 𝑁)) (𝑗) , 𝑗*)\n|𝒮(𝑗*)| (6) 417 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 15 \nThe Two mutations above, 𝑗)  and 𝑗*, were considered correlated if they satisfied either of the 418 \nfollowing criteria:  419 \n• 𝑁)) (𝑗) , 𝑗*) > 0 ∧ 𝐽(𝑗) , 𝑗*) ≥ 0.08, 420 \n• 𝑁)) (𝑗) , 𝑗*) > 0 ∧ 0 < 𝐽(𝑗) , 𝑗*) < 0.08\t ∧ \tmax \t(𝑓(𝑗) , 𝑗*), 𝑓(𝑗*, 𝑗) )) ≥ 0.5. 421 \n2.2) Exclusion of likely germline variants 422 \nTo systematically identify and exclude germline variants misclassified as somatic mutations, 423 \nwe employed an exhaustive search approach that ev aluates the patterns of co -occurrence and 424 \nmutual exclusivity across all mutations. 425 \nA “founder mutation” is defined as a genetic alteration that initiates a clonal expansion and is 426 \nsubsequently present in all cells of that clone and all its subclones. To systematically identify 427 \nsuch mutations and their associated clonal groups, we proceed as follows: 428 \nFor each mutation 𝑟 ∈ {1, … , 𝑚} with a mutant cell fraction (MCF) >5%, we consider it as a 429 \ncandidate founder mutation. For each such candidate 𝑟, we identify the set of mutations 𝒥, 430 \nthat are significantly correlated with 𝑟, as described in Section 2.1. 431 \nThe underlying rationale is that if 𝑟 is a true founder mutation, the set 𝒥, likely represents 432 \nsubsequent mutations that occurred within the clone founded by 𝑟, and thus should exhibit a 433 \nstrong pattern of co-occurrence with 𝑟. 434 \n2.2.1) Genotype imputation of the “founder mutation” 435 \nIn theory, these correlated clones descend from a single founder clone. Consequently, all 436 \ndaughter clones are expected to inherit the founder's mutations. Assuming 𝑟 as the founder 437 \nmutation of these correlated clones, we created a combined muta tion set that includes both 𝑟 438 \nand all mutations significantly correlated with it: 439 \n𝒥,- = {𝑟} ∪ 𝒥, (7) 440 \nFor each cell, we then quantified 𝑞$, the number of mutations from the combined set 𝒥,- that 441 \nare present in that cell: 442 \n𝑞$ = F 𝐼$,&\n&∈𝒥#$\n(8) 443 \nWhere: 444 \n• 𝑞$: The number of mutations from the set 𝒥,- that are present in cell 𝑖. 445 \nImpute founder status: A cell is classified as carrying the founder mutation 𝑟 if it possesses at 446 \nleast two mutations from the set  𝒥,-. This threshold helps mitigate false negatives caused by 447 \nallele dropout or sequencing errors affecting a single mutation. 448 \n𝒰, = A𝑖 ∈ {1, … , 𝑛}DH𝐼$,& = 1I \t ∨ \t (𝑞$ ≥ 2)E (9) 449 \nWhere: 450 \n• 𝒰,: The set of cells inferred to belong to the clone founded by mutation 𝑟. 451 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 16 \nLet 𝒰,____ denote the complement of 𝒰, with respect to covered cells for any mutation  𝑗, i.e., 452 \n𝒰,____ = {𝑖 ∉ 𝒰,, 𝐶$,& ≥ 1}. 453 \nFor each mutation 𝑗 ∈ 𝒥,\n-, we calculated the total number of mutant cells outside the infe rred 454 \nfounder clone: 455 \n𝑁&,,\n0123$45 = DA𝑖 ∈ 𝒞&\\𝒰,ED (10) 456 \nWhere: 457 \n• 𝒞6 denotes the set of j-mutant cells. 458 \nSince all cells carrying mutation 𝑗 ∈ 𝒥,\n- are descendants of the founder clone, no mutant cells 459 \nshould exist outside of it. We then defined a false -positive penalty score for mutation 𝑗 by 460 \nnormalizing 𝑁&,,\n0123$45  by the total number of 𝑗-mutant cells: 461 \n𝑆&,,\n78 = 𝑁&,,\n0123$45\nD𝒞&D\n(11) 462 \nIf r and 𝑗 are both real mosaic mutations, the n 𝑆&,,\n78 will be low. Conversely, a high score 463 \nsuggests otherwise. 464 \n2.2.2) Enumeration over all candidate founder mutations and identification of likely 465 \ngermline variations 466 \nThe above procedure (Sections 2.2.1 and 2.2.2) was repeated for each mutation 𝑟 ∈ {1, … , 𝑚}. 467 \nFor each mutation r, we calculate a summary leakage score: (cv\\std\\mean, fig) 468 \n𝑆,78eeeeee⃗ = 1\n|𝒥,\n-| F 𝑆&,,\n78\n&∈𝒥#$\n(12) 469 \nThis score represents the average phylogenetic inconsistency of all mutations in the clonal set 470 \n𝒥,- founder by 𝑟. 471 \nTo identify germline variants, all candidate mutations 𝑟 are sorted in descending order of 𝑆,78. 472 \nGermline variants form a distinct group at the top of this sorted list. We identify the cutoff 473 \npoint by calculating differences between consecutive sorted scores: 474 \n∆9= 𝑠9 − 𝑠9-) (13) 475 \nwhere 𝑠9 represents the 𝑙-th highest score in the sorted sequence. Mutations preceding a large 476 \ndrop in scores (∆9> 𝜇∆ + 2𝜎∆, where 𝜇∆ and 𝜎∆ are the mean and standar d deviation of all 477 \ndifferences) are classified as germline. 478 \nThese identified germline variants are designated as ℳ;52, placed at the root of the clonal tree, 479 \nand excluded from subsequent lineage reconstruction. This approach effectively separated 480 \ngermline variants from real mosaic mutations. 481 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 17 \n3. Scaffold tree construction using uniformly covered mutations 482 \nFollowing the exclusion of germline mutations, we proceed to identify high -confidence 483 \nsomatic mutations located in genomic regions with ubiquitous expression and uniform 484 \ncoverage. These mutations serve as backbone features for initial phylogenetic reconstruction. 485 \n3.1) Initial filtration based on genotype posterior and mutant reads count 486 \nWe begin with the matrices defined in Section 2.1 and apply additional quality filters to ensure 487 \nreliable detection of ubiquitously expressed mutations. 488 \nCell Filtering: 489 \nThe set of retained cells is defined as those containing sufficient mutation evidence: 490 \n𝒞 = A𝑖 ∈ {1,2, … , 𝑛}DH∃𝑗\t𝑀$,& > 0I ∧ \t H∃𝑗\t𝑃$,& > 0.5IE (14) 491 \nwhere 𝑀 and 𝑃 are the mutant allele frequency and somatic posterior probability matrices. 492 \nMutation filtration: 493 \nFor each mutation 𝑗, we compute its average mutant allele frequency (MAF) across cells where 494 \nit is detected. The set of cells where mutation 𝑗 is detected is defined as: 495 \n𝒞& = A𝑖 ∈ 𝒞D𝑀$,& > 0\t ∧ \t 𝑃$,& > 0.5\tE (15) 496 \nThe average MAF for mutation 𝑗 is: 497 \n𝑀<___ = 1\nD𝒞& D F 𝑀$,&\n$∈𝒞%\n(16) 498 \nThe set of retained mutations is: 499 \nℳ = o𝑗 ∈ {1,2, … , 𝑚}p q𝑚𝑎𝑥\n$\n𝑀$,& ≥ 0.3t \t ∧ \t H𝑀<___ ≥ 0.1Iu (17) 500 \nThe filtered matrices are then: 501 \n𝑃ℳ,𝒞 = 𝑃[ℳ, 𝒞] = A𝑃$,&D𝑖 ∈ 𝒞, 𝑗 ∈ ℳE (18) 502 \n𝑀ℳ,𝒞 = 𝑀[ℳ, 𝒞] = A𝑀$,& D𝑖 ∈ 𝒞, 𝑗 ∈ ℳE (19) 503 \n 504 \n3.2) Selection of uniformly covered mutations 505 \nTo identify high -confidence mutations within ubiquitously expr essed genomic regions, we 506 \napplied two complementary coverage-based filters. 507 \nFor mutation 𝑗 in cell type 𝑡, the proportion of cells lacking sequencing coverage is denoted 508 \nas: 509 \n𝑝?@\n(&,2) = DA𝑖 ∈ 𝒮2|𝐶$,& = 0ED\n𝑁2\n(20) 510 \nWhere: 511 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 18 \n• 𝒮2 denotes the set of all cells that belong to cell type 𝑡; 512 \n• 𝐶$,& denotes the sequencing coverage for mutation 𝑗 in cell 𝑖; 513 \n• 𝑖 denotes a cell 𝑖; 514 \n• 𝑁2 denotes the total number of cells of cell type 𝑡. 515 \nFor each mutation 𝑗, we define 𝑁& as the number of cell types in which a sufficient proportion 516 \nof cells exhibit sequencing coverage: 517 \n𝑁& = F 1 x𝑝?@\n(&,2) ≤ 95%z\n2∈𝒮&\n(21) 518 \nWhere: 519 \n• 𝑝?@\n(&) =\nD𝑖E𝐶$,& = 0F\n? : NA proportion for mutation 𝑗 across all cells. 520 \n• 𝑝?@\n(&,2): NA proportion for mutation 𝑗 in cell type 𝑡; 521 \nBased on measurements from real-world single cell sequencing data, mutations are retained for 522 \nsubsequent analysis as candidate scaffold mutations if they meet:  523 \n• Present in >10% cells across all cells when th ere’s only one single dominant cell type, or 524 \nno cell type annotation information is available. 525 \n• Present in ≥2 cell types (𝑁& ≥ 2). 526 \nFor each mutation 𝑗, we first required that the median sequencing depth across all cells is at 527 \nleast 1x. 528 \nWe then calculate d the coefficient of variation (CV) to assess coverage uniformity. For 529 \nmutation 𝑗 in cell 𝑖, we defined a binarized coverage indicator: 530 \n𝐵$,& = |1, 𝑖𝑓\t𝐶$,& ≥ 1\n0, 𝑖𝑓\t𝐶$,& < 1 (22) 531 \nWhere: 532 \n• 𝐶$,& indicates the read coverage in cell 𝑖 at the genome location of mutation 𝑗; 533 \n𝐵$,& is the binary indicator of whether mutation 𝑗 in cell 𝑖 has sufficient read coverage. 534 \nFor each mutation 𝑗, the mean read coverage along with the standard deviation, and coefficient 535 \nof variation are computed as follows: 536 \n𝐶𝑉& = 𝜎&\n𝐵<~ =\n1\n𝑁 ∑ H𝐵$,& − 𝐵<~ I\n*\n$\n1\n𝑁 ∑ 𝐵$,&$\n(23) 537 \nWhere: 538 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 19 \n• 𝐶$,& denotes the sequencing coverage for mutation 𝑗 in cell 𝑖; 539 \n• 𝑁 is the total number of cells. 540 \n• 𝐵<~  is the mean of 𝐵$,& across all cells for mutation 𝑗; 541 \n• 𝜎& is the standard deviation of 𝐵$,& across all cells for mutation 𝑗; 542 \n• 𝐶𝑉& is the coefficient of variation for the coverage indicator of mutation 𝑗. 543 \nBased on the measurements in real-world data, only mutations with 𝐶𝑉& < 6 were retained as 544 \ncandidate scaffold mutations. 545 \n3.3) Noise-robust quantification of pairwise mutation correlations 546 \nTheoretically, two somatic mutations in an individual can only have one of four phylogenetic 547 \nrelationships: A is ancestral to B, B is ancestral to A, they co-occur on the same node, or they 548 \narise on independent branches. Consequently, if two mutations share an ancestral relationship 549 \nor occur on the same node, they are supposed to co-exist in at least a subset of cells. In contrast, 550 \nmutations on independent branches are not expected to to -exist in any single cell. Therefore, 551 \nthe co-occurrence patterns of mutations across single cells can be used to identify correlated 552 \nmutations. 553 \nHowever, single cell sequencing data are inherently noisy. Sequencing errors and the presence 554 \nof doublet cells can obscure the true phylogenetic relationships between mutations. To establish 555 \nrobust co-occurrence patterns while mitigating stochastic effects, we developed a consensus 556 \napproach based on ensemble clone identification through randomized runs.  557 \nBeginning from each scaffold mutation 𝑗 ∈ 𝑀G, we sequentially evaluated its correlation with 558 \nevery other scaffold mutation using the method described in Section 2.1. Mutations found to 559 \nbe correlated with  𝑗 were assigned to the same clonal population with  𝑗. This process was 560 \nrepeated iteratively until all mutations were assigned to some clone. To ensure robustness, the 561 \nscaffold mutation order was randomly shuffled, and the entire procedure was repeated 100 562 \ntimes.  563 \nThe occurrence frequency of each distinct clone 𝐶 across 100 runs were then calculated: 564 \n𝑓& (𝐶) = 1\n100 F 𝟏 q𝐶 ∈ 𝐶,1!\n(&) t\n)((\n$25,+)\n(24) 565 \nOnly clones with 𝑓&(𝐶) > 0.1 were retained, denote as 𝐶&\n∗. The weight for each clone 𝐶 to 566 \nexist was then computed as: 567 \nΩ(𝐶) = F 𝑓& (𝐶) ∗ 𝟏H𝐶 ∈ 𝐶&\n∗I\n&∈ℳ'\n(25) 568 \nFor each clone with |𝐶| > 2, we extracted all H|J|\n* I unordered mutations pairs. Each pair 569 \n(𝑗, 𝑗K) inherited weight Ω(𝐶) from its parent clone. The consensus weight for each distinct 570 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 20 \nmutation pair was obtained by summing contributions from a ll clones containing both 571 \nmutations: 572 \n𝜔(𝑗, 𝑗K) = F Ω(𝐶)\nJ∈⋃ J%\n∗%\n&,&)∈J\n(26)\n 573 \nFor singleton clones where |𝐶| = 1, although they do not contribute to edge weights, their 574 \nweights Ω(𝐶) are retained and provide valuable information about  mutation independence. 575 \nMutations with high singleton weights typically exhibit weak correlations with other mutations 576 \nand are more likely to form distinct clonal groups in subsequent clustering analysis. 577 \nWe constructed a weighted graph 𝐺M0!35!313 = (𝒱, ℰ, 𝒲) where: 578 \n• 𝒱 = ℳG: vertices represent mutations. 579 \n• ℰ = {(𝑗, 𝑗K)|𝜔(𝑗, 𝑗K) > 0}: edges connect co-occurring mutations.  580 \n• 𝒲(𝑒) = 𝜔(𝑗, 𝑗K): edge weights reflect co-occurrence strength. 581 \nThis graph serves as the input for subsequent community detection. In this way, the relationship 582 \nbetween each mutation pair was robustly quantified, making the approach resistant to stochastic 583 \nnoise. 584 \n3.4) Partitioning of scaffold mutations into clonal groups via weighted graph clustering 585 \nThe weighted consensus correlation graph 𝐺M0!35!313  was partitioned into maximal mutation 586 \ngroups using the Leiden community detection algorithm. This approach identifies 587 \nphylogenetically coherent groups of mutations that exhibit strong internal correlations and 588 \nminimal external correlations, with each group defining the mutational signature of a distinct 589 \nclonal population. 590 \nThe algorithm produces a complete partition of the scaffold mutation set 𝒢 = {𝑔) , 𝑔*, … , 𝑔N }, 591 \nwhere ⋃ 𝑔O = ℳG\nN\nO+)  and 𝑔O ∩ 𝑔O) = ∅ for 𝑘 ≠ 𝑘K. Each mutation group 𝑔O defines the 592 \ncharacteristic mutational profile of a putative cellular subpopulation, representing clones that 593 \nlikely evolved through distinct evolutionary trajectories. 594 \nFrom each phylogenetically coherent mutation gr oup 𝑔O ∈ 𝒢, we selected a single backbone 595 \nmutation to serve as the phylogenetic anchor and operational founder of that clonal group. 596 \nThe founder mutation of a clone is expected to be an early event, present in the vast majority 597 \nof cells within that clonal lineage. Therefore, for each group 𝑔O, we selected the mutation with 598 \nthe highest Mutant Cell Fraction (MCF), defined as the fraction of cells with adequate 599 \nsequencing coverage that harbor the mutation: 600 \n𝑗O\n∗ = 𝑎𝑟𝑔 max\n&∈P*\n𝑀𝐶𝐹(𝑗) (27) 601 \nwhere 𝑀𝐶𝐹(𝑗) =\n|𝒮(&)|\n|J%| . Here, 𝒮(𝑗) = A𝑖 ∈ 𝒞GD𝐼$,&\nG = 1E is the set of scaffold cells supporting 602 \nmutation 𝑗, and |𝐶&| denotes the total number of scaffold cells with non -zero sequencing 603 \ncoverage at locus 𝑗. 604 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 21 \nThe stringent, coverage-based filtration applied to define the scaffold set ℳG (Section 3.3) 605 \nensures that all candidate mutations have low genome -wide missing data rates ( 𝑝?@\n(&) ≤ 90%) 606 \nor are robustly detected across multiple cell types (𝑁& ≥ 2). This pre-processing step mitigates 607 \nthe risk of selecting mutations with spuzriously high MCF due to technical artifacts, thereby 608 \nguaranteeing that a high MCF is a reliable indicator of high cellular prevalence. 609 \nThe final set of backbone mutations is defined as the union of all group representatives: 610 \nℬ = {𝑗)\n∗, 𝑗*\n∗, … , 𝑗N\n∗ } (28) 611 \nThese backbone mutations provide a robust set of phylogenetically informative markers for the 612 \ninitial scaffold tree construction. This selection provides the foundational framework for the 613 \nphylogenetic tree. Notably, the designation of an initial founder mutation is operational for the 614 \ninitial tree building and may be subject to re-evaluation during subsequent steps when placing 615 \nadditional mutations, allowing for the integration of more complex phylogenetic signals. 616 \nFollowing the selection of backbone mutations as phylogenetic anchors, we define the 617 \ncorresponding backbone clones that constitute the fundamental phylogenetic units in our 618 \nreconstruction framework. 619 \nDefinition 3.3.4.1 (Backbone Clone). For each backbone mutation 𝑗O\n∗ ∈ ℬ representing group 620 \n𝑔O, the corresponding backbone clone 𝐶O\n∗ is defined as the set of scaffold cells exhibiting the 621 \ncharacteristic mutation pattern of that phylogenetic lineage: 622 \n𝐶O\n∗ = o𝑖 ∈ 𝐶Gp𝐼$,&*\n∗G = 1u 623 \nTo robustly assign cells to backbone clones while accounting for missing data and technical 624 \ndropouts, we employ a group -wise imputation strategy. For each backbone mutation 𝑗O\n∗, an 625 \nimputed clone vector 𝑉O is constructed by integrating phylogenetic signals from all mutations 626 \nwithin group 𝑔O: 627 \n𝑉O (𝑖) = 1( F 𝐼$,&*\n∗G > 0\n&∈P*\n) 628 \nwhere 1(∙)  is the indicator function. This approach enhances detection sensitivity by 629 \nleveraging phylogenetic coherence within each group. 630 \nCells may exhi bit mutation patterns consistent with multiple backbone clones due to 631 \nphylogenetic relationships or technical artifacts. To resolve such ambiguities, a maximum -632 \nsupport assignment strategy is applied. Let 𝒜(𝑖) = {𝑘|𝑉O(𝑖) = 1} denote the set of candidate 633 \nbackbone clones for cell 𝑖. The final assignment is determined by: 634 \n𝑘∗(𝑖) =\n⎩⎪⎨\n⎪⎧ ∅\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t𝑖𝑓\t|𝒜(𝑖)| = 0\n𝑘\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t𝑖𝑓\t|𝒜(𝑖)| = 1\n𝑎𝑟𝑔 max\nO∈𝒜($)\nF 𝐼$,&\nG\n&∈P*\n\t\t\t\t\t\t𝑖𝑓\t|𝒜(𝑖)| > 1\n 635 \nIn case of ties, the backbone mutation with the highest expression level in the cell is selected: 636 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 22 \n𝑘∗(𝑖) = 𝑎𝑟𝑔 max\nO∈𝒜&+,($)\n𝐸$,&*\n∗  637 \nwhere 𝐸$,& denotes the expression level of mutation 𝑗 in cell 𝑖, and 𝒜2$5 (𝑖) is the set of 638 \nclones with equal maximal support. 639 \nThe backbone clone matrix 𝐵M90!5 is then constructed as: 640 \n𝐵M90!5 = o1\t\t\t\t𝑖𝑓\t𝑘∗(𝑖) = 𝑘\n0\t\t\t\t𝑜𝑡ℎ𝑒𝑟𝑤𝑖𝑠𝑒\t\t 641 \nThis matrix provides a robust binary representation of clonal membership, forming the basis 642 \nfor subsequent phylogenetic tree construction. The assignment procedure ensures that each cell 643 \nis assigned to at most one backbone clone, leveraging phylogenetic coherence to enhance 644 \nrobustness against technical noise while resolving ambiguities using biologically meaningful 645 \ncriteria. 646 \n 647 \n4. A Bayesian penalty approach to place left mutations onto the scaffold tree 648 \nThe scaffold tree reconstruction proceeds iteratively, starting from the backbone tree 𝑇R 649 \ncontaining only the backbone mutations ℬ . Each non -backbone mutation 𝑎 ∈ ℳG ∖ ℬ is 650 \nsequentially integrated into the current tree 𝑇, which evolves during the reconstruction process. 651 \nFor each candidate mutation 𝑎, we evaluate its phylogenetic relationship to every possible 652 \nplacement position within 𝑇  by computing a binary discordance penalty based on co -653 \noccurrence patterns and a comprehensive Bayesian penalty that accounts for phylogenetic 654 \nstructure and uncertainty. 655 \n4.1) Identification of candidate nodes to place the mutation 656 \nFor a given non -backbone mutation 𝑎, we comprehensively evaluate all phylogenetically 657 \npermissible placement positions within the current tree 𝑇. Candidate positions 𝑏 ∈ 𝑇 includes: 658 \n(i) any existing node in 𝑇; (ii) a new node introduced along any existing branch, thereby 659 \nsplitting it into two segments; (iii) a new parent node unifying two or more descendants from 660 \nan existing node; or (iv) a new leaf node attached to any existing node. This exhaustive 661 \nconsideration ensures that all potential evolutionary relationships —including those requiring 662 \nexpansion of the tree topology—are evaluated. 663 \nTo systematically manage these placement options, we employ an anchor-based approach. For 664 \neach candidate node 𝑏 ∈ 𝑇 (which represents a distinct branch point or phylogenetic position 665 \nwithin the tree) and non -backbone mutation  𝑎 ∈ ℳG ∖ ℬ , we compute the standard 666 \ncontingency counts 𝑁)) (𝑏, 𝑎), 𝑁)( (𝑏, 𝑎), 𝑁() (𝑏, 𝑎) and 𝑁(( (𝑏, 𝑎) as defined in Equation 3. 667 \nTo account for uncertainty in mutation calls derived from posterior probabilities 𝑃$,&, we define 668 \nfor each cell in these contingency categories a minimum and maximum possible penalty 669 \ncontribution, based on whether the observed posterior is close to the 0.5 threshold or 670 \nconfidently assigned. 671 \nThe total binary discordance penalty for the pair (𝑎, 𝑏) is defined as: 672 \n𝑃𝑒𝑛𝑎𝑙𝑡𝑦S$!T,U (𝑏, 𝑎) = +𝜔)( ∗ 𝑁)( (𝑏, 𝑎) + 𝜔() ∗ 𝑁() (𝑏, 𝑎) (29) 673 \nwhere the weights are assigned as follows: 674 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 23 \n• 𝜔)( = 1: Full penalty for each cell in 𝑁)( , representing a potential false positive 675 \n(mutation 𝑎 observed but one expected under position 𝑏; 676 \n• 𝜔() = 0.1: A reduced penalty for each cell in 𝑁() , representing a potential false 677 \nnegative (mutation 𝑎 not observed but expected under position 𝑏), reflecting the 678 \nhigher empirical rate of false negatives in single-cell data. 679 \nFor each contingency count, we compute the minimum and maximum possible penalty per cell, 680 \nconsidering the uncertainty in posterior probability values around the 0.5 threshold: 681 \n• For a cell in 𝑁)( (𝑏, 𝑎): the penalty contribution ranges from 0.5 (if  𝑃$,& is just 682 \nabove 0.5) to 1 (if 𝑃$,& = 1). 683 \n• For a cell in 𝑁() (𝑏, 𝑎): the penalty contribution ranges from 0.05 (if  𝑃$,& is just 684 \nbelow 0.5) to 0.1 (if 𝑃$,& = 0). 685 \nThus, the overall minimum and maximum penalties for the pair (𝑏, 𝑎) are: 686 \n𝑃𝑒𝑛𝑎𝑙𝑡𝑦S$!T,U\n#$! (𝑏, 𝑎) = 1 ∗ 0.5 ∗ 𝑁)( (𝑏, 𝑎) + 0.1 ∗ 0.5 ∗ 𝑁() (𝑏, 𝑎)  687 \n𝑃𝑒𝑛𝑎𝑙𝑡𝑦S$!T,U\n#TV (𝑏, 𝑎) = 1 ∗ 1 ∗ 𝑁)( (𝑏, 𝑎) + 0.1 ∗ 1 ∗ 𝑁() (𝑏, 𝑎) (30) 688 \nTo identify candidate anchors, we compare the penalty intervals across all possible 𝑏 ∈ 𝑇. We 689 \nfirst identify the mutation 𝑏∗ with the smallest minimum penalty: 690 \n𝑏∗ = 𝑎𝑟𝑔 min\nW∈X\n𝑃𝑒𝑛𝑎𝑙𝑡𝑦S$!T,U\n#$! (𝑏, 𝑎) (31) 691 \nThe candidate set is then constructed by including all positions 𝑏 for which the minimum 692 \npenalty is no greater than the maximum penalty of this optimal candidate: 693 \nℳT!M;0, = A𝑏 ∈ 𝑇D𝑃𝑒𝑛𝑎𝑙𝑡𝑦S$!T,U\n#$! (𝑏∗, 𝑎) ≤ 𝑃𝑒𝑛𝑎𝑙𝑡𝑦S$!T,U\n#TV (𝑏∗, 𝑎)E (32) 694 \nThis approach ensures that we consider not only the position with the absolute smallest penalty, 695 \nbut also any other positions whose minimum penalty falls within the uncertainty range of the 696 \nbest candidate. 697 \nFor each candidate anchor 𝑏 ∈ ℳT!M;0,, we consider two primary phylogenetic hypotheses: 698 \n• Placement on position  𝑏 : implying an identical phylogenetic origin ( 𝑎 = 𝑏 ), 699 \nsuggesting that observed differences (𝑁)( , 𝑁() ) are due to sequencing errors. 700 \n• Placement as a direct descendant of position 𝑏 : implying a parent -child 701 \nrelationship (𝑎 ⊂ 𝑏), where 𝑎 is a subclone of 𝑏. 702 \nAdditionally, placements that introduce new nodes along the branch leading to 𝑏, or that unify 703 \n𝑏 with other siblings under a new parent node, are evaluated through an extension of this 704 \nanchor-based framework, leveraging the same penalty structure to ensure phylogenetic 705 \nconsistency. 706 \n4.2) A Bayesian approach to penalization 707 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 24 \nWe applied a Bayesian approach to place new mutations onto the scaffold phylogenetic tree. 708 \nThe optimal placement was identified by maximizing the posterior probability of the 709 \nmutation’s location, conditional on the scaffold tree structure and the observed data. This is 710 \nformally equivalent to minimizing a penalty function defined as the negative log -posterior 711 \nprobability. The Bayesian function is as follows: 712 \n 713 \n𝑃(𝑁𝑜𝑑𝑒 = 𝑏|𝑇, 𝐷) = 𝑃(𝑁𝑜𝑑𝑒 = 𝑏|𝑇) ∗ 𝑃(𝐷|𝑇, 𝑁𝑜𝑑𝑒 = 𝑏)\n𝑃(𝐷|𝑇) (33) 714 \nWhere: 715 \n• 𝑁𝑜𝑑𝑒: a candidate phylogenetic position in tree 𝑇 where mutation might be 716 \nplaced. 717 \n• 𝐷 = {𝑃) ,T, 𝑃*,T, … , 𝑃!,T}: the somatic posterior probabilities vector for mutation 718 \nacross all cells. 719 \n• 𝑃$,T: probability that mutation 𝑎 is present in cell 𝑖. 720 \nTaking the negative logarithm yields the penalty function: 721 \n𝑃𝑒𝑛𝑎𝑙𝑡𝑦(𝑎, 𝑏) = −log𝑃(𝑁𝑜𝑑𝑒 = 𝑏|𝑇, 𝐷) (34) 722 \nExpanding this expression: 723 \n𝑃𝑒𝑛𝑎𝑙𝑡𝑦(𝑎, 𝑏) = −log𝑃(𝐷|𝑇, 𝑁𝑜𝑑𝑒 = 𝑏) − log𝑃(𝑁𝑜𝑑𝑒 = 𝑏|𝑇) + log(𝐷|𝑇) (35) 724 \nSince 𝑃(𝐷|𝑇) is constant across candidate placements for the current tree: 725 \n𝑃𝑒𝑛𝑎𝑙𝑡𝑦(𝑎, 𝑏) ∝ −log𝑃(𝐷|𝑇, 𝑁𝑜𝑑𝑒 = 𝑏) − log𝑃H(𝑁𝑜𝑑𝑒 = 𝑏|𝑇)I (36) 726 \nAssuming a uniform prior over 𝑁!0453 possible placement nodes in the current tree 𝑇: 727 \n𝑃(𝑁𝑜𝑑𝑒 = 𝑏|𝑇) = 1\n𝑁!0453\n(37) 728 \n−log𝑃(𝑁𝑜𝑑𝑒 = 𝑏|𝑇) = log𝑁!0453 (38) 729 \nThe data likelihood factors over cells based on the current tree structure: 730 \n𝑃(𝐷|𝑇, 𝑁𝑜𝑑𝑒 = 𝑏) = ¨ 𝑃H𝐷$,TD𝐺$,T I\n!\n$+)\n(39) 731 \n− log 𝑃(𝐷|𝑇, 𝑁𝑜𝑑𝑒 = 𝑏) = F log 𝑃H𝐷$,TD𝐺$,TI\n!\n$+)\n(40) 732 \nWhere: 733 \n• 𝐺$,T: expected genotype (0 or 1) given placement at 𝑁𝑜𝑑𝑒S; 734 \n• 𝐷$,T: observed data (posterior probability 𝑃$,T). 735 \n4.3) Mutation placement by minimizing the penalty function 736 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 25 \nThe theoretical framework translates to the practical penalty function: 737 \n𝑃𝑒𝑛𝑎𝑙𝑡𝑦(𝑎, 𝑏) = Fx𝜔7? ∗ 𝑁( →)\n($) + 𝜔78 ∗ 𝑁) →(\n($) + 𝜔?@→)\n(T) ∗ 𝑁?@→)\n($) + 𝜔?@→(\n(T) ∗ 𝑁?@→(\n($) z\n?\n$+)\n+ log 𝑁!0453 (41) 738 \nBinary discordance indicators (0 or 1): 739 \n𝑁( →)\n($) = 𝟏(𝐺$,T = 1\t⋀\t𝐼$,T = 0) 740 \n𝑁) →(\n($) = 𝟏(𝐺$,T = 0\t⋀\t𝐼$,T = 1) 741 \n𝑁?@→)\n($) = 𝟏(𝐺$,T = 1\t⋀\t𝐶$,T = 0) 742 \n𝑁?@→(\n($) = 𝟏H𝐺$,T = 0\t⋀\t𝐶$,T = 0I (42) 743 \nWhere: 744 \n• Each indicator function 𝟏(∙)  returns 1 if the specified condition is true, 0 745 \notherwise. 746 \n• 𝑁( →)\n($) : indicates if cell 𝑖  has a false negative (expected mutant but observed 747 \nreference). 748 \n• 𝑁) →(\n($) : indicates if cell 𝑖 has a false positive (expected reference but observed 749 \nmutant). 750 \n• 𝑁?@→)\n($) : indicates if cell 𝑖 has missing data imputed as mutant. 751 \n• 𝑁?@→(\n($) : indicates if cell 𝑖 has missing data imputed as reference. 752 \nWeight computation from posterior matrix 𝑃: 753 \n• 𝜔78 = − log(𝑃$,T): penalty for false positives. 754 \n• 𝜔7? = 0.1 ∗ (− logH1 − 𝑃$,TI): penalty for false negatives, scaled by empirical 755 \nerror ratio. 756 \n• 𝜔?@→)\n(T) = − log(𝜔?@ ∗\n!!\n(.)\n!0\n(.)-!!\n(.)): penalty for imputing NA as mutant. 757 \n• 𝜔?@→(\n(T) = − log(𝜔?@ ∗\n!0\n(.)\n!0\n(.)-!!\n(.)): penalty for imputing NA as reference. 758 \nEmpirical error ratio incorporation:  759 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 26 \nBased on extensive analysis of single -cell phylogenetic trees from real datasets, we observed 760 \nthat false negative errors occur approximately 10 times more frequently than false positive 761 \nerrors. This empirical observation ( 𝑓𝑛𝑓𝑝_𝑟𝑎𝑡𝑖𝑜 ≈ 0.1) is incorporated by scaling the false 762 \nnegative penalty, ensuring the penalty function reflects actual error patterns in single cell  763 \nsequencing data. Both penalties maintain their log-probability interpretation while accounting 764 \nfor real-world error distributions. 765 \nWhere: 766 \n• 𝑛)\n(T) = |{𝑖: 𝐼$,T = 1}|: number of cells with observed mutant allele. 767 \n• 𝑛(\n(T) = |{𝑖: 𝐼$,T = 0}|: number of cells with observed reference allele. 768 \n• 𝜔?@ = 0.001: global dropout weight for typical datasets. 769 \nAdditionally, we incorporate a complexity penalty based on the Bayesian Information Criterion 770 \n(BIC) to avoid overfitting: 771 \n𝐵𝐼𝐶85!T92U (𝑏, 𝑎) = 𝜑 ∙ log 𝑛 (43) 772 \n𝑃𝑒𝑛𝑎𝑙𝑡𝑦202T9 (𝑏, 𝑎) = 𝑃𝑒𝑛𝑎𝑙𝑡𝑦(𝑏, 𝑎) + 𝐵𝐼𝐶85!T92U (𝑏, 𝑎) (44) 773 \nWhere: 774 \n• 𝜑: increase in number of parameters when adding 𝑎; 775 \n• 𝜆: regularization parameter for complexity control. 776 \n 777 \nThe mutation 𝑎 is placed at the candidate node 𝑏∗ that minimizes the total penalty: 778 \n𝑏∗ = 𝑎𝑟𝑔 min\nS∈ℳ.1234#\n𝑃𝑒𝑛𝑎𝑙𝑡𝑦202T9 (𝑏, 𝑎) (45) 779 \nIf ℳT!M;0, = ∅, the mutation is placed at the root of the tree, subject to a root-specific penalty 780 \ncalculation. 781 \nAfter processing all non -backbone mutations through this iterative placement procedure, the 782 \nresulting tree structure represents the complete  scaffold tree  𝑇G , containing all scaffold 783 \nmutations at their phylogenetically consistent positions. 784 \nThe scaffold tree 𝑇G provides a robust evolutionary framework grounded in high -confidence 785 \nmutations from ubiquitously expressed regions, serving as the foundation for subsequent 786 \nphylogenetic analysis and placement of additional mutations. 787 \n5. Prune & replace likely artifacts and reconstruct the tree 788 \nFollowing initial tree construction, each mutation was evaluated using five quality metrics to 789 \nidentify potential misplacements: 790 \na) The false-positive flip count of the mutation itself, 791 \nb) The false-positive flip count of its correlated mutations, 792 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 27 \nc) The false-negative flip number of the mutation itself, 793 \nd) The false-negative flip number of its mother mutations in its assigned subclone, 794 \ne) The number of mutant cells carrying the mutation withing its subclone. 795 \nBased on these metrics, mutations were refined as follows: 796 \nMisplaced clusters: If a mutation, along with its correlated mutations, exhibited an excessive 797 \nnumber of false-positive flips (a + b > 0.2), the entire group was considered misplaced. In this 798 \ncase, the cluster was pruned and then re -integrated into the phylogeny at a more plausible 799 \nlocation using the Bayesian penalty approach mentioned. 800 \nRelocation of orphaned mutations : If a mutation displayed a high false -negative flip count 801 \n(c > 2) concurrent with a low mutant cell count in its subclone (e < 2), it was deemed to lack a 802 \nvalid phylogenetic parent. Such mutations were relocated to an independent clonal branch. 803 \nMutations that failed to meet the quality thresholds after one iteration of refinement were 804 \ndiscarded from the final phylogeny. This process was repeated iteratively 1 -5 times until the 805 \ntree topology stabilized. 806 \nRemoval of likely doublet cells: Finally, likely doublet cells contain too many false -positive 807 \nflips ratio (> 0.2) were excluded from the final tree. 808 \nBenchmarking of PhyloSOLID 809 \nFor the benchmark, we used real-world single-cell whole-genome sequencing (WGS) data and 810 \nWGS data from single-cell colonies18–21. The ground-truth mutation lists and phylogenies for 811 \nthese datasets were sourced from their original publications. Germline variants and artifact 812 \nmutations were identified by randomly sampling reads from the raw BAM files 813 \nusing mpileup (Supplementary Table 4). The second simulated dataset used in Fig. 1f was 814 \ngenerated by modeling the allele dropout ratios observed in a realistic scRNA-seq data. 815 \nFor the real -world single-cell RNA-seq data, candidate mosaic mutations were called using 816 \nMosaicSC24, and the ground -truth CRISPR -barcode phylogenies were obtained from the 817 \ncorresponding original papers (Supplementary Table 5-6). 818 \nPhylogenetic trees were constructed from both simulated  and real -world single -cell data 819 \nusing eight software tools, including  HUNTRESS, CellPhy, ScisTree, ScisTree2, BSCITE, 820 \nSiFit, PhISCS and MPBoot, and compared against PhyloSOLID. Default parameters were used 821 \nfor all tools. The tanglegrams and comparison plots between the barcode trees and mutation 822 \ntrees generated by each method were produced using  R packages dendextend33 (v1.17.1) and 823 \nphangorn34 (v2.11.1).  824 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 28 \n 825 \nSupplementary Fig. 1: Framework of PhyloSOLID . a, Overall pipeline. See Methods for 826 \ndetails. b-c, Principal component analysis (PCA) of extracted sequence features distinguishes 827 \nsomatic mutations from germline variants and technical artifacts. In the resulting PCA space, 828 \ntrue somatic mutations form a distinct cluster, separable from both germline variants and 829 \nartifacts. d-e, Application of a logistic regression classifier to discriminate somatic mutations. 830 \nReceiver operating characteristic (ROC) curves demonstrate high classification accuracy on 831 \nboth scDNA -seq and scRNA -seq benchmark datasets . f-g, Identification of correlated 832 \nbackbone mutations via a leaden graph. Representative leiden graphs derived from scDNA-seq 833 \n(f) and scRNA -seq ( g) data are shown.  Benchmark dataset details:  UMB1465, UMB4638, 834 \nUMB4643: scDNA-seq data after whole -genome amplification (sample IDs from Lodato et 835 \nal., Science, 2018). DB6, DB9, DB10: Single-cell colony data (sample IDs from Bae et 836 \nal., Science, 2018, and Park et al.,  Nature, 2021, respectively).  LUAD-1 and LUAD -2 837 \ncorrespond to the original sample IDs 100k and 10k, respectively (Quinn et al., Science, 2021). 838 \ncSCC_1 (original sample ID: P6_cSCC) and normal_skin_1 (original sample ID: P6_normal)  839 \nare from  Andrew et al. , Cell, 2020 . normal_breast_1 (original sample ID: hbca_n03) and 840 \nnormal_breast_2 (original sample ID: hbca_c11) are from Tapsi et al., Nature, 2023. 841 \n  842 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 29 \n 843 \n 844 \nSupplementary Fig. 2: Benchmarking PhyloSOLID on simulated data with high sparsity. 845 \nThe ability of each method to recover the true monoclonal structure under high dropout 846 \nconditions is assessed. For each panel, phylogeny generated by each software was shown on 847 \nthe left and simulated mutation matrix was shown on the right . Rows correspond to cells and 848 \ncolumns to simulated mutations. Mutant cells are shown in red, reference-homozygous (refhom) 849 \nsites in light blue, and sites with no coverage in white.  Execution of PhiSCS, CellPhy, 850 \nScisTree2 and MPBoot failed on this dataset.   851 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 30 \n 852 \nSupplementary Fig.  3: Benchmarking PhyloSOLID on scDNA -seq data . Ground-truth 853 \nphylogeny (leftmost tree, from published data) and the corresponding phylogenies inferred by 854 \neach software.  Each row represents a cell.  Tanglegrams were shown to visualize their 855 \nagreement with the ground-truth tree (the left-most tree). Benchmark dataset details: UMB1465, 856 \nUMB4638, UMB4643: scDNA-seq data after whole-genome amplification (sample IDs from 857 \nLodato et al.,  Science, 2018). TNBC16:  scDNA-seq data after whole -genome amplification 858 \n(sample ID from Wang et al., Nature, 2014). S316, DB10: Single-cell colony data (sample IDs 859 \nfrom Bae et al., Science, 2018, and Park et al., Nature, 2021, respectively). 860 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 31 \n 861 \nSupplementary Fig. 4: Benchmarking phylogeny reconstruction tools on scRNA-seq data 862 \n(LUAD-1) with CRISPR barcode ground truth . For each evaluated tool, the output is 863 \norganized as follows: (Left)  Phylogeny inferred from mosaic mutations called from 864 \nscRNA-seq data. (Middle left)  Cell-by-mutation matrix. Rows: cells; columns: mutations. 865 \nColors denote mutant cells (red), reference -homozygous sites (light blue), and sites with no 866 \ncoverage (white). (Middle right)  Barcode-by-cell matrix. Execution of PhiSCS failed on this 867 \ndataset. Each column represents one CRISPR barcode mutation. (Right)  Published 868 \nground-truth clonal structure based on CRISPR barcodes.  Benchmark dataset details: the 869 \noriginal sample ID of LUAD-1 is “100k” (Quinn et al., Science, 2021).  870 \n871 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 32 \n 872 \nSupplementary Fig. 5: Benchmarking phylogeny reconstruction tools on scRNA-seq data 873 \nwith CRISPR barcode ground truth  (LUAD-2). For each evaluated tool, the output is 874 \norganized as follows: (Left)  Phylogeny inferred from mosaic mutations called from 875 \nscRNA-seq data. (Middle left)  Cell-by-mutation matrix. Rows: cells; columns: mutations. 876 \nColors denote mutant cells (red), reference -homozygous sites (light blue), and sites with no 877 \ncoverage (white). (Middle right)  Barcode-by-cell matrix. Execution of PhiSCS failed on this 878 \ndataset. Each column represents one CRISPR barcode mutation. (Right)  Published 879 \nground-truth clonal structure based on CRISPR barcodes.  Benchmark dataset details: the 880 \noriginal sample ID of LUAD-2 is “10k” (Quinn et al., Science, 2021).  881 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint \n\n 33 \n882 \nSupplementary Fig.  6: Benchmarking PhyloSOLID on scRNA -seq data with CRISPR 883 \nbarcode ground truth. a, b, Analysis of two lung cancer scRNA-seq datasets. For each dataset: 884 \n(Left) Phylogeny reconstructed from mosaic mutations identified in the scRNA -seq data.  885 \n(Middle left)  Cell-by-mutation matrix. Rows represent cells, columns represent mutations. 886 \nMutant cells are shown in red, reference -homozygous (refhom) sites in light blue, and sites 887 \nwith no coverage in white.  (Middle right) Barcode-by-cell matrix. Each column corresponds 888 \nto one CRISPR barcode mutation.  (Right) Ground-truth clonal structure derived from the 889 \npublished CRISPR barcode data. c, d, UMAP visualization of cells colored by their assigned 890 \nclones. Benchmark dataset details: the original sample IDs of LUAD-1 and LUAD-2 are “100k” 891 \nand “10k” (Quinn et al., Science, 2021). 892 \nreuse, remix, or adapt this material for any purpose without crediting the original authors. \nshare,this preprint (which was not certified by peer review) in the Public Domain. It is no longer restricted by copyright. Anyone can legally \nThe copyright holder has placedthis version posted February 6, 2026. ; https://doi.org/10.64898/2026.02.04.703905doi: bioRxiv preprint","source_license":"Public-Domain","license_restricted":false}