Real-Time Course: Reconstruction of cellular diversity and lineage trajectory based on somatic mutational patterns detected from low-pass single-cell transcriptome data

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
⚙ AI-generated deep summary by qwen3.7-flash, 2026-09-17 · read from full text ⓘ

This preprint introduces Real-Time Course Analysis, a computational framework that reconstructs cellular lineage trajectories by identifying somatic mutations within low-pass single-cell RNA sequencing data. The authors applied this method to two human placental tissue datasets, detecting thousands of somatic mutation candidates and constructing phylogenetic trees that aligned with expression-based cell classifications and pseudotime analyses. While the study demonstrates that somatic mutation signatures can determine the temporal order of cells in a lineage, it acknowledges technical challenges such as sequencing errors and the need to distinguish somatic variants from germline mutations. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Single-cell RNA-seq (scRNA-seq) analysis can describe the states of individual cells but cannot clarify direct relationships between cells. Therefore, we developed a method to trace the progenitor cells of cell differentiation pathways through the cell lineage using somatic mutations detected in low-pass scRNA sequences. Using two human placental tissue datasets, we detected 2,172 and 627 somatic mutation candidates. We used only somatic mutation data to classify placental cells and construct a phylogenetic tree; the results were consistent with those of cell classification by expression profiling and estimation of the differentiation process by pseudotime course analysis. Somatic-mutation-based analysis can determine the temporal order of individual cells in the cell lineage, which is not possible with pseudotime analysis. Therefore, somatic mutation signatures in scRNA-seq data are informative for constructing cell lineages, and in combination with expression data obtained simultaneously, they are extremely useful for delineating the processes of human development and pathogenesis.
Full text 112,155 characters · extracted from preprint-html · click to expand
Real-Time Course: Reconstruction of cellular diversity and lineage trajectory based on somatic mutational patterns detected from low-pass single-cell transcriptome data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Article Real-Time Course: Reconstruction of cellular diversity and lineage trajectory based on somatic mutational patterns detected from low-pass single-cell transcriptome data Satoshi Oota, Kuniya Abe, Hideo Yokota, Kazuho Ikeo This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1558242/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Single-cell RNA-seq (scRNA-seq) analysis can describe the states of individual cells but cannot clarify direct relationships between cells. Therefore, we developed a method to trace the progenitor cells of cell differentiation pathways through the cell lineage using somatic mutations detected in low-pass scRNA sequences. Using two human placental tissue datasets, we detected 2,172 and 627 somatic mutation candidates. We used only somatic mutation data to classify placental cells and construct a phylogenetic tree; the results were consistent with those of cell classification by expression profiling and estimation of the differentiation process by pseudotime course analysis. Somatic-mutation-based analysis can determine the temporal order of individual cells in the cell lineage, which is not possible with pseudotime analysis. Therefore, somatic mutation signatures in scRNA-seq data are informative for constructing cell lineages, and in combination with expression data obtained simultaneously, they are extremely useful for delineating the processes of human development and pathogenesis. Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Somatic mutations accumulate during development and aging, resulting in cellular mosaicism among somatic tissues. These mutations are associated with various diseases, including dementia, cardiovascular diseases, and cancers 1,2 , as well as the aging process itself 3 . Somatic mutations play a role in normal development, for example in the acquisition of the diversity of neuronal cell populations in the central nervous system 4 . According to RNA sequence data analyses, a clonal expansion of somatic mutations occurs across normal tissues 5 . During early human development, 2.8 (95% confidence interval, 2.4–3.3) somatic mutations occur per cell per replication event, which is slightly higher than the rate of germline mutations 6 . However, in the later life, somatic mutation rates range from to mutations per base pair per cell division across various cell types 7 . Somatic mutation rates can reach ten times higher than those for the germline 8 . As a result, even monozygotic twins can be genetically diverse owing to somatic mutations 9 . Somatic mutations are thus ubiquitous events in various organisms, including humans 10 ; they can be regarded as counterparts of germline mutations 11 , whose dynamics are subject to the evolutionary process 12 . Similar to germline cells, each somatic cell carries an ‘evolutionary record’ as “as scars (mutations) that accrue” that accrue in the genome with time. Thus, it is theoretically possible to reconstruct the history of a cell using a retrospective cell tracing manner 13,14 , i.e., somatic mutations bear temporal information of the cell lineage. Pioneered by Charles Whitman 15 in the 19 th century for the developmental biology of invertebrates, lineage tracing is now an essential tool to study stem cells in adult mammalian tissues. Somatic mutations have been extensively studied in the context of cancer evolution 16 . Researchers have devised mathematical models to elucidate somatic mutation dynamics from various biological perspectives 17 . In the case of cancer evolution studies, however, we must address the challanges in detecting de novo mutations; the error rates of next-generation sequencing technologies are considerably high, making the detection of rare variants difficult. To handle the intractability, we usually need ultra-deep sequencing 18 to increase accuracy. Genuine somatic mutations must be distinguished from germline mutations 19 . Issues such as these can cause noise in the method of using ‘phylogenetic’ signatures to infer ancestral mutations in progeny cells.. Single-cell sequencing (SCS) aids analyses in various fields, including rare cell types, uncultured microbes, and mosaicism of somatic tissues 20 . Since somatic mutations are attributed to genomic alterations, it is a straightforward approach to use somatic genome sequence data for the detection of rare variants against the zygotic reference genome. However, in SCS, genome data have a shortcoming: A single cell has only two copies of each DNA molecule, which can lead to various technical issues such as coverage nonuniformity, allelic dropout events, false-positive errors, and false-negative errors 21 . In contrast, with RNA sequence sources, we can utilize hundreds or thousands of copies of RNA molecules in a cell. It is important to consider false positives due to biological and technical factors, including RNA editing, sequencing errors (e.g., random errors introduced during reverse transcription and PCR), and potential sampling errors 10 . The overwhelming abundance of single-cell transcriptome data has the potential to compensate for the disadvantages of RNA sequencing data 22 . However, the currently available frameworks rely on the nonlinear transformation of high-dimensional data to human-recognizable low-dimensional representations, which often retain a nonlinear nature and can deviate from physical entities such as absolute time. Here, we propose a new framework to detect cell lineage trajectories using potential somatic mutations detected in single-cell transcriptome data: Real-Time Course Analysis. Unlike pseudotime course analysis using heterogeneously differentiated cells in a single time snapshot 23,24 , the approach in the present study uses somatic mutations to trace cellular lineages back to progeny cells. Thus, our framework can handle the direction of physical time in terms of the expected number of mutations. We focused on gross phylogenetic signatures in single-cell transcriptome data. For example, 60 somatic cells have approximately 10 96 possible lineages 25 ; hence it is virtually impossible to identify only one true cell lineage with an exhaustive search. Our approach is to narrow the window that covers the true lineage using the gross phylogenetic signatures potentially retained in the RNA sequence data. Today, there are several dimension reduction methods to categorize cell types and perform pseudotime course analysis based on machine learning approaches, such as t-distributed stochastic neighbor embedding (t-SNE) 26 and uniform manifold approximation and projection for dimension reduction (UMAP) 27 . However, these methods sometimes lack both biologically relevant interpretations and reproducibility of global data structure 28 . The aims of this study were: (1) reconstructing putative lineage trees of normal placental tissue cells ( Supplementary figure 1 ) using low-coverage single-cell somatic variants obtained from single-cell RNA-seq data, in which we utilize the reference genome sequence to compensate for missing alignments with a particular threshold; (2) developing a proof of concept (PoC) study for a new framework to infer the temporal cellular trajectory using the genetic mosaicism of single-cell populations estimated from the compensated alignments; and (3) comparing and integrating the somatic variant-based cellular trajectory and the corresponding gene expression profiles, dubbed Real-Time Course analysis. Results Detection of somatic variants in the single-cell transcriptome data We mapped 3,088,286,401 bp of Batch1 sequence data onto the reference human genome (GRCh38) 29 . The average coverage percentage was 0.685% [standard deviation (SD) = 0.231]. For Batch2 sequence data, the average coverage percentage was 0.477% (SD: 0.243). The apparent variants (deviations from the reference genome in the mapped transcripts) contained germline and somatic mutations. As no information was provided for haplotype phasing in the data set of Pavličev et al. 30 , we employed a simple method to detect somatic variants in multiallelic sites in a cell population. We categorized possible mutational patterns across the cells, assuming that at most one somatic mutation occurred in a lineage 31 ( Fig. 1 ). Our strategy followed two steps: (1) screening all the sites in the transcripts that were different from the corresponding reference genome sites and (2) comparing the screened data across cells and selection of multiallelic sites. In a biallelic site, we could not determine the lineage in which the variation occurred or whether the site was heterozygous ( Fig. 1 ). However, when we looked into a multiallelic site, the only explanation would be that at least one mutation occurred at some point in the cell lineage ( Fig. 1 c and Fig. 1 e). Therefore, we searched for multiallelic sites. We detected sites that were different from the reference genome during the first screening (pairwise comparison with the reference genome). A total of 1,965,629 sites in Batch1 and 830,905 sites in Batch2 were detected during the first screening. Examples of apparent variant sites are shown in Supplementary figure 2 . Among these, 89 multiallelic sites were found in 54 Batch1 cells. Moreover, 2,083 sites were multiallelic in 43.2 single cells on average (80% of the 54 single cells) in Batch1. Similarly, in Batch2, 53 sites were multiallelic in all 33 single cells. Moreover, 574 sites were multiallelic in 26.4 single cells on average (80% of the 33 single cells). However, it remains unclear in which lineage the mutation occurred. The state of the reference genome can be used as an ancestral state if the observed nucleotides share the reference site. However, multiallelic sites were always heterozygous under our framework, and the reference genome data were not haplotype-phased. Here, we selected the ‘minor allele’ as a derived variant, assuming that minor alleles are newly derived nucleotides in a cell population. We used 2083 variants sites from Batch1 and 574 variants from Batch2 for subsequent analysis. Annotation of variants Next, we examined the locations of these variants in the genic sequences using SnpEff software 45 . Supplementary tables 1 and 2 show the annotations of the estimated variants. Note that the results include alternative annotations, such as nested intronic genes. Some variant counts overlapped between the categories. The number of variants is different from the results based on our criteria; for example, while we detected 2,083 and 574 variants from Batch1 and Batch2 data with the 80% criterion, respectively, SnpEff software estimated 1,903 and 1,398 variants from Batch1 and Batch2 data with the default parameter set, respectively. SnpEff results showed that there were 550 and 216 missense variants and 199 and 135 synonymous variants in Batch1 and Batch2 data, respectively. dN /dS ratio in coding regions We employed an evolutionary framework to assess the quality of the detected somatic variants. Genuine somatic mutations are expected to be subjected to purifying selection in normal tissues 32 . Thus, we calculated the dN/dS ratios from Batch1 and Batch2 sequence data. The overall dN/dS ratios of Batch1 and Batch2 are 0.865 and 0.556, respectively. These results suggest that the somatic variants in the two samples were subject to purifying selection, supporting our expectations. Phylogenetic analysis of placental single-cell data Using the variant data described above, we reconstructed the phylogenetic trees of single cells from two different individual placental tissue samples ( Figs 2 and 3 ). As we used the parsimony method 33 implemented in MEGA X 34 , branch lengths represent the expected numbers of mutations that occurred on the branches. Each leaf (operational taxonomic unit 35 ) represented a cell sampled from each tissue. A model of the differentiation process from villous cytotrophoblasts (vCTB) to extravillous trophoblasts (EVT) has been constructed by Knofler et al. 36 , in which EVT are differentiated from vCTB via cell column trophoblasts (CCT) ( Supplementary figure 1 ). Pavličev et al. used PCA analysis and the expression patterns of marker genes to classify the single-cell sample into five clusters (cell types): intravillous cytotrophoblast (CYT1, CYT2, and CYT3), EVT, and the maternal decidual cell (DC). Using the single-cell classification by Pavličev et al., we mapped cells onto the lineage tree constructed by our variant analysis. As shown in Fig. 2 , most (eight out of nine) of the CYT1 cells are on the "trunk of the tree [ Fig. 2 (c)]" and considered as progenitor cells of CYT2, CYT3, and EVT. These cells can be ordered on trees on the basis of their mutation patterns. Most differentiated cells, that is, CYT2, CYT3, and EVT, emerged after the appearance of SRR4371572 CYT1 cells [ Fig. 2 (d)]. The differentiation of CYT2 and CYT3 suggests that intravillous cytotrophoblasts are a heterogeneous group. The expression profiles of these cells suggest that differentiation proceeds by dynamic gene expression in a different order rather than uniform progression toward a terminal expression profile 30 . Meanwhile, we found a small number of discrepant results regarding the above patterns, as shown in (a) and (b) in Fig. 2 . Supplementary figure 3 shows the relationships between the percentage of coverage and cells that were classified in a discrepant manner. We found that discrepant cells had relatively low coverage values. A similar trend was observed in Batch2 data, demonstrating that most of the CYT1 progenitor cells were mapped before the branching point ( Fig. 3 ). These data suggest that our phylogenetic analysis can classify cells based solely on variant information, and the results show good agreement with cell classification based on expression profiling. More importantly, it is possible to align the progenitor CYT1 cells in a temporal order, which is not possible with cell clustering based on expression profiles, including pseudotime course analysis 37-39 . Pavličev et al. reported that their samples contained maternal DCs. However, the results of the present study suggest that the putative DC cells assumed by Pavličev et al. appear to be differentiated from fetal CYT or stem cells of CYT ( Figs 2 and 3 ). We reanalyzed single-cell expression data using t-SNE and detected three clusters. Both putative DCs (claimed by Pavličev et al.) belong to t-SNE Cluster 1 ( Fig. 4 ). Therefore, our data, both expression and variant data, suggest that these two cells are similar to CYT 1 cells and not DC cells. In fact, the expression of some DC-characteristic genes, such as CLEC4C , THBD , CD1C , CD80 , IL10 , and IL12B , was not found in these cells 30 . Pseudotime course analysis Using our variant data analysis method, it is theoretically possible to perform "real-time" course analysis, and the cells can be aligned in temporal order ( Figs 2 and 3 ). Here, we compared the results of real-time analysis with those of pseudotime course analysis 37-39 . Fig. 5 shows the results of pseudotime course analysis of SRP090944 data Batch1, and the cells mapped onto the pseudotime contours are marked by letters or numbers representing cell types determined by Pavličev et al. 30 , which is more consistent with the real-time course analysis. Pseudotimes of the cells are represented by contours created by a function of Mathematica 40 . The results were generally consistent with the results of our real-time course analysis except for the direction of the pseudotimes. Sixteen cells had pseudotimes greater than 2.5, of which ten correspond to CYT 1 progenitor cells. Cells differentiated from CYT1 cells, i.e., CYT2 and CYT3, are detected in regions with lower pseudotime values, suggesting again that these cells are temporally distant from CYT 1 cells in the cell lineage. Although EVT cells are located at close proximity to CYT1 cells on the pseudotime scale, our real-time analysis indicates that these cells are not closely related to CYT1 cells, which is consistent with previous biological knowledge. Therefore, our real-time analysis can yield similar but more direct results compared to the well-established pseudotime course analysis. Discussion This study proposes a new framework to infer cellular lineage trees using somatic variants detected in low coverage single-cell RNA sequence data. We focused on the gross phylogenetic signature of data from the single-cell rather than on each mutation. We showed that it was possible to reconstruct a cellular lineage tree solely based on variant data, and the lineage thus obtained was consistent with the established biological knowledge. The significance of this study is in inferring the direct relationships between single cells using genetic scars in a cell, that is, somatic mutations. However, conventional analyses provide only indirect relationships based on similarities in expression profiles. In addition, because somatic mutations accumulate over time, our real-time course should be able to trace cellular lineages back to progeny cells, meaning that our approach has the potential to allow us to infer the cellular time course, for example, non-observable past events. In fact, our real-time course provided results more consistent than those of pseudotime-course analysis. The major difference is that the real-time course based on mutation data can infer linear and direct relationships between cells in the lineage in terms of the number of mutations. In contrast, pseudotime course analysis only can infer non-linear and indirect relationships. To compensate for the latter shortcoming, additional studies are warranted to leverage splicing kinetics, i.e., RNA velocity 41 . We also consider that the real-time course innately integrates the cellular trajectories with the corresponding gene expression profiles while the two datasets are separately handled. This means that their relationships can be used to provide insights into the biological system of interest. A potential issue of our approach is that the real-time course relies on the phylogenetic signature retained in low-pass single-cell sequence data. The phylogenetic signature can be obscured by random noise due to sequencing errors as well as variant dropouts due to low coverage. In fact, as shown in Supplementary figure 3 , the discrepant cells had relatively low coverage values (indicated by red bars). To assess the quality of the detected somatic variants, we employed an evolutionary framework in terms of apparent selective pressure. Because somatic mutations have a negative impact on cell survival in the cell population, it is rational to assume that genuine somatic mutations are subjected to purifying selection in normal tissues 32 . According to the overall dN/dS values, somatic mutations were subject to purifying selection, suggesting that the majority of the detected somatic mutations were not likely to be false positives. It is important to identify how frequently somatic mutations occur in normal tissues. According to Tomasetti et al. 7 , the in vivo tissue-specific somatic mutation probability per base per cell division is (SE) in normal lymphocytes. Therefore, stochastic somatic alterations occur at an average rate of three mutations per cell division in the human genome, except for repetitive regions 42,43 . This abundance of somatic mutations makes retrospective cell lineage inference feasible, even in normal tissues. In our single-cell analysis, we need to consider three types of ‘variants’: (1) polymorphism of the population (mutations between individuals): i.e., germline mutations; (2) de novo mutations within the individual: i.e., somatic mutations; and (3) sequencing errors. In bulk sequencing data, we need to employ a statistical approach to handle somatic mutations; typically, we use variant allele frequency (VAF) to infer whether a variant comes from somatic cells or is inherited from parents 44 . However, the distribution of RNA-VAF and DNA-VAF can differ owing to allelic imbalance (AI) 45 . Single-cell analysis allowed us to exclude germline mutations by simply comparing variants in a cell population. It remains challenging to identify the lineage in which a somatic mutation occurs ( Fig. 1 ). We selected ‘minor alleles’ as potential somatic mutations, assuming that the transcriptome allele frequencies reflected the genomic allele frequencies in many genomic regions. This assumption is based on an empirical expectation that the allele frequencies of the genome and transcriptome reads are correlated 46 . We used homologous sites of the reference genome to fill in the missing sites in the inter-cell alignments, assuming that genomic regions that were not covered by transcripts were identical to the reference genome sequences. We used an 80% threshold as a pragmatic solution. In our method, homozygous sites cannot be employed to detect variants because multiallelic sites are required, in which variant sites are consequently chromosome-heterozygous in those progenies. ( Fig. 1 c and Fig. 1 e). Thus, we missed all potential variants in the homologous sites. This limitation can also be mitigated by using biallelic sites derived from homologous sites ( Fig. 1 g) with statistical haplotype phasing 47 . Nevertheless, we believe that our primary objective of providing a PoC for cell lineage inference using transcriptome data has been achieved. Further analyses using other types of mutations will be conducted separately in the future. Ideally, single-cell genome data would be more informative for inferring a cell lineage tree. We employed static and rich information to detect somatic mutations. However, there are technical and practical issues with single-cell genome analysis: We need to deal with only two copies of DNA molecules in a single cell, and the cost for sequencing the entire genome of thousands of cells is huge compared to that for RNA sequencing. Conversely, single-cell transcriptome analysis is currently a well-established technology, and a massive amount of single-cell transcriptome data has been (and will be) accumulated in sequence databases. Genotype (mosaicism) information has not been utilized in single-cell gene expression analyses. We showed that it is possible to extract significant genotypic information from RNA sequencing data without incurring additional costs. With rigorous computational approaches (e.g., haplotype phasing), the current technique can be improved. Collectively, while real-time course analysis of single-cell transcriptome data has limitations in terms of zygosity and coverage at apparent variant sites, it has great potential to delineate cellular trajectories during development and to interpret high-dimensional gene expression data. Furthermore, the biological significance of somatic mutations provides insight into a new perspective of ‘evolution’ within an individual. This approach should be highly effective, especially in human biology, in which experimental manipulation is strictly restricted. In conclusion, we propose a new framework for the estimation of cellular trajectories in complex tissues designated as real-time courses. We showed that it is possible to reconstruct a phylogenetic tree of somatic cells in placental tissues using only low-pass single-cell RNA sequence data. This tree is consistent with the known cell lineage of the placenta in four cell types: CYT 1, CYT 2, CYT 3, and EVT. The findings of this study suggest that robust phylogenetic signals exist in RNA sequencing data, and these signals can be used for the construction of cellular networks operating in human development, as well as in human disease pathogenesis. Our research will contribute to single-cell transcriptome analyses in alternative perspectives, for example, to analyze mosaicism of single cell genomes as a phylogenetic signature to infer cellular trajectories. The real-time course is based on mutation data and can infer linear and direct relationships between cells in the lineage in terms of the number of mutations. In contrast, pseudotime course analysis only can infer non-linear and undirected relationships. To compensate for the latter shortcoming, we need additional experiments to leverage splicing kinetics, i.e., RNA velocity 41 . Furthermore, the real-time course can theoretically infer the dynamics of ancestral (progeny) cells that are impossible to observe directly. In contrast, pseudotime course analysis has no capability to reconstruct the past. We also suggest that the real-time course innately integrates the cellular trajectories with the corresponding gene expression profiles while the two data sets are separately handled. This means that we can use their relationships to increase our insight into the biological system of interest. References 1. Kennedy, S. R., Loeb, L. A. & Herr, A. J. Somatic mutations in aging, cancer and neurodegeneration. Mech. Ageing Dev. 133 , 118-126 (2012). 2. Morley, A. A. The somatic mutation theory of ageing. Mutat. Res. 338 , 19-23 (1995). 3. Kelly, D. P. Cell biology: Ageing theories unified. Nature 470 , 342–343 (2011). 4. Abeliovich, A. et al. On somatic recombination in the central nervous system of transgenic mice. Science 257 , 404-410 (1992). 5. Yizhak, K. et al. RNA sequence analysis reveals macroscopic somatic clonal expansion across normal tissues. Science 364 , eaaw0726 (2019). 6. Ju, Y. S. et al. Somatic mutations reveal asymmetric cellular dynamics in the early human embryo. Nature 543 , 714-718 (2017). 7. Tomasetti, C., Vogelstein, B. & Parmigiani, G. Half or more of the somatic mutations in cancers of self-renewing tissues originate prior to tumor initiation. Proc. Natl Acad. Sci. U. S. A. 110 , 1999–2004 (2013). 8. Lynch, M. Rate, molecular spectrum, and consequences of human mutation. Proc. Natl Acad. Sci. U. S. A. 107 , 961–968 (2010). 9. Jonsson, H. et al. Differences between germline genomes of monozygotic twins. Nat. Genet. 53 , 27-34 (2021). 10. García-Nieto, P. E., Morrison, A. J. & Fraser, H. B. The somatic mutation landscape of the human body. Genome Biol. 20 , 298 (2019). 11. Milholland, B. et al. Differences between germline and somatic mutation rates in humans and mice. Nat. Commun. 8 , 15183 (2017). 12. Rozhok, A. I. & DeGregori, J. Toward an evolutionary model of cancer: Considering the mechanisms that govern the fate of somatic mutations. Proc. Natl Acad. Sci. U. S. A. 112 , 8914-8921 (2015). 13. Woodworth, M. B., Girskis, K. M. & Walsh, C. A. Building a lineage from single cells: Genetic techniques for cell lineage tracking. Nat. Rev. Genet. 18 , 230–244 (2017). 14. Oota, S. Somatic mutations - Evolution within the individual. Methods 176 , 91-98 (2020). 15. Davenport, C. B. The personality, heredity and work of Charles Otis Whitman, 1843-1910. Am. Nat. 51 , 5-30 (1917). 16. McGranahan, N. & Swanton, C. Clonal heterogeneity and tumor evolution: Past, present, and the future. Cell 168 , 613-628 (2017). 17. Altrock, P. M., Liu, L. L. & Michor, F. The mathematics of cancer: Integrating quantitative models. Nat. Rev. Cancer 15 , 730-745 (2015). 18. Rheinbay, E. et al. Recurrent and functional regulatory mutations in breast cancer. Nature 547 , 55-60 (2017). 19. Sun, J. X. et al. A computational approach to distinguish somatic vs. germline origin of genomic alterations from deep sequencing of cancer specimens without a matched normal. PLOS Comput. Biol. 14 , e1005965 (2018). 20. Method of the year 2013. Nat. Methods 11 , 1 (2014). 21. Wang, Y. & Navin, N. E. Advances and applications of single-cell sequencing technologies. Mol. Cell 58 , 598-609 (2015). 22. Tam, P. P. L. & Ho, J. W. K. Cellular diversity and lineage trajectory: Insights from mouse single cell transcriptomes. Development 147 , dev179788 (2020). 23. Ji, Z. & Ji, H. TSCAN: Pseudo-time reconstruction and evaluation in single-cell RNA-seq analysis. Nucleic Acids Res. 44 , e117-e117 (2016). 24. Campbell, K. R. & Yau, C. Uncovering pseudotemporal trajectories with covariates from single cell and bulk expression data. Nat. Commun. 9 , 2442-2442 (2018). 25. Felsenstein, J. The number of evolutionary trees. Syst. Biol. 27 , 27-33 (1978). 26. Hinton, G. & Roweis, S. Stochastic neighbor embedding. Adv. Neural Inf. Process. Syst. 15 , 833--840 (2003). 27. McInnes, L., Healy, J. & Melville, J. UMAP: Uniform manifold approximation and projection for dimension reduction, ArXiv . 1802.03426 (2020) . 28. Kobak, D. & Linderman, G. C. Initialization is critical for preserving global data structure in both t-SNE and UMAP. Nat. Biotechnol. 39 , 156-157 (2021). 29. Schneider, V. A. et al. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. bioRxiv , 072116 (2016). 30. Pavličev, M. et al. Single-cell transcriptomics of the human placenta: Inferring the cell communication network of the maternal-fetal interface. Genome Res. 27 , 349-361 (2017). 31. Kimura, M. The number of heterozygous nucleotide sites maintained in a finite population due to steady flux of mutations. Genetics 61 , 893-903 (1969). 32. Yadav, V. K., DeGregori, J. & De, S. The landscape of somatic mutations in protein coding genes in apparently benign human tissues carries signatures of relaxed purifying selection. Nucleic Acids Res. 44 , 2075-2084 (2016). 33. Fitch, W. M. Toward defining the course of evolution: Minimum change for a specific tree topology. Syst. Zool. 20 , 406-416 (1971). 34. Tamura, K., Stecher, G. & Kumar, S.. MEGA11: Molecular Evolutionary Genetics Analysis Version 11. Mol. Biol. Evol. version 11 38 , 3022-3027 (2021). 35. Long, C. A., Sokal, R. R. & Sneath, P. H. A. Principles of numerical taxonomy. J. Mammal. 46 , 111-112 (1965). 36. Knöfler, M. et al. Human placenta and trophoblast development: Key molecular mechanisms and model systems. Cell. Mol. Life Sci. 76 , 3479-3496 (2019). 37. Trapnell, C. et al. The dynamics and regulators of cell fate decisions are revealed by pseudotemporal ordering of single cells. Nat. Biotechnol. 32 , 381-386 (2014). 38. Qiu, X. et al. Reversed graph embedding resolves complex single-cell trajectories. Nat. Methods 14 , 979-982 (2017). 39. Cao, J. et al. The single-cell transcriptional landscape of mammalian organogenesis. Nature 566 , 496-502 (2019). 40. Wolfram Research & I. 12.2 version (Wolfram Research, Inc. (Champaign, IL, 2020)). 41. La Manno, G. et al. RNA velocity of single cells. Nature 560 , 494-498 (2018). 42. Ortega, M. A. et al. Using single-cell multiple omics approaches to resolve tumor heterogeneity. Clin. Transl. Med. 6 , 46 (2017). 43. Araten, D. J. et al. A quantitative measurement of the human somatic mutation rate. Cancer Res. 65 , 8111-8117 (2005). 44. Dou, Y., Gold, H. D., Luquette, L. J. & Park, P. J. Detecting somatic mutations in normal cells. Trends Genet. 34 , 545-557 (2018). 45. Rhee, J. K., Lee, S., Park, W. Y., Kim, Y. H. & Kim, T. M. Allelic imbalance of somatic mutations in cancer genomes and transcriptomes. Sci. Rep. 7 , 1653 (2017). 46. Ju, Y. S. et al. Extensive genomic and transcriptional diversity identified through massively parallel DNA and RNA sequencing of eighteen Korean individuals. Nat. Genet. 43 , 745-752 (2011). 47. Browning, S. R. & Browning, B. L. Haplotype phasing: Existing methods and new developments. Nat. Rev. Genet. 12 , 703-714 (2011). Methods Single-cell RNA-seq. In this study, we used two public transcriptome datasets obtained from normal placental tissues 30 : SRP090944 Batch1 (54 cells) and Batch2 (33 cells). Pavličev et al. 30 analyzed the placenta data in the context of the cell communication network between two semiallogenic individuals, the mother and the fetus. The authors inferred the cell-cell interactome according to the gene expression of receptor-ligand pairs across cell types and found cell-type-specific expression of G-protein-coupled receptors, suggesting that ligand-receptor profiles could be a reliable tool for cell type identification. The data are registered in the DDBJ Sequence Read Archive as SRR4371525 (GSM2339239)–SRS1732319 (SRX2225328) ( Supplementary tables 3 and 4 ). We assumed that Batch1 and Batch 2 shared no de novo somatic mutations. Pavličev et al. categorized the single-cell data into five clusters (cell types) according to the principal comportment analysis (PCA) on the gene expression profile and 300 marker genes: cytotrophoblast (CYT) 1, CYT2, and CYT3, extravillous trophoblast (EVT), and the putative maternal decidual cell (DC). Supplementary figure 1 shows a model system that integrates the role of Notch/Wnt signaling in trophoblast stemness and extravillous trophoblast differentiation 36 Mapping of sequence data. We intended to construct multiple alignment data that represent single cells by aligning the homologous regions of their transcriptomes. To this end, read sequences (segmented sequences of base pairs corresponding to part of a single DNA fragment, a set of which is expected to cover partial coding regions) were assigned designated coordinates of the reference genome. The entire data analysis pipeline is shown in Supplementary figure 4 . We mapped single-cell transcriptome sequence data to the human genome (GRCh38) 29 using the Burrows-Wheeler Aligner 48 after removing adaptor sequences with trimmomatic 49 . We identified variants by comparing the aligned read sequences with the reference genome using Samtools 50 . In the present study, we used only single-nucleotide variants, excluding all detected indel events. If we encountered incomplete sites across cells by more than 50%, we excluded the corresponding sites. We annotated the detected variant sites using the SnpEff software 51 , which can predict the effects of genetic variants on genes and transcripts. Phylogenetic analysis of single cells. We concatenated all acquired variant sites to generate sequence alignments. We reconstructed cellular lineage trees using the maximum parsimony method 33 implemented in MEGA X 34 with default parameters. The results were output in the Newick tree format 52 for the following process: dN/dS ratio computations with the maximum likelihood method and comparative analysis between the mutation patterns and the gene expression profile. dN/dS ratios in the coding regions. We assembled codon sequences into which the detected variants fell and created codon alignments with exonic variants. We computed the overall dN/dS ratios using the Codeml package 53 . Gene expression reanalysis. For linear dimensionality reduction, we conducted principal component analysis (PCA) on the expression patterns using R (version 3.6.2) 54 , and subsequently applied t-SNE 26 and UMAP 27 to the data for nonlinear dimensionality reduction. The Louvain method was used for the clustering 55 . The pseudotime course analysis. We conducted a pseudotime course analysis using monocle3 (version 1.0.0) 37-39 on R (version 4.1.2) 54 . Pseudotime course analysis estimates the temporal dimension of static single-cell transcriptome data and infers individual gene expression dynamics along with cellular changes. We reduced the data dimension to 26 using PCA. To visualize the results, we used Mathematica 40 . Comparative analysis of mutation patterns and gene expression profiles. We mapped the clustered cells to the cellular lineage trees, reconstructed based on cellular genotypes. To this end, we developed a Mathematica 40 code, AssignCluster2Cell , using two packages: Phylogenetics for Mathematica 56 and Phylogenetics 57 . AssignCluster2Cell reads a Newick 52 formatted tree file and a cluster table of a gene expression profile, and maps cluster ID to the tree topology, evaluating degrees of monophyleticity (DoM) at each node in terms of the proportion of the clusters in its subtree ( Fig. 6 ). We also mapped the clustered cells to the tree to evaluate the DoMs. The pie charts in Fig. 6 show the DoM at each node; for example, if the pie chart at a node has only one slice, the subtree is completely monophyletic; if the pie chart at a node has two slices, the subtree is polyphyletic (subtree B in Fig. 6 ). However, because cell type 2 is dominant in subtree B, we can infer that cell type 1 is derived from cell type 2. In this manner, we determined the relevance of cell types based on gene expression profiles and cell lineages. Theoretically, the root of a cellular lineage tree represents a zygotic cell and the root represents the zygotic genome of the fertilized egg. However, in the present study, the root of an estimated tree represented a progenitor cell of the observed cells. The zygotic cell should reside somewhere between the root and the reference genome. References 48. Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25 , 1754-1760 (2009). 10.1093/bioinformatics/btp324, Pubmed:19451168. 49. Bolger, A. M., Lohse, M. & Usadel, B. Trimmomatic: A flexible trimmer for Illumina sequence data. Bioinformatics 30 , 2114-2120 (2014). 10.1093/bioinformatics/btu170, Pubmed:24695404. 50. Li, H. et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25 , 2078-2079 (2009). 10.1093/bioinformatics/btp352, Pubmed:19505943. 51. Cingolani, P. et al. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly (Austin) 6 , 80-92 (2012). 10.4161/fly.19695, Pubmed:22728672. 52. Felsenstein, J. & PHYLIP. Phylogeny Inference Package . 3.2 version. Cladistics 5 , Vols. 164-166, (1989). 53. Yang, Z. PAML 4: Phylogenetic analysis by maximum likelihood. Mol. Biol. Evol. 24 , 1586-1591 (2007). 10.1093/molbev/msm088, Pubmed:17483113. 54. R Core Team ((Vienna, Austria, 2016)). 55. Blondel, V. D., Guillaume, J., Lambiotte, R. & Lefebvre, E. Fast unfolding of communities in large networks. J. Stat. Mech. Theor. Exp. 2008 , 10008 (2008). 10.1088/1742-5468/2008/10/P10008. 56. Polly, P. D. (Indiana University Bloomington. Indiana ). 57. Zachar, I. GitHub repository. in ed686781c686980dcd686982ff111438bb686935a686986c686983a686951f686982b686984 (GitHub, 2017) , Vol. 2021. 686980. 58. Cock, P. J. A., Fields, C. J., Goto, N., Heuer, M. L. & Rice, P. M. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants. Nucleic Acids Res. 38 , 1767-1771 (2010). 10.1093/nar/gkp1137, Pubmed:20015970. 59. Danecek, P. et al. The variant call format and VCFtools. Bioinformatics 27 , 2156-2158 (2011). 10.1093/bioinformatics/btr330, Pubmed:21653522. Declarations Data Availability Statement The data that support the findings of this study are divided into two sets, SRP090944 Batch1 (54 cells) and Batch2 (33 cells). The both data sets are available in the DDBJ Sequence Read Archive (https://www.ddbj.nig.ac.jp/dra/index-e.html) as SRR4371525 (GSM2339239)–SRS1732319 (SRX2225328). Additional Declarations Yes there is potential Competing Interest. We filed a patent application regarding technologies used in this paper. Supplementary Files Supplementarymaterial.docx Supplementary material Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1558242","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":98761982,"identity":"0e694a41-1ebb-478c-82b2-8b49be3b1eaa","order_by":0,"name":"Satoshi Oota","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAo0lEQVRIiWNgGAWjYLCCDwwSBiDagEj1zAyMM0jWwsxDvGog4GfvP/jZdo+FMb90A0NxATFaJHsOM0vnPJMwk5xzgMF4BjFaDG4kM0jnHJCwMbiRwGDMQ6QW5t8WQC32pGhhk2Y4IGFmIEGsFqBfzCx7DkgYS9xIbCDOL/zsjY9v/DhQZ9g/I/mYMVEhhgQY24xJ1AGM1MckaxkFo2AUjIIRAQBNDipe2bxwKwAAAABJRU5ErkJggg==","orcid":"","institution":"RIKEN","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Satoshi","middleName":"","lastName":"Oota","suffix":""},{"id":98761983,"identity":"008fffbb-3f99-4870-8ddd-268434524a19","order_by":1,"name":"Kuniya Abe","email":"","orcid":"","institution":"RIKEN","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kuniya","middleName":"","lastName":"Abe","suffix":""},{"id":98761984,"identity":"1b9bc778-eed6-4848-bb7b-5768bac6e2a6","order_by":2,"name":"Hideo Yokota","email":"","orcid":"","institution":"","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hideo","middleName":"","lastName":"Yokota","suffix":""},{"id":98761985,"identity":"a4aebb98-7d95-408e-a065-54a1a7866d87","order_by":3,"name":"Kazuho Ikeo","email":"","orcid":"","institution":"National Institute of Genetics","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kazuho","middleName":"","lastName":"Ikeo","suffix":""}],"badges":[],"createdAt":"2022-04-14 13:11:21","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1558242/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1558242/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":20359573,"identity":"a7fb29c9-383a-4065-844d-e86513c79acd","added_by":"auto","created_at":"2022-04-14 18:24:00","extension":"tif","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":2279506,"visible":true,"origin":"","legend":"\u003cp\u003eDetection of somatic mutations via multiallelic variants. (a) An apparent variant site derived from an ancestral heterozygous site without germline and somatic mutations. (b) An apparent variant site derived from an ancestral heterozygous site with a somatic mutation. However, we cannot distinguish this configuration from (a). (c) A variant site derived from an ancestral heterozygous site with a somatic mutation. We can distinguish this configuration from (a) and (b) owing to the third kind of nucleotide G. (d) An apparent variant site derived from an ancestral heterozygous site with a germline mutation. However, we cannot distinguish this configuration from (a) and (b). (e) A variant site derived from an ancestral heterozygous site with a somatic mutation following a germline mutation. We can distinguish this configuration from (a) and (b) owing to the third kind of nucleotide T. (f) A homogenous site derived from a homozygous site. (g) A variant site derived from an ancestral homologous site with a somatic mutation. We cannot distinguish this configuration from (a) and (b). (h) An apparent variant site derived from an ancestral homologous site with a somatic mutation following a germline mutation. However, we cannot distinguish this configuration from (d). We assume that the reference genome is haplotype-phased for convenience. The solid line represents a haplotype-phased reference site. The dashed line represents the homozygous reference site. Only two representative cells derived from their ancestral cells are shown here.\u003c/p\u003e","description":"","filename":"Fig01.tif","url":"https://assets-eu.researchsquare.com/files/rs-1558242/v1/a901c89aa342fa7b5b15d481.tif"},{"id":20359575,"identity":"01978141-1ea2-40d5-98ec-6fdf48c149ed","added_by":"auto","created_at":"2022-04-14 18:24:00","extension":"tif","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":4178356,"visible":true,"origin":"","legend":"\u003cp\u003eMapping of the gene expression profile to a cellular trajectory tree (SRP090944 data; Batch1, 54 cells). The gene expression profile (as clustered cells) is mapped to the tree topology: External nodes represent observed cells, and internal nodes represent inferred cells, which are color-coded according to the gene expression patterns with pie charts. (a) Placenta cells claimed as maternal decidual cell (DC) in a previous study\u003csup\u003e30\u003c/sup\u003e; (b) Placenta cells that were not annotated\u003csup\u003e30\u003c/sup\u003e; (c) A putative stem cell self-renewal phase. Scale bar: expected number of mutations. Each number in the pie charts represents a designated cluster according to the gene expression profile.\u003c/p\u003e","description":"","filename":"Fig02.tif","url":"https://assets-eu.researchsquare.com/files/rs-1558242/v1/f4fdfc21fe1e7a8b8dfa8870.tif"},{"id":20359698,"identity":"12b75d90-aa48-455a-8a7e-5d69e3a1b8f5","added_by":"auto","created_at":"2022-04-14 18:29:00","extension":"tif","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":3784155,"visible":true,"origin":"","legend":"\u003cp\u003eMapping of the gene expression profile to a putative lineage topology (SRP090944 data; Batch2, 33 cells). (a) A putative stem cell self-renewal phase. The other notations are the same as in Figure 3.\u003c/p\u003e","description":"","filename":"Fig03.tif","url":"https://assets-eu.researchsquare.com/files/rs-1558242/v1/59380a66d02b11f9b0af0489.tif"},{"id":20359572,"identity":"3feec094-244d-4129-8117-6db05cdf1913","added_by":"auto","created_at":"2022-04-14 18:24:00","extension":"tif","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1283091,"visible":true,"origin":"","legend":"\u003cp\u003eMapping of the reanalyzed gene expression profile by t-SNE to the cellular trajectory tree (SRP090944 data; Batch1, 54 cells).\u003c/p\u003e","description":"","filename":"Fig04.tif","url":"https://assets-eu.researchsquare.com/files/rs-1558242/v1/c5a67392d41f9c8a19806e68.tif"},{"id":20359860,"identity":"5c4ab766-65b4-4269-92d1-9a39f43df522","added_by":"auto","created_at":"2022-04-14 18:34:00","extension":"tif","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":1958365,"visible":true,"origin":"","legend":"\u003cp\u003eResults of the pseudotime course analysis of SRP090944 data Batch1. Contours were created by \u0026nbsp;function of Mathematica\u003csup\u003e40\u003c/sup\u003e with interpolation order 3. The legend represents pseudotimes with colors. All cells are not labeled in this figure due to Mathematica algorithm. 1: intravillous cytotrophoblast (CYT) 1; 2: CYT 3; 3: CYT3; D: putative decidual cell (DC); E: extravillous cytotrophoblast (EVT); b: non-annotated cell (see Figure 3).\u003c/p\u003e","description":"","filename":"Fig05.tif","url":"https://assets-eu.researchsquare.com/files/rs-1558242/v1/6dff56a21a8b0fdd823141a2.tif"},{"id":20359569,"identity":"187bbba5-abe0-4e47-9277-9292527bdc9f","added_by":"auto","created_at":"2022-04-14 18:24:00","extension":"tif","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":798152,"visible":true,"origin":"","legend":"\u003cp\u003eEvaluation of the degrees of monophyleticity (DoM) at each node in terms of the proportion of the clusters in its subtree. The gene expression profile (a cluster table) is mapped to the tree topology. If consistent with each other, we can observe a set of subtrees, each of which is monophyletic in terms of the clusters, otherwise we will observe para- or polyphyletic subtrees in terms of the clustering. We can evaluate the consistency with a size of a monophyletic subtree over the total number of the subtrees. Code\u0026nbsp;\u003cem\u003eAssignCluster2Cell\u003c/em\u003e\u0026nbsp;calculates the proportion of the number of cells assigned to the clusters, by which we can evaluate the DoM of being monophyletic (monophyleticity) of a subtree. A: ancestral stem cell; B, C, and D: derived stem cells; E, F, G, I, and I: observed differentiated single cells.\u0026nbsp;\u003c/p\u003e","description":"","filename":"Fig06.tif","url":"https://assets-eu.researchsquare.com/files/rs-1558242/v1/bb445def1c93b3a7cfa78e75.tif"},{"id":20548756,"identity":"335e08fb-17ac-4cbb-9ecb-f429529b2d38","added_by":"auto","created_at":"2022-04-20 12:47:13","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":20082399,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1558242/v1/b7885213-1eea-41d6-9468-541c99f14ecb.pdf"},{"id":20359700,"identity":"0fcccb6a-abc6-4606-be4b-3910b776302b","added_by":"auto","created_at":"2022-04-14 18:29:00","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":429216,"visible":true,"origin":"","legend":"Supplementary material","description":"","filename":"Supplementarymaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-1558242/v1/b33c5eb812ed7e66535aa33c.docx"}],"financialInterests":"\u003cb\u003eYes\u003c/b\u003e there is potential Competing Interest.\nWe filed a patent application regarding technologies used in this paper.","formattedTitle":"Real-Time Course: Reconstruction of cellular diversity and lineage trajectory based on somatic mutational patterns detected from low-pass single-cell transcriptome data","fulltext":[{"header":"Introduction","content":"\u003cp\u003eSomatic mutations accumulate during development and aging, resulting in cellular mosaicism among somatic tissues. These mutations are associated with various diseases, including dementia, cardiovascular diseases, and cancers\u003csup\u003e1,2\u003c/sup\u003e, as well as the aging process itself\u003csup\u003e3\u003c/sup\u003e. Somatic mutations play a role in normal development, for example in the acquisition of the diversity of neuronal cell populations in the central nervous system\u003csup\u003e4\u003c/sup\u003e. According to RNA sequence data analyses, a clonal expansion of somatic mutations occurs across normal tissues\u003csup\u003e5\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eDuring early human development, 2.8 (95% confidence interval, 2.4\u0026ndash;3.3) somatic mutations occur per cell per replication event, which is slightly higher than the rate of germline mutations\u003csup\u003e6\u003c/sup\u003e. However, in the later life, somatic mutation rates range from\u0026nbsp;\u0026nbsp;\u003csup\u003e\u0026nbsp;\u003c/sup\u003eto\u0026nbsp;\u0026nbsp;\u0026nbsp;mutations per base pair per cell division across various cell types\u003csup\u003e7\u003c/sup\u003e. Somatic mutation rates can reach ten times higher than those for the germline\u003csup\u003e8\u003c/sup\u003e.\u0026nbsp;As a result, even monozygotic twins can be genetically diverse owing to somatic mutations\u003csup\u003e9\u003c/sup\u003e. Somatic mutations are thus ubiquitous events in various organisms, including humans\u003csup\u003e10\u003c/sup\u003e; they can be regarded as counterparts of germline mutations\u003csup\u003e11\u003c/sup\u003e, whose dynamics are subject to the evolutionary process\u003csup\u003e12\u003c/sup\u003e. Similar to germline cells, each somatic cell carries an \u0026lsquo;evolutionary record\u0026rsquo; as \u0026ldquo;as scars (mutations) that accrue\u0026rdquo; that accrue in the genome with time. Thus, it is theoretically possible to reconstruct the history of a cell using a retrospective cell tracing manner\u003csup\u003e13,14\u003c/sup\u003e, i.e., somatic mutations bear temporal information of the cell lineage.\u003c/p\u003e\n\u003cp\u003ePioneered by Charles Whitman\u003csup\u003e15\u003c/sup\u003e in the 19\u003csup\u003eth\u003c/sup\u003e century for the developmental biology of invertebrates, lineage tracing is now an essential tool to study stem cells in adult mammalian tissues. Somatic mutations have been extensively studied in the context of cancer evolution\u003csup\u003e16\u003c/sup\u003e. Researchers have devised mathematical models to elucidate somatic mutation dynamics from various biological perspectives\u003csup\u003e17\u003c/sup\u003e. In the case of cancer evolution studies, however, we must address the challanges in detecting \u003cem\u003ede novo\u003c/em\u003e mutations; the error rates of next-generation sequencing technologies are considerably high, making the detection of rare variants difficult. To handle the intractability, we usually need ultra-deep sequencing\u003csup\u003e18\u003c/sup\u003e to increase accuracy. Genuine somatic mutations must be distinguished from germline mutations\u003csup\u003e19\u003c/sup\u003e. Issues such as these can cause noise in the method of using \u0026lsquo;phylogenetic\u0026rsquo; signatures to infer ancestral mutations in progeny cells..\u003c/p\u003e\n\u003cp\u003eSingle-cell sequencing (SCS) aids analyses in various fields, including rare cell types, uncultured microbes, and mosaicism of somatic tissues\u003csup\u003e20\u003c/sup\u003e. Since somatic mutations are attributed to genomic alterations, it is a straightforward approach to use somatic genome sequence data for the detection of rare variants against the zygotic reference genome. However, in SCS, genome data have a shortcoming: A single cell has only two copies of each DNA molecule, which can lead to various technical issues such as coverage nonuniformity, allelic dropout events, false-positive errors, and false-negative errors\u003csup\u003e21\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eIn contrast, with RNA sequence sources, we can utilize hundreds or thousands of copies of RNA molecules in a cell. It is important to consider false positives due to biological and technical factors, including RNA editing, sequencing errors (e.g., random errors introduced during reverse transcription and PCR), and potential sampling errors\u003csup\u003e10\u003c/sup\u003e. The overwhelming abundance of single-cell transcriptome data has the potential to compensate for the disadvantages of RNA sequencing data\u003csup\u003e22\u003c/sup\u003e. However, the currently available frameworks rely on the nonlinear transformation of high-dimensional data to human-recognizable low-dimensional representations, which often retain a nonlinear nature and can deviate from physical entities such as absolute time.\u003c/p\u003e\n\u003cp\u003eHere, we propose a new framework to detect cell lineage trajectories using potential somatic mutations detected in single-cell transcriptome data: Real-Time Course Analysis. Unlike pseudotime course analysis using heterogeneously differentiated cells in a single time snapshot\u003csup\u003e23,24\u003c/sup\u003e, the approach in the present study uses somatic mutations to trace cellular lineages back to progeny cells. Thus, our framework can handle the direction of physical time in terms of the expected number of mutations. We focused on gross phylogenetic signatures in single-cell transcriptome data. For example, 60 somatic cells have approximately 10\u003csup\u003e96\u003c/sup\u003e possible lineages\u003csup\u003e25\u003c/sup\u003e; hence it is virtually impossible to identify only one true cell lineage with an exhaustive search. Our approach is to narrow the window that covers the true lineage using the gross phylogenetic signatures potentially retained in the RNA sequence data. Today, there are several dimension reduction methods to categorize cell types and perform pseudotime course analysis based on machine learning approaches, such as t-distributed stochastic neighbor embedding (t-SNE)\u003csup\u003e26\u003c/sup\u003e and uniform manifold approximation and projection for dimension reduction (UMAP)\u003csup\u003e27\u003c/sup\u003e. However, these methods sometimes lack both biologically relevant interpretations and reproducibility of global data structure\u003csup\u003e28\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eThe aims of this study were: (1) reconstructing putative lineage trees of normal placental tissue cells (\u003cstrong\u003eSupplementary figure 1\u003c/strong\u003e) using low-coverage single-cell somatic variants obtained from single-cell RNA-seq data, in which we utilize the reference genome sequence to compensate for missing alignments with a particular threshold; (2) developing a proof of concept (PoC) study for a new framework to infer the temporal cellular trajectory using the genetic mosaicism of single-cell populations estimated from the compensated alignments; and (3) comparing and integrating the somatic variant-based cellular trajectory and the corresponding gene expression profiles, dubbed Real-Time Course analysis.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eDetection of somatic variants in the single-cell transcriptome data\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe mapped 3,088,286,401 bp of Batch1 sequence data onto\u0026nbsp;the reference human genome (GRCh38)\u003csup\u003e29\u003c/sup\u003e. The average coverage percentage\u0026nbsp;was 0.685% [standard deviation (SD) = 0.231]. For Batch2 sequence data, the average coverage percentage was 0.477% (SD: 0.243).\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;The apparent variants (deviations from the reference genome in the mapped transcripts) contained germline and somatic mutations.\u0026nbsp;As no information was provided for haplotype phasing in the data set of\u0026nbsp;Pavličev et al.\u003csup\u003e30\u003c/sup\u003e, we employed a simple method to detect somatic variants in multiallelic sites in a cell population. We categorized possible mutational patterns across the cells, assuming that at most one somatic mutation occurred in a lineage\u003csup\u003e31\u003c/sup\u003e (\u003cstrong\u003eFig. 1\u003c/strong\u003e). Our strategy followed two steps: (1) screening all the sites in the transcripts that were different from the corresponding reference genome sites and (2) comparing the screened data across cells and selection of multiallelic sites.\u003c/p\u003e\n\u003cp\u003eIn a biallelic site, we could not determine the lineage in which the variation occurred or whether the site was heterozygous (\u003cstrong\u003eFig. 1\u003c/strong\u003e). However, when we looked into a multiallelic site, the only explanation would be that at least one mutation occurred at some point in the cell lineage (\u003cstrong\u003eFig.\u003c/strong\u003e \u003cstrong\u003e1\u003c/strong\u003ec and\u0026nbsp;\u003cstrong\u003eFig.\u003c/strong\u003e \u003cstrong\u003e1\u003c/strong\u003ee). Therefore, we searched for multiallelic sites.\u003c/p\u003e\n\u003cp\u003eWe detected sites that were different from the reference genome during the first screening (pairwise comparison with the reference genome). A total of 1,965,629 sites in Batch1 and 830,905 sites in Batch2 were detected during the first screening. Examples of apparent variant sites are shown in \u003cstrong\u003eSupplementary figure \u003cem\u003e2\u003c/em\u003e\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eAmong these, 89 multiallelic sites were found in 54 Batch1 cells. Moreover, 2,083 sites were multiallelic in 43.2 single cells on average (80% of the 54 single cells) in Batch1. Similarly, in Batch2, 53 sites were multiallelic in all 33 single cells. Moreover, 574 sites were multiallelic in 26.4 single cells on average (80% of the 33 single cells). However, it remains unclear in which lineage the mutation occurred. The state of the reference genome can be used as an ancestral state if the observed nucleotides share the reference site. However, multiallelic sites were always heterozygous under our framework, and the reference genome data were not haplotype-phased. Here, we selected the \u0026lsquo;minor allele\u0026rsquo; as a derived variant, assuming that minor alleles are newly derived nucleotides in a cell population. We used 2083 variants sites from Batch1 and 574 variants from Batch2 for subsequent analysis.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAnnotation of variants\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNext, we examined the locations of these variants in the genic sequences using SnpEff software\u003csup\u003e45\u003c/sup\u003e. \u003cstrong\u003eSupplementary tables\u003c/strong\u003e \u003cstrong\u003e1\u003c/strong\u003e and \u003cstrong\u003e2\u003c/strong\u003e show the annotations of the estimated variants. Note that the results include alternative annotations, such as nested intronic genes. Some variant counts overlapped between the categories.\u003cem\u003e\u0026nbsp;\u003c/em\u003eThe number of variants is different from the results based on our criteria; for example, while we detected 2,083 and 574 variants from Batch1 and Batch2 data with the 80% criterion, respectively, SnpEff software estimated 1,903 and 1,398 variants from Batch1 and Batch2 data with the default parameter set, respectively. SnpEff results showed that there were 550 and 216 missense variants and 199 and 135 synonymous variants in Batch1 and Batch2 data, respectively.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003edN\u0026nbsp;/dS ratio in coding regions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe employed an evolutionary framework to assess the quality of the detected somatic variants. Genuine somatic mutations are expected to be subjected to purifying selection in normal tissues\u003csup\u003e32\u003c/sup\u003e.\u0026nbsp;Thus, we calculated the dN/dS ratios from Batch1 and Batch2 sequence data. The overall dN/dS ratios of Batch1 and Batch2 are 0.865 and 0.556, respectively. These results suggest that the somatic variants in the two samples were subject to purifying selection, supporting our expectations.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePhylogenetic analysis of placental single-cell data\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eUsing the variant data described above, we reconstructed the phylogenetic trees of single cells from two different individual placental tissue samples (\u003cstrong\u003eFigs\u003c/strong\u003e \u003cstrong\u003e2\u003c/strong\u003e and \u003cstrong\u003e3\u003c/strong\u003e). As we used the parsimony method\u003csup\u003e33\u003c/sup\u003e implemented in MEGA X\u003csup\u003e34\u003c/sup\u003e, branch lengths represent the expected numbers of mutations that occurred on the branches. Each leaf (operational taxonomic unit\u003csup\u003e35\u003c/sup\u003e) represented a cell sampled from each tissue.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;A model of the differentiation process from villous cytotrophoblasts (vCTB) to extravillous trophoblasts (EVT) has been constructed by Knofler et al.\u003csup\u003e36\u003c/sup\u003e, in which EVT are differentiated from vCTB via cell column trophoblasts (CCT) (\u003cstrong\u003eSupplementary figure \u003cem\u003e1\u003c/em\u003e\u003c/strong\u003e). Pavličev et al. used PCA analysis and the expression patterns of marker genes to classify the single-cell sample into five clusters (cell types): intravillous cytotrophoblast (CYT1, CYT2, and CYT3), EVT, and the maternal decidual cell (DC). Using the single-cell classification by Pavličev et al., we mapped cells onto the lineage tree constructed by our variant analysis. As shown in\u0026nbsp;\u003cstrong\u003eFig. 2\u003c/strong\u003e, most (eight out of nine) of the CYT1 cells are on the \u0026quot;trunk of the tree [\u003cstrong\u003eFig.\u003c/strong\u003e \u003cstrong\u003e2\u003c/strong\u003e (c)]\u0026quot; and considered as progenitor cells of CYT2, CYT3, and EVT. These cells can be ordered on trees on the basis of their mutation patterns. Most differentiated cells, that is, CYT2, CYT3, and EVT, emerged after the appearance of SRR4371572 CYT1 cells [\u003cstrong\u003eFig.\u003c/strong\u003e \u003cstrong\u003e2\u003c/strong\u003e (d)]. The differentiation of CYT2 and CYT3 suggests that\u0026nbsp;intravillous cytotrophoblasts are a heterogeneous group. The expression profiles of these cells suggest that differentiation proceeds by dynamic gene expression in a different order rather than uniform progression toward a terminal expression profile\u003csup\u003e30\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMeanwhile, we found a small number of discrepant results regarding the above patterns, as shown in (a) and (b) in\u0026nbsp;\u003cstrong\u003eFig. 2\u003c/strong\u003e.\u0026nbsp;\u003cstrong\u003eSupplementary figure 3\u003c/strong\u003e shows\u0026nbsp;the relationships between the percentage of coverage and\u0026nbsp;cells that were classified in a\u0026nbsp;discrepant\u0026nbsp;manner. We found that discrepant cells had relatively low coverage values.\u003c/p\u003e\n\u003cp\u003eA similar trend was observed in Batch2 data, demonstrating that most of the CYT1 progenitor cells were mapped before the branching point (\u003cstrong\u003eFig. 3\u003c/strong\u003e). These data suggest that our phylogenetic analysis can classify cells based solely on variant information, and the results show good agreement with cell classification based on expression profiling. More importantly, it is possible to align the progenitor CYT1 cells in a temporal order, which is not possible with cell clustering based on expression profiles, including pseudotime course analysis\u003csup\u003e37-39\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;Pavličev et al. reported that their samples contained maternal DCs. However, the results of the present study suggest that the putative DC cells assumed by Pavličev et al. appear to be differentiated from fetal CYT or stem cells of CYT (\u003cstrong\u003eFigs\u003c/strong\u003e \u003cstrong\u003e2\u003c/strong\u003e and\u0026nbsp;\u003cstrong\u003e3\u003c/strong\u003e). We reanalyzed single-cell expression data using t-SNE and detected three clusters. Both putative DCs (claimed by Pavličev et al.) belong to t-SNE Cluster 1 (\u003cstrong\u003eFig. 4\u003c/strong\u003e). Therefore, our data, both expression and variant data, suggest that these two cells are similar to CYT 1 cells and not DC cells.\u0026nbsp;In fact, the expression of some DC-characteristic genes, such as \u003cem\u003eCLEC4C\u003c/em\u003e,\u003cem\u003e\u0026nbsp;THBD\u003c/em\u003e,\u003cem\u003e\u0026nbsp;CD1C\u003c/em\u003e,\u003cem\u003e\u0026nbsp;CD80\u003c/em\u003e,\u003cem\u003e\u0026nbsp;IL10\u003c/em\u003e, and \u003cem\u003eIL12B\u003c/em\u003e,\u003cem\u003e\u0026nbsp;\u003c/em\u003ewas not found in these cells\u003csup\u003e30\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePseudotime course analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eUsing our variant data analysis method, it is theoretically possible to perform \u0026quot;real-time\u0026quot; course analysis, and the cells can be aligned in temporal order (\u003cstrong\u003eFigs\u003c/strong\u003e \u003cstrong\u003e2\u003c/strong\u003e and \u003cstrong\u003e3\u003c/strong\u003e). Here, we compared the results of real-time analysis with those of pseudotime course analysis\u003csup\u003e37-39\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;\u003cstrong\u003eFig. 5\u003c/strong\u003e shows the results of pseudotime course analysis of SRP090944 data Batch1, and the cells mapped onto the pseudotime contours are marked by letters or numbers representing cell types determined by Pavličev et al.\u003csup\u003e30\u003c/sup\u003e, which is more consistent with the real-time course analysis. Pseudotimes of the cells are represented by contours created by a function of \u003cem\u003eMathematica\u003c/em\u003e\u003csup\u003e40\u003c/sup\u003e. The results were generally consistent with the results of our real-time course analysis except for the direction of the pseudotimes. Sixteen cells had pseudotimes greater than 2.5, of which ten correspond to CYT 1 progenitor cells. Cells differentiated from CYT1 cells, i.e., CYT2 and CYT3, are detected in regions with lower pseudotime values, suggesting again that these cells are temporally distant from CYT 1 cells in the cell lineage. Although EVT cells are located at close proximity to CYT1 cells on the pseudotime scale, our real-time analysis indicates that these cells are not closely related to CYT1 cells, which is consistent with previous biological knowledge. Therefore, our real-time analysis can yield similar but more direct results compared to the well-established pseudotime course analysis.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study proposes a new framework to infer cellular lineage trees using somatic variants detected in\u0026nbsp;low coverage single-cell RNA sequence data. We focused on the gross phylogenetic signature of data from the single-cell rather than on each mutation.\u0026nbsp;We showed that it was possible to reconstruct a cellular lineage tree solely based on variant data, and the lineage thus obtained was consistent with the established biological knowledge.\u003c/p\u003e\n\u003cp\u003eThe significance of this study is in inferring the direct relationships between single cells using genetic scars in a cell, that is, somatic mutations. However, conventional analyses provide only indirect relationships based on similarities in expression profiles. In addition, because somatic mutations accumulate over time, our real-time course should be able to trace cellular lineages back to progeny cells, meaning that our approach has the potential to allow us to infer the cellular time course, for example, non-observable past events. In fact, our real-time course provided results more consistent than those of pseudotime-course analysis. The major difference is that the real-time course based on mutation data can infer linear and direct relationships between cells in the lineage in terms of the number of mutations. In contrast, pseudotime course analysis only can infer non-linear and indirect relationships. To compensate for the latter shortcoming, additional studies are warranted to leverage splicing kinetics, i.e., RNA velocity\u003csup\u003e41\u003c/sup\u003e. We also consider that the real-time course innately integrates the cellular trajectories with the corresponding gene expression profiles while the two datasets are separately handled. This means that their relationships can be used to provide insights into the biological system of interest.\u003c/p\u003e\n\u003cp\u003eA potential issue of our approach is that the real-time course relies on the phylogenetic signature retained in low-pass single-cell sequence data. The phylogenetic signature can be obscured by random noise due to sequencing errors as well as variant dropouts due to low coverage. In fact, as shown in \u003cstrong\u003eSupplementary figure 3\u003c/strong\u003e, the discrepant cells had relatively low coverage values (indicated by red bars).\u0026nbsp;To assess the quality of the detected somatic variants, we employed an evolutionary framework in terms of apparent selective pressure. Because somatic mutations have a negative impact on cell survival in the cell population, it is rational to assume that genuine somatic mutations are subjected to purifying selection in normal tissues\u003csup\u003e32\u003c/sup\u003e. According to the overall dN/dS values, somatic mutations were subject to purifying selection, suggesting that the majority of the detected somatic mutations were not likely to be false positives.\u003c/p\u003e\n\u003cp\u003eIt is important to identify how frequently somatic mutations occur in normal tissues. According to Tomasetti et al.\u003csup\u003e7\u003c/sup\u003e, the in vivo tissue-specific somatic mutation probability per base per cell division is \u003cimg src=\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAOgAAAAoCAYAAAD0QbbMAAAGhUlEQVR4Xu1buUotSxSt8wUiGho5BEYKTokGBg5oYqZiIgjOGCpOkQMqopkTCAaKioEgKGqggSYOiEYGDhgYOv3Be28V7EvfPlXd1ef10Tr2Lrjc4Nap3rXWXrWHqhv7578heDACjICVCMRYoFbywkYxAhIBFig7AiNgMQIsUIvJYdMYARYo+wAjYDECLFCLyWHTGAEWKPsAI2AxAixQi8lh0xgBFij7ACNgMQIs0G8gZ2trS4yMjIjHx8e4rx0cHIi+vj7x9PQkampqxPT0tCgoKPgGq/gTYSPw/v4uVldXxc3NjQDn7jEzMyOmpqbE5+en6OrqEmNjYyIjI8PTDBZo2Cw51js/Pxfj4+MCxF1fXwv3o63n52eRk5Mj9vf3RVlZmejp6ZHzVEJOopm8dAgILC8vi93dXcldbm6uODw8/GtVCLa7u1vyi1FdXS3/LCwssEBDwD+hJe7u7kRWVpa4v78XFRUVcQLFibqzsyOurq7k+hA05t3e3nIUTQjxn/sRuCsvLxfg9OTkJE6gtbW1oqioSExMTEgjKZp+fHwEFyh+PDAw4PlD0ye8iBII+9vb2wLG+BmULIgpzVxbW5NAqgZsHRoakrZiFBcXi/n5ee18U1tJeG7MQBoGnbaItJmZmeLs7CzwN91rmdrmnIf9z87OCvztjgB+61F6hxTupzgmG3+Sa51AY7GYLF/6+/ulmSht6uvr4w5tN87KFNdPoI2Njcoc2704wn5nZ6c0DPXVT9RWRBZqPAyd88PB8vLyZMpJzomUE2J1/wZg6wb26XZuU4FiTaz93QIlYS4uLsptqfag269TmKitMEwPbz/RB/33sLkm3nR2OAVHc0wFqvMJY4G+vLzE5cfkxHt7e74n/PDwsADhp6engYSJtBDRe2NjQ1tAm8zBRgEWRltbm5ibmxOTk5Na56dDCULOzs6Wv6MaMYjDqsi0WaDgtKWlRTax8vPzZQQPst+mpibR0NAg8HdJSYmy1tY5uAmPJnNs4vpbBKoDVPdx93xyyKWlJdHR0RHoIITDONM1d5cLhFVWVkqH8CuwnR8mAeqiEwp7DHeDhpzu7e3Nt+Om26hOoDjEjo+P42pQ5yFhCl4YKS6+hQgeRKBO+2DD0dGRcQT9jVzrNAI/QlPIWYOurKz4NgSNu7gUPdfX10VdXZ2n34AoOHqi3UgdcSRObFTVxvYyykugVPupHJOcLpG0k+yBEBG93c0fEi66uMAUh87X11fg+g/fSUWBwu7fxDVlI/D7i4uLvw50agqhi5uWlia79u3t7X9qUp3vGgsUToZ6zE90lBai2QJD0KWEUenp6dIBTe5+VMS9vr7KyEk1ot/9kXvDXgIloagE6hd5TQ4F9xxnjUbtd9Rv+L5Xau/1rVQV6G/hmg5yJ0duf6KDGnOgD4qmXrwaCZREt7m5KUXmNeBwzc3NUpBIQauqquRJQg0jXNCapqbO0xVpX6LihL0/IVDT9DToPK8mlWqtIE2b70xxnbYy12ovMBIoLlhRK/lFT6cQVB0uiBtROIjDODtpidRmtO3fJFAVlakcQWk/zHU8s74CpeipEpzKUUgIuhY0OrSm9RzVnIicyYyg+E5hYaGyOUKHiqnNQaNhWPNTXaDMdYIRFNETaevDw4NRF5NSXC+BmryUwcGAhwKU1iazBgU0SO2Qlrsv2YN2JsMSXNB1UlmgzLWebc8IGjR64jP0G9VjBrSa0QzxS5X9OnuJ1KJ+zR6KlM7Dg7q7QermoMIKa36qCpS59vYAT4EieuKxgVftp7qGoG4V3YPSaxOkt3SloDPL9G4sqEhJgLpGlzPFQp2MDvTo6KjMHtCFpscLYQkq7HXCECil+sBWdYhSc0rVQwBvuDqAr5j2Cphrfy/wFCgI8YseuntCRCxcxNITO0TU3t5e3xdIJi9HTOZg6zTv8vJSRm4aSJ3RWXY/yaP5uGzHgM2Dg4OBXkL5Q56cGf9HoPRggv6nBSxEul9aWipaW1v/dO5VAsUBhvfNhBntDlcMzsfhql2b8Ggy5zdz7dskSo478aqMACNgggAL1AQly+Yk8x7Usq1G3hwWaORdgAGwGQEWqM3ssG2RR4AFGnkXYABsRoAFajM7bFvkEWCBRt4FGACbEWCB2swO2xZ5BFigkXcBBsBmBFigNrPDtkUeARZo5F2AAbAZARaozeywbZFHgAUaeRdgAGxGgAVqMztsW+QRYIFG3gUYAJsR+BcB7UXjIU1OCwAAAABJRU5ErkJggg==\"\u003e (SE) in normal lymphocytes. Therefore, stochastic somatic alterations occur at an average rate of three mutations per cell division in the human genome, except for repetitive regions\u003csup\u003e42,43\u003c/sup\u003e. This abundance of somatic mutations makes retrospective cell lineage inference feasible, even in normal tissues.\u003c/p\u003e\n\u003cp\u003eIn our single-cell analysis, we need to consider three types of \u0026lsquo;variants\u0026rsquo;: (1) polymorphism of the population (mutations between individuals): i.e., germline mutations; (2) de novo mutations within the individual: i.e., somatic mutations; and (3) sequencing errors. In bulk sequencing data, we need to employ a statistical approach to handle somatic mutations; typically, we use variant allele frequency (VAF)\u0026nbsp;to infer whether a variant comes from somatic cells or is inherited from parents\u003csup\u003e44\u003c/sup\u003e. However, the distribution of RNA-VAF and DNA-VAF can differ owing to allelic imbalance (AI)\u003csup\u003e45\u003c/sup\u003e. Single-cell analysis allowed us to exclude germline mutations by simply comparing variants in a cell population.\u003c/p\u003e\n\u003cp\u003eIt remains challenging to identify the lineage in which a somatic mutation occurs (\u003cstrong\u003eFig. 1\u003c/strong\u003e). We selected \u0026lsquo;minor alleles\u0026rsquo; as potential somatic mutations, assuming that the transcriptome allele frequencies reflected the genomic allele frequencies in many genomic regions. This assumption is based on an empirical expectation that the allele frequencies of the genome and transcriptome reads are correlated\u003csup\u003e46\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eWe used homologous sites of the reference genome to fill in the missing sites in the inter-cell alignments, assuming that genomic regions that were not covered by transcripts were identical to the reference genome sequences. We used an 80% threshold as a pragmatic solution.\u003c/p\u003e\n\u003cp\u003eIn our method, homozygous sites cannot be employed to detect variants because multiallelic sites are required, in which variant sites are consequently chromosome-heterozygous in those progenies. (\u003cstrong\u003eFig.\u003c/strong\u003e \u003cstrong\u003e1\u003c/strong\u003ec and \u003cstrong\u003eFig.\u003c/strong\u003e \u003cstrong\u003e1\u003c/strong\u003ee). Thus, we missed all potential variants in the homologous sites. This limitation can also be mitigated by using biallelic sites derived from homologous sites (\u003cstrong\u003eFig.\u003c/strong\u003e \u003cstrong\u003e1\u003c/strong\u003eg) with statistical haplotype phasing\u003csup\u003e47\u003c/sup\u003e. Nevertheless, we believe that our primary objective of providing a PoC for cell lineage inference using transcriptome data has been achieved. Further analyses using other types of mutations will be conducted separately in the future.\u003c/p\u003e\n\u003cp\u003eIdeally, single-cell genome data would be more informative for inferring a cell lineage tree. We employed static and rich information to detect somatic mutations. However, there are technical and practical issues with single-cell genome analysis: We need to deal with only two copies of DNA molecules in a single cell, and the cost for sequencing the entire genome of thousands of cells is huge compared to that for RNA sequencing. Conversely, single-cell transcriptome analysis is currently a well-established technology, and a massive amount of single-cell transcriptome data has been (and will be) accumulated in sequence databases.\u0026nbsp;Genotype (mosaicism) information has not been utilized in single-cell gene expression analyses. We showed that it is possible to extract significant genotypic information from RNA sequencing data without incurring additional costs. With rigorous computational approaches (e.g., haplotype phasing), the current technique can be improved.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;Collectively, while real-time course analysis of single-cell transcriptome data has limitations in terms of zygosity and coverage at apparent variant sites, it has great potential to delineate cellular trajectories during development and to interpret high-dimensional gene expression data. Furthermore, the biological significance of somatic mutations provides insight into a new perspective of \u0026lsquo;evolution\u0026rsquo; within an individual. This approach should be highly effective, especially in human biology, in which experimental manipulation is strictly restricted.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;In conclusion, we propose a new framework for the estimation of cellular trajectories in complex tissues designated as real-time courses.\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eWe showed that it is possible to reconstruct a phylogenetic tree of somatic cells in placental tissues using only low-pass single-cell RNA sequence data. This tree is consistent with the known cell lineage of the placenta in four cell types: CYT 1, CYT 2, CYT 3, and EVT. The findings of this study suggest that robust phylogenetic signals exist in RNA sequencing data, and these signals can be used for the construction of cellular networks operating in human development, as well as in human disease pathogenesis. Our research will contribute to single-cell transcriptome analyses in alternative perspectives, for example, to analyze mosaicism of single cell genomes as a phylogenetic signature to infer cellular trajectories. The real-time course is based on mutation data and can infer linear and direct relationships between cells in the lineage in terms of the number of mutations. In contrast, pseudotime course analysis only can infer non-linear and undirected relationships. To compensate for the latter shortcoming, we need additional experiments to leverage splicing kinetics, i.e., RNA velocity\u003csup\u003e41\u003c/sup\u003e. Furthermore, the real-time course can theoretically infer the dynamics of ancestral (progeny) cells that are impossible to observe directly. In contrast, pseudotime course analysis has no capability to reconstruct the past.\u003c/p\u003e\n\u003cp\u003eWe also suggest that the real-time course innately integrates the cellular trajectories with the corresponding gene expression profiles while the two data sets are separately handled. This means that we can use their relationships to increase our insight into the biological system of interest.\u003c/p\u003e"},{"header":"References","content":"\u003cp\u003e1.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Kennedy, S. R., Loeb, L. A. \u0026amp; Herr, A. J. Somatic mutations in aging, cancer and neurodegeneration. \u003cem\u003eMech. Ageing Dev.\u003c/em\u003e \u003cstrong\u003e133\u003c/strong\u003e, 118-126 (2012).\u003c/p\u003e\n\u003cp\u003e2.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Morley, A. A. The somatic mutation theory of ageing. \u003cem\u003eMutat. Res.\u003c/em\u003e \u003cstrong\u003e338\u003c/strong\u003e, 19-23 (1995).\u003c/p\u003e\n\u003cp\u003e3.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Kelly, D. P. Cell biology: Ageing theories unified. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e470\u003c/strong\u003e, 342\u0026ndash;343 (2011).\u003c/p\u003e\n\u003cp\u003e4.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Abeliovich, A. et al. On somatic recombination in the central nervous system of transgenic mice. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e257\u003c/strong\u003e, 404-410 (1992).\u003c/p\u003e\n\u003cp\u003e5.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Yizhak, K. et al. RNA sequence analysis reveals macroscopic somatic clonal expansion across normal tissues. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e364\u003c/strong\u003e, eaaw0726 (2019).\u003c/p\u003e\n\u003cp\u003e6.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Ju, Y. S. et al. Somatic mutations reveal asymmetric cellular dynamics in the early human embryo. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e543\u003c/strong\u003e, 714-718 (2017).\u003c/p\u003e\n\u003cp\u003e7.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Tomasetti, C., Vogelstein, B. \u0026amp; Parmigiani, G. Half or more of the somatic mutations in cancers of self-renewing tissues originate prior to tumor initiation. \u003cem\u003eProc. Natl Acad. Sci. U. S. A.\u003c/em\u003e \u003cstrong\u003e110\u003c/strong\u003e, 1999\u0026ndash;2004 (2013).\u003c/p\u003e\n\u003cp\u003e8.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Lynch, M. Rate, molecular spectrum, and consequences of human mutation. \u003cem\u003eProc. Natl Acad. Sci. U. S. A.\u003c/em\u003e \u003cstrong\u003e107\u003c/strong\u003e, 961\u0026ndash;968 (2010).\u003c/p\u003e\n\u003cp\u003e9.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Jonsson, H. et al. Differences between germline genomes of monozygotic twins. \u003cem\u003eNat. Genet.\u003c/em\u003e \u003cstrong\u003e53\u003c/strong\u003e, 27-34 (2021).\u003c/p\u003e\n\u003cp\u003e10.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Garc\u0026iacute;a-Nieto, P. E., Morrison, A. J. \u0026amp; Fraser, H. B. The somatic mutation landscape of the human body. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cstrong\u003e20\u003c/strong\u003e, 298 (2019).\u003c/p\u003e\n\u003cp\u003e11.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Milholland, B. et al. Differences between germline and somatic mutation rates in humans and mice. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, 15183 (2017).\u003c/p\u003e\n\u003cp\u003e12.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Rozhok, A. I. \u0026amp; DeGregori, J. Toward an evolutionary model of cancer: Considering the mechanisms that govern the fate of somatic mutations. \u003cem\u003eProc. Natl Acad. Sci. U. S. A.\u003c/em\u003e \u003cstrong\u003e112\u003c/strong\u003e, 8914-8921 (2015).\u003c/p\u003e\n\u003cp\u003e13.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Woodworth, M. B., Girskis, K. M. \u0026amp; Walsh, C. A. Building a lineage from single cells: Genetic techniques for cell lineage tracking. \u003cem\u003eNat. Rev. Genet.\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 230\u0026ndash;244 (2017).\u003c/p\u003e\n\u003cp\u003e14.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Oota, S. Somatic mutations - Evolution within the individual. \u003cem\u003eMethods\u003c/em\u003e \u003cstrong\u003e176\u003c/strong\u003e, 91-98 (2020).\u003c/p\u003e\n\u003cp\u003e15.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Davenport, C. B. The personality, heredity and work of Charles Otis Whitman, 1843-1910. \u003cem\u003eAm. Nat.\u003c/em\u003e \u003cstrong\u003e51\u003c/strong\u003e, 5-30 (1917).\u003c/p\u003e\n\u003cp\u003e16.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;McGranahan, N. \u0026amp; Swanton, C. Clonal heterogeneity and tumor evolution: Past, present, and the future. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e168\u003c/strong\u003e, 613-628 (2017).\u003c/p\u003e\n\u003cp\u003e17.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Altrock, P. M., Liu, L. L. \u0026amp; Michor, F. The mathematics of cancer: Integrating quantitative models. \u003cem\u003eNat. Rev. Cancer\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, 730-745 (2015).\u003c/p\u003e\n\u003cp\u003e18.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Rheinbay, E. et al. Recurrent and functional regulatory mutations in breast cancer. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e547\u003c/strong\u003e, 55-60 (2017).\u003c/p\u003e\n\u003cp\u003e19.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Sun, J. X. et al. A computational approach to distinguish somatic vs. germline origin of genomic alterations from deep sequencing of cancer specimens without a matched normal. \u003cem\u003ePLOS Comput. Biol.\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, e1005965 (2018).\u003c/p\u003e\n\u003cp\u003e20.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Method of the year 2013. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, 1 (2014).\u003c/p\u003e\n\u003cp\u003e21.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Wang, Y. \u0026amp; Navin, N. E. Advances and applications of single-cell sequencing technologies. \u003cem\u003eMol. Cell\u003c/em\u003e \u003cstrong\u003e58\u003c/strong\u003e, 598-609 (2015).\u003c/p\u003e\n\u003cp\u003e22.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Tam, P. P. L. \u0026amp; Ho, J. W. K. Cellular diversity and lineage trajectory: Insights from mouse single cell transcriptomes. \u003cem\u003eDevelopment\u003c/em\u003e \u003cstrong\u003e147\u003c/strong\u003e, dev179788 (2020).\u003c/p\u003e\n\u003cp\u003e23.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Ji, Z. \u0026amp; Ji, H. TSCAN: Pseudo-time reconstruction and evaluation in single-cell RNA-seq analysis. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e44\u003c/strong\u003e, e117-e117 (2016).\u003c/p\u003e\n\u003cp\u003e24.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Campbell, K. R. \u0026amp; Yau, C. Uncovering pseudotemporal trajectories with covariates from single cell and bulk expression data. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 2442-2442 (2018).\u003c/p\u003e\n\u003cp\u003e25.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Felsenstein, J. The number of evolutionary trees. \u003cem\u003eSyst. Biol.\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 27-33 (1978).\u003c/p\u003e\n\u003cp\u003e26.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Hinton, G. \u0026amp; Roweis, S. Stochastic neighbor embedding. \u003cem\u003eAdv. Neural Inf. Process. Syst.\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, 833--840 (2003).\u003c/p\u003e\n\u003cp\u003e27.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;McInnes, L., Healy, J. \u0026amp; Melville, J. UMAP: Uniform manifold approximation and projection for dimension reduction, \u003cem\u003eArXiv\u003c/em\u003e. \u003cem\u003e1802.03426 (2020)\u003c/em\u003e.\u003c/p\u003e\n\u003cp\u003e28.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Kobak, D. \u0026amp; Linderman, G. C. Initialization is critical for preserving global data structure in both t-SNE and UMAP. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, 156-157 (2021).\u003c/p\u003e\n\u003cp\u003e29.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Schneider, V. A. et al. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. \u003cem\u003ebioRxiv\u003c/em\u003e, 072116 (2016).\u003c/p\u003e\n\u003cp\u003e30.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Pavličev, M. et al. Single-cell transcriptomics of the human placenta: Inferring the cell communication network of the maternal-fetal interface. \u003cem\u003eGenome Res.\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 349-361 (2017).\u003c/p\u003e\n\u003cp\u003e31.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Kimura, M. The number of heterozygous nucleotide sites maintained in a finite population due to steady flux of mutations. \u003cem\u003eGenetics\u003c/em\u003e \u003cstrong\u003e61\u003c/strong\u003e, 893-903 (1969).\u003c/p\u003e\n\u003cp\u003e32.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Yadav, V. K., DeGregori, J. \u0026amp; De, S. The landscape of somatic mutations in protein coding genes in apparently benign human tissues carries signatures of relaxed purifying selection. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e44\u003c/strong\u003e, 2075-2084 (2016).\u003c/p\u003e\n\u003cp\u003e33.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Fitch, W. M. Toward defining the course of evolution: Minimum change for a specific tree topology. \u003cem\u003eSyst. Zool.\u003c/em\u003e \u003cstrong\u003e20\u003c/strong\u003e, 406-416 (1971).\u003c/p\u003e\n\u003cp\u003e34.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Tamura, K., Stecher, G. \u0026amp; Kumar, S.. MEGA11: Molecular Evolutionary Genetics Analysis Version 11. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e version 11 \u003cstrong\u003e38\u003c/strong\u003e, 3022-3027 (2021).\u003c/p\u003e\n\u003cp\u003e35.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Long, C. A., Sokal, R. R. \u0026amp; Sneath, P. H. A. Principles of numerical taxonomy. \u003cem\u003eJ. Mammal.\u003c/em\u003e\u003cstrong\u003e46\u003c/strong\u003e, 111-112 (1965).\u003c/p\u003e\n\u003cp\u003e36.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Kn\u0026ouml;fler, M. et al. Human placenta and trophoblast development: Key molecular mechanisms and model systems. \u003cem\u003eCell. Mol. Life Sci.\u003c/em\u003e \u003cstrong\u003e76\u003c/strong\u003e, 3479-3496 (2019).\u003c/p\u003e\n\u003cp\u003e37.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Trapnell, C. et al. The dynamics and regulators of cell fate decisions are revealed by pseudotemporal ordering of single cells. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e \u003cstrong\u003e32\u003c/strong\u003e, 381-386 (2014).\u003c/p\u003e\n\u003cp\u003e38.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Qiu, X. et al. Reversed graph embedding resolves complex single-cell trajectories. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 979-982 (2017).\u003c/p\u003e\n\u003cp\u003e39.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Cao, J. et al. The single-cell transcriptional landscape of mammalian organogenesis. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e566\u003c/strong\u003e, 496-502 (2019).\u003c/p\u003e\n\u003cp\u003e40.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Wolfram Research \u0026amp; I. 12.2 version (Wolfram Research, Inc. (Champaign, IL, 2020)).\u003c/p\u003e\n\u003cp\u003e41.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;La Manno, G. et al. RNA velocity of single cells. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e560\u003c/strong\u003e, 494-498 (2018).\u003c/p\u003e\n\u003cp\u003e42.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Ortega, M. A. et al. Using single-cell multiple omics approaches to resolve tumor heterogeneity. \u003cem\u003eClin. Transl. Med.\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, 46 (2017).\u003c/p\u003e\n\u003cp\u003e43.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Araten, D. J. et al. A quantitative measurement of the human somatic mutation rate. \u003cem\u003eCancer Res.\u003c/em\u003e \u003cstrong\u003e65\u003c/strong\u003e, 8111-8117 (2005).\u003c/p\u003e\n\u003cp\u003e44.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Dou, Y., Gold, H. D., Luquette, L. J. \u0026amp; Park, P. J. Detecting somatic mutations in normal cells. \u003cem\u003eTrends Genet.\u003c/em\u003e \u003cstrong\u003e34\u003c/strong\u003e, 545-557 (2018).\u003c/p\u003e\n\u003cp\u003e45.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Rhee, J. K., Lee, S., Park, W. Y., Kim, Y. H. \u0026amp; Kim, T. M. Allelic imbalance of somatic mutations in cancer genomes and transcriptomes. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cstrong\u003e7\u003c/strong\u003e, 1653 (2017).\u003c/p\u003e\n\u003cp\u003e46.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Ju, Y. S. et al. Extensive genomic and transcriptional diversity identified through massively parallel DNA and RNA sequencing of eighteen Korean individuals. \u003cem\u003eNat. Genet.\u003c/em\u003e \u003cstrong\u003e43\u003c/strong\u003e, 745-752 (2011).\u003c/p\u003e\n\u003cp\u003e47.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Browning, S. R. \u0026amp; Browning, B. L. Haplotype phasing: Existing methods and new developments. \u003cem\u003eNat. Rev. Genet.\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 703-714 (2011).\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003eSingle-cell RNA-seq.\u0026nbsp;\u003c/strong\u003eIn this study, we used two public transcriptome datasets obtained from normal placental tissues\u003csup\u003e30\u003c/sup\u003e: SRP090944 Batch1 (54 cells) and Batch2 (33 cells).\u0026nbsp;Pavličev et al.\u003csup\u003e30\u003c/sup\u003e analyzed the placenta data in the context of the cell communication network between two semiallogenic individuals, the mother and the fetus.\u0026nbsp;The authors inferred the cell-cell interactome according to the gene expression of receptor-ligand pairs across cell types and found cell-type-specific expression of G-protein-coupled receptors, suggesting that ligand-receptor profiles could be a reliable tool for cell type identification. The data are registered in the DDBJ Sequence Read Archive as SRR4371525 (GSM2339239)\u0026ndash;SRS1732319 (SRX2225328) (\u003cstrong\u003eSupplementary tables\u003c/strong\u003e \u003cstrong\u003e3\u003c/strong\u003e and\u0026nbsp;\u003cstrong\u003e4\u003c/strong\u003e). We assumed that Batch1 and Batch 2 shared no \u003cem\u003ede novo\u003c/em\u003e somatic mutations.\u003c/p\u003e\n\u003cp\u003ePavličev et al. categorized the single-cell data into five clusters (cell types) according to the principal comportment analysis (PCA) on the gene expression profile and 300 marker genes: cytotrophoblast (CYT) 1, CYT2, and CYT3, extravillous trophoblast (EVT), and the putative maternal decidual cell (DC).\u0026nbsp;\u003cstrong\u003eSupplementary figure 1\u003c/strong\u003e shows a\u0026nbsp;model system that integrates the role of Notch/Wnt signaling in trophoblast stemness and extravillous trophoblast differentiation\u003csup\u003e36\u003c/sup\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMapping of sequence data.\u0026nbsp;\u003c/strong\u003eWe intended to\u0026nbsp;construct multiple alignment data that represent single cells by aligning the homologous regions of their transcriptomes. To this end, read sequences (segmented sequences of base pairs corresponding to part of a single DNA fragment, a set of which is expected to cover partial coding regions) were assigned designated coordinates of the reference genome. The entire data analysis pipeline is shown in\u0026nbsp;\u003cstrong\u003eSupplementary figure 4\u003c/strong\u003e.\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eWe mapped single-cell transcriptome sequence data to the human genome (GRCh38)\u003csup\u003e29\u003c/sup\u003e using the Burrows-Wheeler Aligner\u003csup\u003e48\u003c/sup\u003e after removing adaptor sequences with trimmomatic\u003csup\u003e49\u003c/sup\u003e. We identified variants\u0026nbsp;by comparing the aligned read sequences with the reference genome using Samtools\u003csup\u003e50\u003c/sup\u003e. In the present study, we used only single-nucleotide variants, excluding all detected indel events. If we encountered incomplete sites across cells by more than 50%, we excluded the corresponding sites. We annotated the detected variant sites using the SnpEff software\u003csup\u003e51\u003c/sup\u003e, which can predict the effects of genetic variants on genes and transcripts.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePhylogenetic analysis of single cells.\u0026nbsp;\u003c/strong\u003eWe concatenated all acquired variant sites to generate sequence alignments. We reconstructed\u0026nbsp;cellular lineage trees using the maximum parsimony method\u003csup\u003e33\u003c/sup\u003e implemented in MEGA X\u003csup\u003e34\u003c/sup\u003e with default parameters. The results were output in the Newick tree format\u003csup\u003e52\u003c/sup\u003e for the following process: dN/dS ratio computations with the maximum likelihood method and comparative analysis between the mutation patterns and the gene expression profile.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003edN/dS ratios in the coding regions.\u0026nbsp;\u003c/strong\u003eWe assembled codon sequences into which the detected variants fell and created codon alignments with exonic variants. We computed the overall dN/dS ratios using the Codeml package\u003csup\u003e53\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eGene expression reanalysis.\u0026nbsp;\u003c/strong\u003eFor linear dimensionality reduction, we conducted principal component analysis (PCA) on the expression patterns using R (version 3.6.2)\u003csup\u003e54\u003c/sup\u003e, and subsequently applied t-SNE\u003csup\u003e26\u003c/sup\u003eand UMAP\u003csup\u003e27\u003c/sup\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eto the data for nonlinear dimensionality reduction. The Louvain method was used for the clustering\u003csup\u003e55\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eThe pseudotime course analysis.\u0026nbsp;\u003c/strong\u003eWe conducted a pseudotime course analysis using monocle3 (version 1.0.0)\u003csup\u003e37-39\u003c/sup\u003e on R (version 4.1.2)\u003csup\u003e54\u003c/sup\u003e. Pseudotime course analysis estimates the temporal dimension of static single-cell transcriptome data and infers individual gene expression dynamics along with cellular changes. We reduced the data dimension to 26 using PCA. To visualize the results, we used \u003cem\u003eMathematica\u003c/em\u003e\u003csup\u003e40\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eComparative analysis of mutation patterns and gene expression profiles.\u0026nbsp;\u003c/strong\u003eWe mapped the clustered cells to the cellular lineage trees, reconstructed based on cellular genotypes. To this end, we developed a \u003cem\u003eMathematica\u003c/em\u003e\u003csup\u003e40\u003c/sup\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003ecode, \u003cem\u003eAssignCluster2Cell\u003c/em\u003e, using two packages: Phylogenetics for \u003cem\u003eMathematica\u003c/em\u003e\u003csup\u003e56\u003c/sup\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eand Phylogenetics\u003csup\u003e57\u003c/sup\u003e. \u003cem\u003eAssignCluster2Cell\u0026nbsp;\u003c/em\u003ereads a Newick\u003csup\u003e52\u003c/sup\u003e formatted tree file and a cluster table of a gene expression profile, and maps cluster ID to the tree topology, evaluating degrees of monophyleticity (DoM) at each node in terms of the proportion of the clusters in its subtree (\u003cstrong\u003eFig. 6\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eWe also mapped the clustered cells to the tree to evaluate the DoMs. The pie charts in \u003cstrong\u003eFig. 6\u003c/strong\u003e show the DoM at each node; for example, if the pie chart at a node has only one slice, the subtree is completely monophyletic; if the pie chart at a node has two slices, the subtree is polyphyletic (subtree B in \u003cstrong\u003eFig. 6\u003c/strong\u003e). However, because cell type 2 is dominant in subtree B, we can infer that cell type 1 is derived from cell type 2. In this manner, we determined the relevance of cell types based on gene expression profiles and cell lineages.\u003c/p\u003e\n\u003cp\u003eTheoretically, the root of a cellular lineage tree represents a zygotic cell and the root represents the zygotic genome of the fertilized egg. However, in the present study, the root of an estimated tree represented a progenitor cell of the observed cells. The zygotic cell should reside somewhere between the root and the reference genome.\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eReferences\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e\u0026nbsp;\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e48.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Li, H. \u0026amp; Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 1754-1760 (2009). 10.1093/bioinformatics/btp324, Pubmed:19451168.\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e49.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Bolger, A. M., Lohse, M. \u0026amp; Usadel, B. Trimmomatic: A flexible trimmer for Illumina sequence data. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e30\u003c/strong\u003e, 2114-2120 (2014). 10.1093/bioinformatics/btu170, Pubmed:24695404.\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e50.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Li, H. et al. The Sequence Alignment/Map format and SAMtools. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 2078-2079 (2009). 10.1093/bioinformatics/btp352, Pubmed:19505943.\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e51.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Cingolani, P. et al. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. \u003cem\u003eFly (Austin)\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, 80-92 (2012). 10.4161/fly.19695, Pubmed:22728672.\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e52.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Felsenstein, J. \u0026amp; PHYLIP. \u003cem\u003ePhylogeny Inference Package\u003c/em\u003e. 3.2 version. \u003cem\u003eCladistics\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, Vols. 164-166, (1989).\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e53.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Yang, Z. PAML 4: Phylogenetic analysis by maximum likelihood. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e \u003cstrong\u003e24\u003c/strong\u003e, 1586-1591 (2007). 10.1093/molbev/msm088, Pubmed:17483113.\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e54.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;R Core Team ((Vienna, Austria, 2016)).\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e55.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Blondel, V. D., Guillaume, J., Lambiotte, R. \u0026amp; Lefebvre, E. Fast unfolding of communities in large networks. \u003cem\u003eJ. Stat. Mech. Theor. Exp.\u003c/em\u003e \u003cstrong\u003e2008\u003c/strong\u003e, 10008 (2008). 10.1088/1742-5468/2008/10/P10008.\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e56.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Polly, P. D. (Indiana University Bloomington. \u003cem\u003eIndiana\u003c/em\u003e).\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e57.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Zachar, I. GitHub repository. in \u003cem\u003eed686781c686980dcd686982ff111438bb686935a686986c686983a686951f686982b686984 (GitHub, 2017)\u003c/em\u003e, Vol. 2021. 686980.\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e\u003cspan lang=\"\"\u003e58.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Cock, P. J. A., Fields, C. J., Goto, N., Heuer, M. L. \u0026amp; Rice, P. M. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e38\u003c/strong\u003e, 1767-1771 (2010). 10.1093/nar/gkp1137, Pubmed:20015970.\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003e59.\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Danecek, P. et al. The variant call format and VCFtools. \u003cem\u003eBioinformatics\u003c/em\u003e\u003cstrong\u003e27\u003c/strong\u003e, 2156-2158 (2011). 10.1093/bioinformatics/btr330, Pubmed:21653522. \u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData Availability Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data that support the findings of this study are divided into two sets, SRP090944 Batch1 (54 cells) and Batch2 (33 cells). The both data sets are available in the DDBJ Sequence Read Archive (https://www.ddbj.nig.ac.jp/dra/index-e.html) as SRR4371525 (GSM2339239)\u0026ndash;SRS1732319 (SRX2225328).\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-1558242/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1558242/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSingle-cell RNA-seq (scRNA-seq) analysis can describe the states of individual cells but cannot clarify direct relationships between cells. Therefore, we developed a method to trace the progenitor cells of cell differentiation pathways through the cell lineage using somatic mutations detected in low-pass scRNA sequences. Using two human placental tissue datasets, we detected 2,172 and 627 somatic mutation candidates. We used only somatic mutation data to classify placental cells and construct a phylogenetic tree; the results were consistent with those of cell classification by expression profiling and estimation of the differentiation process by pseudotime course analysis. Somatic-mutation-based analysis can determine the temporal order of individual cells in the cell lineage, which is not possible with pseudotime analysis. Therefore, somatic mutation signatures in scRNA-seq data are informative for constructing cell lineages, and in combination with expression data obtained simultaneously, they are extremely useful for delineating the processes of human development and pathogenesis.\u003c/p\u003e","manuscriptTitle":"Real-Time Course: Reconstruction of cellular diversity and lineage trajectory based on somatic mutational patterns detected from low-pass single-cell transcriptome data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-04-14 18:23:58","doi":"10.21203/rs.3.rs-1558242/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"09c25e0e-2ce4-43e0-9a80-7c90c10fa6f3","owner":[],"postedDate":"April 14th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2022-04-30T19:25:42+00:00","versionOfRecord":[],"versionCreatedAt":"2022-04-14 18:23:58","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1558242","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1558242","identity":"rs-1558242","version":["v1"]},"buildId":"369fNeqWncA4NS6XSWjrt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: preprint-html ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-4.0