Inferring Sex, Ethnicity, and Age from RNA-seq Data

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract RNA sequencing provides a comprehensive snapshot of gene expression, reflecting genetic inheritance and dynamic environmental influences. This study explores the predictive power of RNA-seq data combined with advanced machine learning techniques, such as Gradient Boosting Machines, Support Vector Regression, and SHapley Additive exPlanations, to infer complex human traits, including biological sex, age, and ethnicity, across diverse tissues. Using RNA-seq datasets derived from blood, heart, and several brain regions, we achieved near-perfect accuracy in sex determination, emphasizing the critical roles of sex chromosome-linked genes (XIST, KDM5D, EIF1AY). Age prediction demonstrated high tissue-specific precision, identifying transcripts indicative of biological aging, particularly those involved in DNA repair and inflammation, which offer promising biomarkers for aging-related diseases and research. Ethnicity prediction from RNA-seq effectively distinguished closely related populations (e.g., British vs. Utah residents of Northern European descent), surpassing SNP-based approaches by capturing rapid, environment-driven transcriptional adaptations in immune-related genes (IL2RA, FOXO4). Integrating RNA-seq with genomic data further enhanced prediction accuracy, revealing nuanced population-specific transcriptomic signatures shaped by genetic ancestry and environmental factors. Our findings underscore RNA-seq's significant potential for precision medicine, highlighting critical biomarkers and pathways that may guide personalized healthcare, anti-aging strategies, disease risk assessment, and targeted therapeutic interventions.
Full text 229,936 characters · extracted from preprint-html · click to expand
Inferring Sex, Ethnicity, and Age from RNA-seq Data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Inferring Sex, Ethnicity, and Age from RNA-seq Data Tatiana Tatarinova, Arseniy Dokuchaev, Varvara Pozdina, Sergey Gaponov, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7256409/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract RNA sequencing provides a comprehensive snapshot of gene expression, reflecting genetic inheritance and dynamic environmental influences. This study explores the predictive power of RNA-seq data combined with advanced machine learning techniques, such as Gradient Boosting Machines, Support Vector Regression, and SHapley Additive exPlanations, to infer complex human traits, including biological sex, age, and ethnicity, across diverse tissues. Using RNA-seq datasets derived from blood, heart, and several brain regions, we achieved near-perfect accuracy in sex determination, emphasizing the critical roles of sex chromosome-linked genes (XIST, KDM5D, EIF1AY). Age prediction demonstrated high tissue-specific precision, identifying transcripts indicative of biological aging, particularly those involved in DNA repair and inflammation, which offer promising biomarkers for aging-related diseases and research. Ethnicity prediction from RNA-seq effectively distinguished closely related populations (e.g., British vs. Utah residents of Northern European descent), surpassing SNP-based approaches by capturing rapid, environment-driven transcriptional adaptations in immune-related genes (IL2RA, FOXO4). Integrating RNA-seq with genomic data further enhanced prediction accuracy, revealing nuanced population-specific transcriptomic signatures shaped by genetic ancestry and environmental factors. Our findings underscore RNA-seq's significant potential for precision medicine, highlighting critical biomarkers and pathways that may guide personalized healthcare, anti-aging strategies, disease risk assessment, and targeted therapeutic interventions. Biological sciences/Computational biology and bioinformatics/Data mining Biological sciences/Computational biology and bioinformatics/Computational models Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Introduction RNA sequencing (RNA-Seq) provides a high-resolution snapshot of tissue- and cell-specific gene expression in diverse biological contexts. Prior research has linked specific transcriptomic signatures to phenotypic traits such as biological sex, aging, and ethnicity, underscoring the potential of RNA-Seq for precision medicine, population genetics, and aging research. However, significant gaps remain in our understanding of how consistently RNA-Seq can predict these traits across different human tissues and diverse populations, as well as how genetic and environmental influences interact to shape these transcriptomic patterns. Addressing these gaps is critical, as tissue-specific variations and environmental factors such as personalized healthcare, disease prevention strategies, and targeted interventions could profoundly impact clinical applications. In this study, we systematically explore the ML predictive capabilities of RNA-Seq data analysis across tissues (blood, heart, and several brain regions) to robustly infer sex, biological age, and ethnicity, thereby illuminating underlying biological pathways and advancing the clinical relevance of transcriptomics. Inferring Sex Using RNA-Seq Data Biological sex prediction using RNA-Seq data is highly effective, primarily due to distinct gene expression differences driven by the presence of sex chromosomes (X and Y) and hormonal factors. Sex chromosome-linked genes, such as XIST (crucial for female X-chromosome inactivation), and male-specific Y chromosome genes including RPS4Y1, KDM5D, and DDX3Y, exhibit prominent sex-specific expression patterns. Hormonal regulation, particularly involving estrogen and testosterone, further amplifies these sex-specific differences in gene expression, influencing diverse biological processes such as metabolism, immune function, and reproductive health. 1 , 2 For example, sex hormones influence gene expression in the liver, leading to differences in growth hormone signaling and differentially impacting drug metabolism in males and females. 1 , 3 Distinct gene expression signatures between males and females provide a basis for accurately applying machine learning (ML) techniques to infer biological sex using RNA-Seq data. Machine learning models such as random forests (RF) and support vector machines (SVMs) are commonly employed to predict sex using RNA-Seq data. These models are trained on datasets that include sex-labeled samples, enabling them to identify differential expression patterns in key sex-linked and hormone-regulated genes. Tissue specificity is also important; it has been reported that the liver and brain show pronounced sex-biased RNA expression patterns. 4 Sex-specific transcriptomic signatures are clinically significant, influencing drug metabolism, immune responses, and disease susceptibility, reinforcing the role of RNA-Seq in personalized medicine. For instance, females and males often respond differently to cancer immunotherapy; understanding these differences at the transcriptomic level can help optimize treatment strategies. 5 , 6 ML-empowered Inference of Age from RNA-Seq Data Aging is a multifaceted biological phenomenon shaped by genetic predispositions and environmental influences. Distinguishing between chronological age (actual lived time) and biological age (physiological health status) is essential, as biological age more accurately reflects individual health conditions and risk factors for age-related diseases. RNA-Seq data captures age-related transcriptional changes, including genes involved in cellular stress responses, inflammation, and mitochondrial function. Genes such as SIRT1, which decline in expression with age, and inflammatory markers like IL-6, which increase, serve as potential biomarkers of biological aging. Biological age can be inferred from epigenetic markers that indicate cellular damage or regenerative capacity level, offering insights into the individual's biological health and aging rate. 7 Age-related changes in the transcriptome and epigenome have been documented across species and tissues, showing that gene expression can serve as a molecular clock to estimate an individual's age. 8 , 9 Studies have shown that the transcription of genes involved in cellular stress responses, inflammation, and mitochondrial function is age-dependent. For example, sirtuin family members, such as SIRT1, exhibit decreased expression with age, which reduces DNA repair capacity and increases oxidative stress. 10 In contrast, inflammatory genes, such as IL-6, become upregulated with age, contributing to the chronic, low-grade inflammation associated with aging. 11 Specific tissues show distinct aging profiles. For instance, skeletal muscle and the brain exhibit significant changes in mitochondrial and metabolic genes, reflecting the functional decline with age. 12 Conversely, the liver maintains more stable gene expression patterns across the lifespan, reflecting its regenerative capacity. 13 Age prediction models utilize regression-based ML learning algorithms, including support vector regression (SVR) and deep learning models. These models are trained on RNA-Seq datasets that include individuals of varying ages. By identifying patterns of gene expression that correlate with age, the models can predict an individual's age with remarkable precision. Fleischer et al. 10 demonstrated that SVR models trained on fibroblast RNA-Seq data can predict age with a mean error of less than five years. The ability to predict tissue-specific age from RNA-Seq data has far-reaching implications by identifying biomarkers and expression patterns that define biological, rather than chronological, age. This distinction is critical in studies of longevity and age-related diseases, where individuals with "younger" transcriptomic profiles may exhibit lower disease risk. Tissue-specific gene expression patterns observed in tissues like skeletal muscle, brain, and liver provide additional depth to age prediction, identifying individuals at higher risk for diseases like Alzheimer's and cardiovascular disorders, and enabling timely preventive and therapeutic interventions. 14 Inferring Ethnicity Using RNA-Seq Data Inference of ethnicity from RNA-Seq data presents significant complexity due to the intricate interplay between genetic ancestry and diverse environmental and physiological factors, including diet, microbiome, pathogen exposure, and lifestyle. Several studies have demonstrated that population-specific differences in gene expression can provide valuable insights into an individual's ethnic background, highlighting the potential of RNA-Seq in this area. 15 Ethnicity-related differences in gene expression often arise from population-specific genetic variants affecting regulatory pathways, such as immune response genes CYP3A4 and IFN-γ. 16 17 Additionally, environmentally driven epigenetic modifications, such as DNA methylation and histone acetylation, significantly contribute to ethnic transcriptomic differences. Integrative methodologies that combine RNA-Seq with genomic data (e.g., SNP arrays, whole-genome sequencing) substantially enhance the accuracy of ethnicity prediction, with critical implications for pharmacogenomics. Understanding these ethnicity-specific transcriptomic variations enables more precise medication dosing, reducing adverse drug reactions and improving patient outcomes across diverse populations. One of the primary challenges in inferring ethnicity is distinguishing between genetic and environmental influences. For example, differences in gene expression related to immune response could be driven by genetic ancestry or environmental factors, such as pathogen exposure and microbiome. Multi-omics approaches that integrate RNA-Seq data with SNP arrays or whole-genome sequencing have been employed to overcome this challenge. 17 These integrative approaches enhance the accuracy of ethnicity prediction by considering genetic and environmental factors that contribute to gene expression. 18 Ethnicity determination is crucial for precision medicine, particularly in pharmacogenomics. Different ethnic groups often exhibit varying responses to medications due to differences in gene expression related to drug metabolism. For example, individuals of African descent tend to express higher levels of CYP3A4, an enzyme involved in the metabolism of many drugs, compared to those of European or Asian descent. 17 Understanding these differences will enable accurate dosing recommendations and reduce the risks of adverse drug reactions. In summary, RNA-Seq offers robust predictive potential for determining biological sex, age, and ethnicity, providing valuable insights into gene expression patterns influenced by genetic and environmental factors. This study aims to rigorously evaluate RNA-Seq’s predictive capabilities across multiple tissues (blood, heart, and several brain regions), advancing both our biological understanding and the practical application of transcriptomics in personalized healthcare. Results and Discussion Sex determination We analyzed four datasets: BLOOD (329 male and 338 female samples), HEART (9 male and 11 female samples), BRAIN0 (25 male and 9 female samples), and BRAIN1 (150 male and 65 female samples); see Table 1 for details. The largest dataset, BLOOD, served as the training set for developing the sex determination models based on the XGBoost algorithm 19 , 20 , as detailed in the Materials and Methods section. We have tested models with and without sex chromosomes. To identify the most influential transcripts, we utilized the game-based machine learning method SHAP (SHAPLEY Additive exPlanations) 21 . Models derived from the BLOOD dataset were tested on the BLOOD, HEART, BRAIN0, and BRAIN1 datasets using a five-fold cross-validation approach. Performance metrics (ROC AUC, accuracy, precision, recall, and F1 scores) obtained using varying numbers of influential transcripts (1-100) are presented in Figs. 1 – 4 . The number of expressed transcripts for all datasets is in Table 2 . The number of important features for each model is in Table 3 . Supplemental Tables 1–4 describe transcripts used for sex determination. Table 1 Training and Testing RNA-Seq Datasets Project id Dataset name Ethnicity information Age information Organ/tissue Number of samples Male Female ERP001941 BLOOD Provided Not provided Blood (lymphoblastoid cell lines from the 1000 Genomes) 329 338 PRJNA877032 HEART Not provided Provided Heart (atrium and ventricle, left and right) 9 11 PRJNA448069 BRAIN0 Not provided Provided Brain (Frontal cortex) 25 9 PRJNA836496 BRAIN1 Not provided Provided Brain (Putamen, Caudate, Nucleus accumbens) 150 65 Table 2 Number of transcripts in the datasets after filtering. Chr_aXY autosomes Chr_aX Chr_aY BLOOD 15771 15354 15725 15368 HEART 8668 8441 8660 8441 BRAIN0 9703 9448 9671 9460 BRAIN1 9429 9162 9400 9175 Table 3 Number of important transcripts for different tissues, selected based on the SHAP scores Dataset BRAIN0 BRAIN1 HEART BLOOD Chr_aXY 2 2 3 1 Chr_aX 1 1 2 1 Chr_aY 1 1 1 1 Autosomes 46 29 22 40 The following sections discuss the performance outcomes for each dataset individually. BLOOD Models trained with datasets containing sex chromosomes (chr_aXY, chr_aX, and chr_aY) required only ten transcripts to reach nearly 100% accuracy in sex determination (Fig. 1 ). Adding extra features did not affect the model's performance. Among the forty transcripts identified as the most important predictors (Table 3 ), only two (ENST00000647913.2 and MSTRG.36782.15) were on sex chromosomes. Of the top 40 transcripts chosen for sex classification using autosomal genes alone, seven were related to immune functions, while twelve were associated with transcriptional and post-transcriptional regulation (Table S1 ). When we used all chromosomes (dataset chr_aXY), the model correctly identified 335 of 338 females (99%) and 328 of 329 males (99.9%). Using the X chromosome and autosomes (dataset chr_aX), the model accuracy was similar: the model correctly classified 336 females (99%) and 327 males (99%). The model based on the Y chromosome and autosomes (dataset chr_aY) resulted in the accurate identification of 335 females (99%) and 328 males (99.9%). However, inference relying solely on autosomal data yielded significantly lower accuracy, correctly classifying only 246 females (73%) and 240 males (73%) (Figure S1 ). These findings suggest that errors in sex identification probably stem from overlapping expression patterns of the autosomal transcripts between sexes. Therefore, including additional informative predictors may improve model accuracy, especially when the differences in transcript expression between males and females are minor. Since not all differentially expressed transcripts were used to build our model, we were curious to see which genes and gene families show the highest differential expression levels between sexes. We analyzed sex-biased gene expression by ranking differentially expressed genes based on False Discovery Rate (FDR)-corrected t-test p-values. Among the 786 identified sex-biased transcripts, male-specific Y-chromosomal genes ranked highest. They were followed by female-biased genes, mainly located on the X chromosome. Analysis of autosomal features revealed female-biased expression of several pseudogenes (RPS4XP13, ENSG00000253899, RPS4XP2, RPS4XP11, RPS4XP6, RPS4XP3, RPS4XP17) and three protein-coding genes (PPFIA3, EIF2S3B, LMTK3). Top male-biased autosomal loci included LINC01597, ENSG00000291100, MARVELD1, PLEK, IMPACT, ENSG00000280435, ZPBP2, PHETA2, RAB38, AKIRIN1, NQO1, and DDX43. Additionally, several genes showed isoform-level sex biases, notably MSL3, RBM4, and BLOC1S2. Next, we assessed whether a model trained on blood RNA-seq data could predict sex based on gene expression in other tissues. Of the 786 transcripts identified as differentially expressed between sexes in blood, 574 (73%) showed the same direction of sex-biased expression in the frontal cortex (BRAIN0 dataset). Among these, 249 transcripts were significantly differentially expressed in both tissues, indicating consistent sex bias. Conversely, 140 transcripts exhibited discordant sex-biased expression between tissues, including seven transcripts with significant opposite biases. Notably, FRMD4A and a novel transcript overlapping ENSG00000286388 showed male-biased expression in blood but female-biased expression in the frontal cortex. Conversely, genes such as PRDX2, ENSG00000285756, ENSG00000262202, MSTRG.29194, and MSTRG.35408.4 displayed female bias in blood but male bias in the frontal cortex. These results highlight the complexity of sex-biased gene regulation across tissues and emphasize the importance of tissue-specific analyses to accurately interpret sex differences in gene expression (see Supplementary Material Table 1 for detailed gene annotations and references). Sex Determination in Frontal Cortex (BRAIN0 Dataset) We evaluated the performance of the sex determination model (trained on BLOOD RNA-Seq data) using the frontal cortex gene expression dataset with 25 men and 9 women. The model achieved perfect accuracy (1.0) for datasets containing all chromosomes (chr_aXY) using only two transcripts (MSTRG.36782.15, ENST00000611750.1), the X chromosome plus autosomes (chr_aX, one transcript - MSTRG.36020.14), and the Y chromosome plus autosomes (chr_aY, one transcript - MSTRG.36782.15), see Fig. 2 . Overall, the model demonstrated consistently high performance using all chromosomes, attaining near-perfect accuracy. In contrast, predictions based only on autosomal transcripts were less accurate and more variable. For autosomal predictions, performance metrics (ROC AUC, accuracy, F1 score, precision, and recall) were around 0.6 for models with one transcript, gradually increasing with the number of features, and plateauing at approximately 0.76 for 46 transcripts. For the three models with sex chromosomes, the sex of all 25 men and nine women was determined correctly. The model was based solely on autosomes, which identified the sex of 23 men and three women correctly but misclassified two men and six women (see Figure S3). Sex Classification Model Performance using the HEART Dataset When tested on the HEART dataset (comprising 9 men and 11 women), the BLOOD-trained model demonstrated excellent accuracy. Specifically, the model achieved perfect accuracy (1.0) for datasets involving both sex chromosomes (chr_aXY) using two transcripts (ENST00000647913.2, MSTRG.36782.15), the X chromosome plus autosomes (chr_aX, one or two transcripts - ENST00000647913.2), and the Y chromosome plus autosomes (chr_aY, one transcript - MSTRG.36782.15). The prediction accuracy for the autosome-only dataset was significantly lower (0.32) and required 22 transcripts (see Supplemental Table S3, Fig. 3 ). The sex of only four out of nine males (44%) and two out of eleven females (18%) was correctly predicted, resulting in a high misclassification rate (Figure S2). When considering the heart chamber (left/right atrium/ventricle), we observe that the model's accuracy heavily depends on this covariate. For example, all female samples are misclassified except for the two from the left ventricle (Fig. 5 ). In males, the situation is reversed: three left ventricular samples are misclassified, suggesting that in the left ventricle, gene expression in males and females resembles female gene expression in whole blood samples, while other chambers exhibit specific patterns. Sex Classification Performance in the Putamen, Caudate, and Nucleus accumbens (BRAIN1) Dataset The tested BRAIN1 dataset included 150 men and 65 women. The predictive performance of the sex determination model, trained on the BLOOD dataset, was similarly accurate when tested on the BRAIN1 dataset (Fig. 4 ). The highest accuracy (1.0) was achieved for chr_aXY using three transcripts (ENST00000361365.7, ENST00000602495.1, MSTRG.36020.14), chr_aX (two transcripts ENST00000647913.2, MSTRG.36020.14), and chr_aY (one transcript ENST00000361365.7). Predictions relying solely on autosomal transcripts showed reduced accuracy of 0.73 and required 29 transcripts (see the Transcript Description Table S4 in Supplemental Materials). For datasets chr_aXY and chr_aX, the model correctly identified all 65 women and 148 out of 150 men, misclassifying two men as women. The dataset chr_aY correctly identified all 65 women and 147 out of 150 men, with three men misclassified as women. Predictions based solely on autosomes resulted in the highest misclassification rate: only 132 men and eight women were correctly identified, while 18 men and 57 women were incorrectly classified (see Figure S6). The accuracy of the autosomal model prediction for the BRAIN1 dataset (Putamen, Caudate, Nucleus accumbens) was slightly lower than for the BRAIN0 dataset (Frontal cortex), at 0.65 versus 0.76 (p-value = 0.096). This difference is marginally significant; larger sample sizes are needed to examine variation in sex determination across brain regions. If the effect is real, it may be due to the influence of ovarian hormones such as 17β-estradiol (E2), which actively regulate the dopamine system by boosting dopamine cell activity and neurotransmitter release. 22 It is important to note that sex differences have been observed in the prefrontal cortex; for example, females show higher levels of synaptosomal GluA1 and GluA2 subunits in the medial prefrontal cortex, as the striatum appears to be especially sensitive to these hormonal effects. 23 Comparison of Models and Transcript Expression between Tissues It is interesting to see whether the most significant features are consistent across different organs. To compare pairs of organs, we used UpSet plots. 24 We examined the top 50 transcripts for each organ and across different sex chromosome and autosome combinations. Figure 6 shows the UpSet plot, which presents influential transcripts ranked by SHAP importance scores across four chromosome groups: chr_aXY, chr_aX, chr_aY, and autosomes. A subset of genes (DCXR, HIVEP3, MS4A7, PRKCA, TCF15, TIE1, TSFM, and XIST) showed sex-specific expression across all four examined organs (BLOOD, BRAIN0, HEART, and BRAIN1), indicating a common biological signal relevant for sex determination. The persistent presence of these genes in organ-specific models suggests they exhibit sex-specific expression patterns and are key players in conserved regulatory pathways and tissue homeostasis. One of these genes, XIST, encodes a well-established long noncoding RNA (lncRNA) involved in X-chromosome inactivation in female mammals. Its expression remains consistently high across female-derived tissues, making it a reliable marker for sex classification. 25 TSFM, which encodes a mitochondrial elongation factor essential for protein synthesis, emphasizes the housekeeping role of mitochondrial function. 26 Its widespread presence in metabolically active tissues may explain its frequent use as a predictive feature, mainly because mitochondrial activity varies between sexes in metabolic tissues. DCXR encodes a multifunctional reductase involved in carbohydrate metabolism and renal osmoregulation. Its expression has been observed in various tissues, including the kidney, liver, and testis. It may be affected by hormonal regulation or metabolic differences between sexes. 27 Similarly, HIVEP3 is a transcription factor that modulates NF-κB-mediated transcription, a pathway widely linked to sex-biased immune responses and inflammation. This potentially connects this gene to sexually dimorphic gene regulation even outside immune tissues. 28 Genes involved in vascular and developmental signaling are also sex-specific across multiple organs. TIE1, a receptor tyrosine kinase involved in angiogenesis that responds to VEGF signaling, is known to be influenced by sex hormones such as estrogen, particularly in endothelial contexts. 29 PRKCA, which encodes protein kinase C alpha, plays vital roles in cell proliferation, cardiac function, and inflammation. It has been shown to exhibit sex-differentiated expression patterns in cardiac and immune tissues, likely due to differential activation by steroid hormones. 30 TCF15, a helix-loop-helix transcription factor necessary for early mesodermal development, was also found across various tissues. Its involvement in early patterning events like somitogenesis could explain inherent differences in developmental gene networks between male and female embryos. 31 Finally, MS4A7, a member of the membrane-spanning 4-domain family, is enriched in the myeloid lineage and linked to macrophage activation and sex differences related to aging in immune regulation. 32 Sex-specific genes perform core metabolic and structural functions (e.g., TSFM, DCXR), maintain tissue-specific regulatory mechanisms (e.g., TIE1, PRKCA), and play sex-specific developmental and immune roles (e.g., XIST, HIVEP3, MS4A7). Their consistent detection across various tissues indicates they form a biologically meaningful and universal feature set for sex classification across multiple organs and tissue types. Our findings demonstrate that RNA-seq data reliably predicts biological sex across various organs. Models incorporating sex chromosome transcripts consistently show near-perfect accuracy, with sex-linked genes such as XIST, KDM5D, and EIF1AY playing key roles, emphasizing their part in sex determination and X chromosome dosage compensation. Predictions based solely on autosomal transcripts are also possible, albeit less accurate, suggesting sex-specific regulatory effects extend beyond sex chromosomes. Autosomal genes involved in immune response, transcription regulation, and metabolism emerge as important predictors, indicating significant hormonal and epigenetic influences. These results reveal that sex-related gene expression differences extend beyond sex chromosomes, reflecting overall physiological distinctions that could influence clinical approaches, including personalized therapies and disease management. Age determination We employed Support Vector Regression (SVR) with a linear kernel from the Scikit-learn package 33 to predict age in male and female samples in the datasets with age information (HEART, BRAIN0, and BRAIN1). Separate models were trained for each dataset and sex. Influential transcripts were identified based on the absolute value of the Spearman rank correlation coefficient between gene expression and age (p-value < 0.05). To determine the optimal number of transcripts for each dataset and sex, we applied the Elbow method to each dataset and sex. In the male HEART dataset, eight transcripts achieved an RMSE of 2.57. Three transcripts in the female HEART dataset resulted in an RMSE of 1.47. For the male BRAIN0 dataset, ten transcripts yielded an RMSE of 3.49. Five transcripts in the female BRAIN0 dataset led to an RMSE of 2.8. The male BRAIN1 dataset, with fourteen transcripts, showed an RMSE of 4.92, while the female BRAIN1 dataset had ten transcripts and an RMSE of 5.38. These results suggest that predicting female age may require fewer features than predicting male age, possibly due to more pronounced age-dependent hormonal changes in females. Scatter plots illustrating the comparison between predicted and reported ages for each dataset and sex are displayed in the Fig. 7 , with males on the left and females on the right. In the HEART dataset, predicted ages closely match actual ages, with only slight deviations in older individuals. This dataset shows the lowest root mean square error (RMSE), indicating high prediction accuracy: 0.98 years for males and 1.46 years for females. Predictions for the BRAIN0 dataset were less precise, especially for females, with an RMSE of 3.95 years for males and 8.89 years for females. The BRAIN1 dataset showed moderate accuracy, with RMSE values of 6.00 years for males and 6.99 years for females. The BRAIN1 dataset demonstrated intermediate accuracy, with RMSE values of 6.00 years for males and 6.99 years for females. See Table 4 for performance results. Table 4 Table Results of prediction of age by SVR models. MAE - mean absolute error, RMSE - root mean squared error, R 2 - coefficient of determination Dataset Sex #Samples Model #Transcripts MAE RMSE R 2 HEART male 9 SVR, LeaveOneOut CV 9 0.8 0.98 0.93 HEART female 11 SVR, LeaveOneOut CV 5 1.15 1.46 0.92 BRAIN0 male 25 SVR, LeaveOneOut CV 7 3.02 3.95 0.68 BRAIN0 female 9 SVR, LeaveOneOut CV 7 6.41 8.89 0.69 BRAIN1 male 150 SVR, 5-fold CV 17 4.57 6.00 0.45 BRAIN1 female 65 SVR, 5-fold CV 8 5.57 6.99 0.57 Interestingly, some genes show opposite relationships with age in males and females. A notable example is the X-ray repair cross-complementing 4 (XRCC4) gene, which we will discuss in detail. As shown in Fig. 8 , XRCC4 expression decreases with age in males but increases in females. XRCC4 plays a key role in DNA repair by supporting the non-homologous end-joining (NHEJ) repair pathway for double-strand breaks. This repair process is essential for maintaining genomic stability in postmitotic cells, such as neurons, which have limited DNA replication capacity and thus depend heavily on efficient DNA repair. DNA repair ability typically declines with age, and mutations or altered expression of repair genes like XRCC4 have been linked to faster cellular aging and higher risks of cancer and age-related diseases due to genomic instability and mutation buildup 34 . Because XRCC4 is vital in NHEJ-mediated DNA repair, its misregulation can significantly compromise genomic integrity. Misregulation of XRCC4 is connected to neurodegenerative disorders, which become manifested with aging. Since neurons rarely divide, they rely on effective DNA repair pathways to prevent damage accumulation. Faulty XRCC4 repair processes could cause increased neuronal cell death, something seen in neurodegenerative diseases such as Alzheimer's, where aging is a significant risk factor. 35 Additionally, decreased XRCC4 expression may accelerate cellular aging by allowing persistent DNA damage to accumulate, potentially exarcerbating chronic inflammation and contributing to tissue deterioration with age. Aging tissues commonly show an increase in senescent cells; thus, impaired XRCC4 function might hasten this shift to senescence, further driving tissue decline as we age. 36 ​ The observed sex-specific differences in XRCC4 expression with age may result from various factors, including hormonal regulation, epigenetic changes, inflammatory responses, oxidative stress, metabolism, cellular aging, and evolutionary adaptations. Estrogen and androgens influence DNA repair processes differently. Estrogen has been shown to enhance DNA double-strand break (DSB) repair, potentially increasing XRCC4 expression in females. On the other hand, androgens, through the androgen receptor (AR), directly bind to and boost XRCC4 promoter activity, promoting its transcription in males. Age-related declines in these hormones could differently affect XRCC4 expression between the sexes. Sex-specific epigenetic changes, including DNA methylation and histone modifications, can lead to differential gene expression. Studies show that hormonal exposures during critical periods can cause lasting epigenetic modifications, potentially influencing XRCC4 expression differently in males and females. ​ 37 . Chronic inflammation shows sex-specific patterns, with females often exhibiting increased inflammatory responses. These differences can affect gene expression in DNA repair, including XRCC4, which contributes to the observed sex-based variations. 38 Sex-based metabolic differences can lead to varying levels of oxidative stress, which, in turn, impact DNA damage and repair mechanisms. These metabolic differences may lead to different regulation of XRCC4 expression between men and women. Evidence suggests that women might have a higher susceptibility to DNA damage and a greater likelihood of developing senescence. This could result in sex-specific regulation of DNA repair genes, such as XRCC4. 39 Evolutionary pressures have shaped sex-specific strategies for DNA repair and maintenance. These adaptations may have resulted in inherent differences in how XRCC4 is expressed and regulated between males and females. Population Differentiation and Ancestry Prediction from Genomic and Transcriptomic Data The 1000 Genomes Project database contains whole-genome DNA sequencing and blood RNA-seq measurements from multiple populations, including the Finnish (FIN), Yoruba (YRI), Tuscans (TSI), Utah residents with Northern and Western European ancestry (CEU), and British (GBR). Principal Component Analysis (PCA) plots (Fig. 9 ) clearly distinguish Yoruba, Tuscans, and Finnish individuals from British and Utah residents. Admixture profiles demonstrate a significant overlap between CEU and GBR populations (Fig. 10 : panels A, B, and C). Using GPS classification 40 , 34–56% of CEU samples were classified as GBR, and conversely, 33–45% of GBR samples as CEU, depending on the number of admixture components selected. The highest average prediction accuracy across populations was observed with K = 12 admixture components (Table 5 , column "Whole-genome SNPs"). Table 5 Accuracy of Ancestry Inference using different datasets. The whole-genome column corresponds to the prediction accuracy, measured as the proportion of individuals in each population attributed to this population by GPS applied to whole-genome SNPs. The RNA-Seq SNPs column corresponds to the prediction accuracy, measured as the proportion of individuals in each population attributed to this population by GPS applied to RNA-Seq SNPs. POP Whole-genome SNPs RNA-Seq SNPs RNA-Seq Expression RNA-seq Expression SNPs CEU 0.622 0.926 0.986 0.975 FIN 1.000 0.711 0.837 0.877 GBR 0.644 0.470 0.839 0.852 TSI 1.000 0.575 0.771 0.779 YRI 1.000 1.000 0.986 1.000 Average 0.853 0.736 0.856 0.897 We further utilized RNA-seq data from the same individuals to infer ancestry using machine learning with Gradient Boosting Machines (GBM). A cross-validation approach (training with 80% of the samples and testing with the remaining 20%, repeated 500 times) demonstrated high accuracy, with GBM correctly identifying ancestry for 99% of CEU, 86% of FIN, 84% of GBR, 77% of TSI, and 98.6% of YRI samples (Table 6 ). Overall prediction accuracy using RNA expression was 87%, comparable to the 86% achieved using whole-genome SNP data. However, error distributions varied between methods. Whole-genome SNP data accurately distinguished between Finnish, Yoruba, and Tuscan samples but often confused the CEU with the British. Conversely, expression-based analysis better differentiated CEU from GBR, though it was less precise among Europeans from different ethnic origins. Analysis with PyLAE 41 and GPS did not reveal significant genomic regions associated with misclassification, indicating that local ancestry admixture did not significantly affect gene expression differences between misannotated samples. Table 6 Confusion matrix for ancestry inference using RNA-Seq CEU FIN GBR TSI YRI CEU 0.986 0.01 0.002 0.002 0 FIN 0.031 0.838 0.037 0.088 0.006 GBR 0.001 0.049 0.839 0.096 0.014 TSI 0.021 0.098 0.083 0.771 0.027 YRI 0.005 0.001 0.001 0.007 0.986 We evaluated the capability of RNA-seq to detect single-nucleotide polymorphisms (SNPs) for ancestry prediction through PCA and admixture analyses. Among 173 individuals with multiple RNA-seq samples, predictions matched perfectly across samples for 171 individuals. However, different samples yielded conflicting ancestry predictions for two individuals (one CEU and one GBR) due to substantial differences in the number of detected SNPs, exemplified by individual NA12399 (49,302 vs. 15,236 SNPs in the two samples). Overall, SNP-based ancestry RNA-seq predictions demonstrated lower accuracy compared to whole-genome sequencing or RNA expression-based predictions. Combining expression-based and SNP-based methods improved accuracy beyond either method alone (Table 5 ). Next, we investigated which genes contribute most significantly to distinguishing ethnic groups. KEGG and Gene Ontology (GO) annotations (Table 7 , Table 8 , Table 9 , Table 10 ) revealed enrichment in immune-related pathways. This enrichment likely explains why RNA-seq-based methods provided better separation between British and Utah residents than whole-genome sequencing. Individuals relocating from North-Western Europe (e.g., UK) to Utah transitioned from a temperate maritime climate to a semi-arid continental climate, encountering substantial environmental and dietary changes. These included colder and snowier winters (average temperatures dropping from 45°F to 35°F), hotter summers (increasing from 70°F to 80°F), higher elevation (~ 4,200 feet compared to ~ 35 feet), clearer skies, reduced precipitation (by ~ 10%), and calmer wind conditions. These climatic and environmental changes may have driven rapid epigenetic adaptations in immune-related genes, which respond more quickly than genetic mutations to environmental pressures. Table 7 KEGG classification of genes differentially expressed between British and Utah residents. KEGG ID Description P-value Adjusted P-value Number of genes hsa05323 Rheumatoid arthritis 2.04E-07 2.97E-05 11 hsa04640 Hematopoietic cell lineage 3.48E-07 2.97E-05 11 hsa05310 Asthma 3.71E-07 2.97E-05 7 hsa05321 Inflammatory bowel disease 6.41E-07 3.38E-05 9 hsa04672 Intestinal immune network for IgA production 7.52E-07 3.38E-05 8 hsa04659 Th17 cell differentiation 8.46E-07 3.38E-05 11 hsa05330 Allograft rejection 1.62E-06 5.57E-05 7 hsa04612 Antigen processing and presentation 3.82E-06 0.0001 9 hsa04940 Type I diabetes mellitus 3.89E-06 0.0001 7 hsa05332 Graft-versus-host disease 4.56E-06 0.00011 7 hsa05416 Viral myocarditis 1.07E-05 0.00023 8 hsa04658 Th1 and Th2 cell differentiation 1.22E-05 0.00024 9 hsa04141 Protein processing in endoplasmic reticulum 1.28E-05 0.00024 12 hsa05320 Autoimmune thyroid disease 1.63E-05 0.00028 7 hsa05150 Staphylococcus aureus infection 2.40E-05 0.00036 9 hsa05140 Leishmaniasis 2.43E-05 0.00036 8 hsa04145 Phagosome 0.00016 0.00232 10 hsa05166 Human T-cell leukemia virus 1 infection 0.00018 0.00242 12 hsa05169 Epstein-Barr virus infection 0.00032 0.00392 11 hsa05145 Toxoplasmosis 0.00033 0.00392 8 hsa05152 Tuberculosis 0.00049 0.00564 10 hsa05164 Influenza A 0.00144 0.01511 9 hsa05322 Systemic lupus erythematosus 0.00145 0.01511 8 hsa01232 Nucleotide metabolism 0.00206 0.02058 6 hsa04514 Cell adhesion molecules 0.00311 0.02987 8 hsa04061 Viral protein interaction with cytokine and cytokine receptor 0.00464 0.04283 6 Table 8 Top 5 GO categories for biological process ontology GO ID Description P-value Adjusted P-value Number of genes GO:0002377 immunoglobulin production 2.57E-16 8.16E-13 25 GO:0002440 production of molecular mediator of immune response 1.24E-13 1.97E-10 27 GO:0019886 antigen processing and presentation of exogenous peptide antigen via MHC class II 2.63E-10 2.78E-07 9 GO:0002495 antigen processing and presentation of peptide antigen via MHC class II 8.76E-10 5.07E-07 9 GO:0002399 MHC class II protein complex assembly 9.57E-10 5.07E-07 7 Table 9 Top 5 GO categories for cellular component ontology GO ID Description P-value Adjusted P-value Number of genes GO:0009897 external side of the plasma membrane 5.09E-16 1.83E-13 34 GO:0019814 immunoglobulin complex 1.16E-15 2.08E-13 19 GO:0042613 MHC class II protein complex 3.63E-11 4.36E-09 8 GO:0042611 MHC protein complex 1.46E-09 1.32E-07 8 GO:0072562 Blood microparticle 2.20E-08 1.59E-06 14 Table 10 Top 5 GO categories for molecular function ontology GO ID Description P-value Adjusted P-value Number of genes GO:0003823 antigen binding 5.16E-28 2.46E-25 29 GO:0023026 MHC class II protein complex binding 1.85E-12 4.41E-10 10 GO:0023023 MHC protein complex binding 4.99E-11 7.92E-09 10 GO:0032395 MHC class II receptor activity 7.70E-06 9.00E-04 4 GO:0140375 Immune receptor activity 2.00E-04 2.00E-02 9 Sex-specific epigenetic adaptations to environmental factors have also been previously reported 42 . Consistent with this, we identified transcripts with differential expression between CEU and GBR separately for males and females. Fourteen transcripts were expressed at higher levels in Utah females than British females, and eleven transcripts showed higher expression in Utah males than British males. Five transcripts (IGF2BP3, CIBAR1P1, GLDCP1, ITPKB, and PIK3CD-AS1) exhibited higher expression in CEU than in GBR for both sexes. IGF2BP3 is involved in growth regulation through binding insulin-like growth factor 2 mRNA. CIBAR1P1 and GLDCP1 are pseudogenes potentially involved in gene regulation and glucose metabolism. ITPKB participates in immune signaling through inositol phosphate pathways, while PIK3CD-AS1 regulates the immune-associated PIK3CD gene. 78 transcripts showed reduced expression in Utah females compared to British females, and 97 transcripts were lower in Utah males than in British males. Among these, 44 transcripts (including ZNF22, ABTB2, FOXO4, IL2RA, IFITM3, and KCNA3) were consistently lower in Utah males compared to British individuals of both sexes. These genes are involved in vital processes including transcriptional regulation, cellular stress responses, inflammation, immune responses, metabolism, and cellular proliferation. Variations in the expression of these genes may indicate how the immune system adapts to Utah's environment compared to the UK. Rapid epigenetic modifications such as DNA methylation and histone modifications likely facilitated quick adaptation to climatic shifts, dietary changes, and pathogen exposure. For instance, populations like the Yakuts and Inuit demonstrate rapid epigenetic regulation of immune genes in response to Arctic environments 43 , 44 . Genetic drift and natural selection may have further contributed to these adaptations, as observed in high-altitude populations adapting to low oxygen levels 45 . The identified differences in immune gene expression between Utah and British populations may reflect differential adaptations to specific pathogens or stressors unique to each environment. For example, genes such as FOXO4 and IL2RA, crucial regulators of inflammation and immune response, showed significant expression differences, possibly due to varying infectious challenges between regions. We further explored gene expression differences in GO categories between males and females from Europe (GBR) and America (CEU), selecting categories containing five or more genes. The most significant differences were observed in categories related to immune functions, including immunoglobulin complexes (GO:0019814), peptidoglycan binding (GO:0042834), and antigen binding (GO:0003823), with consistently higher expression in GBR compared to CEU. Interestingly, certain genes showed opposite expression changes between sexes across populations. For example, HOMER1 (which inhibits T-cell activation) and SMN1 were more expressed in American females and British males. In contrast, CSNK1A1 (a regulator of mTOR signaling and inflammasome assembly) and FNDC3B displayed opposite patterns, being higher in American males and British females. 46 , 47 These complex expression dynamics highlight sex-specific responses to environmental and evolutionary pressures shaping population differentiation. Conclusions This study highlights the significant potential of RNA-Seq data ML-empowered analysis for accurately predicting complex biological traits, including sex, age, and ethnicity. It provides deep insights into how genetic and environmental factors affect gene expression. We successfully identified key transcripts and biological pathways associated with these traits by employing ML approaches, such as GBM, SVR, and SHAP. This highlights the critical role of transcriptomics in population genetics and personalized medicine. Sex prediction using RNA-Seq data demonstrated remarkable accuracy across multiple tissues, emphasizing the pivotal role of sex chromosomes. Transcripts such as XIST, crucial for X-chromosome inactivation, and Y-linked genes, including KDM5D and EIF1AY, emerged as particularly influential. These findings reinforce RNA-Seq’s capability to distinguish between biological males and females robustly. Notably, autosomal and sex-chromosomal genes significantly contributed to achieving precise sex classifications. Age prediction based on RNA-Seq displayed greater variability across different tissues but produced promising outcomes. Notably, heart tissue provided the most accurate age estimations, particularly in females, suggesting substantial age-related shifts in gene expression within specific tissues. SVR models effectively identified transcripts indicative of biological aging, highlighting their potential as biomarkers for monitoring age-related physiological changes and associated diseases. A significant achievement of our study is the demonstrated capability of RNA-Seq data to accurately differentiate between genetically related populations, such as CEU and GBR. Immune-related pathways played a crucial role in this differentiation, with genes such as IL2RA and FOXO4 being prominently featured. These differences likely represent adaptive responses to distinct environmental pressures, such as climate variations and pathogen exposure, underscoring the responsiveness of immune genes to environmental factors. KEGG and Gene Ontology analyses further highlighted immune function pathways, including Th17 cell differentiation and intestinal immune networks, supporting the hypothesis of rapid epigenetic adaptations driven by environmental pressures. For instance, the distinct climatic conditions of Utah compared to the UK appear to have influenced differential gene expression, especially within immune-related pathways, reflecting adaptive physiological responses. Our findings highlight the benefits of integrating transcriptomic data with other genomic resources, such as whole-genome and whole-transcriptome sequencing, thereby enhancing predictive accuracy. This integrated approach provides a more comprehensive understanding of the genetic and environmental influences, thereby enhancing the precision of ethnicity classification and elucidating the molecular mechanisms underlying population-specific adaptations. Future research should prioritize diversifying datasets by incorporating additional populations exposed to varied environmental contexts, refining predictive models, and broadening their applicability. Additionally, further exploring the interactions between genetic variants and epigenetic modifications will deepen our understanding of gene expression regulation by inherited and environmental factors. Expanding this integrated multi-omics approach has the potential to substantially enhance predictions of disease susceptibility and treatment responses, significantly advancing personalized medicine. In conclusion , our results demonstrate that RNA-Seq data have exceptional potential for predicting complex biological traits and elucidating the intricate relationships between genetics, environment, and gene expression. These predictive capabilities have essential implications for precision medicine, population genetics, and aging studies, enabling the tailoring of therapeutic interventions, clarification of disease risk profiles, and optimization of strategies for healthier aging. While challenges remain, especially in distinguishing genetic and environmental influences in ethnicity prediction, the continued integration of RNA-Seq data with multi-omics resources promises even greater precision in clinical practice and deeper insights into human diversity and adaptive mechanisms. Materials and Methods RNA-Seq Data We used the RNA-Seq dataset from four RNA-Seq experiments, as shown in Table 1 . Whole-genome Data We obtained admixture 48 results from the 1000 Genomes Project 49 for admixture component numbers (K) ranging from 5 to 25, available at the official database of the 1000 Genomes Project 50 . These admixture profiles were derived using 193,634 biallelic markers, including non-singleton single-nucleotide variant (SNV) sites separated by at least 2 kb. From these comprehensive admixture datasets, we selected populations that correspond to our RNA-seq data: Yoruba (YRI), Utah residents with Northern and Western European ancestry (CEU), Finnish (FIN), British (GBR), and Tuscan (TSI). In total, 667 RNA-seq blood samples corresponded to these populations, with 447 samples having matching admixture profiles. Due to multiple RNA-seq measurements available per individual, these 667 samples represented 464 unique individuals. To accurately capture and characterize population structure, we clustered the admixture profiles within each ethnic group, identifying distinct subpopulations, such as FIN_1 and FIN_2. Additionally, outlier detection was performed to remove samples that appeared non-representative or potentially mislabeled. To assess the accuracy of ancestry label reconstruction based on genotyping data, we employed a leave-one-out validation approach combined with the Geographic Population Structure (GPS) algorithm 51 . Predictions were correct if an individual's predicted subpopulation matched at least one closely related subpopulation from their assigned group (for example, an individual from subpopulation FIN_2 correctly predicted as FIN_1 was considered accurately classified). RNA-Seq Data Processing Fastq files were downloaded using the fastq-dump and fasterq-dump programs 52 – 54 . FastQC 55 , 56 was used to assess the quality of the reads, and adapter trimming and removal of low-quality nucleotides were performed with TrimGalore 56 . The reads were aligned to the human genome hg38 (GRCh38.primary_assembly.genome.fa) with STAR 57 , 58 , and duplicated reads were marked with Picard 59 . Transcript expression was quantified with stringtie 60 using Gencode v44 gene annotation. The batch effect was removed with ComBat-seq 61 . Custom R and Python scripts were used to collect TPM matrices and visualize the results. Functional annotation and GO terms were obtained using DAVID 62 . The resulting transcripts were merged to create a superset of 380K transcripts. RNA-seq data from all experiments was mapped to this new reference transcriptome. Visualization The results were visualized using custom R and Python scripts and downloadable packages, such as SuperVenn 63 . Sex determination Data preprocessing This collection contains RNA-seq data from blood generated by the Geuvadis consortium 64 for 464 samples from the 1000 Genomes project, including samples from five of the 1000 Genomes populations (CEU, FIN, GBR, TSI, and YRI). Preprocessing of the BLOOD dataset was performed using the following steps: Remove transcripts with zero median expression in both sexes separately. Select transcripts that are highly correlated with sex (Point-biserial correlation coefficient > 0.8); Apply \(\:{\text{log}}_{2}\left(x+1\right)\) transformation; Exclude transcripts with low variability, as indicated by a coefficient of variation \(\:\frac{\sigma\:}{\mu\:}<0.7\) ; Remove poorly expressed transcripts with mean expression below the median(mean expression) \(\:{\mu\:}_{i}<median\left(\mu\:\right)\) ; Filter out transcripts with a coefficient of variation below the median coefficient of variation across all transcripts. Four datasets were created: Chr_aXY: transcripts from X, Y chromosomes, and autosomes Chr_aX: transcripts from the X chromosome and autosomes Chr_aY: transcripts from X and Y chromosomes and autosomes Autosomes Transcripts located in the pseudoautosomal regions of sex chromosomes (Y chromosome: [10001..2781479], [56887903..57217415] and X chromosome [10001..2781479], [155701383..156030895]) were removed from Chr_aX and Chr_aY. The preprocessing of the HEART, BRAIN0, and BRAIN1 datasets involved two steps: removing transcripts with zero median expression, followed by \(\:{\text{log}}_{2}\left(x+1\right)\) transformation. Table 8 shows the number of transcripts in each dataset. Due to the varying number of transcripts remaining after the pre-processing step in the BLOOD, HEART, BRAIN0, and BRAIN1 datasets, the models were trained using transcripts overlapping with the BLOOD dataset transcripts. ML model training The XGBoost gradient boosting method was chosen as the classification model. 19 All data was standardized (median-centered and scaled using an IQR interval, according to the formulae \(\:z=\frac{x-median\left(x\right)}{IQR\left(x\right)}\) due to the non-normality of expression counts, using the RobustScaler routine from the scikit-learn library 65 . The model was trained using 5-fold cross-validation using the BLOOD dataset. In the evaluation phase (BLOOD, HEART, BRAIN0, and BRAIN1 datasets), the average model prediction of each fold was calculated. The transcripts in each of the 16 datasets were ranked according to their influence on model predictions using the SHapley Additive exPlanations (SHAP) 21 . SHAP is a game-theoretical approach to explaining the output of any machine learning model. SHAP assigns each feature an “importance score” based on its contribution to the model's output. The model was sequentially trained on the most significant transcripts; the number of significant transcripts varied from 1 to 100. ROC AUC, precision, recall, and F1 scores were calculated. Then, we selected the optimal number of transcripts for the final model training, based on maximum of two values: 1) the number from the elbow method using the quality vs. the number of transcript plot, and 2) the number of transcripts for that were in the top 100 SHAP transcripts in the 40 out of 50 cross-validation steps. Due to variations in the number of transcripts retained after this pre-processing stage, we trained models separately for each dataset (HEART, BRAIN0, and BRAIN1) using only those transcripts that overlapped with the BLOOD dataset. This procedure resulted in 16 unique transcript subsets for training machine learning (ML) models. The optimal number of transcripts for each dataset is shown in Table 3 . Importance of transcripts To evaluate the importance of each transcript, we trained models on the BLOOD dataset using each of the 16 transcript subsets. Model performance was validated via repeated stratified 5-fold cross-validation with 10 repetitions, implemented using the RepeatedStratifiedKFold function from the scikit-learn library. This method preserves the class distribution within each fold and ensures robust performance estimates by repeating the process multiple times with different fold partitions. During each cross-validation iteration, we computed SHAP values for feature importance. We selected the top 100 transcripts for every fold with the highest mean absolute SHAP values. Each transcript then received a score from 0 to 50, representing the number of times it appeared in the top 100 across all cross-validation iterations (5 folds × 10 repetitions = 50 iterations). The final list of the top 100 transcripts with the highest cumulative scores was carried forward to the subsequent analysis stage. Selecting the Optimal Number of Transcripts For each of the 16 transcript sets, we obtained the top 100 transcripts ranked by cumulative SHAP importance. We then retrained the models on the BLOOD dataset using only these selected transcripts. Again, we applied repeated stratified 5-fold cross-validation with 10 repetitions for model evaluation. Model performance To assess how model performance depends on the number of input features, we varied the number of transcripts used in training from 1 to 100 in incremental steps. For each model configuration (i.e., each number of top transcripts), we calculated standard classification metrics (Accuracy, Precision, Recall, and F1 Score) on every cross-validation fold. The performance trends were plotted for each set, and the Elbow Method was employed to determine the optimal number of transcripts. The Elbow Method identifies the point at which adding more features results in diminishing returns in performance, thereby enabling the selection of a parsimonious yet informative transcript set. Age determination Data preprocessing BLOOD, HEART, BRAIN0, and BRAIN1 datasets were preprocessed using the same protocol as for sex determination, with modified Step 2 - pre-selection of transcripts highly correlated with age (Spearman rank correlation coefficient R > 0.8, p-value < 0.05. All datasets were then split by sex because of the expected different relationship between gene expression levels and age for males and females. We used the BRAIN1 dataset 66 , which contained 65 female RNA-seq samples (ranging from 21 to 58 years old, mean age 39.77, median age 44) and 150 male samples (ranging from 27 to 62 years old, mean age 46.46, median age 48). The transcripts were ranked based on their absolute correlation value with age, and the top 500 were selected; this dataset was pruned to exclude highly correlated (r > 0.5) transcripts, leaving 307 transcripts for females and 421 for males. ML model training Several linear models were tested: Support Vector Regression (SVR) with linear kernel, Linear Regression, and Elastic Net. SVR demonstrated the best results regarding Root Mean Squared Error (RMSE). Due to the small number of samples in all datasets except BRAIN1, the model was trained using Leave-One-Out cross-validation (for BRAIN1, a 5-fold stratified cross-validation procedure was employed); in the evaluation phase, the average model prediction for each sample was calculated. Random Forest was used to make a model for age prediction: model <- train(AGE~., data=(FE)MALES2, method="rf", preProcess="scale") A 10-fold cross-validation procedure was conducted 30 times, with 80% of the samples used for training and 20% for control. Ancestry determination Expression-based inference We selected the 500 most informative transcripts from the BLOOD dataset showing differential expression between 10 pairs of populations (YRI, TSI, CEU, FIN, GBR). Next, we utilized the Gradient Boosting Machines (GBM) method, which was implemented as an R package gbm. The gbm() function parameters were selected based on brute-force RMSE minimization. The dataset was split 500 times into 80% for training and 20% for testing, and the linear predictors from all runs were averaged for every sample. model_gbm = gbm( train$ethn ~., data = train, distribution = "multinomial" , n.trees = 83, interaction.depth = 5 shrinkage = 0.1, cv.folds = 10 , n.minobsinnode = 10, bag.fraction = .65, n.cores = NULL, verbose = F ) pred = gbm::predict.gbm(model_gbm, test_x)) We averaged the prediction scores for each sample and determined the predicted ancestry as having the highest score. RNA-Seq Data SNP calling BAM files were sorted by coordinates using GATK SortSam, and duplicates were marked with GATK 67 MarkDuplicates. Then bcftools mpileup was used to create a bcf file converted by bcftools to the vcf format. Next, the bcftools 68 view function filtered out low-quality (QUAL < = 10 || DP < 10) positions. Vcftools 67 , 69 was used to remove indels and genotypes with a quality score below the specified threshold (--remove-indels --minGQ 15). VCF files for individual samples were combined using the bcftools merge function and filtered based on a minor allele frequency (MAF) below 0.05 and a missing data threshold of less than 0.9. Ancestry prediction using a combination of SNP data and expression from RNA-Seq Data To combine predictions based on SNPs and expression data derived from RNA-Seq experiments, we used the output of the GBM predict function and converted the GBM output to scores. We identified 50 individuals with the closest admixture vectors (based on Euclidean distance) for each sample to achieve this. Then, we found the relative frequencies of the five ancestry labels (YRI, GBM, TSI, FIN, CEU). The frequencies, \(\:{p}_{i}\) , were converted to scores using the formula \(\:{p}_{i}\text{log}\left(\frac{{p}_{i}}{0.2}\right)\) ; the scores were added to the GBM “predict” scores. The predicted ancestry had the highest total scores. References Trabzuni D, Ramasamy A, Imran S et al (2012) Widespread sex differences in gene expression and splicing in the adult human brain. Nat Commun 4 Kobylińska K, Orłowski T, Adamek M, Biecek P (1926) Explainable machine learning for lung cancer screening models. Appl. Sci. (Basel) 12, (2022) Waxman DJ, Connor C (2006) Growth Hormone Regulation of Sex-Dependent Liver Gene 1142 Expression. Mol Endocrinol 20:2613–2629 Mayne BT et al (2016) Large scale gene expression meta-analysis reveals tissue-specific, sex-biased gene expression in humans. Front Genet 7:183 Klein SL, Morgan R (2020) The impact of sex and gender on immunotherapy outcomes. Biol Sex Differ 11:24 Lopes-Ramos CM, Quackenbush J, DeMeo DL (2020) Genome-wide sex and gender differences in cancer. Front Oncol 10:597788 Moqri M, Poganik JR, Horvath S, Gladyshev V (2025) N. What makes biological age epigenetic clocks tick. Nat Aging. 10.1038/s43587-025-00833-1 Kurochkina NS et al (2024) Age-related changes in human skeletal muscle transcriptome and proteome are more affected by chronic inflammation and physical inactivity than primary aging. Aging Cell 23:e14098 Kim SY, Lee J, Baik M (2024) PSVII-22 Age-related transcriptome changes associated with beef quality in longissimus thoracis muscle of Hanwoo steer. J Anim Sci 102:688–688 Fleischer JG et al (2018) Predicting age from the transcriptome of human dermal fibroblasts. Genome Biol 19:221 Franceschi C, Garagnani P, Parini P, Giuliani C, Santoro A (2018) Inflammaging: a new immune-metabolic viewpoint for age-related diseases. Nat Rev Endocrinol 14:576–590 Johnson ML, Robinson MM, Nair KS (2013) Skeletal muscle aging and the mitochondrion. Trends Endocrinol Metab 24:247–256 Palmer D, Fabris F, Doherty A, Freitas AA (2021) & de Magalhães, J. P. Ageing transcriptome meta-analysis reveals similarities and differences between key mammalian tissues. Aging 13:3313–3341 Horvath S (2013) DNA methylation age of human tissues and cell types. Genome Biol 14:R115 Fachrul M et al (2023) Direct inference and control of genetic population structure from RNA sequencing data. Commun Biol 6:804 Kachuri L et al (2023) Gene expression in African Americans, Puerto Ricans and Mexican Americans reveals ancestry-specific patterns of genetic architecture. Nat Genet 55:952–963 Sun S et al (2017) Differential expression analysis for RNAseq using Poisson mixed models. Nucleic Acids Res 45:e106 Shen X et al (2024) Nonlinear dynamics of multi-omics profiles during human aging. Nat Aging 4:1619–1634 Brownlee J (2016) XGBoost With Python: Gradient Boosted Trees with XGBoost and Scikit-Learn. Machine Learning Mastery Li B et al (2018) Genomic prediction of breeding values using a subset of SNPs identified by three machine learning methods. Front Genet 9:237 Babaei G, Giudici P, InstanceSHAP (2022) An instance-based estimation approach for Shapley values. SSRN Electron J. 10.2139/ssrn.4293269 Tozzi A et al (2015) Endogenous 17Î 2 -estradiol is required for activity-dependent long-term potentiation in the striatum: interaction with the dopaminergic system. Front Cell Neurosci 9:192 Shao S et al (2021) Sex-dependent expression of N-cadherin-GluA1 pathway-related molecules in the prefrontal cortex mediates anxiety-like behavior in male offspring following prenatal stress. Stress 24:612–620 Lex A, Gehlenborg N, Strobelt H, Vuillemot R, Pfister H (2014) UpSet: Visualization of Intersecting Sets. IEEE Trans Vis Comput Graph 20:1983–1992 Cerase A, Pintacuda G, Tattermusch A, Avner P (2015) Xist localization and function: new insights from multiple levels. Genome Biol 16:166 Li X-Y et al (2024) Short-term regulation of TSFM level does not alter amyloidogenesis and mitochondrial function in type-specific cells. Mol Biol Rep 51:484 Yang S et al (2017) Diacetyl/l-xylulose reductase mediates chemical redox cycling in lung epithelial cells. Chem Res Toxicol 30:1406–1418 Hicar MD, Liu Y, Allen CE, Wu LC (2001) Structure of the human zinc finger protein HIVEP3: molecular cloning, expression, exon-intron structure, and comparison with paralogous genes HIVEP1 and HIVEP2. Genomics 71, 89–100 Abu Shtaya A et al (2023) Possible biallelic inheritance in TIE1 in a family with congenital lymphedema, intestinal lymphangiectasia and cutis aplasia. Clin Genet 104:275–276 Li Q et al (2015) Two novel PRKCI polymorphisms and prostate cancer risk in an Eastern Chinese Han population. Mol Carcinog 54:632–641 Yang W-Y et al (2015) Coronary risk in relation to genetic variation in MEOX2 and TCF15 in a Flemish population. BMC Genet 16:116 Ni B et al (2023) The short isoform of MS4A7 is a novel player in glioblastoma microenvironment, M2 macrophage polarization, and tumor progression. J Neuroinflammation 20:80 Pedregosa F et al (2012) Scikit-learn: Machine Learning in Python. arXiv [cs.LG] Gorbunova V, Seluanov A (2016) DNA double strand break repair, aging and the chromatin connection. Mutat Res 788:2–6 Lombard DB et al (2005) DNA repair, genome stability, and aging. Cell 120:497–512 Maynard S, Fang EF, Scheibye-Knudsen M, Croteau DL, Bohr VA (2015) DNA damage, DNA repair, aging, and neurodegeneration. Cold Spring Harb Perspect Med 5 McCarthy MM et al (2009) The epigenetics of sex differences in the brain. J Neurosci 29:12815–12823 Qi X, Ng TKS, Wu B (2023) Sex differences in the mediating role of chronic inflammation on the association between social isolation and cognitive functioning among older adults in the United States. Psychoneuroendocrinology 149:106023 Ng M, Hazrati L-N (2022) Evidence of sex differences in cellular senescence. Neurobiol Aging 120:88–104 Elhaik E et al (2014) Geographic population structure analysis of worldwide human populations infers their biogeographical origins. Nat Commun 5:3513 Smetanin A, Moshkov N, Tatarinova TV (2020) Local Ancestry Prediction with PyLAE. Cold Spring Harbor Laboratory 11.13.380105 (2020) 10.1101/2020.11.13.380105 Svensson EI et al (2018) Sex differences in local adaptation: what can we learn from reciprocal transplant experiments? Philos Trans R Soc Lond B Biol Sci 373:20170420 Kalyakulina A et al (2023) Epigenetics of the far northern Yakutian population. Clin Epigenetics 15:189 Zhou S et al (2019) Genetic architecture and adaptations of Nunavik Inuit. Proc. Natl. Acad. Sci. U. S. A. 116, 16012–16017 Lorenzo FR et al (2014) A genetic mechanism for Tibetan high-altitude adaptation. Nat Genet 46:951–956 Duan S et al (2011) mTOR generates an auto-amplification loop by triggering the βTrCP- and CK1α-dependent degradation of DEPTOR. Mol Cell 44:317–324 Gao D et al (2011) mTOR drives its own activation via SCF(βTrCP)-dependent degradation of the mTOR inhibitor DEPTOR. Mol Cell 44:290–303 Alexander DH, Novembre J, Lange K (2009) Fast model-based estimation of ancestry in unrelated individuals. Genome Res 19:1655–1664 1000 Genomes Project Consortium (2015) A global reference for human genetic variation. Nature 526:68–74 1000 Genomes Project (2015) 1000 Genomes Project https://ftp .1000genomes.ebi.ac.uk/vol1/ftp/release/20130502/supporting/admixture_files/ Elhaik E, Tatarinova TV, Klyosov AA (2014) The ’extremely ancient’chromosome that isn’t: a forensic bioinformatic investigation of Albert Perry’s X-degenerate portion of the Y chromosome. Eur J of Harrington R, Fastq-dump (2020) Bioinf Noteb https://rnnh.github.io/bioinfo-notebook/docs/fastq-dump.html Ismail HD, Bioinformatics (2022) A Practical Guide to NCBI Databases and Sequence Alignments. CRC wraetz (2023) HowTo: fasterq dump. ncbi/sra-tools https://github.com/ncbi/sra-tools/wiki/HowTo:-fasterq-dump Andrews S (2010) & Others. FastQC: a quality control tool for high throughput sequence data Krueger F, TrimGalore, GitHub (2012) https://github.com/FelixKrueger/TrimGalore Dobin A et al (2013) STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29:15–21 Homo sapiens genome assembly GRCh38 NCBI https://www.ncbi.nlm.nih.gov/data-hub/assembly/GCF_000001405.26/ Broad Institute (2019) Picard Toolkit. GitHub Repository http://broadinstitute.github.io/picard Pertea M, Kim D, Pertea GM, Leek JT, Salzberg SL (2016) Transcript-level expression analysis of RNA-seq experiments with HISAT, StringTie and Ballgown. Nat Protoc 11:1650–1667 Zhang Y, Parmigiani G, Johnson WE (2020) ComBat-seq: batch effect adjustment for RNA-seq count data. NAR Genom Bioinform 2:lqaa078 Sherman BT et al (2022) DAVID: a web server for functional enrichment analysis and functional annotation of gene lists (2021 update). Nucleic Acids Res 50:W216–W221 Indukaev F, Supervenn (2020) Precise and Easy-to-Read Multiple Sets Visualization in Python Lappalainen T et al (2013) Transcriptome and genome sequencing uncovers functional variation in humans. Nature 501:506–511 Brownlee WE (2016) The World War I Regime. in Federal Taxation in America 93–123. Cambridge University Press, Cambridge Ketchesin KD et al (2023) Diurnal alterations in gene expression across striatal subregions in psychosis. Biol Psychiatry 93:137–148 McKenna A et al (2010) The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res 20:1297–1303 Li H, Handsaker B, Danecek P, McCarthy S, Marshall J. BCFtools Danecek P et al (2011) The variant call format and VCFtools. Bioinformatics 27:2156–2158 Additional Declarations There is NO Competing Interest. Supplementary Files Supplementv2.docx Supplement for Inferring Sex, Ethnicity, and Age from RNA-seq Data Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7256409","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":493784635,"identity":"4d3c243f-e02f-42e8-8815-dfacbaaa2a78","order_by":0,"name":"Tatiana Tatarinova","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA8ElEQVRIiWNgGAWjYPACCRDB+OABXIANt1oeNgbGBqgWZoMEErRAlEkQpcVevvn5g597LPL4JZKfVSTUHJYzn918gOFD2WE8trAZNvY8kyiWnJFmdiPh2GFjmTvHEhhnnMOnhcGwgeeAROKGMweAWtjSEmdI5Bgw87bh08L+sfEPWMvxbwUJ/9LqZ0jkf2D+i1cLj2Ez2JbjPWYMiW02CRISOQzMjPi0HMspnC0D1DKzvadYIrHPxnCGRJrBwZ5z6Ti1sDcf3/DxzYG6xH5m9o0fPnyTkJeQSH744EeZNU4t2MEBEtWPglEwCkbBKEADAOzBVOBDsYONAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0003-1787-1112","institution":"University of La Verne","correspondingAuthor":true,"prefix":"","firstName":"Tatiana","middleName":"","lastName":"Tatarinova","suffix":""},{"id":493784636,"identity":"b27314a4-f037-48d2-9fc8-99b0ac3645a1","order_by":1,"name":"Arseniy Dokuchaev","email":"","orcid":"","institution":"Department of Mathematics, University of Trento","correspondingAuthor":false,"prefix":"","firstName":"Arseniy","middleName":"","lastName":"Dokuchaev","suffix":""},{"id":493784637,"identity":"0e126f19-ad44-44a3-a0a9-5d69c9f5dfee","order_by":2,"name":"Varvara Pozdina","email":"","orcid":"","institution":"Unaffiliated","correspondingAuthor":false,"prefix":"","firstName":"Varvara","middleName":"","lastName":"Pozdina","suffix":""},{"id":493784638,"identity":"9c3b7f63-5fce-4e12-8198-1f33fee70467","order_by":3,"name":"Sergey Gaponov","email":"","orcid":"","institution":"Novikov Lab","correspondingAuthor":false,"prefix":"","firstName":"Sergey","middleName":"","lastName":"Gaponov","suffix":""},{"id":493784639,"identity":"280e57a4-e330-4a99-8522-8c3944bf5218","order_by":4,"name":"Elizaveta Taranenko","email":"","orcid":"","institution":"University of La Verne","correspondingAuthor":false,"prefix":"","firstName":"Elizaveta","middleName":"","lastName":"Taranenko","suffix":""},{"id":493784640,"identity":"b842ac29-1863-4c0b-9531-569b8838248a","order_by":5,"name":"Igor Efimov","email":"","orcid":"https://orcid.org/0000-0002-1483-5039","institution":"Northwestern University","correspondingAuthor":false,"prefix":"","firstName":"Igor","middleName":"","lastName":"Efimov","suffix":""}],"badges":[],"createdAt":"2025-07-30 21:25:30","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7256409/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7256409/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":88335799,"identity":"68d14476-6a29-4800-9dad-b71157d7c3f4","added_by":"auto","created_at":"2025-08-05 11:56:29","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":90869,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance of four predictive models (all genes, autosomes plus X, autosomes, autosomes plus Y) for the BLOOD dataset as a function of the number of transcripts.\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/d0bc22ca11224121133e02b3.jpg"},{"id":88334771,"identity":"42bfd163-774a-41ad-b1ed-6e878d607f3a","added_by":"auto","created_at":"2025-08-05 11:48:29","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":105841,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance of four predictive models (all genes, autosomes plus X, autosomes, autosomes plus Y) for the BRAIN0 dataset as a function of the number of transcripts.\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/36ab50b1576c59f71d564665.jpg"},{"id":88336087,"identity":"5d46afb1-4897-4490-8708-a05f200f1b29","added_by":"auto","created_at":"2025-08-05 12:04:29","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":94769,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance of four predictive models (all genes, autosomes plus X, autosomes, autosomes plus Y) for the HEART dataset as a function of the number of transcripts.\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/e9a2e5691d81aecfcbf283b8.jpg"},{"id":88336086,"identity":"c38a5e17-f844-414a-8507-07690d332c82","added_by":"auto","created_at":"2025-08-05 12:04:29","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":92273,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance of four predictive models (all genes, autosomes plus X, autosomes plus Y, autosomes) for the BRAIN1 dataset as a function of the number of transcripts.\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/89af2a18ffb1c2963ded26d3.jpg"},{"id":88334775,"identity":"8384f70a-ea73-4e09-b808-622af029cf48","added_by":"auto","created_at":"2025-08-05 11:48:29","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":56836,"visible":true,"origin":"","legend":"\u003cp\u003eHEART dataset. Probability to identify the subject as a male, stratified by heart chamber\u003c/p\u003e","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/496fb4e8e47cb2ed1e0944ef.jpg"},{"id":88334781,"identity":"cd0842e2-d947-4d13-b1c6-6a9ec90420da","added_by":"auto","created_at":"2025-08-05 11:48:29","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":121495,"visible":true,"origin":"","legend":"\u003cp\u003eUpSet plots showing the overlaps between transcripts selected by SHAP importance for ML model training. The bottom left horizontal bar graph labeled the total number of common transcripts in BLOODxBRAIN0/BRAIN1/BLOOD datasets. The circles in each panel matrix represent the unique and overlapping transcripts. Connected circles indicate a certain intersection of transcripts between sets. The top bar summarizes the number of transcripts in each unique or overlapping combination. For example, for the chr_aY group, there are 35 unique transcripts in HEART, three unique transcripts in BRAIN1, and one unique transcript in BLOOD and BRAIN0; one transcripts is common for BRAIN0/BRAIN1/BLOOD, and no common transcripts in all four datasets\u003c/p\u003e","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/fc4b41204c04678e6e766ee8.jpg"},{"id":88335801,"identity":"f5025631-5eb8-4177-b9d9-1c5e1678f8e8","added_by":"auto","created_at":"2025-08-05 11:56:29","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":71927,"visible":true,"origin":"","legend":"\u003cp\u003eComparison between the reported age and the age predicted by the SVR models for HEART, BRAIN0, and BRAIN1 datasets and both sexes\u003c/p\u003e","description":"","filename":"7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/56c90742d1e01edbcee4bfd2.jpg"},{"id":88334794,"identity":"e7a39da5-b38d-4eb0-b338-4a8f06760ce1","added_by":"auto","created_at":"2025-08-05 11:48:30","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":44005,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 10. Sex-dependent age expression trends for gene XRCC4, HEART dataset\u003c/p\u003e","description":"","filename":"8.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/0462fa697ae46a91c7481ceb.jpg"},{"id":88334784,"identity":"f17f7272-5269-4c7a-be7d-cb0f7e82b88a","added_by":"auto","created_at":"2025-08-05 11:48:29","extension":"jpg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":44424,"visible":true,"origin":"","legend":"\u003cp\u003ePCA for 447 samples from the CEU, FIN, GBR, TSI, and YRI populations of the 1000 Genomes dataset, based on whole-genome sequencing using 43,016 positions remaining after LD pruning.\u003c/p\u003e","description":"","filename":"9.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/962c7553ec75a67873b71298.jpg"},{"id":88334780,"identity":"148538b9-9134-4e7e-9d4e-cb7aef274b5b","added_by":"auto","created_at":"2025-08-05 11:48:29","extension":"jpg","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":111908,"visible":true,"origin":"","legend":"\u003cp\u003eA) GBM scores for CEU, FIN, GBR, TSI, and YRI. B) Admixture profiles for K=12, 1000 Genomes dataset C) Admixture profiles for K=12 and CEU, FIN, GBR, TSI, and YRI subset of the 1000 Genomes dataset.\u003c/p\u003e","description":"","filename":"10.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/128e1f569ad5642fd2c8876f.jpg"},{"id":88464738,"identity":"9e724973-3fd7-4fcb-870f-84d34ff5fb58","added_by":"auto","created_at":"2025-08-06 17:14:18","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2276236,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/808acf0d-41b0-4fdd-9b7b-11aaa8990d18.pdf"},{"id":88334779,"identity":"7bec9f5f-f7d1-4325-9872-39687fee6605","added_by":"auto","created_at":"2025-08-05 11:48:29","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2688874,"visible":true,"origin":"","legend":"Supplement for Inferring Sex, Ethnicity, and Age from RNA-seq Data","description":"","filename":"Supplementv2.docx","url":"https://assets-eu.researchsquare.com/files/rs-7256409/v1/e97bd129fdb0f9ef871fa0d4.docx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Inferring Sex, Ethnicity, and Age from RNA-seq Data","fulltext":[{"header":"Introduction","content":"\u003cp\u003eRNA sequencing (RNA-Seq) provides a high-resolution snapshot of tissue- and cell-specific gene expression in diverse biological contexts. Prior research has linked specific transcriptomic signatures to phenotypic traits such as biological sex, aging, and ethnicity, underscoring the potential of RNA-Seq for precision medicine, population genetics, and aging research. However, significant gaps remain in our understanding of how consistently RNA-Seq can predict these traits across different human tissues and diverse populations, as well as how genetic and environmental influences interact to shape these transcriptomic patterns. Addressing these gaps is critical, as tissue-specific variations and environmental factors such as personalized healthcare, disease prevention strategies, and targeted interventions could profoundly impact clinical applications. In this study, we systematically explore the ML predictive capabilities of RNA-Seq data analysis across tissues (blood, heart, and several brain regions) to robustly infer sex, biological age, and ethnicity, thereby illuminating underlying biological pathways and advancing the clinical relevance of transcriptomics.\u003c/p\u003e\u003cp\u003e\u003cb\u003eInferring Sex Using RNA-Seq Data\u003c/b\u003e\u003c/p\u003e\u003cp\u003eBiological sex prediction using RNA-Seq data is highly effective, primarily due to distinct gene expression differences driven by the presence of sex chromosomes (X and Y) and hormonal factors. Sex chromosome-linked genes, such as XIST (crucial for female X-chromosome inactivation), and male-specific Y chromosome genes including RPS4Y1, KDM5D, and DDX3Y, exhibit prominent sex-specific expression patterns. Hormonal regulation, particularly involving estrogen and testosterone, further amplifies these sex-specific differences in gene expression, influencing diverse biological processes such as metabolism, immune function, and reproductive health. \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e For example, sex hormones influence gene expression in the liver, leading to differences in growth hormone signaling and differentially impacting drug metabolism in males and females. \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eDistinct gene expression signatures between males and females provide a basis for accurately applying machine learning (ML) techniques to infer biological sex using RNA-Seq data. Machine learning models such as random forests (RF) and support vector machines (SVMs) are commonly employed to predict sex using RNA-Seq data. These models are trained on datasets that include sex-labeled samples, enabling them to identify differential expression patterns in key sex-linked and hormone-regulated genes. Tissue specificity is also important; it has been reported that the liver and brain show pronounced sex-biased RNA expression patterns. \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eSex-specific transcriptomic signatures are clinically significant, influencing drug metabolism, immune responses, and disease susceptibility, reinforcing the role of RNA-Seq in personalized medicine. For instance, females and males often respond differently to cancer immunotherapy; understanding these differences at the transcriptomic level can help optimize treatment strategies. \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eML-empowered Inference of Age from RNA-Seq Data\u003c/b\u003e\u003c/p\u003e\u003cp\u003eAging is a multifaceted biological phenomenon shaped by genetic predispositions and environmental influences. Distinguishing between chronological age (actual lived time) and biological age (physiological health status) is essential, as biological age more accurately reflects individual health conditions and risk factors for age-related diseases. RNA-Seq data captures age-related transcriptional changes, including genes involved in cellular stress responses, inflammation, and mitochondrial function. Genes such as SIRT1, which decline in expression with age, and inflammatory markers like IL-6, which increase, serve as potential biomarkers of biological aging. Biological age can be inferred from epigenetic markers that indicate cellular damage or regenerative capacity level, offering insights into the individual's biological health and aging rate.\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eAge-related changes in the transcriptome and epigenome have been documented across species and tissues, showing that gene expression can serve as a molecular clock to estimate an individual's age.\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e Studies have shown that the transcription of genes involved in cellular stress responses, inflammation, and mitochondrial function is age-dependent. For example, sirtuin family members, such as SIRT1, exhibit decreased expression with age, which reduces DNA repair capacity and increases oxidative stress. \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e In contrast, inflammatory genes, such as IL-6, become upregulated with age, contributing to the chronic, low-grade inflammation associated with aging. \u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eSpecific tissues show distinct aging profiles. For instance, skeletal muscle and the brain exhibit significant changes in mitochondrial and metabolic genes, reflecting the functional decline with age.\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e Conversely, the liver maintains more stable gene expression patterns across the lifespan, reflecting its regenerative capacity. \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eAge prediction models utilize regression-based ML learning algorithms, including support vector regression (SVR) and deep learning models. These models are trained on RNA-Seq datasets that include individuals of varying ages. By identifying patterns of gene expression that correlate with age, the models can predict an individual's age with remarkable precision. Fleischer et al. \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e demonstrated that SVR models trained on fibroblast RNA-Seq data can predict age with a mean error of less than five years.\u003c/p\u003e\u003cp\u003eThe ability to predict tissue-specific age from RNA-Seq data has far-reaching implications by identifying biomarkers and expression patterns that define biological, rather than chronological, age. This distinction is critical in studies of longevity and age-related diseases, where individuals with \"younger\" transcriptomic profiles may exhibit lower disease risk. Tissue-specific gene expression patterns observed in tissues like skeletal muscle, brain, and liver provide additional depth to age prediction, identifying individuals at higher risk for diseases like Alzheimer's and cardiovascular disorders, and enabling timely preventive and therapeutic interventions.\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eInferring Ethnicity Using RNA-Seq Data\u003c/b\u003e\u003c/p\u003e\u003cp\u003eInference of ethnicity from RNA-Seq data presents significant complexity due to the intricate interplay between genetic ancestry and diverse environmental and physiological factors, including diet, microbiome, pathogen exposure, and lifestyle. Several studies have demonstrated that population-specific differences in gene expression can provide valuable insights into an individual's ethnic background, highlighting the potential of RNA-Seq in this area.\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e Ethnicity-related differences in gene expression often arise from population-specific genetic variants affecting regulatory pathways, such as immune response genes CYP3A4 and IFN-γ. \u003csup\u003e16 17\u003c/sup\u003e Additionally, environmentally driven epigenetic modifications, such as DNA methylation and histone acetylation, significantly contribute to ethnic transcriptomic differences. Integrative methodologies that combine RNA-Seq with genomic data (e.g., SNP arrays, whole-genome sequencing) substantially enhance the accuracy of ethnicity prediction, with critical implications for pharmacogenomics. Understanding these ethnicity-specific transcriptomic variations enables more precise medication dosing, reducing adverse drug reactions and improving patient outcomes across diverse populations.\u003c/p\u003e\u003cp\u003eOne of the primary challenges in inferring ethnicity is distinguishing between genetic and environmental influences. For example, differences in gene expression related to immune response could be driven by genetic ancestry or environmental factors, such as pathogen exposure and microbiome. Multi-omics approaches that integrate RNA-Seq data with SNP arrays or whole-genome sequencing have been employed to overcome this challenge. \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e These integrative approaches enhance the accuracy of ethnicity prediction by considering genetic and environmental factors that contribute to gene expression. \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eEthnicity determination is crucial for precision medicine, particularly in pharmacogenomics. Different ethnic groups often exhibit varying responses to medications due to differences in gene expression related to drug metabolism. For example, individuals of African descent tend to express higher levels of CYP3A4, an enzyme involved in the metabolism of many drugs, compared to those of European or Asian descent. \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e Understanding these differences will enable accurate dosing recommendations and reduce the risks of adverse drug reactions.\u003c/p\u003e\u003cp\u003eIn summary, RNA-Seq offers robust predictive potential for determining biological sex, age, and ethnicity, providing valuable insights into gene expression patterns influenced by genetic and environmental factors. This study aims to rigorously evaluate RNA-Seq\u0026rsquo;s predictive capabilities across multiple tissues (blood, heart, and several brain regions), advancing both our biological understanding and the practical application of transcriptomics in personalized healthcare.\u003c/p\u003e"},{"header":"Results and Discussion","content":"\u003cp\u003e\u003cb\u003eSex determination\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWe analyzed four datasets: BLOOD (329 male and 338 female samples), HEART (9 male and 11 female samples), BRAIN0 (25 male and 9 female samples), and BRAIN1 (150 male and 65 female samples); see Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e for details. The largest dataset, BLOOD, served as the training set for developing the sex determination models based on the XGBoost algorithm \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e,\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e, as detailed in the Materials and Methods section. We have tested models with and without sex chromosomes. To identify the most influential transcripts, we utilized the game-based machine learning method SHAP (SHAPLEY Additive exPlanations) \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. Models derived from the BLOOD dataset were tested on the BLOOD, HEART, BRAIN0, and BRAIN1 datasets using a five-fold cross-validation approach. Performance metrics (ROC AUC, accuracy, precision, recall, and F1 scores) obtained using varying numbers of influential transcripts (1-100) are presented in Figs.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. The number of expressed transcripts for all datasets is in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. The number of important features for each model is in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Supplemental Tables\u0026nbsp;1\u0026ndash;4 describe transcripts used for sex determination.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTraining and Testing RNA-Seq Datasets\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eProject id\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eDataset name\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eEthnicity information\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eAge information\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eOrgan/tissue\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colspan=\"2\" nameend=\"c7\" namest=\"c6\"\u003e\u003cp\u003eNumber of samples\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eMale\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eFemale\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eERP001941\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBLOOD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eProvided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eNot provided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eBlood (lymphoblastoid cell lines from the 1000 Genomes)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e329\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e338\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePRJNA877032\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHEART\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNot provided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eProvided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eHeart (atrium and ventricle, left and right)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e11\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePRJNA448069\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBRAIN0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNot provided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eProvided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eBrain (Frontal cortex)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e25\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePRJNA836496\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBRAIN1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNot provided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eProvided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eBrain (Putamen, Caudate,\u003c/p\u003e\u003cp\u003eNucleus accumbens)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e150\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e65\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eNumber of transcripts in the datasets after filtering.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eChr_aXY\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eautosomes\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eChr_aX\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eChr_aY\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBLOOD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e15771\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e15354\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e15725\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e15368\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHEART\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e8668\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e8441\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e8660\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8441\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBRAIN0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e9703\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e9448\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e9671\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9460\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBRAIN1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e9429\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e9162\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e9400\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9175\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eNumber of important transcripts for different tissues, selected based on the SHAP scores\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDataset\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBRAIN0\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eBRAIN1\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eHEART\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eBLOOD\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChr_aXY\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChr_aX\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChr_aY\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAutosomes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e46\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e29\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e22\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e40\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe following sections discuss the performance outcomes for each dataset individually.\u003c/p\u003e\u003cp\u003e\u003cb\u003eBLOOD\u003c/b\u003e\u003c/p\u003e\u003cp\u003eModels trained with datasets containing sex chromosomes (chr_aXY, chr_aX, and chr_aY) required only ten transcripts to reach nearly 100% accuracy in sex determination (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Adding extra features did not affect the model's performance. Among the forty transcripts identified as the most important predictors (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e), only two (ENST00000647913.2 and MSTRG.36782.15) were on sex chromosomes. Of the top 40 transcripts chosen for sex classification using autosomal genes alone, seven were related to immune functions, while twelve were associated with transcriptional and post-transcriptional regulation (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eWhen we used all chromosomes (dataset chr_aXY), the model correctly identified 335 of 338 females (99%) and 328 of 329 males (99.9%). Using the X chromosome and autosomes (dataset chr_aX), the model accuracy was similar: the model correctly classified 336 females (99%) and 327 males (99%). The model based on the Y chromosome and autosomes (dataset chr_aY) resulted in the accurate identification of 335 females (99%) and 328 males (99.9%). However, inference relying solely on autosomal data yielded significantly lower accuracy, correctly classifying only 246 females (73%) and 240 males (73%) (Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). These findings suggest that errors in sex identification probably stem from overlapping expression patterns of the autosomal transcripts between sexes. Therefore, including additional informative predictors may improve model accuracy, especially when the differences in transcript expression between males and females are minor.\u003c/p\u003e\u003cp\u003eSince not all differentially expressed transcripts were used to build our model, we were curious to see which genes and gene families show the highest differential expression levels between sexes. We analyzed sex-biased gene expression by ranking differentially expressed genes based on False Discovery Rate (FDR)-corrected t-test p-values. Among the 786 identified sex-biased transcripts, male-specific Y-chromosomal genes ranked highest. They were followed by female-biased genes, mainly located on the X chromosome.\u003c/p\u003e\u003cp\u003eAnalysis of autosomal features revealed female-biased expression of several pseudogenes (RPS4XP13, ENSG00000253899, RPS4XP2, RPS4XP11, RPS4XP6, RPS4XP3, RPS4XP17) and three protein-coding genes (PPFIA3, EIF2S3B, LMTK3). Top male-biased autosomal loci included LINC01597, ENSG00000291100, MARVELD1, PLEK, IMPACT, ENSG00000280435, ZPBP2, PHETA2, RAB38, AKIRIN1, NQO1, and DDX43. Additionally, several genes showed isoform-level sex biases, notably MSL3, RBM4, and BLOC1S2.\u003c/p\u003e\u003cp\u003eNext, we assessed whether a model trained on blood RNA-seq data could predict sex based on gene expression in other tissues. Of the 786 transcripts identified as differentially expressed between sexes in blood, 574 (73%) showed the same direction of sex-biased expression in the frontal cortex (BRAIN0 dataset). Among these, 249 transcripts were significantly differentially expressed in both tissues, indicating consistent sex bias. Conversely, 140 transcripts exhibited discordant sex-biased expression between tissues, including seven transcripts with significant opposite biases. Notably, FRMD4A and a novel transcript overlapping ENSG00000286388 showed male-biased expression in blood but female-biased expression in the frontal cortex. Conversely, genes such as PRDX2, ENSG00000285756, ENSG00000262202, MSTRG.29194, and MSTRG.35408.4 displayed female bias in blood but male bias in the frontal cortex. These results highlight the complexity of sex-biased gene regulation across tissues and emphasize the importance of tissue-specific analyses to accurately interpret sex differences in gene expression (see Supplementary Material Table\u0026nbsp;1 for detailed gene annotations and references).\u003c/p\u003e\u003cp\u003e\u003cb\u003eSex Determination in Frontal Cortex (BRAIN0 Dataset)\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWe evaluated the performance of the sex determination model (trained on BLOOD RNA-Seq data) using the frontal cortex gene expression dataset with 25 men and 9 women. The model achieved perfect accuracy (1.0) for datasets containing all chromosomes (chr_aXY) using only two transcripts (MSTRG.36782.15, ENST00000611750.1), the X chromosome plus autosomes (chr_aX, one transcript - MSTRG.36020.14), and the Y chromosome plus autosomes (chr_aY, one transcript - MSTRG.36782.15), see Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. Overall, the model demonstrated consistently high performance using all chromosomes, attaining near-perfect accuracy. In contrast, predictions based only on autosomal transcripts were less accurate and more variable. For autosomal predictions, performance metrics (ROC AUC, accuracy, F1 score, precision, and recall) were around 0.6 for models with one transcript, gradually increasing with the number of features, and plateauing at approximately 0.76 for 46 transcripts. For the three models with sex chromosomes, the sex of all 25 men and nine women was determined correctly. The model was based solely on autosomes, which identified the sex of 23 men and three women correctly but misclassified two men and six women (see Figure S3).\u003c/p\u003e\u003cp\u003e\u003cb\u003eSex Classification Model Performance using the HEART Dataset\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWhen tested on the HEART dataset (comprising 9 men and 11 women), the BLOOD-trained model demonstrated excellent accuracy. Specifically, the model achieved perfect accuracy (1.0) for datasets involving both sex chromosomes (chr_aXY) using two transcripts (ENST00000647913.2, MSTRG.36782.15), the X chromosome plus autosomes (chr_aX, one or two transcripts - ENST00000647913.2), and the Y chromosome plus autosomes (chr_aY, one transcript - MSTRG.36782.15). The prediction accuracy for the autosome-only dataset was significantly lower (0.32) and required 22 transcripts (see Supplemental Table S3, Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eThe sex of only four out of nine males (44%) and two out of eleven females (18%) was correctly predicted, resulting in a high misclassification rate (Figure S2). When considering the heart chamber (left/right atrium/ventricle), we observe that the model's accuracy heavily depends on this covariate. For example, all female samples are misclassified except for the two from the left ventricle (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). In males, the situation is reversed: three left ventricular samples are misclassified, suggesting that in the left ventricle, gene expression in males and females resembles female gene expression in whole blood samples, while other chambers exhibit specific patterns.\u003c/p\u003e\u003cp\u003e\u003cb\u003eSex Classification Performance in the Putamen, Caudate, and Nucleus accumbens (BRAIN1) Dataset\u003c/b\u003e\u003c/p\u003e\u003cp\u003eThe tested BRAIN1 dataset included 150 men and 65 women. The predictive performance of the sex determination model, trained on the BLOOD dataset, was similarly accurate when tested on the BRAIN1 dataset (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). The highest accuracy (1.0) was achieved for chr_aXY using three transcripts (ENST00000361365.7, ENST00000602495.1, MSTRG.36020.14), chr_aX (two transcripts ENST00000647913.2, MSTRG.36020.14), and chr_aY (one transcript ENST00000361365.7). Predictions relying solely on autosomal transcripts showed reduced accuracy of 0.73 and required 29 transcripts (see the Transcript Description Table S4 in Supplemental Materials).\u003c/p\u003e\u003cp\u003eFor datasets chr_aXY and chr_aX, the model correctly identified all 65 women and 148 out of 150 men, misclassifying two men as women. The dataset chr_aY correctly identified all 65 women and 147 out of 150 men, with three men misclassified as women. Predictions based solely on autosomes resulted in the highest misclassification rate: only 132 men and eight women were correctly identified, while 18 men and 57 women were incorrectly classified (see Figure S6).\u003c/p\u003e\u003cp\u003eThe accuracy of the autosomal model prediction for the BRAIN1 dataset (Putamen, Caudate, Nucleus accumbens) was slightly lower than for the BRAIN0 dataset (Frontal cortex), at 0.65 versus 0.76 (p-value\u0026thinsp;=\u0026thinsp;0.096). This difference is marginally significant; larger sample sizes are needed to examine variation in sex determination across brain regions.\u003c/p\u003e\u003cp\u003eIf the effect is real, it may be due to the influence of ovarian hormones such as 17β-estradiol (E2), which actively regulate the dopamine system by boosting dopamine cell activity and neurotransmitter release. \u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e It is important to note that sex differences have been observed in the prefrontal cortex; for example, females show higher levels of synaptosomal GluA1 and GluA2 subunits in the medial prefrontal cortex, as the striatum appears to be especially sensitive to these hormonal effects. \u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eComparison of Models and Transcript Expression between Tissues\u003c/b\u003e\u003c/p\u003e\u003cp\u003eIt is interesting to see whether the most significant features are consistent across different organs. To compare pairs of organs, we used UpSet plots. \u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e We examined the top 50 transcripts for each organ and across different sex chromosome and autosome combinations. Figure\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e shows the UpSet plot, which presents influential transcripts ranked by SHAP importance scores across four chromosome groups: chr_aXY, chr_aX, chr_aY, and autosomes.\u003c/p\u003e\u003cp\u003eA subset of genes (DCXR, HIVEP3, MS4A7, PRKCA, TCF15, TIE1, TSFM, and XIST) showed sex-specific expression across all four examined organs (BLOOD, BRAIN0, HEART, and BRAIN1), indicating a common biological signal relevant for sex determination. The persistent presence of these genes in organ-specific models suggests they exhibit sex-specific expression patterns and are key players in conserved regulatory pathways and tissue homeostasis.\u003c/p\u003e\u003cp\u003eOne of these genes, XIST, encodes a well-established long noncoding RNA (lncRNA) involved in X-chromosome inactivation in female mammals. Its expression remains consistently high across female-derived tissues, making it a reliable marker for sex classification. \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eTSFM, which encodes a mitochondrial elongation factor essential for protein synthesis, emphasizes the housekeeping role of mitochondrial function. \u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e Its widespread presence in metabolically active tissues may explain its frequent use as a predictive feature, mainly because mitochondrial activity varies between sexes in metabolic tissues.\u003c/p\u003e\u003cp\u003eDCXR encodes a multifunctional reductase involved in carbohydrate metabolism and renal osmoregulation. Its expression has been observed in various tissues, including the kidney, liver, and testis. It may be affected by hormonal regulation or metabolic differences between sexes. \u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eSimilarly, HIVEP3 is a transcription factor that modulates NF-κB-mediated transcription, a pathway widely linked to sex-biased immune responses and inflammation. This potentially connects this gene to sexually dimorphic gene regulation even outside immune tissues. \u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eGenes involved in vascular and developmental signaling are also sex-specific across multiple organs. TIE1, a receptor tyrosine kinase involved in angiogenesis that responds to VEGF signaling, is known to be influenced by sex hormones such as estrogen, particularly in endothelial contexts. \u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003ePRKCA, which encodes protein kinase C alpha, plays vital roles in cell proliferation, cardiac function, and inflammation. It has been shown to exhibit sex-differentiated expression patterns in cardiac and immune tissues, likely due to differential activation by steroid hormones. \u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eTCF15, a helix-loop-helix transcription factor necessary for early mesodermal development, was also found across various tissues. Its involvement in early patterning events like somitogenesis could explain inherent differences in developmental gene networks between male and female embryos. \u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e Finally, MS4A7, a member of the membrane-spanning 4-domain family, is enriched in the myeloid lineage and linked to macrophage activation and sex differences related to aging in immune regulation. \u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eSex-specific genes perform core metabolic and structural functions (e.g., TSFM, DCXR), maintain tissue-specific regulatory mechanisms (e.g., TIE1, PRKCA), and play sex-specific developmental and immune roles (e.g., XIST, HIVEP3, MS4A7). Their consistent detection across various tissues indicates they form a biologically meaningful and universal feature set for sex classification across multiple organs and tissue types.\u003c/p\u003e\u003cp\u003eOur findings demonstrate that RNA-seq data reliably predicts biological sex across various organs. Models incorporating sex chromosome transcripts consistently show near-perfect accuracy, with sex-linked genes such as XIST, KDM5D, and EIF1AY playing key roles, emphasizing their part in sex determination and X chromosome dosage compensation. Predictions based solely on autosomal transcripts are also possible, albeit less accurate, suggesting sex-specific regulatory effects extend beyond sex chromosomes. Autosomal genes involved in immune response, transcription regulation, and metabolism emerge as important predictors, indicating significant hormonal and epigenetic influences. These results reveal that sex-related gene expression differences extend beyond sex chromosomes, reflecting overall physiological distinctions that could influence clinical approaches, including personalized therapies and disease management.\u003c/p\u003e\u003cp\u003e\u003cb\u003eAge determination\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWe employed Support Vector Regression (SVR) with a linear kernel from the Scikit-learn package \u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e to predict age in male and female samples in the datasets with age information (HEART, BRAIN0, and BRAIN1). Separate models were trained for each dataset and sex. Influential transcripts were identified based on the absolute value of the Spearman rank correlation coefficient between gene expression and age (p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05). To determine the optimal number of transcripts for each dataset and sex, we applied the Elbow method to each dataset and sex.\u003c/p\u003e\u003cp\u003eIn the male HEART dataset, eight transcripts achieved an RMSE of 2.57. Three transcripts in the female HEART dataset resulted in an RMSE of 1.47. For the male BRAIN0 dataset, ten transcripts yielded an RMSE of 3.49. Five transcripts in the female BRAIN0 dataset led to an RMSE of 2.8. The male BRAIN1 dataset, with fourteen transcripts, showed an RMSE of 4.92, while the female BRAIN1 dataset had ten transcripts and an RMSE of 5.38. These results suggest that predicting female age may require fewer features than predicting male age, possibly due to more pronounced age-dependent hormonal changes in females.\u003c/p\u003e\u003cp\u003eScatter plots illustrating the comparison between predicted and reported ages for each dataset and sex are displayed in the Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e, with males on the left and females on the right. In the HEART dataset, predicted ages closely match actual ages, with only slight deviations in older individuals. This dataset shows the lowest root mean square error (RMSE), indicating high prediction accuracy: 0.98 years for males and 1.46 years for females. Predictions for the BRAIN0 dataset were less precise, especially for females, with an RMSE of 3.95 years for males and 8.89 years for females. The BRAIN1 dataset showed moderate accuracy, with RMSE values of 6.00 years for males and 6.99 years for females. The BRAIN1 dataset demonstrated intermediate accuracy, with RMSE values of 6.00 years for males and 6.99 years for females. See Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e for performance results.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTable Results of prediction of age by SVR models. MAE - mean absolute error, RMSE - root mean squared error, R\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e - coefficient of determination\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"8\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDataset\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSex\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003e#Samples\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003e#Transcripts\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eMAE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eRMSE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eR\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHEART\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003emale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSVR, LeaveOneOut CV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.98\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e\u003cp\u003e0.93\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHEART\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003efemale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e11\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSVR, LeaveOneOut CV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e1.15\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e1.46\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e\u003cp\u003e0.92\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBRAIN0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003emale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e25\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSVR, LeaveOneOut CV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e7\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e3.02\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e3.95\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e\u003cp\u003e0.68\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBRAIN0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003efemale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSVR, LeaveOneOut CV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e7\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e6.41\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e8.89\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e\u003cp\u003e0.69\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBRAIN1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003emale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e150\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSVR, 5-fold CV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e17\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e4.57\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e6.00\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e\u003cp\u003e0.45\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBRAIN1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003efemale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e65\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSVR, 5-fold CV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e5.57\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e6.99\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e\u003cp\u003e0.57\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eInterestingly, some genes show opposite relationships with age in males and females. A notable example is the X-ray repair cross-complementing 4 (XRCC4) gene, which we will discuss in detail. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, XRCC4 expression decreases with age in males but increases in females. XRCC4 plays a key role in DNA repair by supporting the non-homologous end-joining (NHEJ) repair pathway for double-strand breaks. This repair process is essential for maintaining genomic stability in postmitotic cells, such as neurons, which have limited DNA replication capacity and thus depend heavily on efficient DNA repair. DNA repair ability typically declines with age, and mutations or altered expression of repair genes like XRCC4 have been linked to faster cellular aging and higher risks of cancer and age-related diseases due to genomic instability and mutation buildup \u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e. Because XRCC4 is vital in NHEJ-mediated DNA repair, its misregulation can significantly compromise genomic integrity. Misregulation of XRCC4 is connected to neurodegenerative disorders, which become manifested with aging. Since neurons rarely divide, they rely on effective DNA repair pathways to prevent damage accumulation. Faulty XRCC4 repair processes could cause increased neuronal cell death, something seen in neurodegenerative diseases such as Alzheimer's, where aging is a significant risk factor. \u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e Additionally, decreased XRCC4 expression may accelerate cellular aging by allowing persistent DNA damage to accumulate, potentially exarcerbating chronic inflammation and contributing to tissue deterioration with age. Aging tissues commonly show an increase in senescent cells; thus, impaired XRCC4 function might hasten this shift to senescence, further driving tissue decline as we age. \u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003e​ The observed sex-specific differences in XRCC4 expression with age may result from various factors, including hormonal regulation, epigenetic changes, inflammatory responses, oxidative stress, metabolism, cellular aging, and evolutionary adaptations. Estrogen and androgens influence DNA repair processes differently. Estrogen has been shown to enhance DNA double-strand break (DSB) repair, potentially increasing XRCC4 expression in females. On the other hand, androgens, through the androgen receptor (AR), directly bind to and boost XRCC4 promoter activity, promoting its transcription in males. Age-related declines in these hormones could differently affect XRCC4 expression between the sexes. Sex-specific epigenetic changes, including DNA methylation and histone modifications, can lead to differential gene expression. Studies show that hormonal exposures during critical periods can cause lasting epigenetic modifications, potentially influencing XRCC4 expression differently in males and females. ​\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. Chronic inflammation shows sex-specific patterns, with females often exhibiting increased inflammatory responses. These differences can affect gene expression in DNA repair, including XRCC4, which contributes to the observed sex-based variations. \u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e Sex-based metabolic differences can lead to varying levels of oxidative stress, which, in turn, impact DNA damage and repair mechanisms. These metabolic differences may lead to different regulation of XRCC4 expression between men and women. Evidence suggests that women might have a higher susceptibility to DNA damage and a greater likelihood of developing senescence. This could result in sex-specific regulation of DNA repair genes, such as XRCC4. \u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e Evolutionary pressures have shaped sex-specific strategies for DNA repair and maintenance. These adaptations may have resulted in inherent differences in how XRCC4 is expressed and regulated between males and females.\u003c/p\u003e\u003cp\u003e\u003cb\u003ePopulation Differentiation and Ancestry Prediction from Genomic and Transcriptomic Data\u003c/b\u003e\u003c/p\u003e\u003cp\u003eThe 1000 Genomes Project database contains whole-genome DNA sequencing and blood RNA-seq measurements from multiple populations, including the Finnish (FIN), Yoruba (YRI), Tuscans (TSI), Utah residents with Northern and Western European ancestry (CEU), and British (GBR). Principal Component Analysis (PCA) plots (Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e) clearly distinguish Yoruba, Tuscans, and Finnish individuals from British and Utah residents. Admixture profiles demonstrate a significant overlap between CEU and GBR populations (Fig.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e: panels A, B, and C). Using GPS classification \u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e, 34\u0026ndash;56% of CEU samples were classified as GBR, and conversely, 33\u0026ndash;45% of GBR samples as CEU, depending on the number of admixture components selected. The highest average prediction accuracy across populations was observed with K\u0026thinsp;=\u0026thinsp;12 admixture components (Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, column \"Whole-genome SNPs\").\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eAccuracy of Ancestry Inference using different datasets. The whole-genome column corresponds to the prediction accuracy, measured as the proportion of individuals in each population attributed to this population by GPS applied to whole-genome SNPs. The RNA-Seq SNPs column corresponds to the prediction accuracy, measured as the proportion of individuals in each population attributed to this population by GPS applied to RNA-Seq SNPs.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePOP\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eWhole-genome SNPs\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRNA-Seq SNPs\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eRNA-Seq Expression\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eRNA-seq Expression SNPs\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eCEU\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.622\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.926\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.986\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.975\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eFIN\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e1.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.711\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.837\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.877\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGBR\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.644\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.470\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.839\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.852\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eTSI\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e1.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.575\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.771\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.779\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eYRI\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e1.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e1.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.986\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e1.000\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eAverage\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.853\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.736\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.856\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.897\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eWe further utilized RNA-seq data from the same individuals to infer ancestry using machine learning with Gradient Boosting Machines (GBM). A cross-validation approach (training with 80% of the samples and testing with the remaining 20%, repeated 500 times) demonstrated high accuracy, with GBM correctly identifying ancestry for 99% of CEU, 86% of FIN, 84% of GBR, 77% of TSI, and 98.6% of YRI samples (Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). Overall prediction accuracy using RNA expression was 87%, comparable to the 86% achieved using whole-genome SNP data. However, error distributions varied between methods. Whole-genome SNP data accurately distinguished between Finnish, Yoruba, and Tuscan samples but often confused the CEU with the British. Conversely, expression-based analysis better differentiated CEU from GBR, though it was less precise among Europeans from different ethnic origins. Analysis with PyLAE \u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e and GPS did not reveal significant genomic regions associated with misclassification, indicating that local ancestry admixture did not significantly affect gene expression differences between misannotated samples.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eConfusion matrix for ancestry inference using RNA-Seq\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eCEU\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eFIN\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eGBR\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eTSI\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eYRI\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eCEU\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.986\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.01\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.002\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.002\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eFIN\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.031\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.838\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.037\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.088\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.006\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGBR\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.049\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.839\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.096\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.014\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eTSI\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.021\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.098\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.083\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.771\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.027\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eYRI\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.005\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.007\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.986\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eWe evaluated the capability of RNA-seq to detect single-nucleotide polymorphisms (SNPs) for ancestry prediction through PCA and admixture analyses. Among 173 individuals with multiple RNA-seq samples, predictions matched perfectly across samples for 171 individuals. However, different samples yielded conflicting ancestry predictions for two individuals (one CEU and one GBR) due to substantial differences in the number of detected SNPs, exemplified by individual NA12399 (49,302 vs. 15,236 SNPs in the two samples). Overall, SNP-based ancestry RNA-seq predictions demonstrated lower accuracy compared to whole-genome sequencing or RNA expression-based predictions. Combining expression-based and SNP-based methods improved accuracy beyond either method alone (Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eNext, we investigated which genes contribute most significantly to distinguishing ethnic groups. KEGG and Gene Ontology (GO) annotations (Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e7\u003c/span\u003e, Table\u0026nbsp;\u003cspan refid=\"Tab8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, Table\u0026nbsp;\u003cspan refid=\"Tab9\" class=\"InternalRef\"\u003e9\u003c/span\u003e, Table\u0026nbsp;\u003cspan refid=\"Tab10\" class=\"InternalRef\"\u003e10\u003c/span\u003e) revealed enrichment in immune-related pathways. This enrichment likely explains why RNA-seq-based methods provided better separation between British and Utah residents than whole-genome sequencing. Individuals relocating from North-Western Europe (e.g., UK) to Utah transitioned from a temperate maritime climate to a semi-arid continental climate, encountering substantial environmental and dietary changes. These included colder and snowier winters (average temperatures dropping from 45\u0026deg;F to 35\u0026deg;F), hotter summers (increasing from 70\u0026deg;F to 80\u0026deg;F), higher elevation (~\u0026thinsp;4,200 feet compared to ~\u0026thinsp;35 feet), clearer skies, reduced precipitation (by ~\u0026thinsp;10%), and calmer wind conditions. These climatic and environmental changes may have driven rapid epigenetic adaptations in immune-related genes, which respond more quickly than genetic mutations to environmental pressures.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eKEGG classification of genes differentially expressed between British and Utah residents.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eKEGG ID\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDescription\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eP-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAdjusted P-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eNumber of genes\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05323\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRheumatoid arthritis\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.04E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2.97E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e11\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04640\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHematopoietic cell lineage\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.48E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2.97E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e11\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05310\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAsthma\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.71E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2.97E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e7\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05321\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eInflammatory bowel disease\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e6.41E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e3.38E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04672\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIntestinal immune network for IgA production\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e7.52E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e3.38E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04659\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTh17 cell differentiation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e8.46E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e3.38E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e11\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05330\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAllograft rejection\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.62E-06\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e5.57E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e7\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04612\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAntigen processing and presentation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.82E-06\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.0001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04940\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eType I diabetes mellitus\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.89E-06\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.0001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e7\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05332\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eGraft-versus-host disease\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e4.56E-06\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00011\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e7\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05416\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eViral myocarditis\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.07E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00023\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04658\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTh1 and Th2 cell differentiation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.22E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00024\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04141\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eProtein processing in endoplasmic reticulum\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.28E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00024\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05320\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAutoimmune thyroid disease\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.63E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00028\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e7\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05150\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eStaphylococcus aureus infection\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.40E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00036\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05140\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLeishmaniasis\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.43E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00036\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04145\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003ePhagosome\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00016\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00232\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e10\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05166\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHuman T-cell leukemia virus 1 infection\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00018\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00242\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05169\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEpstein-Barr virus infection\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00032\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00392\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e11\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05145\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eToxoplasmosis\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00033\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00392\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05152\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTuberculosis\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00049\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.00564\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e10\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05164\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eInfluenza A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00144\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.01511\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa05322\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSystemic lupus erythematosus\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00145\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.01511\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa01232\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNucleotide metabolism\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00206\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.02058\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e6\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04514\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eCell adhesion molecules\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00311\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.02987\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ehsa04061\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eViral protein interaction with cytokine and cytokine receptor\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.00464\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.04283\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e6\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab8\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTop 5 GO categories for biological process ontology\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO ID\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDescription\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eP-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAdjusted P-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eNumber of genes\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0002377\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eimmunoglobulin\u003c/p\u003e\u003cp\u003eproduction\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.57E-16\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e8.16E-13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e25\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0002440\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eproduction of molecular mediator of immune response\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.24E-13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1.97E-10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e27\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0019886\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eantigen processing and presentation of exogenous peptide antigen via MHC class II\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.63E-10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2.78E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0002495\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eantigen processing and presentation of peptide antigen via MHC class II\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e8.76E-10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e5.07E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0002399\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMHC class II protein complex assembly\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e9.57E-10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e5.07E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e7\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab9\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTop 5 GO categories for cellular component ontology\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO ID\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDescription\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eP-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAdjusted P-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eNumber of genes\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0009897\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eexternal side of the plasma membrane\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e5.09E-16\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1.83E-13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e34\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0019814\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eimmunoglobulin\u003c/p\u003e\u003cp\u003ecomplex\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.16E-15\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2.08E-13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e19\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0042613\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMHC class II\u003c/p\u003e\u003cp\u003eprotein complex\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.63E-11\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e4.36E-09\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0042611\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMHC protein\u003c/p\u003e\u003cp\u003ecomplex\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.46E-09\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1.32E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0072562\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood microparticle\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.20E-08\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1.59E-06\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e14\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab10\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 10\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTop 5 GO categories for molecular function ontology\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO ID\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDescription\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eP-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAdjusted P-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eNumber of genes\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0003823\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eantigen binding\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e5.16E-28\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2.46E-25\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e29\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0023026\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMHC class II protein complex binding\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.85E-12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e4.41E-10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e10\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0023023\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMHC protein complex binding\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e4.99E-11\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e7.92E-09\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e10\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0032395\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMHC class II receptor activity\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e7.70E-06\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e9.00E-04\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e4\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGO:0140375\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eImmune receptor activity\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.00E-04\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2.00E-02\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eSex-specific epigenetic adaptations to environmental factors have also been previously reported \u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e. Consistent with this, we identified transcripts with differential expression between CEU and GBR separately for males and females. Fourteen transcripts were expressed at higher levels in Utah females than British females, and eleven transcripts showed higher expression in Utah males than British males. Five transcripts (IGF2BP3, CIBAR1P1, GLDCP1, ITPKB, and PIK3CD-AS1) exhibited higher expression in CEU than in GBR for both sexes. IGF2BP3 is involved in growth regulation through binding insulin-like growth factor 2 mRNA. CIBAR1P1 and GLDCP1 are pseudogenes potentially involved in gene regulation and glucose metabolism. ITPKB participates in immune signaling through inositol phosphate pathways, while PIK3CD-AS1 regulates the immune-associated PIK3CD gene.\u003c/p\u003e\u003cp\u003e78 transcripts showed reduced expression in Utah females compared to British females, and 97 transcripts were lower in Utah males than in British males. Among these, 44 transcripts (including ZNF22, ABTB2, FOXO4, IL2RA, IFITM3, and KCNA3) were consistently lower in Utah males compared to British individuals of both sexes. These genes are involved in vital processes including transcriptional regulation, cellular stress responses, inflammation, immune responses, metabolism, and cellular proliferation. Variations in the expression of these genes may indicate how the immune system adapts to Utah's environment compared to the UK.\u003c/p\u003e\u003cp\u003eRapid epigenetic modifications such as DNA methylation and histone modifications likely facilitated quick adaptation to climatic shifts, dietary changes, and pathogen exposure. For instance, populations like the Yakuts and Inuit demonstrate rapid epigenetic regulation of immune genes in response to Arctic environments \u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e,\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. Genetic drift and natural selection may have further contributed to these adaptations, as observed in high-altitude populations adapting to low oxygen levels \u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eThe identified differences in immune gene expression between Utah and British populations may reflect differential adaptations to specific pathogens or stressors unique to each environment. For example, genes such as FOXO4 and IL2RA, crucial regulators of inflammation and immune response, showed significant expression differences, possibly due to varying infectious challenges between regions.\u003c/p\u003e\u003cp\u003eWe further explored gene expression differences in GO categories between males and females from Europe (GBR) and America (CEU), selecting categories containing five or more genes. The most significant differences were observed in categories related to immune functions, including immunoglobulin complexes (GO:0019814), peptidoglycan binding (GO:0042834), and antigen binding (GO:0003823), with consistently higher expression in GBR compared to CEU.\u003c/p\u003e\u003cp\u003eInterestingly, certain genes showed opposite expression changes between sexes across populations. For example, HOMER1 (which inhibits T-cell activation) and SMN1 were more expressed in American females and British males. In contrast, CSNK1A1 (a regulator of mTOR signaling and inflammasome assembly) and FNDC3B displayed opposite patterns, being higher in American males and British females. \u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e,\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e These complex expression dynamics highlight sex-specific responses to environmental and evolutionary pressures shaping population differentiation.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eThis study highlights the significant potential of RNA-Seq data ML-empowered analysis for accurately predicting complex biological traits, including sex, age, and ethnicity. It provides deep insights into how genetic and environmental factors affect gene expression. We successfully identified key transcripts and biological pathways associated with these traits by employing ML approaches, such as GBM, SVR, and SHAP. This highlights the critical role of transcriptomics in population genetics and personalized medicine.\u003c/p\u003e\u003cp\u003eSex prediction using RNA-Seq data demonstrated remarkable accuracy across multiple tissues, emphasizing the pivotal role of sex chromosomes. Transcripts such as XIST, crucial for X-chromosome inactivation, and Y-linked genes, including KDM5D and EIF1AY, emerged as particularly influential. These findings reinforce RNA-Seq\u0026rsquo;s capability to distinguish between biological males and females robustly. Notably, autosomal and sex-chromosomal genes significantly contributed to achieving precise sex classifications.\u003c/p\u003e\u003cp\u003eAge prediction based on RNA-Seq displayed greater variability across different tissues but produced promising outcomes. Notably, heart tissue provided the most accurate age estimations, particularly in females, suggesting substantial age-related shifts in gene expression within specific tissues. SVR models effectively identified transcripts indicative of biological aging, highlighting their potential as biomarkers for monitoring age-related physiological changes and associated diseases.\u003c/p\u003e\u003cp\u003eA significant achievement of our study is the demonstrated capability of RNA-Seq data to accurately differentiate between genetically related populations, such as CEU and GBR. Immune-related pathways played a crucial role in this differentiation, with genes such as IL2RA and FOXO4 being prominently featured. These differences likely represent adaptive responses to distinct environmental pressures, such as climate variations and pathogen exposure, underscoring the responsiveness of immune genes to environmental factors.\u003c/p\u003e\u003cp\u003eKEGG and Gene Ontology analyses further highlighted immune function pathways, including Th17 cell differentiation and intestinal immune networks, supporting the hypothesis of rapid epigenetic adaptations driven by environmental pressures. For instance, the distinct climatic conditions of Utah compared to the UK appear to have influenced differential gene expression, especially within immune-related pathways, reflecting adaptive physiological responses.\u003c/p\u003e\u003cp\u003eOur findings highlight the benefits of integrating transcriptomic data with other genomic resources, such as whole-genome and whole-transcriptome sequencing, thereby enhancing predictive accuracy. This integrated approach provides a more comprehensive understanding of the genetic and environmental influences, thereby enhancing the precision of ethnicity classification and elucidating the molecular mechanisms underlying population-specific adaptations.\u003c/p\u003e\u003cp\u003eFuture research should prioritize diversifying datasets by incorporating additional populations exposed to varied environmental contexts, refining predictive models, and broadening their applicability. Additionally, further exploring the interactions between genetic variants and epigenetic modifications will deepen our understanding of gene expression regulation by inherited and environmental factors. Expanding this integrated multi-omics approach has the potential to substantially enhance predictions of disease susceptibility and treatment responses, significantly advancing personalized medicine.\u003c/p\u003e\u003cp\u003e\u003cb\u003eIn conclusion\u003c/b\u003e, our results demonstrate that RNA-Seq data have exceptional potential for predicting complex biological traits and elucidating the intricate relationships between genetics, environment, and gene expression. These predictive capabilities have essential implications for precision medicine, population genetics, and aging studies, enabling the tailoring of therapeutic interventions, clarification of disease risk profiles, and optimization of strategies for healthier aging. While challenges remain, especially in distinguishing genetic and environmental influences in ethnicity prediction, the continued integration of RNA-Seq data with multi-omics resources promises even greater precision in clinical practice and deeper insights into human diversity and adaptive mechanisms.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cp\u003e\u003cb\u003eRNA-Seq Data\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWe used the RNA-Seq dataset from four RNA-Seq experiments, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003cb\u003eWhole-genome Data\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWe obtained admixture \u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e results from the 1000 Genomes Project \u003csup\u003e\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e for admixture component numbers (K) ranging from 5 to 25, available at the official database of the 1000 Genomes Project \u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e. These admixture profiles were derived using 193,634 biallelic markers, including non-singleton single-nucleotide variant (SNV) sites separated by at least 2 kb. From these comprehensive admixture datasets, we selected populations that correspond to our RNA-seq data: Yoruba (YRI), Utah residents with Northern and Western European ancestry (CEU), Finnish (FIN), British (GBR), and Tuscan (TSI). In total, 667 RNA-seq blood samples corresponded to these populations, with 447 samples having matching admixture profiles. Due to multiple RNA-seq measurements available per individual, these 667 samples represented 464 unique individuals.\u003c/p\u003e\u003cp\u003eTo accurately capture and characterize population structure, we clustered the admixture profiles within each ethnic group, identifying distinct subpopulations, such as FIN_1 and FIN_2. Additionally, outlier detection was performed to remove samples that appeared non-representative or potentially mislabeled. To assess the accuracy of ancestry label reconstruction based on genotyping data, we employed a leave-one-out validation approach combined with the Geographic Population Structure (GPS) algorithm \u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e. Predictions were correct if an individual's predicted subpopulation matched at least one closely related subpopulation from their assigned group (for example, an individual from subpopulation FIN_2 correctly predicted as FIN_1 was considered accurately classified).\u003c/p\u003e\u003cp\u003e\u003cb\u003eRNA-Seq Data Processing\u003c/b\u003e\u003c/p\u003e\u003cp\u003eFastq files were downloaded using the fastq-dump and fasterq-dump programs \u003csup\u003e\u003cspan additionalcitationids=\"CR53\" citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e. FastQC \u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e,\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e was used to assess the quality of the reads, and adapter trimming and removal of low-quality nucleotides were performed with TrimGalore \u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e. The reads were aligned to the human genome hg38 (GRCh38.primary_assembly.genome.fa) with STAR \u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e,\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e\u003c/sup\u003e, and duplicated reads were marked with Picard \u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e. Transcript expression was quantified with \u003cem\u003estringtie\u003c/em\u003e \u003csup\u003e\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e\u003c/sup\u003e using Gencode v44 gene annotation.\u003c/p\u003e\u003cp\u003eThe batch effect was removed with ComBat-seq \u003csup\u003e61\u003c/sup\u003e. Custom R and Python scripts were used to collect TPM matrices and visualize the results. Functional annotation and GO terms were obtained using DAVID \u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e\u003c/sup\u003e. The resulting transcripts were merged to create a superset of 380K transcripts. RNA-seq data from all experiments was mapped to this new reference transcriptome.\u003c/p\u003e\u003cp\u003e\u003cb\u003eVisualization\u003c/b\u003e\u003c/p\u003e\u003cp\u003eThe results were visualized using custom R and Python scripts and downloadable packages, such as SuperVenn \u003csup\u003e\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003e\u003cb\u003eSex determination\u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eData preprocessing\u003c/b\u003e\u003c/p\u003e\u003cp\u003eThis collection contains RNA-seq data from blood generated by the Geuvadis consortium \u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e for 464 samples from the 1000 Genomes project, including samples from five of the 1000 Genomes populations (CEU, FIN, GBR, TSI, and YRI). Preprocessing of the BLOOD dataset was performed using the following steps:\u003c/p\u003e\u003cp\u003e\u003col\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eRemove transcripts with zero median expression in both sexes separately.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eSelect transcripts that are highly correlated with sex (Point-biserial correlation coefficient\u0026thinsp;\u0026gt;\u0026thinsp;0.8);\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eApply \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\text{log}}_{2}\\left(x+1\\right)\\)\u003c/span\u003e\u003c/span\u003e transformation;\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eExclude transcripts with low variability, as indicated by a coefficient of variation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\frac{\\sigma\\:}{\\mu\\:}\u0026lt;0.7\\)\u003c/span\u003e\u003c/span\u003e;\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eRemove poorly expressed transcripts with mean expression below the median(mean expression) \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\mu\\:}_{i}\u0026lt;median\\left(\\mu\\:\\right)\\)\u003c/span\u003e\u003c/span\u003e;\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eFilter out transcripts with a coefficient of variation below the median coefficient of variation across all transcripts.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003c/ol\u003e\u003c/p\u003e\u003cp\u003eFour datasets were created:\u003c/p\u003e\u003cp\u003e\u003col\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eChr_aXY: transcripts from X, Y chromosomes, and autosomes\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eChr_aX: transcripts from the X chromosome and autosomes\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eChr_aY: transcripts from X and Y chromosomes and autosomes\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eAutosomes\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003c/ol\u003e\u003c/p\u003e\u003cp\u003eTranscripts located in the pseudoautosomal regions of sex chromosomes (Y chromosome: [10001..2781479], [56887903..57217415] and X chromosome [10001..2781479], [155701383..156030895]) were removed from Chr_aX and Chr_aY. The preprocessing of the HEART, BRAIN0, and BRAIN1 datasets involved two steps: removing transcripts with zero median expression, followed by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\text{log}}_{2}\\left(x+1\\right)\\)\u003c/span\u003e\u003c/span\u003e transformation. Table\u0026nbsp;\u003cspan refid=\"Tab8\" class=\"InternalRef\"\u003e8\u003c/span\u003e shows the number of transcripts in each dataset.\u003c/p\u003e\u003cp\u003eDue to the varying number of transcripts remaining after the pre-processing step in the BLOOD, HEART, BRAIN0, and BRAIN1 datasets, the models were trained using transcripts overlapping with the BLOOD dataset transcripts.\u003c/p\u003e\u003cp\u003e\u003cb\u003eML model training\u003c/b\u003e\u003c/p\u003e\u003cp\u003eThe XGBoost gradient boosting method was chosen as the classification model. \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e All data was standardized (median-centered and scaled using an IQR interval, according to the formulae \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:z=\\frac{x-median\\left(x\\right)}{IQR\\left(x\\right)}\\)\u003c/span\u003e\u003c/span\u003edue to the non-normality of expression counts, using the RobustScaler routine from the scikit-learn library \u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eThe model was trained using 5-fold cross-validation using the BLOOD dataset. In the evaluation phase (BLOOD, HEART, BRAIN0, and BRAIN1 datasets), the average model prediction of each fold was calculated.\u003c/p\u003e\u003cp\u003eThe transcripts in each of the 16 datasets were ranked according to their influence on model predictions using the SHapley Additive exPlanations (SHAP) \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. SHAP is a game-theoretical approach to explaining the output of any machine learning model. SHAP assigns each feature an \u0026ldquo;importance score\u0026rdquo; based on its contribution to the model's output. The model was sequentially trained on the most significant transcripts; the number of significant transcripts varied from 1 to 100. ROC AUC, precision, recall, and F1 scores were calculated. Then, we selected the optimal number of transcripts for the final model training, based on maximum of two values: 1) the number from the elbow method using the quality vs. the number of transcript plot, and 2) the number of transcripts for that were in the top 100 SHAP transcripts in the 40 out of 50 cross-validation steps. Due to variations in the number of transcripts retained after this pre-processing stage, we trained models separately for each dataset (HEART, BRAIN0, and BRAIN1) using only those transcripts that overlapped with the BLOOD dataset. This procedure resulted in 16 unique transcript subsets for training machine learning (ML) models. The optimal number of transcripts for each dataset is shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003cb\u003eImportance of transcripts\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo evaluate the importance of each transcript, we trained models on the BLOOD dataset using each of the 16 transcript subsets. Model performance was validated via repeated stratified 5-fold cross-validation with 10 repetitions, implemented using the \u003cem\u003eRepeatedStratifiedKFold\u003c/em\u003e function from the \u003cem\u003escikit-learn\u003c/em\u003e library. This method preserves the class distribution within each fold and ensures robust performance estimates by repeating the process multiple times with different fold partitions. During each cross-validation iteration, we computed SHAP values for feature importance. We selected the top 100 transcripts for every fold with the highest mean absolute SHAP values. Each transcript then received a score from 0 to 50, representing the number of times it appeared in the top 100 across all cross-validation iterations (5 folds \u0026times; 10 repetitions\u0026thinsp;=\u0026thinsp;50 iterations). The final list of the top 100 transcripts with the highest cumulative scores was carried forward to the subsequent analysis stage.\u003c/p\u003e\u003cp\u003e\u003cb\u003eSelecting the Optimal Number of Transcripts\u003c/b\u003e\u003c/p\u003e\u003cp\u003eFor each of the 16 transcript sets, we obtained the top 100 transcripts ranked by cumulative SHAP importance. We then retrained the models on the BLOOD dataset using only these selected transcripts. Again, we applied repeated stratified 5-fold cross-validation with 10 repetitions for model evaluation.\u003c/p\u003e\u003cp\u003e\u003cb\u003eModel performance\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo assess how model performance depends on the number of input features, we varied the number of transcripts used in training from 1 to 100 in incremental steps. For each model configuration (i.e., each number of top transcripts), we calculated standard classification metrics (Accuracy, Precision, Recall, and F1 Score) on every cross-validation fold. The performance trends were plotted for each set, and the Elbow Method was employed to determine the optimal number of transcripts. The Elbow Method identifies the point at which adding more features results in diminishing returns in performance, thereby enabling the selection of a parsimonious yet informative transcript set.\u003c/p\u003e\u003cp\u003e\u003cb\u003eAge determination\u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eData preprocessing\u003c/b\u003e\u003c/p\u003e\u003cp\u003eBLOOD, HEART, BRAIN0, and BRAIN1 datasets were preprocessed using the same protocol as for sex determination, with modified Step 2 - pre-selection of transcripts highly correlated with age (Spearman rank correlation coefficient R\u0026thinsp;\u0026gt;\u0026thinsp;0.8, p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05. All datasets were then split by sex because of the expected different relationship between gene expression levels and age for males and females.\u003c/p\u003e\u003cp\u003eWe used the BRAIN1 dataset \u003csup\u003e\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e\u003c/sup\u003e, which contained 65 female RNA-seq samples (ranging from 21 to 58 years old, mean age 39.77, median age 44) and 150 male samples (ranging from 27 to 62 years old, mean age 46.46, median age 48). The transcripts were ranked based on their absolute correlation value with age, and the top 500 were selected; this dataset was pruned to exclude highly correlated (r\u0026thinsp;\u0026gt;\u0026thinsp;0.5) transcripts, leaving 307 transcripts for females and 421 for males.\u003c/p\u003e\u003cp\u003e\u003cb\u003eML model training\u003c/b\u003e\u003c/p\u003e\u003cp\u003eSeveral linear models were tested: Support Vector Regression (SVR) with linear kernel, Linear Regression, and Elastic Net. SVR demonstrated the best results regarding Root Mean Squared Error (RMSE). Due to the small number of samples in all datasets except BRAIN1, the model was trained using Leave-One-Out cross-validation (for BRAIN1, a 5-fold stratified cross-validation procedure was employed); in the evaluation phase, the average model prediction for each sample was calculated. Random Forest was used to make a model for age prediction:\u003c/p\u003e\u003cp\u003emodel \u0026lt;- train(AGE~., data=(FE)MALES2, method=\"rf\", preProcess=\"scale\")\u003c/p\u003e\u003cp\u003eA 10-fold cross-validation procedure was conducted 30 times, with 80% of the samples used for training and 20% for control.\u003c/p\u003e\u003cp\u003e\u003cb\u003eAncestry determination\u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eExpression-based inference\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWe selected the 500 most informative transcripts from the BLOOD dataset showing differential expression between 10 pairs of populations (YRI, TSI, CEU, FIN, GBR). Next, we utilized the Gradient Boosting Machines (GBM) method, which was implemented as an R package gbm. The gbm() function parameters were selected based on brute-force RMSE minimization. The dataset was split 500 times into 80% for training and 20% for testing, and the linear predictors from all runs were averaged for every sample.\u003c/p\u003e\u003cp\u003e\u003cem\u003emodel_gbm\u0026thinsp;=\u0026thinsp;gbm( train$ethn ~., data\u0026thinsp;=\u0026thinsp;train, distribution = \"multinomial\"\u003c/em\u003e,\u003c/p\u003e\u003cp\u003e\u003cem\u003en.trees\u0026thinsp;=\u0026thinsp;83, interaction.depth\u0026thinsp;=\u0026thinsp;5 shrinkage\u0026thinsp;=\u0026thinsp;0.1, cv.folds\u0026thinsp;=\u0026thinsp;10\u003c/em\u003e,\u003c/p\u003e\u003cp\u003e\u003cem\u003en.minobsinnode\u0026thinsp;=\u0026thinsp;10, bag.fraction\u0026thinsp;=\u0026thinsp;.65, n.cores\u0026thinsp;=\u0026thinsp;NULL, verbose\u0026thinsp;=\u0026thinsp;F )\u003c/em\u003e\u003c/p\u003e\u003cp\u003e\u003cem\u003epred\u0026thinsp;=\u0026thinsp;gbm::predict.gbm(model_gbm, test_x))\u003c/em\u003e\u003c/p\u003e\u003cp\u003eWe averaged the prediction scores for each sample and determined the predicted ancestry as having the highest score.\u003c/p\u003e\u003cp\u003e\u003cb\u003eRNA-Seq Data SNP calling\u003c/b\u003e\u003c/p\u003e\u003cp\u003eBAM files were sorted by coordinates using GATK SortSam, and duplicates were marked with GATK \u003csup\u003e\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e\u003c/sup\u003e MarkDuplicates. Then bcftools mpileup was used to create a bcf file converted by bcftools to the vcf format. Next, the bcftools \u003csup\u003e\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e\u003c/sup\u003e view function filtered out low-quality (QUAL\u0026thinsp;\u0026lt;\u0026thinsp;=\u0026thinsp;10 || DP\u0026thinsp;\u0026lt;\u0026thinsp;10) positions. Vcftools \u003csup\u003e\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e,\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e\u003c/sup\u003e was used to remove indels and genotypes with a quality score below the specified threshold (--remove-indels --minGQ 15). VCF files for individual samples were combined using the bcftools merge function and filtered based on a minor allele frequency (MAF) below 0.05 and a missing data threshold of less than 0.9.\u003c/p\u003e\u003cp\u003e\u003cb\u003eAncestry prediction using a combination of SNP data and expression from RNA-Seq Data\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo combine predictions based on SNPs and expression data derived from RNA-Seq experiments, we used the output of the GBM predict function and converted the GBM output to scores. We identified 50 individuals with the closest admixture vectors (based on Euclidean distance) for each sample to achieve this. Then, we found the relative frequencies of the five ancestry labels (YRI, GBM, TSI, FIN, CEU). The frequencies, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{p}_{i}\\)\u003c/span\u003e\u003c/span\u003e, were converted to scores using the formula \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{p}_{i}\\text{log}\\left(\\frac{{p}_{i}}{0.2}\\right)\\)\u003c/span\u003e\u003c/span\u003e; the scores were added to the GBM \u0026ldquo;predict\u0026rdquo; scores. The predicted ancestry had the highest total scores.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eTrabzuni D, Ramasamy A, Imran S et al (2012) Widespread sex differences in gene expression and splicing in the adult human brain. Nat Commun 4\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKobylińska K, Orłowski T, Adamek M, Biecek P (1926) Explainable machine learning for lung cancer screening models. \u003cem\u003eAppl. Sci. (Basel)\u003c/em\u003e 12, (2022)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWaxman DJ, Connor C (2006) Growth Hormone Regulation of Sex-Dependent Liver Gene 1142 Expression. Mol Endocrinol 20:2613\u0026ndash;2629\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMayne BT et al (2016) Large scale gene expression meta-analysis reveals tissue-specific, sex-biased gene expression in humans. Front Genet 7:183\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKlein SL, Morgan R (2020) The impact of sex and gender on immunotherapy outcomes. Biol Sex Differ 11:24\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLopes-Ramos CM, Quackenbush J, DeMeo DL (2020) Genome-wide sex and gender differences in cancer. Front Oncol 10:597788\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMoqri M, Poganik JR, Horvath S, Gladyshev V (2025) N. What makes biological age epigenetic clocks tick. Nat Aging. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s43587-025-00833-1\u003c/span\u003e\u003cspan address=\"10.1038/s43587-025-00833-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKurochkina NS et al (2024) Age-related changes in human skeletal muscle transcriptome and proteome are more affected by chronic inflammation and physical inactivity than primary aging. Aging Cell 23:e14098\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKim SY, Lee J, Baik M (2024) PSVII-22 Age-related transcriptome changes associated with beef quality in longissimus thoracis muscle of Hanwoo steer. J Anim Sci 102:688\u0026ndash;688\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFleischer JG et al (2018) Predicting age from the transcriptome of human dermal fibroblasts. Genome Biol 19:221\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFranceschi C, Garagnani P, Parini P, Giuliani C, Santoro A (2018) Inflammaging: a new immune-metabolic viewpoint for age-related diseases. Nat Rev Endocrinol 14:576\u0026ndash;590\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJohnson ML, Robinson MM, Nair KS (2013) Skeletal muscle aging and the mitochondrion. Trends Endocrinol Metab 24:247\u0026ndash;256\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePalmer D, Fabris F, Doherty A, Freitas AA (2021) \u0026amp; de Magalh\u0026atilde;es, J. P. Ageing transcriptome meta-analysis reveals similarities and differences between key mammalian tissues. Aging 13:3313\u0026ndash;3341\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHorvath S (2013) DNA methylation age of human tissues and cell types. Genome Biol 14:R115\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFachrul M et al (2023) Direct inference and control of genetic population structure from RNA sequencing data. Commun Biol 6:804\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKachuri L et al (2023) Gene expression in African Americans, Puerto Ricans and Mexican Americans reveals ancestry-specific patterns of genetic architecture. Nat Genet 55:952\u0026ndash;963\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSun S et al (2017) Differential expression analysis for RNAseq using Poisson mixed models. Nucleic Acids Res 45:e106\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShen X et al (2024) Nonlinear dynamics of multi-omics profiles during human aging. Nat Aging 4:1619\u0026ndash;1634\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBrownlee J (2016) XGBoost With Python: Gradient Boosted Trees with XGBoost and Scikit-Learn. Machine Learning Mastery\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi B et al (2018) Genomic prediction of breeding values using a subset of SNPs identified by three machine learning methods. Front Genet 9:237\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBabaei G, Giudici P, InstanceSHAP (2022) An instance-based estimation approach for Shapley values. SSRN Electron J. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2139/ssrn.4293269\u003c/span\u003e\u003cspan address=\"10.2139/ssrn.4293269\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTozzi A et al (2015) Endogenous 17\u0026Icirc;\u003csup\u003e2\u003c/sup\u003e-estradiol is required for activity-dependent long-term potentiation in the striatum: interaction with the dopaminergic system. Front Cell Neurosci 9:192\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShao S et al (2021) Sex-dependent expression of N-cadherin-GluA1 pathway-related molecules in the prefrontal cortex mediates anxiety-like behavior in male offspring following prenatal stress. Stress 24:612\u0026ndash;620\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLex A, Gehlenborg N, Strobelt H, Vuillemot R, Pfister H (2014) UpSet: Visualization of Intersecting Sets. IEEE Trans Vis Comput Graph 20:1983\u0026ndash;1992\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCerase A, Pintacuda G, Tattermusch A, Avner P (2015) Xist localization and function: new insights from multiple levels. Genome Biol 16:166\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi X-Y et al (2024) Short-term regulation of TSFM level does not alter amyloidogenesis and mitochondrial function in type-specific cells. Mol Biol Rep 51:484\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang S et al (2017) Diacetyl/l-xylulose reductase mediates chemical redox cycling in lung epithelial cells. Chem Res Toxicol 30:1406\u0026ndash;1418\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHicar MD, Liu Y, Allen CE, Wu LC (2001) Structure of the human zinc finger protein HIVEP3: molecular cloning, expression, exon-intron structure, and comparison with paralogous genes HIVEP1 and HIVEP2. \u003cem\u003eGenomics\u003c/em\u003e 71, 89\u0026ndash;100\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAbu Shtaya A et al (2023) Possible biallelic inheritance in TIE1 in a family with congenital lymphedema, intestinal lymphangiectasia and cutis aplasia. Clin Genet 104:275\u0026ndash;276\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi Q et al (2015) Two novel PRKCI polymorphisms and prostate cancer risk in an Eastern Chinese Han population. Mol Carcinog 54:632\u0026ndash;641\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang W-Y et al (2015) Coronary risk in relation to genetic variation in MEOX2 and TCF15 in a Flemish population. BMC Genet 16:116\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNi B et al (2023) The short isoform of MS4A7 is a novel player in glioblastoma microenvironment, M2 macrophage polarization, and tumor progression. J Neuroinflammation 20:80\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePedregosa F et al (2012) Scikit-learn: Machine Learning in Python. arXiv [cs.LG]\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGorbunova V, Seluanov A (2016) DNA double strand break repair, aging and the chromatin connection. Mutat Res 788:2\u0026ndash;6\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLombard DB et al (2005) DNA repair, genome stability, and aging. Cell 120:497\u0026ndash;512\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMaynard S, Fang EF, Scheibye-Knudsen M, Croteau DL, Bohr VA (2015) DNA damage, DNA repair, aging, and neurodegeneration. Cold Spring Harb Perspect Med 5\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMcCarthy MM et al (2009) The epigenetics of sex differences in the brain. J Neurosci 29:12815\u0026ndash;12823\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eQi X, Ng TKS, Wu B (2023) Sex differences in the mediating role of chronic inflammation on the association between social isolation and cognitive functioning among older adults in the United States. Psychoneuroendocrinology 149:106023\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNg M, Hazrati L-N (2022) Evidence of sex differences in cellular senescence. Neurobiol Aging 120:88\u0026ndash;104\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eElhaik E et al (2014) Geographic population structure analysis of worldwide human populations infers their biogeographical origins. Nat Commun 5:3513\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSmetanin A, Moshkov N, Tatarinova TV (2020) Local Ancestry Prediction with PyLAE. Cold Spring Harbor Laboratory 11.13.380105 (2020) \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1101/2020.11.13.380105\u003c/span\u003e\u003cspan address=\"10.1101/2020.11.13.380105\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSvensson EI et al (2018) Sex differences in local adaptation: what can we learn from reciprocal transplant experiments? Philos Trans R Soc Lond B Biol Sci 373:20170420\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKalyakulina A et al (2023) Epigenetics of the far northern Yakutian population. Clin Epigenetics 15:189\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhou S et al (2019) Genetic architecture and adaptations of Nunavik Inuit. \u003cem\u003eProc. Natl. Acad. Sci. U. S. A.\u003c/em\u003e 116, 16012\u0026ndash;16017\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLorenzo FR et al (2014) A genetic mechanism for Tibetan high-altitude adaptation. Nat Genet 46:951\u0026ndash;956\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDuan S et al (2011) mTOR generates an auto-amplification loop by triggering the βTrCP- and CK1α-dependent degradation of DEPTOR. Mol Cell 44:317\u0026ndash;324\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGao D et al (2011) mTOR drives its own activation via SCF(βTrCP)-dependent degradation of the mTOR inhibitor DEPTOR. Mol Cell 44:290\u0026ndash;303\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAlexander DH, Novembre J, Lange K (2009) Fast model-based estimation of ancestry in unrelated individuals. Genome Res 19:1655\u0026ndash;1664\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e1000 Genomes Project Consortium (2015) A global reference for human genetic variation. Nature 526:68\u0026ndash;74\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e1000 Genomes Project (2015) 1000 Genomes Project https://ftp\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e.1000genomes.ebi.ac.uk/vol1/ftp/release/20130502/supporting/admixture_files/\u003c/span\u003e\u003cspan address=\"http://.1000genomes.ebi.ac.uk/vol1/ftp/release/20130502/supporting/admixture_files/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eElhaik E, Tatarinova TV, Klyosov AA (2014) The \u0026rsquo;extremely ancient\u0026rsquo;chromosome that isn\u0026rsquo;t: a forensic bioinformatic investigation of Albert Perry\u0026rsquo;s X-degenerate portion of the Y chromosome. Eur J of\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHarrington R, Fastq-dump (2020) Bioinf Noteb \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://rnnh.github.io/bioinfo-notebook/docs/fastq-dump.html\u003c/span\u003e\u003cspan address=\"https://rnnh.github.io/bioinfo-notebook/docs/fastq-dump.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eIsmail HD, Bioinformatics (2022) A Practical Guide to NCBI Databases and Sequence Alignments. CRC\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ewraetz (2023) HowTo: fasterq dump. ncbi/sra-tools \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/ncbi/sra-tools/wiki/HowTo:-fasterq-dump\u003c/span\u003e\u003cspan address=\"https://github.com/ncbi/sra-tools/wiki/HowTo:-fasterq-dump\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAndrews S (2010) \u0026amp; Others. FastQC: a quality control tool for high throughput sequence data\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKrueger F, TrimGalore, GitHub (2012) \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/FelixKrueger/TrimGalore\u003c/span\u003e\u003cspan address=\"https://github.com/FelixKrueger/TrimGalore\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDobin A et al (2013) STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29:15\u0026ndash;21\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHomo sapiens genome assembly GRCh38 NCBI \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/data-hub/assembly/GCF_000001405.26/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/data-hub/assembly/GCF_000001405.26/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBroad Institute (2019) Picard Toolkit. GitHub Repository \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://broadinstitute.github.io/picard\u003c/span\u003e\u003cspan address=\"http://broadinstitute.github.io/picard\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePertea M, Kim D, Pertea GM, Leek JT, Salzberg SL (2016) Transcript-level expression analysis of RNA-seq experiments with HISAT, StringTie and Ballgown. Nat Protoc 11:1650\u0026ndash;1667\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang Y, Parmigiani G, Johnson WE (2020) ComBat-seq: batch effect adjustment for RNA-seq count data. NAR Genom Bioinform 2:lqaa078\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSherman BT et al (2022) DAVID: a web server for functional enrichment analysis and functional annotation of gene lists (2021 update). Nucleic Acids Res 50:W216\u0026ndash;W221\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eIndukaev F, Supervenn (2020) Precise and Easy-to-Read Multiple Sets Visualization in Python\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLappalainen T et al (2013) Transcriptome and genome sequencing uncovers functional variation in humans. Nature 501:506\u0026ndash;511\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBrownlee WE (2016) The World War I Regime. in Federal Taxation in America 93\u0026ndash;123. Cambridge University Press, Cambridge\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKetchesin KD et al (2023) Diurnal alterations in gene expression across striatal subregions in psychosis. Biol Psychiatry 93:137\u0026ndash;148\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMcKenna A et al (2010) The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res 20:1297\u0026ndash;1303\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi H, Handsaker B, Danecek P, McCarthy S, Marshall J. BCFtools\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDanecek P et al (2011) The variant call format and VCFtools. Bioinformatics 27:2156\u0026ndash;2158\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-7256409/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7256409/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eRNA sequencing provides a comprehensive snapshot of gene expression, reflecting genetic inheritance and dynamic environmental influences. This study explores the predictive power of RNA-seq data combined with advanced machine learning techniques, such as Gradient Boosting Machines, Support Vector Regression, and SHapley Additive exPlanations, to infer complex human traits, including biological sex, age, and ethnicity, across diverse tissues. Using RNA-seq datasets derived from blood, heart, and several brain regions, we achieved near-perfect accuracy in sex determination, emphasizing the critical roles of sex chromosome-linked genes (XIST, KDM5D, EIF1AY). Age prediction demonstrated high tissue-specific precision, identifying transcripts indicative of biological aging, particularly those involved in DNA repair and inflammation, which offer promising biomarkers for aging-related diseases and research. Ethnicity prediction from RNA-seq effectively distinguished closely related populations (e.g., British vs. Utah residents of Northern European descent), surpassing SNP-based approaches by capturing rapid, environment-driven transcriptional adaptations in immune-related genes (IL2RA, FOXO4). Integrating RNA-seq with genomic data further enhanced prediction accuracy, revealing nuanced population-specific transcriptomic signatures shaped by genetic ancestry and environmental factors. Our findings underscore RNA-seq's significant potential for precision medicine, highlighting critical biomarkers and pathways that may guide personalized healthcare, anti-aging strategies, disease risk assessment, and targeted therapeutic interventions.\u003c/p\u003e","manuscriptTitle":"Inferring Sex, Ethnicity, and Age from RNA-seq Data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-05 11:48:24","doi":"10.21203/rs.3.rs-7256409/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b4db86be-ad7d-43e1-9a7f-e7cdc4d06565","owner":[],"postedDate":"August 5th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":52440163,"name":"Biological sciences/Computational biology and bioinformatics/Data mining"},{"id":52440164,"name":"Biological sciences/Computational biology and bioinformatics/Computational models"}],"tags":[],"updatedAt":"2025-08-06T17:50:24+00:00","versionOfRecord":[],"versionCreatedAt":"2025-08-05 11:48:24","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7256409","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7256409","identity":"rs-7256409","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-4.0