Machine learning reveals X chromosome transcriptional signatures that predict systemic lupus erythematosus in females | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine learning reveals X chromosome transcriptional signatures that predict systemic lupus erythematosus in females Lachlan G Gray, Aiden Telfser, Hannah Hu, Alisa Kane, Elissa K Deenick, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8991344/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 13 You are reading this latest preprint version Abstract Background Systemic lupus erythematosus (SLE) has a significant female bias; however, it remains unclear whether X-chromosomal dysregulation increases the risk in females. Furthermore, while the role of X-linked genes in SLE pathogenesis is recognised, the analysis of genetic risks in SLE has largely focussed on autosomal genes. Methods Here, we analysed a public single-cell RNA sequencing (scRNA-seq) dataset of peripheral blood mononuclear cells using machine learning to test whether X-linked gene expression predicts female SLE in a cell type-specific manner. Using pseudobulked expression profiles from 1.1 million cells across 228 female donors, we trained an ensemble of classifiers on all X-chromosomal genes or a literature-derived set of SLE-associated genes. Results In CD4⁺ and CD8⁺ T cells, models based on X-chromosomal genes achieved classification performance comparable to SLE gene models and were significantly enriched for XCI escape genes, with IL2RG, CD40LG, and TKTL1 emerging as key predictors. When applied to an independent paediatric–adult SLE cohort, CD8⁺ T-cell models retained robust performance, indicating a stable and reproducible X-linked signature in this subset. Flow cytometry revealed a modest increase in IL2RG protein expression in SLE T cells, consistent with partial escape from XCI in these cells. Conclusions These findings demonstrate that X-linked transcriptional programs, enriched for known XCI escape genes and concentrated in activated T cells, encode a reproducible transcriptional signal of female SLE and provide a framework for dissecting the contributions of sex chromosomes to autoimmunity. Machine learning Systemic lupus erythematosus X chromosome inactivation Single-cell transcriptomics Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Background Systemic lupus erythematosus (SLE) is a prototypical autoimmune disease characterised by antinuclear antibodies (ANAs), increased type 1 interferon signalling, and disrupted clearance of apoptotic debris[ 1 ]. Early diagnosis is critical yet remains challenging because of the heterogeneous clinical presentation and symptoms that overlap with other inflammatory diseases. Notably, these complexities lead to an average diagnostic delay of 3 years [ 2 ]. SLE, like many autoimmune diseases, shows a marked female predominance (9:1), disproportionately affecting young women, in whom it is the leading cause of mortality [ 3 , 4 ]. With the advent of single-cell RNA sequencing (scRNA-seq), studies of SLE have identified differences in immune cell composition, cell type specific patterns of gene expression, and highlighted novel pathogenic mechanisms, such as interferon-driven activation of B and T cells and the expansion of pathogenic myeloid subsets [ 5 – 8 ]. Despite these advances, few scRNA-seq studies have examined how sex differences, particularly the role of the X chromosome, shape immune dysregulation in patients with SLE. In SLE, sex differences are evident in incidence and severity, with males displaying a lower incidence but more severe lupus nephritis [ 9 ]. Although the mechanisms underlying the female bias in SLE are incompletely understood, sex hormones and sex chromosomes are likely to play critical roles. For instance, sex hormones, including fluctuating oestrogen and progesterone levels, have been linked to increases in disease activity, known as flares [ 10 ]. Sex chromosomes have also been linked to increased risk, with X chromosome aneuploidies such as Klinefelter’s syndrome (XXY) and Trisomy X (XXX) conferring a 14-fold and 2.5-3.5-fold increased risk, respectively [ 11 – 13 ]. These studies suggest that X-linked dosage, particularly genes escaping X chromosome inactivation (XCI), contribute to SLE risk. One prominent example is toll-like receptor 7 ( TLR7 ). TLR7 is biallelically expressed in naïve B cells, monocytes and plasmacytoid dendritic cells (pDC) from females and XXY males [ 14 ]. However, TLR7 likely represents one of several X-linked immune genes whose dosage and regulation may contribute to sex-biased immune phenotypes in patients with SLE. A possible mechanism underlying XCI escape is disruption of the epigenetic maintenance of the inactive X chromosome, a process initiated and coordinated by the long non-coding RNA XIST [ 15 ]. In SLE, epigenetic maintenance of the inactive X chromosome appears dysregulated in lymphocytes, as B cells and T cells from patients exhibit incomplete XCI during differentiation from naïve to effector states[ 16 – 19 ]. Consistent with this lymphocyte specificity, computational analyses of single-cell RNA-seq and chromatin accessibility datasets identify stronger patterns of XCI escape in lymphocytes compared with myeloid cells [ 20 ]. Furthermore, recent work demonstrates that XIST functions both as a female-specific autoantigen and as a ligand for TLR7, thereby triggering interferon production in pDCs [ 21 , 22 ]. While machine learning (ML) approaches have been used to predict SLE and identify key transcriptomic features, most have not considered sex-based differences, as male and female data are often combined [ 23 – 25 ]. Additionally, allele-specific expression, XCI, and X-linked escape were not considered because of technical difficulties in these analyses at single-cell resolution. In this study, we combined scRNA-seq with ML to test whether X-linked gene expression alone is sufficient to discriminate female patients with SLE in a cell type–specific manner. Using pseudobulked expression profiles, our models accurately classified SLE patients within CD4⁺ and CD8⁺ T cells and generalised to an independent paediatric–adult cohort. Notably, models trained on X chromosome genes performed comparably to literature-supported SLE gene sets and were significantly enriched for known XCI escape genes. Together, these results indicate that cell type–specific dysregulation of X-linked transcription encodes a stable discriminative signal for female SLE. Methods Single cell RNA sequencing data collection The scRNA-seq dataset generated by Perez et al. [ 8 ] was downloaded from the CellxGene database [ 26 ] as an ‘.h5ad’ file. The downloaded dataset included cells that passed quality control filters. Importantly, technical variation arising from batch effects were removed using COMBAT. Finally, cell types were annotated using canonical marker gene expression and confirmed by antibody staining. Data processing The dataset was read into Python using the Scanpy package [ 27 ]. To focus this analysis on XCI in females, male samples were removed. To reduce flare-associated transcriptional noise, only samples indicated as controls or managed were included. Finally, only individuals with European or Asian ancestry were selected, as African American, Hispanic, and Latin American comprised too few individuals. The data was then read into R using the Seurat package for dimensionality reduction, clustering, and projection with Uniform Manifold Approximation and Projection (UMAP)[ 28 , 29 ]. Pseudobulking by cell type and gene set Across cell types we aggregated batch corrected counts into a pseudobulked gene expression matrix. Pseudobulking accounts for the shared genetics and environment of individual cells, as these violate assumptions of independence. We evaluate our ML models on two gene sets: all X chromosome genes (1,889 genes) and a positive control set of experimentally validated SLE genes (1,883 genes) obtained from the DisGeNet database [ 30 ]. Training and testing data partitions We constructed training and testing indexes using the Scikit-learn ‘train_test_split’ function with a training size of 80% and a test size of 20% [ 31 ]. The proportions of each class (control and disease) and ancestry (European and Asian) were maintained by incorporating each into the stratify parameter. This process was repeated ten times to account for the heterogeneity among patients with SLE. To prevent data leakage, the training and testing sets were z-score normalised with the Scikit-learn ‘StandardScaler’ function separately. Boruta and ElasticNet feature selection To reduce dimensionality and improve model performance, we applied the feature selection methods Boruta and ElasticNet[ 32 , 33 ]. Boruta uses a random forest wrapper to retain features with consistently high importance, while ElasticNet applies regularised regression to remove redundant features. Feature selection was performed independently within each training partition. Features selected in the training set were then used to filter the corresponding testing set. This process was repeated for each of the ten data partitions to ensure that no information from the test set influenced feature selection. We compared features selected by each method individually, as well as their intersection and combination, and found that using both together improved model performance across cell types. Training machine learning classifiers To evaluate the performance of different machine learning classifiers, we tested five algorithms: logistic regression (logit), random forest (RF), support vector machine (SVM), gradient boosting machine (GBM), and multilayer perceptron (MLP). Each model was implemented using the scikit-learn Python library and optimised using the ‘GridSearchCV’ function with 10-fold cross-validation repeated three times [ 31 ]. When supported by the algorithm, models were trained with class_weight="balanced" to account for class imbalance. For logit, models were trained using ‘LogisticRegression’ with the 'saga' solver and 'elasticnet' penalty. The regularisation strength (C) and L1 ratio (l1_ratio) were tuned across logarithmic scales (C = [0.01, 0.1, 1, 10], l1_ratio = [0, 0.5, 1]). RF classifiers were trained using RandomForestClassifier() and tuning n_estimators = [100, 200, 300], max_depth = [ 5 , 10 , 15 ], min_samples_split = [ 2 , 5 , 8 ], and max_features = ['sqrt', 0.3]. SVM was implemented with ‘SVC’, and the kernel choice and regularisation strength were tuned (kernel = [‘linear’, ‘rbf’, ‘poly’, ‘sigmoid’], C = [0.1, 1, 10, 100, 1000]). GBM used ‘GradientBoostingClassifier’ with tuning of learning_rate = [0.1, 0.05, 0.01], n_estimators = [100, 200, 300], max_features = ['sqrt', 0.3], and max_depth = [ 3 , 5 , 7 ]. MLP were implemented with ‘MLPClassifier’ and tuned for hidden_layer_sizes = [(50,), (100,), (50, 50), (100, 100)], solver = ['sgd', 'adam'], alpha = [0.001, 0.01, 0.1], and learning_rate = ['constant', 'invscaling', 'adaptive']. Soft voting ensemble classifier To take advantage of the diverse strengths of each ML algorithm, the models were combined into a soft voting classifier with the scikit-learn VotingClassifier() function with voting set to ‘soft’. In soft voting, a final prediction is made by summing predicted probabilities from each classifier for each class then selecting the class with the highest summed probability. Evaluating model performance Classes were predicted in the testing data and prediction probabilities calculated to evaluate classification performance. Imbalanced datasets were accounted for by modifying prediction probability thresholds from the typical 0.5 with Youden’s J statistic, a metric that calculates thresholds minimising differences between the true positive rate (TPR) and false positive rate (FPR) from a receiver operating curve (ROC) [ 34 ]. Performance metrics are then calculated from the updated prediction probabilities. Mathew’s Correlation Coefficient (MCC) was selected as the evaluation metric because it demonstrates high performance in unbalanced datasets [ 35 ]. MCC ranges from − 1 to + 1 and is a reliable statistical rate because it considers true positives (TP), false negatives (FN), true negatives (TN), and false positives (FP) proportional to the size of positive and negative elements [ 36 ]. The non-parametric Kruskal-Wallis test was employed to compare model metrics across groups of interest [ 37 ]. When testing across cell types, P values were adjusted to false discovery rates (FDR) to account for multiple testing and reduce the FPR [ 38 ]. Shapley additive explanations (SHAP) The relative contribution of each feature to ensemble model predictions was evaluated using Shapley additive explanations (SHAP) [ 39 ]. Within SHAP, an explainer object is created from KernelExplainer() by providing a function that calculates prediction probabilities for a given model and training dataset. The explainer() function then calculates shapley values from the testing dataset. Additional independent SLE dataset To evaluate model generalisability, we first downloaded the GSE135779 [ 7 ] scRNA-seq dataset from the National Center for Biotechnology Information (NCBI) Gene Expression Omnibus (GEO)[ 40 ]. The data was read into R with Seurat and low-quality cells removed using the R package ddqcR [ 41 ]. Multiple cells encapsulated in the same droplet, known as doublets, were identified and removed using scDblFinder [ 42 ]. The data then underwent dimension reduction, clustering, and UMAP using Seurat. Cell type annotation was performed using CellTypist, an automated model trained on scRNA-seq data from immune cells identified in 20 tissues across 19 studies. CellTypist was run with the ‘Immune_All_Low’ model and ‘majority voting’, where predictions were refined by choosing the dominant cell type in each cluster [ 43 ]. The annotated cell type names were modified to match those in the Perez et al., 2022 dataset [ 8 ]. As sex was missing from the metadata, sex was predicted using a hierarchical clustering approach. First, normalised counts for female specific XIST and male specific RPS4Y1 were pseudobulked with the Seurat AggregateExpression() function to sum expression counts across individuals. Counts were then scaled, and Euclidean distances between genes calculated with the R dist() function and hierarchical clustering with the centroid method with hclust(). Clusters were obtained by cutting the dendrogram into k = 2 groups. Predicted females were then evaluated by confirming the higher expression of XIST . Males were then removed from the analyses. Within adult and paediatric samples, the data was pseudobulked as above, and SLE was predicted by the ensemble models for matching cell types. Flow cytometry validation of IL2RG expression To validate IL2RG protein expression, peripheral blood mononuclear cells (PBMCs) were analysed by flow cytometry from an independent cohort of 10 female patients with SLE and 10 healthy female controls. This experiment was designed to validate the direction and cell-type specificity rather than to provide precise effect size estimates. Venous blood samples were collected at St Vincent’s and Liverpool Hospitals (Sydney, Australia), and PBMCs were isolated by Ficoll density-gradient centrifugation and cryopreserved until use. Thawed PBMCs were stained with a panel of fluorochrome-conjugated antibodies including Brillant Violet 421 CD4 (clone OKT4), Brillant Violet 605 CD8 (clone SK1), and Phycoerythrin IL2RG/CD13 (clone TUGh4), together with Ghost Red 710 viability dye. Cells were acquired on a BD LSR Fortessa and analysed using FlowJo (v10.10). T cells were identified by standard forward and side scatter gating for lymphocytes, followed by exclusion of dead cells, and then gating on CD4⁺ and CD8⁺ populations. IL2RG expression was quantified as geometric mean fluorescence intensity (MFI) within each subset. Statistical analyses were performed in R. Differences in IL2RG expression between SLE and control cohorts were assessed using the Wilcoxon rank-sum test. Results Data curation In this study, we assessed the predictive power of X chromosome genes using an ensemble ML framework (Fig. 1 a). We began with a publicly available scRNA-seq dataset of PBMCs from 261 donors (162 SLE patients and 99 healthy controls) comprising 1,263,676 cells [ 8 ]. To focus our analysis on XCI, all males were excluded from the dataset. Additionally, flare-associated transcriptional noise was avoided by excluding patients experiencing disease flares. To reduce ancestry-related confounders, individuals with Hispanic, Latin American, or African American ancestry were removed as these ancestries comprised too few individuals. The remaining 1,116,410 cells from 96 controls (24 Asian and 72 European) and 132 SLE patients (70 Asian and 62 European) were included in downstream analysis. We used the 24 immune cell annotations from the original work, as these were determined by marker gene expression and confirmed by antibody labelling of cell surface proteins [ 8 ]. To account for within-individual dependence and reduce data sparsity, we pseudobulked batch-corrected gene expression by aggregating counts for each cell type per individual. Following pseudobulking, we trained five ML classifiers using cell type-specific gene expression. Feature selection was performed prior to training, and the classification models were evaluated on held-out test sets. Predictions from each classifier were combined into an ensemble classifier, a strategy that improved robustness over individual models. Model generalisability was then evaluated by predicting an independent scRNA-seq dataset from Nehar-Belaid et al. [ 7 ]. This dataset includes 56 donors: 40 SLE patients (33 paediatric and 7 adult) and 16 healthy controls (11 paediatric and 5 adult) comprising 331,402 PBMCs. Sex was inferred from the expression of XIST (female) and RPS4Y1 (male), and male samples were excluded. After quality control filtering, the validation set contained 52 female donors: 15 controls (10 paediatric and 5 adult) and 37 SLE cases (30 paediatric and 3 adult) comprising 310,310 cells (Fig. 1 b). This framework allowed us to test whether X-linked features provide predicting power beyond established immune signatures. X chromosome models predict SLE in CD4 + and CD8 + T cells Traditional studies perform differential expression analysis on gene expression data by comparing cases to controls. In this study we applied a ML framework to explore potentially more complex and non-linear interactions between genes that differentiate between SLE cases and controls. First, to benchmark the machine learning approach against standard methods, we performed differential expression analysis on pseudobulked data using the edgeR package [ 44 ]. After removing lowly expressed genes, we fitted a generalized negative binomial model with batch and age as covariates. After filtering for significant differentially expressed genes (DEGs) (FDR 0.5), we observe that most DEGs were upregulated in SLE compared to controls (Fig. 2 a). We then applied our ML framework to each cell-type. Pseudobulked expression profiles were filtered for X chromosome genes (N = 1,889) or a set of literature supported SLE-associated genes from the DisGeNET [ 30 ] database (N = 1,883) (Additional file 1: Table S3). We reasoned that models trained on X chromosome genes (chrX) that displayed comparable performance to models trained on SLE-associated genes (SLE) would reveal whether X-linked expression is sufficient to predict SLE. To reduce the feature space and improve model interpretation, we performed feature selection with Elastic Net and Boruta and took the overlap of features selected by both methods [ 32 , 33 ]. Across ten data partitions, the average number of features selected in chrX and SLE models was 4.43 ± 4.14 and 17.15 ± 18.25, respectively, reflecting substantial inter-partition heterogeneity consistent with the known molecular heterogeneity of SLE. After feature selection, we ran classifiers with increasing complexity: logistic regression (logit), random forest (RF), support vector machines (SVM), gradient boosting (GBM), and multilayer perceptron (MLP). Predictions from each individual model were then combined into an ensemble classifier with soft voting to improve robustness. We trained each ML classifier using an 80:20 train-test split, stratified by disease status and ancestry, and repeated over ten data partitions. Performance was evaluated using Matthew’s Correlation Coefficient (MCC), a metric that ranges from − 1 to + 1 and is robust in imbalanced datasets [ 35 , 36 ]. Comprehensive performance metrics across all data splits, models, gene sets, and cell types including MCC, area under the curve (AUC), area under the precision recall curve (AUPRC), and weighted F1-score are reported in ( Additional file 1: Table S1 ). We then compared ensemble models trained on chrX and SLE genes to identify models with similar predictive performance. In three cell types, chrX models achieved comparable classification to models trained on SLE-associated genes (Wilcoxon rank sum test with FDR correction, p > 0.05). These included CD4 + T cells, CD8 + T cells, and CD34 + Progenitor cells (Fig. 2 b). However, Progenitor cells represented a very small fraction of the dataset, and their performance was unstable, with the 95% confidence interval for MCC ranging from 0.05 to 0.44, overlapping chance-level performance. Given the low abundance and lack of robust predictive power, progenitor cells were excluded from further analysis. In contrast, CD4 + and CD8 + T cells showed consistently strong performance with 95% CI MCC values ranging from 0.55–0.73 and 0.49–0.70, respectively. We then asked whether X chromosome features selected across the ten partitions were enriched for known escape genes ( Additional file 1: Table S2 ). Using all X-linked genes as the background, features selected in CD4⁺ T cells (N = 6) and CD8⁺ T cells (N = 9) were found to be significantly enriched for XCI escape genes (χ² test, p < 0.01). Overlap with differential expression results showed that interleukin 2 receptor gamma ( IL2RG ) was significantly upregulated in CD4 + and CD8 + T cells, whereas transketolase like 1 ( TKTL1 ) was significantly downregulated in CD8 + T cells (Fig. 2 C). The absence of statistically significant changes for the remaining features highlights the utility of our ML approach, a framework that leverages subtle, multigene expression patterns to construct predictive decision boundaries. The strong performance of CD4 + T cells and CD8 + T cells is consistent with evidence implicating these cell types in SLE pathogenesis. CD4 + T cells contribute to the pathogenesis of SLE by providing help to autoreactive B cells and through the secretion of inflammatory cytokines. Furthermore, abnormalities in T cell receptor (TCR) signalling pathways lower the activation threshold in SLE to promote autoreactivity [ 45 ]. CD8⁺ T cells, in contrast, often display reduced cytotoxic capacity in SLE, a defect linked to increased susceptibility to infection, one of the leading causes of mortality in SLE [ 46 ]. Notably, memory and effector CD4⁺ and CD8⁺ T cells, but not naïve subsets, showed robust predictive performance, consistent with reports that XCI escape signatures become more detectable following antigen-driven T-cell differentiation (Fig. 2 b)[ 17 , 19 , 47 ]. This suggests that T-cell activation states may modulate X-linked transcription in ways that are informative for distinguishing SLE from controls. Shapley additive explanations reveal key features for model predictions To interpret how individual genes influenced model predictions, we applied Shapley additive explanations (SHAP), a game theory-based approach that estimates the contribution of each feature to the classification output [ 39 ]. SHAP values were calculated using a representative ensemble model retrained on the first data split, with hyperparameters averaged across partitions. In CD4 + T cells, IL2RG showed the strongest and most consistent positive contribution to SLE classification, with higher expression associated with higher predicted probability of SLE. The remaining selected features showed the same direction of effect but with smaller magnitudes, consistent with the differential expression results (Fig. 3 a). In CD8⁺ T cells, IL2RG again exhibited a strong positive influence, but the overall range of SHAP values was larger than in CD4⁺ T cells, indicating greater heterogeneity in how gene expression contributes to classification. TKTL1 showed the opposite pattern, with lower expression driving model predictions toward SLE, in line with its observed downregulation. Additional features, such as brain expressed X-linked 5 ( BEX5 ) and integral membrane protein 2A ( ITM2A ), also displayed clear patterns with low BEX5 expression and high ITM2A expression positively influencing SLE classification (Fig. 3 b). CD8 + T cell model predicts SLE in an independent dataset A significant challenge in ML is overfitting to the training data, resulting in models that fail to generalise to unseen data. To assess generalisability, we applied our ensemble models to an independent scRNA-seq dataset of PBMCs from adult and paediatric SLE patients and healthy donors. Cell types were annotated using CellTypist [ 43 ] and harmonised with the labels used in the training dataset. Pseudobulked expression profiles were generated directly from raw counts without batch correction. As expected, both chrX and SLE models showed reduced performance when transferred to this independent cohort (Fig. 4 ). However, CD8⁺ T cells stood out in adults where chrX and SLE models performed similarly, with a 95% CI for MCC of 0.52–0.72 (Wilcoxon rank-sum test with FDR correction, p > 0.05). Notably, CD8⁺ T cells were the only population to retain a robust and gene-set–independent predictive signal across both datasets. This consistency is unlikely to arise from technical artefact alone, as the two datasets differ in donor age, sequencing chemistry, and batch correction. Instead, it suggests that CD8⁺ T cells carry a stable, reproducible X-linked transcriptional associated with SLE status and preserved across adult cohorts. This finding points to the X chromosome as a source of generalizable, age-related, and cell-type–specific information in female SLE, particularly within CD8⁺ T cells, where the signal appears strongest and most reproducible. IL2RG expression is increased in T cells from patients with SLE To validate our computational findings at the protein level, we quantified IL2RG protein expression in peripheral CD4 + and CD8 + T cells from an independent cohort of 10 female SLE patients, nine of whom were receiving immunosuppressive therapy at the time of sampling, and 10 healthy controls, using flow cytometry. IL2RG was selected for downstream validation based on its central biological role in cytokine signalling, its consistent importance across both CD4⁺ and CD8⁺ T cell models, and the feasibility of orthogonal validation due to its cell surface expression. IL2RG protein expression was significantly elevated in CD4 + T cells from SLE patients (p = 0.043, Wilcoxon rank-sum test) and showed a trend toward elevation in CD8 + T cells (p = 0.052) (Fig. 5 a, b). However, given the low baseline expression of IL2RG and the predominance of immunosuppressed patients in the cohort, the observed differences likely reflect subtle shifts in gene dosage rather than large changes in protein abundance. Together, this finding supports the presence of subtle but detectable protein-level dosage effects consistent with an overall increase in expression, that could be linked to partial escape from XCI. Discussion Using an ensemble of ML classifiers trained on X chromosome genes, we have identified features that accurately classified female SLE from controls, with XCI escape genes significantly overrepresented among predictive features. These results are consistent with existing literature and reinforce the contribution of escape from XCI to the pathogenesis of SLE. We assessed the predictive value of X chromosome models by comparing their performance to models trained on a set of literature-supported SLE genes. Within CD4 + and CD8 + T cells the comparable performance between X chromosome and SLE models suggests that the models learned an X chromosome transcriptomic signal for SLE classification. We then evaluated how overfit the models were to the training data by predicting an independent scRNA-seq dataset of paediatric and adult SLE. Here, CD8 + T cells indicate generalisability in the unseen adult cohort as the X chromosome model had a robust performance comparable to the SLE model. The reduced performance of the CD8 + T cell model in the paediatric cohort suggests that this X-linked signature may be an age-dependent phenomenon, rather than a congenital feature of SLE. Together with recent evidence that aging promotes reactivation of the inactive X chromosome, these data support a model whereby progressive erosion of XCI in lymphocytes contributes to female-biased autoimmune risk over time [ 48 ]. Recent work has shown extensive dysregulation of X-linked transcriptomes in SLE, including differential expression of XIST and genes of the XIST -interactome across adaptive and innate immune cells [ 49 ]. Whereas that study offered a comprehensive descriptive atlas of X-linked expression changes, our approach identifies a predictive subset of X-linked features that generalise across independent cohorts and highlight specific cell states, particularly within CD8⁺ T cells Inspection of the features selected to train the CD4 + and CD8 + T cell models revealed X chromosome genes with known sex differences in their expression. For example, IL2RG , which emerged as a consistently informative feature in CD4 + and CD8 + T cells, encodes a cell surface receptor for the common gamma chain cytokines IL-2, IL-4, IL-7, IL-9, IL-15, and IL-21 [ 50 ]. Moreover, IL2RG expression was previously shown to be influenced by age and sex [ 51 – 53 ]. Recently, Il2rg was found to escape XCI in unstimulated and stimulated T cells from mice [ 19 ]. However, it is currently unclear whether IL2RG exhibits biallelic expression in human CD4 + and CD8 + T cells. In line with its selection in the classification models, IL2RG protein expression showed a subtle but significant increase in CD4⁺ T cells and a similar trend in CD8⁺ T cells from SLE patients. While limited by the detection sensitivity of flow cytometry for low-abundance surface proteins, this finding is consistent with the increased expression of IL2RG observed in the scRNA-seq data. Importantly, the inclusion of predominantly immunosuppressed SLE patients likely contributes to the modest effect sizes observed here. For example, a recent study reported that IL2RG (CD132) expression is significantly upregulated in lymphocytes from untreated SLE patients and positively correlates with disease activity, whereas this relationship is attenuated in treated cohorts [ 54 ]. Together, these results support the presence of subtle protein-level dosage effects consistent with increased gene expression. However, without further allelic expression information, we could not directly link it to partial escape of IL2RG from XCI, but it is a potential mechanism that needs to be further investigated. Another gene CD40 ligand ( CD40LG ), selected in CD4 + T cells, is known to be hypomethylated on the inactive X chromosome in T cells from SLE patients [ 55 ]. Importantly, upregulated CD40LG in CD4 + T cells was demonstrated to stimulate the production of IgG autoantibodies by autologous B cells [ 56 ]. Transketolase like 1 ( TKTL1 ) was the second highest ranked feature in CD8 + T cells and displayed a significant downregulation in SLE. TKTL1 encodes an enzyme involved in glucose metabolism, with roles described in cancer biology. Although its functional role in SLE is unclear, its consistent selection by the models suggests it may contribute to the discriminative signal through subtle shifts in transcriptional programs associated with T cell activation states [ 57 ]. While TKTL1 has not previously been associated with SLE, it was characterised as a variable escape gene [ 58 ]. The low SHAP scores for the remaining features suggest that predictive performance may rely less on strong single-gene effects and more on the subtle, cumulative contribution of multiple features. While these shifts involve known XCI escape genes, they may reflect distinct transcriptional programs associated with T-cell activation states rather than biallelic expression. However, given evidence of epigenetic instability in lymphocytes, reactivation remains a plausible driver. Future allele-specific expression analyses are required to definitively distinguish between dosage compensation failure and activation-induced upregulation, which cannot be resolved from the present data. Our approach can be applied to any sufficiently large scRNA-seq dataset to identify biologically motivated gene sets that discriminate between conditions. Unlike traditional differential expression analysis, our approach makes no assumptions about data distribution, can capture combined effects of multiple genes, and retains lowly expressed genes that might otherwise be filtered out. However, interpretability remains a major challenge. While SHAP analysis highlighted the influence of individual features, the precise decision-making processes of ensemble ML models remain inaccessible. Several limitations should be acknowledged. First, XCI escape status could not be confirmed in the SLE scRNA-seq data. Obtaining allele specific expression information, which would allow us to measure biallelic expression of genes indicating escape, is difficult from single-cell data. A recent re-analysis of this dataset concluded that the power was insufficient to detect escape at single-cell resolution, although several candidates were suggested [ 20 ]. Second, the consequences of the identified expression differences for additional markers, including CD40LG and TKTL1, remain to be validated experimentally. Additional Flow cytometry of extracellular and intracellular proteins could assess changes in expression, while RNA-FISH of nascent transcripts could determine whether these X-linked genes exhibit biallelic expression [ 59 ]. Third, differences between the training and validation cohorts, including sequencing chemistry, ancestry composition, and batch correction, likely contributed to reduced generalisation. Nevertheless, the reproducible performance of CD8⁺ T cells strengthen the confidence that the selected features capture a disease-relevant signal. Additionally, the absence of a well-known SLE X-linked gene, TLR7 , in our results, we believe is due to the filtering criteria we used in our analysis on gene expression levels. This is also consistent with the original Perez et al. work. Finally, given the disproportionately high disease burden in females with African American ancestry their exclusion due to low sample size limits the generalizability of these findings [ 60 ]. This limitation highlights the urgent need for more diverse genomic cohorts to ensure that machine learning-derived diagnostics are equitable across the populations most in need. Despite these limitations, we still observe strong signals of X-linked genes that could be used as proxies for escape or X dysregulation, specific to females where future work will be able to expand on these. Conclusions In summary, this study shows that an ML-based approach identified X chromosome genes, many of which escape XCI, that predict female SLE with an accuracy comparable to established SLE-associated genes. These findings advance our understanding of the X chromosome’s contribution to the female bias in SLE pathogenesis and provide a framework for applying similar ML-based approaches to other autoimmune diseases. Future studies integrating allele-specific expression and protein-level validation will be essential for translating these predictive signatures into mechanistic and therapeutic insights. Abbreviations ANA antinuclear antibodies AUC area under the curve AUPRC area under the precision recall curve BEX5 brain expressed X-linked 5 CD40LG CD40 ligand DEGs differentially expressed genes FDR false discovery rates FN false negatives FP false positives FPR false positive rate GBM gradient boosting machine GEO Gene Expression Omnibus IL2RG interleukin 2 receptor gamma IQR interquartile range ITM2A integral membrane protein 2A logit logistic regression MCC Matthews Correlation Coefficient MFI mean fluorescence intensity ML machine learning MLP multilayer perceptron NCBI National Center for Biotechnology Information PBMCs peripheral blood mononuclear cells pDCs plasmacytoid dendritic cells RF random forest ROC receiver operating curve scRNA-seq single-cell RNA sequencing SHAP Shapley additive explanations SLE systemic lupus erythematosus SVM support vector machine TKTL1 transketolase like 1 TLR7 toll-like receptor 7 TN true negatives TP true positives TPR true positive rate UMAP Uniform Manifold Approximation and Projection XCI X chromosome inactivation chrX X chromosome Declarations Funding L.G.G. was supported by an Australian Government Research Training Program (RTP) Scholarship. H.H. was supported by a University of New South Wales University Postgraduate Award (UPA). E.K.D. was supported by an NHMRC Investigator Grant (APPID 2026131). The SLE cohort study was supported by a SPHERE Triple I Seed Grant awarded to A.K. and E.K.D. T.G.P. was supported by an NHMRC Investigator Grant (APPID 2026122). S.B. was supported by a fellowship from the Magid’s. Data availability The datasets generated and/or analysed during the current study are available in the repository ( 10.5281/zenodo.18853801 ). All scRNA-seq data used in this study are available online from public databases. The raw and batch corrected counts from Perez et al. (2022) are available in the CellxGene database https://cellxgene.cziscience.com/collections/436154da-bcf1-4130-9c8b-120ff9a888f2 and GEO database under accession code GSE137029 ( https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE137029 ). The raw counts from Nehar-Belaid et al. (2020) are available in the GEO database under accession code GSE135779 ( https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135779 ). Author Contribution L.G.G. designed the study, performed all analyses, and drafted the manuscript. A.T. conducted flow cytometry experiments and interpretation. H.H. provided and collected biological samples as part of the SLE cohort study supervised and conceived by A. K. and E.K.D. E.K.D. contributed to flow cytometry design, analysis and interpretation. T.G.P. provided supervision. S.B. conceived the study, supervised all analysis, and provided feedback. All authors read and approved the final manuscript. Acknowledgement L.G.G would like to thank Professor Joseph Powell for his supervision of this project and UNSW Stats Central for feedback on the machine learning framework. Data Availability The datasets generated and/or analysed during the current study are available in the repository (10.5281/zenodo.18853801).All scRNA-seq data used in this study are available online from public databases. The raw and batch corrected counts from Perez et al. (2022) are available in the CellxGene database [https://cellxgene.cziscience.com/collections/436154da-bcf1-4130-9c8b-120ff9a888f2](https:/cellxgene.cziscience.com/collections/436154da-bcf1-4130-9c8b-120ff9a888f2) and GEO database under accession code GSE137029 ( [https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE137029](https:/www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE137029) ). The raw counts from Nehar-Belaid et al. (2020) are available in the GEO database under accession code GSE135779 ( [https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135779](https:/www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135779) ). References Tsokos GC. The immunology of systemic lupus erythematosus. Nat Immunol. 2024;25:1332–43. Kernder A, et al. Delayed diagnosis adversely affects outcome in systemic lupus erythematosus: Cross sectional analysis of the LuLa cohort. Lupus. 2021;30:431–8. Hayter SM, Cook MC. Updated assessment of the prevalence, spectrum and case definition of autoimmune disease. Autoimmun rev. 2012;11:754–65. Yen EY, Singh RR. Lupus—An Unrecognized Leading Cause of Death in Young Females. Arthritis Rheumatol. 2018;70:1251–5. Arazi A, et al. The immune cell landscape in kidneys of patients with lupus nephritis. Nat Immunol. 2019;20:902–14. Mistry P et al. Transcriptomic, epigenetic, and functional analyses implicate neutrophil diversity in the pathogenesis of systemic lupus erythematosus. Proceedings of the National Academy of Sciences, 2019. 116: pp. 25222–25228. Nehar-Belaid D, et al. Mapping systemic lupus erythematosus heterogeneity at the single-cell level. Nat Immunol. 2020;21:1094–106. Perez RK et al. Single-cell RNA-seq reveals cell type-specific molecular and genetic associations to lupus. Science, 2022. 376. Mahmood SB, et al. Evaluating Sex Differences in the Characteristics and Outcomes of Lupus Nephritis: A Systematic Review and Meta-Analysis. Glomerular Dis. 2024;4:19–32. Kim JW et al. Sex hormones affect the pathogenesis and clinical characteristics of systemic lupus erythematosus. Front Med, 2022. 9. Scofield RH, et al. Klinefelter’s syndrome (47,XXY) in male systemic lupus erythematosus patients. Arthr Rhuem. 2008;58:2511–7. Liu K, et al. X Chromosome Dose and Sex Bias in Autoimmune Diseases. Arthritis Rheumatol. 2016;68:1290–300. Palmer A, Tan IJ. Synergistic Effects of Extra X Chromosome on Development of Systemic Lupus Erythematosus and Sjögren Disease. ACR Open Rheumatology; 2025. p. 7. Souyris M et al. TLR7 escapes X chromosome inactivation in immune cells. Sci Immunol, 2018. 3. Forsyth KS, et al. The conneXion between sex and immune responses. Nature Reviews Immunology; 2024. Wang J et al. Unusual maintenance of X chromosome inactivation predisposes female lymphocytes for increased expression from the inactive X. PNAS, 2016. 113: pp. E2029-E2038. Syrett CM et al. Altered X-chromosome inactivation in T cells may promote sex-biased autoimmune diseases. JCI Insight, 2019. 4. Pyfrom S et al. The dynamic epigenetic regulation of the inactive X chromosome in healthy human B cells is dysregulated in lupus patients. PNAS, 2021. 118. Forsyth KS, et al. Maintenance of X chromosome inactivation after T cell activation requires NF-κB signaling. Science Immunology; 2024. Tomofuji Y et al. Quantification of escape from X chromosome inactivation with single-cell omics data. Cell Genomics, 2024. 4(8). Crawford JD et al. The XIST lncRNA is a sex-specific reservoir of TLR7 ligands in SLE. JCI Insight, 2023. 8. Dou DR, et al. Xist ribonucleoproteins promote female sex-biased autoimmunity. Cell. 2024;187:733–49. Figgett WA, et al. Machine learning applied to whole-blood RNA-sequencing data uncovers distinct subsets of patients with systemic lupus erythematosus. Clinical & Translational Immunology; 2019. p. 8. Kegerreis B et al. Machine learning approaches to predict lupus disease activity from gene expression data. Sci Rep, 2019. 9. Ma Y, et al. Accurate Machine Learning Model to Diagnose Chronic Autoimmune Diseases Utilizing Information From B Cells and Monocytes. Frontiers in Immunology; 2022. p. 13. Chan Zuckerberg I. CELLxGENE. 2023. Wolf FA, Angerer P, Theis FJ. SCANPY: large-scale single-cell gene expression data analysis. Genome Biol. 2018;19:15. Hao Y, et al. Integrated analysis of multimodal single-cell data. Cell. 2021;184:3573–87. Team RC. R: A Language and Environment for Statistical Computing. 2023. Piñero J, et al. The DisGeNET knowledge platform for disease genomics: 2019 update. Nucleic Acids Research; 2019. Pedregosa F, et al. Scikit-learn: Machine Learning in Python. J Mach Learn Res. 2011;12:2825–30. Zou H, Hastie T. Regularization and Variable Selection Via the Elastic Net. J Royal Stat Soc Ser B. 2005;67:301–20. Kursa MB, Rudnicki WR. Feature Selection with the Boruta Package. J Stat Softw, 2010. 36. Youden WJ. Index for rating diagnostic tests. Cancer. 1950;3:32–5. Chicco D, Jurman G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics. 2020;21:6. Matthews BW. Comparison of the predicted and observed secondary structure of T4 phage lysozyme. Biochim Biophys Acta. 1975;405:442–51. Kruskal WH, Wallis WA. Use of Ranks in One-Criterion Variance Analysis. J Am Stat Assoc. 1952;47:583. Benjamini Y, Hochberg Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. J Royal Stat Soc Ser B. 1995;57:289–300. Lundberg SM, et al. Explainable machine-learning predictions for the prevention of hypoxaemia during surgery. Nat Biomedical Eng. 2018;2:749–60. Barrett T, et al. NCBI GEO: archive for functional genomics data sets—update. Nucleic Acids Res. 2012;41:D991–5. Subramanian A et al. Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics. Genome Biol, 2022. 23. Germain PL et al. Doublet identification in single-cell sequencing data using scDblFinder. F1000Research, 2021. 10: p. 979. Domínguez Conde C et al. Cross-tissue immune cell analysis reveals tissue-specific features in humans. Science, 2022. 376. Chen Y et al. edgeR v4: powerful differential analysis of sequencing data with expanded functionality and improved support for small counts and larger datasets. Nucleic Acids Res, 2025. 53. Moulton VR, Tsokos GC. T cell signaling abnormalities contribute to aberrant immune cell function and autoimmunity. J Clin Invest. 2015;125:2220–7. Katsuyama E, et al. The CD38/NAD/SIRTUIN1/EZH2 Axis Mitigates Cytotoxic CD8 T Cell Function and Identifies Patients with SLE Prone to Infections. Cell Rep. 2020;30:112–23. Jiwrajka N et al. Impaired dynamic X-chromosome inactivation maintenance in T cells is a feature of spontaneous murine SLE that is exacerbated in female-biased models. J Autoimmun, 2023. 139. Hoelzl S et al. Aging promotes reactivation of the Barr body at distal chromosome regions. Nat Aging, 2025. Soares M, et al. X-linked transcriptome dysregulation across immune cells in systemic lupus erythematosus. Biology Sex Differences. 2025;16:69. Lin JX, Leonard WJ. The Common Cytokine Receptor γ Chain Family of Cytokines. Volume 10. Cold Spring Harbor Perspectives in Biology; 2018. p. a028449. Jansen R, et al. Sex differences in the human peripheral blood transcriptome. BMC Genomics. 2014;15:33. Oliva M et al. The impact of sex on gene expression across human tissues. Science, 2020. 369. Melé M, et al. The human transcriptome across tissues and individuals. Science. 2015;348:660–5. Yin H, et al. 2D4, a humanized monoclonal antibody targeting CD132, is a promising treatment for systemic lupus erythematosus. Signal Transduct Target Therapy. 2024;9:323. Lu Q, et al. Demethylation of CD40LG on the Inactive X in T Cells from Women with Lupus. J Immunol. 2007;179:6352–8. Zhou Y, et al. T cell CD40LG gene expression and the production of IgG by autologous B cells in systemic lupus erythematosus. Clin Immunol. 2009;132:362–70. Patiño-Martinez E, Kaplan MJ. Immunometabolism in systemic lupus erythematosus. Nat Rev Rheumatol. 2025;16:1–19. Cotton AM, et al. Analysis of expressed SNPs identifies variable extents of expression from the human inactive X chromosome. Genome Biol. 2013;14:R122. Chaumeil J et al. Combined Immunofluorescence, RNA Fluorescent In Situ Hybridization, and DNA Fluorescent In Situ Hybridization to Study Chromatin Changes, Transcriptional Activity, Nuclear Organization, and X-Chromosome Inactivation , in X-Chromosome Inactivation . 2008. pp. 297–308. Barber MRW, et al. Global epidemiology of systemic lupus erythematosus. Nat Rev Rheumatol. 2021;17:515–32. Additional Declarations No competing interests reported. Supplementary Files AdditionalFile1.xlsx Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 20 Apr, 2026 Reviews received at journal 09 Apr, 2026 Reviews received at journal 04 Apr, 2026 Reviews received at journal 31 Mar, 2026 Reviewers agreed at journal 30 Mar, 2026 Reviewers agreed at journal 26 Mar, 2026 Reviewers agreed at journal 26 Mar, 2026 Reviewers agreed at journal 19 Mar, 2026 Reviewers invited by journal 19 Mar, 2026 Editor assigned by journal 19 Mar, 2026 Editor invited by journal 04 Mar, 2026 Submission checks completed at journal 03 Mar, 2026 First submitted to journal 03 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8991344","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":609425533,"identity":"fac76216-4268-4f14-92a3-46cd8631b2c9","order_by":0,"name":"Lachlan G Gray","email":"","orcid":"","institution":"UNSW Sydney","correspondingAuthor":false,"prefix":"","firstName":"Lachlan","middleName":"G","lastName":"Gray","suffix":""},{"id":609425535,"identity":"110f5ab9-7a89-403c-b225-93f06983afc6","order_by":1,"name":"Aiden Telfser","email":"","orcid":"","institution":"Garvan Institute of Medical Research","correspondingAuthor":false,"prefix":"","firstName":"Aiden","middleName":"","lastName":"Telfser","suffix":""},{"id":609425537,"identity":"a2108860-5a68-4a60-a23a-23eb7ac20728","order_by":2,"name":"Hannah Hu","email":"","orcid":"","institution":"Garvan Institute of Medical Research","correspondingAuthor":false,"prefix":"","firstName":"Hannah","middleName":"","lastName":"Hu","suffix":""},{"id":609425539,"identity":"c286ef0f-7c6a-425d-890c-58b2da7d1b7d","order_by":3,"name":"Alisa Kane","email":"","orcid":"","institution":"St Vincent's Hospital Sydney","correspondingAuthor":false,"prefix":"","firstName":"Alisa","middleName":"","lastName":"Kane","suffix":""},{"id":609425540,"identity":"8cc70645-e1a0-46dd-8d00-7f912e23057e","order_by":4,"name":"Elissa K Deenick","email":"","orcid":"","institution":"Garvan Institute of Medical Research","correspondingAuthor":false,"prefix":"","firstName":"Elissa","middleName":"K","lastName":"Deenick","suffix":""},{"id":609425544,"identity":"02c49533-202d-49e6-a520-da7a91cefee3","order_by":5,"name":"Tri Giang Phan","email":"","orcid":"","institution":"Garvan Institute of Medical Research","correspondingAuthor":false,"prefix":"","firstName":"Tri","middleName":"Giang","lastName":"Phan","suffix":""},{"id":609425546,"identity":"8412acd2-63b8-4d60-a427-ced26dbe1b6a","order_by":6,"name":"Sara Ballouz","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA6UlEQVRIiWNgGAWjYFACHgaGBwVAmr0BKnCAGC0JBiAappR4LRIJRGqRb+A9+CHBwCZx/sw3Zg9/tjHI8d1IYPzMg0eLwQG+ZIkEg7TExtk55sa8bQzGkjcSmKXxamHgMQBqOZzYLJ1jJs24jSFxw40EBrxa5Bt4jH+AtLRJnjGT/LmNoR6ohfk3Pi0MB3jMwLb0SAAZvNuAQXEjgQ2/ww7zpVkA/WI8gyetTJr3n4ThzDMP2yzn4HNYe+/hGx8qbGTntx/eJvnjjI083/Hkwzfe4HMYMypXAogZG/BpGAWjYBSMglFABAAAzY9HBlgC2wgAAAAASUVORK5CYII=","orcid":"","institution":"UNSW Sydney","correspondingAuthor":true,"prefix":"","firstName":"Sara","middleName":"","lastName":"Ballouz","suffix":""}],"badges":[],"createdAt":"2026-02-27 20:08:17","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8991344/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8991344/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105564579,"identity":"747d265d-fe43-4125-b742-3e3275d53147","added_by":"auto","created_at":"2026-03-27 12:50:06","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":101319,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of the study. (a) Schematic representation of the ML workflow used to predict SLE status from gene expression. Gene expression profiles were generated from pseudobulked scRNA-seq data, with 105 SLE cases and 79 controls included for model training (80%) and 27 cases with 20 controls reserved for model testing (20%) [8]. Feature selection was performed using Boruta and Elastic Net to identify the most informative genes distinguishing cases from controls. Features selected by both methods were included in subsequent analysis. Multiple classifiers, including logistic regression (logit), random forest (RF), support vector machine (SVM), gradient boosting machine (GBM), and multilayer perceptron (MLP), were trained individually and their predictions combined into an ensemble classifier with soft voting. (b) UMAP projection of 1,116,410 PBMCs from 228 donors, clustered and annotated into 24 immune cell types, and the validation dataset of 52 cases and controls split into children and adult cohorts.\u003c/p\u003e","description":"","filename":"Picture1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8991344/v1/142686f50e4f3fe98d207339.jpg"},{"id":105326815,"identity":"25c9ac12-e7a6-4cd5-937d-e7cea7f37c48","added_by":"auto","created_at":"2026-03-24 19:09:54","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":76804,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance of machine learning models. (a) Differential expression results across all cell-types ordered by number of genes. (b) Classification performance for SLE and X chromosome gene sets. Each boxplot shows Matthews Correlation Coefficient (MCC) values across cross-validation splits (Wilcoxon rank sum test with FDR correction, FDR \u0026gt; 0.05). Comparable classification performance to those trained on SLE-associated genes across ten immune cell types are italicised.\u0026nbsp; (c) Overlap of features with DEGs from CD8\u003csup\u003e+ \u003c/sup\u003eT cells and CD4\u003csup\u003e+\u003c/sup\u003e T cells.\u003c/p\u003e","description":"","filename":"Picture2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8991344/v1/9c86e5de0d25d8dc6fa7a60a.jpg"},{"id":105326819,"identity":"1ed7ce0a-0a96-4a90-b231-eb740de82e49","added_by":"auto","created_at":"2026-03-24 19:09:55","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":36348,"visible":true,"origin":"","legend":"\u003cp\u003eSHAP analysis of predictive features in (a) CD4\u003csup\u003e+\u003c/sup\u003e and (b) CD8\u003csup\u003e+\u003c/sup\u003e T cells. \u003cem\u003eIL2RG\u003c/em\u003e was the top predictors in CD4\u003csup\u003e+\u003c/sup\u003e T cells, while \u003cem\u003eIL2RG\u003c/em\u003e, \u003cem\u003eTKTL1\u003c/em\u003e, \u003cem\u003eBEX5, \u003c/em\u003eand \u003cem\u003eITM2A\u003c/em\u003e contributed most in CD8\u003csup\u003e+\u003c/sup\u003e T cells. Each point represents a sample; colour indicates feature value (blue = low, red = high), and position reflects SHAP value (feature impact on classification).\u003c/p\u003e","description":"","filename":"Picture3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8991344/v1/b1ad597db4855367f3922692.jpg"},{"id":105564578,"identity":"a17ad239-7472-4e60-85b7-e1323cbd6c73","added_by":"auto","created_at":"2026-03-27 12:50:06","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":58248,"visible":true,"origin":"","legend":"\u003cp\u003eGeneralisation of X chromosome and SLE gene models to an independent SLE dataset. Boxplots show the distribution of MCC scores for each immune cell type when ensemble models trained on X chromosome genes or SLE-associated genes predicted an external scRNA-seq dataset of PBMCs from adult and paediatric SLE patients and healthy donors. MCC values were calculated separately for adult (top panel) and paediatric (bottom panel) cohorts. Boxplot centre lines indicate medians; box limits indicate interquartile ranges (IQR); whiskers extend to 1.5 x IQR.\u003c/p\u003e","description":"","filename":"Picture4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8991344/v1/da13c59200d81c38cfee9f23.jpg"},{"id":105565972,"identity":"50d03ea7-9803-4400-9337-b95edfb49886","added_by":"auto","created_at":"2026-03-27 12:54:54","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":23810,"visible":true,"origin":"","legend":"\u003cp\u003eIL2RG expression is increased in T cells from patients with SLE. (a) IL2RG protein expression, measured as mean fluorescence intensity (MFI), in peripheral CD4⁺ and (b) CD8⁺ T cells from healthy controls (HC; n = 10) and patients with SLE (n = 10). IL2RG expression was significantly increased in CD4⁺ T cells from SLE patients (p = 0.043, Wilcoxon rank-sum test) and showed a trend toward elevation in CD8⁺ T cells (p = 0.052).\u003c/p\u003e","description":"","filename":"Picture5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8991344/v1/d33addbe91917165a3a74f89.jpg"},{"id":105730383,"identity":"f294f157-1f0c-4641-a360-1fcd31595d71","added_by":"auto","created_at":"2026-03-30 11:24:19","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":861819,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8991344/v1/95646873-625f-4c68-80ce-0e4103dd4d48.pdf"},{"id":105727890,"identity":"99410f4e-535e-4769-8f8a-d4eb98b07ee1","added_by":"auto","created_at":"2026-03-30 11:05:01","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":2124180,"visible":true,"origin":"","legend":"","description":"","filename":"AdditionalFile1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-8991344/v1/b14e13b7a031defe4a7e3114.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine learning reveals X chromosome transcriptional signatures that predict systemic lupus erythematosus in females","fulltext":[{"header":"Background","content":"\u003cp\u003eSystemic lupus erythematosus (SLE) is a prototypical autoimmune disease characterised by antinuclear antibodies (ANAs), increased type 1 interferon signalling, and disrupted clearance of apoptotic debris[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Early diagnosis is critical yet remains challenging because of the heterogeneous clinical presentation and symptoms that overlap with other inflammatory diseases. Notably, these complexities lead to an average diagnostic delay of 3 years [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSLE, like many autoimmune diseases, shows a marked female predominance (9:1), disproportionately affecting young women, in whom it is the leading cause of mortality [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. With the advent of single-cell RNA sequencing (scRNA-seq), studies of SLE have identified differences in immune cell composition, cell type specific patterns of gene expression, and highlighted novel pathogenic mechanisms, such as interferon-driven activation of B and T cells and the expansion of pathogenic myeloid subsets [\u003cspan additionalcitationids=\"CR6 CR7\" citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Despite these advances, few scRNA-seq studies have examined how sex differences, particularly the role of the X chromosome, shape immune dysregulation in patients with SLE.\u003c/p\u003e \u003cp\u003eIn SLE, sex differences are evident in incidence and severity, with males displaying a lower incidence but more severe lupus nephritis [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Although the mechanisms underlying the female bias in SLE are incompletely understood, sex hormones and sex chromosomes are likely to play critical roles. For instance, sex hormones, including fluctuating oestrogen and progesterone levels, have been linked to increases in disease activity, known as flares [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Sex chromosomes have also been linked to increased risk, with X chromosome aneuploidies such as Klinefelter\u0026rsquo;s syndrome (XXY) and Trisomy X (XXX) conferring a 14-fold and 2.5-3.5-fold increased risk, respectively [\u003cspan additionalcitationids=\"CR12\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. These studies suggest that X-linked dosage, particularly genes escaping X chromosome inactivation (XCI), contribute to SLE risk. One prominent example is toll-like receptor 7 (\u003cem\u003eTLR7\u003c/em\u003e). TLR7 is biallelically expressed in na\u0026iuml;ve B cells, monocytes and plasmacytoid dendritic cells (pDC) from females and XXY males [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. However, \u003cem\u003eTLR7\u003c/em\u003e likely represents one of several X-linked immune genes whose dosage and regulation may contribute to sex-biased immune phenotypes in patients with SLE. A possible mechanism underlying XCI escape is disruption of the epigenetic maintenance of the inactive X chromosome, a process initiated and coordinated by the long non-coding RNA \u003cem\u003eXIST\u003c/em\u003e [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. In SLE, epigenetic maintenance of the inactive X chromosome appears dysregulated in lymphocytes, as B cells and T cells from patients exhibit incomplete XCI during differentiation from na\u0026iuml;ve to effector states[\u003cspan additionalcitationids=\"CR17 CR18\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Consistent with this lymphocyte specificity, computational analyses of single-cell RNA-seq and chromatin accessibility datasets identify stronger patterns of XCI escape in lymphocytes compared with myeloid cells [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Furthermore, recent work demonstrates that \u003cem\u003eXIST\u003c/em\u003e functions both as a female-specific autoantigen and as a ligand for TLR7, thereby triggering interferon production in pDCs [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWhile machine learning (ML) approaches have been used to predict SLE and identify key transcriptomic features, most have not considered sex-based differences, as male and female data are often combined [\u003cspan additionalcitationids=\"CR24\" citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Additionally, allele-specific expression, XCI, and X-linked escape were not considered because of technical difficulties in these analyses at single-cell resolution. In this study, we combined scRNA-seq with ML to test whether X-linked gene expression alone is sufficient to discriminate female patients with SLE in a cell type\u0026ndash;specific manner. Using pseudobulked expression profiles, our models accurately classified SLE patients within CD4⁺ and CD8⁺ T cells and generalised to an independent paediatric\u0026ndash;adult cohort. Notably, models trained on X chromosome genes performed comparably to literature-supported SLE gene sets and were significantly enriched for known XCI escape genes. Together, these results indicate that cell type\u0026ndash;specific dysregulation of X-linked transcription encodes a stable discriminative signal for female SLE.\u003c/p\u003e "},{"header":"Methods","content":" \u003cp\u003eSingle cell RNA sequencing data collection\u003c/p\u003e \u003cp\u003eThe scRNA-seq dataset generated by Perez et al. [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] was downloaded from the CellxGene database [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] as an \u0026lsquo;.h5ad\u0026rsquo; file. The downloaded dataset included cells that passed quality control filters. Importantly, technical variation arising from batch effects were removed using COMBAT. Finally, cell types were annotated using canonical marker gene expression and confirmed by antibody staining.\u003c/p\u003e \u003cp\u003eData processing\u003c/p\u003e \u003cp\u003eThe dataset was read into Python using the Scanpy package [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. To focus this analysis on XCI in females, male samples were removed. To reduce flare-associated transcriptional noise, only samples indicated as controls or managed were included. Finally, only individuals with European or Asian ancestry were selected, as African American, Hispanic, and Latin American comprised too few individuals. The data was then read into R using the Seurat package for dimensionality reduction, clustering, and projection with Uniform Manifold Approximation and Projection (UMAP)[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e].\u003c/p\u003e \u003cp\u003ePseudobulking by cell type and gene set\u003c/p\u003e \u003cp\u003eAcross cell types we aggregated batch corrected counts into a pseudobulked gene expression matrix. Pseudobulking accounts for the shared genetics and environment of individual cells, as these violate assumptions of independence. We evaluate our ML models on two gene sets: all X chromosome genes (1,889 genes) and a positive control set of experimentally validated SLE genes (1,883 genes) obtained from the DisGeNet database [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTraining and testing data partitions\u003c/p\u003e \u003cp\u003eWe constructed training and testing indexes using the Scikit-learn \u0026lsquo;train_test_split\u0026rsquo; function with a training size of 80% and a test size of 20% [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. The proportions of each class (control and disease) and ancestry (European and Asian) were maintained by incorporating each into the stratify parameter. This process was repeated ten times to account for the heterogeneity among patients with SLE. To prevent data leakage, the training and testing sets were z-score normalised with the Scikit-learn \u0026lsquo;StandardScaler\u0026rsquo; function separately.\u003c/p\u003e \u003cp\u003eBoruta and ElasticNet feature selection\u003c/p\u003e \u003cp\u003eTo reduce dimensionality and improve model performance, we applied the feature selection methods Boruta and ElasticNet[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Boruta uses a random forest wrapper to retain features with consistently high importance, while ElasticNet applies regularised regression to remove redundant features. Feature selection was performed independently within each training partition. Features selected in the training set were then used to filter the corresponding testing set. This process was repeated for each of the ten data partitions to ensure that no information from the test set influenced feature selection. We compared features selected by each method individually, as well as their intersection and combination, and found that using both together improved model performance across cell types.\u003c/p\u003e \u003cp\u003eTraining machine learning classifiers\u003c/p\u003e \u003cp\u003eTo evaluate the performance of different machine learning classifiers, we tested five algorithms: logistic regression (logit), random forest (RF), support vector machine (SVM), gradient boosting machine (GBM), and multilayer perceptron (MLP). Each model was implemented using the scikit-learn Python library and optimised using the \u0026lsquo;GridSearchCV\u0026rsquo; function with 10-fold cross-validation repeated three times [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. When supported by the algorithm, models were trained with class_weight=\"balanced\" to account for class imbalance. For logit, models were trained using \u0026lsquo;LogisticRegression\u0026rsquo; with the 'saga' solver and 'elasticnet' penalty. The regularisation strength (C) and L1 ratio (l1_ratio) were tuned across logarithmic scales (C = [0.01, 0.1, 1, 10], l1_ratio = [0, 0.5, 1]). RF classifiers were trained using RandomForestClassifier() and tuning n_estimators = [100, 200, 300], max_depth = [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], min_samples_split = [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], and max_features = ['sqrt', 0.3]. SVM was implemented with \u0026lsquo;SVC\u0026rsquo;, and the kernel choice and regularisation strength were tuned (kernel = [\u0026lsquo;linear\u0026rsquo;, \u0026lsquo;rbf\u0026rsquo;, \u0026lsquo;poly\u0026rsquo;, \u0026lsquo;sigmoid\u0026rsquo;], C = [0.1, 1, 10, 100, 1000]). GBM used \u0026lsquo;GradientBoostingClassifier\u0026rsquo; with tuning of learning_rate = [0.1, 0.05, 0.01], n_estimators = [100, 200, 300], max_features = ['sqrt', 0.3], and max_depth = [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. MLP were implemented with \u0026lsquo;MLPClassifier\u0026rsquo; and tuned for hidden_layer_sizes = [(50,), (100,), (50, 50), (100, 100)], solver = ['sgd', 'adam'], alpha = [0.001, 0.01, 0.1], and learning_rate = ['constant', 'invscaling', 'adaptive'].\u003c/p\u003e \u003cp\u003eSoft voting ensemble classifier\u003c/p\u003e \u003cp\u003eTo take advantage of the diverse strengths of each ML algorithm, the models were combined into a soft voting classifier with the scikit-learn VotingClassifier() function with voting set to \u0026lsquo;soft\u0026rsquo;. In soft voting, a final prediction is made by summing predicted probabilities from each classifier for each class then selecting the class with the highest summed probability.\u003c/p\u003e \u003cp\u003eEvaluating model performance\u003c/p\u003e \u003cp\u003eClasses were predicted in the testing data and prediction probabilities calculated to evaluate classification performance. Imbalanced datasets were accounted for by modifying prediction probability thresholds from the typical 0.5 with Youden\u0026rsquo;s J statistic, a metric that calculates thresholds minimising differences between the true positive rate (TPR) and false positive rate (FPR) from a receiver operating curve (ROC) [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Performance metrics are then calculated from the updated prediction probabilities. Mathew\u0026rsquo;s Correlation Coefficient (MCC) was selected as the evaluation metric because it demonstrates high performance in unbalanced datasets [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. MCC ranges from \u0026minus;\u0026thinsp;1 to +\u0026thinsp;1 and is a reliable statistical rate because it considers true positives (TP), false negatives (FN), true negatives (TN), and false positives (FP) proportional to the size of positive and negative elements [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe non-parametric Kruskal-Wallis test was employed to compare model metrics across groups of interest [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. When testing across cell types, P values were adjusted to false discovery rates (FDR) to account for multiple testing and reduce the FPR [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eShapley additive explanations (SHAP)\u003c/p\u003e \u003cp\u003eThe relative contribution of each feature to ensemble model predictions was evaluated using Shapley additive explanations (SHAP) [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. Within SHAP, an explainer object is created from KernelExplainer() by providing a function that calculates prediction probabilities for a given model and training dataset. The explainer() function then calculates shapley values from the testing dataset.\u003c/p\u003e \u003cp\u003eAdditional independent SLE dataset\u003c/p\u003e \u003cp\u003eTo evaluate model generalisability, we first downloaded the GSE135779 [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] scRNA-seq dataset from the National Center for Biotechnology Information (NCBI) Gene Expression Omnibus (GEO)[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. The data was read into R with Seurat and low-quality cells removed using the R package ddqcR [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Multiple cells encapsulated in the same droplet, known as doublets, were identified and removed using scDblFinder [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. The data then underwent dimension reduction, clustering, and UMAP using Seurat. Cell type annotation was performed using CellTypist, an automated model trained on scRNA-seq data from immune cells identified in 20 tissues across 19 studies. CellTypist was run with the \u0026lsquo;Immune_All_Low\u0026rsquo; model and \u0026lsquo;majority voting\u0026rsquo;, where predictions were refined by choosing the dominant cell type in each cluster [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. The annotated cell type names were modified to match those in the Perez et al., 2022 dataset [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. As sex was missing from the metadata, sex was predicted using a hierarchical clustering approach. First, normalised counts for female specific \u003cem\u003eXIST\u003c/em\u003e and male specific \u003cem\u003eRPS4Y1\u003c/em\u003e were pseudobulked with the Seurat AggregateExpression() function to sum expression counts across individuals. Counts were then scaled, and Euclidean distances between genes calculated with the R dist() function and hierarchical clustering with the centroid method with hclust(). Clusters were obtained by cutting the dendrogram into k\u0026thinsp;=\u0026thinsp;2 groups. Predicted females were then evaluated by confirming the higher expression of \u003cem\u003eXIST\u003c/em\u003e. Males were then removed from the analyses. Within adult and paediatric samples, the data was pseudobulked as above, and SLE was predicted by the ensemble models for matching cell types.\u003c/p\u003e \u003cp\u003eFlow cytometry validation of IL2RG expression\u003c/p\u003e \u003cp\u003eTo validate IL2RG protein expression, peripheral blood mononuclear cells (PBMCs) were analysed by flow cytometry from an independent cohort of 10 female patients with SLE and 10 healthy female controls. This experiment was designed to validate the direction and cell-type specificity rather than to provide precise effect size estimates. Venous blood samples were collected at St Vincent\u0026rsquo;s and Liverpool Hospitals (Sydney, Australia), and PBMCs were isolated by Ficoll density-gradient centrifugation and cryopreserved until use.\u003c/p\u003e \u003cp\u003eThawed PBMCs were stained with a panel of fluorochrome-conjugated antibodies including Brillant Violet 421 CD4 (clone OKT4), Brillant Violet 605 CD8 (clone SK1), and Phycoerythrin IL2RG/CD13 (clone TUGh4), together with Ghost Red 710 viability dye. Cells were acquired on a BD LSR Fortessa and analysed using FlowJo (v10.10). T cells were identified by standard forward and side scatter gating for lymphocytes, followed by exclusion of dead cells, and then gating on CD4⁺ and CD8⁺ populations. IL2RG expression was quantified as geometric mean fluorescence intensity (MFI) within each subset. Statistical analyses were performed in R. Differences in IL2RG expression between SLE and control cohorts were assessed using the Wilcoxon rank-sum test.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eData curation\u003c/p\u003e \u003cp\u003eIn this study, we assessed the predictive power of X chromosome genes using an ensemble ML framework (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea). We began with a publicly available scRNA-seq dataset of PBMCs from 261 donors (162 SLE patients and 99 healthy controls) comprising 1,263,676 cells [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. To focus our analysis on XCI, all males were excluded from the dataset. Additionally, flare-associated transcriptional noise was avoided by excluding patients experiencing disease flares. To reduce ancestry-related confounders, individuals with Hispanic, Latin American, or African American ancestry were removed as these ancestries comprised too few individuals. The remaining 1,116,410 cells from 96 controls (24 Asian and 72 European) and 132 SLE patients (70 Asian and 62 European) were included in downstream analysis. We used the 24 immune cell annotations from the original work, as these were determined by marker gene expression and confirmed by antibody labelling of cell surface proteins [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. To account for within-individual dependence and reduce data sparsity, we pseudobulked batch-corrected gene expression by aggregating counts for each cell type per individual. Following pseudobulking, we trained five ML classifiers using cell type-specific gene expression. Feature selection was performed prior to training, and the classification models were evaluated on held-out test sets. Predictions from each classifier were combined into an ensemble classifier, a strategy that improved robustness over individual models. Model generalisability was then evaluated by predicting an independent scRNA-seq dataset from Nehar-Belaid et al. [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. This dataset includes 56 donors: 40 SLE patients (33 paediatric and 7 adult) and 16 healthy controls (11 paediatric and 5 adult) comprising 331,402 PBMCs. Sex was inferred from the expression of \u003cem\u003eXIST\u003c/em\u003e (female) and \u003cem\u003eRPS4Y1\u003c/em\u003e (male), and male samples were excluded. After quality control filtering, the validation set contained 52 female donors: 15 controls (10 paediatric and 5 adult) and 37 SLE cases (30 paediatric and 3 adult) comprising 310,310 cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). This framework allowed us to test whether X-linked features provide predicting power beyond established immune signatures.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eX chromosome models predict SLE in CD4\u003csup\u003e+\u003c/sup\u003e and CD8\u003csup\u003e+\u003c/sup\u003e T cells\u003c/p\u003e \u003cp\u003eTraditional studies perform differential expression analysis on gene expression data by comparing cases to controls. In this study we applied a ML framework to explore potentially more complex and non-linear interactions between genes that differentiate between SLE cases and controls. First, to benchmark the machine learning approach against standard methods, we performed differential expression analysis on pseudobulked data using the edgeR package [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. After removing lowly expressed genes, we fitted a generalized negative binomial model with batch and age as covariates. After filtering for significant differentially expressed genes (DEGs) (FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05; |logFC| \u0026gt; 0.5), we observe that most DEGs were upregulated in SLE compared to controls (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea).\u003c/p\u003e \u003cp\u003eWe then applied our ML framework to each cell-type. Pseudobulked expression profiles were filtered for X chromosome genes (N\u0026thinsp;=\u0026thinsp;1,889) or a set of literature supported SLE-associated genes from the DisGeNET [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] database (N\u0026thinsp;=\u0026thinsp;1,883) (Additional file 1: Table S3). We reasoned that models trained on X chromosome genes (chrX) that displayed comparable performance to models trained on SLE-associated genes (SLE) would reveal whether X-linked expression is sufficient to predict SLE. To reduce the feature space and improve model interpretation, we performed feature selection with Elastic Net and Boruta and took the overlap of features selected by both methods [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Across ten data partitions, the average number of features selected in chrX and SLE models was 4.43\u0026thinsp;\u0026plusmn;\u0026thinsp;4.14 and 17.15\u0026thinsp;\u0026plusmn;\u0026thinsp;18.25, respectively, reflecting substantial inter-partition heterogeneity consistent with the known molecular heterogeneity of SLE. After feature selection, we ran classifiers with increasing complexity: logistic regression (logit), random forest (RF), support vector machines (SVM), gradient boosting (GBM), and multilayer perceptron (MLP). Predictions from each individual model were then combined into an ensemble classifier with soft voting to improve robustness. We trained each ML classifier using an 80:20 train-test split, stratified by disease status and ancestry, and repeated over ten data partitions. Performance was evaluated using Matthew\u0026rsquo;s Correlation Coefficient (MCC), a metric that ranges from \u0026minus;\u0026thinsp;1 to +\u0026thinsp;1 and is robust in imbalanced datasets [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. Comprehensive performance metrics across all data splits, models, gene sets, and cell types including MCC, area under the curve (AUC), area under the precision recall curve (AUPRC), and weighted F1-score are reported in (\u003cb\u003eAdditional file 1: Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eWe then compared ensemble models trained on chrX and SLE genes to identify models with similar predictive performance. In three cell types, chrX models achieved comparable classification to models trained on SLE-associated genes (Wilcoxon rank sum test with FDR correction, p\u0026thinsp;\u0026gt;\u0026thinsp;0.05). These included CD4\u003csup\u003e+\u003c/sup\u003e T cells, CD8\u003csup\u003e+\u003c/sup\u003e T cells, and CD34\u003csup\u003e+\u003c/sup\u003e Progenitor cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb). However, Progenitor cells represented a very small fraction of the dataset, and their performance was unstable, with the 95% confidence interval for MCC ranging from 0.05 to 0.44, overlapping chance-level performance. Given the low abundance and lack of robust predictive power, progenitor cells were excluded from further analysis. In contrast, CD4\u003csup\u003e+\u003c/sup\u003e and CD8\u003csup\u003e+\u003c/sup\u003e T cells showed consistently strong performance with 95% CI MCC values ranging from 0.55\u0026ndash;0.73 and 0.49\u0026ndash;0.70, respectively. We then asked whether X chromosome features selected across the ten partitions were enriched for known escape genes (\u003cb\u003eAdditional file 1: Table S2\u003c/b\u003e). Using all X-linked genes as the background, features selected in CD4⁺ T cells (N\u0026thinsp;=\u0026thinsp;6) and CD8⁺ T cells (N\u0026thinsp;=\u0026thinsp;9) were found to be significantly enriched for XCI escape genes (χ\u0026sup2; test, p\u0026thinsp;\u0026lt;\u0026thinsp;0.01). Overlap with differential expression results showed that interleukin 2 receptor gamma (\u003cem\u003eIL2RG\u003c/em\u003e) was significantly upregulated in CD4\u003csup\u003e+\u003c/sup\u003e and CD8\u003csup\u003e+\u003c/sup\u003e T cells, whereas transketolase like 1 (\u003cem\u003eTKTL1\u003c/em\u003e) was significantly downregulated in CD8\u003csup\u003e+\u003c/sup\u003e T cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC). The absence of statistically significant changes for the remaining features highlights the utility of our ML approach, a framework that leverages subtle, multigene expression patterns to construct predictive decision boundaries.\u003c/p\u003e \u003cp\u003eThe strong performance of CD4\u003csup\u003e+\u003c/sup\u003e T cells and CD8\u003csup\u003e+\u003c/sup\u003e T cells is consistent with evidence implicating these cell types in SLE pathogenesis. CD4\u003csup\u003e+\u003c/sup\u003e T cells contribute to the pathogenesis of SLE by providing help to autoreactive B cells and through the secretion of inflammatory cytokines. Furthermore, abnormalities in T cell receptor (TCR) signalling pathways lower the activation threshold in SLE to promote autoreactivity [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. CD8⁺ T cells, in contrast, often display reduced cytotoxic capacity in SLE, a defect linked to increased susceptibility to infection, one of the leading causes of mortality in SLE [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]. Notably, memory and effector CD4⁺ and CD8⁺ T cells, but not na\u0026iuml;ve subsets, showed robust predictive performance, consistent with reports that XCI escape signatures become more detectable following antigen-driven T-cell differentiation (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb)[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]. This suggests that T-cell activation states may modulate X-linked transcription in ways that are informative for distinguishing SLE from controls.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eShapley additive explanations reveal key features for model predictions\u003c/p\u003e \u003cp\u003eTo interpret how individual genes influenced model predictions, we applied Shapley additive explanations (SHAP), a game theory-based approach that estimates the contribution of each feature to the classification output [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. SHAP values were calculated using a representative ensemble model retrained on the first data split, with hyperparameters averaged across partitions. In CD4\u003csup\u003e+\u003c/sup\u003e T cells, \u003cem\u003eIL2RG\u003c/em\u003e showed the strongest and most consistent positive contribution to SLE classification, with higher expression associated with higher predicted probability of SLE. The remaining selected features showed the same direction of effect but with smaller magnitudes, consistent with the differential expression results (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea). In CD8⁺ T cells, \u003cem\u003eIL2RG\u003c/em\u003e again exhibited a strong positive influence, but the overall range of SHAP values was larger than in CD4⁺ T cells, indicating greater heterogeneity in how gene expression contributes to classification. \u003cem\u003eTKTL1\u003c/em\u003e showed the opposite pattern, with lower expression driving model predictions toward SLE, in line with its observed downregulation. Additional features, such as brain expressed X-linked 5 (\u003cem\u003eBEX5\u003c/em\u003e) and integral membrane protein 2A (\u003cem\u003eITM2A\u003c/em\u003e), also displayed clear patterns with low \u003cem\u003eBEX5\u003c/em\u003e expression and high \u003cem\u003eITM2A\u003c/em\u003e expression positively influencing SLE classification (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eCD8\u003csup\u003e+\u003c/sup\u003e T cell model predicts SLE in an independent dataset\u003c/p\u003e \u003cp\u003eA significant challenge in ML is overfitting to the training data, resulting in models that fail to generalise to unseen data. To assess generalisability, we applied our ensemble models to an independent scRNA-seq dataset of PBMCs from adult and paediatric SLE patients and healthy donors. Cell types were annotated using CellTypist [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e] and harmonised with the labels used in the training dataset. Pseudobulked expression profiles were generated directly from raw counts without batch correction. As expected, both chrX and SLE models showed reduced performance when transferred to this independent cohort (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). However, CD8⁺ T cells stood out in adults where chrX and SLE models performed similarly, with a 95% CI for MCC of 0.52\u0026ndash;0.72 (Wilcoxon rank-sum test with FDR correction, p\u0026thinsp;\u0026gt;\u0026thinsp;0.05). Notably, CD8⁺ T cells were the only population to retain a robust and gene-set\u0026ndash;independent predictive signal across both datasets. This consistency is unlikely to arise from technical artefact alone, as the two datasets differ in donor age, sequencing chemistry, and batch correction. Instead, it suggests that CD8⁺ T cells carry a stable, reproducible X-linked transcriptional associated with SLE status and preserved across adult cohorts. This finding points to the X chromosome as a source of generalizable, age-related, and cell-type\u0026ndash;specific information in female SLE, particularly within CD8⁺ T cells, where the signal appears strongest and most reproducible.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIL2RG expression is increased in T cells from patients with SLE\u003c/p\u003e \u003cp\u003eTo validate our computational findings at the protein level, we quantified IL2RG protein expression in peripheral CD4\u003csup\u003e+\u003c/sup\u003e and CD8\u003csup\u003e+\u003c/sup\u003e T cells from an independent cohort of 10 female SLE patients, nine of whom were receiving immunosuppressive therapy at the time of sampling, and 10 healthy controls, using flow cytometry. IL2RG was selected for downstream validation based on its central biological role in cytokine signalling, its consistent importance across both CD4⁺ and CD8⁺ T cell models, and the feasibility of orthogonal validation due to its cell surface expression. IL2RG protein expression was significantly elevated in CD4\u003csup\u003e+\u003c/sup\u003e T cells from SLE patients (p\u0026thinsp;=\u0026thinsp;0.043, Wilcoxon rank-sum test) and showed a trend toward elevation in CD8\u003csup\u003e+\u003c/sup\u003e T cells (p\u0026thinsp;=\u0026thinsp;0.052) (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea, b). However, given the low baseline expression of IL2RG and the predominance of immunosuppressed patients in the cohort, the observed differences likely reflect subtle shifts in gene dosage rather than large changes in protein abundance. Together, this finding supports the presence of subtle but detectable protein-level dosage effects consistent with an overall increase in expression, that could be linked to partial escape from XCI.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eUsing an ensemble of ML classifiers trained on X chromosome genes, we have identified features that accurately classified female SLE from controls, with XCI escape genes significantly overrepresented among predictive features. These results are consistent with existing literature and reinforce the contribution of escape from XCI to the pathogenesis of SLE. We assessed the predictive value of X chromosome models by comparing their performance to models trained on a set of literature-supported SLE genes. Within CD4\u003csup\u003e+\u003c/sup\u003e and CD8\u003csup\u003e+\u003c/sup\u003e T cells the comparable performance between X chromosome and SLE models suggests that the models learned an X chromosome transcriptomic signal for SLE classification. We then evaluated how overfit the models were to the training data by predicting an independent scRNA-seq dataset of paediatric and adult SLE. Here, CD8\u003csup\u003e+\u003c/sup\u003e T cells indicate generalisability in the unseen adult cohort as the X chromosome model had a robust performance comparable to the SLE model. The reduced performance of the CD8\u003csup\u003e+\u003c/sup\u003e T cell model in the paediatric cohort suggests that this X-linked signature may be an age-dependent phenomenon, rather than a congenital feature of SLE. Together with recent evidence that aging promotes reactivation of the inactive X chromosome, these data support a model whereby progressive erosion of XCI in lymphocytes contributes to female-biased autoimmune risk over time [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. Recent work has shown extensive dysregulation of X-linked transcriptomes in SLE, including differential expression of \u003cem\u003eXIST\u003c/em\u003e and genes of the \u003cem\u003eXIST\u003c/em\u003e-interactome across adaptive and innate immune cells [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e]. Whereas that study offered a comprehensive descriptive atlas of X-linked expression changes, our approach identifies a predictive subset of X-linked features that generalise across independent cohorts and highlight specific cell states, particularly within CD8⁺ T cells\u003c/p\u003e \u003cp\u003eInspection of the features selected to train the CD4\u003csup\u003e+\u003c/sup\u003e and CD8\u003csup\u003e+\u003c/sup\u003e T cell models revealed X chromosome genes with known sex differences in their expression. For example, \u003cem\u003eIL2RG\u003c/em\u003e, which emerged as a consistently informative feature in CD4\u003csup\u003e+\u003c/sup\u003e and CD8\u003csup\u003e+\u003c/sup\u003e T cells, encodes a cell surface receptor for the common gamma chain cytokines IL-2, IL-4, IL-7, IL-9, IL-15, and IL-21 [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]. Moreover, \u003cem\u003eIL2RG\u003c/em\u003e expression was previously shown to be influenced by age and sex [\u003cspan additionalcitationids=\"CR52\" citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. Recently, \u003cem\u003eIl2rg\u003c/em\u003e was found to escape XCI in unstimulated and stimulated T cells from mice [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. However, it is currently unclear whether \u003cem\u003eIL2RG\u003c/em\u003e exhibits biallelic expression in human CD4\u003csup\u003e+\u003c/sup\u003e and CD8\u003csup\u003e+\u003c/sup\u003e T cells. In line with its selection in the classification models, IL2RG protein expression showed a subtle but significant increase in CD4⁺ T cells and a similar trend in CD8⁺ T cells from SLE patients. While limited by the detection sensitivity of flow cytometry for low-abundance surface proteins, this finding is consistent with the increased expression of \u003cem\u003eIL2RG\u003c/em\u003e observed in the scRNA-seq data. Importantly, the inclusion of predominantly immunosuppressed SLE patients likely contributes to the modest effect sizes observed here. For example, a recent study reported that IL2RG (CD132) expression is significantly upregulated in lymphocytes from untreated SLE patients and positively correlates with disease activity, whereas this relationship is attenuated in treated cohorts [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]. Together, these results support the presence of subtle protein-level dosage effects consistent with increased gene expression. However, without further allelic expression information, we could not directly link it to partial escape of IL2RG from XCI, but it is a potential mechanism that needs to be further investigated.\u003c/p\u003e \u003cp\u003eAnother gene CD40 ligand (\u003cem\u003eCD40LG\u003c/em\u003e), selected in CD4\u003csup\u003e+\u003c/sup\u003e T cells, is known to be hypomethylated on the inactive X chromosome in T cells from SLE patients [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e]. Importantly, upregulated \u003cem\u003eCD40LG\u003c/em\u003e in CD4\u003csup\u003e+\u003c/sup\u003e T cells was demonstrated to stimulate the production of IgG autoantibodies by autologous B cells [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. Transketolase like 1 (\u003cem\u003eTKTL1\u003c/em\u003e) was the second highest ranked feature in CD8\u003csup\u003e+\u003c/sup\u003e T cells and displayed a significant downregulation in SLE. \u003cem\u003eTKTL1\u003c/em\u003e encodes an enzyme involved in glucose metabolism, with roles described in cancer biology. Although its functional role in SLE is unclear, its consistent selection by the models suggests it may contribute to the discriminative signal through subtle shifts in transcriptional programs associated with T cell activation states [\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e]. While \u003cem\u003eTKTL1\u003c/em\u003e has not previously been associated with SLE, it was characterised as a variable escape gene [\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e]. The low SHAP scores for the remaining features suggest that predictive performance may rely less on strong single-gene effects and more on the subtle, cumulative contribution of multiple features. While these shifts involve known XCI escape genes, they may reflect distinct transcriptional programs associated with T-cell activation states rather than biallelic expression. However, given evidence of epigenetic instability in lymphocytes, reactivation remains a plausible driver. Future allele-specific expression analyses are required to definitively distinguish between dosage compensation failure and activation-induced upregulation, which cannot be resolved from the present data.\u003c/p\u003e \u003cp\u003eOur approach can be applied to any sufficiently large scRNA-seq dataset to identify biologically motivated gene sets that discriminate between conditions. Unlike traditional differential expression analysis, our approach makes no assumptions about data distribution, can capture combined effects of multiple genes, and retains lowly expressed genes that might otherwise be filtered out. However, interpretability remains a major challenge. While SHAP analysis highlighted the influence of individual features, the precise decision-making processes of ensemble ML models remain inaccessible.\u003c/p\u003e \u003cp\u003eSeveral limitations should be acknowledged. First, XCI escape status could not be confirmed in the SLE scRNA-seq data. Obtaining allele specific expression information, which would allow us to measure biallelic expression of genes indicating escape, is difficult from single-cell data. A recent re-analysis of this dataset concluded that the power was insufficient to detect escape at single-cell resolution, although several candidates were suggested [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Second, the consequences of the identified expression differences for additional markers, including CD40LG and TKTL1, remain to be validated experimentally. Additional Flow cytometry of extracellular and intracellular proteins could assess changes in expression, while RNA-FISH of nascent transcripts could determine whether these X-linked genes exhibit biallelic expression [\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e]. Third, differences between the training and validation cohorts, including sequencing chemistry, ancestry composition, and batch correction, likely contributed to reduced generalisation. Nevertheless, the reproducible performance of CD8⁺ T cells strengthen the confidence that the selected features capture a disease-relevant signal. Additionally, the absence of a well-known SLE X-linked gene, \u003cem\u003eTLR7\u003c/em\u003e, in our results, we believe is due to the filtering criteria we used in our analysis on gene expression levels. This is also consistent with the original Perez et al. work. Finally, given the disproportionately high disease burden in females with African American ancestry their exclusion due to low sample size limits the generalizability of these findings [\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e]. This limitation highlights the urgent need for more diverse genomic cohorts to ensure that machine learning-derived diagnostics are equitable across the populations most in need. Despite these limitations, we still observe strong signals of X-linked genes that could be used as proxies for escape or X dysregulation, specific to females where future work will be able to expand on these.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eIn summary, this study shows that an ML-based approach identified X chromosome genes, many of which escape XCI, that predict female SLE with an accuracy comparable to established SLE-associated genes. These findings advance our understanding of the X chromosome\u0026rsquo;s contribution to the female bias in SLE pathogenesis and provide a framework for applying similar ML-based approaches to other autoimmune diseases. Future studies integrating allele-specific expression and protein-level validation will be essential for translating these predictive signatures into mechanistic and therapeutic insights.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eANA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eantinuclear antibodies\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAUC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003earea under the curve\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAUPRC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003earea under the precision recall curve\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eBEX5\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ebrain expressed X-linked 5\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCD40LG\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eCD40 ligand\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDEGs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003edifferentially expressed genes\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFDR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003efalse discovery rates\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFN\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003efalse negatives\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003efalse positives\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFPR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003efalse positive rate\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eGBM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003egradient boosting machine\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eGEO\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eGene Expression Omnibus\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eIL2RG\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003einterleukin 2 receptor gamma\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eIQR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003einterquartile range\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eITM2A\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eintegral membrane protein 2A\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003elogit\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003elogistic regression\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMCC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMatthews Correlation Coefficient\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMFI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003emean fluorescence intensity\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eML\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003emachine learning\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMLP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003emultilayer perceptron\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eNCBI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eNational Center for Biotechnology Information\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003ePBMCs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eperipheral blood mononuclear cells\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003epDCs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eplasmacytoid dendritic cells\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRF\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003erandom forest\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eROC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ereceiver operating curve\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003escRNA-seq\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003esingle-cell RNA sequencing\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSHAP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eShapley additive explanations\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSLE\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003esystemic lupus erythematosus\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSVM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003esupport vector machine\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTKTL1\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003etransketolase like 1\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTLR7\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003etoll-like receptor 7\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTN\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003etrue negatives\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003etrue positives\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTPR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003etrue positive rate\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eUMAP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eUniform Manifold Approximation and Projection\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eXCI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eX chromosome inactivation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003echrX\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eX chromosome\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eL.G.G. was supported by an Australian Government Research Training Program (RTP) Scholarship. H.H. was supported by a University of New South Wales University Postgraduate Award (UPA). E.K.D. was supported by an NHMRC Investigator Grant (APPID 2026131). The SLE cohort study was supported by a SPHERE Triple I Seed Grant awarded to A.K. and E.K.D. T.G.P. was supported by an NHMRC Investigator Grant (APPID 2026122). S.B. was supported by a fellowship from the Magid\u0026rsquo;s.\u003c/p\u003e \u003cp\u003eData availability\u003c/p\u003e \u003cp\u003eThe datasets generated and/or analysed during the current study are available in the repository \u003cb\u003e(\u003c/b\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5281/zenodo.18853801\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.18853801\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eAll scRNA-seq data used in this study are available online from public databases. The raw and batch corrected counts from Perez et al. (2022) are available in the CellxGene database \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cellxgene.cziscience.com/collections/436154da-bcf1-4130-9c8b-120ff9a888f2\u003c/span\u003e\u003cspan address=\"https://cellxgene.cziscience.com/collections/436154da-bcf1-4130-9c8b-120ff9a888f2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e and GEO database under accession code GSE137029 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE137029\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE137029\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The raw counts from Nehar-Belaid et al. (2020) are available in the GEO database under accession code GSE135779 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135779\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135779\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eL.G.G. designed the study, performed all analyses, and drafted the manuscript. A.T. conducted flow cytometry experiments and interpretation. H.H. provided and collected biological samples as part of the SLE cohort study supervised and conceived by A. K. and E.K.D. E.K.D. contributed to flow cytometry design, analysis and interpretation. T.G.P. provided supervision. S.B. conceived the study, supervised all analysis, and provided feedback. All authors read and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eL.G.G would like to thank Professor Joseph Powell for his supervision of this project and UNSW Stats Central for feedback on the machine learning framework.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets generated and/or analysed during the current study are available in the repository (10.5281/zenodo.18853801).All scRNA-seq data used in this study are available online from public databases. The raw and batch corrected counts from Perez et al. (2022) are available in the CellxGene database [https://cellxgene.cziscience.com/collections/436154da-bcf1-4130-9c8b-120ff9a888f2](https:/cellxgene.cziscience.com/collections/436154da-bcf1-4130-9c8b-120ff9a888f2) and GEO database under accession code GSE137029 ( [https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE137029](https:/www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE137029) ). The raw counts from Nehar-Belaid et al. (2020) are available in the GEO database under accession code GSE135779 ( [https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135779](https:/www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135779) ).\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eTsokos GC. The immunology of systemic lupus erythematosus. Nat Immunol. 2024;25:1332\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKernder A, et al. Delayed diagnosis adversely affects outcome in systemic lupus erythematosus: Cross sectional analysis of the LuLa cohort. Lupus. 2021;30:431\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHayter SM, Cook MC. Updated assessment of the prevalence, spectrum and case definition of autoimmune disease. Autoimmun rev. 2012;11:754\u0026ndash;65.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYen EY, Singh RR. Lupus\u0026mdash;An Unrecognized Leading Cause of Death in Young Females. Arthritis Rheumatol. 2018;70:1251\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArazi A, et al. The immune cell landscape in kidneys of patients with lupus nephritis. Nat Immunol. 2019;20:902\u0026ndash;14.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMistry P et al. \u003cem\u003eTranscriptomic, epigenetic, and functional analyses implicate neutrophil diversity in the pathogenesis of systemic lupus erythematosus.\u003c/em\u003e Proceedings of the National Academy of Sciences, 2019. 116: pp. 25222\u0026ndash;25228.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNehar-Belaid D, et al. Mapping systemic lupus erythematosus heterogeneity at the single-cell level. Nat Immunol. 2020;21:1094\u0026ndash;106.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePerez RK et al. Single-cell RNA-seq reveals cell type-specific molecular and genetic associations to lupus. Science, 2022. 376.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMahmood SB, et al. Evaluating Sex Differences in the Characteristics and Outcomes of Lupus Nephritis: A Systematic Review and Meta-Analysis. Glomerular Dis. 2024;4:19\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim JW et al. Sex hormones affect the pathogenesis and clinical characteristics of systemic lupus erythematosus. Front Med, 2022. 9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eScofield RH, et al. Klinefelter\u0026rsquo;s syndrome (47,XXY) in male systemic lupus erythematosus patients. Arthr Rhuem. 2008;58:2511\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu K, et al. X Chromosome Dose and Sex Bias in Autoimmune Diseases. Arthritis Rheumatol. 2016;68:1290\u0026ndash;300.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePalmer A, Tan IJ. Synergistic Effects of Extra X Chromosome on Development of Systemic Lupus Erythematosus and Sj\u0026ouml;gren Disease. ACR Open Rheumatology; 2025. p. 7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSouyris M et al. TLR7 escapes X chromosome inactivation in immune cells. Sci Immunol, 2018. 3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eForsyth KS, et al. The conneXion between sex and immune responses. Nature Reviews Immunology; 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang J et al. Unusual maintenance of X chromosome inactivation predisposes female lymphocytes for increased expression from the inactive X. PNAS, 2016. 113: pp. E2029-E2038.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSyrett CM et al. Altered X-chromosome inactivation in T cells may promote sex-biased autoimmune diseases. JCI Insight, 2019. 4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePyfrom S et al. The dynamic epigenetic regulation of the inactive X chromosome in healthy human B cells is dysregulated in lupus patients. PNAS, 2021. 118.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eForsyth KS, et al. Maintenance of X chromosome inactivation after T cell activation requires NF-κB signaling. Science Immunology; 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTomofuji Y et al. Quantification of escape from X chromosome inactivation with single-cell omics data. Cell Genomics, 2024. 4(8).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCrawford JD et al. The XIST lncRNA is a sex-specific reservoir of TLR7 ligands in SLE. JCI Insight, 2023. 8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDou DR, et al. Xist ribonucleoproteins promote female sex-biased autoimmunity. Cell. 2024;187:733\u0026ndash;49.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFiggett WA, et al. Machine learning applied to whole-blood RNA-sequencing data uncovers distinct subsets of patients with systemic lupus erythematosus. Clinical \u0026amp; Translational Immunology; 2019. p. 8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKegerreis B et al. Machine learning approaches to predict lupus disease activity from gene expression data. Sci Rep, 2019. 9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMa Y, et al. Accurate Machine Learning Model to Diagnose Chronic Autoimmune Diseases Utilizing Information From B Cells and Monocytes. Frontiers in Immunology; 2022. p. 13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChan Zuckerberg I. CELLxGENE. 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWolf FA, Angerer P, Theis FJ. SCANPY: large-scale single-cell gene expression data analysis. Genome Biol. 2018;19:15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHao Y, et al. Integrated analysis of multimodal single-cell data. Cell. 2021;184:3573\u0026ndash;87.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTeam RC. \u003cem\u003eR: A Language and Environment for Statistical Computing.\u003c/em\u003e 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePi\u0026ntilde;ero J, et al. The DisGeNET knowledge platform for disease genomics: 2019 update. Nucleic Acids Research; 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePedregosa F, et al. Scikit-learn: Machine Learning in Python. J Mach Learn Res. 2011;12:2825\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZou H, Hastie T. Regularization and Variable Selection Via the Elastic Net. J Royal Stat Soc Ser B. 2005;67:301\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKursa MB, Rudnicki WR. Feature Selection with the Boruta Package. J Stat Softw, 2010. 36.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYouden WJ. Index for rating diagnostic tests. Cancer. 1950;3:32\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChicco D, Jurman G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics. 2020;21:6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMatthews BW. Comparison of the predicted and observed secondary structure of T4 phage lysozyme. Biochim Biophys Acta. 1975;405:442\u0026ndash;51.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKruskal WH, Wallis WA. Use of Ranks in One-Criterion Variance Analysis. J Am Stat Assoc. 1952;47:583.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBenjamini Y, Hochberg Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. J Royal Stat Soc Ser B. 1995;57:289\u0026ndash;300.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLundberg SM, et al. Explainable machine-learning predictions for the prevention of hypoxaemia during surgery. Nat Biomedical Eng. 2018;2:749\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarrett T, et al. NCBI GEO: archive for functional genomics data sets\u0026mdash;update. Nucleic Acids Res. 2012;41:D991\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSubramanian A et al. Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics. Genome Biol, 2022. 23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGermain PL et al. Doublet identification in single-cell sequencing data using scDblFinder. F1000Research, 2021. 10: p. 979.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDom\u0026iacute;nguez Conde C et al. Cross-tissue immune cell analysis reveals tissue-specific features in humans. Science, 2022. 376.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen Y et al. edgeR v4: powerful differential analysis of sequencing data with expanded functionality and improved support for small counts and larger datasets. Nucleic Acids Res, 2025. 53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoulton VR, Tsokos GC. T cell signaling abnormalities contribute to aberrant immune cell function and autoimmunity. J Clin Invest. 2015;125:2220\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKatsuyama E, et al. The CD38/NAD/SIRTUIN1/EZH2 Axis Mitigates Cytotoxic CD8 T Cell Function and Identifies Patients with SLE Prone to Infections. Cell Rep. 2020;30:112\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiwrajka N et al. Impaired dynamic X-chromosome inactivation maintenance in T cells is a feature of spontaneous murine SLE that is exacerbated in female-biased models. J Autoimmun, 2023. 139.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoelzl S et al. Aging promotes reactivation of the Barr body at distal chromosome regions. Nat Aging, 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSoares M, et al. X-linked transcriptome dysregulation across immune cells in systemic lupus erythematosus. Biology Sex Differences. 2025;16:69.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin JX, Leonard WJ. The Common Cytokine Receptor γ Chain Family of Cytokines. Volume 10. Cold Spring Harbor Perspectives in Biology; 2018. p. a028449.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJansen R, et al. Sex differences in the human peripheral blood transcriptome. BMC Genomics. 2014;15:33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOliva M et al. The impact of sex on gene expression across human tissues. Science, 2020. 369.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMel\u0026eacute; M, et al. The human transcriptome across tissues and individuals. Science. 2015;348:660\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYin H, et al. 2D4, a humanized monoclonal antibody targeting CD132, is a promising treatment for systemic lupus erythematosus. Signal Transduct Target Therapy. 2024;9:323.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLu Q, et al. Demethylation of CD40LG on the Inactive X in T Cells from Women with Lupus. J Immunol. 2007;179:6352\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou Y, et al. T cell CD40LG gene expression and the production of IgG by autologous B cells in systemic lupus erythematosus. Clin Immunol. 2009;132:362\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePati\u0026ntilde;o-Martinez E, Kaplan MJ. Immunometabolism in systemic lupus erythematosus. Nat Rev Rheumatol. 2025;16:1\u0026ndash;19.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCotton AM, et al. Analysis of expressed SNPs identifies variable extents of expression from the human inactive X chromosome. Genome Biol. 2013;14:R122.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChaumeil J et al. \u003cem\u003eCombined Immunofluorescence, RNA Fluorescent In Situ Hybridization, and DNA Fluorescent In Situ Hybridization to Study Chromatin Changes, Transcriptional Activity, Nuclear Organization, and X-Chromosome Inactivation\u003c/em\u003e, in \u003cem\u003eX-Chromosome Inactivation\u003c/em\u003e. 2008. pp. 297\u0026ndash;308.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarber MRW, et al. Global epidemiology of systemic lupus erythematosus. Nat Rev Rheumatol. 2021;17:515\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-rheumatology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"brhm","sideBox":"Learn more about [BMC Rheumatology](http://bmcrheumatol.biomedcentral.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/brhm/default.aspx","title":"BMC Rheumatology","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Machine learning, Systemic lupus erythematosus, X chromosome inactivation, Single-cell transcriptomics","lastPublishedDoi":"10.21203/rs.3.rs-8991344/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8991344/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSystemic lupus erythematosus (SLE) has a significant female bias; however, it remains unclear whether X-chromosomal dysregulation increases the risk in females. Furthermore, while the role of X-linked genes in SLE pathogenesis is recognised, the analysis of genetic risks in SLE has largely focussed on autosomal genes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHere, we analysed a public single-cell RNA sequencing (scRNA-seq) dataset of peripheral blood mononuclear cells using machine learning to test whether X-linked gene expression predicts female SLE in a cell type-specific manner. Using pseudobulked expression profiles from 1.1 million cells across 228 female donors, we trained an ensemble of classifiers on all X-chromosomal genes or a literature-derived set of SLE-associated genes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn CD4⁺ and CD8⁺ T cells, models based on X-chromosomal genes achieved classification performance comparable to SLE gene models and were significantly enriched for XCI escape genes, with IL2RG, CD40LG, and TKTL1 emerging as key predictors. When applied to an independent paediatric–adult SLE cohort, CD8⁺ T-cell models retained robust performance, indicating a stable and reproducible X-linked signature in this subset. Flow cytometry revealed a modest increase in IL2RG protein expression in SLE T cells, consistent with partial escape from XCI in these cells.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThese findings demonstrate that X-linked transcriptional programs, enriched for known XCI escape genes and concentrated in activated T cells, encode a reproducible transcriptional signal of female SLE and provide a framework for dissecting the contributions of sex chromosomes to autoimmunity.\u003c/p\u003e","manuscriptTitle":"Machine learning reveals X chromosome transcriptional signatures that predict systemic lupus erythematosus in females","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-24 19:09:49","doi":"10.21203/rs.3.rs-8991344/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-04-20T16:09:18+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-09T05:26:38+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-04T20:41:44+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-31T09:48:19+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"43149688477214099895569078727352218973","date":"2026-03-31T03:02:46+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"209028817400703667465321316265298690517","date":"2026-03-26T16:47:31+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"323236140804782052644175635725595388436","date":"2026-03-26T13:31:33+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"336861205938149670537801617059218441757","date":"2026-03-19T14:29:11+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-19T08:17:44+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-19T06:29:16+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-04T18:51:31+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-04T01:53:25+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Rheumatology","date":"2026-03-04T01:48:19+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-rheumatology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"brhm","sideBox":"Learn more about [BMC Rheumatology](http://bmcrheumatol.biomedcentral.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/brhm/default.aspx","title":"BMC Rheumatology","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a27ede09-e4aa-4eb4-a1f0-893f6d5d042f","owner":[],"postedDate":"March 24th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-11T22:53:27+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-24 19:09:49","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8991344","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8991344","identity":"rs-8991344","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.