AI-Driven Synthetic Cohorts to Explore Genetic Associations: Lessons from Testicular Cancer with Relevance to Rare Conditions

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Objective Rare cancers often involve small patient cohorts, which limit statistical power and hinder biomarker discovery. This proof-of-concept study evaluates the feasibility of repurposing the Synthpop R package—originally designed for data anonymization—to generate synthetic datasets and improve statistical inference in rare cancer studies. We demonstrate this approach using the association between the MC4R rs79783591 variant and clinical outcomes in testicular germ cell tumors (TGCT). Materials and Methods A retrospective cohort of 220 Mexican TGCT patients was analyzed, identifying 9 heterozygous carriers of the MC4R variant. To overcome limited sample size, we used Synthpop to generate 234 synthetic carriers, maintaining the original data structure and statistical distributions. Dataset fidelity was validated using machine learning models, principal component analysis, and structural similarity metrics. Cox proportional hazards models and Kaplan–Meier survival analyses assessed associations with overall survival. Results Real carriers were diagnosed at a younger median age (22 vs. 26 years). In the synthetic cohort, the MC4R variant was associated with a threefold increased mortality risk (HR = 3.15; 95% CI: 2.06–4.82; p < 0.001), supporting findings in the real cohort (HR = 5.69; 95% CI: 1.56–20.7; p = 0.008). Synthetic data narrowed confidence intervals and improved effect size estimation. Conclusion Repurposing the Synthpop R package provides a novel approach to enhance statistical power in studies with small sample sizes. This strategy can improve inference reliability and accelerate biomarker discovery in rare cancer research; however, further validation in independent cohorts is required to confirm these findings.
Full text 113,245 characters · extracted from preprint-html · click to expand
AI-Driven Synthetic Cohorts to Explore Genetic Associations: Lessons from Testicular Cancer with Relevance to Rare Conditions | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article AI-Driven Synthetic Cohorts to Explore Genetic Associations: Lessons from Testicular Cancer with Relevance to Rare Conditions Juan Alberto Ríos-Rodríguez, Sylvia Harari-Arakindji, Berenice Cuevas-Estrada, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8099876/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Objective Rare cancers often involve small patient cohorts, which limit statistical power and hinder biomarker discovery. This proof-of-concept study evaluates the feasibility of repurposing the Synthpop R package—originally designed for data anonymization—to generate synthetic datasets and improve statistical inference in rare cancer studies. We demonstrate this approach using the association between the MC4R rs79783591 variant and clinical outcomes in testicular germ cell tumors (TGCT). Materials and Methods A retrospective cohort of 220 Mexican TGCT patients was analyzed, identifying 9 heterozygous carriers of the MC4R variant. To overcome limited sample size, we used Synthpop to generate 234 synthetic carriers, maintaining the original data structure and statistical distributions. Dataset fidelity was validated using machine learning models, principal component analysis, and structural similarity metrics. Cox proportional hazards models and Kaplan–Meier survival analyses assessed associations with overall survival. Results Real carriers were diagnosed at a younger median age (22 vs. 26 years). In the synthetic cohort, the MC4R variant was associated with a threefold increased mortality risk (HR = 3.15; 95% CI: 2.06–4.82; p < 0.001), supporting findings in the real cohort (HR = 5.69; 95% CI: 1.56–20.7; p = 0.008). Synthetic data narrowed confidence intervals and improved effect size estimation. Conclusion Repurposing the Synthpop R package provides a novel approach to enhance statistical power in studies with small sample sizes. This strategy can improve inference reliability and accelerate biomarker discovery in rare cancer research; however, further validation in independent cohorts is required to confirm these findings. Epigenetics & Genomics Biostatistics Artificial Intelligence and Machine Learning Synthpop Synthetic data Testicular germ cell tumors MC4R Cancer genomics Figures Figure 1 Figure 2 Background and Significance Handling small sample sizes in statistical analysis is a significant challenge, as limited data often lacks sufficient power to yield meaningful conclusions ( 1 ). This limitation is particularly pronounced in the study of rare diseases, generally defined as conditions affecting approximately 1 in every 2,000 to 2,500 individuals ( 2 ). In oncology, about one in five cancer patients has a rare malignancy, posing substantial challenges for epidemiological research. Methodological adjustments are therefore necessary to reliably detect statistically significant differences within these small patient groups ( 1 ). Testicular cancer (TCa) exemplifies this challenge, being a rare malignancy worldwide with an age-standardized incidence rate (ASIR) of 1.8 and a mortality rate (ASMR) of 0.2 per 100,000 men, according to GLOBOCAN 2020 ( 3 ). However, these rates vary notably across regions. In Mexico, the ASIR is 5.1 per 100,000 men ( 4 ), considerably above the global average, and the ASMR is 0.94 per 100,000 men between 2015 and 2019, ranking third highest in Latin America ( 3 ). Approximately 95–98% of testicular tumors are germ cell tumors (TGCT), which are typically highly curable, with 5-year overall survival rates exceeding 80%. Despite this favorable prognosis, TCa remains a public health concern in Mexico, partly due to limited research on TGCT within the population ( 5 – 7 ). Notably, valuable insights have been provided into both the somatic genetic alterations associated with chemoresistance and the social determinants that influence treatment outcomes. ( 5 – 7 ). Germline genetic variants could potentially explain both the increased incidence and mortality of TCa. To date, no high-penetrance variant has been directly associated with risk, and the low frequency of the disease makes variant analysis challenging( 8 ). This is particularly evident for genes like MC4R , known for its role in autosomal dominant obesity ( 9 ) and is expressed in fetal germ cell tests ( 10 , 11 ). Missense variants in MC4R have been shown to promote teratoma formation in knock-in mice, potentially through modulation of anti-apoptotic signaling, which may favor transdifferentiation into somatic cells or increased parthenogenesis ( 10 , 11 ). Large-scale data analyses have also suggested associations between Single Nucleotide Polymorphisms (SNP) near MC4R and other cancers, including endometrial and breast cancer, underscoring the need to investigate its potential role in gonadal cancer. ( 11 ) Given the challenges of studying rare diseases, advanced computational methods have gained importance. Synthetic data generation offers a versatile solution by using techniques like Classification and Regression Trees (CARTs) to accurately replicate key oncology endpoints, including overall and progression-free survival. ( 12 , 13 ). The synthpop package is a powerful non-parametric tool that sequentially generates synthetic datasets preserving the statistical properties and correlations of the original data, without containing real individual records ( 14 ). Comparative analyses highlight its superior performance over other methods, including deep learning, particularly for small datasets ( 15 ) This capability is crucial as it allows researchers to conduct analyses and obtain results consistent with those from the real data ( 16 , 17 ). Building on these findings, this study evaluates CART-based synthetic data generation using synthpop as an artificial intelligence (AI) approach to strengthen epidemiological analysis. As a proof of concept, it focuses on TCa —a rare disease— and investigates the MC4R rs79783591 genetic variant, which could be a critical factor in the Mexican population, in relation to disease risk and prognosis, addressing the challenges posed by a small patient cohort. Materials and Methods The manuscript was prepared following the STROBE guidelines for cohort studies and MI-CLAIM. A representative cohort of Mexican patients with TGCT, treated between 2007 and 2020 at the Instituto Nacional de Cancerología (INCan), a national referral center in Mexico City, was included after prior informed consent (022/068/OMI; CEI/052/22). The sample size was determined for exploratory purposes based on previously reported prevalence in an external cohort( 7 ). Patients ≥ 15 years with histologically confirmed TGCT, complete clinical records, and at least 5 years of follow-up were eligible, while those with uncertain diagnoses or incomplete records were excluded. Initial exome sequencing was performed on 40 patients from an external cohort previously published ( 15 ), which were reanalyzed to detect variants in the MC4R gene. Genotyping of the rs79783591 variant was then extended to the entire INCan cohort using real-time polymerase chain reaction (qPCR) with TaqMan probes. Demographic and clinical information was collected from medical records, and logistic regression was performed to estimate the association between genetic variants and TGCT risk, calculating Odds Ratios (OR) with 95% confidence intervals. To address the limited sample size, synthetic data were generated using the synthpop package in R. Datasets were created from the subpopulation of mutation carriers using a CART model, preserving the distributions, correlations, and multivariate structure of the original data. Multiple synthetic datasets were generated across simulations, with seeds set sequentially according to the number of groups, aiming to match the size of the non-carrier population. Validation of the synthetic datasets included statistical comparisons with the original data (Kolmogorov-Smirnov, t-test or Mann-Whitney U, Chi-square or Fisher’s exact tests with Benjamini-Hochberg adjustment), assessment of structural similarity (Spearman correlation matrices, principal component analysis with permutation testing). Application of machine learning models (Random Forest, XGBoost, SVM, and logistic regression) was performed to confirm that real and synthetic data could not be distinguished. Hyperparameters were optimized by grid search where applicable, and model performance was evaluated using an 80/20 train-test split with 5-fold cross-validation. Performance metrics included area under the curve (AUC) with confidence intervals, sensitivity, specificity, F1 score, and Cohen’s kappa. Summary statistics are reported as mean ± standard deviation or median with 25th and 75th percentiles for continuous variables, and as percentages for categorical variables. Qualitative variables were compared using Chi-square or Fisher’s exact tests, while quantitative variables were assessed for normality (Shapiro-Wilk) and compared with Student’s t-test or Mann-Whitney U test as appropriate. Survival analyses employed Kaplan-Meier curves, with group comparisons assessed using log-rank, Peto-Peto, and Tarone-Ware tests. Variance of key clinical variables (BMI and follow-up time) was compared using F-tests. Cox regression models assessed the variant’s effect on mortality risk, with standard errors of coefficients and C-index calculated to evaluate precision and discriminative ability; a synthetic cohort was employed to increase statistical power. These additional tests were included to account for potential differences in early- and late-event survival between carriers and non-carriers. Statistical significance was set at p < 0.05. All analyses were conducted in RStudio (version 2024.12.0). Results Based on a prevalence of 0.125 from the previously published external cohort, a finite population of 2,700, a 5% margin of error, and 97% confidence, the initial sample size was 244, with 24 excluded by eligibility criteria, leaving 220 TGCT patients. Germline DNA analysis identified 9 (4.1%) heterozygous carriers of the rs79783591 variant in the MC4R gene (Fig. 1). Sixteen patients initially reviewed were excluded for not meeting the eligibility criteria. The carrier frequency was compared to 0.0062 in the general Mexican population, as reported by the Population Architecture using Genomics and Epidemiology (PAGE) Study database via ClinVar ( 18 ). Logistic regression showed a statistically significant association, with an Odds Ratio (OR) of 3.42 (95% CI: 1.07–8.5; p = 0.02), suggesting a potential link between the variant and increased TGCT risk. A bivariate analysis was performed on the real cohort to compare the clinical characteristics between patients who were carriers and non-carriers of the MC4R genetic variant (Table 1 ). The analysis revealed a statistically significant difference in age at diagnosis. Carriers were diagnosed at a median age of 22 years, which was notably younger than the non-carriers (median of 26 years; p = 0.028). No significant differences were observed in Body Mass Index (BMI) or the presence of distant metastases. Regarding tumor histology, non-seminoma was the most common subtype. Furthermore, the study found a higher mortality rate among carriers (33.33%) compared to non-carriers (14.7%). However, this association did not reach statistical significance (p = 0.147). Table 1 Clinical Characteristics Comparison Between MC4R Variant Carriers and Non-Carriers in the Cohort Variable Total N = 220 (%) Non-carriers n = 211 (95.9%) Carriers n = 9 (4.1%) p-value a Age Median (p25-p75) 26 (22–31) 26 ( 23 – 26 ) 22 ( 21 – 25 ) 0.028* BMI Median (p25-p75) 24.3 (22.7–27.6) 24.3 (22.7–27.7) 24.2 (21.2–26.3) 0.348 Histology seminoma Non-seminoma 84 (38.2) 136 (61.8) 82 (38.9) 136 (61.1) 2 (22.22) 7 (77.78) 0.488 Distant metastasis Absent Present 144 (65.5) 76 (34.5) 137 (64.9) 74 (35.1) 7 (77.78) 2 (22.22) 0.722 Vital status Alive Deceased 186 (84.5) 34 (15.5) 180 (85.3) 31 (14.7) 6 (66.67) 3 (33.33) 0.147 BMI: Body Mass Index; a Mann-Whitney U test, b Chi-square test, significance at p-value < 0.05* The synthetic carrier cohort was generated to preserve the key clinical characteristics observed in the original dataset, consisting of 234 patients created across 26 sets of 9 individuals each, which was confirmed through a series of rigorous validation tests (Table 2 ). Statistical comparisons showed strong consistency between the synthetic and real data, with no significant differences in most variables. Furthermore, no observable differences in the Variance Inflation Factor (VIF) were found among the continuous variables, suggesting the absence of multicollinearity in the synthetic data. To evaluate structural similarity, Spearman correlation matrices were generated for both datasets. A permutation test confirmed that the overall correlation structures were not significantly different (p > 0.05). However, a subtle inconsistency was noted for the Age variable, where the correlation with Surveillance time had a reversed sign (real: -0.0183, synthetic: 0.0317). To ensure the reliability of the results, the Age variable was therefore omitted from subsequent analyses (Supplementary Fig. 1). Principal Component Analysis (PCA) was also used to assess structural similarity. The PCA did not show significant differences when comparing PC2 vs. PC3 and PC1 vs. PC3. However, when comparing PC1 vs. PC2, a permutation test was found to be significant despite a clear visual overlap (Supplementary Fig. 2). Finally, machine learning models (Logistic Regression, Random Forest, XGBoost, and SVM) were employed to test their ability to distinguish between real and synthetic data. The models' performance was poor across the board, with the highest AUC being 0.674 for Random Forest and a logistic regression AUC of 0.553 (Supplementary Table 1). Table 2 Comparison of clinical characteristics between the synthetic carrier ohort and the real carrier cohort Variable Synthetic Carrier n = 234 (%) Real Carriers n = 9 (%) p-value a Age Median (p25-p75) 22 ( 21 – 25 ) 22 ( 21 – 25 ) 0.88 b BMI Median (p25-p75) 24.2 (21.2–26.3) 24.2 (21.2–26.3) 0.96 b Histology Seminoma Non-seminoma 62 (26.5) 172 (73.5) 2 (22.22) 7 (77.78) 1 Distant Metastasis Absent Present 185 (79.1) 49 (20.9) 7 (77.78) 2 (22.22) 1 Vital Status Alive Deceased 144 (61.5) 90 (38.5) 6 (66.7) 3 (33.3) 1 BMI: Body Mass Index; a Chi-square test , b Mann-Whitney U test, significance at p-value < 0.05* The Kaplan-Meier survival analysis with a log-rank test was performed to assess time to mortality in both real and synthetic carrier cohorts (Fig. 2). In the real cohort, the survival curve for carriers visually separated below that of the non-carriers. In the real cohort, survival curve separation was not significant (log-rank, Peto-Peto, Tarone-Ware p ≈ 0.09–0.12), likely due to few carriers (N = 9). In the synthetic cohort, separation remained but was highly significant (p < 0.0001). This cohort also exhibited notably narrower and more stable confidence intervals around the carriers' survival curve. Table 3 illustrates the results of the multivariate analysis using both logistic regression (OR) and Cox regression (HR) models to identify factors associated with mortality in the real and synthetic MC4R cohorts. In both cohorts, BMI and histology did not show a significant association with mortality in either model (p > 0.05). In contrast, the presence of distant metastasis was a highly significant predictor of mortality in both the real cohort (OR = 44.4, 95% CI: 11.7–296; HR = 22.9, 95% CI: 6.50–80.5) and the synthetic cohort (OR = 4.37, 95% CI: 2.62–7.45; HR = 2.92, 95% CI: 2.01–4.24), with p-values consistently below 0.001 for all four models. The presence of the MC4R variant was also a significant predictor of mortality. In the real cohort, it showed a significant association in both the logistic regression (OR = 15.1, 95% CI: 1.65–158, p = 0.016) and the Cox regression models (HR = 5.69, 95% CI: 1.56–20.7, p = 0.008). These findings were confirmed and strengthened in the synthetic cohort, where the variant was a highly significant predictor in both the logistic regression (OR = 5.11, 95% CI: 3.06–8.84, p < 0.001) and Cox regression models (HR = 3.15, 95% CI: 2.06–4.82, p < 0.001). Notably, the synthetic models yielded significantly narrower and more precise confidence intervals for both the OR and HR, with lower standard errors for all Cox coefficients (Carrier: 0.614 → 0.216; BMI: 0.049 → 0.029; Histology: 0.610 → 0.211) and a minor reduction in C-index (0.675 → 0.636), providing a more stable estimate of the variant's effect. Table 3 Regression models of mortality in real and synthetic MC4R Cohorts Variable Real OR (IC) p-value Simulated OR (IC) p-value Real HR (IC) p-value Simulated HR (IC) p-value BMI 1.08 (0.94, 1.23) 0.272 0.97 (0.90, 1.05) 0.487 1.07 (0.96, 1.18) 0.214 0.99 (0.93, 1.05) 0.768 Histology Seminoma Non-seminoma Ref. 3.24 (0.94, 15.1) 0.0874 Ref. 1.10 (0.66, 1.86) 0.714 Ref. 2.56 (0.76, 8.69) 0.131 Ref. 1.17 (0.78, 1.78) 0.449 Distant Metastasis Absent Present Ref. 44.4 (11.7, 296) < 0.001 Ref. 4.37 (2.62, 7.45) < 0.001 Ref. 22.9 (6.50, 80.5) < 0.001 Ref. 2.92 (2.01, 4.24) < 0.001 Mutation Carrier Absent Present Ref. 15.1 (1.65, 158) 0.016 Ref. 5.11 (3.06, 8.84) < 0.001 Ref. 5.69 (1.56, 20.7) 0.008 Ref. 3.15 (2.06, 4.82) < 0.001 BMI: Body Mass Index; significance at p-value < 0.05* Discussion This study provides the first evidence of an association between the MC4R rs79783591 variant and cancer, suggesting a potential role as a risk and prognostic factor in TGCT. Carriers showed an increased mortality risk, adjusted for BMI, disease location, and histology, and were diagnosed at a median age four years earlier than non-carriers. These findings are consistent with observations in Hispanic populations, who often present with more aggressive and chemoresistant disease at younger ages; these patterns do not appear to be fully explained by sociodemographic factors ( 6 , 7 , 19 , 20 ). Although statistical significance was limited by sample size, simulated cases confirmed a threefold increase in mortality. The variant (rs79783591, c.806T > A; p.Ile269Asn) lies in MC4R (18q21.32), with a gnomAD frequency of 0.00011 and 0.00629 in the Mexican PAGE cohort. Despite its conflicting ClinVar classification ( 18 ), the variant has been associated with obesity in Mexican children and adults ( 9 ); however, carriers here exhibited relatively normal BMI values. TGCT has high heritability, but no high-penetrance germline variants or validated prognostic genetic markers have been established ( 21 – 23 ). Functional hypotheses suggest that MC4R variants may influence germ cell survival via anti-apoptotic pathways, potentially disrupting the balance between proliferation and apoptosis ( 11 ). Confirmation of these associations and clarification of underlying mechanisms will require larger, multi-center studies. Additionally, we were unable to calculate the haplotype as the presence of another potentially associated variant could not be excluded. To address the limitations of small sample size, we implemented synthetic data generated with synthpop, which improved model performance, narrowed confidence intervals, and enabled detection of significant Kaplan–Meier differences that were only marginal in the real dataset, while preserving significance in multivariate analyses. This approach preserves the trends of the original cohort while enhancing statistical power and has been shown to outperform other methods, such as deep learning-based Generative Adversarial Networks (GANs), particularly in small, tabular clinical datasets ( 14 , 15 , 24 – 26 ). In fact, given the extremely limited number of carriers in our study (n = 9), the use of GANs would be impractical due to overfitting and instability in training, further supporting the choice of CART-based synthetic data generation as a more robust and transparent alternative. While synthetic data provide valuable insights and help identify potential prognostic factors, they do not perfectly reflect real populations and may introduce bias if original data variability is limited. Complementary tests sensitive to early- and late-event differences (Peto-Peto and Tarone-Ware) confirmed trends in the real cohort and detected significant differences in the synthetic cohort, supporting the utility of synthetic data to increase statistical power. While synthetic data provide valuable insights, limited variability in the original cohort may introduce bias, potentially distorting observed clinical patterns. As shown in our Kaplan–Meier analysis, a “fall” was observed due to such variability, which might not reflect true clinical behavior. Permutation analyses of PCA components further highlighted that some relationships between variables may be incompletely captured. These limitations emphasize that even within small cohorts, increasing the number of cases and its diversity can improve robustness and reduce potential distortions. Importantly, the synthetic cohort produced lower standard errors and slightly reduced C-index values compared to the real cohort, indicating more precise and stable estimates of the variant’s effect while preserving discriminative ability. Despite the small cohort size, INCan is a national referral center, and its cases may reasonably reflect TGCT patients across Mexico ( 6 ). Compared to more complex methods such as GANs, synthpop remains a transparent and accessible approach that preserves essential statistical properties, enables a novel dataset expansion without discarding variables based on arbitrary thresholds, and facilitates the identification of potential prognostic markers that might otherwise be overlooked. Although synthpop provides a valuable approach for increasing statistical power in small cohorts, its limitations should be acknowledged. The limited number of carriers in the original cohort may restrict detection of statistically significant associations, and synthetic data do not perfectly reflect a real population; rather, they model what would occur in a population with similar characteristics to the observed cases. Nonetheless, synthpop is an accessible, transparent, and easily applicable method that preserves essential statistical properties, allows dataset expansion without discarding variables based on p-value thresholds, and facilitates identification of potential biological risk factors. Additionally, the synthetic data benefit from the package’s built-in prospective design, which anonymizes patient information, improving ethical considerations and compliance. Importantly, integrating synthetic data generation with machine learning validation represents a novel framework for augmenting small cohorts, including in TCa, and is generalizable to other rare conditions where limited sample sizes hinder robust statistical inference. Despite the strengthened findings provided by synthetic data, further research in slightly expanded multi-center cohorts is needed to confirm these results, explore additional genetic variants, and assess the robustness of this approach, ultimately supporting the prioritization of candidate biomarkers and clinical factors for translational applications. Conclusion The MC4R rs79783591 variant may represent both a risk and prognostic factor for TGCT in Mexican patients. This study provides preliminary evidence of this association, addressing the scarce knowledge of genetic susceptibility to this cancer type in Hispanic populations. Importantly, this research illustrates the utility of artificial intelligence–driven approaches to mitigate limitations posed by small sample sizes, a frequent obstacle in rare disease studies. By generating synthetic data through a repurposing of synthpop, the analysis gained statistical power, uncovering associations that conventional methods might have overlooked. Future studies with more diverse cohorts are needed to validate these findings and investigate additional genetic variants. Despite its limitations, the use of AI tools like synthpop opens promising avenues for personalized medicine and research in data-limited populations. Declarations CRediT authorship contribution statement: Conceptualization: Juan Alberto Ríos-Rodríguez, Berenice Cuevas-Estrada, Sylvia Harari-Arakindji, Rodrigo González‐Barrios, Talia Wegman-Ostrosky; Data curation: Berenice Cuevas-Estrada, Juan Alberto Ríos-Rodríguez, José Antonio García-Pacheco, Sylvia Harari-Arakindji: Formal analysis: Juan Alberto Ríos-Rodríguez, Berenice Cuevas-Estrada, José Antonio García-Pacheco; Investigation: Berenice Cuevas-Estrada, Juan Alberto Ríos-Rodríguez, José Antonio García-Pacheco, Sylvia Harari-Arakindji: Methodology: Juan Alberto Ríos-Rodríguez, Berenice Cuevas-Estrada; Project administration: Rodrigo González‐Barrios, Talia Wegman-Ostrosky; Resources: Rodrigo González‐Barrios, Talia Wegman-Ostrosky, Nora Sobrevilla-Moreno, Miguel A Jiménez-Ríos; Supervision: Rodrigo González‐Barrios, Talia Wegman-Ostrosky; Visualization: Juan Alberto Ríos-Rodríguez, Berenice Cuevas-Estrada, José Antonio García-Pacheco; Writing – original draft: Juan Alberto Ríos-Rodríguez, Sylvia Harari-Arakindji, Talia Wegman-Ostrosky, José Antonio García-Pacheco, Berenice Cuevas-Estrada; Writing – review & editing: Patricia Ostrosky-Wegman, Alejandro Mohar-Betancourt, Rodrigo González‐Barrios, Berenice Cuevas-Estrada, Juan Alberto Ríos-Rodríguez, Talia Wegman-Ostrosky, José Antonio García-Pacheco, Sylvia Harari-Arakindji Declaration of Competing Interest The authors declare no conflicts of interest. Funding statement This work was supported by Instituto Nacional de Cancerología, Consejo Nacional de Humanidades, Ciencias y Tecnologías (RG-B, 2017-2-290041) and Financiamiento de Proyectos de Investigación para la Salud (FPIS-INCAN-4352) from the Dirección General de Políticas de Investigación en Salud (DGPIS), Secretaría de Salud, México. Acknowledgments J.A.R.-R acknowledges the Master’s in Biomedical Sciences (MBC) program by the School of Medicine and Health Sciences, Tecnologico de Monterrey, and was supported by Secretaría de Ciencia, Humanidades, Tecnología e Innovación (SECIHTI) fellowship 1247799. Appreciation is also extended to the Instituto Potosino de Investigación Científica y Tecnológica A.C. (IPICYT) for its academic collaboration through the diploma program in Artificial Intelligence for Health Informatics. B.C.-E was supported by SECIHTI fellowship 1145899 and J.A.G.-P fellowship 1328330. Ethics statement The study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board of the Instituto Nacional de Cancerología (012/031/ICI, 2012–2022) (022/068/OMI) (CEI/052/22). All participants provided written informed consent prior to inclusion in the study, following institutional guidelines. Data and Code Availability Statement The data and code supporting the findings of this study are available upon reasonable request from the corresponding author. References Whicher D, Philbin S, Aronson N. An overview of the impact of rare disease characteristics on research methodology. Orphanet J Rare Dis [Internet]. 2018 Jan 19 [cited 2025 Feb 16];13(1):14. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC5775563/ Hee SW, Willis A, Tudur Smith C, Day S, Miller F, Madan J, et al. Does the low prevalence affect the sample size of interventional clinical trials of rare diseases? An analysis of data from the aggregate analysis of clinicaltrials.gov. Orphanet J Rare Dis [Internet]. 2017 Mar 2 [cited 2025 Feb 16];12(1):1–20. Available from: https://ojrd.biomedcentral.com/articles/10.1186/s13023-017-0597-1 Crispin-Rios Y, Faura-Gonzales M, Torres-Roman JS, Quispe-Vicuña C, Franco-Jimenez US, Valcarcel B, et al. Testicular cancer mortality in Latin America and the Caribbean: trend analysis from 1997 to 2019. BMC Cancer [Internet]. 2023 Dec 1 [cited 2025 Jun 18];23(1):1–8. Available from: https://bmccancer.biomedcentral.com/articles/10.1186/s12885-023-11511-z Surveillance, Epidemiology, and End Results (SEER) Program [Internet]. 2025 [cited 2025 Jun 16]. Cancer Stat Facts: Testicular Cancer. Available from: https://seer.cancer.gov/statfacts/html/testis.html?utm_source=chatgpt.com Cuevas-Estrada B, Montalvo-Casimiro M, Munguia-Garza P, Ríos-Rodríguez JA, González-Barrios R, Herrera LA. Breaking the Mold: Epigenetics and Genomics Approaches Addressing Novel Treatments and Chemoresponse in TGCT Patients. International Journal of Molecular Sciences 2023, Vol 24, Page 7873 [Internet]. 2023 Apr 26 [cited 2024 Dec 11];24(9):7873. Available from: https://www.mdpi.com/1422-0067/24/9/7873/htm Ríos-Rodríguez JA, Montalvo-Casimiro M, Álvarez-López DI, Reynoso-Noverón N, Cuevas-Estrada B, Mendoza-Pérez J, et al. Understanding Sociodemographic Factors among Hispanics Through a Population-Based Study on Testicular Cancer in Mexico. J Racial Ethn Health Disparities. 2025 Feb;12:148–60. González-Barrios R, Alcaraz N, Montalvo-Casimiro M, Cervera A, Arriaga-Canon C, Munguia-Garza P, et al. Genomic Profile in a Non-Seminoma Testicular Germ-Cell Tumor Cohort Reveals a Potential Biomarker of Sensitivity to Platinum-Based Therapy. Cancers (Basel) [Internet]. 2022 May 1 [cited 2025 Aug 19];14(9):2065. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC9101377/ Islam R, Hansen A, Liesen A, Schorle H. An update on the genetic predisposition of testicular germ cell tumors. Transl Androl Urol [Internet]. 2024 [cited 2025 Jun 16];13(3):476–8. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC10999021/ Vázquez-Moreno M, Zeng H, Locia-Morales D, Peralta-Romero J, Asif H, Maharaj A, et al. The Melanocortin 4 Receptor p.Ile269Asn Mutation Is Associated with Childhood and Adult Obesity in Mexicans. J Clin Endocrinol Metab [Internet]. 2020 Apr 1 [cited 2025 Feb 16];105(4):E1468–77. Available from: https://pubmed.ncbi.nlm.nih.gov/31841602/ Naser AA, Miyazaki T, Wang J, Takabayashi S, Pachoensuk T, Tokumoto T. MC4R mutant mice develop ovarian teratomas. Sci Rep [Internet]. 2021 Dec 1 [cited 2025 Aug 16];11(1). Available from: https://pubmed.ncbi.nlm.nih.gov/33568756/ Seki S, Ohura K, Miyazaki T, Naser AA, Takabayashi S, Tsutsumi E, et al. The Mc4r gene is responsible for the development of experimentally induced testicular teratomas. Sci Rep [Internet]. 2023 Dec 1 [cited 2025 Aug 16];13(1). Available from: https://pubmed.ncbi.nlm.nih.gov/37127675/ Rujas M, Martín Gómez del Moral Herranz R, Fico G, Merino-Barbancho B. Synthetic data generation in healthcare: A scoping review of reviews on domains, motivations, and future applications. Int J Med Inform [Internet]. 2025 Mar 1 [cited 2025 Aug 16];195:105763. Available from: https://www.sciencedirect.com/science/article/pii/S138650562400426X?via%3Dihub Akiya I, Ishihara T, Yamamoto K. Comparison of Synthetic Data Generation Techniques for Control Group Survival Data in Oncology Clinical Trials: Simulation Study. JMIR Med Inform [Internet]. 2024 [cited 2025 Aug 16];12. Available from: https://pubmed.ncbi.nlm.nih.gov/38889082/ Quintana DS. A synthetic dataset primer for the biobehavioural sciences to promote reproducibility and hypothesis generation. Elife [Internet]. 2020 Mar 1 [cited 2025 Aug 16];9. Available from: https://pubmed.ncbi.nlm.nih.gov/32159513/ Miletic M, Sariyar M. Utility-based Analysis of Statistical Approaches and Deep Learning Models for Synthetic Data Generation With Focus on Correlation Structures: Algorithm Development and Validation. 2025 Mar; Braddon AE, Robinson S, Alati R, Betts KS. Exploring the utility of synthetic data to extract more value from sensitive health data assets: A focused example in perinatal epidemiology. Paediatr Perinat Epidemiol [Internet]. 2023 May 1 [cited 2025 Aug 19];37(4):292–300. Available from: https://pubmed.ncbi.nlm.nih.gov/36482827/ Major-Smith D, Kwong ASF, Timpson NJ, Heron J, Northstone K. Releasing synthetic data from the Avon Longitudinal Study of Parents and Children (ALSPAC): Guidelines and applied examples. Wellcome Open Res [Internet]. 2024 Dec 24 [cited 2025 Aug 19];9:57. Available from: https://pubmed.ncbi.nlm.nih.gov/39931104/ Landrum MJ, Lee JM, Riley GR, Jang W, Rubinstein WS, Church DM, et al. ClinVar: Public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res [Internet]. 2014 Jan 1 [cited 2025 Jun 22];42(D1). Available from: https://pubmed.ncbi.nlm.nih.gov/24234437/ Chertack N, Ghandour RA, Singla N, Freifeld Y, Hutchinson RC, Courtney K, et al. Overcoming sociodemographic factors in the care of patients with testicular cancer at a safety net hospital. Cancer [Internet]. 2020 Oct 10 [cited 2025 Aug 16];126(19):4362–70. Available from: https://acsjournals.onlinelibrary.wiley.com/doi/10.1002/cncr.33076 Rezaee ME, Elias R, Li HL, Agrawal P, Pallauf M, Enikeev D, et al. Survival outcomes and molecular drivers of testicular cancer in hispanic men. Urologic Oncology: Seminars and Original Investigations [Internet]. 2024 Sep 1 [cited 2025 Aug 16];42(9):293.e1-293.e7. Available from: https://pubmed.ncbi.nlm.nih.gov/38821727/ Cuevas-Estrada B, Montalvo-Casimiro M, Munguia-Garza P, Ríos-Rodríguez JA, González-Barrios R, Herrera LA. Breaking the Mold: Epigenetics and Genomics Approaches Addressing Novel Treatments and Chemoresponse in TGCT Patients. Int J Mol Sci [Internet]. 2023 May 1 [cited 2025 Aug 19];24(9). Available from: https://pubmed.ncbi.nlm.nih.gov/37175579/ Stang A, McMaster ML, Sesterhenn IA, Rapley E, Huddart R, Heimdal K, et al. Histological features of sporadic and familial testicular germ cell tumors compared and analysis of age-related changes of histology. Cancers (Basel) [Internet]. 2021 Apr 1 [cited 2025 Aug 19];13(7). Available from: https://pubmed.ncbi.nlm.nih.gov/33916078/ Islam R, Hansen A, Liesen A, Schorle H. An update on the genetic predisposition of testicular germ cell tumors. Transl Androl Urol [Internet]. 2024 [cited 2025 Aug 19];13(3):476–8. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC10999021/ Fedurek P, Lacroix L, Lehmann J, Aktipis A, Cronk L, Townsend C, et al. Status does not predict stress: Women in an egalitarian hunter-gatherer society. Evol Hum Sci [Internet]. 2020 [cited 2025 Aug 16];2. Available from: https://pubmed.ncbi.nlm.nih.gov/37588349/ Naughton M, Weaving D, Scott T, Compton H. Synthetic Data as a Strategy to Resolve Data Privacy and Confidentiality Concerns in the Sport Sciences: Practical Examples and an R Shiny Application. Int J Sports Physiol Perform [Internet]. 2023 Jul 18 [cited 2025 Aug 16];18(10):1213–8. Available from: https://journals.humankinetics.com/view/journals/ijspp/18/10/article-p1213.xml Warmenhoven J, Impellizzeri FM, Shrier I, Vigotsky AD, Lolli L, Menaspà P, et al. Synthetic Data for Sharing and Exploration in High-Performance Sport: Considerations for Application. Sports Medicine [Internet]. 2025 [cited 2025 Aug 16]; Available from: https://pubmed.ncbi.nlm.nih.gov/40569343/ Additional Declarations The authors declare no competing interests. Supplementary Files Supplementaryfigure1.png Supplementary figure 1. Variance inflation factors and Spearman correlation matrices of real and synthetic carrier data This figure displays the VIFs for the variables Surveillance time, BMI, and Age in the real carrier dataset, indicating minimal multicollinearity (a). It also presents heatmap matrices showing the Spearman correlation coefficients for the real carrier data (b) and the synthetic carrier data (c). The structural similarity between the correlation matrices was confirmed by a permutation test, which resulted in a non-significant p-value (p > 0.05), indicating that the synthetic dataset effectively preserves the dependency structures of the original data. Supplementaryfigure2.png Supplementary figure 2. PCA of real and simulated carriers Principal Component Analysis (PCA) comparing the "Real Carriers Group" (blue) with the "Simulated Carriers Group" (red). Both groups are complemented with the same set of real patient controls. The two-dimensional scatter plots show the projections of the data points onto the planes formed by the first three principal components (PC1, PC2, and PC3), which explain 39.8%, 31.4%, and 28.8% of the total variance, respectively. The p-values at the top of each graph indicate the statistical significance of the difference between the two group distributions on that plane. A p-value > 0.05 suggests no statistically significant difference. (a) PC1 vs PC2: This plot shows a strong overlap between both groups. However, the p-value of 0.042, just below the 0.05 threshold, indicates a subtle but statistically significant difference in the distributions of the two groups. (b) PC1 vs PC3 and (c) PC2 vs PC3: These plots demonstrate a high degree of overlap between the groups. With p-values greater than 0.05, there is no statistically significant difference between the distributions of the real and simulated groups in these planes. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8099876","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":547740250,"identity":"621b1311-0837-40fc-a4d8-cb69c9c2d062","order_by":0,"name":"Juan Alberto Ríos-Rodríguez","email":"","orcid":"","institution":"Tecnologico de Monterrey","correspondingAuthor":false,"prefix":"","firstName":"Juan","middleName":"Alberto","lastName":"Ríos-Rodríguez","suffix":""},{"id":547740251,"identity":"099db987-0288-435b-8bbe-9a5e3a591c58","order_by":1,"name":"Sylvia Harari-Arakindji","email":"","orcid":"","institution":"Universidad Anáhuac","correspondingAuthor":false,"prefix":"","firstName":"Sylvia","middleName":"","lastName":"Harari-Arakindji","suffix":""},{"id":547740252,"identity":"0a641e01-45be-4bea-b6e9-65fe714b1569","order_by":2,"name":"Berenice Cuevas-Estrada","email":"","orcid":"","institution":"Instituto Nacional de Cancerología","correspondingAuthor":false,"prefix":"","firstName":"Berenice","middleName":"","lastName":"Cuevas-Estrada","suffix":""},{"id":547740253,"identity":"056c32a2-70b4-4da3-bfd5-18d9e68e67b6","order_by":3,"name":"José Antonio García-Pacheco","email":"","orcid":"","institution":"Instituto Nacional de Cancerología","correspondingAuthor":false,"prefix":"","firstName":"José","middleName":"Antonio","lastName":"García-Pacheco","suffix":""},{"id":547740254,"identity":"e8167a5c-b2ea-4600-8d60-112e82e60376","order_by":4,"name":"Patricia Ostrosky-Wegman","email":"","orcid":"","institution":"Universidad Nacional Autónoma de México","correspondingAuthor":false,"prefix":"","firstName":"Patricia","middleName":"","lastName":"Ostrosky-Wegman","suffix":""},{"id":547740255,"identity":"790aa444-bbd7-4a9e-8fba-5a151c014651","order_by":5,"name":"Nora Sobrevilla-Moreno","email":"","orcid":"","institution":"Instituto Nacional de Cancerología","correspondingAuthor":false,"prefix":"","firstName":"Nora","middleName":"","lastName":"Sobrevilla-Moreno","suffix":""},{"id":547740256,"identity":"dea0040f-1a4e-4c3e-951b-41c63c4e2904","order_by":6,"name":"Miguel A Jiménez-Ríos","email":"","orcid":"","institution":"Instituto Nacional de Cancerología","correspondingAuthor":false,"prefix":"","firstName":"Miguel","middleName":"A","lastName":"Jiménez-Ríos","suffix":""},{"id":547740257,"identity":"6534fd17-b8b0-452a-ad13-7ae83f4330fe","order_by":7,"name":"Alejandro Mohar-Betancourt","email":"","orcid":"","institution":"Universidad Nacional Autónoma de México","correspondingAuthor":false,"prefix":"","firstName":"Alejandro","middleName":"","lastName":"Mohar-Betancourt","suffix":""},{"id":547740258,"identity":"8ed9e705-1294-44dc-ab9f-f1b7909f2b04","order_by":8,"name":"Rodrigo González‐Barrios","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABAUlEQVRIiWNgGAWjYHACA4YEIMnPwNgApCRI0CLZQJIWMHmAWFeZtzdv+/Dgl02e8fHDjY8Laizk5Rt4zB78zGGIxmWIzJljxTMS+9KKzc4kNhvPOCZhuOEAj7lh7zaG3JkN2LVISOQYMyT2HE7cdoOxTZq3QYJxAwOPmQQvUEs/DofBtWyewdj+G6jFfj7QYZJ/gVra8GlJ+HE4cYMEYxszUEtiwwEeM2m8tvAcK2ZIbEgrlgD6RZrnmETyhsNsZdKy2yRw+4W9eTPjjz82efztxx9+5qmps50PDEPJt9tscjfgCDEwYGwDxyYUMEPMwqMeBP4gaxkFo2AUjIJRgAYA8mRXNaFPyL4AAAAASUVORK5CYII=","orcid":"","institution":"Instituto Nacional de Cancerología","correspondingAuthor":true,"prefix":"","firstName":"Rodrigo","middleName":"","lastName":"González‐Barrios","suffix":""},{"id":547740259,"identity":"f75b0fea-79c1-41b7-ad2f-968e66ca809a","order_by":9,"name":"Talia Wegman-Ostrosky","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAvElEQVRIiWNgGAWjYFAC5oYDDAw2QAZj4wGiNPAwMIK0pIG0NBCvBUgdBnOI02LPfrDxwM+283Zr2w8DbamxiSZsC09iw8HettvJ284kArUcS8ttIOwwoErebbeTzQ4AGYwNh4nQwv+w4eDfbeeSzc4/JFaLRGLDYd5tB+zMbhBty42HDYdl/yUnmAEZBxKI8Qt7f/Lhj2/O2NmbnU9/+OBDjQ1hLTCQCFaZQKxyELAnRfEoGAWjYBSMMAAApwhMNIhVqHUAAAAASUVORK5CYII=","orcid":"","institution":"Instituto Nacional de Cancerología","correspondingAuthor":true,"prefix":"","firstName":"Talia","middleName":"","lastName":"Wegman-Ostrosky","suffix":""}],"badges":[],"createdAt":"2025-11-12 22:18:40","currentVersionCode":1,"declarations":{"humanSubjects":true,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":true,"humanSubjectConsent":true,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-8099876/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8099876/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":96566000,"identity":"7361a0fa-e83c-40b5-8887-78ff8afe9a67","added_by":"auto","created_at":"2025-11-23 15:50:21","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":105389,"visible":true,"origin":"","legend":"","description":"","filename":"AIMC4RManuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/cbc18355a962e1daf08501ea.docx"},{"id":96604982,"identity":"e283cc40-f198-45b6-ab97-f99330a2ee0c","added_by":"auto","created_at":"2025-11-24 09:17:04","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":342,"visible":true,"origin":"","legend":"","description":"","filename":"rs8099876.json","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/95657d31fcc6de741af282b7.json"},{"id":96566006,"identity":"7c83c2b9-1c00-418e-a2c6-f627447d8181","added_by":"auto","created_at":"2025-11-23 15:50:21","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":91025,"visible":true,"origin":"","legend":"","description":"","filename":"rs80998760enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/00c3bad8e21b9548e07d456d.xml"},{"id":96566008,"identity":"c03588e9-b021-454c-9fbf-60ee4d0f1726","added_by":"auto","created_at":"2025-11-23 15:50:21","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":89583,"visible":true,"origin":"","legend":"","description":"","filename":"rs80998760structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/a71ba6130d56195948a487d3.xml"},{"id":96604939,"identity":"913f78e0-234c-4dfe-b2eb-f4e0344a21b6","added_by":"auto","created_at":"2025-11-24 09:16:26","extension":"html","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":99551,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/677813360bb05d11a7f977e1.html"},{"id":96566003,"identity":"a4ed1eb4-5c78-41ed-8225-e20180e37259","added_by":"auto","created_at":"2025-11-23 15:50:21","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":241656,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAllelic Discrimination and Amplification Curves for the Detection of the Variant in TCGT Patients.\u003c/strong\u003e (a) Allelic discrimination plot showing the distribution of patients carrying the variant in a heterozygous state (green), non-carriers (blue), and three undetermined samples (black). No individuals were identified as homozygous for the variant; therefore, appropriate adjustments were made to the analysis. (b) Amplification curve of a heterozygous patient for the variant (T/A), highlighting the dual amplification signal in the presence of the T allele. (c) Amplification curve of a non-carrier patient (AA), demonstrating amplification of a single line corresponding to the absence of the T allele. The analysis was performed using PCR, where amplification of both lines indicated the presence of the T allele, while amplification of a single line confirmed the absence of the variant.\u003c/p\u003e","description":"","filename":"Figure1..png","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/d51313b097998ec3aa4081b2.png"},{"id":96566002,"identity":"c39b8360-d634-445d-8b46-51eb70b7e9f7","added_by":"auto","created_at":"2025-11-23 15:50:21","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":432581,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eKaplan-Meier survival analysis comparing carriers and non-carriers in real and synthetic cohorts\u003c/strong\u003e. a) Real Cohort: Survival curves for carriers (red) and non-carriers (blue) with 95% confidence intervals. The p-value of 0.1 indicates no significant difference in survival. (b) Synthetic Cohort: Survival curves for carriers (red) and non-carriers (blue) with 95% confidence intervals. This AI-enhanced, computationally generated cohort reveals a statistically significant lower survival probability for carriers (p \u0026lt; 0.0001), providing a visually clearer representation of the survival differences.\u003c/p\u003e","description":"","filename":"Figure2..png","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/9e926acfb5910d383ff99bf9.png"},{"id":97248402,"identity":"dbb93ea6-0401-4e7f-b40a-3359fceee05e","added_by":"auto","created_at":"2025-12-02 12:57:15","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1387011,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/383360c2-9cf6-47a6-84fc-9eb6bce857cd.pdf"},{"id":96604987,"identity":"3520ab9f-4a03-4a81-975a-cac1fc8a6c98","added_by":"auto","created_at":"2025-11-24 09:17:04","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":68733,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupplementary figure 1. Variance inflation factors and Spearman correlation matrices of real and synthetic carrier data\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis figure displays the VIFs for the variables Surveillance time, BMI, and Age in the real carrier dataset, indicating minimal multicollinearity (a). It also presents heatmap matrices showing the Spearman correlation coefficients for the real carrier data (b) and the synthetic carrier data (c). The structural similarity between the correlation matrices was confirmed by a permutation test, which resulted in a non-significant p-value (p \u0026gt; 0.05), indicating that the synthetic dataset effectively preserves the dependency structures of the original data.\u003c/p\u003e","description":"","filename":"Supplementaryfigure1.png","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/d6be85b8439affeeb6e80185.png"},{"id":96566001,"identity":"7c568e64-b305-42f3-ab8e-31e293624a9d","added_by":"auto","created_at":"2025-11-23 15:50:21","extension":"png","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":90083,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupplementary figure 2. PCA of real and simulated carriers\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePrincipal Component Analysis (PCA) comparing the \"Real Carriers Group\" (blue) with the \"Simulated Carriers Group\" (red). Both groups are complemented with the same set of real patient controls. The two-dimensional scatter plots show the projections of the data points onto the planes formed by the first three principal components (PC1, PC2, and PC3), which explain 39.8%, 31.4%, and 28.8% of the total variance, respectively. The p-values at the top of each graph indicate the statistical significance of the difference between the two group distributions on that plane. A p-value \u0026gt; 0.05 suggests no statistically significant difference. (a) PC1 vs PC2: This plot shows a strong overlap between both groups. However, the p-value of 0.042, just below the 0.05 threshold, indicates a subtle but statistically significant difference in the distributions of the two groups. (b) PC1 vs PC3 and (c) PC2 vs PC3: These plots demonstrate a high degree of overlap between the groups. With p-values greater than 0.05, there is no statistically significant difference between the distributions of the real and simulated groups in these planes.\u003c/p\u003e","description":"","filename":"Supplementaryfigure2.png","url":"https://assets-eu.researchsquare.com/files/rs-8099876/v1/201d7404c1daa09e39bee163.png"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eAI-Driven Synthetic Cohorts to Explore Genetic Associations: Lessons from Testicular Cancer with Relevance to Rare Conditions\u003c/p\u003e","fulltext":[{"header":"Background and Significance","content":"\u003cp\u003eHandling small sample sizes in statistical analysis is a significant challenge, as limited data often lacks sufficient power to yield meaningful conclusions (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). This limitation is particularly pronounced in the study of rare diseases, generally defined as conditions affecting approximately 1 in every 2,000 to 2,500 individuals (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). In oncology, about one in five cancer patients has a rare malignancy, posing substantial challenges for epidemiological research. Methodological adjustments are therefore necessary to reliably detect statistically significant differences within these small patient groups (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eTesticular cancer (TCa) exemplifies this challenge, being a rare malignancy worldwide with an age-standardized incidence rate (ASIR) of 1.8 and a mortality rate (ASMR) of 0.2 per 100,000 men, according to GLOBOCAN 2020 (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). However, these rates vary notably across regions. In Mexico, the ASIR is 5.1 per 100,000 men (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e), considerably above the global average, and the ASMR is 0.94 per 100,000 men between 2015 and 2019, ranking third highest in Latin America (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eApproximately 95\u0026ndash;98% of testicular tumors are germ cell tumors (TGCT), which are typically highly curable, with 5-year overall survival rates exceeding 80%. Despite this favorable prognosis, TCa remains a public health concern in Mexico, partly due to limited research on TGCT within the population (\u003cspan additionalcitationids=\"CR6\" citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). Notably, valuable insights have been provided into both the somatic genetic alterations associated with chemoresistance and the social determinants that influence treatment outcomes. (\u003cspan additionalcitationids=\"CR6\" citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eGermline genetic variants could potentially explain both the increased incidence and mortality of TCa. To date, no high-penetrance variant has been directly associated with risk, and the low frequency of the disease makes variant analysis challenging(\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). This is particularly evident for genes like \u003cem\u003eMC4R\u003c/em\u003e, known for its role in autosomal dominant obesity (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e) and is expressed in fetal germ cell tests (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e). Missense variants in \u003cem\u003eMC4R\u003c/em\u003e have been shown to promote teratoma formation in knock-in mice, potentially through modulation of anti-apoptotic signaling, which may favor transdifferentiation into somatic cells or increased parthenogenesis (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e). Large-scale data analyses have also suggested associations between Single Nucleotide Polymorphisms (SNP) near \u003cem\u003eMC4R\u003c/em\u003e and other cancers, including endometrial and breast cancer, underscoring the need to investigate its potential role in gonadal cancer. (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e)\u003c/p\u003e\u003cp\u003eGiven the challenges of studying rare diseases, advanced computational methods have gained importance. Synthetic data generation offers a versatile solution by using techniques like Classification and Regression Trees (CARTs) to accurately replicate key oncology endpoints, including overall and progression-free survival. (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). The synthpop package is a powerful non-parametric tool that sequentially generates synthetic datasets preserving the statistical properties and correlations of the original data, without containing real individual records (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e). Comparative analyses highlight its superior performance over other methods, including deep learning, particularly for small datasets (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e) This capability is crucial as it allows researchers to conduct analyses and obtain results consistent with those from the real data (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eBuilding on these findings, this study evaluates CART-based synthetic data generation using synthpop as an artificial intelligence (AI) approach to strengthen epidemiological analysis. As a proof of concept, it focuses on TCa \u0026mdash;a rare disease\u0026mdash; and investigates the \u003cem\u003eMC4R rs79783591\u003c/em\u003e genetic variant, which could be a critical factor in the Mexican population, in relation to disease risk and prognosis, addressing the challenges posed by a small patient cohort.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cp\u003eThe manuscript was prepared following the STROBE guidelines for cohort studies and MI-CLAIM. A representative cohort of Mexican patients with TGCT, treated between 2007 and 2020 at the Instituto Nacional de Cancerolog\u0026iacute;a (INCan), a national referral center in Mexico City, was included after prior informed consent (022/068/OMI; CEI/052/22). The sample size was determined for exploratory purposes based on previously reported prevalence in an external cohort(\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). Patients\u0026thinsp;\u0026ge;\u0026thinsp;15 years with histologically confirmed TGCT, complete clinical records, and at least 5 years of follow-up were eligible, while those with uncertain diagnoses or incomplete records were excluded. Initial exome sequencing was performed on 40 patients from an external cohort previously published (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e), which were reanalyzed to detect variants in the MC4R gene. Genotyping of the rs79783591 variant was then extended to the entire INCan cohort using real-time polymerase chain reaction (qPCR) with TaqMan probes. Demographic and clinical information was collected from medical records, and logistic regression was performed to estimate the association between genetic variants and TGCT risk, calculating Odds Ratios (OR) with 95% confidence intervals.\u003c/p\u003e\u003cp\u003eTo address the limited sample size, synthetic data were generated using the synthpop package in R. Datasets were created from the subpopulation of mutation carriers using a CART model, preserving the distributions, correlations, and multivariate structure of the original data. Multiple synthetic datasets were generated across simulations, with seeds set sequentially according to the number of groups, aiming to match the size of the non-carrier population.\u003c/p\u003e\u003cp\u003eValidation of the synthetic datasets included statistical comparisons with the original data (Kolmogorov-Smirnov, t-test or Mann-Whitney U, Chi-square or Fisher\u0026rsquo;s exact tests with Benjamini-Hochberg adjustment), assessment of structural similarity (Spearman correlation matrices, principal component analysis with permutation testing). Application of machine learning models (Random Forest, XGBoost, SVM, and logistic regression) was performed to confirm that real and synthetic data could not be distinguished. Hyperparameters were optimized by grid search where applicable, and model performance was evaluated using an 80/20 train-test split with 5-fold cross-validation. Performance metrics included area under the curve (AUC) with confidence intervals, sensitivity, specificity, F1 score, and Cohen\u0026rsquo;s kappa.\u003c/p\u003e\u003cp\u003eSummary statistics are reported as mean\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation or median with 25th and 75th percentiles for continuous variables, and as percentages for categorical variables. Qualitative variables were compared using Chi-square or Fisher\u0026rsquo;s exact tests, while quantitative variables were assessed for normality (Shapiro-Wilk) and compared with Student\u0026rsquo;s t-test or Mann-Whitney U test as appropriate. Survival analyses employed Kaplan-Meier curves, with group comparisons assessed using log-rank, Peto-Peto, and Tarone-Ware tests. Variance of key clinical variables (BMI and follow-up time) was compared using F-tests. Cox regression models assessed the variant\u0026rsquo;s effect on mortality risk, with standard errors of coefficients and C-index calculated to evaluate precision and discriminative ability; a synthetic cohort was employed to increase statistical power. These additional tests were included to account for potential differences in early- and late-event survival between carriers and non-carriers. Statistical significance was set at p\u0026thinsp;\u0026lt;\u0026thinsp;0.05. All analyses were conducted in RStudio (version 2024.12.0).\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eBased on a prevalence of 0.125 from the previously published external cohort, a finite population of 2,700, a 5% margin of error, and 97% confidence, the initial sample size was 244, with 24 excluded by eligibility criteria, leaving 220 TGCT patients. Germline DNA analysis identified 9 (4.1%) heterozygous carriers of the rs79783591 variant in the MC4R gene (Fig.\u0026nbsp;1). Sixteen patients initially reviewed were excluded for not meeting the eligibility criteria. The carrier frequency was compared to 0.0062 in the general Mexican population, as reported by the Population Architecture using Genomics and Epidemiology (PAGE) Study database via ClinVar (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). Logistic regression showed a statistically significant association, with an Odds Ratio (OR) of 3.42 (95% CI: 1.07\u0026ndash;8.5; p\u0026thinsp;=\u0026thinsp;0.02), suggesting a potential link between the variant and increased TGCT risk.\u003c/p\u003e\u003cp\u003eA bivariate analysis was performed on the real cohort to compare the clinical characteristics between patients who were carriers and non-carriers of the \u003cem\u003eMC4R\u003c/em\u003e genetic variant (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The analysis revealed a statistically significant difference in age at diagnosis. Carriers were diagnosed at a median age of 22 years, which was notably younger than the non-carriers (median of 26 years; p\u0026thinsp;=\u0026thinsp;0.028). No significant differences were observed in Body Mass Index (BMI) or the presence of distant metastases. Regarding tumor histology, non-seminoma was the most common subtype. Furthermore, the study found a higher mortality rate among carriers (33.33%) compared to non-carriers (14.7%). However, this association did not reach statistical significance (p\u0026thinsp;=\u0026thinsp;0.147).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eClinical Characteristics Comparison Between \u003cem\u003eMC4R\u003c/em\u003e Variant Carriers and Non-Carriers in the Cohort\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eVariable\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTotal\u003c/p\u003e\u003cp\u003eN\u0026thinsp;=\u0026thinsp;220 (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNon-carriers\u003c/p\u003e\u003cp\u003en\u0026thinsp;=\u0026thinsp;211 (95.9%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eCarriers\u003c/p\u003e\u003cp\u003en\u0026thinsp;=\u0026thinsp;9 (4.1%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003ep-value\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003cp\u003eMedian (p25-p75)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e26 (22\u0026ndash;31)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e26 (\u003cspan additionalcitationids=\"CR24 CR25\" citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e22 (\u003cspan additionalcitationids=\"CR22 CR23 CR24\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.028*\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBMI\u003c/p\u003e\u003cp\u003eMedian (p25-p75)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e24.3 (22.7\u0026ndash;27.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e24.3 (22.7\u0026ndash;27.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e24.2 (21.2\u0026ndash;26.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.348\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHistology\u003c/p\u003e\u003cp\u003eseminoma\u003c/p\u003e\u003cp\u003eNon-seminoma\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e84 (38.2)\u003c/p\u003e\u003cp\u003e136 (61.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e82 (38.9)\u003c/p\u003e\u003cp\u003e136 (61.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2 (22.22)\u003c/p\u003e\u003cp\u003e7 (77.78)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.488\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDistant metastasis\u003c/p\u003e\u003cp\u003eAbsent\u003c/p\u003e\u003cp\u003ePresent\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e144 (65.5)\u003c/p\u003e\u003cp\u003e76 (34.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e137 (64.9)\u003c/p\u003e\u003cp\u003e74 (35.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e7 (77.78)\u003c/p\u003e\u003cp\u003e2 (22.22)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.722\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eVital status\u003c/p\u003e\u003cp\u003eAlive\u003c/p\u003e\u003cp\u003eDeceased\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e186 (84.5)\u003c/p\u003e\u003cp\u003e34 (15.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e180 (85.3)\u003c/p\u003e\u003cp\u003e31 (14.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e6 (66.67)\u003c/p\u003e\u003cp\u003e3 (33.33)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.147\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"5\"\u003eBMI: Body Mass Index; \u003csup\u003ea\u003c/sup\u003eMann-Whitney U test, \u003csup\u003eb\u003c/sup\u003eChi-square test, significance at p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05*\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe synthetic carrier cohort was generated to preserve the key clinical characteristics observed in the original dataset, consisting of 234 patients created across 26 sets of 9 individuals each, which was confirmed through a series of rigorous validation tests (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Statistical comparisons showed strong consistency between the synthetic and real data, with no significant differences in most variables. Furthermore, no observable differences in the Variance Inflation Factor (VIF) were found among the continuous variables, suggesting the absence of multicollinearity in the synthetic data.\u003c/p\u003e\u003cp\u003eTo evaluate structural similarity, Spearman correlation matrices were generated for both datasets. A permutation test confirmed that the overall correlation structures were not significantly different (p\u0026thinsp;\u0026gt;\u0026thinsp;0.05). However, a subtle inconsistency was noted for the Age variable, where the correlation with Surveillance time had a reversed sign (real: -0.0183, synthetic: 0.0317). To ensure the reliability of the results, the Age variable was therefore omitted from subsequent analyses (Supplementary Fig.\u0026nbsp;1).\u003c/p\u003e\u003cp\u003ePrincipal Component Analysis (PCA) was also used to assess structural similarity. The PCA did not show significant differences when comparing PC2 vs. PC3 and PC1 vs. PC3. However, when comparing PC1 vs. PC2, a permutation test was found to be significant despite a clear visual overlap (Supplementary Fig.\u0026nbsp;2). Finally, machine learning models (Logistic Regression, Random Forest, XGBoost, and SVM) were employed to test their ability to distinguish between real and synthetic data. The models' performance was poor across the board, with the highest AUC being 0.674 for Random Forest and a logistic regression AUC of 0.553 (Supplementary Table\u0026nbsp;1).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eComparison of clinical characteristics between the synthetic carrier ohort and the real carrier cohort\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eVariable\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSynthetic Carrier\u003c/p\u003e\u003cp\u003en\u0026thinsp;=\u0026thinsp;234 (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal\u003c/p\u003e\u003cp\u003eCarriers\u003c/p\u003e\u003cp\u003en\u0026thinsp;=\u0026thinsp;9 (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003ep-value\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003cp\u003eMedian (p25-p75)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e22 (\u003cspan additionalcitationids=\"CR22 CR23 CR24\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e22 (\u003cspan additionalcitationids=\"CR22 CR23 CR24\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.88\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBMI\u003c/p\u003e\u003cp\u003eMedian (p25-p75)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e24.2 (21.2\u0026ndash;26.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e24.2 (21.2\u0026ndash;26.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.96\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHistology\u003c/p\u003e\u003cp\u003eSeminoma\u003c/p\u003e\u003cp\u003eNon-seminoma\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e62 (26.5)\u003c/p\u003e\u003cp\u003e172 (73.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2 (22.22)\u003c/p\u003e\u003cp\u003e7 (77.78)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDistant Metastasis\u003c/p\u003e\u003cp\u003eAbsent\u003c/p\u003e\u003cp\u003ePresent\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e185 (79.1)\u003c/p\u003e\u003cp\u003e49 (20.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e7 (77.78)\u003c/p\u003e\u003cp\u003e2 (22.22)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eVital Status\u003c/p\u003e\u003cp\u003eAlive\u003c/p\u003e\u003cp\u003eDeceased\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e144 (61.5)\u003c/p\u003e\u003cp\u003e90 (38.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e6 (66.7)\u003c/p\u003e\u003cp\u003e3 (33.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"4\"\u003eBMI: Body Mass Index; \u003csup\u003ea\u003c/sup\u003eChi-square test \u003csup\u003e, b\u003c/sup\u003eMann-Whitney U test, significance at p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05*\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe Kaplan-Meier survival analysis with a log-rank test was performed to assess time to mortality in both real and synthetic carrier cohorts (Fig.\u0026nbsp;2). In the real cohort, the survival curve for carriers visually separated below that of the non-carriers. In the real cohort, survival curve separation was not significant (log-rank, Peto-Peto, Tarone-Ware p\u0026thinsp;\u0026asymp;\u0026thinsp;0.09\u0026ndash;0.12), likely due to few carriers (N\u0026thinsp;=\u0026thinsp;9). In the synthetic cohort, separation remained but was highly significant (p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001). This cohort also exhibited notably narrower and more stable confidence intervals around the carriers' survival curve.\u003c/p\u003e\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e illustrates the results of the multivariate analysis using both logistic regression (OR) and Cox regression (HR) models to identify factors associated with mortality in the real and synthetic \u003cem\u003eMC4R\u003c/em\u003e cohorts.\u003c/p\u003e\u003cp\u003eIn both cohorts, BMI and histology did not show a significant association with mortality in either model (p\u0026thinsp;\u0026gt;\u0026thinsp;0.05). In contrast, the presence of distant metastasis was a highly significant predictor of mortality in both the real cohort (OR\u0026thinsp;=\u0026thinsp;44.4, 95% CI: 11.7\u0026ndash;296; HR\u0026thinsp;=\u0026thinsp;22.9, 95% CI: 6.50\u0026ndash;80.5) and the synthetic cohort (OR\u0026thinsp;=\u0026thinsp;4.37, 95% CI: 2.62\u0026ndash;7.45; HR\u0026thinsp;=\u0026thinsp;2.92, 95% CI: 2.01\u0026ndash;4.24), with p-values consistently below 0.001 for all four models.\u003c/p\u003e\u003cp\u003eThe presence of the MC4R variant was also a significant predictor of mortality. In the real cohort, it showed a significant association in both the logistic regression (OR\u0026thinsp;=\u0026thinsp;15.1, 95% CI: 1.65\u0026ndash;158, p\u0026thinsp;=\u0026thinsp;0.016) and the Cox regression models (HR\u0026thinsp;=\u0026thinsp;5.69, 95% CI: 1.56\u0026ndash;20.7, p\u0026thinsp;=\u0026thinsp;0.008). These findings were confirmed and strengthened in the synthetic cohort, where the variant was a highly significant predictor in both the logistic regression (OR\u0026thinsp;=\u0026thinsp;5.11, 95% CI: 3.06\u0026ndash;8.84, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) and Cox regression models (HR\u0026thinsp;=\u0026thinsp;3.15, 95% CI: 2.06\u0026ndash;4.82, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Notably, the synthetic models yielded significantly narrower and more precise confidence intervals for both the OR and HR, with lower standard errors for all Cox coefficients (Carrier: 0.614 \u0026rarr; 0.216; BMI: 0.049 \u0026rarr; 0.029; Histology: 0.610 \u0026rarr; 0.211) and a minor reduction in C-index (0.675 \u0026rarr; 0.636), providing a more stable estimate of the variant's effect.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eRegression models of mortality in real and synthetic \u003cem\u003eMC4R\u003c/em\u003e Cohorts\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eVariable\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eReal OR (IC)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003ep-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSimulated OR (IC)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003ep-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eReal HR (IC)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003ep-value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eSimulated HR (IC)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003ep-value\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBMI\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e1.08 (0.94, 1.23)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.272\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.97 (0.90, 1.05)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.487\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e1.07 (0.96, 1.18)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.214\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e\u003cp\u003e0.99 (0.93, 1.05)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e\u003cp\u003e0.768\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHistology\u003c/p\u003e\u003cp\u003eSeminoma\u003c/p\u003e\u003cp\u003eNon-seminoma\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e3.24 (0.94, 15.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.0874\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e1.10 (0.66, 1.86)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.714\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e2.56 (0.76, 8.69)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.131\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e1.17 (0.78, 1.78)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.449\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDistant Metastasis\u003c/p\u003e\u003cp\u003eAbsent\u003c/p\u003e\u003cp\u003ePresent\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e44.4 (11.7, 296)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e4.37 (2.62, 7.45)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e22.9 (6.50, 80.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e2.92 (2.01, 4.24)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMutation Carrier\u003c/p\u003e\u003cp\u003eAbsent\u003c/p\u003e\u003cp\u003ePresent\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e15.1 (1.65, 158)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.016\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e5.11 (3.06, 8.84)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e5.69 (1.56, 20.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.008\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eRef.\u003c/p\u003e\u003cp\u003e3.15 (2.06, 4.82)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"9\"\u003eBMI: Body Mass Index; significance at p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05*\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study provides the first evidence of an association between the \u003cem\u003eMC4R rs79783591\u003c/em\u003e variant and cancer, suggesting a potential role as a risk and prognostic factor in TGCT. Carriers showed an increased mortality risk, adjusted for BMI, disease location, and histology, and were diagnosed at a median age four years earlier than non-carriers. These findings are consistent with observations in Hispanic populations, who often present with more aggressive and chemoresistant disease at younger ages; these patterns do not appear to be fully explained by sociodemographic factors (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). Although statistical significance was limited by sample size, simulated cases confirmed a threefold increase in mortality.\u003c/p\u003e\u003cp\u003eThe variant (rs79783591, c.806T\u0026thinsp;\u0026gt;\u0026thinsp;A; p.Ile269Asn) lies in \u003cem\u003eMC4R\u003c/em\u003e (18q21.32), with a gnomAD frequency of 0.00011 and 0.00629 in the Mexican PAGE cohort. Despite its conflicting ClinVar classification (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e), the variant has been associated with obesity in Mexican children and adults (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e); however, carriers here exhibited relatively normal BMI values. TGCT has high heritability, but no high-penetrance germline variants or validated prognostic genetic markers have been established (\u003cspan additionalcitationids=\"CR22\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e). Functional hypotheses suggest that MC4R variants may influence germ cell survival via anti-apoptotic pathways, potentially disrupting the balance between proliferation and apoptosis (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e). Confirmation of these associations and clarification of underlying mechanisms will require larger, multi-center studies. Additionally, we were unable to calculate the haplotype as the presence of another potentially associated variant could not be excluded.\u003c/p\u003e\u003cp\u003eTo address the limitations of small sample size, we implemented synthetic data generated with synthpop, which improved model performance, narrowed confidence intervals, and enabled detection of significant Kaplan\u0026ndash;Meier differences that were only marginal in the real dataset, while preserving significance in multivariate analyses. This approach preserves the trends of the original cohort while enhancing statistical power and has been shown to outperform other methods, such as deep learning-based Generative Adversarial Networks (GANs), particularly in small, tabular clinical datasets (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan additionalcitationids=\"CR25\" citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e). In fact, given the extremely limited number of carriers in our study (n\u0026thinsp;=\u0026thinsp;9), the use of GANs would be impractical due to overfitting and instability in training, further supporting the choice of CART-based synthetic data generation as a more robust and transparent alternative. While synthetic data provide valuable insights and help identify potential prognostic factors, they do not perfectly reflect real populations and may introduce bias if original data variability is limited.\u003c/p\u003e\u003cp\u003eComplementary tests sensitive to early- and late-event differences (Peto-Peto and Tarone-Ware) confirmed trends in the real cohort and detected significant differences in the synthetic cohort, supporting the utility of synthetic data to increase statistical power.\u003c/p\u003e\u003cp\u003eWhile synthetic data provide valuable insights, limited variability in the original cohort may introduce bias, potentially distorting observed clinical patterns. As shown in our Kaplan\u0026ndash;Meier analysis, a \u0026ldquo;fall\u0026rdquo; was observed due to such variability, which might not reflect true clinical behavior. Permutation analyses of PCA components further highlighted that some relationships between variables may be incompletely captured. These limitations emphasize that even within small cohorts, increasing the number of cases and its diversity can improve robustness and reduce potential distortions. Importantly, the synthetic cohort produced lower standard errors and slightly reduced C-index values compared to the real cohort, indicating more precise and stable estimates of the variant\u0026rsquo;s effect while preserving discriminative ability.\u003c/p\u003e\u003cp\u003eDespite the small cohort size, INCan is a national referral center, and its cases may reasonably reflect TGCT patients across Mexico (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e). Compared to more complex methods such as GANs, synthpop remains a transparent and accessible approach that preserves essential statistical properties, enables a novel dataset expansion without discarding variables based on arbitrary thresholds, and facilitates the identification of potential prognostic markers that might otherwise be overlooked.\u003c/p\u003e\u003cp\u003eAlthough synthpop provides a valuable approach for increasing statistical power in small cohorts, its limitations should be acknowledged. The limited number of carriers in the original cohort may restrict detection of statistically significant associations, and synthetic data do not perfectly reflect a real population; rather, they model what would occur in a population with similar characteristics to the observed cases. Nonetheless, synthpop is an accessible, transparent, and easily applicable method that preserves essential statistical properties, allows dataset expansion without discarding variables based on p-value thresholds, and facilitates identification of potential biological risk factors. Additionally, the synthetic data benefit from the package\u0026rsquo;s built-in prospective design, which anonymizes patient information, improving ethical considerations and compliance.\u003c/p\u003e\u003cp\u003eImportantly, integrating synthetic data generation with machine learning validation represents a novel framework for augmenting small cohorts, including in TCa, and is generalizable to other rare conditions where limited sample sizes hinder robust statistical inference. Despite the strengthened findings provided by synthetic data, further research in slightly expanded multi-center cohorts is needed to confirm these results, explore additional genetic variants, and assess the robustness of this approach, ultimately supporting the prioritization of candidate biomarkers and clinical factors for translational applications.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThe \u003cem\u003eMC4R rs79783591\u003c/em\u003e variant may represent both a risk and prognostic factor for TGCT in Mexican patients. This study provides preliminary evidence of this association, addressing the scarce knowledge of genetic susceptibility to this cancer type in Hispanic populations. Importantly, this research illustrates the utility of artificial intelligence\u0026ndash;driven approaches to mitigate limitations posed by small sample sizes, a frequent obstacle in rare disease studies. By generating synthetic data through a repurposing of synthpop, the analysis gained statistical power, uncovering associations that conventional methods might have overlooked.\u003c/p\u003e\u003cp\u003eFuture studies with more diverse cohorts are needed to validate these findings and investigate additional genetic variants. Despite its limitations, the use of AI tools like synthpop opens promising avenues for personalized medicine and research in data-limited populations.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eCRediT authorship contribution statement:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptualization: Juan Alberto R\u0026iacute;os-Rodr\u0026iacute;guez, Berenice Cuevas-Estrada, Sylvia Harari-Arakindji, Rodrigo Gonz\u0026aacute;lez‐Barrios, Talia Wegman-Ostrosky; Data curation: Berenice Cuevas-Estrada, Juan Alberto R\u0026iacute;os-Rodr\u0026iacute;guez, Jos\u0026eacute; Antonio Garc\u0026iacute;a-Pacheco, Sylvia Harari-Arakindji: Formal analysis: Juan Alberto R\u0026iacute;os-Rodr\u0026iacute;guez, Berenice Cuevas-Estrada, Jos\u0026eacute; Antonio Garc\u0026iacute;a-Pacheco; Investigation: Berenice Cuevas-Estrada, Juan Alberto R\u0026iacute;os-Rodr\u0026iacute;guez, Jos\u0026eacute; Antonio Garc\u0026iacute;a-Pacheco, Sylvia Harari-Arakindji: Methodology: Juan Alberto R\u0026iacute;os-Rodr\u0026iacute;guez, Berenice Cuevas-Estrada; Project administration: Rodrigo Gonz\u0026aacute;lez‐Barrios, Talia Wegman-Ostrosky; Resources: Rodrigo Gonz\u0026aacute;lez‐Barrios, Talia Wegman-Ostrosky, Nora Sobrevilla-Moreno, Miguel A Jim\u0026eacute;nez-R\u0026iacute;os; Supervision: Rodrigo Gonz\u0026aacute;lez‐Barrios, Talia Wegman-Ostrosky; Visualization: Juan Alberto R\u0026iacute;os-Rodr\u0026iacute;guez, Berenice Cuevas-Estrada, Jos\u0026eacute; Antonio Garc\u0026iacute;a-Pacheco; Writing \u0026ndash; original draft: Juan Alberto R\u0026iacute;os-Rodr\u0026iacute;guez, Sylvia Harari-Arakindji, Talia Wegman-Ostrosky, Jos\u0026eacute; Antonio Garc\u0026iacute;a-Pacheco, Berenice Cuevas-Estrada; Writing \u0026ndash; review \u0026amp; editing: Patricia Ostrosky-Wegman, Alejandro Mohar-Betancourt, Rodrigo Gonz\u0026aacute;lez‐Barrios, Berenice Cuevas-Estrada, Juan Alberto R\u0026iacute;os-Rodr\u0026iacute;guez, Talia Wegman-Ostrosky, Jos\u0026eacute; Antonio Garc\u0026iacute;a-Pacheco, Sylvia Harari-Arakindji\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDeclaration of Competing Interest\u003c/strong\u003e\u003c/p\u003e\n\u003ch5\u003eThe authors declare no conflicts of interest.\u003c/h5\u003e\n\u003cp\u003e\u003cstrong\u003eFunding statement\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by Instituto Nacional de Cancerolog\u0026iacute;a, Consejo Nacional de Humanidades, Ciencias y Tecnolog\u0026iacute;as (RG-B, 2017-2-290041) and Financiamiento de Proyectos de Investigaci\u0026oacute;n para la Salud (FPIS-INCAN-4352) from the Direcci\u0026oacute;n General de Pol\u0026iacute;ticas de Investigaci\u0026oacute;n en Salud (DGPIS), Secretar\u0026iacute;a de Salud, M\u0026eacute;xico.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eJ.A.R.-R acknowledges the Master\u0026rsquo;s in Biomedical Sciences (MBC) program by the School of Medicine and Health Sciences, Tecnologico de Monterrey, and was supported by Secretar\u0026iacute;a de Ciencia, Humanidades, Tecnolog\u0026iacute;a e Innovaci\u0026oacute;n (SECIHTI) fellowship 1247799. Appreciation is also extended to the Instituto Potosino de Investigaci\u0026oacute;n Cient\u0026iacute;fica y Tecnol\u0026oacute;gica A.C. (IPICYT) for its academic collaboration through the diploma program in Artificial Intelligence for Health Informatics. B.C.-E was supported by SECIHTI fellowship\u0026nbsp;1145899 and J.A.G.-P fellowship\u0026nbsp;1328330.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board of the Instituto Nacional de Cancerolog\u0026iacute;a (012/031/ICI, 2012\u0026ndash;2022) (022/068/OMI) (CEI/052/22). All participants provided written informed consent prior to inclusion in the study, following institutional guidelines.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData and Code Availability Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data and code supporting the findings of this study are available upon reasonable request from the corresponding author.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eWhicher D, Philbin S, Aronson N. An overview of the impact of rare disease characteristics on research methodology. Orphanet J Rare Dis [Internet]. 2018 Jan 19 [cited 2025 Feb 16];13(1):14. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC5775563/\u003c/li\u003e\n\u003cli\u003eHee SW, Willis A, Tudur Smith C, Day S, Miller F, Madan J, et al. Does the low prevalence affect the sample size of interventional clinical trials of rare diseases? An analysis of data from the aggregate analysis of clinicaltrials.gov. Orphanet J Rare Dis [Internet]. 2017 Mar 2 [cited 2025 Feb 16];12(1):1\u0026ndash;20. Available from: https://ojrd.biomedcentral.com/articles/10.1186/s13023-017-0597-1\u003c/li\u003e\n\u003cli\u003eCrispin-Rios Y, Faura-Gonzales M, Torres-Roman JS, Quispe-Vicu\u0026ntilde;a C, Franco-Jimenez US, Valcarcel B, et al. Testicular cancer mortality in Latin America and the Caribbean: trend analysis from 1997 to 2019. BMC Cancer [Internet]. 2023 Dec 1 [cited 2025 Jun 18];23(1):1\u0026ndash;8. Available from: https://bmccancer.biomedcentral.com/articles/10.1186/s12885-023-11511-z\u003c/li\u003e\n\u003cli\u003eSurveillance, Epidemiology, and End Results (SEER) Program [Internet]. 2025 [cited 2025 Jun 16]. Cancer Stat Facts: Testicular Cancer. Available from: https://seer.cancer.gov/statfacts/html/testis.html?utm_source=chatgpt.com\u003c/li\u003e\n\u003cli\u003eCuevas-Estrada B, Montalvo-Casimiro M, Munguia-Garza P, R\u0026iacute;os-Rodr\u0026iacute;guez JA, Gonz\u0026aacute;lez-Barrios R, Herrera LA. Breaking the Mold: Epigenetics and Genomics Approaches Addressing Novel Treatments and Chemoresponse in TGCT Patients. International Journal of Molecular Sciences 2023, Vol 24, Page 7873 [Internet]. 2023 Apr 26 [cited 2024 Dec 11];24(9):7873. Available from: https://www.mdpi.com/1422-0067/24/9/7873/htm\u003c/li\u003e\n\u003cli\u003eR\u0026iacute;os-Rodr\u0026iacute;guez JA, Montalvo-Casimiro M, \u0026Aacute;lvarez-L\u0026oacute;pez DI, Reynoso-Nover\u0026oacute;n N, Cuevas-Estrada B, Mendoza-P\u0026eacute;rez J, et al. Understanding Sociodemographic Factors among Hispanics Through a Population-Based Study on Testicular Cancer in Mexico. J Racial Ethn Health Disparities. 2025 Feb;12:148\u0026ndash;60. \u003c/li\u003e\n\u003cli\u003eGonz\u0026aacute;lez-Barrios R, Alcaraz N, Montalvo-Casimiro M, Cervera A, Arriaga-Canon C, Munguia-Garza P, et al. Genomic Profile in a Non-Seminoma Testicular Germ-Cell Tumor Cohort Reveals a Potential Biomarker of Sensitivity to Platinum-Based Therapy. Cancers (Basel) [Internet]. 2022 May 1 [cited 2025 Aug 19];14(9):2065. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC9101377/\u003c/li\u003e\n\u003cli\u003eIslam R, Hansen A, Liesen A, Schorle H. An update on the genetic predisposition of testicular germ cell tumors. Transl Androl Urol [Internet]. 2024 [cited 2025 Jun 16];13(3):476\u0026ndash;8. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC10999021/\u003c/li\u003e\n\u003cli\u003eV\u0026aacute;zquez-Moreno M, Zeng H, Locia-Morales D, Peralta-Romero J, Asif H, Maharaj A, et al. The Melanocortin 4 Receptor p.Ile269Asn Mutation Is Associated with Childhood and Adult Obesity in Mexicans. J Clin Endocrinol Metab [Internet]. 2020 Apr 1 [cited 2025 Feb 16];105(4):E1468\u0026ndash;77. Available from: https://pubmed.ncbi.nlm.nih.gov/31841602/\u003c/li\u003e\n\u003cli\u003eNaser AA, Miyazaki T, Wang J, Takabayashi S, Pachoensuk T, Tokumoto T. MC4R mutant mice develop ovarian teratomas. Sci Rep [Internet]. 2021 Dec 1 [cited 2025 Aug 16];11(1). Available from: https://pubmed.ncbi.nlm.nih.gov/33568756/\u003c/li\u003e\n\u003cli\u003eSeki S, Ohura K, Miyazaki T, Naser AA, Takabayashi S, Tsutsumi E, et al. The Mc4r gene is responsible for the development of experimentally induced testicular teratomas. Sci Rep [Internet]. 2023 Dec 1 [cited 2025 Aug 16];13(1). Available from: https://pubmed.ncbi.nlm.nih.gov/37127675/\u003c/li\u003e\n\u003cli\u003eRujas M, Mart\u0026iacute;n G\u0026oacute;mez del Moral Herranz R, Fico G, Merino-Barbancho B. Synthetic data generation in healthcare: A scoping review of reviews on domains, motivations, and future applications. Int J Med Inform [Internet]. 2025 Mar 1 [cited 2025 Aug 16];195:105763. Available from: https://www.sciencedirect.com/science/article/pii/S138650562400426X?via%3Dihub\u003c/li\u003e\n\u003cli\u003eAkiya I, Ishihara T, Yamamoto K. Comparison of Synthetic Data Generation Techniques for Control Group Survival Data in Oncology Clinical Trials: Simulation Study. JMIR Med Inform [Internet]. 2024 [cited 2025 Aug 16];12. Available from: https://pubmed.ncbi.nlm.nih.gov/38889082/\u003c/li\u003e\n\u003cli\u003eQuintana DS. A synthetic dataset primer for the biobehavioural sciences to promote reproducibility and hypothesis generation. Elife [Internet]. 2020 Mar 1 [cited 2025 Aug 16];9. Available from: https://pubmed.ncbi.nlm.nih.gov/32159513/\u003c/li\u003e\n\u003cli\u003eMiletic M, Sariyar M. Utility-based Analysis of Statistical Approaches and Deep Learning Models for Synthetic Data Generation With Focus on Correlation Structures: Algorithm Development and Validation. 2025 Mar; \u003c/li\u003e\n\u003cli\u003eBraddon AE, Robinson S, Alati R, Betts KS. Exploring the utility of synthetic data to extract more value from sensitive health data assets: A focused example in perinatal epidemiology. Paediatr Perinat Epidemiol [Internet]. 2023 May 1 [cited 2025 Aug 19];37(4):292\u0026ndash;300. Available from: https://pubmed.ncbi.nlm.nih.gov/36482827/\u003c/li\u003e\n\u003cli\u003eMajor-Smith D, Kwong ASF, Timpson NJ, Heron J, Northstone K. Releasing synthetic data from the Avon Longitudinal Study of Parents and Children (ALSPAC): Guidelines and applied examples. Wellcome Open Res [Internet]. 2024 Dec 24 [cited 2025 Aug 19];9:57. Available from: https://pubmed.ncbi.nlm.nih.gov/39931104/\u003c/li\u003e\n\u003cli\u003eLandrum MJ, Lee JM, Riley GR, Jang W, Rubinstein WS, Church DM, et al. ClinVar: Public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res [Internet]. 2014 Jan 1 [cited 2025 Jun 22];42(D1). Available from: https://pubmed.ncbi.nlm.nih.gov/24234437/\u003c/li\u003e\n\u003cli\u003eChertack N, Ghandour RA, Singla N, Freifeld Y, Hutchinson RC, Courtney K, et al. Overcoming sociodemographic factors in the care of patients with testicular cancer at a safety net hospital. Cancer [Internet]. 2020 Oct 10 [cited 2025 Aug 16];126(19):4362\u0026ndash;70. Available from: https://acsjournals.onlinelibrary.wiley.com/doi/10.1002/cncr.33076\u003c/li\u003e\n\u003cli\u003eRezaee ME, Elias R, Li HL, Agrawal P, Pallauf M, Enikeev D, et al. Survival outcomes and molecular drivers of testicular cancer in hispanic men. Urologic Oncology: Seminars and Original Investigations [Internet]. 2024 Sep 1 [cited 2025 Aug 16];42(9):293.e1-293.e7. Available from: https://pubmed.ncbi.nlm.nih.gov/38821727/\u003c/li\u003e\n\u003cli\u003eCuevas-Estrada B, Montalvo-Casimiro M, Munguia-Garza P, R\u0026iacute;os-Rodr\u0026iacute;guez JA, Gonz\u0026aacute;lez-Barrios R, Herrera LA. Breaking the Mold: Epigenetics and Genomics Approaches Addressing Novel Treatments and Chemoresponse in TGCT Patients. Int J Mol Sci [Internet]. 2023 May 1 [cited 2025 Aug 19];24(9). Available from: https://pubmed.ncbi.nlm.nih.gov/37175579/\u003c/li\u003e\n\u003cli\u003eStang A, McMaster ML, Sesterhenn IA, Rapley E, Huddart R, Heimdal K, et al. Histological features of sporadic and familial testicular germ cell tumors compared and analysis of age-related changes of histology. Cancers (Basel) [Internet]. 2021 Apr 1 [cited 2025 Aug 19];13(7). Available from: https://pubmed.ncbi.nlm.nih.gov/33916078/\u003c/li\u003e\n\u003cli\u003eIslam R, Hansen A, Liesen A, Schorle H. An update on the genetic predisposition of testicular germ cell tumors. Transl Androl Urol [Internet]. 2024 [cited 2025 Aug 19];13(3):476\u0026ndash;8. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC10999021/\u003c/li\u003e\n\u003cli\u003eFedurek P, Lacroix L, Lehmann J, Aktipis A, Cronk L, Townsend C, et al. Status does not predict stress: Women in an egalitarian hunter-gatherer society. Evol Hum Sci [Internet]. 2020 [cited 2025 Aug 16];2. Available from: https://pubmed.ncbi.nlm.nih.gov/37588349/\u003c/li\u003e\n\u003cli\u003eNaughton M, Weaving D, Scott T, Compton H. Synthetic Data as a Strategy to Resolve Data Privacy and Confidentiality Concerns in the Sport Sciences: Practical Examples and an R Shiny Application. Int J Sports Physiol Perform [Internet]. 2023 Jul 18 [cited 2025 Aug 16];18(10):1213\u0026ndash;8. Available from: https://journals.humankinetics.com/view/journals/ijspp/18/10/article-p1213.xml\u003c/li\u003e\n\u003cli\u003eWarmenhoven J, Impellizzeri FM, Shrier I, Vigotsky AD, Lolli L, Menasp\u0026agrave; P, et al. Synthetic Data for Sharing and Exploration in High-Performance Sport: Considerations for Application. Sports Medicine [Internet]. 2025 [cited 2025 Aug 16]; Available from: https://pubmed.ncbi.nlm.nih.gov/40569343/\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"National Autonomous University of Mexico","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Synthpop, Synthetic data, Testicular germ cell tumors, MC4R, Cancer genomics","lastPublishedDoi":"10.21203/rs.3.rs-8099876/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8099876/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eObjective\u003c/h2\u003e\u003cp\u003eRare cancers often involve small patient cohorts, which limit statistical power and hinder biomarker discovery. This proof-of-concept study evaluates the feasibility of repurposing the Synthpop R package\u0026mdash;originally designed for data anonymization\u0026mdash;to generate synthetic datasets and improve statistical inference in rare cancer studies. We demonstrate this approach using the association between the MC4R rs79783591 variant and clinical outcomes in testicular germ cell tumors (TGCT).\u003c/p\u003e\u003ch2\u003eMaterials and Methods\u003c/h2\u003e\u003cp\u003eA retrospective cohort of 220 Mexican TGCT patients was analyzed, identifying 9 heterozygous carriers of the MC4R variant. To overcome limited sample size, we used Synthpop to generate 234 synthetic carriers, maintaining the original data structure and statistical distributions. Dataset fidelity was validated using machine learning models, principal component analysis, and structural similarity metrics. Cox proportional hazards models and Kaplan\u0026ndash;Meier survival analyses assessed associations with overall survival.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e\u003cp\u003eReal carriers were diagnosed at a younger median age (22 vs. 26 years). In the synthetic cohort, the MC4R variant was associated with a threefold increased mortality risk (HR\u0026thinsp;=\u0026thinsp;3.15; 95% CI: 2.06\u0026ndash;4.82; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001), supporting findings in the real cohort (HR\u0026thinsp;=\u0026thinsp;5.69; 95% CI: 1.56\u0026ndash;20.7; p\u0026thinsp;=\u0026thinsp;0.008). Synthetic data narrowed confidence intervals and improved effect size estimation.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e\u003cp\u003eRepurposing the Synthpop R package provides a novel approach to enhance statistical power in studies with small sample sizes. This strategy can improve inference reliability and accelerate biomarker discovery in rare cancer research; however, further validation in independent cohorts is required to confirm these findings.\u003c/p\u003e","manuscriptTitle":"AI-Driven Synthetic Cohorts to Explore Genetic Associations: Lessons from Testicular Cancer with Relevance to Rare Conditions","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-23 15:50:16","doi":"10.21203/rs.3.rs-8099876/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2bf69494-5c20-4521-b153-68e5b578f7d4","owner":[],"postedDate":"November 23rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":58385244,"name":"Epigenetics \u0026 Genomics"},{"id":58385245,"name":"Biostatistics"},{"id":58385246,"name":"Artificial Intelligence and Machine Learning"}],"tags":[],"updatedAt":"2025-11-23T15:50:16+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-23 15:50:16","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8099876","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8099876","identity":"rs-8099876","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00