Large-scale sequencing study of de novo regulatory Tandem Repeats (TRs) identifies new ASD (Autism Spectrum Disorders) candidate genes integrating gene expression mapping, brain scRNA-seq and organoid models.

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract In this study, we performed an integrative analysis of de novo tandem repeats (TRs) to unravel the missing heritability that may be hidden in 85,394 active cis-regulatory elements (cCREs) from ENCODE through target sequencing in a Spanish cohort of 200 ASD trios, using a robust bioinformatic pipeline. For the integrative analysis, we use data from 1,637 ASD simplex quad families from the Simons Simplex Collection (SSC). We then incorporated multiple layers of functional annotation, including predicted transcription factor (TF) binding sites, gene mapping based on physical proximity and expression correlation, pathogenicity scoring, single-cell RNA-seq data from human brain in ASD cases and controls and cortical organoid expression data. Together, our analyses identified multiple ASD-relevant candidate genes supported by convergent lines of evidence. Notably, ECHS1 emerged as a strong candidate, affected by several de novo TRs in both the Spanish cohort and the SSC. It was also identified as the most significantly associated gene through expression-based gene mapping (T-Gene) and showed consistent differential expression in excitatory neurons of the cerebral cortex at the single-cell level along with increased expression in late-stage cortical organoids. These findings remark the value of integrating genetic and transcriptomic information to improve the identification of potential risk genes for ASD, particularly within non-coding regions. Our approach also highlights the importance of identifying complex genetic variation, such as de novo TRs, that are typically missed in conventional exome or whole-genome analyses, and require specialized bioinformatic strategies for accurate detection and interpretation.
Full text 159,552 characters · extracted from preprint-html · click to expand
Large-scale sequencing study of de novo regulatory Tandem Repeats (TRs) identifies new ASD (Autism Spectrum Disorders) candidate genes integrating gene expression mapping, brain scRNA-seq and organoid models. | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Large-scale sequencing study of de novo regulatory Tandem Repeats (TRs) identifies new ASD (Autism Spectrum Disorders) candidate genes integrating gene expression mapping, brain scRNA-seq and organoid models. Maria Cristina Rodriguez Fontenla, Pablo Carballo-Pacoret, Sara Dominguez-Alonso, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8374597/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 12 You are reading this latest preprint version Abstract In this study, we performed an integrative analysis of de novo tandem repeats (TRs) to unravel the missing heritability that may be hidden in 85,394 active cis-regulatory elements (cCREs) from ENCODE through target sequencing in a Spanish cohort of 200 ASD trios, using a robust bioinformatic pipeline. For the integrative analysis, we use data from 1,637 ASD simplex quad families from the Simons Simplex Collection (SSC). We then incorporated multiple layers of functional annotation, including predicted transcription factor (TF) binding sites, gene mapping based on physical proximity and expression correlation, pathogenicity scoring, single-cell RNA-seq data from human brain in ASD cases and controls and cortical organoid expression data. Together, our analyses identified multiple ASD-relevant candidate genes supported by convergent lines of evidence. Notably, ECHS1 emerged as a strong candidate, affected by several de novo TRs in both the Spanish cohort and the SSC. It was also identified as the most significantly associated gene through expression-based gene mapping (T-Gene) and showed consistent differential expression in excitatory neurons of the cerebral cortex at the single-cell level along with increased expression in late-stage cortical organoids. These findings remark the value of integrating genetic and transcriptomic information to improve the identification of potential risk genes for ASD, particularly within non-coding regions. Our approach also highlights the importance of identifying complex genetic variation, such as de novo TRs, that are typically missed in conventional exome or whole-genome analyses, and require specialized bioinformatic strategies for accurate detection and interpretation. Health sciences/Diseases/Psychiatric disorders/Autism spectrum disorders Biological sciences/Genetics Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 INTRODUCTION Autism Spectrum Disorders (ASD) are neurodevelopmental disorders (NDDs), characterised by repetitive or restrictive behaviours, atypical social function and communication deficits (1,2). ASD have a high genetic component with an estimated heritability of around 80% (3). Although several autism-associated genes and variants have been identified through various genetic approaches—including GWAS (4), WES (5), WGS (6), and integrative omic analyses such as single-cell RNA sequencing (scRNA-seq) (7), these studies still fall short of accounting for the full genetic contribution estimated in family and twin studies. Thus, a substantial fraction of the so-called “missing heritability” remains unexplained. Several factors may contribute to missing heritability in ASD, including sample heterogeneity, undetected ultra rare variants, variants lying within regulatory regions, somatic and postzygotic mutations and complex structural variation such as tandem repeats and mobile elements. Detecting many of these variants often requires specialized sequencing technologies and bioinformatic tools. Thus, de novo and inherited single-base mutations, small insertions and deletions (indels) or large copy number variations (CNVs) can be identified under conventional sequencing and routine genetic diagnosis screening procedures. However, part of this missing heritability may stem from the fact that current bioinformatics pipelines often discard information from repetitive regions which are challenging to interpret using short-read sequencing methods (8,9). Tandem repeats (TRs) are sequences of bp motifs that are repeated consecutively a specific number of times in the DNA sequence. They vary both in the sequence of the repeated motif and in the number of repeat units. Motif repeats shorter than 100 bp represent approximately 3% of the human genome (10), with over a million known TRs distributed throughout the genome. The classification of TRs has always been diverse, as they are usually classified according to their length (expansion or contraction of the repeat motif), location and/or frequency of occurrence: microsatellites (STRs) with repeat motifs shorter than 5 bp, are usually present throughout the genome. Minisatellites, with repeat motifs ranging from 5 bp to 100 bp, are less common. Satellite DNA, found at centromeres, consists of 6 bp motifs repeated between 300 and 8,000 times, while tandem repeats also consisting of 6 bp motifs, are repeated between 2,000 and 50,000 times at the ends of chromosomes. In addition, motifs longer than 5bp, often called Variant Number Tandem Repeats (VNTRs), have also been studied for their role in human traits (11). As we can see, the classification of TRs is complex. They are found in many parts of the genome, although they are predominantly located within non-coding regions. TRs accelerate evolution and adaptation due to their mutation rates, which are between 10 and 100,000 times higher than those of the rest of the genome (12). These high mutation rates create new copy lengths and copy numbers, contributing to genomic diversity and potentially giving rise to novel traits. In addition, TRs are involved in cellular processes such as cell cycle stability, telomere maintenance, or the regulation of gene expression. They are known to contribute to disease etiology by altering TF binding and three-dimensional chromatin structure, affecting differential gene expression (13). For example, well-known repeats of TRs lead to diseases, such as the CAG repeat motif, which causes Huntington's Disease (14), as well as other genetic disorders (15). In addition to autism, where the expansion of the CGG motif in the FMR1 gene is a direct cause of Fragile X syndrome, a condition closely related to ASD (16), TRs are linked to several other diseases of the nervous system, such as ataxia, schizophrenia, amyotrophic lateral sclerosis (ALS) or dementia (17–20). It is known that the number of copies of TRs (21) and structural variations are relevant in the aetiology of ASD (22,23), as it is also known that there is an increased variation in TRs in regulatory regions (11). There are studies that have investigated genome-wide tandem repeats (24,25), and others who have studied mutations in regulatory regions (26), underscoring the need for future research on the identification of regulatory target genes and neuronal populations involved using single-cell resolution, which has not yet been addressed. In this study, we set out to investigate the role of regulatory de novo tandem repeats (TRs) in ASD risk by integrating genomic, epigenomic, and transcriptomic data. Specifically, our objectives were: (1) to identify de novo TR mutations within active cis-regulatory elements (cCREs) in a Spanish ASD cohort of 200 ASD trios. and validate their contribution using large-scale data from the SSC (1,637 ASD simplex quad families); and (2) to functionally prioritize these variants through multi-layered annotation, including predicted disruption of transcription factor binding, expression-based gene mapping, pathogenicity scoring, biological pathways, single-cell transcriptomic signatures from the ASD human brain vs controls and expression patterns of target genes in cortical brain organoids. This approach aims to uncover novel ASD risk genes within the regulatory genome and to highlight the contribution of complex regulatory variation that is often missed by conventional sequencing analyses. MATERIAL AND METHODS Spanish Cohort The analysis presented in this study was conducted on a subset of 200 trios from the Spanish ASD cohort described in Alonso-Gonzalez et al. 2021 (27), which comprises 360 trios, each consisting of unaffected parents and an affected proband, from Fundación Pública de Medicina Xenómica, Santiago de Compostela and Hospital Universitario Gregorio Marañon, Madrid, Spain. Patients involved in the study were diagnosed by psychiatrists or neurologists, following the criteria of both the Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition Text Revision (DSM-IV-TR) and Fifth Edition (DSM-5). When deemed necessary, the Autism Diagnostic Observation Schedule (ADOS) and the Autism Diagnostic Interview-Revised (ADI-R) were also administered. Patients with syndromic autism were excluded. All participants (probands, parents or legal representatives) gave their written consent and the study was conducted under the Declaration of Helsinki. The Galician Committee of Research Ethics (Xunta de Galicia) has approved this study under registration number 2020/400. From the 360 trios forming the Spanish cohort with exome data, Copy Number Variation (CNV) array and phenotypic information, 200 trios were selected among those with negative results in the CNV array, following the procedure outlined in Alonso-Gonzalez et al., 2021 (27). Selection of regulatory regions Targeted sequencing (see next section) has been carried out on cis-regulatory regions (cCREs) from the ENCODE project (28) (version 2, https://screenv2.wenglab.org/ ) , corresponding to approximately 1% of the total human genome, classified in 5 different categories: promoter-like signatures (PLS), proximal enhancer-like signatures (pELS), distal enhancer-like signatures (dELS), high DNase and H3K4me3 signals only (DNase–H3K4me3) and high DNase and CTCF signals only (CTCF-only). Brain and GI tissues with DNAse-Seq data available were selected, both from adult (n = 14) and embryonic tissue (n = 59). For each tissue, cCREs labeled as “Low-Dnase” were excluded, as they are inactive in the given tissue. From the total of 926,535 human cCREs, we selected those that exhibit activity in a greater number of the interrogated tissues. Thus, we selected cCREs that were active in 36 or more tissues (n = 85,394 cCREs). Sequencing of regulatory regions Targeted sequencing was done at the National Center for Genomic Analysis (CNAG) using the KAPA HyperChoice Target Enrichment custom probes and with an average coverage of 30⋅. Samples with sex discrepancies when compared to reported pedigrees were dropped and replaced, along with all other samples from the same family. Moreover, samples which failed CNAG’s quality control (DNA < 100 ng / critical degradation (genomic quality number (GQN) < 3.3)) were also removed and consequently substituted, leaving 71 trios from Santiago and 129 from Madrid. Sequencing reads were aligned to GRCh38/hg38 using the Burrows-Wheeler Aligner. Raw CRAM and vcf sequencing files have been transferred, stored and handled at the CESGA, Centro de Supercomputación de Galicia ( https://www.cesga.es ). Sequencing data are available at: EGA repository number: EGAS50000001395 https://ega-archive.org/studies/EGAS50000001395 Simons Simplex Collection cohort To carry out a gene-level meta-analysis and a single-cell mapping analysis, we requested the STR dataset from the Simons Simplex Collection via the Simons Foundation Autism Research Initiative (SFARI) ( https://base.sfari.org ). We retrieved a CSV file containing detailed information on 175,291 mutations identified in 1,593 quad families (also used in Mitra et al.(29)). Genomic positions for STRs were overlapped with the cCREs analysed in our study using bedtools ( https://github.com/arq5x/bedtools2 ) , creating a new custom file only with cCREs that we had sequenced that had mutations in SSC cohort. The subsequent gene annotation was done with this file as input. It should be noted that the TRs workflow applied in our cohort and in Mitra et al. is largely identical ensuring a high degree of methodological homogeneity. Tandem Repeat (TR) Mutation Detection Workflow The workflow (including file types, tools used, input and output formats, applied filters, and quality control steps performed in the analysis of the Spanish cohort), is summarized in Fig. 1 . Custom TR catalogue elaboration and file transformation, calling of TRs, call and locus-level filters of variants and in-silico detection of TRs de novo variants, are found in Supplementary Information . A bed file with 107 mutations remained ( Supplementary Table 1 ). Prioritizing pathogenic TR mutations SISTR (Selection Inference at Short TRs) method described in Mitra et al., considers mutation, genetic drift and negative natural selection to calculate selection coefficients for each TRs. The parameter ‘ s ’ can be interpreted as the reduction in reproductive fitness for each repeat unit copy number, relative to the TR model allele in the population. To estimate selection coefficients in the overlapping cCREs- TRs with the GangSTR catalogue, SISTR was run on the parental cohort (400 individuals). Pathogenicity coefficients, which are allele-specific selection coefficients, have been calculated as follows: (|a - opt|)s , where a is the number of repeats for the de novo allele, opt is the optimal (or modal) number of repeats for that TR, and s is the selection coefficient for the TR calculated by SISTR in our parental cohort. The input for SISTR was carried out using the statSTR tool from TR Tools package ( https://github.com/gymrek-lab/TRTools ). A single GangSTR vcf file was created, and statSTR was run with the filters --acount and --numcalled, to extract allele frequencies and total alleles. The resulting output was then merged with the GangSTR TR catalogue to obtain the input for running SISTR on the parental cohort. Functional Annotation of TF Binding Sites To determine whether de novo TR mutations disrupt transcription factor binding sites (TFBS) in the Spanish cohort, we have run the FIMO tool (30) (also from MEME Suite), designed to scan nucleotide sequences, to search for TF binding motifs. We selected the binding motifs of vertebrates from the JASPAR database ( https://jaspar.elixir.no/downloads/ ) (31), that provides a catalog of TF motifs, obtained from validated experiments such as ChIP-Seq. Two fasta files included sequences centered on the TR loci within or near SFARI genes, with one file representing the reference genome and the other incorporating the corresponding d e novo TR mutation sequences were generated. To this end, we have extended the mutated regions by ±15 bp to improve possible TF binding to overlapping regions between the flanking region and the mutation itself. Integrative genetic analysis of de novo TRs from SSC collection and Spanish cohort We aim to conduct a joint meta-analysis of microsatellite datasets from the Simons Simplex Collection (SSC) (25), (downloaded from base.sfari.org) and our Spanish cohort to identify overlapping genomic regions associated with ASD. The methods used in that study ensure homogeneity in the detection and annotation of TRs across both cohorts, allowing for reliable integration with our dataset. Integrating these datasets increase statistical power, allow us to assess the consistency of observed signals across cohorts, and strengthen the robustness of candidate loci with potential functional relevance. Thus, a total of 200 trios of Spanish cohort + 1,637 quads from SSC and 107 de novo TRs mutations + 1305 from SSC were analysed. Overlapping cCREs between both cohorts were identified using bedtools intersect . Gene Mapping of Regulatory TRs using physical distance and expression data To identify target genes based on genomic physical distance, we used the -closest function from the bedtools suite ( https://bedtools.readthedocs.io/en/latest/content/tools/closest.html ). To identify target genes for TRs within cCREs, we used the T-Gene tool (32), which infers regulatory relationships by combining genomic closeness and ENCODE ChIP-seq expression data (hg19). As input for both T-Gene and bedtools, we used the 57 cCREs carrying the 107 high-confidence de novo regulatory TRs from the Spanish cohort ( Supplementary Table 2 ). The same was done with the SSC dataset (565 cCREs carrying 1305 mutations, Supplementary Table 3 ). A GTF annotation file ( https://www.gencodegenes.org/human/release_7.html ) corresponding to the hg19 genome assembly, was used as input in bedtools to ensure consistency between both analysis physical and expression. To match this reference genome, TR coordinates from both the Spanish cohort and the SSC dataset from Mitra et al. were converted to hg19 using UCSC Table Browser ( https://genome.ucsc.edu/cgi-bin/hgTables ). We customized T-Gene output to include all potential regulatory links by selecting the “Link” option, which avoids discarding cCREs that may share target genes. The output of T-gene containing regulatory genes for both cohorts is in Supplementary Table 4 (Spanish cohort) and 5 (SSC cohort). For the Spanish cohort, genes were ordered to each analysed cCRE ( Supplementary Table 6 ). Subsequently, gene lists were filtered considering Correlation and Distance p-value (CnD p-value) < 0.05 of T-Gene. These filtered gene expression lists were used for distance vs expression analysis, GO enrichment and single-cell analysis. For distance vs expression gene comparison, we filtered T-gene output taking in account only the most linked gene (lowest CnD p-value) for each cCRE for both cohorts. In the Spanish cohort, we looked for genes related with SFARI database, to retain genes highly linked with ASD. GO enrichment analysis of TR associated genes Gene Ontology enrichment analysis has been carried out on two platforms: Enrichr ( https://maayanlab.cloud/Enrichr/ ) and ReVIGO ( http://revigo.irb.hr ). Enrichr is an online tool for enrichment analysis of gene sets; it was used to identify biological processes associated with genes related to our regulatory TRs by expression (T-Gene). To carry out the GO enrichment analysis in the Spanish cohort, the output from T-Gene ( Supplementary Table 4 ) was filtered to retain only those links with a CnD p-value < 0.05. The same was done for SSC cohort. To carry out the GO enrichment analysis in both cohorts, we retain genes from T-gene output ( Supplementary Tables 4 and 5 ) with CnD p-value < 0.5 without duplicated genes shared between cohorts. For both cohorts and the metaanalysis with data from Mitra et al., we applied two thresholds: q-value < 0.1 and Bonferroni correction for multiple tests. Moreover, we have selected the top 30 Enrichr-enriched GO biological processes based on their p-value. Visualization was further refined using a custom R script adapted from ReVIGO. Single-Cell Gene Differential Gene Expression Analysis In order to carry out a single cell differential gene expression analysis in neuronal and glial cells, we used ASD scGENE Portal ( http://solo.bmap.ucla.edu/asdscgene/ ) , resource that uses RNA-seq cell data from brain (33 ASD individuals and 31 control subjects) (33), comprising nearly 600,000 cell nuclei. We used this resource to explore whether any of the significant associated genes from both cohorts were expressed in specific neuronal or glial subtypes, and whether their expression was altered in ASD compared to controls. There were 35 cell types identified by single-nucleus RNA sequencing (snRNA-seq), divided in 8 major cell types: oligodendrocyte progenitor cells (OPCs); astrocyte (ASTRO); microglia (MG); endothelial cells (ENDO); inhibitory neurons (INT); excitatory neurons (EXT); oligodendrocyte (ODC) and blood-brain barrier (BBB). To carry out this analysis in the Spanish cohort, we retain genes from the T-gene output ( Supplementary Table 4 ) with a CnD p-value < 0.05, and we looked for genes showing significant differential expression between ASD cases and controls defined by an FDR 0.3. For the dataset from Mitra et al., due to the larger number of candidate genes, we applied a slightly more stringent initial filter, selecting genes from the T-gene output ( Supplementary Table 5 ) with a CnD p-value < 0.05 and a q-value < 0.1. These genes were then queried in the single-cell portal using the same differential expression thresholds as described above. Gene expression trajectories across the differentiation of human cortical organoids in vitro and in vivo human brain development. To explore the developmental trajectories of candidate gene expression in BrainSpan and brain cortical organoids, we used the Gene Expression tool (GECO, http://solo.bmap.ucla.edu/shiny/GECO/# ), developed based on Gordon et al. (34). Visualization and Statistical Analysis Tool RStudio v2024.12.0 + 467 was used to create the graphs presented, and the circlize library for RStudio was used for customized circo plot. Microsoft Excel 16.97 was used for the directionality bias graph. RESULTS SPANISH COHORT Identification, Distribution, and Patterns of Regulatory de novo Tandem Repeats (TRs) MonSTR identified 1,067 de novo non-strict de novo TRs mutations in the 200 families analysed from the Spanish cohort ( Supplementary Table 7 ). After QC, 3 children and 6 TR loci were discarded and a list of 821 non-strict de novo TR mutations without outliers was obtained. After MonSTR prioritization, a total of 107 strict de novo TR mutations remained in 57 different cCREs, 83 different probands, out of the 200 analysed (after using GangSTR and MonSTR pipeline. See Methods section and Supplementary Table 1 ). The mean number of mutations per proband was 0.51, which is consistent with the genome-wide average of 53.9 mutations per individual reported by Mitra et al., taking in account that our analysis covers approximately 1% of the genome. Analysing candidate cis-regulatory elements (cCREs), we found an enrichment in one particular cCRE category: proximal enhancer-like signature (pELS), with a total of 38 mutations out of 107 mutations (Fig. 2 a). Regarding the patterns of expansion, the most frequent pattern in the strict de novo TRs, involved the motif being repeated just once more than in the parent sequence. In contrast, the most common contraction mutation pattern was the deletion of three consecutive copies of the motif (Fig. 2 b). When analyzing the motif size of TRs in our cohort, we found that 38 out of the 107 mutations occurred in TRs with 2 bp motifs, making this the most frequent motif size in our dataset (Fig. 2 c). In our study, a significant bias towards expansions over contractions in de novo TRs was found. Specifically, 81 out of 107 (76%) were expansions while 26/107 (24%) were contractions (binomial test, n = 107, k = 81; p = 9.37x10 − 8 ) ( See Supplementary Table 1 ). In addition, there is a directionality bias in mutation size that is observed in other studies as well. When the parental allele is short, the mutation tends to expand, and when the parental allele is large, it tends to contract. (Fig. 2 d). Prioritizing pathogenic TR mutations Due to SISTR’s current limitation to repeat motifs of 2–4 base pairs, not all regions could be analyzed. Of the total regions considered, 7,561 regions were analysed using SISTR. 134 regions (1.8% could not be analysed, indicating that SISTR may assume certain patterns in the TRs that are not fulfilled at these loci). As a result, selection coefficients (‘s’) were successfully estimated for 7,427 loci ( Supplementary Table 8 ). Of the 107 de novo TR mutations identified, SISTR was able to calculate pathogenicity scores for 66. Among these, 13 mutations in 16 probands were in the top quartile of pathogenicity (score > 0.002655), and one mutation was linked with a SFARI gene (see Gene Mapping in the next section), CASZ1 (See Supplementary Table 9 ). Detailed information for all 66 mutations analyzed with SISTR is provided in Supplementary Table 10 . Gene Mapping of Regulatory TRs by physical distance and expression data As we mentioned in Material and Methods section, 57 cCREs with TRs mutations in the Spanish cohort were associated with their putative target genes using expression data of T-Gene ( Supplementary Table 4 ). For the distance vs expression comparison, T-gene links with CnD p-value > 0.05 were discarded, so 53 cCREs finally remain, because 4 cCREs did not have any link with CnD p-value < 0.05 (see Material and Methods and Supplementary Table 11 ). Closest genes by physical distance using bedtools are listed in Supplementary Table 12 . These 53 cCREs were used to compare whether the regulated gene was the closest physically or not. In nearly half of the cCREs (21 out of 53; 39.6%), the gene most significantly associated with the regulatory region was not the closest gene in terms of physical distance ( See Supplementary Table 13 ). Moreover, in the Spanish cohort, 9 mutations affecting 8 different cCREs were linked to genes listed in the Simons Foundation Autism Research Initiative (SFARI) database (Banerjee-Basu & Packer, 2010), the most comprehensive and reliable resource for ASD-associated genes (see Supplementary Table 14 ). Functional Annotation of TF (Transcription Factor) Binding Sites The results of looking at the cCREs with a linked SFARI gene have shown that in most cases (80%), TFs predicted to have the strongest binding to the mutated regions are the same as those predicted to bind the unmutated regions. The majority of these top-binding TFs belong to the ZNF (zinc finger) family. However, we identified two de novo TRs that result in a change in the TF with the highest predicted binding affinity. The first TR affects TF binding by replacing KLF9 with Nrf1 , a gene previously studied in the context of neurodevelopmental disorders which belongs to a family of TFs previously related to neuronal functions. In the second case, a de novo TR alters binding from ZNF384 to EWSR1-FLI1 ( Supplementary Table 15 ). GO enrichment analysis of TR associated genes GO enrichment analysis was performed using the genes linked to cCREs with a CnD p-value < 0.05 identified in the T-Gene output for Spanish cohort (n = 122) (See Supplementary Table 16 ). No GO term was significant (q-value < 0.1). However, when applying the Bonferroni correction for multiple testing (adjusted significance threshold = 0.05/121 ≈ 0.00041), two GO terms were significantly enriched: Regulation of Cell Differentiation (GO:0045595, p-value = 0.00024) and Negative Regulation of Astrocyte Differentiation (GO:0048712, p-value = 0.00036) ( Supplementary Table 17 ). The top 30 biological processes by p-value from GO enrichment analysis of Enrichr, using ReVIGO for visualizing Spanish cohort genes are shown in Supplementary Fig. 1 . Single-Cell Gene Differential Gene Expression Analysis scRNA-seq analysis in the Spanish cohort was done with the 122 genes linked to cCRE with CnD p-value < 0.05, the same used for the GO analysis ( Supplementary Table 16 ). Single cell gene differential gene expression analysis revealed differential expression of several genes ( Supplementary Table 18 ). Among them, ECHS1,CALY, KTN1 and THRA were overexpressed in cases, while NR4A1, PHLDB2, UBSK, KLHL32, PBX1, GLI3, NRG1 and FARP1 were underexpressed in cases. C9orf3 was underexpressed in cases in oligodendrocytes and overexpressed in excitatory neurons. NRG1 shows the greatest degree of underexpression in cases from the Spanish cohort ( logFC=- 1,016; FDR = 5.69 e − 14 ; cell type = EXT_9_L6) . GLI3 was underexpressed in astrocytes (logFC=-0.397, FDR = 1.26 x 10 − 4 ), THRA was overexpressed in excitatory neurons (logFC = 0.397, FDR = 1.97 x 10 − 3 ) and PBX1 was underexpressed in several cell types, being astrocytes the most underexpressed cell type (logFC=-0.484, FDR = 5.33 x 10 − 5 ). These three genes met the thresholds for differential expression and are present in the SFARI database ( Supplementary Table 5 ). Genes that did not met these thresholds were excluded in subsequent analysis. Integration of results in the Spanish cohort This integrative circoplot (Fig. 3 ) highlights the convergence of de novo TRs mutations, regulatory genomic context (non coding regions), and gene expression signatures at the single-cell level. Notably, several high-priority genes (e.g., ECHS1, CASZ1, PBX1) appear in multiple layers of the analysis, suggesting that they are strong candidates supported at genetic and functional annotation levels. Furthermore, predicted TF binding alterations driven by TR mutations, offering a potential mechanistic link between non-coding variation and transcriptional dysregulation were also represented. The integration of single-cell transcriptomic data strengthens the relevance of these findings by providing cell-type-resolved evidence of dysregulation. It is important to remark that ECHS1 is not only affected by TR mutations in regulatory regions, but also shows altered expression in ASD cases versus controls, reinforcing its candidacy as a neurodevelopmental risk gene. INTEGRATIVE ANALYSIS : SIMONS SIMPLEX COLLECTION AND THE SPANISH COHORT Integrative genetic analysis of de novo TRs from SSC collection and Spanish cohort Seven cCREs were overlapped between both cohorts considering that Mitra et al. 2021 carried out a WGS analysis ( Supplementary Table 19 ). Gene Mapping of Regulatory TRs by physical distance and expression data As we mentioned in Material and Methods section, 565 cCREs with TRs mutations in the SSC cohort were associated with their putative target genes using expression data of T-Gene ( Supplementary Table 5 ). For the distance vs expression comparison, T-gene links with CnD p-value > 0.05 were discarded, so 487 cCREs finally remain, because 78 cCREs did not have any link with CnD p-value < 0.05 (see Material and Methods and Supplementary Table 20 ). Closest genes by physical distance using bedtools are listed in Supplementary Table 21 . These 487 cCREs were used to compare whether the regulated gene was the closest physically or not. Among the most significantly associated genes for these 487 cCREs, 162 (33.3%) were not the nearest gene based on physical distance ( Supplementary Table 22 ). If we remain with the most significant gene linked to each cCRE both of our cohort and from Mitra et al., 179/530 of the genes weren't the closest to the cCRE (33.8%). Ten regulatory genes were shared between the Spanish cohort and the SSC collection, taking in account the most significant genes linked with cCREs with CnD p-value < 0.05): ECHS1, C9orf3, C9orf72, FARP1, KIAA0922, PBX1, SOCS3, UBE2K, THNSL2 and WIPF2 ( Supplementary Table 23 ). GO enrichment analysis of TR associated genes As we commented in the Spanish cohort, GO enrichment analysis was performed using the genes linked to cCREs with a CnD p-value < 0.05 identified in the T-Gene output of SSC cohort (n = 1125) ( Supplementary Table 24 ). In the SSC cohort, no GO term was significant (q-value < 0.1). When applying the Bonferroni correction for multiple testing (adjusted significance threshold = 0.05/3268 ≈ 0.00002), no GO term was significant either ( Supplementary Table 25 ). When the Spanish cohort was meta-analysed together with the SSC data, without duplicated genes shared between cohorts (n = 1220 genes) ( Supplementary Table 26 ), the significant GO terms (q-value < 0.1) were: Positive Regulation of Transcription by RNA Polymerase II (GO:0045944, q-value = 0.0099); Positive Regulation of DNA-templated Transcription (GO:0045893, q-value = 0.3475) and Neural Tube Closure (GO: 0001843, q-value = 0.0567) ( Supplementary Table 27 ). If we apply Bonferroni correction (0.05/212 = 0.00024), the significant GO terms were the same as above, plus Primary Neural Tube Formation (GO:0014020, p-value = 0.00019) ( Supplementary Table 27 ). In this analysis, 27 genes were shared between cohorts ( Supplementary Table 28 ), including ECHS1 . The top 30 biological processes by p-value from GO enrichment analysis of Enrichr, using ReVIGO for visualizing genes from both cohorts (integrative analysis) are shown in Fig. 4 . Single-Cell Gene Differential Gene Expression Analysis The same analysis of scRNA-seq done in the Spanish cohort, was done with linked genes from the SSC cohort with CnD p-value < 0.05 and q-value < 0,1 (n = 202), as we discussed above ( Supplementary Table 28 ). Results have shown that the most upregulated gene was PTGES3 ( logFC = 0.855; FDR = 1.60x10 − 10 ; cell type = EXT_9_L6 ), while the most downregulated was RERE (logFC= -0,647; FDR = 5.87 x10 − 8 ; cell type = EXT_9_L6 and l ogFC= -0,934; FDR = 3.93 x10 − 6 ; cell type = ODC_1) ( Supplementary Table 29 ). When comparing gene mapping lists from the Spanish cohort (CnD p-value 0.05 and q-value < 0.1), 5 genes are shared between cohorts, ECHS1, THNSL2, KIAA0922, PRRC2A and C9orf72 ( Supplementary Table 31 ). However single cell gene expression meta-analysis showed that only ECHS1 was significantly differentially expressed in both our cohort and the SSC cohort, specifically in excitatory neurons from the cerebral cortex ( logFC = 0,343; FDR = 5.92x10 − 8 ; cell type = EXT_4_L56 ). Despite being identified as pathogenic by SISTR and supported by expression data in both cohorts and SFARI classification, CASZ1 did not exhibit significant differential expression between ASD cases and controls, nor cell-type-specific expression. Integration of SSC and region in chromosome 10 emerges as the top candidate region ECHS1 (enoyl-CoA hydratase, short chain 1) is the only differentially expressed gene between ASD cases and controls when the targeted genes from de novo TRs in the Spanish and SSC cohorts are considered together in a single-cell analysis (Supplementary Tables 18 and 29 ). We found one individual with a mutation in the cCRE regulating ECHS1 in the Spanish cohort and 11 individuals with 11 different mutations in the same cCRE in the SSC cohort (see Supplementary Table 32 ). Moreover, we discovered that this region is close to another mutated cCRE in the Spanish cohort linked to another gene, CALY (Calcyon Neuron Specific Vesicular Protein) (Fig. 6 ). We would like to point out that, as will be discussed in the next section, genes were assigned to mutated cCRES using T-Gene, an in silico method that does not employ ASD-specific expression data. Single-cell brain data from ASD cases and controls identified ECHS1 as the unique gene, but CALY ,located approximately 25000 bp upstream, was the most overexpressed when genes mutated in the Spanish cohort were considered (in interneurons and excitatory neurons) (see Supplementary Table 2 ). Together with the described function of CALY , the encoded protein interacts with the D1 dopamine receptor, this data suggests that the entire region would be an excellent candidate for future ASD functional studies. We observed that ECHS1 expression decreases during the early stages of differentiation of cortical organoids (~ day 100). After this point, expression progressively increases and remains elevated throughout later stages of neurodevelopment. This biphasic pattern suggests that ECHS1 may have a role both in early neurodevelopmental transitions and in later maturation processes. Across human brain development (Brain Span data) ECHS1 shows relatively stable expression with a modest dip around stage 6 (late fetal/early postnatal), followed by an increase during later postnatal stages. In contrast, CALY expression markedly rose during early and mid-differentiation (100–300 days), and subsequently declined, suggesting predominant involvement in early neurodevelopmental processes. Together, these findings highlight distinct yet complementary temporal profiles: whereas CALY is most active during early neuronal specification and connectivity, ECHS1 becomes increasingly relevant at later stages, underscoring how disruption of either process may contribute to ASD. Discussion In our cohort, 57 different candidate cCREs and 107 high-confidence strict de novo TRs were identified. TRs demonstrated a large bias toward expansions (76%) over contractions, which is consistent with previous genome-wide studies. Interestingly, in contrast to the more prevalent single-repeat deletions previously documented, the most common contraction pattern in our dataset featured the deletion of three repeats (25). In addition, it was also found that there was a clear bias toward expansions among de novo TRs mutations. This pattern is consistent with prior genome-wide studies, which have also reported a higher prevalence of expansions over contractions among de novo TR mutations (25). From a functional perspective, this bias may have important biological implications. Expansions in regulatory regions could disrupt transcription factor binding motifs, alter chromatin accessibility, or interfere with the timing and levels of gene expression mechanisms during neurodevelopment (35). Importantly, the observed bias toward expansions was detected in a set of high-confidence, strictly de novo TRs, suggesting that the pattern is unlikely to result from technical artifacts. Instead, it likely reflects mutational patterns in the germline, potentially linked to chromatin architecture in regulatory regions. Moreover, it is important to note that this cohort consists exclusively of cases that have not been genetically diagnosed with autism through exome sequencing or CNV arrays but exhibit phenotypic traits consistent with ASD. This suggests that there must be some form of gene dysregulation that remains undetectable by the genetic diagnostic methods currently used. Supporting this, a recent genome-wide study identified over 2,000 VNTR loci enriched in regulatory regions, highlighting a potential source of genetic variation that escapes conventional detection (11). In this Spanish cohort, a total of 107 de novo TRs mutations were identified. Using SISTR pathogenicity scores several genes were highlighted. TM4SF1 (36) and C12orf44 (37), which exhibit the highest pathogenicity scores, have been previously linked to neuronal defects but are not listed in the SFARI database. In addition, CASZ1 , identified in a large-scale targeted sequencing study of neurodevelopmental disorders (38), is involved in neurodevelopment, regulating neurogenesis and the transition to gliogenesis (39), and is classified as a high-confidence ASD gene in SFARI. Other relevant genes also emerged, such as TM4SF1 (40,41); UBE2K (42–44); and NR4A1 (45,46). The results from gene mapping analyses, based on physical proximity or expression data, underscore the importance of considering expression data when assigning genes to regulatory regions for further study. In nearly 40% of cases, the gene regulated by a cCRE was not the physically closest gene, indicating that the nearest gene is not always the one controlled by a regulatory region. These results remark the need to integrate three-dimensional genomic data (e.g., Hi-C, Capture-C) and expression-based approaches, rather than relying solely on linear distance, to assign regulatory elements to their target genes. When the our cohort was analyzed independently, GO enrichment highlighted processes such as negative regulation of astrocyte differentiation and regulation of cell differentiation are especially relevant, underscoring the role of astrocytes in the etiology of ASD and pointing to the possible potential functional consequences of these mutations in early neural development and reinforcing the role of astrocytes in ASD etiology (47). Moreover, in the meta-analysis integrating both cohorts, several biological processes of early neurodevelopment were significantly overrepresented. However, none of the genes classified as pathogenic by SISTR in the Spanish cohort were shared with the SSC collection. Using gene mapping with expression data to link cCREs carrying de novo TRs in both cohorts, we identified only ten genes with de novo TRs shared between both cohorts: ECHS1, C9orf3, C9orf72, FARP1, KIAA0922, PBX1, SOCS3, UBE2K, THNSL2, and WIPF2 . Among these, only ECHS1 (enoyl-CoA hydratase, short chain 1) is differentially expressed at the single cell level, showing overexpression in excitatory neurons in ASD brain. Taken together with the findings of Mitra et al., our results point to the involvement of novel genes that can only be uncovered through the application of TRs detection algorithms and subsequent analyses. Importantly, we also identify genes not represented in SFARI, suggesting the presence of complex structural repeat variants as well as rare, inherited variants in additional genes, further underscoring the genetic complexity underlying ASD. E CHS1 , as we will discuss later, emerges as the main candidate gene because it is the only one shared between the Spanish cohort and the large SSC cohort at multiple analysis levels. ECHS1 is located near CALY , one of the target genes identified for de novo TRs in the Spanish cohort, which also showed overexpression in excitatory neurons in scRNAseq data. We found one individual carrying a mutation in the cCRE regulating ECHS1 in the Spanish cohort, as well as 11 additional individuals with 11 distinct mutations in the SSC (29) which remarks the relevance of the region. It should be noted that genes were assigned to mutated cCRES using an in silico method that integrates generic ChIP-seq data from ENCODE and single-cell brain data from ASD cases vs controls, adding greater specificity to the ASD phenotype. We hypothesize that TRs within the cCREs analyzed in this study, or even in additional cCREs not yet included (as ENCODE v3 has identified a large number of novel cCREs in this region than the ENCODE v2 used in this study), may act as regulators of nearby genes independently CALY or ECSH1. Among these two genes, PRAP1 (proline-rich acidic protein 1) and FUOM (fucose mutarotase) are also present; however, no mutations in our study targeted cCRES within these genes, and they are not expressed at the single-cell level in ASD brains nor controls. Since the largest number of mutations are within ECHS1 , we will thoroughly discuss its role without losing sight of the role of CALY , which encodes a dopamine receptor D1-interacting protein (48). The role of CALY in ASD or other NDDs has not been described or studied. However, dopamine receptors have been extensively investigated in ASD at the genetic and functional level. In particular, dopamine receptors D2/D3 in the striatum and D1 in the prefrontal cortex appear to be dysregulated in ASD (49). ECHS1 encodes a mitochondrial enzyme with a key role in b-oxidation, as well as in the metabolic pathways of isoleucine and valine (50). Mutations in ECHS1 are linked to a wide spectrum of clinical phenotypes, ranging from neonatal death to adult survival (51). The most common phenotype manifests like the Leigh syndrome (52) and related encephalopathies, characterised by different neurological impairments, arrhythmia and neonatal seizures (50). A second group of individuals manifest developmental regression resulting in severe developmental delay. Another group of related phenotypes include individuals with a normal development with paroxysmal dystonia that may be exacerbated by illness or exertion (53).In the context of ASD etiology, the overexpression of ECHS1 observed in excitatory neurons may suggest that even subtle dysregulation of this mitochondrial pathway could contribute to altered neuronal activity. Given its essential role in energy metabolism, such changes may be particularly relevant to excitatory/inhibitory imbalance, a mechanism increasingly recognized in autism pathophysiology (54,55). We aim to explore ECHS1 and CALY together in cortical organoids at different times of differentiation to determine whether their regulation might be coordinated during neurodevelopment. Based on cortical organoid and Brain Span data, their timing of functional relevance differs: CALY appears to be relevant during early differentiation and synaptic formation while ECHS1 seems to be more involved in later neuronal maturation and energy metabolism, consistent with its mitochondrial function. In terms of potential disease impact, dysregulation of CALY could disrupt early synaptic development and connectivity, impairing circuit formation, whereas dysregulation of ECHS1 might compromise neuronal maturation processes, particularly through metabolic stress at later stages. We hypothesize that if de novo TRs occur in regulatory regions in individuals with ASD, they may contribute to milder or intermediate phenotypes, in contrast to mutations in the coding region that are associated with severe syndromes as those described above. This suggests that noncoding TRs in regulatory regions of ECHS1 might represent a mechanism underlying less severe, but clinically relevant, ASD-associated phenotypes. The statistical strength of the SSC cohort further supports this gene as a promising target for genetic and functional studies aimed at clarifying how its dysregulation contributes to ASD pathophysiology. Limitations and Future Directions While our study provides novel insights into the role of de novo TRs in regulatory regions and highlights ECHS1 as a promising candidate gene, it should be noted that there are several limitations. First, confirming the impact of TR mutations would require utilizing human-derived neuronal and glial models, such as brain organoids or iPSC-derived cells from the patients harboring the mutation. Second, although we leveraged large cohorts such as the Spanish cohort and SSC, there are important differences between them that may affect interpretation. For example, the Spanish cohort includes individuals with a prior specific exclusion/selection criteria, whereas the SSC cohort does not. These differences may slightly limit the comparability of our findings across populations, even when the methods used to identify variants are the same. Third, the assignment of regulatory elements to target genes relied on current annotations, which may not fully capture the complexity of three-dimensional genome architecture or long-range regulatory interactions. It is also important to note that TRs larger than 150 bp are difficult to detect with established short-read sequencing platforms, as these are generally unable to accurately genotype large or complex repeat expansions (56). While specific short-read TR programs exist, they are currently limited in detecting repeats with very large motif sizes. Furthermore, genes were linked to cCREs containing mutations without accounting for TR characteristics, which may limit prioritization of certain regulatory elements over others. Future studies should integrate multi-omics data, including epigenomic, proteomic, and functional assays, to better understand how TR-mediated dysregulation contributes to ASD. Targeted functional studies of ECHS1 region will be important to explore how subtle changes in expression affect neuronal excitability and glial-neuronal interactions. Longitudinal analyses and studies in more diverse populations could further clarify how regulatory TRs contribute to the heterogeneity of ASD phenotypes, from mild to severe manifestations. Overall, addressing these limitations will provide a more comprehensive understanding of the genetic and cellular mechanisms underlying ASD and may identify novel therapeutic targets. In conclusion, our study highlights the importance of de novo TRs in regulatory regions as contributors to ASD risk, uncovering candidate gene regions such as ECHS1 that would have remained undetected using conventional approaches. By integrating genetic, transcriptomic, single-cell and cortical organoid data across cohorts, we provide new insights of these genes in neurodevelopment. Data availability Sequencing data from the Spanish cohort are transferred to EGA ( European Genome-phenome Archive ) with Study Accession:EGAS50000001395. Declarations Acknowledgements This study has been funded by Instituto de Salud Carlos III (ISCIII) through the project “PI24/00595” and co-funded by the European Union. C. Arango was supported by the Spanish Ministry of Science and Innovation, Instituto de Salud Carlos III (ISCIII), co-financed by the European Union, ERDF Funds from the European Commission, “A way of making Europe”, financed by the European Union – NextGenerationEU (PMP21/00051), PI19/01024. PI22/01824 CIBERSAM, Madrid Regional Government (B2017/BMD-3740 AGES-CM-2), European Union Structural Funds, European Union Seventh Framework Program, European Union H2020 Program under the Innovative Medicines Initiative 2 Joint Undertaking: Project PRISM-2 (Grant agreement No.101034377), Project AIMS-2-TRIALS (Grant agreement No 777394), Horizon Europe, the National Institute of Mental Health of the National Institutes of Health under Award Number 1U01MH124639-01 (Project ProNET) and Award Number 5P50MH115846-03 (project FEP-CAUSAL), Fundación Familia Alonso, and Fundación Alicia Koplowitz. We would like to thank Virginia Sestelo Prado by the work she carried out during her internship as part of her external master practicum working with T-Gene tool. Conflict of interests Dr. Arango has been a consultant to or has received honoraria or grants from Abbot, Acadia, Ambrosetti, Angelini, Biogen, BMS, Boehringer, Carnot, Gedeon Richter, Janssen Cilag, Lundbeck, Medscape, Menarini, Minerva, Otsuka, Pfizer, Roche, Rovi, Sage, Servier, Shire, Schering Plough, Sumitomo Dainippon Pharma, Sunovion, Takeda and Teva. Contributions P Carballo-Pacoret has carried out the analyses and wrote the paper. S.Dominguez-Alonso has carried part of the analyses regarding the sequencing and selection of regulatory regions and managing of sequencing files. J Gonzalez-Peñas, A. Carracedo, C. Arango and M.Parellada participated in the selection and recruitment of samples. A. Carracedo and C.Rodriguez-Fontenla participated in the design and coordination of this study. C.Rodriguez-Fontenla critically revised the work, wrote the paper and approved the final content. References American Psychiatric Association, American Psychiatric Association, editors. Diagnostic and statistical manual of mental disorders: DSM-5. 5th ed. Washington, D.C: American Psychiatric Association; 2013. 947 p. Rosti RO, Sadek AA, Vaux KK, Gleeson JG. The genetic landscape of autism spectrum disorders. Dev Med Child Neurol. 2014 Jan;56(1):12–8. Sandin S, Lichtenstein P, Kuja-Halkola R, Hultman C, Larsson H, Reichenberg A. The Heritability of Autism Spectrum Disorder. JAMA. 2017 Sept 26;318(12):1182. Autism Spectrum Disorder Working Group of the Psychiatric Genomics Consortium, BUPGEN, Major Depressive Disorder Working Group of the Psychiatric Genomics Consortium, 23andMe Research Team, Grove J, Ripke S, et al. Identification of common genetic risk variants for autism spectrum disorder. Nat Genet. 2019 Mar;51(3):431–44. Gaugler T, Klei L, Sanders SJ, Bodea CA, Goldberg AP, Lee AB, et al. Most genetic risk for autism resides with common variation. Nat Genet. 2014 Aug;46(8):881–5. Sanders SJ, Murtha MT, Gupta AR, Murdoch JD, Raubeson MJ, Willsey AJ, et al. De novo mutations revealed by whole-exome sequencing are strongly associated with autism. Nature. 2012 May;485(7397):237–41. Quesnel-Vallières M, Weatheritt RJ, Cordes SP, Blencowe BJ. Autism spectrum disorder: insights into convergent mechanisms from transcriptomics. Nat Rev Genet. 2019 Jan;20(1):51–63. Bahlo M, Bennett MF, Degorski P, Tankard RM, Delatycki MB, Lockhart PJ. Recent advances in the detection of repeat expansions with short-read next-generation sequencing. F1000Research. 2018 June 13;7:736. Hannan AJ. Tandem repeat polymorphisms: modulators of disease susceptibility and candidates for ‘missing heritability’. Trends Genet. 2010 Feb;26(2):59–65. International Human Genome Sequencing Consortium, Whitehead Institute for Biomedical Research, Center for Genome Research:, Lander ES, Linton LM, Birren B, Nusbaum C, et al. Initial sequencing and analysis of the human genome. Nature. 2001 Feb 15;409(6822):860–921. Zhang S, Song Q, Zhang P, Wang X, Guo R, Li Y, et al. Genome-wide investigation of VNTR motif polymorphisms in 8,222 genomes: Implications for biological regulation and human traits. Cell Genomics. 2024 Dec;4(12):100699. Gymrek M, Willems T, Reich D, Erlich Y. Interpreting short tandem repeat variations in humans using mutational constraint. Nat Genet. 2017 Oct 1;49(10):1495–501. Liao X, Zhu W, Zhou J, Li H, Xu X, Zhang B, et al. Repetitive DNA sequence detection and its role in the human genome. Commun Biol. 2023 Sept 19;6(1):954. Macdonald M. A novel gene containing a trinucleotide repeat that is expanded and unstable on Huntington’s disease chromosomes. Cell. 1993 Mar;72(6):971–83. Usdin K. The biological effects of simple tandem repeats: Lessons from the repeat expansion diseases: Table 1. Genome Res. 2008 July;18(7):1011–9. Fyke W, Velinov M. FMR1 and Autism, an Intriguing Connection Revisited. Genes. 2021 Aug 6;12(8):1218. Cortese A, Simone R, Sullivan R, Vandrovcova J, Tariq H, Yau WY, et al. Biallelic expansion of an intronic repeat in RFC1 is a common cause of late-onset ataxia. Nat Genet. 2019 Apr;51(4):649–58. Malik I, Kelley CP, Wang ET, Todd PK. Molecular mechanisms underlying nucleotide repeat expansion disorders. Nat Rev Mol Cell Biol. 2021 Sept;22(9):589–607. Mojarad BA, Engchuan W, Trost B, Backstrom I, Yin Y, Thiruvahindrapuram B, et al. Genome-wide tandem repeat expansions contribute to schizophrenia risk. Mol Psychiatry. 2022 Sept;27(9):3692–8. Trost B, Thiruvahindrapuram B, Chan AJS, Engchuan W, Higginbotham EJ, Howe JL, et al. Genomic architecture of autism from comprehensive whole-genome sequence annotation. Cell. 2022 Nov;185(23):4409–4427.e18. Hannan AJ. Tandem repeats mediating genetic plasticity in health and disease. Nat Rev Genet. 2018 May;19(5):286–98. Brandler WM, Antaki D, Gujral M, Kleiber ML, Whitney J, Maile MS, et al. Paternally inherited cis-regulatory structural variants are associated with autism. Science. 2018 Apr 20;360(6386):327–31. Marshall CR, Noor A, Vincent JB, Lionel AC, Feuk L, Skaug J, et al. Structural Variation of Chromosomes in Autism Spectrum Disorder. Am J Hum Genet. 2008 Feb;82(2):477–88. Trost B, Engchuan W, Nguyen CM, Thiruvahindrapuram B, Dolzhenko E, Backstrom I, et al. Genome-wide detection of tandem DNA repeats that are expanded in autism. Nature. 2020 Oct 1;586(7827):80–6. Mitra I, Huang B, Mousavi N, Ma N, Lamkin M, Yanicky R, et al. Patterns of de novo tandem repeat mutations and their role in autism. Nature. 2021 Jan 14;589(7841):246–50. Werling D, Brand H, An JY, Stone M, Glessner J, Zhu L, et al. LIMITED CONTRIBUTION OF RARE, NONCODING VARIATION TO AUTISM SPECTRUM DISORDER FROM SEQUENCING OF 2,076 GENOMES IN QUARTET FAMILIES. Eur Neuropsychopharmacol. 2019;29:S784–5. Alonso-Gonzalez A, Calaza M, Amigo J, González-Peñas J, Martínez-Regueiro R, Fernández-Prieto M, et al. Exploring the biological role of postzygotic and germinal de novo mutations in ASD. Sci Rep. 2021 Jan 11;11(1):319. The ENCODE Project Consortium, Abascal F, Acosta R, Addleman NJ, Adrian J, Afzal V, et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature. 2020 July 30;583(7818):699–710. Mitra I, Huang B, Mousavi N, Ma N, Lamkin M, Yanicky R, et al. Patterns of de novo tandem repeat mutations and their role in autism. Nature. 2021 Jan 14;589(7841):246–50. Grant CE, Bailey TL, Noble WS. FIMO: scanning for occurrences of a given motif. Bioinformatics. 2011 Apr 1;27(7):1017–8. Rauluseviciute I, Riudavets-Puig R, Blanc-Mathieu R, Castro-Mondragon JA, Ferenc K, Kumar V, et al. JASPAR 2024: 20th anniversary of the open-access database of transcription factor binding profiles. Nucleic Acids Res. 2024 Jan 5;52(D1):D174–82. O’Connor T, Grant CE, Bodén M, Bailey TL. T-Gene: improved target gene prediction. Wren J, editor. Bioinformatics. 2020 June 1;36(12):3902–4. Wamsley B, Bicks L, Cheng Y, Kawaguchi R, Quintero D, Margolis M, et al. Molecular cascades and cell type–specific signatures in ASD revealed by single-cell genomics. Science. 2024 May 24;384(6698):eadh2602. Gordon A, Yoon SJ, Tran SS, Makinson CD, Park JY, Andersen J, et al. Long-term maturation of human cortical organoids matches key early postnatal transitions. Nat Neurosci. 2021 Mar;24(3):331–42. Sun JX, Helgason A, Masson G, Ebenesersdóttir SS, Li H, Mallick S, et al. A direct characterization of human mutation based on microsatellites. Nat Genet. 2012 Oct;44(10):1161–5. Shih SC, Zukauskas A, Li D, Liu G, Ang LH, Nagy JA, et al. The L6 Protein TM4SF1 Is Critical for Endothelial Cell Function and Tumor Angiogenesis. Cancer Res. 2009 Apr 15;69(8):3272–7. Guo T, Nan Z, Miao C, Jin X, Yang W, Wang Z, et al. The autophagy-related gene Atg101 in Drosophila regulates both neuron and midgut homeostasis. J Biol Chem. 2019 Apr;294(14):5666–76. Wang T, Hoekzema K, Vecchio D, Wu H, Sulovari A, Coe BP, et al. Large-scale targeted sequencing identifies risk genes for neurodevelopmental disorders. Nat Commun. 2020 Oct 1;11(1):4932. Liu T, Li T, Ke S. Role of the CASZ1 transcription factor in tissue development and disease. Eur J Med Res. 2023 Dec 5;28(1):562. Kitajima H, Maruyama R, Niinuma T, Yamamoto E, Takasawa A, Takasawa K, et al. TM4SF1-AS1 inhibits apoptosis by promoting stress granule formation in cancer cells. Cell Death Dis. 2023 July 13;14(7):424. Tang Q, Chen J, Di Z, Yuan W, Zhou Z, Liu Z, et al. TM4SF1 promotes EMT and cancer stemness via the Wnt/β-catenin/SOX2 pathway in colorectal cancer. J Exp Clin Cancer Res. 2020 Dec;39(1):232. Cai Y, Ji Y, Liu Y, Zhang D, Gong Z, Li L, et al. Microglial circ-UBE2K exacerbates depression by regulating parental gene UBE2K via targeting HNRNPU. Theranostics. 2024;14(10):4058–75. Filatova EV, Shadrina MI, Alieva AKh, Kolacheva AA, Slominsky PA, Ugrumov MV. Expression analysis of genes of ubiquitin-proteasome protein degradation system in MPTP-induced mice models of early stages of Parkinson’s disease. Dokl Biochem Biophys. 2014 May;456(1):116–8. Meiklejohn H, Mostaid MS, Luza S, Mancuso SG, Kang D, Atherton S, et al. Blood and brain protein levels of ubiquitin-conjugating enzyme E2K (UBE2K) are elevated in individuals with schizophrenia. J Psychiatr Res. 2019 June;113:51–7. Li H, Zhao P, Xu Q, Shan S, Hu C, Qiu Z, et al. The autism-related gene SNRPN regulates cortical and spine development via controlling nuclear receptor Nr4a1. Sci Rep. 2016 July 19;6(1):29878. Okay K, Varış PÜ, Miral S, Ekinci B, Yaraş T, Karakülah G, et al. Alternative splicing and gene co-expression network-based analysis of dizygotic twins with autism-spectrum disorder and their parents. Genomics. 2021 July;113(4):2561–71. Allen M, Huang BS, Notaras MJ, Lodhi A, Barrio-Alonso E, Lituma PJ, et al. Astrocytes derived from ASD individuals alter behavior and destabilize neuronal activity through aberrant Ca2 + signaling. Mol Psychiatry. 2022 May;27(5):2470–84. Ha CM, Park D, Han JK, Jang J ill, Park JY, Hwang EM, et al. Calcyon Forms a Novel Ternary Complex with Dopamine D1 Receptor through PSD-95 Protein and Plays a Role in Dopamine Receptor Internalization. J Biol Chem. 2012 Sept;287(38):31813–22. Su P, Lai TKY, Lee FHF, Abela AR, Fletcher PJ, Liu F. Disruption of SynGAP–dopamine D1 receptor complexes alters actin and microtubule dynamics and impairs GABAergic interneuron migration. Sci Signal. 2019 Aug 6;12(593):eaau9122. Masnada S, Parazzini C, Bini P, Barbarini M, Alberti L, Valente M, et al. Phenotypic spectrum of short-chain enoyl-Coa hydratase-1 (ECHS1) deficiency. Eur J Paediatr Neurol. 2020 Sept;28:151–8. Muntean C, Tripon F, Bogliș A, Bănescu C. Pathogenic Biallelic Mutations in ECHS1 in a Case with Short-Chain Enoyl-CoA Hydratase (SCEH) Deficiency-Case Report and Literature Review. Int J Environ Res Public Health. 2022 Feb 13;19(4):2088. Sakai C, Yamaguchi S, Sasaki M, Miyamoto Y, Matsushima Y, Goto Y ichi. ECHS1 Mutations Cause Combined Respiratory Chain Deficiency Resulting in Leigh Syndrome. Hum Mutat. 2015 Feb;36(2):232–9. Olgiati S, Skorvanek M, Quadri M, Minneboo M, Graafland J, Breedveld GJ, et al. Paroxysmal exercise-induced dystonia within the phenotypic spectrum of ECHS1 deficiency: ECHS1 Mutations, Dystonia, and PED. Mov Disord. 2016 July;31(7):1041–8. Satterstrom FK, Kosmicki JA, Wang J, Breen MS, De Rubeis S, An JY, et al. Large-Scale Exome Sequencing Study Implicates Both Developmental and Functional Changes in the Neurobiology of Autism. Cell. 2020 Feb;180(3):568–584.e23. Yizhar O, Fenno LE, Prigge M, Schneider F, Davidson TJ, O’Shea DJ, et al. Neocortical excitation/inhibition balance in information processing and social dysfunction. Nature. 2011 Sept;477(7363):171–8. Chintalaphani SR, Pineda SS, Deveson IW, Kumar KR. An update on the neurological short tandem repeat expansion disorders and the emergence of long-read sequencing diagnostics. Acta Neuropathol Commun. 2021 Dec;9(1):98. Additional Declarations Yes Dr. Arango has been a consultant to or has received honoraria or grants from Abbot, Acadia, Ambrosetti, Angelini, Biogen, BMS, Boehringer, Carnot, Gedeon Richter, Janssen Cilag, Lundbeck, Medscape, Menarini, Minerva, Otsuka, Pfizer, Roche, Rovi, Sage, Servier, Shire, Schering Plough, Sumitomo Dainippon Pharma, Sunovion, Takeda and Teva. Supplementary Files SupplementaryTables1.xlsx Supplementary Tables Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Reject after peer review 13 Mar, 2026 Review # 1 received at journal 08 Mar, 2026 Review # 2 received at journal 05 Mar, 2026 Review # 3 received at journal 19 Feb, 2026 Reviewer # 3 agreed at journal 18 Feb, 2026 Reviewer # 2 agreed at journal 18 Feb, 2026 Reviewer # 1 agreed at journal 17 Feb, 2026 Reviewers invited by journal 09 Feb, 2026 Editor assigned by journal 04 Jan, 2026 Submission checks completed at journal 04 Jan, 2026 First submitted to journal 18 Dec, 2025 Unknown event 17 Dec, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8374597","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":588618879,"identity":"4f3e67c4-8b2b-4b31-8211-33f5bd72dfe1","order_by":0,"name":"Maria Cristina Rodriguez Fontenla","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABb0lEQVRIie2RMWvCQBiGv0NIlotdLyjmFxQuBKTQFv9KQkEXNxeHICmBc1G6KhT6F4RC6RgJxCXVVSiIImRphrhZqNLTxFAFOxeaB44kl3vu/e4+gIyMP4rDBwZAVvKZvIjx4wouzinCqYKd/S4EZOtMlvBzhyOFOkcLL9tvC2dtzorKk83ma5iV8u+u/YGbLcBidzRfvRJFm9YWEfCZmLJfo8OO18DUQ221Aw1NHhvsGvsCYDw27vs+UV+mdY2A7x4UpwqOZOmYCogRDLox8BHTJIahQuqqLTGCuCIAYofyypMAhhuuKAwx+StVtgTwTtkwUnnu1QJA27SwaRXcXQp4iBWSFHspWTRWeLQxIHqZ30QuVQJwix4vzDNYoUh1TeYp6JHPYOyr/S4jdz0/1IjupWeZVHOr0NQrij0K5LCpl/K+uIxCs1XCYodGn+zm9qHNbywyD4WdQOMWkX1zj9DPtDMhF/3+PyMjI+Of8Q1m2okNqYGu1AAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-9829-6393","institution":"University of Santiago de Compostela","correspondingAuthor":true,"prefix":"","firstName":"Maria","middleName":"Cristina Rodriguez","lastName":"Fontenla","suffix":""},{"id":588618880,"identity":"e23bf458-e71f-4aa4-acea-1403c56ae381","order_by":1,"name":"Pablo Carballo-Pacoret","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Pablo","middleName":"","lastName":"Carballo-Pacoret","suffix":""},{"id":588618881,"identity":"bec7effe-201b-47bd-9b39-57b36ba3f156","order_by":2,"name":"Sara Dominguez-Alonso","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Sara","middleName":"","lastName":"Dominguez-Alonso","suffix":""},{"id":588618882,"identity":"95fafdd0-d723-4865-b225-8c23501810a5","order_by":3,"name":"Javier Gonzalez-Peñas","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Javier","middleName":"","lastName":"Gonzalez-Peñas","suffix":""},{"id":588618883,"identity":"b04c3a96-52f3-4573-beb5-da3fc07bc4c8","order_by":4,"name":"Mara Parellada","email":"","orcid":"https://orcid.org/0000-0001-7977-3601","institution":"Hospital General Universitario Gregorio Marañón, IiSGM, CIBERSAM, School of Medicine, UCM","correspondingAuthor":false,"prefix":"","firstName":"Mara","middleName":"","lastName":"Parellada","suffix":""},{"id":588618884,"identity":"57f54569-0bee-403f-8424-6593d6e44178","order_by":5,"name":"Celso Arango","email":"","orcid":"https://orcid.org/0000-0003-3382-4754","institution":"Department of Chid Psychiatry. Hospital Universitario La Paz. IdiPAZ. School of Medicine. Universidad Autónoma de Madrid. CIBERSAM","correspondingAuthor":false,"prefix":"","firstName":"Celso","middleName":"","lastName":"Arango","suffix":""},{"id":588618885,"identity":"a77e8b06-1c1e-466b-8de5-b0a9f3c7b3f7","order_by":6,"name":"Angel Carracedo","email":"","orcid":"","institution":"Universidad de Santiago de Compostela (CIMUS) and SERGAS","correspondingAuthor":false,"prefix":"","firstName":"Angel","middleName":"","lastName":"Carracedo","suffix":""}],"badges":[],"createdAt":"2025-12-16 09:55:10","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8374597/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8374597/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":102597957,"identity":"a15235e3-4bff-4a72-9184-46a1d5bec0b6","added_by":"auto","created_at":"2026-02-13 12:26:57","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":261072,"visible":true,"origin":"","legend":"\u003cp\u003eTRs mutations detection in Spanish cohort workflow. Bam, bed (custom TR catalogue) and fasta files were used as input for variant calling with GangSTR. The vcf output was used as input for call and locus-level filters with DumpSTR, and the vcf output was used as input for de novo TR mutations detection. Then, a quality control and filters were applied, resulting in a filtered bed file with TRs mutations.\u003c/p\u003e","description":"","filename":"11.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8374597/v1/41692b552302859ef7736200.jpg"},{"id":102598057,"identity":"5eef45c6-35f5-4e93-925c-a97b5045495d","added_by":"auto","created_at":"2026-02-13 12:27:10","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":135782,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003ea)Bar plot displaying the ENCODE cCRE (candidate cis-regulatory elements) categories in which TR mutations were located in the Spanish cohort.b) Bar plot showing the distribution of de novo TRs by size, expressed as the number of repeat units gained or lost in the Spanish cohort; c) Bar plot showing the distribution of TR mutations according to motif length (e.g., mono-, di-, tri-nucleotide repeats...);d) Scatter plot illustrating the parental allele bias, with a tendency for shorter alleles to expand and longer alleles to contract\u003c/em\u003e\u003c/p\u003e","description":"","filename":"12.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8374597/v1/d46080e529d2c5a3825efe74.jpg"},{"id":102597767,"identity":"ac421f17-5b46-49b0-90cb-faa271f8beea","added_by":"auto","created_at":"2026-02-13 12:26:30","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":669209,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eIntegration of results from the Spanish cohort (genetic and functional annotations). From outside to inside: a) Sequenced cCREs; b) de novo TRs; c) ENCODE categories of cCREs; d) Genes with top pathogenicity scores; e) de novo TRs predicted to alter TF binding changes; f) SFARI genes; g) Genes differentially expressed between ASD cases and controls at scRNA-seq analysis.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"13.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8374597/v1/4ab2ac93bcf731b48e5306d5.jpg"},{"id":102597961,"identity":"5266255d-a07d-496e-a326-a8a0b5fdf2e4","added_by":"auto","created_at":"2026-02-13 12:26:58","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":440205,"visible":true,"origin":"","legend":"\u003cp\u003eDot plot showing the top 30 most significant biological processes related with both genes linked to Spanish TRs mutations and genes linked to TRs mutations from SSC cohort.\u003c/p\u003e","description":"","filename":"14.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8374597/v1/44258e410e59e1c95a16d1db.jpg"},{"id":102597839,"identity":"768f717c-d1bc-4105-bad4-881c20d9c8ac","added_by":"auto","created_at":"2026-02-13 12:26:38","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":306635,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eSingle-cell differential expression dot plot displaying genes on the Y-axis analyzed in the Spanish cohort : CALY, ECHS1, NR4A1, THRA, PHLDB2, UBE2K, KLHL32, PBX1, KTN1, GLI3, NRG1, C9orf3, FARP1 (expression values provided in Supplementary Material), and in the SSC cohort (remaining genes on the Y-axis). NRG1 is the most downregulated gene in the Spanish cohort, while PTGSE3 is the most upregulated associated gene in the Simons Simplex Collection. Only ECHS1 shows significant association in both cohorts. Dot size corresponds to statistical significance (-log10 p-value), and color intensity reflects the magnitude of fold change between ASD and control samples, with red indicating upregulation in ASD and blue indicating downregulation.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"15.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8374597/v1/2ddc7c5a01caeddfa373e2c3.jpg"},{"id":102597784,"identity":"5bf53e48-0b19-4273-b93e-623b3d5930bb","added_by":"auto","created_at":"2026-02-13 12:26:33","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":182075,"visible":true,"origin":"","legend":"\u003cp\u003ea) Region in chr10 with two mutated cCREs affecting regulation of CALY and ECHS1 genes (only mutated ccREs are shown). Yellow thunders represent Spanish cohort mutations, and red thunders SSC mutations. b) Developmental trajectories of ECHS1 and CALY expression in cortical organoids and BrainSpan. Transition from prenatal to post-natal stage with horizontal grey area/line.\u003c/p\u003e","description":"","filename":"16.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8374597/v1/19a5e2cfd4dd4f3fbb1f5bd4.jpg"},{"id":104781119,"identity":"03234dfd-57cf-470d-8e62-2dcb2c4cf376","added_by":"auto","created_at":"2026-03-17 07:54:51","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3327833,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8374597/v1/4066213d-edb6-4f03-b0f2-5c0909039b5e.pdf"},{"id":102598007,"identity":"68df2d0a-dca4-498b-876c-f42b7180856a","added_by":"auto","created_at":"2026-02-13 12:26:59","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":40923129,"visible":true,"origin":"","legend":"Supplementary Tables","description":"","filename":"SupplementaryTables1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-8374597/v1/0373b69e2cf7fc71feb29953.xlsx"}],"financialInterests":"\u003cb\u003eYes\u003c/b\u003e\nDr. Arango has been a consultant to or has received honoraria or grants from Abbot, Acadia, Ambrosetti, Angelini, Biogen, BMS, Boehringer, Carnot, Gedeon Richter, Janssen Cilag, Lundbeck, Medscape, Menarini, Minerva, Otsuka, Pfizer, Roche, Rovi, Sage, Servier, Shire, Schering Plough, Sumitomo Dainippon Pharma, Sunovion, Takeda and Teva.","formattedTitle":"Large-scale sequencing study of de novo regulatory Tandem Repeats (TRs) identifies new ASD (Autism Spectrum Disorders) candidate genes integrating gene expression mapping, brain scRNA-seq and organoid models.","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eAutism Spectrum Disorders (ASD) are neurodevelopmental disorders (NDDs), characterised by repetitive or restrictive behaviours, atypical social function and communication deficits (1,2). ASD have a high genetic component with an estimated heritability of around 80% (3). Although several autism-associated genes and variants have been identified through various genetic approaches\u0026mdash;including GWAS (4), WES (5), WGS (6), and integrative omic analyses such as single-cell RNA sequencing (scRNA-seq) (7), these studies still fall short of accounting for the full genetic contribution estimated in family and twin studies. Thus, a substantial fraction of the so-called \u0026ldquo;missing heritability\u0026rdquo; remains unexplained. Several factors may contribute to missing heritability in ASD, including sample heterogeneity, undetected ultra rare variants, variants lying within regulatory regions, somatic and postzygotic mutations and complex structural variation such as tandem repeats and mobile elements. Detecting many of these variants often requires specialized sequencing technologies and bioinformatic tools.\u003c/p\u003e \u003cp\u003eThus, de novo and inherited single-base mutations, small insertions and deletions (indels) or large copy number variations (CNVs) can be identified under conventional sequencing and routine genetic diagnosis screening procedures. However, part of this missing heritability may stem from the fact that current bioinformatics pipelines often discard information from repetitive regions which are challenging to interpret using short-read sequencing methods (8,9).\u003c/p\u003e \u003cp\u003eTandem repeats (TRs) are sequences of bp motifs that are repeated consecutively a specific number of times in the DNA sequence. They vary both in the sequence of the repeated motif and in the number of repeat units. Motif repeats shorter than 100 bp represent approximately 3% of the human genome (10), with over a million known TRs distributed throughout the genome.\u003c/p\u003e \u003cp\u003eThe classification of TRs has always been diverse, as they are usually classified according to their length (expansion or contraction of the repeat motif), location and/or frequency of occurrence: microsatellites (STRs) with repeat motifs shorter than 5 bp, are usually present throughout the genome. Minisatellites, with repeat motifs ranging from 5 bp to 100 bp, are less common. Satellite DNA, found at centromeres, consists of 6 bp motifs repeated between 300 and 8,000 times, while tandem repeats also consisting of 6 bp motifs, are repeated between 2,000 and 50,000 times at the ends of chromosomes. In addition, motifs longer than 5bp, often called Variant Number Tandem Repeats (VNTRs), have also been studied for their role in human traits (11).\u003c/p\u003e \u003cp\u003eAs we can see, the classification of TRs is complex. They are found in many parts of the genome, although they are predominantly located within non-coding regions. TRs accelerate evolution and adaptation due to their mutation rates, which are between 10 and 100,000 times higher than those of the rest of the genome (12). These high mutation rates create new copy lengths and copy numbers, contributing to genomic diversity and potentially giving rise to novel traits. In addition, TRs are involved in cellular processes such as cell cycle stability, telomere maintenance, or the regulation of gene expression. They are known to contribute to disease etiology by altering TF binding and three-dimensional chromatin structure, affecting differential gene expression (13).\u003c/p\u003e \u003cp\u003eFor example, well-known repeats of TRs lead to diseases, such as the CAG repeat motif, which causes Huntington's Disease (14), as well as other genetic disorders (15). In addition to autism, where the expansion of the CGG motif in the \u003cem\u003eFMR1\u003c/em\u003e gene is a direct cause of Fragile X syndrome, a condition closely related to ASD (16), TRs are linked to several other diseases of the nervous system, such as ataxia, schizophrenia, amyotrophic lateral sclerosis (ALS) or dementia (17\u0026ndash;20).\u003c/p\u003e \u003cp\u003eIt is known that the number of copies of TRs (21) and structural variations are relevant in the aetiology of ASD (22,23), as it is also known that there is an increased variation in TRs in regulatory regions (11). There are studies that have investigated genome-wide tandem repeats (24,25), and others who have studied mutations in regulatory regions (26), underscoring the need for future research on the identification of regulatory target genes and neuronal populations involved using single-cell resolution, which has not yet been addressed.\u003c/p\u003e \u003cp\u003eIn this study, we set out to investigate the role of regulatory de novo tandem repeats (TRs) in ASD risk by integrating genomic, epigenomic, and transcriptomic data. Specifically, our objectives were: (1) to identify de novo TR mutations within active cis-regulatory elements (cCREs) in a Spanish ASD cohort of 200 ASD trios. and validate their contribution using large-scale data from the SSC (1,637 ASD simplex quad families); and (2) to functionally prioritize these variants through multi-layered annotation, including predicted disruption of transcription factor binding, expression-based gene mapping, pathogenicity scoring, biological pathways, single-cell transcriptomic signatures from the ASD human brain vs controls and expression patterns of target genes in cortical brain organoids.\u003c/p\u003e \u003cp\u003eThis approach aims to uncover novel ASD risk genes within the regulatory genome and to highlight the contribution of complex regulatory variation that is often missed by conventional sequencing analyses.\u003c/p\u003e"},{"header":"MATERIAL AND METHODS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eSpanish Cohort\u003c/h2\u003e \u003cp\u003eThe analysis presented in this study was conducted on a subset of 200 trios from the Spanish ASD cohort described in Alonso-Gonzalez et al. 2021 (27), which comprises 360 trios, each consisting of unaffected parents and an affected proband, from Fundaci\u0026oacute;n P\u0026uacute;blica de Medicina Xen\u0026oacute;mica, Santiago de Compostela and Hospital Universitario Gregorio Mara\u0026ntilde;on, Madrid, Spain. Patients involved in the study were diagnosed by psychiatrists or neurologists, following the criteria of both the Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition Text Revision (DSM-IV-TR) and Fifth Edition (DSM-5). When deemed necessary, the Autism Diagnostic Observation Schedule (ADOS) and the Autism Diagnostic Interview-Revised (ADI-R) were also administered. Patients with syndromic autism were excluded. All participants (probands, parents or legal representatives) gave their written consent and the study was conducted under the Declaration of Helsinki. The Galician Committee of Research Ethics (Xunta de Galicia) has approved this study under registration number 2020/400.\u003c/p\u003e \u003cp\u003eFrom the 360 trios forming the Spanish cohort with exome data, Copy Number Variation (CNV) array and phenotypic information, 200 trios were selected among those with negative results in the CNV array, following the procedure outlined in Alonso-Gonzalez et al., 2021 (27).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eSelection of regulatory regions\u003c/h3\u003e\n\u003cp\u003eTargeted sequencing (see next section) has been carried out on cis-regulatory regions (cCREs) from the ENCODE project (28) (version 2, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://screenv2.wenglab.org/\u003c/span\u003e\u003cspan address=\"https://screenv2.wenglab.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e, corresponding to approximately 1% of the total human genome, classified in 5 different categories: promoter-like signatures (PLS), proximal enhancer-like signatures (pELS), distal enhancer-like signatures (dELS), high DNase and H3K4me3 signals only (DNase\u0026ndash;H3K4me3) and high DNase and CTCF signals only (CTCF-only). Brain and GI tissues with DNAse-Seq data available were selected, both from adult (n\u0026thinsp;=\u0026thinsp;14) and embryonic tissue (n\u0026thinsp;=\u0026thinsp;59). For each tissue, cCREs labeled as \u0026ldquo;Low-Dnase\u0026rdquo; were excluded, as they are inactive in the given tissue. From the total of 926,535 human cCREs, we selected those that exhibit activity in a greater number of the interrogated tissues. Thus, we selected cCREs that were active in 36 or more tissues (n\u0026thinsp;=\u0026thinsp;85,394 cCREs).\u003c/p\u003e\n\u003ch3\u003eSequencing of regulatory regions\u003c/h3\u003e\n\u003cp\u003eTargeted sequencing was done at the National Center for Genomic Analysis (CNAG) using the KAPA HyperChoice Target Enrichment custom probes and with an average coverage of 30\u0026sdot;. Samples with sex discrepancies when compared to reported pedigrees were dropped and replaced, along with all other samples from the same family. Moreover, samples which failed CNAG\u0026rsquo;s quality control (DNA\u0026thinsp;\u0026lt;\u0026thinsp;100 ng / critical degradation (genomic quality number (GQN)\u0026thinsp;\u0026lt;\u0026thinsp;3.3)) were also removed and consequently substituted, leaving 71 trios from Santiago and 129 from Madrid. Sequencing reads were aligned to GRCh38/hg38 using the Burrows-Wheeler Aligner. Raw CRAM and vcf sequencing files have been transferred, stored and handled at the CESGA, Centro de Supercomputaci\u0026oacute;n de Galicia (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.cesga.es\u003c/span\u003e\u003cspan address=\"https://www.cesga.es\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Sequencing data are available at: EGA repository number: EGAS50000001395 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ega-archive.org/studies/EGAS50000001395\u003c/span\u003e\u003cspan address=\"https://ega-archive.org/studies/EGAS50000001395\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n\u003ch3\u003eSimons Simplex Collection cohort\u003c/h3\u003e\n\u003cp\u003eTo carry out a gene-level meta-analysis and a single-cell mapping analysis, we requested the STR dataset from the Simons Simplex Collection via the Simons Foundation Autism Research Initiative (SFARI) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://base.sfari.org\u003c/span\u003e\u003cspan address=\"https://base.sfari.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e We retrieved a CSV file containing detailed information on 175,291 mutations identified in 1,593 quad families (also used in Mitra et al.(29)). Genomic positions for STRs were overlapped with the cCREs analysed in our study using bedtools (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/arq5x/bedtools2\u003c/span\u003e\u003cspan address=\"https://github.com/arq5x/bedtools2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e, creating a new custom file only with cCREs that we had sequenced that had mutations in SSC cohort.\u003c/p\u003e \u003cp\u003eThe subsequent gene annotation was done with this file as input. It should be noted that the TRs workflow applied in our cohort and in Mitra et al. is largely identical ensuring a high degree of methodological homogeneity.\u003c/p\u003e\n\u003ch3\u003eTandem Repeat (TR) Mutation Detection Workflow\u003c/h3\u003e\n\u003cp\u003eThe workflow (including file types, tools used, input and output formats, applied filters, and quality control steps performed in the analysis of the Spanish cohort), is summarized in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Custom TR catalogue elaboration and file transformation, calling of TRs, call and locus-level filters of variants and \u003cem\u003ein-silico\u003c/em\u003e detection of TRs de novo variants, are found in \u003cem\u003eSupplementary Information\u003c/em\u003e. A bed file with 107 mutations remained (\u003cem\u003eSupplementary Table\u0026nbsp;1\u003c/em\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003ePrioritizing pathogenic TR mutations\u003c/h2\u003e \u003cp\u003e \u003cem\u003eSISTR (Selection Inference at Short TRs)\u003c/em\u003e method described in Mitra et al., considers mutation, genetic drift and negative natural selection to calculate selection coefficients for each TRs. The parameter \u0026lsquo;\u003cem\u003es\u003c/em\u003e\u0026rsquo; can be interpreted as the reduction in reproductive fitness for each repeat unit copy number, relative to the TR model allele in the population.\u003c/p\u003e \u003cp\u003eTo estimate selection coefficients in the overlapping cCREs- TRs with the \u003cem\u003eGangSTR\u003c/em\u003e catalogue, SISTR was run on the parental cohort (400 individuals).\u003c/p\u003e \u003cp\u003ePathogenicity coefficients, which are allele-specific selection coefficients, have been calculated as follows: \u003cem\u003e(|a - opt|)s\u003c/em\u003e, where \u003cem\u003ea\u003c/em\u003e is the number of repeats for the de novo allele, \u003cem\u003eopt\u003c/em\u003e is the optimal (or modal) number of repeats for that TR, and \u003cem\u003es\u003c/em\u003e is the selection coefficient for the TR calculated by SISTR in our parental cohort.\u003c/p\u003e \u003cp\u003eThe input for SISTR was carried out using the statSTR tool from TR Tools package (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/gymrek-lab/TRTools\u003c/span\u003e\u003cspan address=\"https://github.com/gymrek-lab/TRTools\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e A single GangSTR vcf file was created, and statSTR was run with the filters --acount and --numcalled, to extract allele frequencies and total alleles. The resulting output was then merged with the GangSTR TR catalogue to obtain the input for running SISTR on the parental cohort.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eFunctional Annotation of TF Binding Sites\u003c/h3\u003e\n\u003cp\u003eTo determine whether de novo TR mutations disrupt transcription factor binding sites (TFBS) in the Spanish cohort, we have run the \u003cem\u003eFIMO\u003c/em\u003e tool (30) (also from MEME Suite), designed to scan nucleotide sequences, to search for TF binding motifs. We selected the binding motifs of vertebrates from the JASPAR database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://jaspar.elixir.no/downloads/\u003c/span\u003e\u003cspan address=\"https://jaspar.elixir.no/downloads/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e (31), that provides a catalog of TF motifs, obtained from validated experiments such as ChIP-Seq.\u003c/p\u003e \u003cp\u003eTwo fasta files included sequences centered on the TR loci within or near SFARI genes, with one file representing the reference genome and the other incorporating the corresponding d\u003cem\u003ee novo\u003c/em\u003e TR mutation sequences were generated. To this end, we have extended the mutated regions by \u0026plusmn;15 bp to improve possible TF binding to overlapping regions between the flanking region and the mutation itself.\u003c/p\u003e\n\u003ch3\u003eIntegrative genetic analysis of de novo TRs from SSC collection and Spanish cohort\u003c/h3\u003e\n\u003cp\u003eWe aim to conduct a joint meta-analysis of microsatellite datasets from the Simons Simplex Collection (SSC) (25), (downloaded from base.sfari.org) and our Spanish cohort to identify overlapping genomic regions associated with ASD. The methods used in that study ensure homogeneity in the detection and annotation of TRs across both cohorts, allowing for reliable integration with our dataset.\u003c/p\u003e \u003cp\u003eIntegrating these datasets increase statistical power, allow us to assess the consistency of observed signals across cohorts, and strengthen the robustness of candidate loci with potential functional relevance. Thus, a total of 200 trios of Spanish cohort\u0026thinsp;+\u0026thinsp;1,637 quads from SSC and 107 de novo TRs mutations\u0026thinsp;+\u0026thinsp;1305 from SSC were analysed. Overlapping cCREs between both cohorts were identified using \u003cem\u003ebedtools intersect\u003c/em\u003e.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eGene Mapping of Regulatory TRs using physical distance and expression data\u003c/h2\u003e \u003cp\u003eTo identify target genes based on genomic physical distance, we used the -closest function from the bedtools suite (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://bedtools.readthedocs.io/en/latest/content/tools/closest.html\u003c/span\u003e\u003cspan address=\"https://bedtools.readthedocs.io/en/latest/content/tools/closest.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e\u003c/p\u003e \u003cp\u003eTo identify target genes for TRs within cCREs, we used the T-Gene tool (32), which infers regulatory relationships by combining genomic closeness and ENCODE ChIP-seq expression data (hg19).\u003c/p\u003e \u003cp\u003eAs input for both T-Gene and bedtools, we used the 57 cCREs carrying the 107 high-confidence de novo regulatory TRs from the Spanish cohort (\u003cem\u003eSupplementary Table\u0026nbsp;2\u003c/em\u003e). The same was done with the SSC dataset (565 cCREs carrying 1305 mutations, \u003cem\u003eSupplementary Table\u0026nbsp;3\u003c/em\u003e). A GTF annotation file (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.gencodegenes.org/human/release_7.html\u003c/span\u003e\u003cspan address=\"https://www.gencodegenes.org/human/release_7.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e corresponding to the hg19 genome assembly, was used as input in bedtools to ensure consistency between both analysis physical and expression.\u003c/p\u003e \u003cp\u003eTo match this reference genome, TR coordinates from both the Spanish cohort and the SSC dataset from Mitra et al. were converted to hg19 using UCSC Table Browser (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://genome.ucsc.edu/cgi-bin/hgTables\u003c/span\u003e\u003cspan address=\"https://genome.ucsc.edu/cgi-bin/hgTables\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e\u003c/p\u003e \u003cp\u003eWe customized T-Gene output to include all potential regulatory links by selecting the \u0026ldquo;Link\u0026rdquo; option, which avoids discarding cCREs that may share target genes. The output of T-gene containing regulatory genes for both cohorts is in \u003cem\u003eSupplementary Table\u0026nbsp;4\u003c/em\u003e (Spanish cohort) and \u003cem\u003e5\u003c/em\u003e (SSC cohort). For the Spanish cohort, genes were ordered to each analysed cCRE (\u003cem\u003eSupplementary Table\u0026nbsp;6\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eSubsequently, gene lists were filtered considering Correlation and Distance p-value (CnD p-value)\u0026thinsp;\u0026lt;\u0026thinsp;0.05 of T-Gene. These filtered gene expression lists were used for distance vs expression analysis, GO enrichment and single-cell analysis.\u003c/p\u003e \u003cp\u003eFor distance vs expression gene comparison, we filtered T-gene output taking in account only the most linked gene (lowest CnD p-value) for each cCRE for both cohorts.\u003c/p\u003e \u003cp\u003eIn the Spanish cohort, we looked for genes related with SFARI database, to retain genes highly linked with ASD.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eGO enrichment analysis of TR associated genes\u003c/h2\u003e \u003cp\u003eGene Ontology enrichment analysis has been carried out on two platforms: Enrichr (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://maayanlab.cloud/Enrichr/\u003c/span\u003e\u003cspan address=\"https://maayanlab.cloud/Enrichr/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e and ReVIGO (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://revigo.irb.hr\u003c/span\u003e\u003cspan address=\"http://revigo.irb.hr\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e Enrichr is an online tool for enrichment analysis of gene sets; it was used to identify biological processes associated with genes related to our regulatory TRs by expression (T-Gene).\u003c/p\u003e \u003cp\u003eTo carry out the GO enrichment analysis in the Spanish cohort, the output from T-Gene (\u003cem\u003eSupplementary Table\u0026nbsp;4\u003c/em\u003e) was filtered to retain only those links with a CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05. The same was done for SSC cohort.\u003c/p\u003e \u003cp\u003eTo carry out the GO enrichment analysis in both cohorts, we retain genes from T-gene output (\u003cem\u003eSupplementary Tables\u0026nbsp;4 and 5\u003c/em\u003e) with CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.5 without duplicated genes shared between cohorts.\u003c/p\u003e \u003cp\u003eFor both cohorts and the metaanalysis with data from Mitra et al., we applied two thresholds: q-value\u0026thinsp;\u0026lt;\u0026thinsp;0.1 and Bonferroni correction for multiple tests. Moreover, we have selected the top 30 Enrichr-enriched GO biological processes based on their p-value. Visualization was further refined using a custom R script adapted from ReVIGO.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eSingle-Cell Gene Differential Gene Expression Analysis\u003c/h2\u003e \u003cp\u003eIn order to carry out a single cell differential gene expression analysis in neuronal and glial cells, we used ASD scGENE Portal (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://solo.bmap.ucla.edu/asdscgene/\u003c/span\u003e\u003cspan address=\"http://solo.bmap.ucla.edu/asdscgene/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e, resource that uses RNA-seq cell data from brain (33 ASD individuals and 31 control subjects) (33), comprising nearly 600,000 cell nuclei. We used this resource to explore whether any of the significant associated genes from both cohorts were expressed in specific neuronal or glial subtypes, and whether their expression was altered in ASD compared to controls.\u003c/p\u003e \u003cp\u003eThere were 35 cell types identified by single-nucleus RNA sequencing (snRNA-seq), divided in 8 major cell types: oligodendrocyte progenitor cells (OPCs); astrocyte (ASTRO); microglia (MG); endothelial cells (ENDO); inhibitory neurons (INT); excitatory neurons (EXT); oligodendrocyte (ODC) and blood-brain barrier (BBB).\u003c/p\u003e \u003cp\u003eTo carry out this analysis in the Spanish cohort, we retain genes from the T-gene output (\u003cem\u003eSupplementary Table\u0026nbsp;4\u003c/em\u003e) with a CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05, and we looked for genes showing significant differential expression between ASD cases and controls defined by an FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05, and an absolute log fold change (|logFC|)\u0026thinsp;\u0026gt;\u0026thinsp;0.3.\u003c/p\u003e \u003cp\u003eFor the dataset from Mitra et al., due to the larger number of candidate genes, we applied a slightly more stringent initial filter, selecting genes from the T-gene output (\u003cem\u003eSupplementary Table\u0026nbsp;5\u003c/em\u003e) with a CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 and a q-value\u0026thinsp;\u0026lt;\u0026thinsp;0.1. These genes were then queried in the single-cell portal using the same differential expression thresholds as described above.\u003c/p\u003e \u003cp\u003e \u003cb\u003eGene expression trajectories across the differentiation of human cortical organoids in vitro and in vivo human brain development.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo explore the developmental trajectories of candidate gene expression in BrainSpan and brain cortical organoids, we used the Gene Expression tool (GECO, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://solo.bmap.ucla.edu/shiny/GECO/#\u003c/span\u003e\u003cspan address=\"http://solo.bmap.ucla.edu/shiny/GECO/#\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), developed based on Gordon et al. (34).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eVisualization and Statistical Analysis Tool\u003c/h2\u003e \u003cp\u003eRStudio v2024.12.0\u0026thinsp;+\u0026thinsp;467 was used to create the graphs presented, and the \u003cem\u003ecirclize\u003c/em\u003e library for RStudio was used for customized circo plot. Microsoft Excel 16.97 was used for the directionality bias graph.\u003c/p\u003e \u003c/div\u003e"},{"header":"RESULTS","content":"\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eSPANISH COHORT\u003c/h2\u003e \u003cdiv id=\"Sec17\" class=\"Section3\"\u003e \u003ch2\u003eIdentification, Distribution, and Patterns of Regulatory de novo Tandem Repeats (TRs)\u003c/h2\u003e \u003cp\u003eMonSTR identified 1,067 de novo non-strict de novo TRs mutations in the 200 families analysed from the Spanish cohort (\u003cem\u003eSupplementary Table\u0026nbsp;7\u003c/em\u003e). After QC, 3 children and 6 TR loci were discarded and a list of 821 non-strict de novo TR mutations without outliers was obtained. After MonSTR prioritization, a total of 107 strict de novo TR mutations remained in 57 different cCREs, 83 different probands, out of the 200 analysed (after using GangSTR and MonSTR pipeline. See Methods section and \u003cem\u003eSupplementary Table\u0026nbsp;1\u003c/em\u003e). The mean number of mutations per proband was 0.51, which is consistent with the genome-wide average of 53.9 mutations per individual reported by Mitra et al., taking in account that our analysis covers approximately 1% of the genome.\u003c/p\u003e \u003cp\u003eAnalysing candidate cis-regulatory elements (cCREs), we found an enrichment in one particular cCRE category: proximal enhancer-like signature (pELS), with a total of 38 mutations out of 107 mutations (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eRegarding the patterns of expansion, the most frequent pattern in the strict de novo TRs, involved the motif being repeated just once more than in the parent sequence. In contrast, the most common contraction mutation pattern was the deletion of three consecutive copies of the motif (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003eWhen analyzing the motif size of TRs in our cohort, we found that 38 out of the 107 mutations occurred in TRs with 2 bp motifs, making this the most frequent motif size in our dataset (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec).\u003c/p\u003e \u003cp\u003eIn our study, a significant bias towards expansions over contractions in de novo TRs was found. Specifically, 81 out of 107 (76%) were expansions while 26/107 (24%) were contractions (binomial test, n\u0026thinsp;=\u0026thinsp;107, k\u0026thinsp;=\u0026thinsp;81; p\u0026thinsp;=\u0026thinsp;9.37x10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e) (\u003cem\u003eSee Supplementary Table\u0026nbsp;1\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eIn addition, there is a directionality bias in mutation size that is observed in other studies as well. When the parental allele is short, the mutation tends to expand, and when the parental allele is large, it tends to contract. (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed).\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003ePrioritizing pathogenic TR mutations\u003c/h2\u003e \u003cp\u003eDue to SISTR\u0026rsquo;s current limitation to repeat motifs of 2\u0026ndash;4 base pairs, not all regions could be analyzed. Of the total regions considered, 7,561 regions were analysed using SISTR. 134 regions (1.8% could not be analysed, indicating that SISTR may assume certain patterns in the TRs that are not fulfilled at these loci). As a result, selection coefficients (\u0026lsquo;s\u0026rsquo;) were successfully estimated for 7,427 loci (\u003cem\u003eSupplementary Table\u0026nbsp;8\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eOf the 107 de novo TR mutations identified, SISTR was able to calculate pathogenicity scores for 66. Among these, 13 mutations in 16 probands were in the top quartile of pathogenicity (score\u0026thinsp;\u0026gt;\u0026thinsp;0.002655), and one mutation was linked with a SFARI gene (see Gene Mapping in the next section), \u003cem\u003eCASZ1\u003c/em\u003e (See Supplementary \u003cem\u003eTable\u0026nbsp;9\u003c/em\u003e). Detailed information for all 66 mutations analyzed with SISTR is provided in \u003cem\u003eSupplementary Table\u0026nbsp;10\u003c/em\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eGene Mapping of Regulatory TRs by physical distance and expression data\u003c/h2\u003e \u003cp\u003eAs we mentioned in Material and Methods section, 57 cCREs with TRs mutations in the Spanish cohort were associated with their putative target genes using expression data of T-Gene (\u003cem\u003eSupplementary Table\u0026nbsp;4\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eFor the distance vs expression comparison, T-gene links with CnD p-value\u0026thinsp;\u0026gt;\u0026thinsp;0.05 were discarded, so 53 cCREs finally remain, because 4 cCREs did not have any link with CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 (see Material and Methods and \u003cem\u003eSupplementary Table\u0026nbsp;11\u003c/em\u003e). Closest genes by physical distance using bedtools are listed in \u003cem\u003eSupplementary Table\u0026nbsp;12\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eThese 53 cCREs were used to compare whether the regulated gene was the closest physically or not. In nearly half of the cCREs (21 out of 53; 39.6%), the gene most significantly associated with the regulatory region was not the closest gene in terms of physical distance (\u003cem\u003eSee Supplementary Table\u0026nbsp;13\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eMoreover, in the Spanish cohort, 9 mutations affecting 8 different cCREs were linked to genes listed in the Simons Foundation Autism Research Initiative (SFARI) database (Banerjee-Basu \u0026amp; Packer, 2010), the most comprehensive and reliable resource for ASD-associated genes (see \u003cem\u003eSupplementary Table\u0026nbsp;14\u003c/em\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eFunctional Annotation of TF (Transcription Factor) Binding Sites\u003c/h2\u003e \u003cp\u003eThe results of looking at the cCREs with a linked SFARI gene have shown that in most cases (80%), TFs predicted to have the strongest binding to the mutated regions are the same as those predicted to bind the unmutated regions. The majority of these top-binding TFs belong to the \u003cem\u003eZNF (zinc finger)\u003c/em\u003e family.\u003c/p\u003e \u003cp\u003eHowever, we identified two \u003cem\u003ede novo\u003c/em\u003e TRs that result in a change in the TF with the highest predicted binding affinity. The first TR affects TF binding by replacing \u003cem\u003eKLF9\u003c/em\u003e with \u003cem\u003eNrf1\u003c/em\u003e, a gene previously studied in the context of neurodevelopmental disorders which belongs to a family of TFs previously related to neuronal functions. In the second case, a de novo TR alters binding from \u003cem\u003eZNF384\u003c/em\u003e to \u003cem\u003eEWSR1-FLI1\u003c/em\u003e (\u003cem\u003eSupplementary Table 15\u003c/em\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eGO enrichment analysis of TR associated genes\u003c/h2\u003e \u003cp\u003eGO enrichment analysis was performed using the genes linked to cCREs with a CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 identified in the T-Gene output for Spanish cohort (n\u0026thinsp;=\u0026thinsp;122) (See \u003cem\u003eSupplementary Table\u0026nbsp;16\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eNo GO term was significant (q-value\u0026thinsp;\u0026lt;\u0026thinsp;0.1). However, when applying the Bonferroni correction for multiple testing (adjusted significance threshold\u0026thinsp;=\u0026thinsp;0.05/121\u0026thinsp;\u0026asymp;\u0026thinsp;0.00041), two GO terms were significantly enriched: \u003cem\u003eRegulation of Cell Differentiation\u003c/em\u003e (GO:0045595, p-value\u0026thinsp;=\u0026thinsp;0.00024) and \u003cem\u003eNegative Regulation of Astrocyte Differentiation\u003c/em\u003e (GO:0048712, p-value\u0026thinsp;=\u0026thinsp;0.00036) (\u003cem\u003eSupplementary Table\u0026nbsp;17\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eThe top 30 biological processes by p-value from GO enrichment analysis of Enrichr, using ReVIGO for visualizing Spanish cohort genes are shown in \u003cem\u003eSupplementary Fig.\u0026nbsp;1\u003c/em\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eSingle-Cell Gene Differential Gene Expression Analysis\u003c/h2\u003e \u003cp\u003escRNA-seq analysis in the Spanish cohort was done with the 122 genes linked to cCRE with CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05, the same used for the GO analysis (\u003cem\u003eSupplementary Table\u0026nbsp;16\u003c/em\u003e). Single cell gene differential gene expression analysis revealed differential expression of several genes (\u003cem\u003eSupplementary Table\u0026nbsp;18\u003c/em\u003e). Among them, \u003cem\u003eECHS1,CALY, KTN1\u003c/em\u003e and \u003cem\u003eTHRA\u003c/em\u003e were overexpressed in cases, while \u003cem\u003eNR4A1, PHLDB2, UBSK, KLHL32, PBX1, GLI3, NRG1 and FARP1\u003c/em\u003e were underexpressed in cases. \u003cem\u003eC9orf3\u003c/em\u003e was underexpressed in cases in oligodendrocytes and overexpressed in excitatory neurons. \u003cem\u003eNRG1\u003c/em\u003e shows the greatest degree of underexpression in cases from the Spanish cohort ( \u003cem\u003elogFC=-\u003c/em\u003e1,016; \u003cem\u003eFDR\u0026thinsp;=\u003c/em\u003e\u0026thinsp;5.69\u003csup\u003ee\u0026thinsp;\u0026minus;\u0026thinsp;14\u003c/sup\u003e; \u003cem\u003ecell type\u0026thinsp;=\u0026thinsp;EXT_9_L6)\u003c/em\u003e.\u003c/p\u003e \u003cp\u003e \u003cem\u003eGLI3\u003c/em\u003e was underexpressed in astrocytes \u003cem\u003e(logFC=-0.397, FDR\u0026thinsp;=\u0026thinsp;1.26 x 10\u003c/em\u003e\u003csup\u003e\u003cem\u003e\u0026minus;\u0026thinsp;4\u003c/em\u003e\u003c/sup\u003e\u003cem\u003e), THRA\u003c/em\u003e was overexpressed in excitatory neurons \u003cem\u003e(logFC\u0026thinsp;=\u0026thinsp;0.397, FDR\u0026thinsp;=\u0026thinsp;1.97 x 10\u003c/em\u003e\u003csup\u003e\u003cem\u003e\u0026minus;\u0026thinsp;3\u003c/em\u003e\u003c/sup\u003e) and \u003cem\u003ePBX1\u003c/em\u003e was underexpressed in several cell types, being astrocytes the most underexpressed cell type (logFC=-0.484, FDR\u0026thinsp;=\u0026thinsp;\u003cem\u003e5.33 x 10\u003c/em\u003e\u003csup\u003e\u003cem\u003e\u0026minus;\u0026thinsp;5\u003c/em\u003e\u003c/sup\u003e\u003cem\u003e).\u003c/em\u003e These three genes met the thresholds for differential expression and are present in the SFARI database (\u003cem\u003eSupplementary Table\u0026nbsp;5\u003c/em\u003e). Genes that did not met these thresholds were excluded in subsequent analysis.\u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003eIntegration of results in the Spanish cohort\u003c/h2\u003e \u003cp\u003eThis integrative circoplot (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003e) highlights the convergence of de novo TRs mutations, regulatory genomic context (non coding regions), and gene expression signatures at the single-cell level. Notably, several high-priority genes (e.g., \u003cem\u003eECHS1, CASZ1, PBX1)\u003c/em\u003e appear in multiple layers of the analysis, suggesting that they are strong candidates supported at genetic and functional annotation levels. Furthermore, predicted TF binding alterations driven by TR mutations, offering a potential mechanistic link between non-coding variation and transcriptional dysregulation were also represented. The integration of single-cell transcriptomic data strengthens the relevance of these findings by providing cell-type-resolved evidence of dysregulation. It is important to remark that \u003cem\u003eECHS1\u003c/em\u003e is not only affected by TR mutations in regulatory regions, but also shows altered expression in ASD cases versus controls, reinforcing its candidacy as a neurodevelopmental risk gene.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003eINTEGRATIVE ANALYSIS : SIMONS SIMPLEX COLLECTION AND THE SPANISH COHORT\u003c/h2\u003e \u003cdiv id=\"Sec25\" class=\"Section3\"\u003e \u003ch2\u003eIntegrative genetic analysis of de novo TRs from SSC collection and Spanish cohort\u003c/h2\u003e \u003cp\u003eSeven cCREs were overlapped between both cohorts considering that Mitra et al. 2021 carried out a WGS analysis (\u003cem\u003eSupplementary Table\u0026nbsp;19\u003c/em\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section3\"\u003e \u003ch2\u003eGene Mapping of Regulatory TRs by physical distance and expression data\u003c/h2\u003e \u003cp\u003eAs we mentioned in Material and Methods section, 565 cCREs with TRs mutations in the SSC cohort were associated with their putative target genes using expression data of T-Gene (\u003cem\u003eSupplementary Table\u0026nbsp;5\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eFor the distance vs expression comparison, T-gene links with CnD p-value\u0026thinsp;\u0026gt;\u0026thinsp;0.05 were discarded, so 487 cCREs finally remain, because 78 cCREs did not have any link with CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 (see Material and Methods and \u003cem\u003eSupplementary Table\u0026nbsp;20\u003c/em\u003e). Closest genes by physical distance using bedtools are listed in \u003cem\u003eSupplementary Table\u0026nbsp;21\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eThese 487 cCREs were used to compare whether the regulated gene was the closest physically or not. Among the most significantly associated genes for these 487 cCREs, 162 (33.3%) were not the nearest gene based on physical distance (\u003cem\u003eSupplementary Table\u0026nbsp;22\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eIf we remain with the most significant gene linked to each cCRE both of our cohort and from Mitra et al., 179/530 of the genes weren't the closest to the cCRE (33.8%).\u003c/p\u003e \u003cp\u003eTen regulatory genes were shared between the Spanish cohort and the SSC collection, taking in account the most significant genes linked with cCREs with CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05): \u003cem\u003eECHS1, C9orf3, C9orf72, FARP1, KIAA0922, PBX1, SOCS3, UBE2K, THNSL2\u003c/em\u003e and \u003cem\u003eWIPF2\u003c/em\u003e (\u003cem\u003eSupplementary Table\u0026nbsp;23\u003c/em\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section3\"\u003e \u003ch2\u003eGO enrichment analysis of TR associated genes\u003c/h2\u003e \u003cp\u003eAs we commented in the Spanish cohort, GO enrichment analysis was performed using the genes linked to cCREs with a CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 identified in the T-Gene output of SSC cohort (n\u0026thinsp;=\u0026thinsp;1125) (\u003cem\u003eSupplementary Table\u0026nbsp;24\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eIn the SSC cohort, no GO term was significant (q-value\u0026thinsp;\u0026lt;\u0026thinsp;0.1). When applying the Bonferroni correction for multiple testing (adjusted significance threshold\u0026thinsp;=\u0026thinsp;0.05/3268\u0026thinsp;\u0026asymp;\u0026thinsp;0.00002), no GO term was significant either (\u003cem\u003eSupplementary Table\u0026nbsp;25\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eWhen the Spanish cohort was meta-analysed together with the SSC data, without duplicated genes shared between cohorts (n = 1220 genes) (\u003cem\u003eSupplementary Table 26\u003c/em\u003e), the significant GO terms (q-value \u0026lt; 0.1) were: \u003cem\u003ePositive Regulation of Transcription by RNA Polymerase II\u003c/em\u003e (GO:0045944, q-value = 0.0099); \u003cem\u003ePositive Regulation of DNA-templated Transcription\u003c/em\u003e (GO:0045893, q-value = 0.3475) and \u003cem\u003eNeural Tube Closure\u003c/em\u003e (GO: 0001843, q-value = 0.0567) (\u003cem\u003eSupplementary Table 27\u003c/em\u003e). \u003c/p\u003e \u003cp\u003eIf we apply Bonferroni correction (0.05/212\u0026thinsp;=\u0026thinsp;0.00024), the significant GO terms were the same as above, plus \u003cem\u003ePrimary Neural Tube Formation\u003c/em\u003e (GO:0014020, p-value\u0026thinsp;=\u0026thinsp;0.00019) (\u003cem\u003eSupplementary Table\u0026nbsp;27\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eIn this analysis, 27 genes were shared between cohorts (\u003cem\u003eSupplementary Table 28\u003c/em\u003e), including \u003cem\u003eECHS1\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eThe top 30 biological processes by p-value from GO enrichment analysis of Enrichr, using ReVIGO for visualizing genes from both cohorts (integrative analysis) are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec28\" class=\"Section2\"\u003e \u003ch2\u003eSingle-Cell Gene Differential Gene Expression Analysis\u003c/h2\u003e \u003cp\u003eThe same analysis of scRNA-seq done in the Spanish cohort, was done with linked genes from the SSC cohort with CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 and q-value\u0026thinsp;\u0026lt;\u0026thinsp;0,1 (n\u0026thinsp;=\u0026thinsp;202), as we discussed above (\u003cem\u003eSupplementary Table\u0026nbsp;28\u003c/em\u003e). Results have shown that the most upregulated gene was \u003cem\u003ePTGES3\u003c/em\u003e (\u003cem\u003elogFC\u0026thinsp;=\u0026thinsp;0.855; FDR\u0026thinsp;=\u0026thinsp;1.60x10\u003c/em\u003e\u003csup\u003e\u003cem\u003e\u0026minus;\u0026thinsp;10\u003c/em\u003e\u003c/sup\u003e; \u003cem\u003ecell type\u0026thinsp;=\u0026thinsp;EXT_9_L6\u003c/em\u003e), while the most downregulated was \u003cem\u003eRERE (logFC= -0,647; FDR\u0026thinsp;=\u0026thinsp;5.87 x10\u003c/em\u003e\u003csup\u003e\u003cem\u003e\u0026minus;\u0026thinsp;8\u003c/em\u003e\u003c/sup\u003e ; \u003cem\u003ecell type\u0026thinsp;=\u0026thinsp;EXT_9_L6\u003c/em\u003e and l\u003cem\u003eogFC=\u003c/em\u003e -0,934; \u003cem\u003eFDR\u0026thinsp;=\u003c/em\u003e\u0026thinsp;3.93 \u003cem\u003ex10\u003c/em\u003e\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e ; \u003cem\u003ecell type\u0026thinsp;=\u0026thinsp;ODC_1)\u003c/em\u003e (\u003cem\u003eSupplementary Table\u0026nbsp;29\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eWhen comparing gene mapping lists from the Spanish cohort (CnD p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05, and the SSC cohort (CnD p-value\u0026thinsp;\u0026gt;\u0026thinsp;0.05 and q-value\u0026thinsp;\u0026lt;\u0026thinsp;0.1), 5 genes are shared between cohorts, \u003cem\u003eECHS1, THNSL2, KIAA0922, PRRC2A\u003c/em\u003e and \u003cem\u003eC9orf72\u003c/em\u003e (\u003cem\u003eSupplementary Table\u0026nbsp;31\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eHowever single cell gene expression meta-analysis showed that only \u003cem\u003eECHS1\u003c/em\u003e was significantly differentially expressed in both our cohort and the SSC cohort, specifically in excitatory neurons from the cerebral cortex (\u003cem\u003elogFC\u0026thinsp;=\u0026thinsp;0,343; FDR\u0026thinsp;=\u0026thinsp;5.92x10\u003c/em\u003e\u003csup\u003e\u003cem\u003e\u0026minus;\u0026thinsp;8\u003c/em\u003e\u003c/sup\u003e; \u003cem\u003ecell type\u0026thinsp;=\u0026thinsp;EXT_4_L56\u003c/em\u003e).\u003c/p\u003e \u003cp\u003eDespite being identified as pathogenic by SISTR and supported by expression data in both cohorts and SFARI classification, \u003cem\u003eCASZ1\u003c/em\u003e did not exhibit significant differential expression between ASD cases and controls, nor cell-type-specific expression.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec29\" class=\"Section2\"\u003e \u003ch2\u003eIntegration of SSC and region in chromosome 10 emerges as the top candidate region\u003c/h2\u003e \u003cp\u003e \u003cem\u003eECHS1 (enoyl-CoA hydratase, short chain 1)\u003c/em\u003e is the only differentially expressed gene between ASD cases and controls when the targeted genes from de novo TRs in the Spanish and SSC cohorts are considered together in a single-cell analysis (Supplementary \u003cem\u003eTables\u0026nbsp;18 and 29\u003c/em\u003e). We found one individual with a mutation in the cCRE regulating \u003cem\u003eECHS1\u003c/em\u003e in the Spanish cohort and 11 individuals with 11 different mutations in the same cCRE in the SSC cohort (see \u003cem\u003eSupplementary Table\u0026nbsp;32\u003c/em\u003e). Moreover, we discovered that this region is close to another mutated cCRE in the Spanish cohort linked to another gene, \u003cem\u003eCALY (Calcyon Neuron Specific Vesicular Protein)\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e6\u003c/span\u003e). We would like to point out that, as will be discussed in the next section, genes were assigned to mutated cCRES using T-Gene, an in silico method that does not employ ASD-specific expression data. Single-cell brain data from ASD cases and controls identified \u003cem\u003eECHS1\u003c/em\u003e as the unique gene, but \u003cem\u003eCALY\u003c/em\u003e,located approximately 25000 bp upstream, was the most overexpressed when genes mutated in the Spanish cohort were considered (in interneurons and excitatory neurons) (see \u003cem\u003eSupplementary Table\u0026nbsp;2\u003c/em\u003e). Together with the described function of \u003cem\u003eCALY\u003c/em\u003e, the encoded protein interacts with the D1 dopamine receptor, this data suggests that the entire region would be an excellent candidate for future ASD functional studies.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe observed that \u003cem\u003eECHS1\u003c/em\u003e expression decreases during the early stages of differentiation of cortical organoids (~\u0026thinsp;day 100). After this point, expression progressively increases and remains elevated throughout later stages of neurodevelopment. This biphasic pattern suggests that \u003cem\u003eECHS1\u003c/em\u003e may have a role both in early neurodevelopmental transitions and in later maturation processes.\u003c/p\u003e \u003cp\u003eAcross human brain development (Brain Span data) \u003cem\u003eECHS1\u003c/em\u003e shows relatively stable expression with a modest dip around stage 6 (late fetal/early postnatal), followed by an increase during later postnatal stages. In contrast, \u003cem\u003eCALY\u003c/em\u003e expression markedly rose during early and mid-differentiation (100\u0026ndash;300 days), and subsequently declined, suggesting predominant involvement in early neurodevelopmental processes. Together, these findings highlight distinct yet complementary temporal profiles: whereas \u003cem\u003eCALY\u003c/em\u003e is most active during early neuronal specification and connectivity, \u003cem\u003eECHS1\u003c/em\u003e becomes increasingly relevant at later stages, underscoring how disruption of either process may contribute to ASD.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn our cohort, 57 different candidate cCREs and 107 high-confidence strict de novo TRs were identified. TRs demonstrated a large bias toward expansions (76%) over contractions, which is consistent with previous genome-wide studies. Interestingly, in contrast to the more prevalent single-repeat deletions previously documented, the most common contraction pattern in our dataset featured the deletion of three repeats (25). In addition, it was also found that there was a clear bias toward expansions among de novo TRs mutations. This pattern is consistent with prior genome-wide studies, which have also reported a higher prevalence of expansions over contractions among de novo TR mutations (25). From a functional perspective, this bias may have important biological implications. Expansions in regulatory regions could disrupt transcription factor binding motifs, alter chromatin accessibility, or interfere with the timing and levels of gene expression mechanisms during neurodevelopment (35). Importantly, the observed bias toward expansions was detected in a set of high-confidence, strictly de novo TRs, suggesting that the pattern is unlikely to result from technical artifacts. Instead, it likely reflects mutational patterns in the germline, potentially linked to chromatin architecture in regulatory regions. Moreover, it is important to note that this cohort consists exclusively of cases that have not been genetically diagnosed with autism through exome sequencing or CNV arrays but exhibit phenotypic traits consistent with ASD. This suggests that there must be some form of gene dysregulation that remains undetectable by the genetic diagnostic methods currently used. Supporting this, a recent genome-wide study identified over 2,000 VNTR loci enriched in regulatory regions, highlighting a potential source of genetic variation that escapes conventional detection (11).\u003c/p\u003e \u003cp\u003eIn this Spanish cohort, a total of 107 de novo TRs mutations were identified. Using SISTR pathogenicity scores several genes were highlighted. \u003cem\u003eTM4SF1 (36)\u003c/em\u003e and \u003cem\u003eC12orf44\u003c/em\u003e (37), which exhibit the highest pathogenicity scores, have been previously linked to neuronal defects but are not listed in the SFARI database. In addition, \u003cem\u003eCASZ1\u003c/em\u003e, identified in a large-scale targeted sequencing study of neurodevelopmental disorders (38), is involved in neurodevelopment, regulating neurogenesis and the transition to gliogenesis (39), and is classified as a high-confidence ASD gene in SFARI. Other relevant genes also emerged, such as \u003cem\u003eTM4SF1\u003c/em\u003e (40,41); \u003cem\u003eUBE2K\u003c/em\u003e (42\u0026ndash;44); and \u003cem\u003eNR4A1\u003c/em\u003e (45,46).\u003c/p\u003e \u003cp\u003eThe results from gene mapping analyses, based on physical proximity or expression data, underscore the importance of considering expression data when assigning genes to regulatory regions for further study. In nearly 40% of cases, the gene regulated by a cCRE was not the physically closest gene, indicating that the nearest gene is not always the one controlled by a regulatory region. These results remark the need to integrate three-dimensional genomic data (e.g., Hi-C, Capture-C) and expression-based approaches, rather than relying solely on linear distance, to assign regulatory elements to their target genes.\u003c/p\u003e \u003cp\u003eWhen the our cohort was analyzed independently, GO enrichment highlighted processes such as negative regulation of astrocyte differentiation and regulation of cell differentiation are especially relevant, underscoring the role of astrocytes in the etiology of ASD and pointing to the possible potential functional consequences of these mutations in early neural development and reinforcing the role of astrocytes in ASD etiology (47). Moreover, in the meta-analysis integrating both cohorts, several biological processes of early neurodevelopment were significantly overrepresented.\u003c/p\u003e \u003cp\u003eHowever, none of the genes classified as pathogenic by SISTR in the Spanish cohort were shared with the SSC collection. Using gene mapping with expression data to link cCREs carrying de novo TRs in both cohorts, we identified only ten genes with de novo TRs shared between both cohorts: \u003cem\u003eECHS1, C9orf3, C9orf72, FARP1, KIAA0922, PBX1, SOCS3, UBE2K, THNSL2, and WIPF2\u003c/em\u003e. Among these, only \u003cem\u003eECHS1 (enoyl-CoA hydratase, short chain 1)\u003c/em\u003e is differentially expressed at the single cell level, showing overexpression in excitatory neurons in ASD brain.\u003c/p\u003e \u003cp\u003eTaken together with the findings of Mitra et al., our results point to the involvement of novel genes that can only be uncovered through the application of TRs detection algorithms and subsequent analyses. Importantly, we also identify genes not represented in SFARI, suggesting the presence of complex structural repeat variants as well as rare, inherited variants in additional genes, further underscoring the genetic complexity underlying ASD. E\u003cem\u003eCHS1\u003c/em\u003e, as we will discuss later, emerges as the main candidate gene because it is the only one shared between the Spanish cohort and the large SSC cohort at multiple analysis levels.\u003c/p\u003e \u003cp\u003e \u003cem\u003eECHS1\u003c/em\u003e is located near \u003cem\u003eCALY\u003c/em\u003e, one of the target genes identified for de novo TRs in the Spanish cohort, which also showed overexpression in excitatory neurons in scRNAseq data. We found one individual carrying a mutation in the cCRE regulating \u003cem\u003eECHS1\u003c/em\u003e in the Spanish cohort, as well as 11 additional individuals with 11 distinct mutations in the SSC (29) which remarks the relevance of the region. It should be noted that genes were assigned to mutated cCRES using an \u003cem\u003ein silico\u003c/em\u003e method that integrates generic ChIP-seq data from ENCODE and single-cell brain data from ASD cases vs controls, adding greater specificity to the ASD phenotype. We hypothesize that TRs within the cCREs analyzed in this study, or even in additional cCREs not yet included (as ENCODE v3 has identified a large number of novel cCREs in this region than the ENCODE v2 used in this study), may act as regulators of nearby genes independently \u003cem\u003eCALY\u003c/em\u003e or \u003cem\u003eECSH1.\u003c/em\u003e\u003c/p\u003e \u003cp\u003eAmong these two genes, \u003cem\u003ePRAP1 (proline-rich acidic protein 1)\u003c/em\u003e and \u003cem\u003eFUOM (fucose mutarotase)\u003c/em\u003e are also present; however, no mutations in our study targeted cCRES within these genes, and they are not expressed at the single-cell level in ASD brains nor controls.\u003c/p\u003e \u003cp\u003eSince the largest number of mutations are within \u003cem\u003eECHS1\u003c/em\u003e, we will thoroughly discuss its role without losing sight of the role of \u003cem\u003eCALY\u003c/em\u003e, which encodes a dopamine receptor D1-interacting protein (48). The role of \u003cem\u003eCALY\u003c/em\u003e in ASD or other NDDs has not been described or studied. However, dopamine receptors have been extensively investigated in ASD at the genetic and functional level. In particular, dopamine receptors D2/D3 in the striatum and D1 in the prefrontal cortex appear to be dysregulated in ASD (49). \u003cem\u003eECHS1\u003c/em\u003e encodes a mitochondrial enzyme with a key role in b-oxidation, as well as in the metabolic pathways of isoleucine and valine (50). Mutations in \u003cem\u003eECHS1\u003c/em\u003e are linked to a wide spectrum of clinical phenotypes, ranging from neonatal death to adult survival (51). The most common phenotype manifests like the Leigh syndrome (52) and related encephalopathies, characterised by different neurological impairments, arrhythmia and neonatal seizures (50). A second group of individuals manifest developmental regression resulting in severe developmental delay. Another group of related phenotypes include individuals with a normal development with paroxysmal dystonia that may be exacerbated by illness or exertion (53).In the context of ASD etiology, the overexpression of \u003cem\u003eECHS1\u003c/em\u003e observed in excitatory neurons may suggest that even subtle dysregulation of this mitochondrial pathway could contribute to altered neuronal activity. Given its essential role in energy metabolism, such changes may be particularly relevant to excitatory/inhibitory imbalance, a mechanism increasingly recognized in autism pathophysiology (54,55).\u003c/p\u003e \u003cp\u003eWe aim to explore \u003cem\u003eECHS1\u003c/em\u003e and \u003cem\u003eCALY\u003c/em\u003e together in cortical organoids at different times of differentiation to determine whether their regulation might be coordinated during neurodevelopment. Based on cortical organoid and Brain Span data, their timing of functional relevance differs: \u003cem\u003eCALY\u003c/em\u003e appears to be relevant during early differentiation and synaptic formation while \u003cem\u003eECHS1\u003c/em\u003e seems to be more involved in later neuronal maturation and energy metabolism, consistent with its mitochondrial function. In terms of potential disease impact, dysregulation of \u003cem\u003eCALY\u003c/em\u003e could disrupt early synaptic development and connectivity, impairing circuit formation, whereas dysregulation of \u003cem\u003eECHS1\u003c/em\u003e might compromise neuronal maturation processes, particularly through metabolic stress at later stages.\u003c/p\u003e \u003cp\u003eWe hypothesize that if de novo TRs occur in regulatory regions in individuals with ASD, they may contribute to milder or intermediate phenotypes, in contrast to mutations in the coding region that are associated with severe syndromes as those described above. This suggests that noncoding TRs in regulatory regions of \u003cem\u003eECHS1\u003c/em\u003e might represent a mechanism underlying less severe, but clinically relevant, ASD-associated phenotypes.\u003c/p\u003e \u003cp\u003eThe statistical strength of the SSC cohort further supports this gene as a promising target for genetic and functional studies aimed at clarifying how its dysregulation contributes to ASD pathophysiology.\u003c/p\u003e \u003cdiv id=\"Sec31\" class=\"Section2\"\u003e \u003ch2\u003eLimitations and Future Directions\u003c/h2\u003e \u003cp\u003eWhile our study provides novel insights into the role of de novo TRs in regulatory regions and highlights \u003cem\u003eECHS1\u003c/em\u003e as a promising candidate gene, it should be noted that there are several limitations.\u003c/p\u003e \u003cp\u003eFirst, confirming the impact of TR mutations would require utilizing human-derived neuronal and glial models, such as brain organoids or iPSC-derived cells from the patients harboring the mutation. Second, although we leveraged large cohorts such as the Spanish cohort and SSC, there are important differences between them that may affect interpretation. For example, the Spanish cohort includes individuals with a prior specific exclusion/selection criteria, whereas the SSC cohort does not. These differences may slightly limit the comparability of our findings across populations, even when the methods used to identify variants are the same. Third, the assignment of regulatory elements to target genes relied on current annotations, which may not fully capture the complexity of three-dimensional genome architecture or long-range regulatory interactions.\u003c/p\u003e \u003cp\u003eIt is also important to note that TRs larger than 150 bp are difficult to detect with established short-read sequencing platforms, as these are generally unable to accurately genotype large or complex repeat expansions (56). While specific short-read TR programs exist, they are currently limited in detecting repeats with very large motif sizes. Furthermore, genes were linked to cCREs containing mutations without accounting for TR characteristics, which may limit prioritization of certain regulatory elements over others.\u003c/p\u003e \u003cp\u003eFuture studies should integrate multi-omics data, including epigenomic, proteomic, and functional assays, to better understand how TR-mediated dysregulation contributes to ASD. Targeted functional studies of \u003cem\u003eECHS1\u003c/em\u003e region will be important to explore how subtle changes in expression affect neuronal excitability and glial-neuronal interactions. Longitudinal analyses and studies in more diverse populations could further clarify how regulatory TRs contribute to the heterogeneity of ASD phenotypes, from mild to severe manifestations. Overall, addressing these limitations will provide a more comprehensive understanding of the genetic and cellular mechanisms underlying ASD and may identify novel therapeutic targets.\u003c/p\u003e \u003cp\u003eIn conclusion, our study highlights the importance of de novo TRs in regulatory regions as contributors to ASD risk, uncovering candidate gene regions such as \u003cem\u003eECHS1\u003c/em\u003e that would have remained undetected using conventional approaches. By integrating genetic, transcriptomic, single-cell and cortical organoid data across cohorts, we provide new insights of these genes in neurodevelopment.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec32\" class=\"Section2\"\u003e \u003ch2\u003eData availability\u003c/h2\u003e \u003cp\u003eSequencing data from the Spanish cohort are transferred to EGA ( European Genome-phenome Archive ) with Study Accession:EGAS50000001395.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAcknowledgements\u003c/h2\u003e \u003cp\u003eThis study has been funded by Instituto de Salud Carlos III (ISCIII) through the project \u0026ldquo;PI24/00595\u0026rdquo; and co-funded by the European Union.\u003c/p\u003e \u003cp\u003eC. Arango was supported by the Spanish Ministry of Science and Innovation, Instituto de Salud Carlos III (ISCIII), co-financed by the European Union, ERDF Funds from the European Commission, \u0026ldquo;A way of making Europe\u0026rdquo;, financed by the European Union \u0026ndash; NextGenerationEU (PMP21/00051), PI19/01024. PI22/01824 CIBERSAM, Madrid Regional Government (B2017/BMD-3740 AGES-CM-2), European Union Structural Funds, European Union Seventh Framework Program, European Union H2020 Program under the Innovative Medicines Initiative 2 Joint Undertaking: Project PRISM-2 (Grant agreement No.101034377), Project AIMS-2-TRIALS (Grant agreement No 777394), Horizon Europe, the National Institute of Mental Health of the National Institutes of Health under Award Number 1U01MH124639-01 (Project ProNET) and Award Number 5P50MH115846-03 (project FEP-CAUSAL), Fundaci\u0026oacute;n Familia Alonso, and Fundaci\u0026oacute;n Alicia Koplowitz.\u003c/p\u003e \u003cp\u003eWe would like to thank Virginia Sestelo Prado by the work she carried out during her internship as part of her external master practicum working with T-Gene tool.\u003c/p\u003e \u003cp\u003e \u003cspan type=\"BoldUnderline\" class=\"BoldUnderline\" name=\"Emphasis\"\u003eConflict of interests\u003c/span\u003e \u003c/p\u003e \u003cp\u003eDr. Arango has been a consultant to or has received honoraria or grants from Abbot, Acadia, Ambrosetti, Angelini, Biogen, BMS, Boehringer, Carnot, Gedeon Richter, Janssen Cilag, Lundbeck, Medscape, Menarini, Minerva, Otsuka, Pfizer, Roche, Rovi, Sage, Servier, Shire, Schering Plough, Sumitomo Dainippon Pharma, Sunovion, Takeda and Teva.\u003c/p\u003e \u003cp\u003e \u003cspan type=\"BoldUnderline\" class=\"BoldUnderline\" name=\"Emphasis\"\u003eContributions\u003c/span\u003e \u003c/p\u003e \u003cp\u003eP Carballo-Pacoret has carried out the analyses and wrote the paper. S.Dominguez-Alonso has carried part of the analyses regarding the sequencing and selection of regulatory regions and managing of sequencing files. J Gonzalez-Pe\u0026ntilde;as, A. Carracedo, C. Arango and M.Parellada participated in the selection and recruitment of samples. A. Carracedo and C.Rodriguez-Fontenla participated in the design and coordination of this study. C.Rodriguez-Fontenla critically revised the work, wrote the paper and approved the final content.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAmerican Psychiatric Association, American Psychiatric Association, editors. Diagnostic and statistical manual of mental disorders: DSM-5. 5th ed. Washington, D.C: American Psychiatric Association; 2013. 947 p.\u003c/li\u003e\n\u003cli\u003eRosti RO, Sadek AA, Vaux KK, Gleeson JG. The genetic landscape of autism spectrum disorders. Dev Med Child Neurol. 2014 Jan;56(1):12–8.\u003c/li\u003e\n\u003cli\u003eSandin S, Lichtenstein P, Kuja-Halkola R, Hultman C, Larsson H, Reichenberg A. The Heritability of Autism Spectrum Disorder. JAMA. 2017 Sept 26;318(12):1182.\u003c/li\u003e\n\u003cli\u003eAutism Spectrum Disorder Working Group of the Psychiatric Genomics Consortium, BUPGEN, Major Depressive Disorder Working Group of the Psychiatric Genomics Consortium, 23andMe Research Team, Grove J, Ripke S, et al. Identification of common genetic risk variants for autism spectrum disorder. Nat Genet. 2019 Mar;51(3):431–44.\u003c/li\u003e\n\u003cli\u003eGaugler T, Klei L, Sanders SJ, Bodea CA, Goldberg AP, Lee AB, et al. Most genetic risk for autism resides with common variation. Nat Genet. 2014 Aug;46(8):881–5.\u003c/li\u003e\n\u003cli\u003eSanders SJ, Murtha MT, Gupta AR, Murdoch JD, Raubeson MJ, Willsey AJ, et al. De novo mutations revealed by whole-exome sequencing are strongly associated with autism. Nature. 2012 May;485(7397):237–41.\u003c/li\u003e\n\u003cli\u003eQuesnel-Vallières M, Weatheritt RJ, Cordes SP, Blencowe BJ. Autism spectrum disorder: insights into convergent mechanisms from transcriptomics. Nat Rev Genet. 2019 Jan;20(1):51–63.\u003c/li\u003e\n\u003cli\u003eBahlo M, Bennett MF, Degorski P, Tankard RM, Delatycki MB, Lockhart PJ. Recent advances in the detection of repeat expansions with short-read next-generation sequencing. F1000Research. 2018 June 13;7:736.\u003c/li\u003e\n\u003cli\u003eHannan AJ. Tandem repeat polymorphisms: modulators of disease susceptibility and candidates for ‘missing heritability’. Trends Genet. 2010 Feb;26(2):59–65.\u003c/li\u003e\n\u003cli\u003eInternational Human Genome Sequencing Consortium, Whitehead Institute for Biomedical Research, Center for Genome Research:, Lander ES, Linton LM, Birren B, Nusbaum C, et al. Initial sequencing and analysis of the human genome. Nature. 2001 Feb 15;409(6822):860–921.\u003c/li\u003e\n\u003cli\u003eZhang S, Song Q, Zhang P, Wang X, Guo R, Li Y, et al. Genome-wide investigation of VNTR motif polymorphisms in 8,222 genomes: Implications for biological regulation and human traits. Cell Genomics. 2024 Dec;4(12):100699.\u003c/li\u003e\n\u003cli\u003eGymrek M, Willems T, Reich D, Erlich Y. Interpreting short tandem repeat variations in humans using mutational constraint. Nat Genet. 2017 Oct 1;49(10):1495–501.\u003c/li\u003e\n\u003cli\u003eLiao X, Zhu W, Zhou J, Li H, Xu X, Zhang B, et al. Repetitive DNA sequence detection and its role in the human genome. Commun Biol. 2023 Sept 19;6(1):954.\u003c/li\u003e\n\u003cli\u003eMacdonald M. A novel gene containing a trinucleotide repeat that is expanded and unstable on Huntington’s disease chromosomes. Cell. 1993 Mar;72(6):971–83.\u003c/li\u003e\n\u003cli\u003eUsdin K. The biological effects of simple tandem repeats: Lessons from the repeat expansion diseases: Table 1. Genome Res. 2008 July;18(7):1011–9.\u003c/li\u003e\n\u003cli\u003eFyke W, Velinov M. FMR1 and Autism, an Intriguing Connection Revisited. Genes. 2021 Aug 6;12(8):1218.\u003c/li\u003e\n\u003cli\u003eCortese A, Simone R, Sullivan R, Vandrovcova J, Tariq H, Yau WY, et al. Biallelic expansion of an intronic repeat in RFC1 is a common cause of late-onset ataxia. Nat Genet. 2019 Apr;51(4):649–58.\u003c/li\u003e\n\u003cli\u003eMalik I, Kelley CP, Wang ET, Todd PK. Molecular mechanisms underlying nucleotide repeat expansion disorders. Nat Rev Mol Cell Biol. 2021 Sept;22(9):589–607.\u003c/li\u003e\n\u003cli\u003eMojarad BA, Engchuan W, Trost B, Backstrom I, Yin Y, Thiruvahindrapuram B, et al. Genome-wide tandem repeat expansions contribute to schizophrenia risk. Mol Psychiatry. 2022 Sept;27(9):3692–8.\u003c/li\u003e\n\u003cli\u003eTrost B, Thiruvahindrapuram B, Chan AJS, Engchuan W, Higginbotham EJ, Howe JL, et al. Genomic architecture of autism from comprehensive whole-genome sequence annotation. Cell. 2022 Nov;185(23):4409–4427.e18.\u003c/li\u003e\n\u003cli\u003eHannan AJ. Tandem repeats mediating genetic plasticity in health and disease. Nat Rev Genet. 2018 May;19(5):286–98.\u003c/li\u003e\n\u003cli\u003eBrandler WM, Antaki D, Gujral M, Kleiber ML, Whitney J, Maile MS, et al. Paternally inherited cis-regulatory structural variants are associated with autism. Science. 2018 Apr 20;360(6386):327–31.\u003c/li\u003e\n\u003cli\u003eMarshall CR, Noor A, Vincent JB, Lionel AC, Feuk L, Skaug J, et al. Structural Variation of Chromosomes in Autism Spectrum Disorder. Am J Hum Genet. 2008 Feb;82(2):477–88.\u003c/li\u003e\n\u003cli\u003eTrost B, Engchuan W, Nguyen CM, Thiruvahindrapuram B, Dolzhenko E, Backstrom I, et al. Genome-wide detection of tandem DNA repeats that are expanded in autism. Nature. 2020 Oct 1;586(7827):80–6.\u003c/li\u003e\n\u003cli\u003eMitra I, Huang B, Mousavi N, Ma N, Lamkin M, Yanicky R, et al. Patterns of de novo tandem repeat mutations and their role in autism. Nature. 2021 Jan 14;589(7841):246–50.\u003c/li\u003e\n\u003cli\u003eWerling D, Brand H, An JY, Stone M, Glessner J, Zhu L, et al. LIMITED CONTRIBUTION OF RARE, NONCODING VARIATION TO AUTISM SPECTRUM DISORDER FROM SEQUENCING OF 2,076 GENOMES IN QUARTET FAMILIES. Eur Neuropsychopharmacol. 2019;29:S784–5.\u003c/li\u003e\n\u003cli\u003eAlonso-Gonzalez A, Calaza M, Amigo J, González-Peñas J, Martínez-Regueiro R, Fernández-Prieto M, et al. Exploring the biological role of postzygotic and germinal de novo mutations in ASD. Sci Rep. 2021 Jan 11;11(1):319.\u003c/li\u003e\n\u003cli\u003eThe ENCODE Project Consortium, Abascal F, Acosta R, Addleman NJ, Adrian J, Afzal V, et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature. 2020 July 30;583(7818):699–710.\u003c/li\u003e\n\u003cli\u003eMitra I, Huang B, Mousavi N, Ma N, Lamkin M, Yanicky R, et al. Patterns of de novo tandem repeat mutations and their role in autism. Nature. 2021 Jan 14;589(7841):246–50.\u003c/li\u003e\n\u003cli\u003eGrant CE, Bailey TL, Noble WS. FIMO: scanning for occurrences of a given motif. Bioinformatics. 2011 Apr 1;27(7):1017–8.\u003c/li\u003e\n\u003cli\u003eRauluseviciute I, Riudavets-Puig R, Blanc-Mathieu R, Castro-Mondragon JA, Ferenc K, Kumar V, et al. JASPAR 2024: 20th anniversary of the open-access database of transcription factor binding profiles. Nucleic Acids Res. 2024 Jan 5;52(D1):D174–82.\u003c/li\u003e\n\u003cli\u003eO’Connor T, Grant CE, Bodén M, Bailey TL. T-Gene: improved target gene prediction. Wren J, editor. Bioinformatics. 2020 June 1;36(12):3902–4.\u003c/li\u003e\n\u003cli\u003eWamsley B, Bicks L, Cheng Y, Kawaguchi R, Quintero D, Margolis M, et al. Molecular cascades and cell type–specific signatures in ASD revealed by single-cell genomics. Science. 2024 May 24;384(6698):eadh2602.\u003c/li\u003e\n\u003cli\u003eGordon A, Yoon SJ, Tran SS, Makinson CD, Park JY, Andersen J, et al. Long-term maturation of human cortical organoids matches key early postnatal transitions. Nat Neurosci. 2021 Mar;24(3):331–42.\u003c/li\u003e\n\u003cli\u003eSun JX, Helgason A, Masson G, Ebenesersdóttir SS, Li H, Mallick S, et al. A direct characterization of human mutation based on microsatellites. Nat Genet. 2012 Oct;44(10):1161–5.\u003c/li\u003e\n\u003cli\u003eShih SC, Zukauskas A, Li D, Liu G, Ang LH, Nagy JA, et al. The L6 Protein TM4SF1 Is Critical for Endothelial Cell Function and Tumor Angiogenesis. Cancer Res. 2009 Apr 15;69(8):3272–7.\u003c/li\u003e\n\u003cli\u003eGuo T, Nan Z, Miao C, Jin X, Yang W, Wang Z, et al. The autophagy-related gene Atg101 in Drosophila regulates both neuron and midgut homeostasis. J Biol Chem. 2019 Apr;294(14):5666–76.\u003c/li\u003e\n\u003cli\u003eWang T, Hoekzema K, Vecchio D, Wu H, Sulovari A, Coe BP, et al. Large-scale targeted sequencing identifies risk genes for neurodevelopmental disorders. Nat Commun. 2020 Oct 1;11(1):4932.\u003c/li\u003e\n\u003cli\u003eLiu T, Li T, Ke S. Role of the CASZ1 transcription factor in tissue development and disease. Eur J Med Res. 2023 Dec 5;28(1):562.\u003c/li\u003e\n\u003cli\u003eKitajima H, Maruyama R, Niinuma T, Yamamoto E, Takasawa A, Takasawa K, et al. TM4SF1-AS1 inhibits apoptosis by promoting stress granule formation in cancer cells. Cell Death Dis. 2023 July 13;14(7):424.\u003c/li\u003e\n\u003cli\u003eTang Q, Chen J, Di Z, Yuan W, Zhou Z, Liu Z, et al. TM4SF1 promotes EMT and cancer stemness via the Wnt/β-catenin/SOX2 pathway in colorectal cancer. J Exp Clin Cancer Res. 2020 Dec;39(1):232.\u003c/li\u003e\n\u003cli\u003eCai Y, Ji Y, Liu Y, Zhang D, Gong Z, Li L, et al. Microglial circ-UBE2K exacerbates depression by regulating parental gene UBE2K via targeting HNRNPU. Theranostics. 2024;14(10):4058–75.\u003c/li\u003e\n\u003cli\u003eFilatova EV, Shadrina MI, Alieva AKh, Kolacheva AA, Slominsky PA, Ugrumov MV. Expression analysis of genes of ubiquitin-proteasome protein degradation system in MPTP-induced mice models of early stages of Parkinson’s disease. Dokl Biochem Biophys. 2014 May;456(1):116–8.\u003c/li\u003e\n\u003cli\u003eMeiklejohn H, Mostaid MS, Luza S, Mancuso SG, Kang D, Atherton S, et al. Blood and brain protein levels of ubiquitin-conjugating enzyme E2K (UBE2K) are elevated in individuals with schizophrenia. J Psychiatr Res. 2019 June;113:51–7.\u003c/li\u003e\n\u003cli\u003eLi H, Zhao P, Xu Q, Shan S, Hu C, Qiu Z, et al. The autism-related gene SNRPN regulates cortical and spine development via controlling nuclear receptor Nr4a1. Sci Rep. 2016 July 19;6(1):29878.\u003c/li\u003e\n\u003cli\u003eOkay K, Varış PÜ, Miral S, Ekinci B, Yaraş T, Karakülah G, et al. Alternative splicing and gene co-expression network-based analysis of dizygotic twins with autism-spectrum disorder and their parents. Genomics. 2021 July;113(4):2561–71.\u003c/li\u003e\n\u003cli\u003eAllen M, Huang BS, Notaras MJ, Lodhi A, Barrio-Alonso E, Lituma PJ, et al. Astrocytes derived from ASD individuals alter behavior and destabilize neuronal activity through aberrant Ca2 + signaling. Mol Psychiatry. 2022 May;27(5):2470–84.\u003c/li\u003e\n\u003cli\u003eHa CM, Park D, Han JK, Jang J ill, Park JY, Hwang EM, et al. Calcyon Forms a Novel Ternary Complex with Dopamine D1 Receptor through PSD-95 Protein and Plays a Role in Dopamine Receptor Internalization. J Biol Chem. 2012 Sept;287(38):31813–22.\u003c/li\u003e\n\u003cli\u003eSu P, Lai TKY, Lee FHF, Abela AR, Fletcher PJ, Liu F. Disruption of SynGAP–dopamine D1 receptor complexes alters actin and microtubule dynamics and impairs GABAergic interneuron migration. Sci Signal. 2019 Aug 6;12(593):eaau9122.\u003c/li\u003e\n\u003cli\u003eMasnada S, Parazzini C, Bini P, Barbarini M, Alberti L, Valente M, et al. Phenotypic spectrum of short-chain enoyl-Coa hydratase-1 (ECHS1) deficiency. Eur J Paediatr Neurol. 2020 Sept;28:151–8.\u003c/li\u003e\n\u003cli\u003eMuntean C, Tripon F, Bogliș A, Bănescu C. Pathogenic Biallelic Mutations in ECHS1 in a Case with Short-Chain Enoyl-CoA Hydratase (SCEH) Deficiency-Case Report and Literature Review. Int J Environ Res Public Health. 2022 Feb 13;19(4):2088.\u003c/li\u003e\n\u003cli\u003eSakai C, Yamaguchi S, Sasaki M, Miyamoto Y, Matsushima Y, Goto Y ichi. ECHS1 Mutations Cause Combined Respiratory Chain Deficiency Resulting in Leigh Syndrome. Hum Mutat. 2015 Feb;36(2):232–9.\u003c/li\u003e\n\u003cli\u003eOlgiati S, Skorvanek M, Quadri M, Minneboo M, Graafland J, Breedveld GJ, et al. Paroxysmal exercise-induced dystonia within the phenotypic spectrum of \u003cem\u003eECHS1\u003c/em\u003e deficiency: \u003cem\u003eECHS1\u003c/em\u003e Mutations, Dystonia, and PED. Mov Disord. 2016 July;31(7):1041–8.\u003c/li\u003e\n\u003cli\u003eSatterstrom FK, Kosmicki JA, Wang J, Breen MS, De Rubeis S, An JY, et al. Large-Scale Exome Sequencing Study Implicates Both Developmental and Functional Changes in the Neurobiology of Autism. Cell. 2020 Feb;180(3):568–584.e23.\u003c/li\u003e\n\u003cli\u003eYizhar O, Fenno LE, Prigge M, Schneider F, Davidson TJ, O’Shea DJ, et al. Neocortical excitation/inhibition balance in information processing and social dysfunction. Nature. 2011 Sept;477(7363):171–8.\u003c/li\u003e\n\u003cli\u003eChintalaphani SR, Pineda SS, Deveson IW, Kumar KR. An update on the neurological short tandem repeat expansion disorders and the emergence of long-read sequencing diagnostics. Acta Neuropathol Commun. 2021 Dec;9(1):98.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"molecular-psychiatry","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"mp","sideBox":"Learn more about [Molecular Psychiatry](http://www.nature.com/mp/)","snPcode":"41380","submissionUrl":"https://mts-mp.nature.com/cgi-bin/main.plex","title":"Molecular Psychiatry","twitterHandle":"@molpsychiatry","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-8374597/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8374597/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIn this study, we performed an integrative analysis of de novo tandem repeats (TRs) to unravel the missing heritability that may be hidden in 85,394 active cis-regulatory elements (cCREs) from ENCODE through target sequencing in a Spanish cohort of 200 ASD trios, using a robust bioinformatic pipeline. For the integrative analysis, we use data from 1,637 ASD simplex quad families from the Simons Simplex Collection (SSC). We then incorporated multiple layers of functional annotation, including predicted transcription factor (TF) binding sites, gene mapping based on physical proximity and expression correlation, pathogenicity scoring, single-cell RNA-seq data from human brain in ASD cases and controls and cortical organoid expression data.\u003c/p\u003e \u003cp\u003eTogether, our analyses identified multiple ASD-relevant candidate genes supported by convergent lines of evidence. Notably, \u003cem\u003eECHS1\u003c/em\u003e emerged as a strong candidate, affected by several de novo TRs in both the Spanish cohort and the SSC. It was also identified as the most significantly associated gene through expression-based gene mapping (T-Gene) and showed consistent differential expression in excitatory neurons of the cerebral cortex at the single-cell level along with increased expression in late-stage cortical organoids.\u003c/p\u003e \u003cp\u003eThese findings remark the value of integrating genetic and transcriptomic information to improve the identification of potential risk genes for ASD, particularly within non-coding regions. Our approach also highlights the importance of identifying complex genetic variation, such as de novo TRs, that are typically missed in conventional exome or whole-genome analyses, and require specialized bioinformatic strategies for accurate detection and interpretation.\u003c/p\u003e","manuscriptTitle":"Large-scale sequencing study of de novo regulatory Tandem Repeats (TRs) identifies new ASD (Autism Spectrum Disorders) candidate genes integrating gene expression mapping, brain scRNA-seq and organoid models.","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-13 12:25:57","doi":"10.21203/rs.3.rs-8374597/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Reject after peer review","date":"2026-03-13T12:39:50+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"This content is not available.","date":"2026-03-09T01:28:31+00:00","index":1,"fulltext":"This content is not available."},{"type":"editorInvitedReview","content":"This content is not available.","date":"2026-03-06T00:58:48+00:00","index":2,"fulltext":"This content is not available."},{"type":"editorInvitedReview","content":"This content is not available.","date":"2026-02-19T19:27:48+00:00","index":3,"fulltext":"This content is not available."},{"type":"reviewerAgreed","content":"This content is not available.","date":"2026-02-18T15:51:13+00:00","index":3,"fulltext":"This content is not available."},{"type":"reviewerAgreed","content":"This content is not available.","date":"2026-02-18T11:10:40+00:00","index":2,"fulltext":"This content is not available."},{"type":"reviewerAgreed","content":"This content is not available.","date":"2026-02-18T02:41:06+00:00","index":1,"fulltext":"This content is not available."},{"type":"reviewersInvited","content":"","date":"2026-02-10T01:11:32+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-01-04T15:26:47+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-04T15:07:36+00:00","index":"","fulltext":""},{"type":"submitted","content":"Molecular Psychiatry","date":"2025-12-18T08:48:27+00:00","index":"","fulltext":""},{"type":"checksFailed","content":"","date":"2025-12-17T15:22:35+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"molecular-psychiatry","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"mp","sideBox":"Learn more about [Molecular Psychiatry](http://www.nature.com/mp/)","snPcode":"41380","submissionUrl":"https://mts-mp.nature.com/cgi-bin/main.plex","title":"Molecular Psychiatry","twitterHandle":"@molpsychiatry","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"d99f35f9-2b09-4a84-a497-b0375c479ce8","owner":[],"postedDate":"February 13th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":62622941,"name":"Health sciences/Diseases/Psychiatric disorders/Autism spectrum disorders"},{"id":62622942,"name":"Biological sciences/Genetics"}],"tags":[],"updatedAt":"2026-04-24T11:02:21+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-13 12:25:57","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8374597","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8374597","identity":"rs-8374597","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00