Effective regulatory relationships between genes differ from direct transcription factor-target interactions, and exhibit a distinct statistical signature

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

Effective gene regulation is diffusely distributed throughout transcriptional networks and can be identified by a distinct statistical signature rather than direct transcription factor-target interactions.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

This preprint studied whether effective gene regulatory relationships in gene regulatory networks align with direct transcription factor (TF)–target interactions, using a simple computational GRN model alongside human expression and perturbation datasets. Across simulated networks of 1,000 genes and 400 time points, and in human data spanning 796 TFs with confirmed targets and CMAP/LINCS perturbation experiments, the authors found widespread lack of statistical (Spearman correlation) and causal association between TF activity and their direct targets, with knockdowns of direct regulators not producing stronger effects on targets than random genes. They define “effective regulators” as genes whose experimental down-regulation most strongly impacts a target, showing these effective regulators form a smaller set with a distinct statistical signature, and networks built from them show greater clustering/connectivity than canonical TF-based networks. A key caveat is that the causal inferences depend on their simplified GRN model and on available TF–target annotations and perturbation coverage, which may not capture all biological intermediates. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Gene expression patterns are determined by the interactions between thousands of transcriptional regulators and their target genes. However, mounting evidence of an apparent lack of causal and statistical association between transcription factors and their targets, raises the question of where the effective regulation of any particular gene lies. Using both a simple computational gene regulatory network model together with human expression and perturbation data for 796 known transcription factors and their confirmed targets, we demonstrate a widespread and generalized lack of statistical and causal association between the activity of regulators and direct targets. We introduce the concept of effective regulators: genes that, upon experimental down-regulation, have the strongest impact on a given target. Effective regulators form a small set of genes that display a distinct statistical signature, and networks constructed from effective regulators exhibit greater clustering and connectivity than canonical TF-based networks. Our results show that, while at the transcriptome-level, actual regulatory relationships between genes do not coincide with direct transcription factor–target interactions and are, instead, diffusely distributed throughout the transcriptional network, they can however be readily identified based on their distinct statistical properties.
Full text 134,910 characters · extracted from preprint-html · click to expand
Effective regulatory relationships between genes differ from direct transcription factor-target interactions, and exhibit a distinct statistical signature | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Effective regulatory relationships between genes differ from direct transcription factor-target interactions, and exhibit a distinct statistical signature Diego Arturo Velázquez-Trejo, Jorge Adrián Islas-Ortíz, Bertha Rueda-Zarazúa, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8643144/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 17 You are reading this latest preprint version Abstract Gene expression patterns are determined by the interactions between thousands of transcriptional regulators and their target genes. However, mounting evidence of an apparent lack of causal and statistical association between transcription factors and their targets, raises the question of where the effective regulation of any particular gene lies. Using both a simple computational gene regulatory network model together with human expression and perturbation data for 796 known transcription factors and their confirmed targets, we demonstrate a widespread and generalized lack of statistical and causal association between the activity of regulators and direct targets. We introduce the concept of effective regulators: genes that, upon experimental down-regulation, have the strongest impact on a given target. Effective regulators form a small set of genes that display a distinct statistical signature, and networks constructed from effective regulators exhibit greater clustering and connectivity than canonical TF-based networks. Our results show that, while at the transcriptome-level, actual regulatory relationships between genes do not coincide with direct transcription factor–target interactions and are, instead, diffusely distributed throughout the transcriptional network, they can however be readily identified based on their distinct statistical properties. Biological sciences/Computational biology and bioinformatics Biological sciences/Genetics Biological sciences/Systems biology Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Introduction Regulatory networks can be represented as a set of interactions between regulatory elements and their individual targets 1 – 3 . Natural transcriptomes are dynamical systems where the activity of a particular gene in the network at a given time is dynamically driven by the activity of its immediate transcriptional regulators, and therefore a level of statistical association, and a definite causal association is to be expected between regulators and regulatory targets 3 – 5 . While transcriptional networks are a particular instance of regulatory network, where the complete regulator-target interaction map is not fully known, attempts to infer actual regulatory interactions between transcription factors and targets based on their potential statistical association have led to mixed results 3 , 4 . Along these lines, although numerous approaches to detect regulatory interactions based on coexpression measurements have been proposed, a sizable number of studies have reported a general lack of statistical association between regulators and targets, with a more restricted number of reports documenting a significant regulator-target association, when compared with random pairs of background genes 2 , 4 , 6 . In addition, several independent studies have reported no significant changes in gene expression of target genes, in response to experimental knock out/down of a growing number of specific transcription factors 7 , 8 . These discrepancies have been ascribed to the particular complexity of the cellular transcription process, involving a number of intermediate events such as protein synthesis, chromatin accessibility, transcription factor paralog redundancy, co-factors and complex signaling cascades up-stream of transcriptional initiation, among others 6 . Thus, in general, whether trivial statistical and causal associations, at the level of direct regulatory interactions are to be expected in regulatory networks, remains unclear. Here we used a simple computational gene regulatory network model to test general assumptions regarding regulatory interactions, and specifically ask whether statistical and causal associations are to be expected between regulators and direct targets in regulatory networks in general 9 . We next contrast our findings with empirical human expression and gene perturbation data, and find that in both cases, regulators tend to display little to no correlation with their direct targets and that, knocking down their activity has, on average, no stronger effect on their targets than background genes. We then introduce the concept of effective regulators, defined as those genes that exert the strongest causal influence on a given gene and characterize their key general statistical properties. Methods Description of the model We used a simple GRN model developed by 9 in which the regulatory structure is represented by a set of interaction pairs composed where nodes represent genes and directed edges linking them represent regulatory positive and negative interactions. We specifically used networks consisting of 1000 nodes, each node being regulated by a variable number of randomly chosen nodes, drawn from the entire network, with a minimum ( \(\:min\) ) number of 22 regulators per node. We set this minimum because the simulated networks, on average, exhibited a density closer to that estimated from transcriptomic networks. The number of regulators ( \(\:k\:=\:kneg\:+\:kpos\) ) per node ( \(\:i\) ) was set to follow a power law distribution defined as: \(\:{k}_{i}\:=\:min\:+\:round\left(\frac{1}{{x}^{2}}\right)\) \(\:\) Where \(\:min\:=\:22\) is the minimum number of regulators and \(\:x\:\sim\:Unif(0,\:1)\) is a uniformly distributed random number. The expression state of a gene/node is represented by a real number and the initial state of the network was set to be a random number for each gene/node drawn from a uniform distribution between 0 and 10. The rate of expression of a gene under the influence of a single regulator is given by a sigmoid function of the level of expression of that single regulator (Fig. 1 A). $$\:{f}_{i}\left(x\right)\:=\:\frac{A}{1+{e}^{B-\:C\cdot\:x}}+1$$ Where x is the level of expression of the regulator. A, B and C are parameters controlling the maximum value, the inflection point (threshold) and slope at that point for this function. in the present study, we selected \(\:A\:=\:5,\:B\:=\:5,\:C\:=\:1\:(Yin\:et\:al.,\:2021)\:\) . If \(\:{f}_{i}\) is the rate of expression induced by the regulator \(\:i\) , the overall expression of the target ( \(\:{E}_{x}\) ) under the influence of \(\:k\) regulators will be given by: $$\:{E}_{x}=\frac{\:{\prod\:}_{i}^{kpos}{f}_{i}\left({E}_{i}\right)}{\:{\prod\:}_{j}^{kneg}{f}_{j}\left({E}_{j}\right)}$$ Where \(\:kpos\) is the number of positive direct regulators and \(\:kneg\) is the number of direct negative regulators. The expression state of each gene was simulated over a succession for 400 discrete time points, with the state of a given node at time \(\:t\) being a function of the state of its regulators at time \(\:t-1\) . In order to conduct knocking-down simulations, for a specific gene, its expression was multiplied by 0.01 across the whole dynamics. We used robust Z-Score to quantify the expression changes between control and mutated conditions, defined as: $$\:{Z}_{i}\:=\:\frac{\stackrel{\sim}{X}{{i\:}_{mutated}}_{}-\:{\stackrel{\sim}{Xi}}_{control}}{{qi}_{75\left(control\right)\:}-\:{qi}_{25\left(control\right)}}$$ where \(\:{\stackrel{\sim}{Xi}}_{mutated}\) is the median of the expression levels of the \(\:ith\) gene when it is knocked–out, \(\:{\stackrel{\sim}{Xi}}_{control}\) is the median of the expression levels of the \(\:ith\) gene in control conditions, \(\:q{i}_{75\:\left(control\right)},\:q{i}_{25\left(control\right)}\) are the 75th and 25th quantiles, respectively, of the expression levels of the \(\:ith\) gene in control conditions. In both the computer model and in natural transcriptome data, Spearman correlations were used as a measure of statistical associations between regulators and target. Statistical significance of these associations were calculated numerically by generating null (expected) distributions based on 10 thousand random background gene pairs, and multiple testing corrections done using the Benjamini–Hochberg method. Significance of individually knocking–out direct regulators over their corresponding targets, we first calculated the target’s robust z‑score after individually knocking out each direct regulator. We then obtained an equally sized set of robust z‑scores derived from knockdowns of random background (unrelated) genes. These two distributions were compared using a Wilcoxon rank‑sum test to determine whether perturbations caused by direct‑regulator knock‑outs differed significantly from those produced by random knock‑outs. P‑values were adjusted for multiple testing with the Benjamini–Hochberg procedure. Human transcriptome and gene perturbation data Transcriptome-wide gene expression profiles from human tissues were obtained from the FANTOM5 dataset (Functional ANnoTation Of the Mammalian genome) 10 . Annotated targets, for a total of 1792 human transcription factors compiled from the TransFact database and TFLink were used 11 , 12 . Gene perturbation data were obtained from consensus perturbation signatures derived from the CMAP database project 13 based on the LINCS L1000 platform, compiling transcriptional responses on 979 hallmark genes to 4323 gene perturbations experiments 14 , 15 . Model implementation, simulations, in addition to network measurements, experimental and synthetic expression data analysis and statistical tests, were carried out in Python version 3.1.1. Results General Lack of statistical association between regulators and regulatory targets in GRN In order to assess the degree of statistical association between direct regulators and their targets in GRNs, we used a simple computational network model of regulatory interactions capable of simulating gene expression dynamics based on simple kinetic properties (see methods, 9 ). Using this simple model, we simulated expression data in a network of 1,000 nodes (genes) for 400 time points (iterations), and retrieved the list of direct regulator–target pairs from the underlying network interaction architecture, which explicitly specifies signed (positive and negative) regulatory interactions between individual pairs of genes. Spearman correlation coefficients were computed between the resulting expression level of each regulator and its annotated targets across the 400 simulated time points. Figure 1 A shows a typical scatter plot illustrating a complete lack of association between the activity of a regulator and its direct target. Consistent with this, the correlation distribution for all regulator-target pairs with positive regulatory signs, revealed no apparent positive association between their expression profiles (Fig. 1 B). Similarly, for negative regulatory interactions, expression levels of regulators and their targets showed no trace of a bias toward negative correlations (Fig. 1 C and D). These results show that, even in a simple regulatory network, where regulatory interactions are fully known and well-defined, no statistical association is observed between individual regulators and their direct targets. A similar lack of statistical association has been reported in natural transcriptional networks between the expression level of transcription factors and their direct known targets 5 , 16 – 20 . In order to confirm and further characterize this lack of association between regulators and their targets in natural transcriptomes, we examined human expression data from the FANTOM5 dataset (Functional ANnoTation Of the Mammalian genome), which compiles RNA expression profiles across a broad range of tissues and cell types 10 . We next identified, in this dataset, annotated targets for a total of 1792 human transcription factors compiled in the TransFact database 21 , 22 . Figure 1 E shows the example of a well known transcription factor (MYC) and one of its known targets (Snurportin 1), where a clear lack of association is apparent. A similar lack of association is observed across the whole set of Myc's currently annotated targets (Fig. 1 F). To assess if this lack of statistical association between regulators and targets is in any way different to random expectations both in the synthetic regulatory network model as well as in natural transcriptomes, we obtained, first for the computational model, the distribution of absolute Spearman correlation coefficients (as a measure of absolute statistical association independent of sign) between expression levels of twenty thousand regulator-target pairs, and compared it with the observed distribution of the same number of random gene pairs. As shown in Fig. 2 A, virtually identical distributions are obtained, suggesting that statistical associations between direct regulators and targets are indistinguishable from random expectations. To directly determine the statistical significance of each individual regulator-target pair association, we used the random gene pair distribution as the null distribution to compute the individual probability value for each actual regulator-target pair. Figure 2 B shows the ordered adjusted p values (FDR), for each of the actual regulator-target pairs in the simulated model, effectively revealing a complete lack of statistical significance for virtually all direct regulatory interactions when compared to random expectations. To ascertain if the same lack of statistical significance is observed in natural transcriptomes, we extracted expression data for 1000 random human genes and a random sample of 20 experimentally confirmed transcriptional regulators targeting each of these genes (totalling 20 thousand regulator-target pairs, encompassing all 1792 annotated human transcription factors) and computed the Spearman correlation coefficient for all these regulatory interactions. Figure 2 C shows the resulting distribution of absolute Spearman correlation coefficients between expression levels of all regulator-target pairs, compared with the distribution obtained from the same number of random gene pairs. As was the case for the synthetic gene regulatory model, identical distributions were obtained. Following the same approach used in Fig. 2 B, we determined the statistical significance of each individual regulator-target pair association by using the random gene pair distribution as the null distribution to compute the individual probability value for each actual regulator-target pair. Figure 2 D reveals again a complete lack of statistical significance for all direct regulatory interactions when compared to random expectations. Taken together, these results confirm and demonstrate a general lack of statistical association between the expression of individual regulators and that of their direct targets in gene regulatory networks. Moreover, the fact that the exact same lack of statistical association between direct regulators and targets is observed in a simple computer model of regulatory network as well as in natural transcriptomes, strongly suggest that this lack of association is a general property of regulatory networks rather than the result of the particular complexity of the cellular transcriptional process specifically taking place in biological systems. Knocking down direct regulators has no stronger effect on their targets than random background genes in GRNs. Having established a generalized lack of statistical association between regulators and targets in regulatory networks, we next asked if there is an effective causal association between regulators and their direct targets. To this end, in the synthetic model, we systematically performed knock-down experiments by individually perturbing each gene in the network (see methods). Following each knock-down, we compared the expression levels of all genes in the network, between the mutated condition and the unperturbed control and expressed this difference using robust z-score (see methods), thereby obtaining a perturbation matrix compiling the effects on every single gene after individually knocking-down all genes in the model. To determine whether direct regulators produce stronger perturbation effects on their targets than unrelated genes, we compared – for each gene – the distribution of absolute Z-scores from 20 random direct regulators against those of an equally-sized random sample of unrelated background genes. As shown in Fig. 3 A, only 8.1% of target genes showed statistically significant levels of perturbation (FDR < 0.05, dashed line), indicating that for most genes, even in a simplistic GRN model, direct regulators do not exert, on average, a stronger influence on their direct targets than random background genes (Fig. 3 A). We then asked if the same lack of effective causal association is observed in natural transcriptomes. In this regard, several independent studies have reported, for a number of specific genes, a lack of significant changes in gene expression in response to experimental knock out/down of some or various of their associated transcription factors 7 , 8 . In order to ascertain the generality of these findings, we used consensus perturbation data for 796 transcription factors and their effects on 978 LINCs landmark genes compiled in the CMAP dataset 13, across a variety of human cell lines 14 , 15 . As in the previous analysis, for each landmark gene we compared the distribution of absolute perturbation Z-scores resulting from knocking down 20 random known transcriptional regulators of that particular gene, with an equally-sized random sample of unrelated background genes. As shown in Fig. 3 B, only 0.1% of target genes showed statistically significant levels of average perturbation (FDR < 0.05, dashed line) after knocking down their direct transcriptional regulators, demonstrating that the average effect of direct regulators on a given target gene is no different than the effect of random background genes. Again, the fact that the exact same lack of causal association between direct regulators and targets is observed in a simple computer model of regulatory network as well as in actual natural transcriptomes, strongly suggest that this lack of association is a general property of regulatory networks rather than the result of the particular complexity of the cellular transcriptional process specifically taking place in biological systems. It is worth inspecting the perturbation profile of any given target gene after individually knocking down all other genes in the system. This is easily done in the computational model where all genes in the system can be experimentally knocked down. When, for a given gene, the perturbation effects of knocking down all other genes are sorted in descending order of impact (three representative examples are shown in Fig. 4 A, B and C), it becomes clear that direct regulators (arrows in Fig. 4 ) do not cluster exclusively at the top of the resulting ranking but are instead regularly dispersed throughout the ranked distribution. Moreover, for every direct regulator there is a large number of seemingly unrelated background genes that have a much stronger effect on the target. This exact same pattern is observed in both the computational model and the actual experimental perturbation data compiled in the CMAP database, bearing in mind that the effective regulatory strength of any given perturbation can only be measured relative to a defined set of 4,324 individual gene perturbations. In the latter case only the top 1000 out of 4,324 perturbations are shown in Fig. 4 D-F, for all the 978 LINCs landmark “target” genes present in this dataset. Taken together these results show that at the whole network level, actual regulatory relationships between genes cannot be trivially ascribed to direct transcription factor–target interactions but are, instead, diffusely distributed throughout the transcriptional network. Having established that direct regulators (i.e., transcription factors) do not generally exert stronger effects on the expression of their target genes than background genes, we sought to define a data-driven notion of regulatory influence. We thus introduce the term effective regulators to refer to those genes that, after knocking down all genes in the system, exert the strongest absolute perturbation effect upon a given target, regardless of whether or not they include direct transcriptional regulators (i. e., transcription factors). Accordingly, knocking down every gene in the system would necessarily have a graded effect on any target gene (as illustrated in Fig. 4 ), and we can therefore define the top K effective regulators of a given gene as the top K genes exerting the strongest impact on it. This circumstance is of particular importance when only partial perturbation data is available as is the case for the CMAP database, where only a subset of 4,323 genes were assessed against a total of 978 landmark “target” genes, and effective regulators can only be defined relative to this particular subset of putative regulators. We can evaluate the expected contrast between direct and effective regulators by directly comparing their respective relative perturbation effects on their corresponding targets. As shown in Fig. 5 A, we found that effective regulators in the simulated model consistently exerted a stronger regulatory effect on their targets than direct regulators (Mann-Whitney test, p = 3.77e-160). Using the experimental consensus perturbation data compiled in the CMAP dataset, contrasting the median relative perturbation effect of annotated transcription factors per gene with an equal number of top relative effective regulators (relative to the set of 4,323 included in that dataset) we find an even stronger contrast between relative effective and direct regulators (Mann-Whitney test, p = 2.67e-153. Figure 5 B) These results demonstrate that effective regulators, as defined here, display a significantly more robust causal control on gene expression than canonical transcriptional regulators. Topological analysis of regulatory networks One way to characterize the contrasting properties of effective regulatory interactions is by examining the resulting network representation composed of only these interactions in contrast to the network resulting from direct regulatory interactions ( i. e., transcription factor-target canonical interactions). To this end, we simulated 20 independent synthetic GRNs with identical parameters and network densities. For each network, we identified both the top-1 effective and direct regulator for each of the 1,000 genes, yielding on average a significantly lower number of unique effective regulators than unique direct regulators across all independent simulations (see Fig. 6 A, B and C). A significantly shorter average path length, and much larger average clustering was found in effective regulatory networks, suggesting a more structured organization and a larger number of lnter-cluster interactions in effective networks than in direct regulatory networks (Fig. 6 D). This latter feature is in line with a much larger number of strongly connected components in effective regulatory networks than those found in direct networks (Fig. 6 F). These results show that effective regulatory networks display distinct topological features that distinguish them from the expected structure based on direct regulatory interactions. To further characterize the level of divergence (and convergence) between effective and direct regulatory interactions and their associated network representations, we looked at the overlap between the unique top 1 effective and direct regulators observed in the synthetic model. As shown in Fig. 7 A, there is a strong but not complete overlap between these two sets of regulators. When all the genes in the system are ordered from the highest to the lowest regulatory influence for all their corresponding target genes, we can indicate in the resulting graphic matrix the relative position of all direct and non direct regulators (black dots vs grey dots in Fig. 7 B), on their target genes (matrix columns). In line with the strong but not complete overlap between the top 1 effective and direct regulators, we can observe a graded enrichment of direct regulators among those genes exerting the strongest regulatory influence. This overlap, however, rapidly fades as we descend along the list of ordered genes. In order to quantify this divergence we generated networks including only the top1, top2, top3 and so on, direct regulators on the one hand, together with networks only including the top 1, 2, 3 and so on, effective regulators on the other. For every top-k value we measured the Frobenius distance between the corresponding adjacency matrices (which quantifies the topological differences between networks), as well as the Jaccard index of the overlap between the top K direct and effective regulators. As shown in Figs. 7 C and D, a rapid increase in divergence (Frobenius distance) and decrease in overlap of regulators (Jaccard index) is observed as we increase the number of top regulators included in the networks. This result shows that although a strong but not complete overlap is observed between effective and direct regulatory networks when the top regulators in either case are considered, this overlap rapidly decreases as successive top regulators (effective on the one hand, and direct on the other) are added to each network. In order to test if this structural divergence is also observed in natural transcriptomes, we again used CMAP´s consensus perturbation data to identify top regulators among the 4323 perturbed genes included in that study. Because this set of tested genes represent a relatively small subset of all protein coding genes in the human genome, only a small proportion of them are likely to be actual top regulators (either direct or effective) and we would, therefore, expect substantial noise in the expected behaviour of these genes. Still if the properties we have so far described are general to regulatory networks, we would expect a statistical bias towards an enrichment of direct regulators among top ranking genes with the strongest regulatory influence, even within this restricted set of 4323 experimentally perturbed genes. As shown in Fig. 7 E when unique top 1 regulators (relative to this set of 4323 perturbed genes) are extracted and compared with the top 1 transcription factors, a highly statistically significant overlap is observed, demonstrating a statistical enrichment of direct regulators among top effective (relative) regulators. When all 4323 perturbed genes are ordered from the highest to the lowest regulatory influence for all their corresponding LINCs landmark target genes (Matrix columns, in Fig. 7 F, n = 978), we can again indicate in the resulting graphic matrix the relative position of all direct and non direct relative regulators (black dots vs grey dots in Fig. 7 F). In line with the results shown in Fig. 4 D-F, direct regulators are thoroughly dispersed along the ordering of regulatory influences for each target gene (columns). In this case the enrichment of direct regulators among top (relative) effective regulators, is still present but much less obvious. In order to determine if the gradual divergence in the network structure observed in the simple synthetic model is also parallelled in natural transcriptomes, we conducted the same analysis of Figs. 7 C-D, and generated the corresponding networks based on the top k relative direct regulators on the one hand, and again contrasted them with networks based on the top k relative effective regulators on the one hand. As shown in Figs. 7 G and H, a rapid increase in divergence (Frobenius distance) and decrease in overlap of regulators (Jaccard index) is also observed as we increase the number of top regulators of either class included in each network. Taken together, these results show that in both the simulated model as well as in natural transcriptomes a gradual divergence in network structure is observed as successive top regulators (effective on the one hand, and direct on the other) are added to the network. Again, the fact that the exact same topological divergence is observed between the direct and effective regulatory network in both a simple computer model of regulatory network as well as in actual natural transcriptomes, strongly suggest that this divergence is a general property of regulatory networks rather than the result of the particular complexity of the cellular transcriptional process specifically taking place in biological systems. Statistical signatures of effective regulators Given that effective regulators constitute a reduced and functionally enriched subset of genes, we next investigated whether they exhibit distinct properties under control (non-perturbed) conditions, Using the synthetic gene expression data, we computed the mean expression and standard deviation across 200 samples in our simulations. As shown in Figs. 8 A, effective regulators tend to have expression levels and variances that cluster near the global median, as opposed to adopting extreme values or being evenly distributed through the corresponding distributions, this suggests that the strongest effective regulators tend to cluster around moderate levels of expression and expression variance. This association between level of expression and variance is readily evident throughout the whole expression and variance spectrum as a strong correlation between these two statistics, suggesting an emerging coupling between activity levels and variance in regulatory networks in general (right scatter plot in Fig. 8 A). To test whether the variance and expression level clustering of effective regulators could be leveraged as an identifying feature of effective regulators, we used a very small random set of them, and asked if they could be used to retrieve the whole set of 214 top 3 unique regulators with the highest regulatory influence (absolute perturbation value) on all 1000 genes in the synthetic system. To this end, exploiting our previous finding that top 1 effective regulators tend to be enriched in direct regulators, we randomly selected 20 direct regulators included among the top 1 regulators and obtained their centroid (mean coordinates in the expression-variance space). Next, we identified the 214 genes top genes with the shortest euclidean distance to this centroid (indicated as blue dots in Fig. 8 B, dashed lines indicating the resulting centroid). As shown in Fig. 8 C, the overlap between this set of estimated regulators (based on a small set of known effective regulators) and the actual set of 214 top 3 effective regulators was over 99%. This result shows that the expression and variance clustering of effective regulators represents a distinct statistical feature of these genes, to the point that this feature can be used to retrieve and identify potential effective regulators in the absence of an exhaustive knock-down screening. In order to ascertain if this distinct statistical feature is also present in natural transcriptomes, we used independent expression data for 20 000 protein coding genes compiled in the Fantom 5 dataset and obtained, for each of them their mean expression level and variance. Using the consensus perturbation data derived from the CMAP dataset (absolute perturbation index, from all 4323 perturbed genes) we identified all 675 unique top 3 genes with the highest (relative) regulatory influence on all 978 LINCS landmark target genes. As shown in Fig. 8 D, we also find a strong association between expression and variance across the whole human transcriptome as well as a distinct tight clustering of the 675 unique top 3 regulators in the expression level and variance space, even bearing in mind that only a small proportion of these of 675 genes are likely to be actual top effective regulators. As we did in the synthetic system, we used a very small random set of them as anchor genes, and asked if they could be used to retrieve a statistically significant proportion this set of top regulators, bearing in mind that these 675 are the top regulators relative to the limited sample of 4323 perturbed genes included in the CMAP study. As in our previous analysis we randomly selected 20 direct regulators included among the top 1 regulators, obtained their centroid (mean coordinates in the expression-variance space) and identified the 3244 genes top genes with the shortest euclidean distance to this centroid (indicated as blue dots in Fig. 8 E, dashed lines indicating the resulting centroid). 3244 genes closest to the centroid were selected because only about one fifth of them would also be present among the 3244 genes included in the CMAP dataset. As shown in Fig. 8 F, the overlap between this set of estimated regulators (based on a small set of known anchor regulators) and the actual set of 675 top 3 relative regulators was 209 out of 655 (excluding the anchor genes). A numerical simulation of the expected overlap using over 5000 random samplings demonstrates that the observed overlap was highly significant, demonstrating an enrichment of relative top regulators among those genes closest to the centroid in the expression and variance space, thereby demonstrating the validity of this statistical feature in natural transcriptomes. These results confirm that effective regulators in real systems also exhibit a distinct statistical feature, (namely a tight clustering around moderate values of expression and variance), and that this expression-derived metrics can be leveraged to identify and retrieve effective regulators even in partial or sparsely sampled datasets, without the need of an exhaustive whole-genome knock-down screening. Discussion Natural transcriptomes consist of dynamical systems where the activity of a particular gene in the network at a given time is dynamically driven by the activity of its immediate transcriptional regulators. Given the resulting network of regulatory interactions, it is widely assumed that a level of statistical association, and a definite causal association is to be expected between regulators and regulatory 3 – 5 . Here we confirm previous findings revealing that direct regulators and their targets display no stronger statistical associations than random gene pairs in both a synthetic GRN with known signed edges (positive and negative direct regulation) as well as the human transcriptome, as regulator–target Spearman correlations are no stronger than the statistical association observed in random background gene pairs. Aside from the fact that statistical associations do not reveal causal directionality 23 , 24 , simple expression correlations consistently fail to identify true regulatory links independently of the causal directionality 25 , 17 , 16 . Even in much simpler bacterial models, both activating and repressing interactions have been found to show a positive correlation with their targets, and transcription factor-target networks are no more consistent with measured expression than random networks, revealing a widespread mismatch between annotated regulatory influence and the sign of the statistical association actually measured 18 . These discrepancies have been ascribed to the intricate complexity of the cellular transcription process, involving a number of intermediate events such as protein synthesis, chromatin accessibility, transcription factor paralog redundancy, co-factors and complex signaling cascades up-stream of transcriptional initiation, among others in addition to uncertainties regarding transcription factor target annotations 6 – 8 , 19 , 21 , 26 – 28 . It is, however important to stress that the exact same lack of statistical association between direct regulators and targets is observed in a simple computer model of regulatory network as well as in natural transcriptomes, strongly suggesting that this lack of association is a general property of regulatory networks rather than the result of the particular complexity of the cellular transcriptional process specific to biological systems. We assessed the causal association between regulators and targets in regulatory networks, using first the synthetic model, where we systematically performed knock-down experiments by individually knocking down direct regulators and looking at the effect on their targets, and found that the average effect of direct regulators on a given target gene is no different than the effect of random background genes. We confirmed this exact same lack of causal association in the human transcriptome using consensus perturbation data compiled in the CMAP database, where we could show for a total of 796 transcription factors that direct transcriptional regulators had no stronger causal regulatory effect on their targets than random background genes. This result is consistent with a growing number of studies reporting, for a number of specific genes, a lack of significant changes in gene expression in response to experimental knock out/down of some, or various, of their associated transcription factors 7 , 8 .. Based on this fact, we can conclude that at the whole network level, actual regulatory relationships between genes cannot be trivially ascribed to direct transcription factor–target interactions (direct regulators) but are, instead, diffusely distributed throughout the transcriptional network. Furthermore, the fact that the exact same lack of causal association between direct regulators and targets is observed in a simple computer model of regulatory network, as well as in actual natural transcriptomes, again suggests that this lack of association is a general property of regulatory networks rather than the result of the distinct complexity of the cellular transcriptional process. Motivated by these findings, we introduced the definition of effective regulators—the gene or set of genes that cause the strongest absolute perturbation effect on a given target— as a more functional and evidence-based unit of regulatory influence. By definition the top ranking members of this class of regulators consistently outperform the full set of annotated direct regulators for any given gene (Figs. 4 and 5 ). However, this definition assigns a regulatory weight to all and every gene in the network regardless of whether or not they include direct regulators. As shown in Fig. 4 , the regulatory weight of all genes in the network quickly decays as the ranking list grows suggesting that the bulk of the regulatory influence is concentrated in a small number of top regulators. As effective regulators implies (for a defined number of top k regulators) an “effective” regulatory network, a preliminary exploration of the contrasting network structure between direct and effective regulatory networks reveal substantial statistical differences and points of convergence. Of note is the fact that the top regulators (both in the synthetic system and in the human transcriptome) tend to be statistically enriched in direct regulators (i.e transcription factors), the opposite, however, not being true. An obvious difficulty of this definition of effective regulatory relationships is that their full and exact identification and pairwise annotation would necessarily require an exhaustive perturbation screening of all the genes in the system and their full transcriptome wide effects across a number of cell types or other conditions, which at the moment constitutes a hugely impractical task in any model species. Given this important difficulty it is of particular relevance that effective regulators in real systems as well as in the synthetic GRN model, exhibit a distinct statistical feature, (namely a tight clustering around moderate values of expression and variance), and that this expression-derived metrics can be leveraged to identify and retrieve effective regulators even in partial or sparsely sampled datasets, possibly without the need of an exhaustive genome-wide knock-down screening. In summary, In this study, we compared the behavior of a natural transcriptome to that of a generic regulatory network model, and draw qualitative and quantitative parallels that reveal general regulatory network properties. In both the synthetic GRN and transcriptomic data we can observe similar qualitative and quantitative behaviours. The fact that these concordant patterns arise in a deliberately simplistic model as well as in natural transcriptomes strongly suggest that the lack of statistical and causal association between direct regulators and targets, the existence of effective regulators, their network properties and their distinct statistical behavior, are not an exclusive feature of the biological transcriptional machinery, but rather an emergent set of general properties of regulatory networks. Declarations Funding: Diego Arturo Velázquez Trejo received a studentship from the National Ministry of Science, Technology and Humanities (SECHITI). Author Contribution Diego Velazquez and Humberto Gutiérrez conceived and designed the study, wrote the main manuscript text and prepared all the figures. Jorge Islas, Bertha Rueda, Daniela Toro, Leticia Ramirez y Angel Olascoaga conducted data compilation and processing. Acknowledgement This paper serves as a fulfullmet of D.A.V-T for obtaining a M.Sc. degree in the Posgrado en Ciencias Biológicas UNAM. We thank the Secretaría de Ciencia, Humanidades, Tecnología e Innovación (SECIHTI) for funding and for the support of this research through a graduate scholarship. Data Availability Computing consensus transcriptional profiles for LINCS L1000 perturbations Daniel Himmelstein, Caty Chung ThinkLab (2015-03-26) https://doi.org/f3mqwc DOI: 10.15363/thinklab.d43FANTOM 5 Data: Lizio, M. et al. Gateways to the FANTOM5 promoter level mammalian expression atlas. Genome Biol 16, 22 (2015). References Islam, S. & Bhattacharya, S. Dynamical systems theory as an organizing principle for single-cell biology. NPJ Syst. Biol. Appl. 11 , 85 (2025). Li, R. et al. Inferring gene regulatory networks using transcriptional profiles as dynamical attractors. PLoS Comput. Biol. 19 , e1010991 (2023). Wang, Y. et al. NetREX-CF integrates incomplete transcription factor data with gene expression to reconstruct gene regulatory networks. Commun. Biol. 5 , 1282 (2022). Lu, J. et al. Causal network inference from gene transcriptional time-series response to glucocorticoids. PLoS Comput. Biol. 17 , e1008223 (2021). Zaborowski, A. B. & Walther, D. Determinants of correlated expression of transcription factors and their target genes. Nucleic Acids Res. 48 , 11347–11369 (2020). Kernfeld, E., Keener, R., Cahan, P. & Battle, A. Transcriptome data are insufficient to control false discoveries in regulatory network inference. Cells 15 , 709–724e13 (2024). Le, P. Glucocorticoid receptor-dependent gene regulatory networks. PLoS Genet. 1 , e16 (2005). Hu, Z., Killion, P. J. & Iyer, V. R. Genetic reconstruction of a functional transcriptional regulatory network. Nat. Genet. 39 , 683–687 (2007). Yin, W., Mendoza, L., Monzon-Sandoval, J., Urrutia, A. O. & Gutierrez, H. Emergence of co-expression in gene regulatory networks. PLOS ONE . 16 , e0247671 (2021). Lizio, M. et al. Gateways to the FANTOM5 promoter level mammalian expression atlas. Genome Biol. 16 , 22 (2015). Plaisier, C. L. et al. Causal Mechanistic Regulatory Network for Glioblastoma Deciphered Using Systems Genetics Network Analysis. Cell. Syst. 3 , 172–186 (2016). Liska, O. et al. TFLink: an integrated gateway to access transcription factor-target gene interactions for multiple species. Database (Oxford) baac083 (2022). (2022). Subramanian, A. et al. A Next Generation Connectivity Map: L1000 Platform and the First 1,000,000 Profiles. Cell 171 , 1437–1452e17 (2017). Himmelstein, D. S. et al. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. eLife 6 , e26726 (2017). Himmelstein, D. & Chung, C. Computing consensus transcriptional profiles for LINCS L1000 perturbations. Thinklab https://doi.org/ (2015). 10.15363/thinklab.d43 Muley, V. Y. & König, R. Human transcriptional gene regulatory network compiled from 14 data resources. Biochimie 193 , 115–125 (2022). Marbach, D. et al. Wisdom of crowds for robust gene network inference. Nat. Methods . 9 , 796–804 (2012). Larsen, S. J., Röttger, R., Schmidt, H. H. H. W. & Baumbach J. E. coli gene regulatory networks are inconsistent with gene expression data. Nucleic Acids Res. 47 , 85–92 (2019). Gitter, A. et al. Backup in gene regulatory networks explains differences between binding and knockout results. Mol. Syst. Biol. 5 , 276 (2009). Lambert, S. A. et al. Hum. Transcription Factors Cell 172 , 650–665 (2018). Dai, Z., Dai, X., Xiang, Q. & Feng, J. Robustness of transcriptional regulatory program influences gene expression variability. BMC Genom. 10 , 573 (2009). Plaisier, C. L. et al. Causal Mechanistic Regulatory Network for Glioblastoma Deciphered Using Systems Genetics Network Analysis. cels 3, 172–186 (2016). Stuart, J. M., Segal, E., Koller, D. & Kim, S. K. A gene-coexpression network for global discovery of conserved genetic modules. Science 302 , 249–255 (2003). Aldrich, J. Correlations Genuine and Spurious in Pearson and Yule. Stat. Sci. 10 , 364–376 (1995). Faith, J. J. et al. Large-scale mapping and validation of Escherichia coli transcriptional regulation from a compendium of expression profiles. PLoS Biol. 5 , e8 (2007). Kim, H. D. & O’Shea, E. K. A quantitative model of transcription factor–activated gene expression. Nat. Struct. Mol. Biol. 15 , 1192–1198 (2008). Haynes, B. C. et al. Mapping functional transcription factor networks from gene expression data. Genome Res. 23 , 1319–1328 (2013). Cusanovich, D. A., Pavlovic, B., Pritchard, J. K. & Gilad, Y. The functional consequences of variation in transcription factor binding. PLoS Genet. 10 , e1004226 (2014). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 04 May, 2026 Reviews received at journal 24 Apr, 2026 Reviews received at journal 20 Apr, 2026 Reviews received at journal 13 Apr, 2026 Reviews received at journal 11 Apr, 2026 Reviewers agreed at journal 10 Apr, 2026 Reviewers agreed at journal 09 Apr, 2026 Reviews received at journal 07 Apr, 2026 Reviewers agreed at journal 07 Apr, 2026 Reviewers agreed at journal 07 Apr, 2026 Reviewers agreed at journal 07 Apr, 2026 Reviewers agreed at journal 07 Apr, 2026 Reviewers invited by journal 07 Apr, 2026 Editor invited by journal 27 Mar, 2026 Editor assigned by journal 24 Jan, 2026 Submission checks completed at journal 24 Jan, 2026 First submitted to journal 19 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8643144","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":621519658,"identity":"7ad6b38c-0e1e-41f5-96bd-d3aa6ca54d90","order_by":0,"name":"Diego Arturo Velázquez-Trejo","email":"","orcid":"","institution":"National Institute of Genomic Medicine","correspondingAuthor":false,"prefix":"","firstName":"Diego","middleName":"Arturo","lastName":"Velázquez-Trejo","suffix":""},{"id":621519666,"identity":"7bfa18e3-9f23-4e16-a365-e2296acd8e91","order_by":1,"name":"Jorge Adrián Islas-Ortíz","email":"","orcid":"","institution":"National Institute of Genomic Medicine","correspondingAuthor":false,"prefix":"","firstName":"Jorge","middleName":"Adrián","lastName":"Islas-Ortíz","suffix":""},{"id":621519675,"identity":"75227934-50bc-4b73-a3a4-161c0726f48e","order_by":2,"name":"Bertha Rueda-Zarazúa","email":"","orcid":"","institution":"European Bioinformatics Institute","correspondingAuthor":false,"prefix":"","firstName":"Bertha","middleName":"","lastName":"Rueda-Zarazúa","suffix":""},{"id":621519678,"identity":"5b80a0f5-0c2f-4cf7-855e-4a9ec063a863","order_by":3,"name":"Daniela del-Toro","email":"","orcid":"","institution":"National Institute of Genomic Medicine","correspondingAuthor":false,"prefix":"","firstName":"Daniela","middleName":"","lastName":"del-Toro","suffix":""},{"id":621519680,"identity":"8781130f-7e91-4709-b855-18aba2e6c6ad","order_by":4,"name":"Olascoaga-Del Angel","email":"","orcid":"","institution":"National Autonomous University of Mexico","correspondingAuthor":false,"prefix":"","firstName":"Olascoaga-Del","middleName":"","lastName":"Angel","suffix":""},{"id":621519681,"identity":"bd706785-78a3-49e4-a14e-5b0cbffd0bb9","order_by":5,"name":"Leticia Ramirez-Lugo","email":"","orcid":"","institution":"National Institute of Genomic Medicine","correspondingAuthor":false,"prefix":"","firstName":"Leticia","middleName":"","lastName":"Ramirez-Lugo","suffix":""},{"id":621519683,"identity":"69973231-4e50-4291-a2b8-e06fbcb7e570","order_by":6,"name":"Humberto Gutiérrez","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABFUlEQVRIiWNgGAWjYBACAxCRAGYyHzgAF2ZswKuFsQGihS2BBC0QJo8BQhifFnP+NeYPHlTcy+fnP/PxcGGOjV2/dAPzi487GKINDmDXYjnjjWFDwpliy5kNZzccnrktLXnmnANsljPPMOTOxGGTwY0zhg2JbQkGBgd7Nxzm3XY42eBGApsxbxtDbj8uv4C1/EswsD/M8wCo5X+yPUxLGy4t53uAWhqAtrDxMAC1HLAzkEhgfozfFrbCGQnHEgwkzrAZALUkJ0jcSGxjnHlGArdfzh/e8PFHTYIBf//hx595t9nZ889IPvzh4w6b3A04QoxBIgGVn9jAwNgmARTHoR4I+NHMsgdi5g+41Y+CUTAKRsEIBAAAV2XML3I5OQAAAABJRU5ErkJggg==","orcid":"","institution":"National Institute of Genomic Medicine","correspondingAuthor":true,"prefix":"","firstName":"Humberto","middleName":"","lastName":"Gutiérrez","suffix":""}],"badges":[],"createdAt":"2026-01-19 21:53:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8643144/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8643144/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106778906,"identity":"77dbb90a-5e1b-40d8-a67c-624dc24fe244","added_by":"auto","created_at":"2026-04-13 11:16:20","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":289971,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLack of correlation between target genes and their direct regulators. \u003c/strong\u003eA) Scatter plot showing the log-transformed expression levels of a representative positively acting regulator and its target gene in the synthetic gene regulatory system. B) Distribution of Spearman correlation coefficients computed between all target genes and their corresponding positive direct regulators across simulated conditions. C) Scatter plot of a representative negatively acting regulator and its target gene in the synthetic system. D) Distribution of Spearman correlation coefficients computed between all target genes and their corresponding negative direct regulators in the synthetic model. E) Representative scatter plot of log-transformed expression levels of the transcription factor MYC and one of its annotated targets (SNUPN). F) Distribution of Spearman correlation coefficients for all known human MYC targets.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8643144/v1/3b1b708b97c50e26c383b812.png"},{"id":107480388,"identity":"f583dbee-a9c0-4d07-80a4-50675cda7b16","added_by":"auto","created_at":"2026-04-22 02:09:35","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":361485,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCorrelations between direct regulators and their target genes are no stronger than random gene pairs in both synthetic and human expression data\u003c/strong\u003e A) Correlation distribution of absolute Spearman correlation coefficients for 20 000 identified regulator-target gene pairs (blue) and an equally sized number of randomly selected gene pairs (red) in the synthetic regulatory model. B) Statistical significance of each individual regulator-target pair association was estimated using the random gene pair distribution as the null distribution to compute the individual probability value for each actual regulator-target pair. Chart shows the ordered adjusted p values (FDR, Benjamini-Hochberg correction), for each of the actual regulator-target pairs in the simulated model. C) Correlation distribution of absolute Spearman correlation coefficients for 20000 regulator-target gene pairs (20 annotated regulators per target, for a total of 1000 targets, encompassing 1792 known transcription factors) (blue) and an equally sized number of randomly selected gene pairs (red), measured in human expression data D) Following the same procedure as in panel B, to assess the statistical significance of each individual regulator-target pair, the graph shows the ordered adjusted p values (Log (FDR), Benjamini-Hochberg correction), for each of the actual regulator-target pairs in for 20000 regulator target pairs measured in human expression data.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8643144/v1/141d4788ed8b1c1e6a78c92d.png"},{"id":106960188,"identity":"55d448f6-cd38-4340-b74d-474fe6460a64","added_by":"auto","created_at":"2026-04-15 09:19:15","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":303676,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eKnock-down of direct transcriptional regulators does not produce stronger effects on target gene expression than random gene perturbations in the synthetic model. \u003c/strong\u003eWe systematically performed knock-down experiments by individually perturbing each gene in the network. Following each knock-down, we compared the expression levels of all genes in the network, between the mutated condition and the unperturbed control and expressed this difference using robust z-score. For each gene, we compared the distribution of absolute Z-scores from 20 random direct regulators (20 random annotated transcription factors in the case of human perturbation data) against those of an equally-sized random sample of unrelated background genes followed by a Mann-Whitney test, to compute the corresponding the adjusted p value (FDR) . A) Chart showing the ordered adjusted p values for all 1000 target genes in the synthetic system. B) Identical analysis carried on on 1000 human genes, using consensus perturbation signatures derived from the CMAP project. The chart shows the ordered adjusted p values for all 979 LINCS hallmark target genes included in the CMAP study target genes in the synthetic system. Insets show comparisons between the effect of direct regulators (Dreg, blue) and random background genes (Bckg, red) for genes in the top, median and bottom ranking (human target gene symbols are indicated for panel B).\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8643144/v1/4a61432ab403b6011b1e9d49.png"},{"id":106778907,"identity":"15c09e9d-7874-4176-a14b-c84b46daecf9","added_by":"auto","created_at":"2026-04-13 11:16:20","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":823963,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTranscriptional regulators are not enriched among the most impactful perturbations to their targets. A–F)\u003c/strong\u003e For each of the six representative selected target genes (Gene 391, Gene 650 and Gene 556, from the synthetic system, and RBM15B, LPGAT1, and PRKCQ from Human perturbation data ), the plots show the ordered perturbation index (absolute Z-scores) resulting from the individual knock-down of all 1000 genes in a synthetic system (or top 1000 out of 4323 perturbed genes in the CMAP study). Perturbed genes are sorted in descending order according to the magnitude of their perturbation effect on the target gene's expression. Each point corresponds to one perturbed gene. The green dot marks the gene whose knock-down produced the maximum absolute Z-score (top‑1 effective perturbation). Arrows indicate in each case the positions of direct transcriptional regulators of the annotated target.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8643144/v1/9b903b62b8a36e0173944167.png"},{"id":106778912,"identity":"a6599c0d-aac2-4cc0-93c9-144b6a7b5e67","added_by":"auto","created_at":"2026-04-13 11:16:20","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":226149,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEffective regulators outperform direct regulators in both the synthetic model and natural transcriptomes\u003c/strong\u003e. For each gene or node in the system, the median of the absolute perturbation effect (robust Z-score) of all its direct regulators was obtained alongside that of an equal number of top effective regulators and the ratio of these medians to the background median effect was calculated. A) Violin plots showing the distribution of log-perturbation ratios for direct and effective regulators (left and right plot respectively) in the synthetic model. B) Same analysis as in panel A, but using human gene expression perturbation data. DR median: Direct regulator median; ER media: effective regulator median, bckgrnd median: background median.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8643144/v1/257aae3c5eec6e7885cd16db.png"},{"id":106778910,"identity":"92fe9bd3-2de7-43d7-bc46-085350ce480e","added_by":"auto","created_at":"2026-04-13 11:16:20","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":899154,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEffective regulatory networks exhibit distinct topological organization compared with direct regulatory networks. \u003c/strong\u003e(A–B) Representative layouts of a direct regulatory network (A, blue) and its corresponding effective regulatory network (B, red) reconstructed from the same simulation, with nodes representing genes and arrows represent regulatory interactions. (C–F) Violin plots summarizing topological the indicated network topological metrics across 20 independently simulated networks. For each system, two directed graphs were reconstructed: an effective network obtained from the top three most effective regulators per gene (red) and a direct network obtained from the top three known direct regulators (blue). Metrics shown are (C) total number of unique regulators, (D) average global clustering coefficient, (E) average shortest path length, and (F) number of strongly connected components. Boxes indicate medians and interquartile ranges. P-values correspond to Mann–Whitney–Wilcoxon tests comparing effective versus direct networks.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8643144/v1/859857b497c22cfa20741ec0.png"},{"id":106778913,"identity":"1cf46497-1410-4087-8960-c5faec38fd0e","added_by":"auto","created_at":"2026-04-13 11:16:20","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":798880,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStructural divergence between effective and direct regulatory networks derived from both synthetic and human transcriptome.\u003c/strong\u003e A) Venn diagram comparing the set of top‑1 effective regulators and top‑1 direct regulators (i.e., known direct regulators) for each gene in a representative synthetic regulatory network. B) Matrix visualization of the rank positions of direct regulators (purple dots) and background genes (grey dots), within the sorted Z-score effects for 100 randomly selected target genes (columns) in the synthetic system (matrix columns). C) Frobenius distance between adjacency matrices of effective and direct regulatory networks as a function of k (top‑k direct or effective regulators per gene) in the synthetic model. D) Jaccard index between the same direct and effective networks, showing a decline of the overlap between direct and effective regulators with increasing k. E) Venn diagram comparing, for Human gene perturbation data, the set of relative top‑1 effective regulators and relative top‑1 unique direct regulators (relative to the limited set of 4323 perturbed genes included in the CMAP screening). F) Matrix visualization of the rank positions of relative direct regulators (purple dots) and background genes (grey dots), within the sorted Z-score effects for all 926 LINCS landmark target genes (matrix columns). G) Frobenius distance between adjacency matrices of relative effective and relative direct regulatory networks as a function of k (top‑k direct or effective regulators per gene) for networks constructed from human perturbation data. H) Jaccard index between the same direct and effective networks, showing a decline of the overlap between relative direct and effective regulators with increasing k.\u003c/p\u003e","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8643144/v1/17adf81d74244264e0198918.png"},{"id":106778909,"identity":"efba0a0b-ceae-4c49-add7-0802019525d2","added_by":"auto","created_at":"2026-04-13 11:16:20","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":460515,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEffective regulators display a distinct statistical signature in both synthetic regulatory networks and human transcriptome-derived networks. \u003c/strong\u003e(A) Distributions of log10-transformed mean expression (top histogram) and log10-transformed expression variance (bottom histogram) under control conditions for effective regulators (red) and all other genes (grey), right scatter plot:Two-dimensional plot of log10-transformed expression standard deviation versus log10-transformed mean expression, showing in red the clustered position of effective regulators for data derived from the synthetic model. (B) Two-dimensional plot of log10-transformed expression standard deviation versus log10-transformed mean expression showing the estimated centroid of effective regulation in the synthetic system based on twenty random anchor points (yellow dots). The centroid coordinates were estimated using a subset of effective–direct regulators (n = 20 out of N = 1,000), and predicted effective regulators were obtained by selecting the top 241 genes closest to the centroid (measured as euclidean distance). (C) Venn diagram illustrating the enrichment of effective regulators among the genes closest to the estimated centroid in the synthetic system; the corresponding null distribution of this enrichment and actual overlap is shown in the bottom chart. (D) Two-dimensional plot of log10-transformed expression standard deviation versus log10-transformed mean expression in human expression data, with the 655 relative top 3 effective regulators, as derived from the CMAP study, highlighted in red. (E) Predicted regulators (indicated in blue) based on a sample of 20 random top 1 relative direct regulators (yellow dots) derived from the CMAP perturbation study. The estimated centroid of effective regulation based on those 20 anchor points is indicated by dashed red lines. (F) Venn diagram illustrating the enrichment of top 3 relative effective regulators among the genes closest to the estimated centroid in the combined as estimated from human expression data; the corresponding null distribution is shown and actual overlap (red line) is shown in the bottom chart.\u003c/p\u003e","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8643144/v1/65991195b7f301ef98e4bd34.png"},{"id":107485515,"identity":"7969facc-57e8-4609-b56e-508e8a21982a","added_by":"auto","created_at":"2026-04-22 02:35:17","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3492090,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8643144/v1/8b993e83-bbf9-478e-944c-80a61793c2ce.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Effective regulatory relationships between genes differ from direct transcription factor-target interactions, and exhibit a distinct statistical signature","fulltext":[{"header":"Introduction","content":"\u003cp\u003eRegulatory networks can be represented as a set of interactions between regulatory elements and their individual targets \u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Natural transcriptomes are dynamical systems where the activity of a particular gene in the network at a given time is dynamically driven by the activity of its immediate transcriptional regulators, and therefore a level of statistical association, and a definite causal association is to be expected between regulators and regulatory targets \u003csup\u003e\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. While transcriptional networks are a particular instance of regulatory network, where the complete regulator-target interaction map is not fully known, attempts to infer actual regulatory interactions between transcription factors and targets based on their potential statistical association have led to mixed results \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e,\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. Along these lines, although numerous approaches to detect regulatory interactions based on coexpression measurements have been proposed, a sizable number of studies have reported a general lack of statistical association between regulators and targets, with a more restricted number of reports documenting a significant regulator-target association, when compared with random pairs of background genes \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. In addition, several independent studies have reported no significant changes in gene expression of target genes, in response to experimental knock out/down of a growing number of specific transcription factors \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. These discrepancies have been ascribed to the particular complexity of the cellular transcription process, involving a number of intermediate events such as protein synthesis, chromatin accessibility, transcription factor paralog redundancy, co-factors and complex signaling cascades up-stream of transcriptional initiation, among others \u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. Thus, in general, whether trivial statistical and causal associations, at the level of direct regulatory interactions are to be expected in regulatory networks, remains unclear.\u003c/p\u003e \u003cp\u003eHere we used a simple computational gene regulatory network model to test general assumptions regarding regulatory interactions, and specifically ask whether statistical and causal associations are to be expected between regulators and direct targets in regulatory networks in general \u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. We next contrast our findings with empirical human expression and gene perturbation data, and find that in both cases, regulators tend to display little to no correlation with their direct targets and that, knocking down their activity has, on average, no stronger effect on their targets than background genes. We then introduce the concept of effective regulators, defined as those genes that exert the strongest causal influence on a given gene and characterize their key general statistical properties.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eDescription of the model\u003c/h2\u003e \u003cp\u003eWe used a simple GRN model developed by \u003csup\u003e9\u003c/sup\u003e in which the regulatory structure is represented by a set of interaction pairs composed where nodes represent genes and directed edges linking them represent regulatory positive and negative interactions. We specifically used networks consisting of 1000 nodes, each node being regulated by a variable number of randomly chosen nodes, drawn from the entire network, with a minimum (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:min\\)\u003c/span\u003e\u003c/span\u003e) number of 22 regulators per node. We set this minimum because the simulated networks, on average, exhibited a density closer to that estimated from transcriptomic networks. The number of regulators (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:k\\:=\\:kneg\\:+\\:kpos\\)\u003c/span\u003e\u003c/span\u003e) per node (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:i\\)\u003c/span\u003e\u003c/span\u003e) was set to follow a power law distribution defined as:\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:{k}_{i}\\:=\\:min\\:+\\:round\\left(\\frac{1}{{x}^{2}}\\right)\\)\u003c/span\u003e \u003c/span\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:\\)\u003c/span\u003e \u003c/span\u003e \u003c/p\u003e \u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:min\\:=\\:22\\)\u003c/span\u003e\u003c/span\u003e is the minimum number of regulators and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\:\\sim\\:Unif(0,\\:1)\\)\u003c/span\u003e\u003c/span\u003e is a uniformly distributed random number.\u003c/p\u003e \u003cp\u003eThe expression state of a gene/node is represented by a real number and the initial state of the network was set to be a random number for each gene/node drawn from a uniform distribution between 0 and 10. The rate of expression of a gene under the influence of a single regulator is given by a sigmoid function of the level of expression of that single regulator (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003cdiv id=\"Equa\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:{f}_{i}\\left(x\\right)\\:=\\:\\frac{A}{1+{e}^{B-\\:C\\cdot\\:x}}+1$$\u003c/div\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003eWhere x is the level of expression of the regulator. A, B and C are parameters controlling the maximum value, the inflection point (threshold) and slope at that point for this function. in the present study, we selected \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:A\\:=\\:5,\\:B\\:=\\:5,\\:C\\:=\\:1\\:(Yin\\:et\\:al.,\\:2021)\\:\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eIf \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{f}_{i}\\)\u003c/span\u003e\u003c/span\u003e is the rate of expression induced by the regulator \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:i\\)\u003c/span\u003e\u003c/span\u003e, the overall expression of the target (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{E}_{x}\\)\u003c/span\u003e\u003c/span\u003e) under the influence of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:k\\)\u003c/span\u003e\u003c/span\u003e regulators will be given by:\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:{E}_{x}=\\frac{\\:{\\prod\\:}_{i}^{kpos}{f}_{i}\\left({E}_{i}\\right)}{\\:{\\prod\\:}_{j}^{kneg}{f}_{j}\\left({E}_{j}\\right)}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:kpos\\)\u003c/span\u003e\u003c/span\u003e is the number of positive direct regulators and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:kneg\\)\u003c/span\u003e\u003c/span\u003e is the number of direct negative regulators.\u003c/p\u003e \u003cp\u003eThe expression state of each gene was simulated over a succession for 400 discrete time points, with the state of a given node at time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:t\\)\u003c/span\u003e\u003c/span\u003e being a function of the state of its regulators at time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:t-1\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eIn order to conduct knocking-down simulations, for a specific gene, its expression was multiplied by 0.01 across the whole dynamics. We used robust Z-Score to quantify the expression changes between control and mutated conditions, defined as:\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:{Z}_{i}\\:=\\:\\frac{\\stackrel{\\sim}{X}{{i\\:}_{mutated}}_{}-\\:{\\stackrel{\\sim}{Xi}}_{control}}{{qi}_{75\\left(control\\right)\\:}-\\:{qi}_{25\\left(control\\right)}}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\stackrel{\\sim}{Xi}}_{mutated}\\)\u003c/span\u003e\u003c/span\u003e is the median of the expression levels of the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:ith\\)\u003c/span\u003e\u003c/span\u003e gene when it is knocked\u0026ndash;out, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\stackrel{\\sim}{Xi}}_{control}\\)\u003c/span\u003e\u003c/span\u003e is the median of the expression levels of the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:ith\\)\u003c/span\u003e\u003c/span\u003e gene in control conditions, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:q{i}_{75\\:\\left(control\\right)},\\:q{i}_{25\\left(control\\right)}\\)\u003c/span\u003e\u003c/span\u003e are the 75th and 25th quantiles, respectively, of the expression levels of the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:ith\\)\u003c/span\u003e\u003c/span\u003e gene in control conditions.\u003c/p\u003e \u003cp\u003eIn both the computer model and in natural transcriptome data, Spearman correlations were used as a measure of statistical associations between regulators and target.\u003c/p\u003e \u003cp\u003eStatistical significance of these associations were calculated numerically by generating null (expected) distributions based on 10 thousand random background gene pairs, and multiple testing corrections done using the Benjamini\u0026ndash;Hochberg method.\u003c/p\u003e \u003cp\u003eSignificance of individually knocking\u0026ndash;out direct regulators over their corresponding targets, we first calculated the target\u0026rsquo;s robust z‑score after individually knocking out each direct regulator. We then obtained an equally sized set of robust z‑scores derived from knockdowns of random background (unrelated) genes. These two distributions were compared using a Wilcoxon rank‑sum test to determine whether perturbations caused by direct‑regulator knock‑outs differed significantly from those produced by random knock‑outs. P‑values were adjusted for multiple testing with the Benjamini\u0026ndash;Hochberg procedure.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eHuman transcriptome and gene perturbation data\u003c/h3\u003e\n\u003cp\u003eTranscriptome-wide gene expression profiles from human tissues were obtained from the FANTOM5 dataset (Functional ANnoTation Of the Mammalian genome) \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. Annotated targets, for a total of 1792 human transcription factors compiled from the TransFact database and TFLink were used \u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e,\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Gene perturbation data were obtained from consensus perturbation signatures derived from the CMAP database project \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e based on the LINCS L1000 platform, compiling transcriptional responses on 979 hallmark genes to 4323 gene perturbations experiments \u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eModel implementation, simulations, in addition to network measurements, experimental and synthetic expression data analysis and statistical tests, were carried out in Python version 3.1.1.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eGeneral Lack of statistical association between regulators and regulatory targets in GRN\u003c/h2\u003e \u003cp\u003eIn order to assess the degree of statistical association between direct regulators and their targets in GRNs, we used a simple computational network model of regulatory interactions capable of simulating gene expression dynamics based on simple kinetic properties (see methods, \u003csup\u003e9\u003c/sup\u003e). Using this simple model, we simulated expression data in a network of 1,000 nodes (genes) for 400 time points (iterations), and retrieved the list of direct regulator\u0026ndash;target pairs from the underlying network interaction architecture, which explicitly specifies signed (positive and negative) regulatory interactions between individual pairs of genes. Spearman correlation coefficients were computed between the resulting expression level of each regulator and its annotated targets across the 400 simulated time points. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA shows a typical scatter plot illustrating a complete lack of association between the activity of a regulator and its direct target. Consistent with this, the correlation distribution for all regulator-target pairs with positive regulatory signs, revealed no apparent positive association between their expression profiles (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB). Similarly, for negative regulatory interactions, expression levels of regulators and their targets showed no trace of a bias toward negative correlations (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC and D). These results show that, even in a simple regulatory network, where regulatory interactions are fully known and well-defined, no statistical association is observed between individual regulators and their direct targets.\u003c/p\u003e \u003cp\u003eA similar lack of statistical association has been reported in natural transcriptional networks between the expression level of transcription factors and their direct known targets \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan additionalcitationids=\"CR17 CR18 CR19\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. In order to confirm and further characterize this lack of association between regulators and their targets in natural transcriptomes, we examined human expression data from the FANTOM5 dataset (Functional ANnoTation Of the Mammalian genome), which compiles RNA expression profiles across a broad range of tissues and cell types \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. We next identified, in this dataset, annotated targets for a total of 1792 human transcription factors compiled in the TransFact database \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE shows the example of a well known transcription factor (MYC) and one of its known targets (Snurportin 1), where a clear lack of association is apparent. A similar lack of association is observed across the whole set of Myc's currently annotated targets (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eF).\u003c/p\u003e \u003cp\u003eTo assess if this lack of statistical association between regulators and targets is in any way different to random expectations both in the synthetic regulatory network model as well as in natural transcriptomes, we obtained, first for the computational model, the distribution of absolute Spearman correlation coefficients (as a measure of absolute statistical association independent of sign) between expression levels of twenty thousand regulator-target pairs, and compared it with the observed distribution of the same number of random gene pairs. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA, virtually identical distributions are obtained, suggesting that statistical associations between direct regulators and targets are indistinguishable from random expectations. To directly determine the statistical significance of each individual regulator-target pair association, we used the random gene pair distribution as the null distribution to compute the individual probability value for each actual regulator-target pair. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB shows the ordered adjusted p values (FDR), for each of the actual regulator-target pairs in the simulated model, effectively revealing a complete lack of statistical significance for virtually all direct regulatory interactions when compared to random expectations. To ascertain if the same lack of statistical significance is observed in natural transcriptomes, we extracted expression data for 1000 random human genes and a random sample of 20 experimentally confirmed transcriptional regulators targeting each of these genes (totalling 20 thousand regulator-target pairs, encompassing all 1792 annotated human transcription factors) and computed the Spearman correlation coefficient for all these regulatory interactions. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC shows the resulting distribution of absolute Spearman correlation coefficients between expression levels of all regulator-target pairs, compared with the distribution obtained from the same number of random gene pairs. As was the case for the synthetic gene regulatory model, identical distributions were obtained. Following the same approach used in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB, we determined the statistical significance of each individual regulator-target pair association by using the random gene pair distribution as the null distribution to compute the individual probability value for each actual regulator-target pair. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD reveals again a complete lack of statistical significance for all direct regulatory interactions when compared to random expectations. Taken together, these results confirm and demonstrate a general lack of statistical association between the expression of individual regulators and that of their direct targets in gene regulatory networks. Moreover, the fact that the exact same lack of statistical association between direct regulators and targets is observed in a simple computer model of regulatory network as well as in natural transcriptomes, strongly suggest that this lack of association is a general property of regulatory networks rather than the result of the particular complexity of the cellular transcriptional process specifically taking place in biological systems.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eKnocking down direct regulators has no stronger effect on their targets than random background genes in GRNs.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eHaving established a generalized lack of statistical association between regulators and targets in regulatory networks, we next asked if there is an effective causal association between regulators and their direct targets. To this end, in the synthetic model, we systematically performed knock-down experiments by individually perturbing each gene in the network (see methods). Following each knock-down, we compared the expression levels of all genes in the network, between the mutated condition and the unperturbed control and expressed this difference using robust z-score (see methods), thereby obtaining a perturbation matrix compiling the effects on every single gene after individually knocking-down all genes in the model.\u003c/p\u003e \u003cp\u003eTo determine whether direct regulators produce stronger perturbation effects on their targets than unrelated genes, we compared \u0026ndash; for each gene \u0026ndash; the distribution of absolute Z-scores from 20 random direct regulators against those of an equally-sized random sample of unrelated background genes. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA, only 8.1% of target genes showed statistically significant levels of perturbation (FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05, dashed line), indicating that for most genes, even in a simplistic GRN model, direct regulators do not exert, on average, a stronger influence on their direct targets than random background genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe then asked if the same lack of effective causal association is observed in natural transcriptomes. In this regard, several independent studies have reported, for a number of specific genes, a lack of significant changes in gene expression in response to experimental knock out/down of some or various of their associated transcription factors \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. In order to ascertain the generality of these findings, we used consensus perturbation data for 796 transcription factors and their effects on 978 LINCs landmark genes compiled in the CMAP dataset 13, across a variety of human cell lines \u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. As in the previous analysis, for each landmark gene we compared the distribution of absolute perturbation Z-scores resulting from knocking down 20 random known transcriptional regulators of that particular gene, with an equally-sized random sample of unrelated background genes. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB, only 0.1% of target genes showed statistically significant levels of average perturbation (FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05, dashed line) after knocking down their direct transcriptional regulators, demonstrating that the average effect of direct regulators on a given target gene is no different than the effect of random background genes. Again, the fact that the exact same lack of causal association between direct regulators and targets is observed in a simple computer model of regulatory network as well as in actual natural transcriptomes, strongly suggest that this lack of association is a general property of regulatory networks rather than the result of the particular complexity of the cellular transcriptional process specifically taking place in biological systems.\u003c/p\u003e \u003cp\u003eIt is worth inspecting the perturbation profile of any given target gene after individually knocking down all other genes in the system. This is easily done in the computational model where all genes in the system can be experimentally knocked down. When, for a given gene, the perturbation effects of knocking down all other genes are sorted in descending order of impact (three representative examples are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA, B and C), it becomes clear that direct regulators (arrows in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e) do not cluster exclusively at the top of the resulting ranking but are instead regularly dispersed throughout the ranked distribution. Moreover, for every direct regulator there is a large number of seemingly unrelated background genes that have a much stronger effect on the target. This exact same pattern is observed in both the computational model and the actual experimental perturbation data compiled in the CMAP database, bearing in mind that the effective regulatory strength of any given perturbation can only be measured relative to a defined set of 4,324 individual gene perturbations. In the latter case only the top 1000 out of 4,324 perturbations are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD-F, for all the 978 LINCs landmark \u0026ldquo;target\u0026rdquo; genes present in this dataset.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTaken together these results show that at the whole network level, actual regulatory relationships between genes cannot be trivially ascribed to direct transcription factor\u0026ndash;target interactions but are, instead, diffusely distributed throughout the transcriptional network.\u003c/p\u003e \u003cp\u003eHaving established that direct regulators (i.e., transcription factors) do not generally exert stronger effects on the expression of their target genes than background genes, we sought to define a data-driven notion of regulatory influence. We thus introduce the term \u003cem\u003eeffective regulators\u003c/em\u003e to refer to those genes that, after knocking down all genes in the system, exert the strongest absolute perturbation effect upon a given target, regardless of whether or not they include direct transcriptional regulators (i. e., transcription factors). Accordingly, knocking down every gene in the system would necessarily have a graded effect on any target gene (as illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e), and we can therefore define the top K effective regulators of a given gene as the top K genes exerting the strongest impact on it. This circumstance is of particular importance when only partial perturbation data is available as is the case for the CMAP database, where only a subset of 4,323 genes were assessed against a total of 978 landmark \u0026ldquo;target\u0026rdquo; genes, and effective regulators can only be defined relative to this particular subset of putative regulators.\u003c/p\u003e \u003cp\u003eWe can evaluate the expected contrast between direct and effective regulators by directly comparing their respective relative perturbation effects on their corresponding targets. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA, we found that effective regulators in the simulated model consistently exerted a stronger regulatory effect on their targets than direct regulators (Mann-Whitney test, p\u0026thinsp;=\u0026thinsp;3.77e-160). Using the experimental consensus perturbation data compiled in the CMAP dataset, contrasting the median relative perturbation effect of annotated transcription factors per gene with an equal number of top relative effective regulators (relative to the set of 4,323 included in that dataset) we find an even stronger contrast between relative effective and direct regulators (Mann-Whitney test, p\u0026thinsp;=\u0026thinsp;2.67e-153. Figure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB)\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThese results demonstrate that effective regulators, as defined here, display a significantly more robust causal control on gene expression than canonical transcriptional regulators.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eTopological analysis of regulatory networks\u003c/h3\u003e\n\u003cp\u003eOne way to characterize the contrasting properties of effective regulatory interactions is by examining the resulting network representation composed of only these interactions in contrast to the network resulting from direct regulatory interactions ( i. e., transcription factor-target canonical interactions). To this end, we simulated 20 independent synthetic GRNs with identical parameters and network densities. For each network, we identified both the top-1 effective and direct regulator for each of the 1,000 genes, yielding on average a significantly lower number of unique effective regulators than unique direct regulators across all independent simulations (see Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eA, B and C). A significantly shorter average path length, and much larger average clustering was found in effective regulatory networks, suggesting a more structured organization and a larger number of lnter-cluster interactions in effective networks than in direct regulatory networks (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eD). This latter feature is in line with a much larger number of strongly connected components in effective regulatory networks than those found in direct networks (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eF). These results show that effective regulatory networks display distinct topological features that distinguish them from the expected structure based on direct regulatory interactions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo further characterize the level of divergence (and convergence) between effective and direct regulatory interactions and their associated network representations, we looked at the overlap between the unique top 1 effective and direct regulators observed in the synthetic model. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eA, there is a strong but not complete overlap between these two sets of regulators. When all the genes in the system are ordered from the highest to the lowest regulatory influence for all their corresponding target genes, we can indicate in the resulting graphic matrix the relative position of all direct and non direct regulators (black dots vs grey dots in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eB), on their target genes (matrix columns). In line with the strong but not complete overlap between the top 1 effective and direct regulators, we can observe a graded enrichment of direct regulators among those genes exerting the strongest regulatory influence. This overlap, however, rapidly fades as we descend along the list of ordered genes. In order to quantify this divergence we generated networks including only the top1, top2, top3 and so on, direct regulators on the one hand, together with networks only including the top 1, 2, 3 and so on, effective regulators on the other. For every top-k value we measured the Frobenius distance between the corresponding adjacency matrices (which quantifies the topological differences between networks), as well as the Jaccard index of the overlap between the top K direct and effective regulators. As shown in Figs.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eC and D, a rapid increase in divergence (Frobenius distance) and decrease in overlap of regulators (Jaccard index) is observed as we increase the number of top regulators included in the networks. This result shows that although a strong but not complete overlap is observed between effective and direct regulatory networks when the top regulators in either case are considered, this overlap rapidly decreases as successive top regulators (effective on the one hand, and direct on the other) are added to each network.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn order to test if this structural divergence is also observed in natural transcriptomes, we again used CMAP\u0026acute;s consensus perturbation data to identify top regulators among the 4323 perturbed genes included in that study. Because this set of tested genes represent a relatively small subset of all protein coding genes in the human genome, only a small proportion of them are likely to be actual top regulators (either direct or effective) and we would, therefore, expect substantial noise in the expected behaviour of these genes. Still if the properties we have so far described are general to regulatory networks, we would expect a statistical bias towards an enrichment of direct regulators among top ranking genes with the strongest regulatory influence, even within this restricted set of 4323 experimentally perturbed genes.\u003c/p\u003e \u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eE when unique top 1 regulators (relative to this set of 4323 perturbed genes) are extracted and compared with the top 1 transcription factors, a highly statistically significant overlap is observed, demonstrating a statistical enrichment of direct regulators among top effective (relative) regulators. When all 4323 perturbed genes are ordered from the highest to the lowest regulatory influence for all their corresponding LINCs landmark target genes (Matrix columns, in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eF, n\u0026thinsp;=\u0026thinsp;978), we can again indicate in the resulting graphic matrix the relative position of all direct and non direct relative regulators (black dots vs grey dots in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eF). In line with the results shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD-F, direct regulators are thoroughly dispersed along the ordering of regulatory influences for each target gene (columns). In this case the enrichment of direct regulators among top (relative) effective regulators, is still present but much less obvious.\u003c/p\u003e \u003cp\u003eIn order to determine if the gradual divergence in the network structure observed in the simple synthetic model is also parallelled in natural transcriptomes, we conducted the same analysis of Figs.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eC-D, and generated the corresponding networks based on the top k relative direct regulators on the one hand, and again contrasted them with networks based on the top k relative effective regulators on the one hand. As shown in Figs.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eG and H, a rapid increase in divergence (Frobenius distance) and decrease in overlap of regulators (Jaccard index) is also observed as we increase the number of top regulators of either class included in each network. Taken together, these results show that in both the simulated model as well as in natural transcriptomes a gradual divergence in network structure is observed as successive top regulators (effective on the one hand, and direct on the other) are added to the network. Again, the fact that the exact same topological divergence is observed between the direct and effective regulatory network in both a simple computer model of regulatory network as well as in actual natural transcriptomes, strongly suggest that this divergence is a general property of regulatory networks rather than the result of the particular complexity of the cellular transcriptional process specifically taking place in biological systems.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eStatistical signatures of effective regulators\u003c/h2\u003e \u003cp\u003eGiven that effective regulators constitute a reduced and functionally enriched subset of genes, we next investigated whether they exhibit distinct properties under control (non-perturbed) conditions,\u003c/p\u003e \u003cp\u003eUsing the synthetic gene expression data, we computed the mean expression and standard deviation across 200 samples in our simulations. As shown in Figs.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eA, effective regulators tend to have expression levels and variances that cluster near the global median, as opposed to adopting extreme values or being evenly distributed through the corresponding distributions, this suggests that the strongest effective regulators tend to cluster around moderate levels of expression and expression variance. This association between level of expression and variance is readily evident throughout the whole expression and variance spectrum as a strong correlation between these two statistics, suggesting an emerging coupling between activity levels and variance in regulatory networks in general (right scatter plot in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo test whether the variance and expression level clustering of effective regulators could be leveraged as an identifying feature of effective regulators, we used a very small random set of them, and asked if they could be used to retrieve the whole set of 214 top 3 unique regulators with the highest regulatory influence (absolute perturbation value) on all 1000 genes in the synthetic system. To this end, exploiting our previous finding that top 1 effective regulators tend to be enriched in direct regulators, we randomly selected 20 direct regulators included among the top 1 regulators and obtained their centroid (mean coordinates in the expression-variance space). Next, we identified the 214 genes top genes with the shortest euclidean distance to this centroid (indicated as blue dots in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eB, dashed lines indicating the resulting centroid). As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eC, the overlap between this set of estimated regulators (based on a small set of known effective regulators) and the actual set of 214 top 3 effective regulators was over 99%. This result shows that the expression and variance clustering of effective regulators represents a distinct statistical feature of these genes, to the point that this feature can be used to retrieve and identify potential effective regulators in the absence of an exhaustive knock-down screening. In order to ascertain if this distinct statistical feature is also present in natural transcriptomes, we used independent expression data for 20 000 protein coding genes compiled in the Fantom 5 dataset and obtained, for each of them their mean expression level and variance. Using the consensus perturbation data derived from the CMAP dataset (absolute perturbation index, from all 4323 perturbed genes) we identified all 675 unique top 3 genes with the highest (relative) regulatory influence on all 978 LINCS landmark target genes. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eD, we also find a strong association between expression and variance across the whole human transcriptome as well as a distinct tight clustering of the 675 unique top 3 regulators in the expression level and variance space, even bearing in mind that only a small proportion of these of 675 genes are likely to be actual top effective regulators. As we did in the synthetic system, we used a very small random set of them as anchor genes, and asked if they could be used to retrieve a statistically significant proportion this set of top regulators, bearing in mind that these 675 are the top regulators relative to the limited sample of 4323 perturbed genes included in the CMAP study. As in our previous analysis we randomly selected 20 direct regulators included among the top 1 regulators, obtained their centroid (mean coordinates in the expression-variance space) and identified the 3244 genes top genes with the shortest euclidean distance to this centroid (indicated as blue dots in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eE, dashed lines indicating the resulting centroid). 3244 genes closest to the centroid were selected because only about one fifth of them would also be present among the 3244 genes included in the CMAP dataset. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eF, the overlap between this set of estimated regulators (based on a small set of known anchor regulators) and the actual set of 675 top 3 relative regulators was 209 out of 655 (excluding the anchor genes). A numerical simulation of the expected overlap using over 5000 random samplings demonstrates that the observed overlap was highly significant, demonstrating an enrichment of relative top regulators among those genes closest to the centroid in the expression and variance space, thereby demonstrating the validity of this statistical feature in natural transcriptomes.\u003c/p\u003e \u003cp\u003eThese results confirm that effective regulators in real systems also exhibit a distinct statistical feature, (namely a tight clustering around moderate values of expression and variance), and that this expression-derived metrics can be leveraged to identify and retrieve effective regulators even in partial or sparsely sampled datasets, without the need of an exhaustive whole-genome knock-down screening.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eNatural transcriptomes consist of dynamical systems where the activity of a particular gene in the network at a given time is dynamically driven by the activity of its immediate transcriptional regulators. Given the resulting network of regulatory interactions, it is widely assumed that a level of statistical association, and a definite causal association is to be expected between regulators and regulatory \u003csup\u003e\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. Here we confirm previous findings revealing that direct regulators and their targets display no stronger statistical associations than random gene pairs in both a synthetic GRN with known signed edges (positive and negative direct regulation) as well as the human transcriptome, as regulator\u0026ndash;target Spearman correlations are no stronger than the statistical association observed in random background gene pairs. Aside from the fact that statistical associations do not reveal causal directionality \u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e, simple expression correlations consistently fail to identify true regulatory links independently of the causal directionality \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e, \u003csup\u003e17\u003c/sup\u003e, \u003csup\u003e16\u003c/sup\u003e. Even in much simpler bacterial models, both activating and repressing interactions have been found to show a positive correlation with their targets, and transcription factor-target networks are no more consistent with measured expression than random networks, revealing a widespread mismatch between annotated regulatory influence and the sign of the statistical association actually measured \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. These discrepancies have been ascribed to the intricate complexity of the cellular transcription process, involving a number of intermediate events such as protein synthesis, chromatin accessibility, transcription factor paralog redundancy, co-factors and complex signaling cascades up-stream of transcriptional initiation, among others in addition to uncertainties regarding transcription factor target annotations \u003csup\u003e\u003cspan additionalcitationids=\"CR7\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan additionalcitationids=\"CR27\" citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. It is, however important to stress that the exact same lack of statistical association between direct regulators and targets is observed in a simple computer model of regulatory network as well as in natural transcriptomes, strongly suggesting that this lack of association is a general property of regulatory networks rather than the result of the particular complexity of the cellular transcriptional process specific to biological systems.\u003c/p\u003e \u003cp\u003eWe assessed the causal association between regulators and targets in regulatory networks, using first the synthetic model, where we systematically performed knock-down experiments by individually knocking down direct regulators and looking at the effect on their targets, and found that the average effect of direct regulators on a given target gene is no different than the effect of random background genes. We confirmed this exact same lack of causal association in the human transcriptome using consensus perturbation data compiled in the CMAP database, where we could show for a total of 796 transcription factors that direct transcriptional regulators had no stronger causal regulatory effect on their targets than random background genes. This result is consistent with a growing number of studies reporting, for a number of specific genes, a lack of significant changes in gene expression in response to experimental knock out/down of some, or various, of their associated transcription factors \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e..\u003c/p\u003e \u003cp\u003eBased on this fact, we can conclude that at the whole network level, actual regulatory relationships between genes cannot be trivially ascribed to direct transcription factor\u0026ndash;target interactions (direct regulators) but are, instead, diffusely distributed throughout the transcriptional network. Furthermore, the fact that the exact same lack of causal association between direct regulators and targets is observed in a simple computer model of regulatory network, as well as in actual natural transcriptomes, again suggests that this lack of association is a general property of regulatory networks rather than the result of the distinct complexity of the cellular transcriptional process.\u003c/p\u003e \u003cp\u003eMotivated by these findings, we introduced the definition of effective regulators\u0026mdash;the gene or set of genes that cause the strongest absolute perturbation effect on a given target\u0026mdash; as a more functional and evidence-based unit of regulatory influence. By definition the top ranking members of this class of regulators consistently outperform the full set of annotated direct regulators for any given gene (Figs.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e and \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). However, this definition assigns a regulatory weight to all and every gene in the network regardless of whether or not they include direct regulators. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, the regulatory weight of all genes in the network quickly decays as the ranking list grows suggesting that the bulk of the regulatory influence is concentrated in a small number of top regulators. As effective regulators implies (for a defined number of top k regulators) an \u0026ldquo;effective\u0026rdquo; regulatory network, a preliminary exploration of the contrasting network structure between direct and effective regulatory networks reveal substantial statistical differences and points of convergence. Of note is the fact that the top regulators (both in the synthetic system and in the human transcriptome) tend to be statistically enriched in direct regulators (i.e transcription factors), the opposite, however, not being true.\u003c/p\u003e \u003cp\u003eAn obvious difficulty of this definition of effective regulatory relationships is that their full and exact identification and pairwise annotation would necessarily require an exhaustive perturbation screening of all the genes in the system and their full transcriptome wide effects across a number of cell types or other conditions, which at the moment constitutes a hugely impractical task in any model species. Given this important difficulty it is of particular relevance that effective regulators in real systems as well as in the synthetic GRN model, exhibit a distinct statistical feature, (namely a tight clustering around moderate values of expression and variance), and that this expression-derived metrics can be leveraged to identify and retrieve effective regulators even in partial or sparsely sampled datasets, possibly without the need of an exhaustive genome-wide knock-down screening.\u003c/p\u003e \u003cp\u003eIn summary, In this study, we compared the behavior of a natural transcriptome to that of a generic regulatory network model, and draw qualitative and quantitative parallels that reveal general regulatory network properties. In both the synthetic GRN and transcriptomic data we can observe similar qualitative and quantitative behaviours. The fact that these concordant patterns arise in a deliberately simplistic model as well as in natural transcriptomes strongly suggest that the lack of statistical and causal association between direct regulators and targets, the existence of effective regulators, their network properties and their distinct statistical behavior, are not an exclusive feature of the biological transcriptional machinery, but rather an emergent set of general properties of regulatory networks.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eFunding:\u003c/h2\u003e \u003cp\u003eDiego Arturo Vel\u0026aacute;zquez Trejo received a studentship from the National Ministry of Science, Technology and Humanities (SECHITI).\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eDiego Velazquez and Humberto Guti\u0026eacute;rrez conceived and designed the study, wrote the main manuscript text and prepared all the figures. Jorge Islas, Bertha Rueda, Daniela Toro, Leticia Ramirez y Angel Olascoaga conducted data compilation and processing.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThis paper serves as a fulfullmet of D.A.V-T for obtaining a M.Sc. degree in the Posgrado en Ciencias Biol\u0026oacute;gicas UNAM. We thank the Secretar\u0026iacute;a de Ciencia, Humanidades, Tecnolog\u0026iacute;a e Innovaci\u0026oacute;n (SECIHTI) for funding and for the support of this research through a graduate scholarship.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eComputing consensus transcriptional profiles for LINCS L1000 perturbations Daniel Himmelstein, Caty Chung ThinkLab (2015-03-26) https://doi.org/f3mqwc DOI: 10.15363/thinklab.d43FANTOM 5 Data: Lizio, M. et al. Gateways to the FANTOM5 promoter level mammalian expression atlas. Genome Biol 16, 22 (2015).\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eIslam, S. \u0026amp; Bhattacharya, S. Dynamical systems theory as an organizing principle for single-cell biology. \u003cem\u003eNPJ Syst. Biol. Appl.\u003c/em\u003e \u003cb\u003e11\u003c/b\u003e, 85 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, R. et al. Inferring gene regulatory networks using transcriptional profiles as dynamical attractors. \u003cem\u003ePLoS Comput. Biol.\u003c/em\u003e \u003cb\u003e19\u003c/b\u003e, e1010991 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, Y. et al. NetREX-CF integrates incomplete transcription factor data with gene expression to reconstruct gene regulatory networks. \u003cem\u003eCommun. Biol.\u003c/em\u003e \u003cb\u003e5\u003c/b\u003e, 1282 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLu, J. et al. Causal network inference from gene transcriptional time-series response to glucocorticoids. \u003cem\u003ePLoS Comput. Biol.\u003c/em\u003e \u003cb\u003e17\u003c/b\u003e, e1008223 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZaborowski, A. B. \u0026amp; Walther, D. Determinants of correlated expression of transcription factors and their target genes. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e48\u003c/b\u003e, 11347\u0026ndash;11369 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKernfeld, E., Keener, R., Cahan, P. \u0026amp; Battle, A. Transcriptome data are insufficient to control false discoveries in regulatory network inference. \u003cem\u003eCells\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e, 709\u0026ndash;724e13 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLe, P. Glucocorticoid receptor-dependent gene regulatory networks. \u003cem\u003ePLoS Genet.\u003c/em\u003e \u003cb\u003e1\u003c/b\u003e, e16 (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu, Z., Killion, P. J. \u0026amp; Iyer, V. R. Genetic reconstruction of a functional transcriptional regulatory network. \u003cem\u003eNat. Genet.\u003c/em\u003e \u003cb\u003e39\u003c/b\u003e, 683\u0026ndash;687 (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYin, W., Mendoza, L., Monzon-Sandoval, J., Urrutia, A. O. \u0026amp; Gutierrez, H. Emergence of co-expression in gene regulatory networks. \u003cem\u003ePLOS ONE\u003c/em\u003e. \u003cb\u003e16\u003c/b\u003e, e0247671 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLizio, M. et al. Gateways to the FANTOM5 promoter level mammalian expression atlas. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cb\u003e16\u003c/b\u003e, 22 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePlaisier, C. L. et al. Causal Mechanistic Regulatory Network for Glioblastoma Deciphered Using Systems Genetics Network Analysis. \u003cem\u003eCell. Syst.\u003c/em\u003e \u003cb\u003e3\u003c/b\u003e, 172\u0026ndash;186 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiska, O. et al. TFLink: an integrated gateway to access transcription factor-target gene interactions for multiple species. \u003cem\u003eDatabase (Oxford)\u003c/em\u003e baac083 (2022). (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSubramanian, A. et al. A Next Generation Connectivity Map: L1000 Platform and the First 1,000,000 Profiles. \u003cem\u003eCell\u003c/em\u003e \u003cb\u003e171\u003c/b\u003e, 1437\u0026ndash;1452e17 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHimmelstein, D. S. et al. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. \u003cem\u003eeLife\u003c/em\u003e \u003cb\u003e6\u003c/b\u003e, e26726 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHimmelstein, D. \u0026amp; Chung, C. Computing consensus transcriptional profiles for LINCS L1000 perturbations. Thinklab https://doi.org/ (2015). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.15363/thinklab.d43\u003c/span\u003e\u003cspan address=\"10.15363/thinklab.d43\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMuley, V. Y. \u0026amp; K\u0026ouml;nig, R. Human transcriptional gene regulatory network compiled from 14 data resources. \u003cem\u003eBiochimie\u003c/em\u003e \u003cb\u003e193\u003c/b\u003e, 115\u0026ndash;125 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarbach, D. et al. Wisdom of crowds for robust gene network inference. \u003cem\u003eNat. Methods\u003c/em\u003e. \u003cb\u003e9\u003c/b\u003e, 796\u0026ndash;804 (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLarsen, S. J., R\u0026ouml;ttger, R., Schmidt, H. H. H. W. \u0026amp; Baumbach J. E. coli gene regulatory networks are inconsistent with gene expression data. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e47\u003c/b\u003e, 85\u0026ndash;92 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGitter, A. et al. Backup in gene regulatory networks explains differences between binding and knockout results. \u003cem\u003eMol. Syst. Biol.\u003c/em\u003e \u003cb\u003e5\u003c/b\u003e, 276 (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLambert, S. A. et al. \u003cem\u003eHum. Transcription Factors Cell\u003c/em\u003e \u003cb\u003e172\u003c/b\u003e, 650\u0026ndash;665 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDai, Z., Dai, X., Xiang, Q. \u0026amp; Feng, J. Robustness of transcriptional regulatory program influences gene expression variability. \u003cem\u003eBMC Genom.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, 573 (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePlaisier, C. L. et al. Causal Mechanistic Regulatory Network for Glioblastoma Deciphered Using Systems Genetics Network Analysis. \u003cem\u003ecels\u003c/em\u003e 3, 172\u0026ndash;186 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStuart, J. M., Segal, E., Koller, D. \u0026amp; Kim, S. K. A gene-coexpression network for global discovery of conserved genetic modules. \u003cem\u003eScience\u003c/em\u003e \u003cb\u003e302\u003c/b\u003e, 249\u0026ndash;255 (2003).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAldrich, J. Correlations Genuine and Spurious in Pearson and Yule. \u003cem\u003eStat. Sci.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, 364\u0026ndash;376 (1995).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFaith, J. J. et al. Large-scale mapping and validation of Escherichia coli transcriptional regulation from a compendium of expression profiles. \u003cem\u003ePLoS Biol.\u003c/em\u003e \u003cb\u003e5\u003c/b\u003e, e8 (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim, H. D. \u0026amp; O\u0026rsquo;Shea, E. K. A quantitative model of transcription factor\u0026ndash;activated gene expression. \u003cem\u003eNat. Struct. Mol. Biol.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e, 1192\u0026ndash;1198 (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHaynes, B. C. et al. Mapping functional transcription factor networks from gene expression data. \u003cem\u003eGenome Res.\u003c/em\u003e \u003cb\u003e23\u003c/b\u003e, 1319\u0026ndash;1328 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCusanovich, D. A., Pavlovic, B., Pritchard, J. K. \u0026amp; Gilad, Y. The functional consequences of variation in transcription factor binding. \u003cem\u003ePLoS Genet.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, e1004226 (2014).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-8643144/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8643144/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eGene expression patterns are determined by the interactions between thousands of transcriptional regulators and their target genes. However, mounting evidence of an apparent lack of causal and statistical association between transcription factors and their targets, raises the question of where the effective regulation of any particular gene lies. Using both a simple computational gene regulatory network model together with human expression and perturbation data for 796 known transcription factors and their confirmed targets, we demonstrate a widespread and generalized lack of statistical and causal association between the activity of regulators and direct targets. We introduce the concept of effective regulators: genes that, upon experimental down-regulation, have the strongest impact on a given target. Effective regulators form a small set of genes that display a distinct statistical signature, and networks constructed from effective regulators exhibit greater clustering and connectivity than canonical TF-based networks. Our results show that, while at the transcriptome-level, actual regulatory relationships between genes do not coincide with direct transcription factor\u0026ndash;target interactions and are, instead, diffusely distributed throughout the transcriptional network, they can however be readily identified based on their distinct statistical properties.\u003c/p\u003e","manuscriptTitle":"Effective regulatory relationships between genes differ from direct transcription factor-target interactions, and exhibit a distinct statistical signature","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-13 11:16:15","doi":"10.21203/rs.3.rs-8643144/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-05-04T10:43:42+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-24T15:26:48+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-20T14:38:38+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-14T01:00:23+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-11T19:42:02+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"53613228871143556512127799280047235725","date":"2026-04-10T16:57:20+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"168989771744306899472334472654112552032","date":"2026-04-09T05:00:03+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-07T22:24:26+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"105393831155331358518030284464731401696","date":"2026-04-07T19:12:31+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"70449713044061307356084366201876782520","date":"2026-04-07T11:17:19+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"308104471370243317689067471663642794907","date":"2026-04-07T06:47:17+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"300973840835648595773995261156302981170","date":"2026-04-07T05:05:16+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-07T04:55:54+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-27T10:00:27+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-01-24T05:13:05+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-24T05:11:46+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-01-19T21:36:57+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0e061138-b01a-47f6-bda9-ee22a7e8a17b","owner":[],"postedDate":"April 13th, 2026","published":true,"recentEditorialEvents":[{"type":"decision","content":"Revision requested","date":"2026-05-04T10:43:42+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[{"id":66143549,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":66143550,"name":"Biological sciences/Genetics"},{"id":66143551,"name":"Biological sciences/Systems biology"}],"tags":[],"updatedAt":"2026-05-04T10:54:02+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-13 11:16:15","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8643144","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8643144","identity":"rs-8643144","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-23T02:00:01.238055+00:00
License: CC-BY-4.0