Human Alu elements promote the establishment and enhancement of piRNA-protein-coding gene targeting relationships

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Background: PIWI-interacting RNAs (piRNAs) are the most diverse category of small RNAs in animals. Recent evidence suggests that transposable elements (TEs) incorporated into protein-coding genes (PCGs) can be targeted by piRNAs. Thus, TEs might have a piRNA-mediated influence on organisms. In human PCGs, the extent to which TEs contribute to the presence of piRNA target sites remains to be assessed. Moreover, related evolutionary forces remain to be explored. Results: We found that the presence of Alu elements, a class of primate-specific TEs, in human PCGs almost always results in potential piRNA target sites. Additionally, we observed that Alu elements can exert a secondary influence on piRNAs and their potential target sites via interlocus gene conversion (IGC). This mutagenic process can homogenize piRNAs and their potential target sites, resulting in an excess of single nucleotide variants (SNVs) that increase piRNA-PCG targeting affinity in the genome. Although Aluelements facilitate the occurrence of SNVs that increase piRNA-PCG targeting affinity, these SNVs tend to show low allele frequencies in the human population. This footprint suggests that natural selection opposes the promotion effect of Alu elements on the formation of piRNA-PCG targeting relationships. Conclusions: Human Alu elements promote both the establishment and enhancement of piRNA-PCG targeting relationships. In addition, piRNA-PCG targeting relationships impose a piRNA-related selective constraint on the evolution of human PCGs. Our work suggests that the interplay between Alu elements and piRNAs is an important factor that influences the evolutionary trajectory of human PCGs.
Full text 147,907 characters · extracted from preprint-html · click to expand
Human Alu elements promote the establishment and enhancement of piRNA-protein-coding gene targeting relationships | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Human Alu elements promote the establishment and enhancement of piRNA-protein-coding gene targeting relationships Chong He, Hao Zhu This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2222130/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: PIWI-interacting RNAs (piRNAs) are the most diverse category of small RNAs in animals. Recent evidence suggests that transposable elements (TEs) incorporated into protein-coding genes (PCGs) can be targeted by piRNAs. Thus, TEs might have a piRNA-mediated influence on organisms. In human PCGs, the extent to which TEs contribute to the presence of piRNA target sites remains to be assessed. Moreover, related evolutionary forces remain to be explored. Results: We found that the presence of Alu elements, a class of primate-specific TEs, in human PCGs almost always results in potential piRNA target sites. Additionally, we observed that Alu elements can exert a secondary influence on piRNAs and their potential target sites via interlocus gene conversion (IGC). This mutagenic process can homogenize piRNAs and their potential target sites, resulting in an excess of single nucleotide variants (SNVs) that increase piRNA-PCG targeting affinity in the genome. Although Alu elements facilitate the occurrence of SNVs that increase piRNA-PCG targeting affinity, these SNVs tend to show low allele frequencies in the human population. This footprint suggests that natural selection opposes the promotion effect of Alu elements on the formation of piRNA-PCG targeting relationships. Conclusions: Human Alu elements promote both the establishment and enhancement of piRNA-PCG targeting relationships. In addition, piRNA-PCG targeting relationships impose a piRNA-related selective constraint on the evolution of human PCGs. Our work suggests that the interplay between Alu elements and piRNAs is an important factor that influences the evolutionary trajectory of human PCGs. Alu element transposable element piRNA interlocus gene conversion natural selection genome evolution protein-coding gene Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction PIWI-interacting RNAs (piRNAs) constitute the most diverse category of small RNAs in animals. These small RNAs guide Argonaute proteins of the PIWI clade (e.g., HIWI in humans, PIWI in fruit flies) to corresponding targets, triggering a series of reactions at the posttranscriptional and/or transcriptional levels, leading to the repression of target genes [ 1 – 5 ]. In general, the target recognition of piRNAs follows base-pairing rules, in which RNA segments that are substantially similar (generally allowing ≤ 4 mismatches) and reverse complementary to piRNA members will become target sites [ 2 , 6 , 7 ]. Therefore, genes that produce transcripts carrying these piRNA target sites could be repressed by the piRNA pathway. The genomes of humans and other eukaryotes have been widely shaped by various transposable elements (TEs) [ 8 – 10 ]. TEs are usually regarded as 'molecular parasites' because they are able to relocate or amplify themselves within the genome. It is well established that the piRNA pathway can serve as the 'immune system' of the genome to suppress the transposition of TEs [ 2 , 11 ]. Hence, the interplay between TEs and piRNAs plays an important role in shaping the genome [ 2 , 11 , 12 ]. Some TEs can integrate into protein-coding genes (PCGs). These TEs are retained in PCGs because either they do not strongly disrupt the normal functions of their host genes or they confer beneficial characteristics upon their host genes [ 10 , 13 , 14 ]. Although there seems to be no need to suppress these TEs, recent evidence suggests that TEs incorporated into PCGs are also targeted by piRNAs [ 3 , 4 ]. Hence, TEs incorporated into PCGs could influence their host genes via interaction with piRNAs. To assess the significance of such a piRNA-mediated influence, the first question that needs to be addressed is how likely TEs are to cause piRNA target sites to appear in PCGs. On the other hand, factors that can affect the evolutionary trajectory of these piRNA target sites also remain to be explored. In the present work, we carried out a series of analyses using human piRNA sequencing data and population genetic data [ 15 – 19 ]. We analyzed potential piRNA target sites in human PCGs and single nucleotide variants (SNVs) in these regions. We further recognized that the presence of a class of primate-specific TEs, Alu elements, is highly likely to cause piRNA target sites to appear in human PCGs. Additionally, we found that a mutagenic process induced by Alu elements, interlocus gene conversion (IGC), tends to increase sequence similarity between piRNAs and their potential target sites. Hence, the integration of Alu elements into human PCGs promotes the establishment of the piRNA-PCG targeting relationship, and IGC induced by Alu elements promotes the enhancement of the piRNA-PCG targeting relationship. Moreover, we recognized a footprint indicating that piRNA-PCG targeting relationships impose a piRNA-related selective constraint on the evolution of human PCGs. In brief, this work reveals how the interplay between Alu elements and piRNAs influences the evolution of human PCGs. Results Our analyses were begun by mapping unique human piRNA sequences to the longest mRNA sequences of human PCGs using BWA allowing ≤ 4 mismatches. Through this procedure, a fraction of sequences in PCGs were found to share similarity with piRNAs. Here, we refer to these sequences as 'piRNA-similar sequences' (Fig. 1 a). piRNA/piRNA-similar sequence pairs can be further divided into two classes: sense and antisense (Fig. 1 a). For a sense piRNA/piRNA-similar sequence pair, the piRNA strand is oriented in the same direction as the mRNA strand of the piRNA-similar sequence. Sense piRNA-similar sequences are merely sequences that are similar to piRNAs (Fig. 1 a). In contrast, for an antisense piRNA/piRNA-similar sequence pair, the piRNA strand is complementary to the mRNA strand of the piRNA-similar sequence. As piRNAs recognize targets according to base-pairing rules, antisense piRNA-similar sequences have the potential to form targeting relationships with their related piRNAs (Fig. 1 a). Additionally, we refer to SNVs that can reduce mismatches between piRNA-similar sequences and their related piRNAs as 'attracting SNVs' (Fig. 1 a). Furthermore, we refer to attracting SNVs in sense and antisense piRNA/piRNA-similar sequence pairs as 'sense attracting SNVs' and 'antisense attracting SNVs', respectively. Antisense attracting SNVs can increase piRNA-PCG targeting affinity, whereas sense attracting SNVs have no effect on the piRNA-PCG targeting relationship (Fig. 1 a). We mainly analyzed piRNA-similar sequences and attracting SNVs in this study. Two aspects of attracting SNVs were considered: their occurrence across piRNA/piRNA-similar sequence pairs and their allele frequencies in the human population (Fig. 1 b). The presence of Alu elements almost always results in potential piRNA target sites in human PCGs Antisense piRNA-similar sequences in PCGs have the potential to form targeting relationships with piRNAs. In this section, we show that Alu elements significantly contribute to the presence of antisense piRNA-similar sequences in human PCGs. First, we analyzed how many pairing relationships between antisense piRNA-similar sequences and their related piRNAs are contributed by TEs and what kinds of TEs are major contributors. We found that in 88.88% of antisense piRNA/piRNA-similar sequence pairs, the antisense piRNA-similar sequence was TE-derived. Notably, in 85.88% of antisense piRNA/piRNA-similar sequence pairs, the antisense piRNA-similar sequence was Alu derived (Fig. 2 a), accounting for 96.66% of pairing relationships related to TE-derived piRNA-similar sequences. We also investigated how many pairing relationships were related to TE-derived piRNAs. piRNAs were mapped to the reference genome using BWA allowing no mismatches. Thus, information about their possible genomic locations was obtained. We found that in 98.5% of antisense piRNA/piRNA-similar sequence pairs, the piRNA could be mapped to TE-derived regions. Additionally, in 96.2% of antisense piRNA/piRNA-similar sequence pairs, the piRNA could be mapped to an Alu -derived region (Fig. 2 a), accounting for 97.66% of pairing relationships related to TE-derived piRNAs. This result suggests that TEs, especially Alu elements, contribute significantly to the emergence of the targeting relationship between piRNAs and PCGs. In contrast, among all of the PCG sequences analyzed in this study, only 9.56% were TE-derived, and among all of the piRNAs analyzed in this study, only 37.26% were TE-derived. Moreover, we calculated the following conditional probabilities to further investigate the contribution of Alu elements to the targeting relationship between piRNAs and human PCGs. We define P (antisense piRNA-similar| Alu ) as the probability that a nucleotide is located in at least one antisense piRNA-similar sequence, given that it is known that the nucleotide is located in an Alu -derived region. P (antisense piRNA-similar| Alu ) can be calculated as the total length of the intersecting regions of antisense piRNA-similar sequences and Alu -derived regions divided by the total length of Alu -derived regions (Fig. 2 b). P (antisense piRNA-similar|non- Alu ) can also be calculated in a similar manner. Strikingly, we found that the value of P (antisense piRNA-similar| Alu ) was 0.9819 (Fig. 2 b), indicating that the presence of Alu elements showed a 0.9819 probability of resulting in antisense piRNA-similar sequences. In other words, the presence of Alu elements almost always results in antisense piRNA-similar sequences in human PCGs. In contrast, the value of P (antisense piRNA-similar|non- Alu ) only reached 0.0882 (Fig. 2 b). Additionally, in a PCG, a region can be present in multiple antisense piRNA/piRNA-similar sequence pairs, as it can share sequence similarity with multiple piRNAs. We calculated the average count of antisense piRNA/piRNA-similar sequence pairs across Alu- derived regions and across non- Alu- derived regions in human PCGs (Fig. 2 b). The average count of antisense piRNA/piRNA-similar sequence pairs across Alu- derived regions was 28.37, while the average count of antisense piRNA/piRNA-similar sequence pairs across non- Alu- derived regions was 0.0918. The former value was 309 times greater than the latter. This indicates that over the same sequence length, Alu -derived sequences can develop many more piRNA-PCG targeting relationships than non- Alu- derived sequences. For sense piRNA/piRNA-similar sequence pairs, a close association with Alu elements was also detected (Additional file 1: Figure. S1). Collectively, these results suggest that Alu elements are major contributors to the sequence similarity shared between piRNAs and PCGs. As a consequence of this effect, the presence of Alu elements has a high probability resulting in antisense piRNA-similar sequences in human PCGs. Hence, Alu elements significantly contribute to the presence of potential piRNA target sites in human PCGs. IGC induced by Alu elements tends to increase the sequence similarity between piRNAs and their potential target sites Above, we have shown that Alu elements are major contributors to the sequence similarity shared between piRNAs and PCGs. TEs can induce IGC, which is a mutagenic process in which a genomic region mistakenly utilizes its homologs as templates to 'repair' its DNA [ 20 – 23 ]. Hence, Alu elements might exert a secondary influence on piRNA-similar sequences and their related piRNAs via IGC. In this section, we provide evidence of the influence of IGC. In a piRNA/piRNA-similar sequence pair, whether an SNV is considered an attracting SNV depends on the nucleotide paired with the SNV. If the occurrence of SNVs is mutually independent in piRNA-similar sequences and their related piRNAs, an SNV occurring in a mismatch site should show a 1/3 probability of being an attracting SNV (Fig. 3 a). Under the influence of IGC, the occurrence of genetic variants will not be mutually independent in piRNA-similar sequences and their related piRNAs. Hence, we first evaluated the influence of IGC by examining whether the occurrence of attracting SNVs deviates from the expectation of 1/3. We analyzed SNVs occurring in mismatch sites. In each piRNA/piRNA-similar sequence pair, for SNVs occurring in mismatch sites, we counted the occurrences of attracting SNVs and the occurrences of nonattracting SNVs. Then, we summed the count results for each piRNA/piRNA-similar sequence pair and calculated the proportions of occurrences of attracting SNVs and nonattracting SNVs. Interestingly, the observed proportion of the occurrences of attracting SNVs was approximately 60% rather than 1/3. In both SNVs in piRNA-similar sequences and SNVs in their related piRNAs, this phenomenon was consistently observed (Fig. 3 b and Additional file 1: Figure. S2; p < 0.05, binomial test). Additionally, this phenomenon was consistently observed using both 1000 Genomes data and UK10K data (Fig. 3 b and Additional file 1: Figure. S2). Notably, for both sense and antisense piRNA/piRNA-similar sequence pairs, there was an excess of attracting SNVs (Fig. 3 b and Additional file 1: Figure. S2). Sense piRNA/piRNA-similar sequence pairs do not reflect any piRNA-PCG targeting relationship. As the excess of attracting SNVs was observed in sense piRNA/piRNA-similar sequence pairs, this phenomenon is unlikely to reflect the impact of natural selection. In other words, the observed phenomenon is unlikely to be caused by an effect on the persistence of genetic variants. Alternatively, it is likely to be caused by an effect on the occurrence of genetic variants—i.e., mutation. Thus, the observed phenomenon suggests that a mutagenic process homogenizes piRNA-similar sequences and their related piRNAs. To date, IGC is the only recognized mutagenic process that has such an effect [ 20 – 23 ]. To provide further evidence of IGC, we conducted the analysis with another indicator. An excess of shared nucleotide changes (divergences or polymorphisms) between different loci has been proposed as an indicator of IGC [ 24 – 26 ]. We analyzed SNVs occurring in identical sites (above, we analyzed SNVs occurring in mismatch sites). If we observed that a pair of SNVs were located in a pair of identical sites and that this pair of SNVs were also identical, we referred to this SNV pair as a 'shared SNV' (Fig. 4 a). In each piRNA/piRNA-similar sequence pair, we counted the occurrences of shared SNVs. Then, we summed the count results for each piRNA/piRNA-similar sequence pair. For comparison, we built a control dataset by shuffling the one-to-one relationships of the identical sites in each piRNA/piRNA-similar sequence pair (see Materials and Methods). We found that the number of occurrences of shared SNVs in the real data was significantly higher than that in the control dataset, indicating an excess of shared SNVs (Fig. 4 b; see the logic of significance testing in Materials and Methods). Similar to the above analysis, in both sense and antisense piRNA/piRNA-similar sequence pairs, an excess of shared SNVs was detected (Fig. 4 b). Thus, the excess of shared SNVs was also caused by a mutagenic process that homogenizes piRNA-similar sequences and their related piRNAs. Collectively, the results provide two parallel lines of evidence, the excess of attracting SNVs and the excess of shared SNVs, suggesting that piRNA-similar sequences and their related piRNAs are under the influence of IGC. The above results have shown that piRNA-similar sequences and their related piRNAs are influenced by IGC. Nevertheless, there is still a question to be addressed. piRNA-similar sequences and their related piRNAs are 25–32 bp, and such short homologous sequences are unlikely to induce IGC [ 20 , 23 , 27 , 28 ]. Hence, if piRNA-similar sequences and their related piRNAs are under the influence of IGC, regions that share sequence similarity should not be restricted to the regions of piRNA-similar sequences and their related piRNAs. We expanded the genomic sequences of piRNA-similar sequences and their related piRNAs by 150 bp in both the upstream and downstream directions (total length of the expanded region = piRNA length + 150 bp × 2). In each expanded region, we analyzed the sequence similarity of the most similar block (maximum similarity) (Fig. 5 ). Our analysis included 30 repeats, and each repeat included 100 expanded regions of randomly sampled piRNA/piRNA-similar sequence pairs. For each repeat, we calculated the median value of maximum similarity. The most similar block was set to different lengths in the analysis. When the length of the most similar block was set to 200 bp, in all 30 repeats, the median value of maximum similarity was ≥ 80% (Fig. 5 and Additional file 1: Figure. S3). When the length of the most similar block was set to a shorter length (100 or 150 bp), the median value of maximum similarity was higher (Fig. 5 and Additional file 1: Figure. S3). This result indicates that if a piRNA shares sequence similarity with a segment of a PCG, their flanking regions will also share sequence similarity. IGC at a high rate usually requires completely identical regions that are longer than 200 bp, referred to as the 'minimal efficient processing segment' (MEPS) [ 20 , 29 ]. Although the sequence similarity between flanking regions was reasonably high in the above analysis, it did not satisfy the MEPS criterion. Nevertheless, there is evidence suggesting that IGC can still occur between sequences that are not completely identical and shorter than the MEPS criterion, although the rate is relatively low [ 20 , 26 , 30 , 31 ]. In addition, if sequence similarity satisfies the MEPS criterion, the rate of IGC will be higher than the spontaneous mutation rate by orders of magnitude [ 27 , 28 , 32 ]. With such a high rate of IGC, we should observe that the allele frequencies of attracting SNVs are higher than those of nonattracting SNVs by orders of magnitude. In our analyses, we did observe that the allele frequencies of attracting SNVs were higher than those of nonattracting SNVs (compare Fig. 6 and Additional file 1: Figure. S4), which was consistent with a scenario in which a mutagenic process was additionally introduced. However, the difference did not reach orders of magnitude (compare Fig. 6 and Additional file 1: Figure. S4). Thus, our results were consistent with a scenario in which although an influence of IGC existed, the rate of IGC was relatively low. Moreover, as genomic analyses utilize genetic information that accumulates in DNA over time, even though the rate of IGC is relatively low, IGC can still leave a footprint for the analysis [ 20 , 26 ]. Additionally, increasing evidence suggests that IGC can occur among Alu elements [ 33 – 35 ]. Overall, as the sequence similarity shared between piRNAs and PCGs is mainly contributed by Alu elements, sequences surrounding piRNA-similar sequences and their related piRNAs also share extensive similarity. These sequences sharing extensive similarity allow the influence of IGC to be exerted upon piRNA-similar sequences and their related piRNAs. Consequently, an excess of attracting SNVs occurs in piRNA-similar sequences and their related piRNAs. Antisense attracting SNVs are a subset of attracting SNVs that can increase piRNA-PCG targeting affinity. IGC promotes the occurrence of antisense attracting SNVs, thus constituting an evolutionary force that reinforces the piRNA-PCG targeting relationship. Natural selection opposes the promotion effect of human Alu elements on the formation of piRNA-PGC targeting relationships. In the above sections, we show that human Alu elements promote both the establishment and enhancement of piRNA-PCG targeting relationships. The presence of piRNA target sites in PCGs should be relevant to the fitness of organisms. piRNA-related natural selection thereby can influence the evolution of human PCGs. If so, piRNA-related natural selection should leave a signature in piRNA target sites. As antisense attracting SNVs can increase piRNA-PCG targeting affinity (Fig. 1 ), analyzing whether they tend to show high or low allele frequencies can provide information about the action of piRNA-related natural selection. If antisense attracting SNVs tend to show high allele frequencies, it indicates that natural selection favors the influence of antisense attracting SNVs. Hence, with regard to the targeting relationship between piRNAs and PCGs, natural selection acts as an 'attractive force'. Otherwise, the opposite scenario is indicated, in which natural selection acts as a 'repulsive force', against the targeting relationship between piRNAs and PCGs. As mentioned in the last section, IGC could cause a systematic difference in allele frequencies between attracting SNVs and nonattracting SNVs (compare Fig. 6 and Additional file 1: Figure. S4). Hence, IGC could be a confounding factor for inferring the footprint of natural selection. IGC influences both sense attracting SNVs and antisense attracting SNVs, but sense attracting SNVs have no effect on the piRNA-PCG targeting relationship. Therefore, we compared antisense attracting SNVs with sense attracting SNVs, rather than nonattracting SNVs, to control the confounding effect of IGC. We observed that antisense attracting SNVs tended to show lower allele frequencies than sense attracting SNVs, which supports the 'repulsive force' hypothesis rather than the 'attractive force' hypothesis (Fig. 6 and Additional file 1: Figure. S5; p < 0.05, one-tailed Mann‒Whitney U test). In particular, when we restricted the allele frequencies of SNVs to ≥ 0.001, the tendency toward low allele frequencies became more prominent (Fig. 6 ). Moreover, we examined different categories of SNVs in PCGs (Fig. 6 and Additional file 1: Figure. S5). Without restricting allele frequencies, antisense attracting SNVs showed a tendency toward low allele frequencies in 5'UTR, 3'UTR and synonymous SNVs ( p < 0.05, Mann‒Whitney U test; Additional file 1: Figure. S5), but we did not detect a significant tendency in nonsynonymous SNVs (Additional file 1: Figure. S5). One possible explanation is that nonsynonymous SNVs can alter the amino acid sequence, and this effect could obscure the influence of nonsynonymous SNVs on the piRNA targeting relationship. Nevertheless, when allele frequencies were restricted to ≥ 0.001, for all categories of SNVs, including nonsynonymous SNVs, we detected a pattern in which antisense attracting SNVs tended to show low frequencies in comparison with sense attracting SNVs (Fig. 6 ; p < 0.05, one-tailed Mann‒Whitney U test). Additionally, we repeated the above analyses using another piRNA targeting rule to identify piRNA-similar sequences [ 7 ]. This rule considers the 'seed region' of the piRNA target sites, rather than simply requiring ≤ 4 mismatches (see Materials and Methods) [ 7 ]. Using this rule for identifying potential piRNA target sites, we observed a pattern very similar to that described above (Additional file 1: Figure. S6). Furthermore, in both the 1000 Genomes data and the UK10K data, the tendency toward low allele frequencies for antisense attracting SNVs was consistently observed (Fig. 6 ). The observation of a consistent pattern under different conditions suggests that this pattern is not caused by chance. Finally, we did not detect a tendency toward low allele frequencies for antisense nonattracting SNVs in comparison with sense nonattracting SNVs (Additional file 1: Figure. S4). Collectively, our results suggest that antisense piRNA-similar sequences are not merely sequences that share similarity with piRNAs. Instead, the presence of these regions in human PCGs imposes a piRNA-related selective constraint on the evolution of human PCGs. Furthermore, the phenomenon that antisense attracting SNVs tend to show low allele frequencies suggests that the action of natural selection is likely to function as a 'repulsive force' rather than an 'attractive force'. Hence, the promotion effect of human Alu elements on the formation of piRNA-PGC targeting relationships is opposed by natural selection. Discussion It is well established that the transposition of TEs can disrupt genome stability, leading to deleterious effects. In addition to the direct influence of transposition, the homologous sequences produced by TEs can further induce posttransposition influence on the genome [ 23 , 29 , 36 ]. The target recognition of piRNAs depends on the sequence similarity between piRNAs and their target sites. Hence, homologous sequences produced by TEs could have an influence on the target relationship between piRNAs and their target sites. Previously, researchers recognized that mRNAs harboring TEs tend to be upregulated in Piwil1 mutant mice [ 4 ]. This evidence suggests that TEs could contribute to the establishment of piRNA-PCG targeting relationships [ 3 , 4 ]. Although previous studies have provided insight into the roles of TEs in the formation of piRNA targeting relationships, evolutionary factors that influence the evolutionary trajectories of piRNA targeting relationships remain to be explored. To elucidate the impact of TEs on the establishment of piRNA-PCG targeting relationships, it is better to quantify such an impact. In this work, we calculated the conditional probability, P (antisense piRNA-similar| Alu ). Through this calculation, the promotion effect of Alu elements on the establishment of piRNA target sites can be quantified as a probability. Currently, it remains elusive whether there exists a class of TEs whose presence is highly likely to result in piRNAs target sites. Although calculating P (antisense piRNA-similar| Alu ) is not complicated, our result explicitly suggests that the presence of Alu elements is highly likely to cause piRNA target sites to appear in human PCGs. This approach allowed us to reveal the significant role of Alu elements in the establishment of piRNA-PCG targeting relationships. IGC can affect the fitness of organisms via various mechanisms, such as altering amino acid sequences and rewiring regulatory networks [ 23 , 29 ]. However, its influence on the piRNA-PCG targeting relationship has not been explored. In this work, we found that human Alu elements allow the influence of IGC to be exerted upon piRNA-similar sequences and their related piRNAs. In particular, IGC induced by Alu elements facilitates the occurrence of attracting SNVs. Thus, IGC induced by Alu elements makes the sequence similarity between piRNAs and PCGs prone to increase, which indicates that in addition to promoting the establishment of piRNA-PCG targeting relationships, human Alu elements also promote the enhancement of piRNA-PCG targeting relationships. Furthermore, we observed that antisense attracting SNVs tend to show lower allele frequencies than sense attracting SNVs. This footprint of piRNA-related natural selection suggests that the presence of antisense piRNA-similar sequences in human PCGs imposes a piRNA-related selective constraint on human PCGs. Meanwhile, this footprint of natural selection suggests that natural selection opposes the promotion effect of Alu elements on the formation of piRNA-PCG targeting relationships. Hence, the uncovered promotion effect of Alu elements on the formation of piRNA-PCG targeting relationships is unlikely to be a beneficial feature in the evolution of the human genome. It may be a little counterintuitive that evolution allows an adverse feature to appear in our genomes. Nevertheless, our observation is in accordance with current knowledge about the evolution of TEs and piRNAs. TEs undergo self-amplification or relocation in the genome. As this process continues, some TEs eventually integrate into piRNA clusters (regions that produce piRNAs). Once TEs integrate into piRNA clusters, they can become templates of piRNAs. Thus, the genome evolves the ability to generate piRNAs suppressing these TEs [ 1 , 11 , 37 ]. This evolutionary model is widely accepted to explain how the genome continuously adjusts the piRNA repertoire in response to the activity of TEs [ 1 , 11 , 37 ]. However, before TEs integrate into piRNA clusters, they might have integrated into other genomic regions, including PCGs. Thus, as the genome evolves the ability to suppress TEs, nonfunctional or deleterious piRNA-PCG targeting relationships can also appear. As the advantage of suppressing TEs can outweigh the disadvantage of forming nonfunctional or deleterious piRNA-PCG targeting relationships, these piRNA-PCG targeting relationships are retained in the genome. This might explain why the presence of Alu elements in PCGs promotes the establishment of piRNA-PCG targeting relationships. In addition, Alu elements can exert a secondary influence on piRNA-PCG targeting relationships via IGC. IGC is a mutagenic process that allows the sequence similarity shared between piRNAs and PCGs to increase. Thus, even though a large number of piRNA-PCG targeting relationships are nonfunctional or deleterious, there is a tendency to enhance them in the genome. Together, the promotion effect of Alu elements on the formation of piRNA-PCG targeting relationships is likely to be a bitter pill to swallow in the evolution of the human genome. Therefore, the response of natural selection is more likely to oppose, rather than favor, such an effect. Overall, the presence of Alu elements in human PCGs allows multiple piRNA-related factors to exert influence on the evolution of human PCGs. The interplay between Alu elements and piRNAs adds an additional complexity to the evolution of human PCGs. Recognizing the piRNA-mediated influence of Alu elements reported here extends our understanding of how TEs shape our genomes via various evolutionary forces. At the end, we would like to discuss some treatments adopted in our analyses. In the present work, we mainly analyzed piRNA-similar sequences rather than piRNA target sites predicted by more complicated rules. A potential concern is that this approach could result in the misidentification of piRNA target sites. However, we note that the misidentification of piRNA target sites would not seriously undermine the validity of the results. First, piRNA-similar sequences include bona fide piRNA target sites. If piRNA-similar sequences mainly originate from Alu elements, bona fide piRNA target sites will also mainly originate from Alu elements. Similarly, if piRNA-similar sequences are under the influence of IGC, bona fide piRNA target sites will also be under the influence of IGC. Second, even if some piRNA target sites were misidentified, they would not exert a piRNA-mediated fitness effect on PCGs. Thus, they would not cause a shift in the distribution of allele frequencies (as shown in Fig. 6 ). In other words, misidentified piRNA target sites may result in some random errors but cannot cause systematic errors. Moreover, we obtained similar results using a different rule for identifying potential piRNA target sites (Additional file 1: Figure. S6). In addition, the use of a relaxed rule for identifying piRNA target sites can reduce the probability of failing to identify true piRNA target sites. Together, analyzing piRNA-similar sequences is a simple yet prudential strategy for elucidating the evolutionary interplay between Alu elements and piRNAs. Conclusions In the human genome, Alu elements significantly contribute to the sequence similarity shared between piRNAs and PCGs. Consequently, Alu elements promote the establishment of the targeting relationship between piRNAs and PCGs. Additionally, Alu elements can induce IGC, which constitutes an evolutionary force that enhances piRNA-PCG targeting relationships. Furthermore, piRNA-PCG targeting relationships impose a piRNA-related selective constraint on human PCGs. Overall, the interplay between Alu elements and piRNAs is an important factor that participates in the shaping of human PCGs. Materials And Methods Data resources The human piRNA data used in this study were published in three previous reports [ 17 – 19 ]. These data were generated via periodate oxidation treatment followed by small RNA sequencing. The SRA accession numbers of these data were SRR2156539, SRR2156540 [ 17 ], SRR835324, SRR950451 [ 18 ], SRR8575349, SRR8575350, SRR8575385, SRR8575386, SRR8575387, SRR8575388, SRR8575409, SRR8575410, SRR8575352, and SRR8575351 [ 19 ]. SRA files in format were converted to files in FASTQ format by fastq-dump in the SRA Toolkit. Reference human genome (hg38) sequences were obtained from the USCS Genome Browser website. The population genetics data of the high-coverage sequencing version of the 1000 Genomes project were obtained from http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/working/20201028_3202_phased/ . The population genetics data of the UK10K project were obtained from https://www.uk10k.org/data.html . The coordinates of the UK10K data were converted to correspond to the hg38 assembly by using Picard LiftoverVcf. The mRNA sequences and the genomic positions of the 5'UTRs, coding sequences and 3'UTRs of PCGs were obtained from Ensembl BioMart. For each gene, we retained the longest mRNA sequence as the representative transcript for subsequent analyses. The genomic positions of TE-derived sequences were obtained from the RepeatMasker website ( http://repeatmasker.org/species/hg.html ). Preprocessing Of Pirna Sequencing Data The downloaded human piRNA sequencing data were processed with Minion [ 38 ] and Cutadapt [ 39 ] to remove adaptors. Sequences lacking adaptors that had lengths of 25–32 bp and appeared more than once in the sequencing data were retained for subsequent analyses. We removed redundant piRNA sequences and retained only one sequence for each unique piRNA in subsequent analyses. Sequences that could be mapped to mature microRNAs, microRNA hairpins (downloaded from miRbase [ 40 ]), snoRNAs, rRNAs or vtRNAs (downloaded from Ensembl BioMart) were excluded from the analyses. Identification Of Pirna-similar Sequences In Protein-coding Genes And Possible Genomic Locations Of Pirnas To identify piRNA-similar sequences, we mapped piRNAs to cDNAs using BWA [ 41 ], allowing up to 4 mismatches (set -n 4, in bwa aln). To identify possible genomic locations of piRNAs, piRNA sequences were mapped to the reference genome using BWA, requiring perfect matching (set -n 0, in bwa aln). Because of the short length of piRNAs, the BWA-backtrack algorithm was used when performing BWA mapping (including three steps: bwa index, bwa aln and bwa samse). The mapping results (BAM files) were parsed by Pysam to obtain location information of piRNA-similar sequences in mRNAs and location information of piRNA sequences in the reference genome. The location information of piRNA-similar sequences in mRNAs was transformed into location information in the reference genome according to the corresponding relationship between these two kinds of location information. According to the strand relationship between mRNAs and related piRNAs, piRNA-similar sequences were divided into two categories: sense and antisense. When identifying the genomic locations of piRNAs, there were some piRNAs that could not be mapped to the reference genome. In the following analyses, if information about piRNA genomic locations was needed, piRNAs that could not be mapped to the reference genome were excluded. Identification Of Attracting Snvs Information on SNVs was parsed from VCF files using cyvcf2 [ 42 ]. piRNA-similar sequences were identified from reference genomic data. Hence, in the VCF file, the 'reference allele' (the REF item) presented an allelic state identical to that shown in piRNA/piRNA-similar sequence pairs. In contrast, the 'alternate allele' (the ALT item) could alter the sequence similarity between piRNAs and piRNA-similar sequences. For each piRNA/piRNA-similar sequence pair, if an SNV was located in a mismatch site and the ALT item eliminated that mismatch site, we referred to the allele of the ALT item as an attracting SNV. Otherwise, the allele of the ALT item was referred to as a nonattracting SNV (Fig. 1 a). Analysis of the Alu associations of piRNA-similar sequences and their related piRNAs piRNA-similar sequences that overlapped with Alu elements by more than 70% were defined as Alu -derived piRNA-similar sequences. As some piRNAs could be mapped to multiple genomic locations, as long as a piRNA could be mapped to a genomic location that overlapped with an Alu element by more than 70%, that piRNA was defined as an Alu -derived piRNA. We analyzed the proportion of piRNA/piRNA-similar sequence pairs in which the piRNA was Alu -derived and the proportion of piRNA/piRNA-similar sequence pairs in which the piRNA-similar sequence was Alu -derived. For comparison, the Alu -derived fractions of PCGs and piRNAs were also calculated. To further elucidate the contribution of Alu elements to the sequence similarity shared between piRNAs and PCGs, the following indices were calculated. These indices were calculated for sense and antisense piRNA/piRNA-similar sequence pairs separately (the results for antisense pairs are shown in Fig. 2 , and the results for sense pairs are shown in Additional file 1: Figure. S1). P (piRNA-similar| Alu ) , the total length of intersecting regions of piRNA-similar sequences and Alu -derived regions (light blue regions in Fig. 2 ), divided by the total length of Alu -derived regions in PCGs (gray box in Fig. 2 ). This index indicates the probability that a nucleotide is located in at least one piRNA-similar sequence, given that it is known that the nucleotide is located in an Alu -derived region in PCGs. P (piRNA-similar|non- Alu ) , the total length of intersecting regions of piRNA-similar sequences and non- Alu regions (light gray regions in Fig. 2 ) divided by the total length of non- Alu regions in PCGs (gray line in Fig. 2 ). This index indicates the probability that a nucleotide is located in at least one piRNA-similar sequence, given that it is known that the nucleotide is located in a non- Alu region in PCGs. The values of P (piRNA-similar| Alu ) and P (piRNA-similar|non- Alu ) were compared. Average count of piRNA/piRNA-similar sequence pairs across Alu regions , the total number of piRNA/piRNA-similar sequence pairs whose piRNA-similar sequences reside in Alu regions divided by the total length of Alu regions in PCGs. Average count of piRNA/piRNA-similar sequence pairs across non- Alu regions , the total number of piRNA/piRNA-similar sequence pairs whose piRNA-similar sequences reside in non- Alu regions divided by the total length of non- Alu regions in PCGs. The average counts of piRNA/piRNA-similar sequence pairs across Alu and non- Alu regions were compared. Analysis Of The Igc Footprint For each piRNA/piRNA-similar sequence pair, we analyzed SNVs occurring in mismatch sites and the nucleotides paired with them, examining whether they are attracting SNVs. SNVs in four types of data were analyzed separately: 1) sense piRNA-similar sequences, 2) antisense piRNA-similar sequences (shown in Fig. 3 ), 3) piRNAs related to sense piRNA-similar sequences, and 4) piRNAs related to antisense piRNA-similar sequences (shown in Additional file 1: Figure. S2). Some piRNAs can be mapped to multiple genomic locations. In the analysis of piRNAs related to piRNA-similar sequences, all possible piRNA genomic locations were examined. We counted the occurrences of attracting SNVs and the occurrences of nonattracting SNVs in each piRNA/piRNA-similar sequence pair. Then, we summed the count results for each piRNA/piRNA-similar sequence pair and calculated the proportion of occurrences of attracting SNVs and the proportion of occurrences of nonattracting SNVs. Theoretically, if the occurrence of SNVs is mutually independent in piRNA-similar sequences and their related piRNAs, the proportion of occurrences of attracting SNVs should be approximately 1/3 (Fig. 3 ). We examined whether the observation significantly deviated from 1/3 through the binomial test (implemented by Scipy.stats.binom_test). On the other hand, we explored further evidence from identical sites in piRNA/piRNA-similar sequence pairs. In each piRNA/piRNA-similar sequence pair, if a pair of identical sites both contained SNVs and the ALT items of the two SNV sites were identical, we referred to this SNV pair as a 'shared SNV' (Fig. 4 ). For comparison, we built a control dataset by shuffling one-to-one site pairing relationships in piRNA/piRNA-similar sequence pairs. For example, suppose that in a piRNA/piRNA-similar sequence pair, site pairs 1–2, 4–15, 17–23, and 25–28 are identical. Among these 25 site pairs, there are a SNV sites in the piRNA sequence and b SNV sites in the piRNA-similar sequence. If X of the a SNV sites in the piRNA sequence and X of the b SNV sites in the piRNA-similar sequence are paired and show identical allelic states in their ALT items, the number of shared SNVs will be X . In the shuffling procedure, the 25 identical sites are randomly paired. Thus, the a SNV sites in the piRNA sequence and the b SNV sites in the piRNA-similar sequence will be paired anew. If after shuffling, X' of the a SNV sites in the piRNA sequence and X' of the b SNV sites in the piRNA-similar sequence are paired and show identical allelic states in their ALT items, then X' shared SNVs are randomly generated without accounting for any site homology. For each piRNA/piRNA-similar sequence pair, we performed the shuffling procedure described above 30 times. Hence, 30 repeats of X' were generated, which constituted the control dataset. If the occurrence of shared SNVs is determined by chance, X should not significantly deviate from X' . To perform significance testing, the sample mean and the sample standard deviation of the 30 generated X' values (denoted by \(\stackrel{-}{X{\prime }}\) and s , respectively) were calculated to estimate the mean and the standard deviation of the random distribution of X' (denoted by µ and σ , respectively). According to Chebyshev's inequality, P {| X – µ | ≥ ε } ≤ σ 2 / ε 2 ; when ε = 5 σ , P { X – µ ≥ 5 σ } ≤ 1/25 = 0.04. Therefore, if the observed number of shared SNVs in the real data ( X ) is larger than the mean of the control dataset ( \(\stackrel{-}{X{\prime }}\) ) by 5 × the standard deviation, the observation is considered significantly different from the expectation of random effects, suggesting the influence of IGC exists. Analysis Of The Sequence Similarity Between Regions Surrounding Pirnas And Their Pirna-similar Sequences The analysis of flanking regions included 30 repeats. Each repeat included 100 piRNA/piRNA-similar sequence pairs that were randomly sampled from the complete set. In the analysis, we randomly assigned one of the possible piRNA locations for each piRNA/piRNA-similar sequence pair. Hence, for each piRNA/piRNA-similar sequence pair, a pair of genomic locations corresponding to the piRNA and the piRNA-similar sequence were obtained. Then, both of these genomic locations were expanded by 150 bp in the upstream and downstream directions. The expanded regions were aligned using MAFFT [ 43 ]. For each pair of expanded regions, the sequence similarity of the most similar block (maximum similarity) was analyzed. Three different block lengths were used in the analysis: 100, 150 and 200 bp. In each repeat, we calculated the median value of maximum similarity across the 100 expanded regions. Corresponding boxplots were also generated. The results are shown in Fig. 5 and Additional file 1: Figure S3. Analysis Of The Footprint Of Natural Selection To plot the distribution of allele frequencies, if an SNV could eliminate a mismatch in any piRNA/piRNA-similar sequence pair, it was used to plot the distribution of allele frequencies for attracting SNVs; otherwise, it was used to plot the distribution of allele frequencies for nonattracting SNVs. Additionally, if an SNV was counted as an attracting SNV in any antisense piRNA/piRNA-similar sequence pair, it was used to plot the distribution of allele frequencies for antisense attracting SNVs. Otherwise, it was used to plot the distribution of allele frequencies for sense attracting SNVs. Thus, those SNVs that allowed the similarity of any piRNA/piRNA-similar sequence pair to increase were used to plot the distribution of allele frequencies for attracting SNVs, and those SNVs that had the potential to reinforce any piRNA-PCG targeting relationship were used to plot the distribution of allele frequencies for antisense attracting SNVs. The distribution of allele frequencies for antisense attracting SNVs was compared to the distribution of allele frequencies for sense attracting SNVs. A one-tailed Mann‒Whitney U test was used to examine whether the distribution of allele frequencies for antisense attracting SNVs was significantly skewed toward low frequencies in comparison with that of sense attracting SNVs (implemented by using Scipy.stats.mannwhitneyu, with the setting alternative='greater'). Information about allele frequencies was obtained directly from the 'AF' items in VCF files. For the 1000 Genomes data, variants with unconfident allele frequency estimates were indicated by 'AF = 0.0'. Such variants were excluded from our analysis. Furthermore, SNVs were categorized as those located in 5'UTRs and 3'UTR and as synonymous and nonsynonymous variants. The allele frequencies of these categories were further analyzed separately (Fig. 6 , bottom panels). In addition, the allele frequencies of nonattracting SNVs in sense and antisense pairing relationships were also compared (Additional file 1: Figure. S4). On the other hand, we changed the rule for identifying piRNA-similar sequences. A recent study proposed the following piRNA targeting rules (the 'stringent' rules of [ 7 ]): a mismatch at the first nucleotide of the piRNA is not counted, up to one GU wobble pair is allowed in the first 2–7 bp (seed region), and only two mismatches plus an additional GU mismatch are allowed overall. In our analysis, piRNAs were first mapped to mRNAs using BWA, allowing up to 5 mismatches. Segments that satisfied the above-described rule and segments that contained one more mismatch than the above-described rule were identified as piRNA-similar sequences. Allele frequencies in these segments were analyzed as described above to examine whether the footprint of natural selection was robust to the rule for identifying piRNA-similar sequences. The results are shown in Additional file 1: Figure. S6. All of the above analyses were implemented by using in-house Python scripts. The Matplotlib [ 44 ], SciPy [ 45 ], NumPy [ 46 ], Pandas [ 47 ] and seaborn [ 48 ] packages were used for statistical analyses. The scripts have been deposited at https://github.com/he-chong/pirna_pcg_evol.git . Abbreviations piRNA: PIWI-interacting RNA; TE: Transposable element; PCG: Protein-coding gene; IGC: Interlocus gene conversion; SNV: Single nucleotide variant; MEPS: Minimal efficient processing segment; UTR: Untranslated region. Declarations Acknowledgments We thank Meng-Yun Chen for the repeated and in-depth discussion. Author's contribution CH and HZ conceived the study and wrote the manuscript. CH carried out the analyses. All authors have read and approved the manuscript for publication. Funding HZ was supported by National Natural Science Foundation of China (31771456). No funding body played a role in the study design, analysis or interpretation. Availability of data and materials In-house Python scripts are available in GitHub at https://github.com/he-chong/pirna_pcg_evol.git. Ethics approval and consent to participate Not applicable Consent for publication Not applicable Competing interests The authors declare that they have no competing interests. Author details 1 Bioinformatics Section, School of Basic Medical Sciences, Southern Medical University, Guangzhou, 510515, China 2 Guangdong-Hong Kong-Macao Greater Bay Area Center for Brain Science and Brain-Inspired Intelligence, Southern Medical University, Guangzhou, 510515, China 3 Guangdong Provincial Key Lab of Single Cell Technology and Application, Southern Medical University, Guangzhou, 510515, China References Aravin AA, Hannon GJ, Brennecke J. The Piwi-piRNA pathway provides an adaptive defense in the transposon arms race. Science. 2007;318:761–4. Czech B, Munafò M, Ciabrelli F, Eastwood EL, Fabry MH, Kneuss E, et al. piRNA-guided genome defense: From biogenesis to silencing. Annu Rev Genet. 2018;52:131–57. Wang C, Lin H. Roles of piRNAs in transposon and pseudogene regulation of germline mRNAs and lncRNAs. Genome Biol. 2021;22:1–21. Watanabe T, Cheng EC, Zhong M, Lin H. Retrotransposons and pseudogenes regulate mRNAs and lncRNAs via the piRNA pathway in the germline. Genome Res. 2015;25:368–80. Grimson A, Srivastava M, Fahey B, Woodcroft BJ, Chiang HR, King N, et al. Early origins and evolution of microRNAs and Piwi-interacting RNAs in animals. Nature. 2008;455:1193–7. Zhang P, Kang JY, Gou LT, Wang J, Xue Y, Skogerboe G, et al. MIWI and piRNA-mediated cleavage of messenger RNAs in mouse testes. Cell Res. 2015;25:193–207. Zhang D, Tu S, Stubna M, Wu WS, Huang WC, Weng Z, et al. The piRNA targeting rules and the resistance to piRNA silencing in endogenous genes. Science. 2018;359:587–92. Bourque G, Burns KH, Gehring M, Gorbunova V, Seluanov A, Hammell M, et al. Ten things you should know about transposable elements. Genome Biol. 2018;19:199. Li WH, Gu Z, Wang H, Nekrutenko A. Evolutionary analyses of the human genome. Nature. 2001;409:847–9. Jangam D, Feschotte C, Betrán E. Transposable element domestication as an adaptation to evolutionary conflicts. Trends Genet. 2017;33:817–31. Ophinni Y, Palatini U, Hayashi Y, Parrish NF. piRNA-guided CRISPR-like immunity in eukaryotes. Trends Immunol. 2019;40:998–1010. Brookfield JFY. The ecology of the genome - Mobile DNA elements and their hosts. Nat Rev Genet. 2005;6:128–36. Nekrutenko A, Li WH. Transposable elements are found in a large number of human protein-coding genes. Trends Genet. 2001;17:619–21. Chuong EB, Elde NC, Feschotte C. Regulatory activities of transposable elements: from conflicts to benefits. Nat Rev Genet. 2017;18:71–86. The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature. 2015;526:68–74. Walter K, Min JL, Huang J, Crooks L, Memari Y, McCarthy S, et al. The UK10K project identifies rare variants in health and disease. Nature. 2015;526:82–9. Williams Z, Morozov P, Mihailovic A, Lin C, Puvvula PK, Juranek S, et al. Discovery and characterization of piRNAs in the human fetal ovary. Cell Rep. 2015;13:854–63. Ha H, Song J, Wang S, Kapusta A, Feschotte C, Chen KC, et al. A comprehensive analysis of piRNAs from adult human testis and their relationship with genes and mobile elements. BMC Genomics. 2014;15:1–16. Özata DM, Yu T, Mou H, Gainetdinov I, Colpan C, Cecchini K, et al. Evolutionarily conserved pachytene piRNA loci are highly divergent among modern humans. Nat Ecol Evol. 2020;4:156–68. Mansai SP, Kado T, Innan H. The rate and tract length of gene conversion between duplicated genes. Genes. 2011;2:313–31. Roman Arguello J, Connallon T. Gene duplication and ectopic gene conversion in Drosophila. Genes. 2011;2:131–51. Hastings PJ. Mechanisms of ectopic gene conversion. Genes. 2010;1:427–39. Fawcett JA, Innan H. The role of gene conversion between transposable elements in rewiring regulatory networks. Genome Biol Evol. 2019;11:1723–9. Dumont BL. Interlocus gene conversion explains at least 2.7 % of single nucleotide variants in human segmental duplications. BMC Genomics. 2015;16:1–11. Mansai SP, Innan H. The power of the methods for detecting interlocus gene conversion. Genetics. 2010;184:517–27. Kijima TE, Innan H. On the estimation of the insertion time of LTR retrotransposable elements. Mol Biol Evol. 2010;27:896–904. Liskay RM, Letsou A, Stachelek JL. Homology requirement for efficient gene conversion between duplicated chromosomal sequences in mammalian cells. Genetics. 1987;115:161–7. Mehta A, Beach A, Haber JE. Homology requirements and competition between gene conversion and break-induced replication during double-strand break repair. Mol Cell. 2017;65:515-526.e3. Chen JM, Cooper DN, Chuzhanova N, Férec C, Patrinos GP. Gene conversion: Mechanisms, evolution and human disease. Nat Rev Genet. 2007;8:762–75. Mézard C, Pompon D, Nicolas A. Recombination between similar but not identical DNA sequences during yeast transformation occurs within short stretches of identity. Cell. 1992;70:659–70. Ahn BY, Dornfeld KJ, Fagrelius TJ, Livingston DM. Effect of limited homology on gene conversion in a Saccharomyces cerevisiae plasmid recombination system. Mol Cell Biol. 1988;8:2442–8. Reiter LT, Hastings PJ, Nelis E, De Jonghe P, Van Broeckhoven C, Lupski JR. Human meiotic recombination products revealed by sequencing a hotspot for homologous strand exchange in multiple HNPP deletion patients. Am J Hum Genet. 1998;62:1023–33. Zhi D. Sequence correlation between neighboring Alu instances suggests post-retrotransposition sequence exchange due to Alu gene conversion. Gene. 2007;390:117–21. Doronina L, Reising O, Schmitz J. Gene conversion amongst Alu SINE elements. Genes. 2021;12:905. Roy AM, Carroll ML, Nguyen S V., Salem AH, Oldridge M, Wilkie AOM, et al. Potential gene conversion and source genes for recently integrated Alu elements. Genome Res. 2000;10:1485–95. Deininger P. Alu elements: know the SINEs. Genome Biol. 2011;12:236. Czech B, Hannon GJ. One loop to rule them all: the ping-pong cycle and piRNA-guided silencing. Trends Biochem Sci. 2016;41:324–37. Davis MPA, van Dongen S, Abreu-Goodger C, Bartonicek N, Enright AJ. Kraken: A set of tools for quality control and analysis of high-throughput sequence data. Methods. 2013;63:41–9. Martin M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet.journal. 2011;17:10. Kozomara A, Birgaoanu M, Griffiths-Jones S. miRBase: from microRNA sequences to function. Nucleic Acids Res. 2019;47:D155–62. Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25:1754–60. Pedersen BS, Quinlan AR. cyvcf2: fast, flexible variant analysis with Python. Hancock J, editor. Bioinformatics. 2017;33:1867–9. Katoh K, Misawa K, Kuma K ichi, Miyata T. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res. 2002;30:3059–66. Hunter JD. Matplotlib: A 2D Graphics Environment. Comput Sci Eng. 2007;9:90–5. Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17:261–72. Stéfan van der Walt SCC and GV. The NumPy array: a structure for efficient numerical computation. Comput Sci Eng. 2011;13:22–30. McKinney W. Data structures for statistical computing in Python. In: Proceedings of the 9th Python in Science Conference. 2010. p. 51–6. Waskom M. seaborn: statistical data visualization. J Open Source Softw. 2021;6:3021. Supplementary Files Additionalfile1.pdf Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2222130","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":154103133,"identity":"50db84a6-bf81-4a88-b154-eed9ca0b6e8c","order_by":0,"name":"Chong He","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAtklEQVRIiWNgGAWjYBACA2YGNgaGCjkw58AD4rWcMYZoSSBKCwNQC2MbRAsDUVrM2ZmPPfg4z0DO4Nrhh0Bb7OR0GwhosWxmSzecuc3AWHJ2mgFQS7Kx2QFCDjvMYybNu+1PYr90AkjLgcRtRGn5O8cgsU06/QMJWhgbDIC25BBpC9AvaZI9x0B+ySk4kGBAhF/M+Q8fk/hRAwyx2+mbP3yosJMjqAXdnaQpHwWjYBSMglGAAwAAyVU/pGAPBoIAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-6194-3374","institution":"Southern Medical University School of Basic Medical Sciences","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Chong","middleName":"","lastName":"He","suffix":""},{"id":154103134,"identity":"84e7e667-78e1-4d0f-8cf2-c174f39bc228","order_by":1,"name":"Hao Zhu","email":"","orcid":"https://orcid.org/0000-0001-7384-3840","institution":"Southern Medical University School of Basic Medical Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hao","middleName":"","lastName":"Zhu","suffix":""}],"badges":[],"createdAt":"2022-10-31 13:41:20","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2222130/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2222130/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":29659280,"identity":"cb1ce6b5-04f6-467c-984b-b28d39842b99","added_by":"auto","created_at":"2022-11-29 15:31:56","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":120076,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ea \u003c/strong\u003eIllustration of piRNA-similar sequences and 'attracting SNVs'.\u003cstrong\u003e \u003c/strong\u003epiRNA-similar sequences are segments in protein-coding genes (PCGs) that share similarity with piRNAs. piRNA-similar sequences can be categorized according to the strand relationship between mRNA and piRNA. Antisense piRNA-similar sequences are potential target sites of piRNAs, while sense piRNA-similar sequences are not. Attracting SNVs are those single nucleotide variants (SNVs) that can reduce mismatches in piRNA/piRNA-similar sequence pairs. Antisense attracting SNVs can increase the piRNA-PCG targeting affinity, while sense attracting SNVs have no effect on the piRNA-PCG targeting relationship. \u003cstrong\u003eb\u003c/strong\u003e Illustration of the information provided by attracting SNVs. The occurrence of attracting SNVs across piRNA/piRNA-similar sequence pairs and the allele frequencies of attracting SNVs in the human population provide information for our analyses.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2222130/v1/cefb98b2184e8fbc61ffb9c6.jpeg"},{"id":29659283,"identity":"d15f8da8-9325-4842-887a-eda61ead00f2","added_by":"auto","created_at":"2022-11-29 15:31:56","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":160458,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eAlu\u003c/em\u003e elements significantly contribute to the presence of potential piRNA target sites in human PCGs.\u003cstrong\u003e a\u003c/strong\u003eThe blue pie charts show the proportion of antisense piRNA/piRNA-similar sequence pairs that have piRNA-similar sequences derived from TEs (left) or \u003cem\u003eAlu\u003c/em\u003eelements (right). The pink pie charts show the proportion of antisense piRNA/piRNA-similar sequence pairs that have piRNAs derived from TEs (left) or \u003cem\u003eAlu\u003c/em\u003eelements (right). TEs, especially \u003cem\u003eAlu\u003c/em\u003e elements, are major contributors to antisense piRNA/piRNA-similar sequence pairs. \u003cstrong\u003eb \u003c/strong\u003e\u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|\u003cem\u003eAlu\u003c/em\u003e) is much larger than \u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|non-\u003cem\u003eAlu\u003c/em\u003e), which indicates that antisense piRNA-similar sequences are more likely to be present in \u003cem\u003eAlu\u003c/em\u003e-derived regions than in non-\u003cem\u003eAlu\u003c/em\u003e-derived regions. Additionally, the average count of piRNA-similar sequences across \u003cem\u003eAlu\u003c/em\u003e-derived regions is much larger than that across non-\u003cem\u003eAlu\u003c/em\u003e-derived regions, which indicates that \u003cem\u003eAlu\u003c/em\u003e-derived regions can share sequence similarity with more piRNAs than non-\u003cem\u003eAlu\u003c/em\u003e-derived regions. These indicators collectively suggest that \u003cem\u003eAlu\u003c/em\u003e elements contribute significantly to potential piRNA target sites in human PCGs. Specifically, \u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|\u003cem\u003eAlu\u003c/em\u003e) is 0.9819, which indicates that the presence of \u003cem\u003eAlu\u003c/em\u003e elements almost always results in antisense piRNA-similar sequences in human PCGs.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2222130/v1/4846c14dbe227351f1c294fb.jpeg"},{"id":29660006,"identity":"526ec9b5-702e-4a6c-8f02-122300373ef7","added_by":"auto","created_at":"2022-11-29 15:39:56","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":93252,"visible":true,"origin":"","legend":"\u003cp\u003eThe occurrence of SNVs tends to increase the sequence similarity between piRNA-similar sequences and their related piRNAs. The results of analyzing SNVs in piRNA-similar sequences are shown here. \u003cstrong\u003ea\u003c/strong\u003e If the occurrence of SNVs is mutually independent in piRNA-similar sequences and their related piRNAs, an SNV occurring in a mismatch site should be irrelevant to the nucleotide paired with it. Hence, it is expected that the proportion of occurrences of attracting SNVs will be 1/3. \u003cstrong\u003eb\u003c/strong\u003e The observed proportion of occurrences of attracting SNVs. The numbers of occurrences of attracting SNVs and nonattracting SNVs are shown in parentheses. A pink bar and a blue bar indicate the proportions of occurrences of attracting and nonattracting SNVs, respectively. In both sense and antisense piRNA-similar sequences, the proportion of occurrences of attracting SNVs was observed to be significantly higher than 1/3 (indicated by an asterisk, \u003cem\u003ep \u003c/em\u003e\u0026lt; 0.05, one-tailed binomial test) using both 1000 Genomes data and UK10K data.\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2222130/v1/075e586c9e0846a3aa004401.jpeg"},{"id":29659286,"identity":"0f98ec94-b956-43c5-af4d-f43c250c3aba","added_by":"auto","created_at":"2022-11-29 15:31:56","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":158749,"visible":true,"origin":"","legend":"\u003cp\u003eAn excess of shared SNVs exists in piRNA/piRNA-similar sequence pairs. \u003cstrong\u003ea \u003c/strong\u003eIllustration of a 'shared SNV'. In a piRNA/piRNA-similar sequence pair, if a pair of identical sites both contain SNVs and the two SNVs are also identical, the piRNA/piRNA-similar sequence pair has a shared SNV. \u003cstrong\u003eb\u003c/strong\u003e In both sense and antisense piRNA/piRNA-similar sequence pairs, the number of occurrences of shared SNVs in the real data is significantly larger than that in the control dataset (larger than the mean by more than 5 × the standard deviation (SD)), indicating the impact of interlocus gene conversion.\u003c/p\u003e","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2222130/v1/65cbd4a1af88004cc8300b67.jpeg"},{"id":29660007,"identity":"df98de8b-f9b5-4b20-b28a-aabf4de0c619","added_by":"auto","created_at":"2022-11-29 15:39:56","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":248283,"visible":true,"origin":"","legend":"\u003cp\u003eSequences surrounding piRNA-similar sequences and their related piRNAs also share sequence similarity. The two panels show the results of sense and antisense piRNA/piRNA-similar sequence pairs. In each repeat, 100 piRNA/piRNA-similar sequence pairs are randomly sampled and analyzed. For each piRNA/piRNA-similar sequence pair, sequences are expanded by 150 bp in the upstream and downstream directions, and the sequence similarity of the most similar block (maximum similarity) is calculated. The median value of maximum similarity is calculated for each repeat. In all 30 repeats, the median value of maximum similarity is ≥ 80%.\u003c/p\u003e","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2222130/v1/11729ee9bace2d68706b139b.jpeg"},{"id":29660008,"identity":"850a0774-caa9-4bc8-9c71-7860a3b0e2e2","added_by":"auto","created_at":"2022-11-29 15:39:56","extension":"jpeg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":173163,"visible":true,"origin":"","legend":"\u003cp\u003eAntisense attracting SNVs tend to show low allele frequencies.We compared attracting SNVs in antisense piRNA-similar sequences (antisense attracting SNVs) with attracting SNVs in sense piRNA-similar sequences (sense attracting SNVs) to examine whether the distribution of allele frequencies for antisense attracting SNVs is skewed toward high or low frequencies. The results with allele frequencies restricted to ≥ 0.001 are shown here. An asterisk indicates that the allele frequencies of antisense attracting SNVs are significantly lower than the allele frequencies of sense attracting SNVs (\u003cem\u003ep\u003c/em\u003e \u0026lt; 0.05, one-tailed Mann‒Whitney U test). In the top panels, each histogram shows the overall distribution of allele frequencies for sense/antisense attracting SNVs, pooling all categories of SNVs together (at a log10 scale). In comparison with the distribution of allele frequencies for sense attracting SNVs (blue), the distribution of allele frequencies for antisense attracting SNVs (red) is skewed toward low allele frequencies. In the bottom panels, the distribution of allele frequencies for sense/antisense attracting SNVs is shown separately for each category of SNVs (at a log10 scale). In all categories, the distribution of allele frequencies for antisense attracting SNVs is skewed toward low allele frequencies.\u003c/p\u003e","description":"","filename":"floatimage6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2222130/v1/d4759ac3f3e7e2aacdd9ee0e.jpeg"},{"id":32712343,"identity":"b6841725-dd39-4fe6-a988-2ae568356cf8","added_by":"auto","created_at":"2023-02-09 15:35:31","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":988459,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2222130/v1/5beec413-5e73-4cfe-ba55-caa3590c52e8.pdf"},{"id":29659281,"identity":"f33b7ccf-e064-40fb-94ab-b3bbf037b167","added_by":"auto","created_at":"2022-11-29 15:31:56","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":4095931,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2222130/v1/ef9cdf60beaca9e85138a11f.pdf"}],"financialInterests":"","formattedTitle":"Human Alu elements promote the establishment and enhancement of piRNA-protein-coding gene targeting relationships","fulltext":[{"header":"Introduction","content":"\u003cp\u003ePIWI-interacting RNAs (piRNAs) constitute the most diverse category of small RNAs in animals. These small RNAs guide Argonaute proteins of the PIWI clade (e.g., HIWI in humans, PIWI in fruit flies) to corresponding targets, triggering a series of reactions at the posttranscriptional and/or transcriptional levels, leading to the repression of target genes [\u003cspan additionalcitationids=\"CR2 CR3 CR4\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. In general, the target recognition of piRNAs follows base-pairing rules, in which RNA segments that are substantially similar (generally allowing\u0026thinsp;\u0026le;\u0026thinsp;4 mismatches) and reverse complementary to piRNA members will become target sites [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Therefore, genes that produce transcripts carrying these piRNA target sites could be repressed by the piRNA pathway.\u003c/p\u003e \u003cp\u003eThe genomes of humans and other eukaryotes have been widely shaped by various transposable elements (TEs) [\u003cspan additionalcitationids=\"CR9\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. TEs are usually regarded as 'molecular parasites' because they are able to relocate or amplify themselves within the genome. It is well established that the piRNA pathway can serve as the 'immune system' of the genome to suppress the transposition of TEs [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Hence, the interplay between TEs and piRNAs plays an important role in shaping the genome [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Some TEs can integrate into protein-coding genes (PCGs). These TEs are retained in PCGs because either they do not strongly disrupt the normal functions of their host genes or they confer beneficial characteristics upon their host genes [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Although there seems to be no need to suppress these TEs, recent evidence suggests that TEs incorporated into PCGs are also targeted by piRNAs [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Hence, TEs incorporated into PCGs could influence their host genes via interaction with piRNAs. To assess the significance of such a piRNA-mediated influence, the first question that needs to be addressed is how likely TEs are to cause piRNA target sites to appear in PCGs. On the other hand, factors that can affect the evolutionary trajectory of these piRNA target sites also remain to be explored.\u003c/p\u003e \u003cp\u003eIn the present work, we carried out a series of analyses using human piRNA sequencing data and population genetic data [\u003cspan additionalcitationids=\"CR16 CR17 CR18\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. We analyzed potential piRNA target sites in human PCGs and single nucleotide variants (SNVs) in these regions. We further recognized that the presence of a class of primate-specific TEs, \u003cem\u003eAlu\u003c/em\u003e elements, is highly likely to cause piRNA target sites to appear in human PCGs. Additionally, we found that a mutagenic process induced by \u003cem\u003eAlu\u003c/em\u003e elements, interlocus gene conversion (IGC), tends to increase sequence similarity between piRNAs and their potential target sites. Hence, the integration of \u003cem\u003eAlu\u003c/em\u003e elements into human PCGs promotes the establishment of the piRNA-PCG targeting relationship, and IGC induced by \u003cem\u003eAlu\u003c/em\u003e elements promotes the enhancement of the piRNA-PCG targeting relationship. Moreover, we recognized a footprint indicating that piRNA-PCG targeting relationships impose a piRNA-related selective constraint on the evolution of human PCGs. In brief, this work reveals how the interplay between \u003cem\u003eAlu\u003c/em\u003e elements and piRNAs influences the evolution of human PCGs.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eOur analyses were begun by mapping unique human piRNA sequences to the longest mRNA sequences of human PCGs using BWA allowing\u0026thinsp;\u0026le;\u0026thinsp;4 mismatches. Through this procedure, a fraction of sequences in PCGs were found to share similarity with piRNAs. Here, we refer to these sequences as \u0026apos;piRNA-similar sequences\u0026apos; (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ea). piRNA/piRNA-similar sequence pairs can be further divided into two classes: sense and antisense (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ea). For a sense piRNA/piRNA-similar sequence pair, the piRNA strand is oriented in the same direction as the mRNA strand of the piRNA-similar sequence. Sense piRNA-similar sequences are merely sequences that are similar to piRNAs (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ea). In contrast, for an antisense piRNA/piRNA-similar sequence pair, the piRNA strand is complementary to the mRNA strand of the piRNA-similar sequence. As piRNAs recognize targets according to base-pairing rules, antisense piRNA-similar sequences have the potential to form targeting relationships with their related piRNAs (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ea). Additionally, we refer to SNVs that can reduce mismatches between piRNA-similar sequences and their related piRNAs as \u0026apos;attracting SNVs\u0026apos; (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ea). Furthermore, we refer to attracting SNVs in sense and antisense piRNA/piRNA-similar sequence pairs as \u0026apos;sense attracting SNVs\u0026apos; and \u0026apos;antisense attracting SNVs\u0026apos;, respectively. Antisense attracting SNVs can increase piRNA-PCG targeting affinity, whereas sense attracting SNVs have no effect on the piRNA-PCG targeting relationship (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ea). We mainly analyzed piRNA-similar sequences and attracting SNVs in this study. Two aspects of attracting SNVs were considered: their occurrence across piRNA/piRNA-similar sequence pairs and their allele frequencies in the human population (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003eb).\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eThe presence of\u003c/strong\u003e \u003cspan class=\"BoldItalic\" name=\"Emphasis\" type=\"BoldItalic\"\u003eAlu\u003c/span\u003e \u003cstrong\u003eelements almost always results in potential piRNA target sites in human PCGs\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eAntisense piRNA-similar sequences in PCGs have the potential to form targeting relationships with piRNAs. In this section, we show that \u003cem\u003eAlu\u003c/em\u003e elements significantly contribute to the presence of antisense piRNA-similar sequences in human PCGs. First, we analyzed how many pairing relationships between antisense piRNA-similar sequences and their related piRNAs are contributed by TEs and what kinds of TEs are major contributors. We found that in 88.88% of antisense piRNA/piRNA-similar sequence pairs, the antisense piRNA-similar sequence was TE-derived. Notably, in 85.88% of antisense piRNA/piRNA-similar sequence pairs, the antisense piRNA-similar sequence was \u003cem\u003eAlu\u003c/em\u003e derived (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003ea), accounting for 96.66% of pairing relationships related to TE-derived piRNA-similar sequences. We also investigated how many pairing relationships were related to TE-derived piRNAs. piRNAs were mapped to the reference genome using BWA allowing no mismatches. Thus, information about their possible genomic locations was obtained. We found that in 98.5% of antisense piRNA/piRNA-similar sequence pairs, the piRNA could be mapped to TE-derived regions. Additionally, in 96.2% of antisense piRNA/piRNA-similar sequence pairs, the piRNA could be mapped to an \u003cem\u003eAlu\u003c/em\u003e-derived region (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003ea), accounting for 97.66% of pairing relationships related to TE-derived piRNAs. This result suggests that TEs, especially \u003cem\u003eAlu\u003c/em\u003e elements, contribute significantly to the emergence of the targeting relationship between piRNAs and PCGs. In contrast, among all of the PCG sequences analyzed in this study, only 9.56% were TE-derived, and among all of the piRNAs analyzed in this study, only 37.26% were TE-derived.\u003c/p\u003e\n\u003cp\u003eMoreover, we calculated the following conditional probabilities to further investigate the contribution of \u003cem\u003eAlu\u003c/em\u003e elements to the targeting relationship between piRNAs and human PCGs. We define \u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|\u003cem\u003eAlu\u003c/em\u003e) as the probability that a nucleotide is located in at least one antisense piRNA-similar sequence, given that it is known that the nucleotide is located in an \u003cem\u003eAlu\u003c/em\u003e-derived region. \u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|\u003cem\u003eAlu\u003c/em\u003e) can be calculated as the total length of the intersecting regions of antisense piRNA-similar sequences and \u003cem\u003eAlu\u003c/em\u003e-derived regions divided by the total length of \u003cem\u003eAlu\u003c/em\u003e-derived regions (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eb). \u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|non-\u003cem\u003eAlu\u003c/em\u003e) can also be calculated in a similar manner. Strikingly, we found that the value of \u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|\u003cem\u003eAlu\u003c/em\u003e) was 0.9819 (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eb), indicating that the presence of \u003cem\u003eAlu\u003c/em\u003e elements showed a 0.9819 probability of resulting in antisense piRNA-similar sequences. In other words, the presence of \u003cem\u003eAlu\u003c/em\u003e elements almost always results in antisense piRNA-similar sequences in human PCGs. In contrast, the value of \u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|non-\u003cem\u003eAlu\u003c/em\u003e) only reached 0.0882 (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eb).\u003c/p\u003e\n\u003cp\u003eAdditionally, in a PCG, a region can be present in multiple antisense piRNA/piRNA-similar sequence pairs, as it can share sequence similarity with multiple piRNAs. We calculated the average count of antisense piRNA/piRNA-similar sequence pairs across \u003cem\u003eAlu-\u003c/em\u003ederived regions and across non-\u003cem\u003eAlu-\u003c/em\u003ederived regions in human PCGs (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eb). The average count of antisense piRNA/piRNA-similar sequence pairs across \u003cem\u003eAlu-\u003c/em\u003ederived regions was 28.37, while the average count of antisense piRNA/piRNA-similar sequence pairs across non-\u003cem\u003eAlu-\u003c/em\u003ederived regions was 0.0918. The former value was 309 times greater than the latter. This indicates that over the same sequence length, \u003cem\u003eAlu\u003c/em\u003e-derived sequences can develop many more piRNA-PCG targeting relationships than non-\u003cem\u003eAlu-\u003c/em\u003ederived sequences. For sense piRNA/piRNA-similar sequence pairs, a close association with \u003cem\u003eAlu\u003c/em\u003e elements was also detected (Additional file 1: Figure. S1). Collectively, these results suggest that \u003cem\u003eAlu\u003c/em\u003e elements are major contributors to the sequence similarity shared between piRNAs and PCGs. As a consequence of this effect, the presence of \u003cem\u003eAlu\u003c/em\u003e elements has a high probability resulting in antisense piRNA-similar sequences in human PCGs. Hence, \u003cem\u003eAlu\u003c/em\u003e elements significantly contribute to the presence of potential piRNA target sites in human PCGs.\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eIGC induced by\u003c/strong\u003e \u003cspan class=\"BoldItalic\" name=\"Emphasis\" type=\"BoldItalic\"\u003eAlu\u003c/span\u003e \u003cstrong\u003eelements tends to increase the sequence similarity between piRNAs and their potential target sites\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eAbove, we have shown that \u003cem\u003eAlu\u003c/em\u003e elements are major contributors to the sequence similarity shared between piRNAs and PCGs. TEs can induce IGC, which is a mutagenic process in which a genomic region mistakenly utilizes its homologs as templates to \u0026apos;repair\u0026apos; its DNA [\u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e\u0026ndash;\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e]. Hence, \u003cem\u003eAlu\u003c/em\u003e elements might exert a secondary influence on piRNA-similar sequences and their related piRNAs via IGC. In this section, we provide evidence of the influence of IGC.\u003c/p\u003e\n\u003cp\u003eIn a piRNA/piRNA-similar sequence pair, whether an SNV is considered an attracting SNV depends on the nucleotide paired with the SNV. If the occurrence of SNVs is mutually independent in piRNA-similar sequences and their related piRNAs, an SNV occurring in a mismatch site should show a 1/3 probability of being an attracting SNV (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003ea). Under the influence of IGC, the occurrence of genetic variants will not be mutually independent in piRNA-similar sequences and their related piRNAs. Hence, we first evaluated the influence of IGC by examining whether the occurrence of attracting SNVs deviates from the expectation of 1/3. We analyzed SNVs occurring in mismatch sites. In each piRNA/piRNA-similar sequence pair, for SNVs occurring in mismatch sites, we counted the occurrences of attracting SNVs and the occurrences of nonattracting SNVs. Then, we summed the count results for each piRNA/piRNA-similar sequence pair and calculated the proportions of occurrences of attracting SNVs and nonattracting SNVs. Interestingly, the observed proportion of the occurrences of attracting SNVs was approximately 60% rather than 1/3. In both SNVs in piRNA-similar sequences and SNVs in their related piRNAs, this phenomenon was consistently observed (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eb and Additional file 1: Figure. S2; \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05, binomial test). Additionally, this phenomenon was consistently observed using both 1000 Genomes data and UK10K data (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eb and Additional file 1: Figure. S2). Notably, for both sense and antisense piRNA/piRNA-similar sequence pairs, there was an excess of attracting SNVs (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eb and Additional file 1: Figure. S2). Sense piRNA/piRNA-similar sequence pairs do not reflect any piRNA-PCG targeting relationship. As the excess of attracting SNVs was observed in sense piRNA/piRNA-similar sequence pairs, this phenomenon is unlikely to reflect the impact of natural selection. In other words, the observed phenomenon is unlikely to be caused by an effect on the persistence of genetic variants. Alternatively, it is likely to be caused by an effect on the occurrence of genetic variants\u0026mdash;i.e., mutation. Thus, the observed phenomenon suggests that a mutagenic process homogenizes piRNA-similar sequences and their related piRNAs. To date, IGC is the only recognized mutagenic process that has such an effect [\u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e\u0026ndash;\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e]. To provide further evidence of IGC, we conducted the analysis with another indicator. An excess of shared nucleotide changes (divergences or polymorphisms) between different loci has been proposed as an indicator of IGC [\u003cspan class=\"CitationRef\"\u003e24\u003c/span\u003e\u0026ndash;\u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e]. We analyzed SNVs occurring in identical sites (above, we analyzed SNVs occurring in mismatch sites). If we observed that a pair of SNVs were located in a pair of identical sites and that this pair of SNVs were also identical, we referred to this SNV pair as a \u0026apos;shared SNV\u0026apos; (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003ea). In each piRNA/piRNA-similar sequence pair, we counted the occurrences of shared SNVs. Then, we summed the count results for each piRNA/piRNA-similar sequence pair. For comparison, we built a control dataset by shuffling the one-to-one relationships of the identical sites in each piRNA/piRNA-similar sequence pair (see Materials and Methods). We found that the number of occurrences of shared SNVs in the real data was significantly higher than that in the control dataset, indicating an excess of shared SNVs (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eb; see the logic of significance testing in Materials and Methods). Similar to the above analysis, in both sense and antisense piRNA/piRNA-similar sequence pairs, an excess of shared SNVs was detected (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eb). Thus, the excess of shared SNVs was also caused by a mutagenic process that homogenizes piRNA-similar sequences and their related piRNAs. Collectively, the results provide two parallel lines of evidence, the excess of attracting SNVs and the excess of shared SNVs, suggesting that piRNA-similar sequences and their related piRNAs are under the influence of IGC.\u003c/p\u003e\n\u003cp\u003eThe above results have shown that piRNA-similar sequences and their related piRNAs are influenced by IGC. Nevertheless, there is still a question to be addressed. piRNA-similar sequences and their related piRNAs are 25\u0026ndash;32 bp, and such short homologous sequences are unlikely to induce IGC [\u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e28\u003c/span\u003e]. Hence, if piRNA-similar sequences and their related piRNAs are under the influence of IGC, regions that share sequence similarity should not be restricted to the regions of piRNA-similar sequences and their related piRNAs. We expanded the genomic sequences of piRNA-similar sequences and their related piRNAs by 150 bp in both the upstream and downstream directions (total length of the expanded region\u0026thinsp;=\u0026thinsp;piRNA length\u0026thinsp;+\u0026thinsp;150 bp \u0026times; 2). In each expanded region, we analyzed the sequence similarity of the most similar block (maximum similarity) (Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e). Our analysis included 30 repeats, and each repeat included 100 expanded regions of randomly sampled piRNA/piRNA-similar sequence pairs. For each repeat, we calculated the median value of maximum similarity. The most similar block was set to different lengths in the analysis. When the length of the most similar block was set to 200 bp, in all 30 repeats, the median value of maximum similarity was \u0026ge;\u0026thinsp;80% (Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e and Additional file 1: Figure. S3). When the length of the most similar block was set to a shorter length (100 or 150 bp), the median value of maximum similarity was higher (Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e and Additional file 1: Figure. S3). This result indicates that if a piRNA shares sequence similarity with a segment of a PCG, their flanking regions will also share sequence similarity.\u003c/p\u003e\n\u003cp\u003eIGC at a high rate usually requires completely identical regions that are longer than 200 bp, referred to as the \u0026apos;minimal efficient processing segment\u0026apos; (MEPS) [\u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e]. Although the sequence similarity between flanking regions was reasonably high in the above analysis, it did not satisfy the MEPS criterion. Nevertheless, there is evidence suggesting that IGC can still occur between sequences that are not completely identical and shorter than the MEPS criterion, although the rate is relatively low [\u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e]. In addition, if sequence similarity satisfies the MEPS criterion, the rate of IGC will be higher than the spontaneous mutation rate by orders of magnitude [\u003cspan class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e32\u003c/span\u003e]. With such a high rate of IGC, we should observe that the allele frequencies of attracting SNVs are higher than those of nonattracting SNVs by orders of magnitude. In our analyses, we did observe that the allele frequencies of attracting SNVs were higher than those of nonattracting SNVs (compare Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e and Additional file 1: Figure. S4), which was consistent with a scenario in which a mutagenic process was additionally introduced. However, the difference did not reach orders of magnitude (compare Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e and Additional file 1: Figure. S4). Thus, our results were consistent with a scenario in which although an influence of IGC existed, the rate of IGC was relatively low. Moreover, as genomic analyses utilize genetic information that accumulates in DNA over time, even though the rate of IGC is relatively low, IGC can still leave a footprint for the analysis [\u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e]. Additionally, increasing evidence suggests that IGC can occur among \u003cem\u003eAlu\u003c/em\u003e elements [\u003cspan class=\"CitationRef\"\u003e33\u003c/span\u003e\u0026ndash;\u003cspan class=\"CitationRef\"\u003e35\u003c/span\u003e]. Overall, as the sequence similarity shared between piRNAs and PCGs is mainly contributed by \u003cem\u003eAlu\u003c/em\u003e elements, sequences surrounding piRNA-similar sequences and their related piRNAs also share extensive similarity. These sequences sharing extensive similarity allow the influence of IGC to be exerted upon piRNA-similar sequences and their related piRNAs. Consequently, an excess of attracting SNVs occurs in piRNA-similar sequences and their related piRNAs. Antisense attracting SNVs are a subset of attracting SNVs that can increase piRNA-PCG targeting affinity. IGC promotes the occurrence of antisense attracting SNVs, thus constituting an evolutionary force that reinforces the piRNA-PCG targeting relationship.\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eNatural selection opposes the promotion effect of human\u003c/strong\u003e \u003cspan class=\"BoldItalic\" name=\"Emphasis\" type=\"BoldItalic\"\u003eAlu\u003c/span\u003e \u003cstrong\u003eelements on the formation of piRNA-PGC targeting relationships.\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eIn the above sections, we show that human \u003cem\u003eAlu\u003c/em\u003e elements promote both the establishment and enhancement of piRNA-PCG targeting relationships. The presence of piRNA target sites in PCGs should be relevant to the fitness of organisms. piRNA-related natural selection thereby can influence the evolution of human PCGs. If so, piRNA-related natural selection should leave a signature in piRNA target sites. As antisense attracting SNVs can increase piRNA-PCG targeting affinity (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e), analyzing whether they tend to show high or low allele frequencies can provide information about the action of piRNA-related natural selection. If antisense attracting SNVs tend to show high allele frequencies, it indicates that natural selection favors the influence of antisense attracting SNVs. Hence, with regard to the targeting relationship between piRNAs and PCGs, natural selection acts as an \u0026apos;attractive force\u0026apos;. Otherwise, the opposite scenario is indicated, in which natural selection acts as a \u0026apos;repulsive force\u0026apos;, against the targeting relationship between piRNAs and PCGs. As mentioned in the last section, IGC could cause a systematic difference in allele frequencies between attracting SNVs and nonattracting SNVs (compare Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e and Additional file 1: Figure. S4). Hence, IGC could be a confounding factor for inferring the footprint of natural selection. IGC influences both sense attracting SNVs and antisense attracting SNVs, but sense attracting SNVs have no effect on the piRNA-PCG targeting relationship. Therefore, we compared antisense attracting SNVs with sense attracting SNVs, rather than nonattracting SNVs, to control the confounding effect of IGC.\u003c/p\u003e\n\u003cp\u003eWe observed that antisense attracting SNVs tended to show lower allele frequencies than sense attracting SNVs, which supports the \u0026apos;repulsive force\u0026apos; hypothesis rather than the \u0026apos;attractive force\u0026apos; hypothesis (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e and Additional file 1: Figure. S5; \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05, one-tailed Mann‒Whitney U test). In particular, when we restricted the allele frequencies of SNVs to \u0026ge;\u0026thinsp;0.001, the tendency toward low allele frequencies became more prominent (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e). Moreover, we examined different categories of SNVs in PCGs (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e and Additional file 1: Figure. S5). Without restricting allele frequencies, antisense attracting SNVs showed a tendency toward low allele frequencies in 5\u0026apos;UTR, 3\u0026apos;UTR and synonymous SNVs (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05, Mann‒Whitney U test; Additional file 1: Figure. S5), but we did not detect a significant tendency in nonsynonymous SNVs (Additional file 1: Figure. S5). One possible explanation is that nonsynonymous SNVs can alter the amino acid sequence, and this effect could obscure the influence of nonsynonymous SNVs on the piRNA targeting relationship. Nevertheless, when allele frequencies were restricted to \u0026ge;\u0026thinsp;0.001, for all categories of SNVs, including nonsynonymous SNVs, we detected a pattern in which antisense attracting SNVs tended to show low frequencies in comparison with sense attracting SNVs (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e; \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05, one-tailed Mann‒Whitney U test).\u003c/p\u003e\n\u003cp\u003eAdditionally, we repeated the above analyses using another piRNA targeting rule to identify piRNA-similar sequences [\u003cspan class=\"CitationRef\"\u003e7\u003c/span\u003e]. This rule considers the \u0026apos;seed region\u0026apos; of the piRNA target sites, rather than simply requiring\u0026thinsp;\u0026le;\u0026thinsp;4 mismatches (see Materials and Methods) [\u003cspan class=\"CitationRef\"\u003e7\u003c/span\u003e]. Using this rule for identifying potential piRNA target sites, we observed a pattern very similar to that described above (Additional file 1: Figure. S6). Furthermore, in both the 1000 Genomes data and the UK10K data, the tendency toward low allele frequencies for antisense attracting SNVs was consistently observed (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e). The observation of a consistent pattern under different conditions suggests that this pattern is not caused by chance. Finally, we did not detect a tendency toward low allele frequencies for antisense nonattracting SNVs in comparison with sense nonattracting SNVs (Additional file 1: Figure. S4). Collectively, our results suggest that antisense piRNA-similar sequences are not merely sequences that share similarity with piRNAs. Instead, the presence of these regions in human PCGs imposes a piRNA-related selective constraint on the evolution of human PCGs. Furthermore, the phenomenon that antisense attracting SNVs tend to show low allele frequencies suggests that the action of natural selection is likely to function as a \u0026apos;repulsive force\u0026apos; rather than an \u0026apos;attractive force\u0026apos;. Hence, the promotion effect of human \u003cem\u003eAlu\u003c/em\u003e elements on the formation of piRNA-PGC targeting relationships is opposed by natural selection.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIt is well established that the transposition of TEs can disrupt genome stability, leading to deleterious effects. In addition to the direct influence of transposition, the homologous sequences produced by TEs can further induce posttransposition influence on the genome [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. The target recognition of piRNAs depends on the sequence similarity between piRNAs and their target sites. Hence, homologous sequences produced by TEs could have an influence on the target relationship between piRNAs and their target sites. Previously, researchers recognized that mRNAs harboring TEs tend to be upregulated in \u003cem\u003ePiwil1\u003c/em\u003e mutant mice [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. This evidence suggests that TEs could contribute to the establishment of piRNA-PCG targeting relationships [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Although previous studies have provided insight into the roles of TEs in the formation of piRNA targeting relationships, evolutionary factors that influence the evolutionary trajectories of piRNA targeting relationships remain to be explored.\u003c/p\u003e \u003cp\u003eTo elucidate the impact of TEs on the establishment of piRNA-PCG targeting relationships, it is better to quantify such an impact. In this work, we calculated the conditional probability, \u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|\u003cem\u003eAlu\u003c/em\u003e). Through this calculation, the promotion effect of \u003cem\u003eAlu\u003c/em\u003e elements on the establishment of piRNA target sites can be quantified as a probability. Currently, it remains elusive whether there exists a class of TEs whose presence is highly likely to result in piRNAs target sites. Although calculating \u003cem\u003eP\u003c/em\u003e(antisense piRNA-similar|\u003cem\u003eAlu\u003c/em\u003e) is not complicated, our result explicitly suggests that the presence of \u003cem\u003eAlu\u003c/em\u003e elements is highly likely to cause piRNA target sites to appear in human PCGs. This approach allowed us to reveal the significant role of \u003cem\u003eAlu\u003c/em\u003e elements in the establishment of piRNA-PCG targeting relationships.\u003c/p\u003e \u003cp\u003eIGC can affect the fitness of organisms via various mechanisms, such as altering amino acid sequences and rewiring regulatory networks [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. However, its influence on the piRNA-PCG targeting relationship has not been explored. In this work, we found that human \u003cem\u003eAlu\u003c/em\u003e elements allow the influence of IGC to be exerted upon piRNA-similar sequences and their related piRNAs. In particular, IGC induced by \u003cem\u003eAlu\u003c/em\u003e elements facilitates the occurrence of attracting SNVs. Thus, IGC induced by \u003cem\u003eAlu\u003c/em\u003e elements makes the sequence similarity between piRNAs and PCGs prone to increase, which indicates that in addition to promoting the establishment of piRNA-PCG targeting relationships, human \u003cem\u003eAlu\u003c/em\u003e elements also promote the enhancement of piRNA-PCG targeting relationships.\u003c/p\u003e \u003cp\u003eFurthermore, we observed that antisense attracting SNVs tend to show lower allele frequencies than sense attracting SNVs. This footprint of piRNA-related natural selection suggests that the presence of antisense piRNA-similar sequences in human PCGs imposes a piRNA-related selective constraint on human PCGs. Meanwhile, this footprint of natural selection suggests that natural selection opposes the promotion effect of \u003cem\u003eAlu\u003c/em\u003e elements on the formation of piRNA-PCG targeting relationships. Hence, the uncovered promotion effect of \u003cem\u003eAlu\u003c/em\u003e elements on the formation of piRNA-PCG targeting relationships is unlikely to be a beneficial feature in the evolution of the human genome.\u003c/p\u003e \u003cp\u003eIt may be a little counterintuitive that evolution allows an adverse feature to appear in our genomes. Nevertheless, our observation is in accordance with current knowledge about the evolution of TEs and piRNAs. TEs undergo self-amplification or relocation in the genome. As this process continues, some TEs eventually integrate into piRNA clusters (regions that produce piRNAs). Once TEs integrate into piRNA clusters, they can become templates of piRNAs. Thus, the genome evolves the ability to generate piRNAs suppressing these TEs [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. This evolutionary model is widely accepted to explain how the genome continuously adjusts the piRNA repertoire in response to the activity of TEs [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. However, before TEs integrate into piRNA clusters, they might have integrated into other genomic regions, including PCGs. Thus, as the genome evolves the ability to suppress TEs, nonfunctional or deleterious piRNA-PCG targeting relationships can also appear. As the advantage of suppressing TEs can outweigh the disadvantage of forming nonfunctional or deleterious piRNA-PCG targeting relationships, these piRNA-PCG targeting relationships are retained in the genome. This might explain why the presence of \u003cem\u003eAlu\u003c/em\u003e elements in PCGs promotes the establishment of piRNA-PCG targeting relationships. In addition, \u003cem\u003eAlu\u003c/em\u003e elements can exert a secondary influence on piRNA-PCG targeting relationships via IGC. IGC is a mutagenic process that allows the sequence similarity shared between piRNAs and PCGs to increase. Thus, even though a large number of piRNA-PCG targeting relationships are nonfunctional or deleterious, there is a tendency to enhance them in the genome. Together, the promotion effect of \u003cem\u003eAlu\u003c/em\u003e elements on the formation of piRNA-PCG targeting relationships is likely to be a bitter pill to swallow in the evolution of the human genome. Therefore, the response of natural selection is more likely to oppose, rather than favor, such an effect. Overall, the presence of \u003cem\u003eAlu\u003c/em\u003e elements in human PCGs allows multiple piRNA-related factors to exert influence on the evolution of human PCGs. The interplay between \u003cem\u003eAlu\u003c/em\u003e elements and piRNAs adds an additional complexity to the evolution of human PCGs. Recognizing the piRNA-mediated influence of \u003cem\u003eAlu\u003c/em\u003e elements reported here extends our understanding of how TEs shape our genomes via various evolutionary forces.\u003c/p\u003e \u003cp\u003eAt the end, we would like to discuss some treatments adopted in our analyses. In the present work, we mainly analyzed piRNA-similar sequences rather than piRNA target sites predicted by more complicated rules. A potential concern is that this approach could result in the misidentification of piRNA target sites. However, we note that the misidentification of piRNA target sites would not seriously undermine the validity of the results. First, piRNA-similar sequences include \u003cem\u003ebona fide\u003c/em\u003e piRNA target sites. If piRNA-similar sequences mainly originate from \u003cem\u003eAlu\u003c/em\u003e elements, \u003cem\u003ebona fide\u003c/em\u003e piRNA target sites will also mainly originate from \u003cem\u003eAlu\u003c/em\u003e elements. Similarly, if piRNA-similar sequences are under the influence of IGC, \u003cem\u003ebona fide\u003c/em\u003e piRNA target sites will also be under the influence of IGC. Second, even if some piRNA target sites were misidentified, they would not exert a piRNA-mediated fitness effect on PCGs. Thus, they would not cause a shift in the distribution of allele frequencies (as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). In other words, misidentified piRNA target sites may result in some random errors but cannot cause systematic errors. Moreover, we obtained similar results using a different rule for identifying potential piRNA target sites (Additional file 1: Figure. S6). In addition, the use of a relaxed rule for identifying piRNA target sites can reduce the probability of failing to identify true piRNA target sites. Together, analyzing piRNA-similar sequences is a simple yet prudential strategy for elucidating the evolutionary interplay between \u003cem\u003eAlu\u003c/em\u003e elements and piRNAs.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eIn the human genome, \u003cem\u003eAlu\u003c/em\u003e elements significantly contribute to the sequence similarity shared between piRNAs and PCGs. Consequently, \u003cem\u003eAlu\u003c/em\u003e elements promote the establishment of the targeting relationship between piRNAs and PCGs. Additionally, \u003cem\u003eAlu\u003c/em\u003e elements can induce IGC, which constitutes an evolutionary force that enhances piRNA-PCG targeting relationships. Furthermore, piRNA-PCG targeting relationships impose a piRNA-related selective constraint on human PCGs. Overall, the interplay between \u003cem\u003eAlu\u003c/em\u003e elements and piRNAs is an important factor that participates in the shaping of human PCGs.\u003c/p\u003e"},{"header":"Materials And Methods","content":"\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eData resources\u003c/h2\u003e \u003cp\u003eThe human piRNA data used in this study were published in three previous reports [\u003cspan additionalcitationids=\"CR18\" citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. These data were generated via periodate oxidation treatment followed by small RNA sequencing. The SRA accession numbers of these data were SRR2156539, SRR2156540 [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], SRR835324, SRR950451 [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e], SRR8575349, SRR8575350, SRR8575385, SRR8575386, SRR8575387, SRR8575388, SRR8575409, SRR8575410, SRR8575352, and SRR8575351 [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. SRA files in format were converted to files in FASTQ format by fastq-dump in the SRA Toolkit. Reference human genome (hg38) sequences were obtained from the USCS Genome Browser website. The population genetics data of the high-coverage sequencing version of the 1000 Genomes project were obtained from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/working/20201028_3202_phased/\u003c/span\u003e\u003cspan address=\"http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/working/20201028_3202_phased/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. The population genetics data of the UK10K project were obtained from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.uk10k.org/data.html\u003c/span\u003e\u003cspan address=\"https://www.uk10k.org/data.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. The coordinates of the UK10K data were converted to correspond to the hg38 assembly by using Picard LiftoverVcf. The mRNA sequences and the genomic positions of the 5'UTRs, coding sequences and 3'UTRs of PCGs were obtained from Ensembl BioMart. For each gene, we retained the longest mRNA sequence as the representative transcript for subsequent analyses. The genomic positions of TE-derived sequences were obtained from the RepeatMasker website (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://repeatmasker.org/species/hg.html\u003c/span\u003e\u003cspan address=\"http://repeatmasker.org/species/hg.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePreprocessing Of Pirna Sequencing Data\u003c/h3\u003e\n\u003cp\u003eThe downloaded human piRNA sequencing data were processed with Minion [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e] and Cutadapt [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e] to remove adaptors. Sequences lacking adaptors that had lengths of 25\u0026ndash;32 bp and appeared more than once in the sequencing data were retained for subsequent analyses. We removed redundant piRNA sequences and retained only one sequence for each unique piRNA in subsequent analyses. Sequences that could be mapped to mature microRNAs, microRNA hairpins (downloaded from miRbase [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]), snoRNAs, rRNAs or vtRNAs (downloaded from Ensembl BioMart) were excluded from the analyses.\u003c/p\u003e\n\u003ch3\u003eIdentification Of Pirna-similar Sequences In Protein-coding Genes And Possible Genomic Locations Of Pirnas\u003c/h3\u003e\n\u003cp\u003eTo identify piRNA-similar sequences, we mapped piRNAs to cDNAs using BWA [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e], allowing up to 4 mismatches (set -n 4, in bwa aln). To identify possible genomic locations of piRNAs, piRNA sequences were mapped to the reference genome using BWA, requiring perfect matching (set -n 0, in bwa aln). Because of the short length of piRNAs, the BWA-backtrack algorithm was used when performing BWA mapping (including three steps: bwa index, bwa aln and bwa samse). The mapping results (BAM files) were parsed by Pysam to obtain location information of piRNA-similar sequences in mRNAs and location information of piRNA sequences in the reference genome. The location information of piRNA-similar sequences in mRNAs was transformed into location information in the reference genome according to the corresponding relationship between these two kinds of location information. According to the strand relationship between mRNAs and related piRNAs, piRNA-similar sequences were divided into two categories: sense and antisense. When identifying the genomic locations of piRNAs, there were some piRNAs that could not be mapped to the reference genome. In the following analyses, if information about piRNA genomic locations was needed, piRNAs that could not be mapped to the reference genome were excluded.\u003c/p\u003e\n\u003ch3\u003eIdentification Of Attracting Snvs\u003c/h3\u003e\n\u003cp\u003eInformation on SNVs was parsed from VCF files using cyvcf2 [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. piRNA-similar sequences were identified from reference genomic data. Hence, in the VCF file, the 'reference allele' (the REF item) presented an allelic state identical to that shown in piRNA/piRNA-similar sequence pairs. In contrast, the 'alternate allele' (the ALT item) could alter the sequence similarity between piRNAs and piRNA-similar sequences. For each piRNA/piRNA-similar sequence pair, if an SNV was located in a mismatch site and the ALT item eliminated that mismatch site, we referred to the allele of the ALT item as an attracting SNV. Otherwise, the allele of the ALT item was referred to as a nonattracting SNV (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea).\u003c/p\u003e \u003cp\u003e \u003cb\u003eAnalysis of the\u003c/b\u003e \u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eAlu\u003c/span\u003e \u003cb\u003eassociations of piRNA-similar sequences and their related piRNAs\u003c/b\u003e\u003c/p\u003e \u003cp\u003epiRNA-similar sequences that overlapped with \u003cem\u003eAlu\u003c/em\u003e elements by more than 70% were defined as \u003cem\u003eAlu\u003c/em\u003e-derived piRNA-similar sequences. As some piRNAs could be mapped to multiple genomic locations, as long as a piRNA could be mapped to a genomic location that overlapped with an \u003cem\u003eAlu\u003c/em\u003e element by more than 70%, that piRNA was defined as an \u003cem\u003eAlu\u003c/em\u003e-derived piRNA. We analyzed the proportion of piRNA/piRNA-similar sequence pairs in which the piRNA was \u003cem\u003eAlu\u003c/em\u003e-derived and the proportion of piRNA/piRNA-similar sequence pairs in which the piRNA-similar sequence was \u003cem\u003eAlu\u003c/em\u003e-derived. For comparison, the \u003cem\u003eAlu\u003c/em\u003e-derived fractions of PCGs and piRNAs were also calculated. To further elucidate the contribution of \u003cem\u003eAlu\u003c/em\u003e elements to the sequence similarity shared between piRNAs and PCGs, the following indices were calculated. These indices were calculated for sense and antisense piRNA/piRNA-similar sequence pairs separately (the results for antisense pairs are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, and the results for sense pairs are shown in Additional file 1: Figure. S1).\u003c/p\u003e \u003cp\u003e \u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eP\u003c/span\u003e \u003cb\u003e(piRNA-similar|\u003c/b\u003e \u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eAlu\u003c/span\u003e \u003cb\u003e)\u003c/b\u003e, the total length of intersecting regions of piRNA-similar sequences and \u003cem\u003eAlu\u003c/em\u003e-derived regions (light blue regions in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), divided by the total length of \u003cem\u003eAlu\u003c/em\u003e-derived regions in PCGs (gray box in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). This index indicates the probability that a nucleotide is located in at least one piRNA-similar sequence, given that it is known that the nucleotide is located in an \u003cem\u003eAlu\u003c/em\u003e-derived region in PCGs.\u003c/p\u003e \u003cp\u003e \u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eP\u003c/span\u003e \u003cb\u003e(piRNA-similar|non-\u003c/b\u003e \u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eAlu\u003c/span\u003e \u003cb\u003e)\u003c/b\u003e, the total length of intersecting regions of piRNA-similar sequences and non-\u003cem\u003eAlu\u003c/em\u003e regions (light gray regions in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) divided by the total length of non-\u003cem\u003eAlu\u003c/em\u003e regions in PCGs (gray line in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). This index indicates the probability that a nucleotide is located in at least one piRNA-similar sequence, given that it is known that the nucleotide is located in a non-\u003cem\u003eAlu\u003c/em\u003e region in PCGs. The values of \u003cem\u003eP\u003c/em\u003e(piRNA-similar|\u003cem\u003eAlu\u003c/em\u003e) and \u003cem\u003eP\u003c/em\u003e(piRNA-similar|non-\u003cem\u003eAlu\u003c/em\u003e) were compared.\u003c/p\u003e \u003cp\u003e \u003cb\u003eAverage count of piRNA/piRNA-similar sequence pairs across\u003c/b\u003e \u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eAlu\u003c/span\u003e \u003cb\u003eregions\u003c/b\u003e, the total number of piRNA/piRNA-similar sequence pairs whose piRNA-similar sequences reside in \u003cem\u003eAlu\u003c/em\u003e regions divided by the total length of \u003cem\u003eAlu\u003c/em\u003e regions in PCGs.\u003c/p\u003e \u003cp\u003e \u003cb\u003eAverage count of piRNA/piRNA-similar sequence pairs across non-\u003c/b\u003e \u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eAlu\u003c/span\u003e \u003cb\u003eregions\u003c/b\u003e, the total number of piRNA/piRNA-similar sequence pairs whose piRNA-similar sequences reside in non-\u003cem\u003eAlu\u003c/em\u003e regions divided by the total length of non-\u003cem\u003eAlu\u003c/em\u003e regions in PCGs. The average counts of piRNA/piRNA-similar sequence pairs across \u003cem\u003eAlu\u003c/em\u003e and non-\u003cem\u003eAlu\u003c/em\u003e regions were compared.\u003c/p\u003e\n\u003ch3\u003eAnalysis Of The Igc Footprint\u003c/h3\u003e\n\u003cp\u003eFor each piRNA/piRNA-similar sequence pair, we analyzed SNVs occurring in mismatch sites and the nucleotides paired with them, examining whether they are attracting SNVs. SNVs in four types of data were analyzed separately: 1) sense piRNA-similar sequences, 2) antisense piRNA-similar sequences (shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e), 3) piRNAs related to sense piRNA-similar sequences, and 4) piRNAs related to antisense piRNA-similar sequences (shown in Additional file 1: Figure. S2). Some piRNAs can be mapped to multiple genomic locations. In the analysis of piRNAs related to piRNA-similar sequences, all possible piRNA genomic locations were examined. We counted the occurrences of attracting SNVs and the occurrences of nonattracting SNVs in each piRNA/piRNA-similar sequence pair. Then, we summed the count results for each piRNA/piRNA-similar sequence pair and calculated the proportion of occurrences of attracting SNVs and the proportion of occurrences of nonattracting SNVs. Theoretically, if the occurrence of SNVs is mutually independent in piRNA-similar sequences and their related piRNAs, the proportion of occurrences of attracting SNVs should be approximately 1/3 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). We examined whether the observation significantly deviated from 1/3 through the binomial test (implemented by Scipy.stats.binom_test).\u003c/p\u003e \u003cp\u003eOn the other hand, we explored further evidence from identical sites in piRNA/piRNA-similar sequence pairs. In each piRNA/piRNA-similar sequence pair, if a pair of identical sites both contained SNVs and the ALT items of the two SNV sites were identical, we referred to this SNV pair as a 'shared SNV' (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). For comparison, we built a control dataset by shuffling one-to-one site pairing relationships in piRNA/piRNA-similar sequence pairs. For example, suppose that in a piRNA/piRNA-similar sequence pair, site pairs 1\u0026ndash;2, 4\u0026ndash;15, 17\u0026ndash;23, and 25\u0026ndash;28 are identical. Among these 25 site pairs, there are \u003cem\u003ea\u003c/em\u003e SNV sites in the piRNA sequence and \u003cem\u003eb\u003c/em\u003e SNV sites in the piRNA-similar sequence. If \u003cem\u003eX\u003c/em\u003e of the \u003cem\u003ea\u003c/em\u003e SNV sites in the piRNA sequence and \u003cem\u003eX\u003c/em\u003e of the \u003cem\u003eb\u003c/em\u003e SNV sites in the piRNA-similar sequence are paired and show identical allelic states in their ALT items, the number of shared SNVs will be \u003cem\u003eX\u003c/em\u003e. In the shuffling procedure, the 25 identical sites are randomly paired. Thus, the \u003cem\u003ea\u003c/em\u003e SNV sites in the piRNA sequence and the \u003cem\u003eb\u003c/em\u003e SNV sites in the piRNA-similar sequence will be paired anew. If after shuffling, \u003cem\u003eX'\u003c/em\u003e of the \u003cem\u003ea\u003c/em\u003e SNV sites in the piRNA sequence and \u003cem\u003eX'\u003c/em\u003e of the \u003cem\u003eb\u003c/em\u003e SNV sites in the piRNA-similar sequence are paired and show identical allelic states in their ALT items, then \u003cem\u003eX'\u003c/em\u003e shared SNVs are randomly generated without accounting for any site homology. For each piRNA/piRNA-similar sequence pair, we performed the shuffling procedure described above 30 times. Hence, 30 repeats of \u003cem\u003eX'\u003c/em\u003e were generated, which constituted the control dataset. If the occurrence of shared SNVs is determined by chance, \u003cem\u003eX\u003c/em\u003e should not significantly deviate from \u003cem\u003eX'\u003c/em\u003e. To perform significance testing, the sample mean and the sample standard deviation of the 30 generated \u003cem\u003eX'\u003c/em\u003e values (denoted by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\stackrel{-}{X{\\prime }}\\)\u003c/span\u003e\u003c/span\u003e and \u003cem\u003es\u003c/em\u003e, respectively) were calculated to estimate the mean and the standard deviation of the random distribution of \u003cem\u003eX'\u003c/em\u003e (denoted by \u003cem\u003e\u0026micro;\u003c/em\u003e and \u003cem\u003eσ\u003c/em\u003e, respectively). According to Chebyshev's inequality, \u003cem\u003eP\u003c/em\u003e{|\u003cem\u003eX\u003c/em\u003e \u0026ndash; \u003cem\u003e\u0026micro;\u003c/em\u003e| \u0026ge; \u003cem\u003eε\u003c/em\u003e} \u0026le; \u003cem\u003eσ\u003c/em\u003e\u003csup\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sup\u003e/\u003cem\u003eε\u003c/em\u003e\u003csup\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sup\u003e; when \u003cem\u003eε\u003c/em\u003e\u0026thinsp;=\u0026thinsp;5\u003cem\u003eσ\u003c/em\u003e, \u003cem\u003eP\u003c/em\u003e{\u003cem\u003eX\u003c/em\u003e \u0026ndash; \u003cem\u003e\u0026micro;\u003c/em\u003e\u0026thinsp;\u0026ge;\u0026thinsp;5\u003cem\u003eσ\u003c/em\u003e} \u0026le; 1/25\u0026thinsp;=\u0026thinsp;0.04. Therefore, if the observed number of shared SNVs in the real data (\u003cem\u003eX\u003c/em\u003e) is larger than the mean of the control dataset (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\stackrel{-}{X{\\prime }}\\)\u003c/span\u003e\u003c/span\u003e) by 5 \u0026times; the standard deviation, the observation is considered significantly different from the expectation of random effects, suggesting the influence of IGC exists.\u003c/p\u003e\n\u003ch3\u003eAnalysis Of The Sequence Similarity Between Regions Surrounding Pirnas And Their Pirna-similar Sequences\u003c/h3\u003e\n\u003cp\u003eThe analysis of flanking regions included 30 repeats. Each repeat included 100 piRNA/piRNA-similar sequence pairs that were randomly sampled from the complete set. In the analysis, we randomly assigned one of the possible piRNA locations for each piRNA/piRNA-similar sequence pair. Hence, for each piRNA/piRNA-similar sequence pair, a pair of genomic locations corresponding to the piRNA and the piRNA-similar sequence were obtained. Then, both of these genomic locations were expanded by 150 bp in the upstream and downstream directions.\u003c/p\u003e \u003cp\u003eThe expanded regions were aligned using MAFFT [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. For each pair of expanded regions, the sequence similarity of the most similar block (maximum similarity) was analyzed. Three different block lengths were used in the analysis: 100, 150 and 200 bp. In each repeat, we calculated the median value of maximum similarity across the 100 expanded regions. Corresponding boxplots were also generated. The results are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e and Additional file 1: Figure S3.\u003c/p\u003e\n\u003ch3\u003eAnalysis Of The Footprint Of Natural Selection\u003c/h3\u003e\n\u003cp\u003eTo plot the distribution of allele frequencies, if an SNV could eliminate a mismatch in any piRNA/piRNA-similar sequence pair, it was used to plot the distribution of allele frequencies for attracting SNVs; otherwise, it was used to plot the distribution of allele frequencies for nonattracting SNVs. Additionally, if an SNV was counted as an attracting SNV in any antisense piRNA/piRNA-similar sequence pair, it was used to plot the distribution of allele frequencies for antisense attracting SNVs. Otherwise, it was used to plot the distribution of allele frequencies for sense attracting SNVs. Thus, those SNVs that allowed the similarity of any piRNA/piRNA-similar sequence pair to increase were used to plot the distribution of allele frequencies for attracting SNVs, and those SNVs that had the potential to reinforce any piRNA-PCG targeting relationship were used to plot the distribution of allele frequencies for antisense attracting SNVs. The distribution of allele frequencies for antisense attracting SNVs was compared to the distribution of allele frequencies for sense attracting SNVs. A one-tailed Mann‒Whitney U test was used to examine whether the distribution of allele frequencies for antisense attracting SNVs was significantly skewed toward low frequencies in comparison with that of sense attracting SNVs (implemented by using Scipy.stats.mannwhitneyu, with the setting alternative='greater'). Information about allele frequencies was obtained directly from the 'AF' items in VCF files. For the 1000 Genomes data, variants with unconfident allele frequency estimates were indicated by 'AF\u0026thinsp;=\u0026thinsp;0.0'. Such variants were excluded from our analysis. Furthermore, SNVs were categorized as those located in 5'UTRs and 3'UTR and as synonymous and nonsynonymous variants. The allele frequencies of these categories were further analyzed separately (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e, bottom panels). In addition, the allele frequencies of nonattracting SNVs in sense and antisense pairing relationships were also compared (Additional file 1: Figure. S4).\u003c/p\u003e \u003cp\u003eOn the other hand, we changed the rule for identifying piRNA-similar sequences. A recent study proposed the following piRNA targeting rules (the 'stringent' rules of [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]): a mismatch at the first nucleotide of the piRNA is not counted, up to one GU wobble pair is allowed in the first 2\u0026ndash;7 bp (seed region), and only two mismatches plus an additional GU mismatch are allowed overall. In our analysis, piRNAs were first mapped to mRNAs using BWA, allowing up to 5 mismatches. Segments that satisfied the above-described rule and segments that contained one more mismatch than the above-described rule were identified as piRNA-similar sequences. Allele frequencies in these segments were analyzed as described above to examine whether the footprint of natural selection was robust to the rule for identifying piRNA-similar sequences. The results are shown in Additional file 1: Figure. S6.\u003c/p\u003e \u003cp\u003eAll of the above analyses were implemented by using in-house Python scripts. The Matplotlib [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e], SciPy [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e], NumPy [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e], Pandas [\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e] and seaborn [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e] packages were used for statistical analyses. The scripts have been deposited at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/he-chong/pirna_pcg_evol.git\u003c/span\u003e\u003cspan address=\"https://github.com/he-chong/pirna_pcg_evol.git\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003epiRNA: PIWI-interacting RNA; TE: Transposable element; PCG: Protein-coding gene; IGC: Interlocus gene conversion; SNV: Single nucleotide variant; MEPS: Minimal efficient processing segment; UTR: Untranslated region.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe thank Meng-Yun Chen for the repeated and in-depth discussion.\u003cstrong\u003e \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor\u0026apos;s contribution\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCH and HZ conceived the study and wrote the manuscript. CH carried out the analyses. All authors have read and approved the manuscript for publication.\u003cstrong\u003e \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHZ was supported by National Natural Science Foundation of China (31771456). No funding body played a role in the study design, analysis or interpretation. \u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn-house Python scripts are available in GitHub at https://github.com/he-chong/pirna_pcg_evol.git. \u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable \u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable \u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests. \u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor details\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e1 \u003c/sup\u003eBioinformatics Section, School of Basic Medical Sciences, Southern Medical University, Guangzhou, 510515, China\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e2 \u003c/sup\u003eGuangdong-Hong Kong-Macao Greater Bay Area Center for Brain Science and Brain-Inspired Intelligence, Southern Medical University, Guangzhou, 510515, China\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e3 \u003c/sup\u003eGuangdong Provincial Key Lab of Single Cell Technology and Application, Southern Medical University, Guangzhou, 510515, China\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAravin AA, Hannon GJ, Brennecke J. The Piwi-piRNA pathway provides an adaptive defense in the transposon arms race. Science. 2007;318:761\u0026ndash;4. \u003c/li\u003e\n\u003cli\u003eCzech B, Munaf\u0026ograve; M, Ciabrelli F, Eastwood EL, Fabry MH, Kneuss E, et al. piRNA-guided genome defense: From biogenesis to silencing. Annu Rev Genet. 2018;52:131\u0026ndash;57. \u003c/li\u003e\n\u003cli\u003eWang C, Lin H. Roles of piRNAs in transposon and pseudogene regulation of germline mRNAs and lncRNAs. Genome Biol. 2021;22:1\u0026ndash;21. \u003c/li\u003e\n\u003cli\u003eWatanabe T, Cheng EC, Zhong M, Lin H. Retrotransposons and pseudogenes regulate mRNAs and lncRNAs via the piRNA pathway in the germline. Genome Res. 2015;25:368\u0026ndash;80. \u003c/li\u003e\n\u003cli\u003eGrimson A, Srivastava M, Fahey B, Woodcroft BJ, Chiang HR, King N, et al. Early origins and evolution of microRNAs and Piwi-interacting RNAs in animals. Nature. 2008;455:1193\u0026ndash;7. \u003c/li\u003e\n\u003cli\u003eZhang P, Kang JY, Gou LT, Wang J, Xue Y, Skogerboe G, et al. MIWI and piRNA-mediated cleavage of messenger RNAs in mouse testes. Cell Res. 2015;25:193\u0026ndash;207. \u003c/li\u003e\n\u003cli\u003eZhang D, Tu S, Stubna M, Wu WS, Huang WC, Weng Z, et al. The piRNA targeting rules and the resistance to piRNA silencing in endogenous genes. Science. 2018;359:587\u0026ndash;92. \u003c/li\u003e\n\u003cli\u003eBourque G, Burns KH, Gehring M, Gorbunova V, Seluanov A, Hammell M, et al. Ten things you should know about transposable elements. Genome Biol. 2018;19:199. \u003c/li\u003e\n\u003cli\u003eLi WH, Gu Z, Wang H, Nekrutenko A. Evolutionary analyses of the human genome. Nature. 2001;409:847\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eJangam D, Feschotte C, Betr\u0026aacute;n E. Transposable element domestication as an adaptation to evolutionary conflicts. Trends Genet. 2017;33:817\u0026ndash;31. \u003c/li\u003e\n\u003cli\u003eOphinni Y, Palatini U, Hayashi Y, Parrish NF. piRNA-guided CRISPR-like immunity in eukaryotes. Trends Immunol. 2019;40:998\u0026ndash;1010. \u003c/li\u003e\n\u003cli\u003eBrookfield JFY. The ecology of the genome - Mobile DNA elements and their hosts. Nat Rev Genet. 2005;6:128\u0026ndash;36. \u003c/li\u003e\n\u003cli\u003eNekrutenko A, Li WH. Transposable elements are found in a large number of human protein-coding genes. Trends Genet. 2001;17:619\u0026ndash;21. \u003c/li\u003e\n\u003cli\u003eChuong EB, Elde NC, Feschotte C. Regulatory activities of transposable elements: from conflicts to benefits. Nat Rev Genet. 2017;18:71\u0026ndash;86. \u003c/li\u003e\n\u003cli\u003eThe 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature. 2015;526:68\u0026ndash;74. \u003c/li\u003e\n\u003cli\u003eWalter K, Min JL, Huang J, Crooks L, Memari Y, McCarthy S, et al. The UK10K project identifies rare variants in health and disease. Nature. 2015;526:82\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eWilliams Z, Morozov P, Mihailovic A, Lin C, Puvvula PK, Juranek S, et al. Discovery and characterization of piRNAs in the human fetal ovary. Cell Rep. 2015;13:854\u0026ndash;63. \u003c/li\u003e\n\u003cli\u003eHa H, Song J, Wang S, Kapusta A, Feschotte C, Chen KC, et al. A comprehensive analysis of piRNAs from adult human testis and their relationship with genes and mobile elements. BMC Genomics. 2014;15:1\u0026ndash;16. \u003c/li\u003e\n\u003cli\u003e\u0026Ouml;zata DM, Yu T, Mou H, Gainetdinov I, Colpan C, Cecchini K, et al. Evolutionarily conserved pachytene piRNA loci are highly divergent among modern humans. Nat Ecol Evol. 2020;4:156\u0026ndash;68. \u003c/li\u003e\n\u003cli\u003eMansai SP, Kado T, Innan H. The rate and tract length of gene conversion between duplicated genes. Genes. 2011;2:313\u0026ndash;31. \u003c/li\u003e\n\u003cli\u003eRoman Arguello J, Connallon T. Gene duplication and ectopic gene conversion in Drosophila. Genes. 2011;2:131\u0026ndash;51. \u003c/li\u003e\n\u003cli\u003eHastings PJ. Mechanisms of ectopic gene conversion. Genes. 2010;1:427\u0026ndash;39. \u003c/li\u003e\n\u003cli\u003eFawcett JA, Innan H. The role of gene conversion between transposable elements in rewiring regulatory networks. Genome Biol Evol. 2019;11:1723\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eDumont BL. Interlocus gene conversion explains at least 2.7 % of single nucleotide variants in human segmental duplications. BMC Genomics. 2015;16:1\u0026ndash;11. \u003c/li\u003e\n\u003cli\u003eMansai SP, Innan H. The power of the methods for detecting interlocus gene conversion. Genetics. 2010;184:517\u0026ndash;27. \u003c/li\u003e\n\u003cli\u003eKijima TE, Innan H. On the estimation of the insertion time of LTR retrotransposable elements. Mol Biol Evol. 2010;27:896\u0026ndash;904. \u003c/li\u003e\n\u003cli\u003eLiskay RM, Letsou A, Stachelek JL. Homology requirement for efficient gene conversion between duplicated chromosomal sequences in mammalian cells. Genetics. 1987;115:161\u0026ndash;7. \u003c/li\u003e\n\u003cli\u003eMehta A, Beach A, Haber JE. Homology requirements and competition between gene conversion and break-induced replication during double-strand break repair. Mol Cell. 2017;65:515-526.e3. \u003c/li\u003e\n\u003cli\u003eChen JM, Cooper DN, Chuzhanova N, F\u0026eacute;rec C, Patrinos GP. Gene conversion: Mechanisms, evolution and human disease. Nat Rev Genet. 2007;8:762\u0026ndash;75. \u003c/li\u003e\n\u003cli\u003eM\u0026eacute;zard C, Pompon D, Nicolas A. Recombination between similar but not identical DNA sequences during yeast transformation occurs within short stretches of identity. Cell. 1992;70:659\u0026ndash;70. \u003c/li\u003e\n\u003cli\u003eAhn BY, Dornfeld KJ, Fagrelius TJ, Livingston DM. Effect of limited homology on gene conversion in a Saccharomyces cerevisiae plasmid recombination system. Mol Cell Biol. 1988;8:2442\u0026ndash;8. \u003c/li\u003e\n\u003cli\u003eReiter LT, Hastings PJ, Nelis E, De Jonghe P, Van Broeckhoven C, Lupski JR. Human meiotic recombination products revealed by sequencing a hotspot for homologous strand exchange in multiple HNPP deletion patients. Am J Hum Genet. 1998;62:1023\u0026ndash;33. \u003c/li\u003e\n\u003cli\u003eZhi D. Sequence correlation between neighboring Alu instances suggests post-retrotransposition sequence exchange due to Alu gene conversion. Gene. 2007;390:117\u0026ndash;21. \u003c/li\u003e\n\u003cli\u003eDoronina L, Reising O, Schmitz J. Gene conversion amongst Alu SINE elements. Genes. 2021;12:905. \u003c/li\u003e\n\u003cli\u003eRoy AM, Carroll ML, Nguyen S V., Salem AH, Oldridge M, Wilkie AOM, et al. Potential gene conversion and source genes for recently integrated Alu elements. Genome Res. 2000;10:1485\u0026ndash;95. \u003c/li\u003e\n\u003cli\u003eDeininger P. Alu elements: know the SINEs. Genome Biol. 2011;12:236. \u003c/li\u003e\n\u003cli\u003eCzech B, Hannon GJ. One loop to rule them all: the ping-pong cycle and piRNA-guided silencing. Trends Biochem Sci. 2016;41:324\u0026ndash;37. \u003c/li\u003e\n\u003cli\u003eDavis MPA, van Dongen S, Abreu-Goodger C, Bartonicek N, Enright AJ. Kraken: A set of tools for quality control and analysis of high-throughput sequence data. Methods. 2013;63:41\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eMartin M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet.journal. 2011;17:10. \u003c/li\u003e\n\u003cli\u003eKozomara A, Birgaoanu M, Griffiths-Jones S. miRBase: from microRNA sequences to function. Nucleic Acids Res. 2019;47:D155\u0026ndash;62. \u003c/li\u003e\n\u003cli\u003eLi H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25:1754\u0026ndash;60. \u003c/li\u003e\n\u003cli\u003ePedersen BS, Quinlan AR. cyvcf2: fast, flexible variant analysis with Python. Hancock J, editor. Bioinformatics. 2017;33:1867\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eKatoh K, Misawa K, Kuma K ichi, Miyata T. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res. 2002;30:3059\u0026ndash;66. \u003c/li\u003e\n\u003cli\u003eHunter JD. Matplotlib: A 2D Graphics Environment. Comput Sci Eng. 2007;9:90\u0026ndash;5. \u003c/li\u003e\n\u003cli\u003eVirtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17:261\u0026ndash;72. \u003c/li\u003e\n\u003cli\u003eSt\u0026eacute;fan van der Walt SCC and GV. The NumPy array: a structure for efficient numerical computation. Comput Sci Eng. 2011;13:22\u0026ndash;30. \u003c/li\u003e\n\u003cli\u003eMcKinney W. Data structures for statistical computing in Python. In: Proceedings of the 9th Python in Science Conference. 2010. p. 51\u0026ndash;6. \u003c/li\u003e\n\u003cli\u003eWaskom M. seaborn: statistical data visualization. J Open Source Softw. 2021;6:3021. \u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Alu element, transposable element, piRNA, interlocus gene conversion, natural selection, genome evolution, protein-coding gene","lastPublishedDoi":"10.21203/rs.3.rs-2222130/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2222130/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground: \u003c/strong\u003ePIWI-interacting RNAs (piRNAs) are the most diverse category of small RNAs in animals. Recent evidence suggests that transposable elements (TEs) incorporated into protein-coding genes (PCGs) can be targeted by piRNAs. Thus, TEs might have a piRNA-mediated influence on organisms. In human PCGs, the extent to which TEs contribute to the presence of piRNA target sites remains to be assessed. Moreover, related evolutionary forces remain to be explored.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults: \u003c/strong\u003eWe found that the presence of \u003cem\u003eAlu\u003c/em\u003e elements, a class of primate-specific TEs,\u003cem\u003e \u003c/em\u003ein human PCGs almost always results in potential piRNA target sites. Additionally, we observed that\u003cem\u003e Alu\u003c/em\u003e elements can exert a secondary influence on piRNAs and their potential target sites via interlocus gene conversion (IGC). This mutagenic process can homogenize piRNAs and their potential target sites, resulting in an excess of single nucleotide variants (SNVs) that increase piRNA-PCG targeting affinity in the genome. Although \u003cem\u003eAlu\u003c/em\u003eelements facilitate the occurrence of SNVs that increase piRNA-PCG targeting affinity, these SNVs tend to show low allele frequencies in the human population. This footprint suggests that natural selection opposes the promotion effect of \u003cem\u003eAlu\u003c/em\u003e elements on the formation of piRNA-PCG targeting relationships.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions:\u003c/strong\u003e Human \u003cem\u003eAlu\u003c/em\u003e elements promote both the establishment and enhancement of piRNA-PCG targeting relationships. In addition, piRNA-PCG targeting relationships impose a piRNA-related selective constraint on the evolution of human PCGs. Our work suggests that the interplay between \u003cem\u003eAlu\u003c/em\u003e elements and piRNAs is an important factor that influences the evolutionary trajectory of human PCGs.\u003c/p\u003e","manuscriptTitle":"Human Alu elements promote the establishment and enhancement of piRNA-protein-coding gene targeting relationships","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-11-29 15:31:51","doi":"10.21203/rs.3.rs-2222130/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"340e5a25-bb02-4cb1-b76e-6e535ff7b2d6","owner":[],"postedDate":"November 29th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-02-09T15:35:23+00:00","versionOfRecord":[],"versionCreatedAt":"2022-11-29 15:31:51","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-2222130","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2222130","identity":"rs-2222130","version":["v1"]},"buildId":"ApUGefWb6u5IBVtyqm6d5","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-27T02:00:06.600101+00:00
License: CC-BY-4.0