Unveiling mirtron-triggered gene silencing mechanisms in soybean seed coat pigmentation

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background The diversity of soybean seed coat colors is largely determined by anthocyanin accumulation, regulated by chalcone synthase (CHS) genes at the I locus. While RNA interference (RNAi) has been implicated in this process, the underlying mechanisms remain incompletely understood. This study investigates the role of mirtrons (miRNA precursor derived from intronic sequences) in gene silencing related to seed coat pigmentation in black and yellow soybean varieties. Results To validate the presence and function of these mirtrons, we conducted small RNA sequencing (sRNA-seq) for each allele at the I locus. Our analysis reveals a complex, multi-layered regulatory mechanism involving the genomic architecture the I locus, mRNA, and sRNA interactions. We identified mirtron-triggered miRNAs (MT-miRNAs) and their amplification into secondary miRNAs, which collectively mediate genome-wide silencing of CHS genes. Conclusion This study elucidates a cascade of mirtron-triggered gene silencing (MTGS) that regulates seed coat pigmentation soybeans. These findings provide novel insights into RNAi-mediated control of anthocyanin biosynthesis and highlight the significance of mirtrons in plant gene regulation.
Full text 145,400 characters · extracted from preprint-html · click to expand
Unveiling mirtron-triggered gene silencing mechanisms in soybean seed coat pigmentation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Unveiling mirtron-triggered gene silencing mechanisms in soybean seed coat pigmentation Yoram Choi, Joo-Seok Park, Jin-Hyun Kim, Min-Gyun Jeong, Yeong-Il Jeong, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6817785/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background The diversity of soybean seed coat colors is largely determined by anthocyanin accumulation, regulated by chalcone synthase (CHS) genes at the I locus. While RNA interference (RNAi) has been implicated in this process, the underlying mechanisms remain incompletely understood. This study investigates the role of mirtrons (miRNA precursor derived from intronic sequences) in gene silencing related to seed coat pigmentation in black and yellow soybean varieties. Results To validate the presence and function of these mirtrons, we conducted small RNA sequencing (sRNA-seq) for each allele at the I locus. Our analysis reveals a complex, multi-layered regulatory mechanism involving the genomic architecture the I locus, mRNA, and sRNA interactions. We identified mirtron-triggered miRNAs (MT-miRNAs) and their amplification into secondary miRNAs, which collectively mediate genome-wide silencing of CHS genes. Conclusion This study elucidates a cascade of mirtron-triggered gene silencing (MTGS) that regulates seed coat pigmentation soybeans. These findings provide novel insights into RNAi-mediated control of anthocyanin biosynthesis and highlight the significance of mirtrons in plant gene regulation. chalcone synthase mirtron triggered gene silencing RNA interference seed coat color soybean Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Background Soybeans contain a range of secondary metabolites, such as anthocyanins, lignin, phytoalexins, and flavones, which provide various advantages in interacting with the environment and other organisms, influencing their phenotypes ( 1 ). Among these compounds, anthocyanins and other polyphenolic compounds produced via the phenylpropanoid pathway contribute to stress resistance and seed storage capacity ( 2 ). They particularly impact seed phenotypes by depositing pigments in the seed coat ( 3 , 4 ). The accumulation of anthocyanin in the seed coat through the phenylpropanoid pathway results in a darker coloration, facilitating anthocyanin deposition in the outer layer of soybean seeds ( 5 ). In the flavonoid biosynthesis pathway, chalcone synthase (CHS) is responsible for producing naringenin chalcone from 4-coumaroyl-CoA and 3 malonyl-CoA ( 6 ). Subsequently, downstream enzymes including chalcone isomerase (CHI), flavanone 3-hydroxylase (F3H), flavonoid 3’-hydoxylase (F3’H), dihydroflavonol-4-reductase (DFR), anthocyanidin synthase (ANS), flavonoid 2-O-glycosyltransferase (UF3GT), and others facilitate the deposition of pigments like anthocyanins and proanthocyanins in the epidermal layer ( 7 – 9 ). Until now, eight genetic loci ( I , R , T , W1 , O , D1 , D2 , and qSC1 ) involved in the seed coat pigmentation of soybean have been reported ( 10 – 14 ). Of these, the I , T and W1 loci, which function within the phenylpropanoid pathway, are associated with CHS, F3’H, and F3’5’H, respectively ( 15 ). The I locus, which operates at the highest level of the pathway, exhibits dominance in the following order: I > i i > i k > i . Each allele is associated with specific seed coat and hilum phenotypes: I results in a yellow coat and a yellow or grey hilum; i i leads to a yellow coat and a pigmented hilum; i k produces a partially pigmented coat and a pigmented hilum; and i results in a fully pigmented coat and hilum ( 16 , 17 ). The I locus consists of two clusters of CHS genes: cluster A includes ‘ CHS4 - CHS3 - CHS1 ’, while cluster B includes ‘ CHS1-CHS3 - CHS4 ’ ( 1 ). Research on the I locus has revealed that the formation of a large hairpin loop by inverted repeat sequences in cluster A and cluster B triggers miRNA production, leading to decreased CHS gene expression in yellow soybeans compared to black soybeans ( 18 – 20 ). Analysis of soybean Expressed Sequence Tag (EST) suggested the presence of a small exon between CHS3 and CHS4 in cluster B ( 1 ). Subsequent mRNA read mapping and k -mer analysis of the BAC77G7-a BAC clone indicated the production of chimeric transcripts within cluster B of the I , i i , and i k alleles ( 8 ). Further investigations, using mutations in the DICER-LIKE2 (DCL2) responsible for miRNA generation, substantiated the role of RNA interference (RNAi) in seed coat pigmentation. This research infers that the seed coat phenotype in soybean results from RNAi ( 21 ). However, there is still a lack of experimental evidence confirming the existence of chimeric transcripts originating from the I locus and verifying that RNAi in CHS expression is initiated by these chimeric transcripts. In this study, our primary objective was to uncover and provide experimental evidence for the molecular mechanism previously suggested ( 8 ) to regulate the pigmentation of soybean seed coat. We began by validating the precise positioning of the CHS genes and chimeric transcript using the latest Glycine max reference genome sequence (version 4). Notably, the chimeric transcript exhibited an inverted orientation of CHS1 and CHS3 within its pre-mRNA intron region. The remarkable sequence similarity of 98.54% between these two genes likely facilitate the formation of a hairpin structure during the splicing process of the chimeric transcript. This structure induces the production of intron-derived pre-miRNA, where the primary miRNA targets the entire CHS gene family with high sequence similarity, leading to the generation of secondary miRNAs. Consequently, this process reduces the expression of the CHS gene family, thereby suppressing seed coat pigmentation. We experimentally validated each step of the CHS transcription inhibition process mediated by miRNAs derived from the chimeric transcript at the I locus. To confirm the origin of small RNAs (sRNAs) and their targets, we conducted sRNA sequencing (sRNA-seq) on samples with I, i i , i k , and i alleles. Through comparative analyses of CHS expression levels, we provide comprehensive and robust evidence elucidating the process of seed coat pigmentation based on the allele type at the I locus. Methods Plant materials and growth conditions A total of 25 soybean varieties were examined in this study: twelve black soybean varieties carrying the I allele, eleven yellow soybean varieties with the i i allele, one yellow soybean variety with the I allele, and one saddle-patterned brown soybean variety with the i k allele (Fig. 1 and Supplementary Table 1). All specimens until the seedling stage were cultivated in a controlled plant growth chamber with a photoperiod of 16 hours of light and 8 hours of darkness, at 28℃ during the day and 26℃ at night, maintaining a relative humidity of 60%. Cultivation was carried out from the V1 stage (first trifoliate leaf emergence) to the R6.5 stage (full seed with the onset of coloration) in a glass greenhouse under the same conditions, except for temperature, which was maintained at 32℃ during the day and 28℃ at night. All plants were grown in a mixture of universal potting soil and vermiculite at a 20:1 ratio. DNA/ RNA extraction and cDNA synthesis Plant tissue samples were collected at different developmental stages — VC (unifoliate leaf emergence), V2–V3 (second and third trifoliate leaf emergence), and R6.5 (full seed with the onset of coloration) — across all cultivars, and all samples were collected between 3 PM and 6 PM. Leaf samples were collected during the VC stage, while both leaf and root samples were collected at the V2-V3 stage, and leaf and seed coat samples were obtained at the R6.5 stage. All samples were rapidly frozen in liquid nitrogen and ground into a fine powder. DNA and total RNA were then extracted from all samples, except for seed coat samples, using the yesG™ plant DNA extraction kit (GenesGen, Busan, ROK), and the yesR™ Total RNA plus kit (GenesGen, Busan, ROK), respectively. For seed coat samples with high polysaccharide content, the cetyltrimethylammonium bromide (CTAB) method( 35 ) was employed for DNA extraction, and total RNA was isolated using Plant RNA Purification reagent (Invitrogen, MA, USA). Polysaccharides were removed by treating the samples with 10 M LiCl just before the washing step during the nucleic acid extraction. The quality and integrity of the extracted nucleic acids were assessed by evaluating the 18S/25S rRNA ratio (1:1) via gel electrophoresis and measuring the A260/A280 (1.7–2.1) using a spectrophotometer. Total RNA was subsequently used for cDNA synthesis, following the supplier’s protocol, using SuPrimeScript RT Premix (with oligo dT, 2X) (GenetBio, Nonsan, ROK). PCR amplification Polymerase Chain Reaction (PCR) was conducted to amplify DNA using 10ng of extracted DNA and DNA polymerase (Coregen, Busan, ROK). The PCR protocol began with an initial pre-denaturation at 95°C for 2 minutes, followed by 30 cycles of denaturation at 95°C for 20 seconds, annealing at 58°C for 40 seconds, and elongation at 72°C for 1 minute. A final extension step was carried out at 72°C for 5 minutes. All primers used in this study were designed using a Python script based on the nearest-neighbor thermodynamics approach proposed by SantaLucia ( 36 ) for melting temperature (Tm) calculation. (Supplementary Tables 2 and 3). Small RNA sequencing and mapping Small RNA (sRNA) sequencing was performed on twelve samples, each with three biological replicates, from four distinct soybean varieties: Daeheuk (BL12, allele i ), Hwangkeum (YL12, allele I ), PI475822C (BR01, allele i k ), and Williams82 (YL07, allele i i ). For each sample, 3 µg of total RNA was extracted from seed coats, followed by small RNA selection, library preparation and small RNA sequencing, conducted by Macrogen, Seoul, ROK. The quality of the single-end reads generated was first assessed using FastQC (v0.11.7; https://www.bioinformatics.babraham.ac.uk/progects/fastqc ). Adapter-contaminated and low-quality reads were removed using the Trimmomatic ( 37 ), with the options LEADING:3, TRAILING:3, SLIDINGWINDOW:10, resulting in high-quality reads. These reads were then aligned to the soybean reference genome sequence (Wm82.a4.v1) using the Bowtie aligner ( 38 ) with the options –a, --best, and --strata. sRNA read counts for each cultivar were determined using the HTSeq program ( 39 ), with the options –m union, -s no, -f bam, -r pos, and –t transcript, and the read counts were normalized using the DESeq2 R package (v1.42.1)( 40 ) with the DESeq2 normalization method for comprehensive sample-to-sample comparisons. Quantitative real-time PCR analysis Quantitative real-time PCR (qRT-PCR) was conducted using 10 ng of cDNA as the template for each reaction, with TB Green Premix Ex Taq II (TaKaRa, Shiga, Japan). The PCR procedure included the following steps: Step1, 30 seconds at 95°C; step 2, 5 seconds at 95°C and 20 seconds at 60°C for 40 cycles; step3, 60°C to 95°C with a 0.5°C increment every 5 seconds for melt curve analysis. The delta-delta Ct method ( 41 ) was employed to analyze expression levels, normalizing the expression values with Tubulin beta 3 (TUBB3) as a house keeping gene ( 42 ). All CHS and chimeric transcript-targeting primers used in this study were designed using a Python script based on the nearest-neighbor thermodynamics approach proposed by SantaLucia ( 36 ) for Tm calculation (Supplementary Table 6). Phylogenetic analysis A phylogenetic analysis was conducted to elucidate the relationships among CHS family genes. CHS gene sequences were extracted from the G. max genome sequence and aligned using the MUSCLE algorithm ( 43 ). A neighbor-joining tree was then constructed with 1000 bootstrap replications using MEGA11 ( https://www.megasoftware.net/)(44) . Additionally, to assess sequence similarity within the CHS family, pairwise sequence alignments were performed using BLASTN (v2.12.0+)( 45 ). Sequence similarity analysis of exons and introns in CHS family genes To assess sequence similarity within the CHS gene family, we initially extracted the exon and intron sequences from the soybean reference genome sequence (Wm82.a4.v1). We focused on the sequences corresponding to the first exon, second exon, and intron, which were analyzed using a self-BLASTN search with the blastn-short algorithm to calculate sequence similarity ( 46 ). However, the results from BLASTN, a local alignment tool, were fragmented across the sequences. To address this limitation and obtain an overall measure of sequence similarity, we developed an in-house Python program to concatenate BLAST (available at GitHub: wntjr3364/concatenateBLAST), to concatenate the fragmented alignment results, providing a comprehensive assessment of sequence similarity across the entire length of the sequences. Analysis of sequence variations To investigate the genomic structure of soybean (c.v Hwangkem), we followed a comprehensive analysis pipeline. Initial quality assessment was performed using FastQC for the four resequencing accessions from a public database (ENA, https://www.ebi.ac.uk/ena/browser/home ), followed by the removal of low-quality and adapter sequences with the Trimmomatic program. The paired-end reads were then aligned to the G. max cv. Williams82 assembly version 4.0 (Gmax.v.4.0) using the BWA-MEM algorithm within Burrows-Wheeler Aligner (BWA) with default parameters ( 47 ). To improve mapping accuracy, duplicated reads were removed from the aligned data using Picard's MarkDuplicates tool (v2.22.1) ( https://broadinstitute.github.io/picard ). Statistical analysis All statistical analyses were conducted using the R software (version 4.2.3). One-way and Two-way Analysis of Variance (ANOVA) (α = 0.05) were utilized to identify genes with statistically significant changes in gene expression. Subsequently, the significance of difference between sample means was evaluated using Tukey’s honestly significant difference (HSD) post hoc test (α = 0.05). Results Divergence in genomic structure of the I locus between black and yellow soybean Analysis of black soybean (allele i ) and yellow soybean (allele i i ) varieties (Fig. 1 and Supplementary Table 1) revealed structural differences in the genomic region of the I locus cluster B on chromosome 8, specifically at coordinates 8,499,335-8,499,069 (on the – strand) (Fig. 2 A). Sequences upstream of cluster B encompassed the subtilisin promoter and the coding sequences for four exons of subtilisin. PCR amplification of this region showed the absence of expected band in black soybeans with the i allele. In contrast, this band was present in the yellow soybeans with the i i allele and in BR01 (PI475822C), which exhibits the saddle-patterned brown phenotype with the i k allele. However, YL12 (Hwangkeum), despite being a yellow soybean with the I allele, did not exhibit amplification (Fig. 2 B). Sequencing analysis of the amplified region confirmed a perfect conservation of the translocated segment (Supplementary Fig. 1A). Generation of chimeric transcript through genomic structural variation To gain insights into the origin of the mirtron, a crucial intermediate molecule in CHS gene silencing, the genomic structure at the I locus was investigated. The translocation of the subtilisin promoter and its four exons results in the creation of a new open reading frame (ORF) downstream. This leads to the encoding of the fifth exon located 9,077 base pairs (bp) downstream from the fourth exon, resulting in a novel 10,079 bp chimeric pre-mRNA. Subsequent splicing removes a long intron in this transcript, producing a final 723 bp mature mRNA (Fig. 2 A). Quantification of mature chimeric transcripts was performed using qRT-PCR analysis on leaves at the VC stage for ten black soybean varieties (allele i ), six yellow soybean varieties (allele i i ), and one saddle type soybean variety (allele i k ) (Fig. 2 C). All black soybeans (allele i ) exhibited expression values below the threshold (i.e., not expressed). In contrast, yellow soybean varieties (allele i i ) and saddle type variety (allele i k ) showed detectable expression. Notably, YL08 (Haman) exhibited the lowest expression, while YL04 (Paldal) displayed the highest expression. To confirm the elimination of long intron between the fourth and fifth exons of the chimeric transcript, PCR amplification was performed on cDNA from leaves and roots at the VC, V2-V3, and R6.5 stages. Results demonstrated that the chimeric transcript was amplified in both leaves and roots at all developmental stages in the i i and i k allele types (Fig. 2 D). Sanger sequencing validated the connection between the fourth and fifth exons (Supplementary Fig. 1B). Investigation of total small RNAs using next-generation sequencing analysis To explore the expression patterns of miRNAs targeting CHS genes based on the I locus and to identify miRNAs originating from CHS1 and CHS3 within the intron region of the chimeric transcript, small RNA sequencing was conducted on the following samples: Daeheuk (BL12, allele i ), PI475822C (BR01, allele i k ), Williams82 (YL07, allele i i ), and Hwangkeum (YL12, allele I ). The small RNA sequencing results revealed that all samples exhibited GC contents within the range of 50–60%, Q20, Q30 scores of 90% or higher, and generated approximately 10 million reads or more (Supplementary Figs. 2A and 2B). The reads mapped to the intergenic region in Daeheuk (BL12, allele i ), Hwangkeum (YL12, allele I ), PI475822C (BR01, allele i k ), and Williams82 (YL07, allele i i ) at percentages of 88.66%, 89.90%, 90.93%, and 90%, respectively, and to the genic region at percentages of 11.34%, 10.1%, 9.07%, and 10% (Supplementary Fig. 2C and Supplementary Table 4). PCA analysis resulted in a coverage of 82% in PC1 and PC2, with the samples clustering into three distinct groups (Supplementary Fig. 2D). Considering that small RNAs in plants are predominantly within the range of 20 ~ 24 bp( 22 ), our investigation showed that the size of small RNAs fell within the range of 18 ~ 25 bp for all cultivars. Additionally, GC contents ranged from 50–60%, and all reads exhibited Q20 and Q30 scores surpassing 90% (Supplementary Fig. 3A). Sequencing data were classified separately for reads in the range of 18–25 bp for subsequent analysis. The number of produced reads varied, with a maximum of 2.1 million, a minimum of 0.3 million, and an average of 1.1 million reads (Supplementary Fig. 3B). Among the produced reads, 21 bp reads were the most abundant, accounting for 15.75 % i Daeheuk, 25 bp reads accounted for 18.17 % i Hwangkeum, 24 bp reads accounted for 18.34 % i PI475822C, and 22 bp reads accounted for 15.25 % i Williams82 (Supplementary Fig. 3C). Verification of CHS-targeting small RNA originating from the chimeric transcript The production of small RNAs derived from mirtron and their targeting of CHS genes indicate the mirtron-triggered gene silencing (MTGS) mechanism. To substantiate this mechanism, we mapped small RNAs at the I locus. Cluster B within the I locus, where the chimeric transcript originates, comprises CHS1 , CHS3 , CHS4 and a translocated subtilisin genomic segment. We aimed to confirm the generation of CHS -targeting small RNA, particularly in CHS1 and CHS3 , which share intron sequences within the chimeric transcript, by analyzing small RNA patterns in both exon and intron regions of CHS1 , CHS3 , and CHS4 . To differentiate between initially generated small RNAs and those subsequently amplified during the small RNA generation process, we categorized them into 22 bp small RNAs (primary) and 21 bp small RNAs (secondary). Upon examining the overall small RNA expression pattern in the cluster B region of the I locus, we found that small RNAs containing sequences from CHS1 , CHS3 , and CHS4 were expressed in all soybean varieties except for black soybean (Daeheuk) (Fig. 3 A). The expression levels of chimeric transcript were measured in the seed coats of each allele, revealing expression in YL12 (allele I ), BR01 (allele \(\:{i}^{k}\) ), and YL07 (allele \(\:{i}^{i}\) ), but not in BL12 (allele i ) (Fig. 3 B). Notably, sRNA reads were confirmed in the intron regions of CHS1 and CHS3 , whereas there were very few reads in the intron region of CHS4 (Fig. 3 C). To differentiate between primary and secondary small RNAs, we categorized them into 22 bp and 21 bp and counted the sequencing reads mapped to exon and intron regions (Fig. 3 D). Small RNAs mapping to exon regions were generally expressed in all cultivars except BL12, with significantly higher expression levels for 21 bp small RNAs compared to 22 bp small RNAs. In contrast, small RNA read mapping to the intron regions was significantly lower in quantity. Interestingly, while there were no specific patterns in 21 bp small RNAs, 22 bp small RNAs exhibited a distinctive count of reads specifically in CHS1 and CHS3 . Validation of the influence of small RNAs on CHS gene expression The impact of small RNAs on CHS gene expression was investigated using real-time PCR quantitative analysis. Assessment of the expression levels of small RNAs targeting the entire CHS gene family revealed that BL12 (allele i ) exhibited overall lower expression levels, while YL12 (allele I ), BR01 (allele \(\:{i}^{k}\) ) and YL07 (allele \(\:{i}^{i}\) ) displayed higher expression levels. Specifically, CHS7 and CHS8 showed notably high expression (Fig. 4 A). These small RNAs, functioning as miRNAs, have the potential to repress CHS gene expression due to shared sequence composition. To elucidate this process, CHS transcripts were quantified in soybean seed coats at the R6.5 stage. BL12, BL10, and BL11 with the i allele displayed overall high expression levels, whereas YL12 with the I allele and YL07 with the \(\:{i}^{i}\) allele exhibited very low expression. High expression levels of CHS7/CHS8 were observed in most black soybeans harboring the i allele (Fig. 4 B). Closer examination of the sRNA-seq reads mapped to CHS7 and CHS8 , which showed high expression at both the mRNA and miRNA levels, revealed that they were concentrated in the 2nd exon rather than the 1st exon, with a preference for the 21 bp size over the 22 bp size (Fig. 4 C). At the R6.5 stage, the expression of the CHS gene in leaves mirrored the pattern observed in the seed coat (Fig. 5 A). In the \(\:{i}^{i}\) , I , and \(\:{i}^{k}\) allele cultivars, the expression levels were low, while in i allele cultivars, they showed statistically higher expression. When CHS expression levels were examined in the leaves of BL12 (allele i ), YL07 (allele \(\:{i}^{i}\) ), YL12 (allele I ), and BR01 (allele \(\:{i}^{k}\) ) at the VC, V2–V3, and R6.5 stages, they exhibited similarly low expression in the VC and V2–V3 stages but statistically high expression in the R6.5 stage (Fig. 5 B) Analysis of genomic structure of Hwangkeum using whole genome resequencing data To explore the reasons behind the challenge of confirming the genomic structure and chimeric transcript in Hwangkeum, despite its similar sRNA and CHS transcript expression patterns to Williams82 and PI475822C, we conducted a structural variation analysis using whole genome resequencing data. Four whole-genome resequencing datasets for the Hwangkeum (YL12) variety were obtained from the European Nucleotide Archive ( https://www.ebi.ac.uk/ena/browser/home ) and aligned to the Williams82 reference genome sequence (Supplementary Fig. 7). The resequencing reads showed good alignment with the reference genome sequence. In the I locus region (chromosome 8; 8,498,281–8,527,803), the Hwangkeum sequence closely resembled that of Williams82, though several single mismatches were identified in specific sequences. These mismatches were primarily concentrated in the region between cluster A and B. Discussion Our study aimed to uncover the mechanisms contributing to phenotypic variations in soybean seed coat color, focusing on the differences between black (allele i ) and yellow (allele I, i i , i k ) soybeans. Inspired by a comparative analysis between Glycine soja (W05, black wild soybean) and G. max (Williams82, yellow soybean) ( 23 ), we investigated the genomic variation in the I locus cluster B region. Our research confirmed the expression of chimeric transcripts and identified primary miRNAs derived from these transcripts. Our findings demonstrated that the repression of CHS gene expression accounts for the differences in seed coat color between black and yellow soybeans (Fig. 7 ). To validate these results, we analyzed twelve cultivars carrying the i allele, eleven with i i allele, one with I allele, and one with i k allele (Fig. 1 and Supplementary Table 1). This comprehensive analysis provided strong evidence linking the observed phenotypic variations to specific allelic differences. Initially, the I locus was identified to harbor an inverted repeat of cluster A and B through a genomic comparison of G. soja W05 and G. max Williams82. Upstream of the cluster B region, sequences include the subtilisin promoter and the coding sequences for four exons of subtilisin. Previous studies ( 8 , 23 ) reported genomic structural changes leading to the translocation of this region. Through PCR amplification and Sanger sequencing, we confirmed that this region remains intact in the \(\:{i}^{i}\) and \(\:{i}^{k}\) alleles (Figs. 2 A, 2 B, and Supplementary Fig. 1A). The translocation of the subtilisin promoter and its four exons caused a rearrangement, leading to the creation of a new open reading frame (ORF). Consequently, a previously nonexistent 10,079 bp chimeric pre-mRNA is formed, with the final fifth exon encoded 9,077 bp downstream from the fourth exon. Secondly, the long 9,077 bp intron within this chimeric transcript is spliced to produce a fully mature 723 bp mRNA (Fig. 2 A). This transcript was verified through PCR amplification spanning the first and fifth exons (Fig. 2 D). Expression of of the chimeric transcript was detected in VC stage leaves, V2–V3 stage leaves and roots, and R6.5 stage leaves. Notably, in VC stage leaves, it was expressed in cultivars carrying both the i i and i k alleles (Fig. 2 C), supporting the transcription of a 10,079 bp pre-mRNA driven by the relocated subtilisin promoter on chromosome 8. Sequencing confirmed the excision of the 9,077 bp intron during splicing (Supplementary Fig. 1B). The splicing mechanism involves the 2’ OH group of an adenine within the branch point sequence (BPS) attacking the phosphate backbone at the 5’ splice site, leading to the formation of a lariat structure (24–26). This process brings CHS1 and CHS3 , which are embedded in an inverted orientation within the intron of chimeric transcript and share 98.54% sequence similarity, into close proximity. Their alignment likely promotes the formation of a complementary base-pairing region, resulting in a hairpin structure known as a mirtron (a short intron-derived pre-miRNA)( 27 – 29 ). This structural feature suggests the potential for mirtron formation, serving as a precursor to miRNA through subsequent processing. To investigate miRNA derived from the intron of chimeric transcript, we conducted small RNA (sRNA) sequencing across different alleles at the I locus, including Daeheuk (BL12, allele i ), Hwangkeum (YL12, allele I ), Williams82 (YL07, allele i i ), PI475822C (BR01, allele i k ) (Fig. 3 A and Supplementary Fig. 2). While the CHS gene family exhibits high sequence similarity in exon regions, except for CHS13 and CHS14 (Supplementary Fig. 4), their introns are more divergent. Notably, only CHS1 and CHS3 share similar intronic sequences (Figs. 6 A, 6 B, and 6 C). This unique feature enables CHS1 and CHS3 , which are positioned in an inverted orientation within the intron of chimeric transcripts from cluster B of the I locus, to potentially form hairpin structures during splicing. These hairpins are processed into primary miRNAs containing CHS1 and CHS3 sequences. As a result, the primary miRNAs can target other CHS genes with high sequence similarity, triggering the production of secondary miRNAs. Analysis of the overall sRNA expression pattern in the cluster B region of the I locus, along with the expression of chimeric transcripts responsible for primary miRNA plroduction, revealed that CHS1 , CHS3 , and CHS4 —including their mapped sRNAs and associated mRNA transcripts— are expressed in all soybean varieties except the black soybean (Daeheuk, BL12) (Figs. 3 A and 3 B). Notably, miRNAs expressed in soybean seed coats predominantly exist in a 22-nt primary form, while the secondary form is 21-nt in length ( 20 ). To confirm that CHS1 and CHS3 are the sources of primary miRNAs targeting CHS genes, we closely examined their intron regions, along with that of CHS4. This analysis revealed that 22-nt sRNA reads specifically mapped to the introns of CHS1 and CHS3 , but not in CHS4 (Fig. 3 C). These 22-nt sRNAs (primary miRNAs) are clearly mapped to CHS1 and CHS3 , whereas other CHS genes show much lower mapping (Fig. 3 D). Interestingly, 21-nt sRNAs (secondary miRNAs) are distributed more evenly across all CHS genes, suggesting that the amplification of primary miRNAs into secondary miRNAs is limited by sequence divergence (Figs. 3 D and 6 C). In contrast, the exon regions of i i (YL07) and i k (BR01) alleles exhibited higher levels of 21-nt sRNAs compared to 22-nt sRNAs, indicating that extensive amplification of secondary miRNAs occurs due to high exon sequence similarity among CHS genes. Notably, the amplification of secondary miRNAs from primary miRNAs varies across CHS genes, ranging from a 2.37-fold increase ( CHS1 ) to a 10.48-fold increase ( CHS10 ), with Williams82 used as the reference (Fig. 3 D). This observed pattern is attributable to differences in the intronic sequences of CHS1/3 and CHS4 , despite substantial sequence similarity in the exon regions. sRNA read mapping further supports this observation, as reads aligned specifically to regions with high sequence similarity to CHS1/3 (Fig. 6 ). For CHS2 , CHS4 , CHS5 , CHS6 , CHS9 , CHS10 , and CHS11 , which share homology with CHS1 / 3 primarily in the first and second exons, sRNA reads were mapped exclusively to those exonic regions. Similarly, for CHS7 / 8 , which shares high similarity only in the second exon, reads were confined to that region. These findings support the conclusion that CHS -targeting primary miRNAs originate within the intron of the chimeric transcript, explaining the widespread presence of both primary miRNAs and secondary miRNAs in all examined varieties except black soybeans (Fig. 3 A). The amplification of primary miRNAs into secondary miRNAs is particularly notable in CHS7 and CHS8 . sRNA-seq reads mapping to CHS genes show that these reads are nearly absent in black soybeans, whereas I , i i , and i k genotypes display similar expression patterns (Fig. 4 A), with CHS7 and CHS8 exhibiting significantly higher sRNA expression levels. In the seed coat of black soybeans, where CHS gene silencing does not occur, mRNA levels of CHS7 and CHS8 are notably elevated (Fig. 4 B). Further analysis of sRNA-seq reads mapped to CHS7 and CHS8 , categorized by length, reveals that both 22-nt and 21-nt sRNAs are concentrated in the second exon, which shares a high sequence similarity with CHS1 and CHS3 (Fig. 4 C). Notably, 21-nt reads are more abundant than 22-nt reads, indicating that CHS7 and CHS8 transcripts are initially targeted by 22-nt primary miRNAs, which are subsequently amplified into 21-nt secondary miRNAs. This provides clear evidence of mirtron-triggered miRNA amplification process. This pattern of 22-nt primary miRNAs giving rise to 21-nt secondary miRNAs is consistently observed across the exons of highly sequence-similar CHS1-11 genes (Fig. 6 and Supplementary Fig. 4). Ultimately, primary miRNAs derived from CHS1 and CHS3 initiate a robust amplification of secondary miRNAs by targeting other CHS transcripts with high sequence similarity. This mechanism affects a broad spectrum of CHS genes, including three CHS1 s, one CHS2 , four CHS3 s, and one each of CHS5 , CHS6 , CHS7 , CHS8 , CHS9 , CHS10 , and CHS11 , extending the silencing effect beyond chromosome 8 to the entire genome (Supplementary Fig. 5 and Supplementary Table 5). The abundant production of these primary and secondary miRNAs constitutes a potent post-transcriptional regulatory mechanism that controls CHS gene expression across the genome (Fig. 5 A). This regulatory silencing is not restricted to the seed coat but also occurs in leaves and roots. Notably, the suppression of CHS gene is most pronounced during the reproductive growth stage — characterized by rapid increase in CHS gene — and continues through seed maturation. In contrast, minimal silencing is observed during early developmental or vegetative stages (Fig. 5 B). The robust production of primary and secondary miRNAs via this mechanism broadly suppresses the expression of CHS genes across diverse soybean cultivars, including Hwangkeum ( I ), PI475822C ( i k ), and Williams82 ( i i ) (Figs. 4 B and 5 A). This suppression effectively inhibits CHS expression, thereby preventing the accumulation of anthocyanin and proanthocyanin in the seed coat ( 8 , 16 , 30 , 31 ). While specific data on CHS13 and CHS14 was not presented, these genes were excluded from all experimental analyses due to their consistently low expression levels observed across multiple developmental stages, including VC stage leaves, V2–V3 stage leaves and roots, and R6.5 stage leaves and seed coats of Williams82. Their exclusion was further supported by their low sequence similarity to the broader CHS gene family (Supplementary Fig. 4). sRNA-seq read mapping to CHS13 and CHS14 revealed only minimal read counts, irrespective of the I locus allele or sRNA size. Notably, sequences mapped to the intron region of CHS13 were identified as simple ‘TA’ sequence repeats (Supplementary Fig. 6), suggesting that these reads likely represent background sRNA commonly found throughout the soybean genome, coincidently aligning to a repetitive region rather than originating from the I locus. The i k allele shares structural features at the I locus with the i i and I alleles but is associated with the k1 mutant, which disrupts the function of Argonaute5, which is a key component of RNA interference (RNAi) pathway. This deficiency in Argonaute5 impairs the silencing of CHS genes specifically around the hilum, leading to a characteristic saddle-shaped pigmentation pattern on the seed coat. Notably, in the hilum-adjacent region where pigment accumulates, the expression levels of CHS7 and CHS8 are approximately 26-fold higher than in more distal regions ( 32 – 34 ). However, in areas where Argonaut5 remains functional, CHS gene silencing proceeds similarly to that observed in the i i allele background (Figs. 2 – 5 ). For the I allele (Hwangkeum, YL12), PCR amplification targeting the chimeric transcript region and translocated sequences failed to yield detectable products (Fig. 2 ). This discrepancy is likely due to the primers being designed based solely on the Williams82 reference genome, which may not align effectively with the genome sequence of Hwangkeum (YL12, allele I ) (Supplementary Tables 2 and 3). As a result, no PCR amplification was observed for regions with low-abundance transcripts (Figs. 2 B and 2 D). Despite the PCR failure, sRNA-seq analysis revealed read alignment patterns consistent with those observed in Williams82 ( i i ), including CHS targeting and the production of 22-nt primary miRNAs and 21-nt secondary miRNAs. These findings suggest that sequence mismatches, rather than major structural differences, are responsible for the failed PCR amplification and altered read alignment. Resequencing data from Hwangkeum ( I ) confirmed that reads predominantly mapped to the I locus (Supplementary Fig. 7), although single nucleotide mismatches were frequent, underscoring the limited specificity of the William82-based primers. Nevertheless, the presence of both primary and secondary miRNAs (Figs. 3 and 4 ) indicates that the core silencing mechanism is conserved. Consequently, suppression of seed coat pigmentation in Hwangkeum resembles the results observed in Williams82 ( i i ) and PI475822C ( i k ) is observed. Conclusions This study presents a comprehensive analysis of genomic variations associated with alleles at the I locus, confirming the expression of the chimeric transcript arising from these structural differences. Using small RNA sequencing, we identified the generation of primary miRNAs originating from the intronic regions of CHS1 and CHS3 and demonstrated the extensive production of secondary miRNAs targeting the broader CHS gene family. Our findings elucidate the functional role of mirtrons derived from chimeric transcript, resolving a longstanding question in the field regarding the RNA interference (RNAi) mechanism at the I locus. Furthermore, by quantifying CHS gene expression across multiple developmental stages and tissues in cultivars carrying different I locus alleles, we provide strong evidence supporting the mechanism underlying seed coat pigmentation suppression in soybeans. This step-by-step experimental validation integrates previously fragmented insights, leading to the establishment of the mirtron-triggered gene silencing (MTGS) model (Fig. 7 ). Declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Availability of data and materials The datasets generated in this study are available in the NCBI BioProject repository under accession number PRJNA1112103 (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1112103). Competing interests The authors declare that they have no competing interests. Funding This research was supported by Basic Science Research Program through the National Research Foundation of Korea (NRF) funded by the Ministry of Education (grant number RS-2020-NR049596), Biomaterials Specialized Graduate Program through the Korea Environmental Industry & Technology Institute (KEITI) funded by the Ministry of Environment (MOE). Authors’ contributions YC: conceptualization, investigation, experiment, and writing – original draft; J-SP: data curation and analysis; J-HK: conceptualization and analysis; M-GJ: data curation and validation; Y-IJ: conceptualization and data curation; H-KC: conceptualization, funding acquisition, supervision, project administration, writing – original draft, and writing – review & editing. All authors read and approved the final manuscript. Acknowledgments Not applicable. References Clough SJ, Tuteja JH, Li M, Marek LF, Shoemaker RC, Vodkin LO. Features of a 103-kb gene-rich region in soybean include an inverted perfect repeat cluster of CHS genes comprising the I locus. Genome. 2004;47(5):819–31. Jiang L, Yang X, Gao X, Yang H, Ma S, Huang S, et al. Multiomics Analyses Reveal the Dual Role of Flavonoids in Pigmentation and Abiotic Stress Tolerance of Soybean Seeds. J Agric Food Chem. 2024;72(6):3231–43. Wu K, Xiao S, Chen Q, Wang Q, Zhang Y, Li K, et al. Changes in the Activity and Transcription of Antioxidant Enzymes in Response to Al Stress in Black Soybeans. Plant Mol Biol Rep. 2013;31(1):141–50. ZHANG T, KAWABATA K, KITANO R. Preventive Effects of Black Soybean Seed Coat Polyphenols against DNA Damage in Salmonella typhimurium. Food Sci Technol Res. 2013;19(4):685–90. Wu K, Xiao S, Chen Q, Wang Q, Zhang Y, Li K, et al. Changes in the Activity and Transcription of Antioxidant Enzymes in Response to Al Stress in Black Soybeans. Plant Mol Biol Rep. 2013;31(1):141–50. Zabala G, Zou J, Tuteja J, Gonzalez DO, Clough SJ, Vodkin LO. Transcriptome changes in the phenylpropanoid pathway of Glycine max in response to Pseudomonas syringaeinfection. BMC Plant Biol. 2006;6(1):26. BRUGLIERA TANAKAY, KALC F, SENIOR G, NAKAMURA MDYSONB. Flower Color Modification by Engineering of the Flavonoid Biosynthetic Pathway: Practical Perspectives. Biosci Biotechnol Biochem. 2010;74(9):1760–9. Kim JH, Park JS, Lee CY, Jeong MG, Xu JL, Choi Y et al. Dissecting seed pigmentation-associated genomic loci and genes by employing dual approaches of reference-based and k-mer-based GWAS with 438 Glycine accessions. Rajcan I, editor. PLoS One. 2020;15(12):e0243085. Kim JM, Lee JW, Seo JS, Ha BK, Kwon SJ. Differentially Expressed Genes Related to Isoflavone Biosynthesis in a Soybean Mutant Revealed by a Comparative Transcriptomic Analysis. Plants. 2024;13(5):584. Kohzuma K, Sato Y, Ito H, Okuzaki A, Watanabe M, Kobayashi H, et al. The non-mendelian green cotyledon gene in soybean encodes a small subunit of photosystem II. Plant Physiol. 2017;173(4):2138–47. Senda M, Kasai A, Yumoto S, Akada S, Ishikawa R, Harada T, et al. Sequence divergence at chalcone synthase gene in pigmented seed coat soybean mutants of the Inhibitor locus. Genes Genet Syst. 2002;77(5):341–50. Senda M, Jumonji A, Yumoto S, Ishikawa R, Harada T, Niizeki M, et al. Analysis of the duplicated CHS1 gene related to the suppression of the seed coat pigmentation in yellow soybeans. Theor Appl Genet. 2002;104(6–7):1086–91. Senda M, Yamaguchi N, Hiraoka M, Kawada S, Iiyoshi R, Yamashita K, et al. Accumulation of proanthocyanidins and/or lignin deposition in buff-pigmented soybean seed coats may lead to frequent defective cracking. Planta. 2017;245(3):659–70. Yuan B, Yuan C, Wang Y, Liu X, Qi G, Wang Y, et al. Identification of genetic loci conferring seed coat color based on a high-density map in soybean. Front Plant Sci. 2022;13(August):1–11. Zabala G, Vodkin LO. A rearrangement resulting in small tandem repeats in the F3′5′H gene of white flower genotypes is associated with the soybean W1 locus. Crop Sci. 2007;47(SUPPL 2):113–24. Todd JJ, Vodkin LO. Pigmented Soybean. Plant Physiol. 1993;102:663–70. Senda M, Masuta C, Ohnishi S, Goto K, Kasai A, Sano T, et al. Patterning of virus-infected Glycine max seed coat is associated with suppression of endogenous silencing of chalcone synthase genes. Plant Cell. 2004;16(4):807–18. Tuteja JH, Zabala G, Varala K, Hudson M, Vodkin LO, Endogenous. Tissue-Specific Short Interfering RNAs Silence the Chalcone Synthase Gene Family in Glycine max Seed Coats. Plant Cell. 2009;21(10):3063–77. Tuteja JH, Clough SJ, Chan WC, Vodkin LO. Tissue-Specific Gene Silencing Mediated by a Naturally Occurring Chalcone Synthase Gene Cluster in Glycine max [W]. Plant Cell. 2004;16(4):819–35. Cho YB, Jones SI, Vodkin L. The Transition from Primary siRNAs to Amplified Secondary siRNAs That Regulate Chalcone Synthase During Development of Glycine max Seed Coats. Freitag M, editor. PLoS One. 2013;8(10):e76954. Jia J, Ji R, Li Z, Yu Y, Nakano M, Long Y, et al. Soybean DICER-LIKE2 Regulates Seed Coat Color via Production of Primary 22-Nucleotide Small Interfering RNAs from Long Inverted Repeats. Plant Cell. 2020;32(12):3662–73. Gao Z, Nie J, Wang H. MicroRNA biogenesis in plant. Plant Growth Regul. 2021;93(1):1–12. Xie M, Chung CYL, Li MW, Wong FL, Wang X, Liu A, et al. A reference-grade wild soybean genome. Nat Commun. 2019;10(1):1–12. Drift R, Velocity R. Encyclopedia of Astrobiology. Encyclopedia of Astrobiology. 2011. Fica SM, Tuttle N, Novak T, Li NS, Lu J, Koodathingal P, et al. RNA catalyses nuclear pre-mRNA splicing. Nature. 2013;503(7475):229–34. Hang J, Wan R, Yan C. Structural basis of pre-mRNA splicing. 2015;(August):4–7. Curtis HJ, Sibley CR, Wood MJA. Mirtrons, an emerging class of atypical miRNA. Wiley Interdiscip Rev RNA. 2012;3(5):617–32. Flynt AS, Greimann JC, Chung WJ, Lima CD, Lai EC. MicroRNA Biogenesis via Splicing and Exosome-Mediated Trimming in Drosophila. Mol Cell. 2010;38(6):900–7. Okamura K, Hagen JW, Duan H, Tyler DM, Lai EC. The Mirtron Pathway Generates microRNA-Class Regulatory RNAs in Drosophila. Cell. 2007;130(1):89–100. Zabala G, Vodkin L. Cloning of the Pleiotropic T Locus in Soybean and Two Recessive Alleles That Differentially Affect Structure and Expression of the Encoded Flavonoid 3′ Hydroxylase. Genetics. 2003;163(1):295–309. Zabala G, Zou J, Tuteja J, Gonzalez DO, Clough SJ, Vodkin LO. Transcriptome changes in the phenylpropanoid pathway of Glycine max in response to Pseudomonas syringae infection. BMC Plant Biol. 2006;6(1):26. Palmer RG, Pfeiffer TW, Buss GR, Kilen TC. Qualitative Genetics. In: Soybeans: Improvement, Production, and Uses. 2016. pp. 137–233. Cho YB, Jones SI, Vodkin LO. Mutations in Argonaute5 Illuminate Epistatic Interactions of the K1 and I Loci Leading to Saddle Seed Color Patterns in Glycine max. Plant Cell. 2017;29(4):708–25. Song J, Liu Z, Hong H, Ma Y, Tian L, Li X, et al. Identification and validation of loci governing seed coat color by combining association mapping and bulk segregation analysis in soybean. PLoS ONE. 2016;11(7):1–19. Stefanova P, Taseva M, Georgieva T, Gotcheva V, Angelov A. A Modified CTAB Method for DNA Extraction from Soybean and Meat Products. Biotechnol Biotechnol Equip. 2013;27(3):3803–10. SantaLucia J, Allawi HT, Seneviratne PA. Improved Nearest-Neighbor Parameters for Predicting DNA Duplex Stability. Biochemistry. 1996;35(11):3555–62. Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014;30(15):2114–20. Langmead B, Trapnell C, Pop M, Salzberg SL. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol. 2009;10(3):R25. Anders S, Pyl PT, Huber W. HTSeq–a Python framework to work with high-throughput sequencing data. Bioinformatics. 2015;31(2):166–9. Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15(12):550. Livak KJ, Schmittgen TD. Analysis of Relative Gene Expression Data Using Real-Time Quantitative PCR and the 2 – ∆∆CT Method. Methods. 2001;25(4):402–8. Gao M, Liu Y, Ma X, Shuai Q, Gai J, Li Y. Evaluation of reference genes for normalization of gene expression using quantitative RT-PCR under aluminum, cadmium, and heat stresses in soybean. PLoS ONE. 2017;12(1):1–15. Edgar RC. MUSCLE: Multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004;32(5):1792–7. Tamura K, Stecher G, Kumar S. MEGA11: Molecular Evolutionary Genetics Analysis Version 11. Battistuzzi FU, editor. Mol Biol Evol. 2021;38(7):3022–7. Camacho C, Coulouris G, Avagyan V, Ma N, Papadopoulos J, Bealer K, et al. BLAST+: architecture and applications. BMC Bioinformatics. 2009;10(1):421. Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990;215(3):403–10. Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25(14):1754–60. Additional Declarations No competing interests reported. Supplementary Files SupplementaryFigure1.pdf SupplementaryFigure2.pdf SupplementaryFigure3.pdf SupplementaryFigure4.pdf SupplementaryFigure5.pdf SupplementaryFigure6.pdf SupplementaryFigure7.pdf SupplementaryTable1.xlsx SupplementaryTable2.xlsx SupplementaryTable3.xlsx SupplementaryTable4.xlsx SupplementaryTable5.xlsx SupplementaryTable6.xlsx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6817785","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":476578538,"identity":"b686efdd-3fef-4a7f-a769-86ed4a45422e","order_by":0,"name":"Yoram Choi","email":"","orcid":"","institution":"Dong-A University","correspondingAuthor":false,"prefix":"","firstName":"Yoram","middleName":"","lastName":"Choi","suffix":""},{"id":476578539,"identity":"8d15574e-619d-45ba-bb22-bde91467420e","order_by":1,"name":"Joo-Seok Park","email":"","orcid":"","institution":"Gyeongsang National University","correspondingAuthor":false,"prefix":"","firstName":"Joo-Seok","middleName":"","lastName":"Park","suffix":""},{"id":476578540,"identity":"2425db0a-7798-47e5-92d4-d18cdcc42233","order_by":2,"name":"Jin-Hyun Kim","email":"","orcid":"","institution":"National Institute of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Jin-Hyun","middleName":"","lastName":"Kim","suffix":""},{"id":476578541,"identity":"9757d945-a293-4528-9481-48bb96c544b4","order_by":3,"name":"Min-Gyun Jeong","email":"","orcid":"","institution":"Dong-A University","correspondingAuthor":false,"prefix":"","firstName":"Min-Gyun","middleName":"","lastName":"Jeong","suffix":""},{"id":476578542,"identity":"c2e6efa5-6026-4c3a-82c4-cbacedd2c9e9","order_by":4,"name":"Yeong-Il Jeong","email":"","orcid":"","institution":"Dong-A University","correspondingAuthor":false,"prefix":"","firstName":"Yeong-Il","middleName":"","lastName":"Jeong","suffix":""},{"id":476578543,"identity":"bf7eba22-6875-4867-9bc6-0f30e62a2470","order_by":5,"name":"Hong-Kyu Choi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAtElEQVRIiWNgGAWjYBACAx4GNiBpw9gA4ScQrSWNZC0Mh0nQYs5z+NmDDwXnZftnJDB++MGQlk9Qi2Vvm7nhDIPbxjNuJDBL9jDkWDYQdNh5HjZpHoPbiQ03EhikGRgqDAjaAtVyLnE+0JbfxGk52wPSciBxw40ENqAtOURoOXPMTHKGQbLxxjMP2yx7DNKI0ZL8TOLDHzvZeceTD9/4UZFMWAsSAEUNSRpGwSgYBaNgFOAEAJXKOSvX/7rxAAAAAElFTkSuQmCC","orcid":"","institution":"Dong-A University","correspondingAuthor":true,"prefix":"","firstName":"Hong-Kyu","middleName":"","lastName":"Choi","suffix":""}],"badges":[],"createdAt":"2025-06-04 07:53:37","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6817785/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6817785/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":85564875,"identity":"fd234046-3684-46bf-bb26-23157715a04a","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":382212,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePhenotypes of soybean seeds associated with the \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003elocus.\u003c/strong\u003e Yellow-green, yellow, red, and brown boxes indicate soybeans with the \u003cem\u003ei\u003c/em\u003e, i\u003csup\u003ei \u003c/sup\u003e, \u003cem\u003eI\u003c/em\u003e, i\u003csup\u003ek\u003c/sup\u003e \u0026nbsp;genotypes, respectively. Scale bar: 10mm\u003c/p\u003e","description":"","filename":"Figure15.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/9d03eae6b70c0978ff82c5c3.jpg"},{"id":85564876,"identity":"58171dce-36bb-4847-b36b-380116eb42e0","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":279681,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGenomic structure of the \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e locus and expression of chimeric transcripts in relation to soybean seed coat color. \u003c/strong\u003e(A) Schematic representation of the genomic structure of the \u003cem\u003eI\u003c/em\u003e locus cluster B in non-pigmented soybean. Green, blue, and red boxes on the black line represent exons of \u003cem\u003eSubtilisin\u003c/em\u003e, \u003cem\u003eCHS3\u003c/em\u003e, and \u003cem\u003eCHS1\u003c/em\u003e, respectively. The light green box with a red arrow marks the boundary of translocated region. mRNA transcripts are shown as boxes placed outside the black line. (B) PCR amplification at the translocation site (red arrow in panel a) based on allele in the \u003cem\u003eI\u003c/em\u003e locus. The amplified bands are approximately 267 base pairs (bp) in size. (C) Quantification of chimeric transcript based on the \u003cem\u003eI\u003c/em\u003elocus allele in leaves at VC stage with three biological replicates. Data was normalized with reference to the YL02 value. Results are presented are as mean \u0026nbsp;± \u0026nbsp;s.d., with different letters indicating significant differences according to ANOVA Tukey’s test (\u003cem\u003ep\u003c/em\u003e \u0026lt; 0.05, \u003cem\u003ef\u003c/em\u003e = 8.286, \u003cem\u003edf\u003c/em\u003e = 6, \u003cem\u003en\u003c/em\u003e = 3). Error bars that do not share letters indicate significant differences. (D) PCR amplification of chimeric transcript in leaves and roots at various developmental stages. The amplified bands are approximately 613 bp in size. CHS\u003cem\u003e: \u003c/em\u003echalcone synthase\u003cem\u003e,\u003c/em\u003e BL: black-colored soybean, YL: yellow-colored soybean, BR: brown-colored soybean with a saddle pattern, VC: vegetative stage with unifoliate leaves, V2–V3: vegetative stage with two to three trifoliate leaves, R6.5: reproductive stage with maturing of full seeds, and NTC: non-template control.\u003c/p\u003e","description":"","filename":"Figure22.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/8046ca3e8fa56e036e82b78a.jpg"},{"id":85566006,"identity":"bab79fc3-958d-49e6-bc36-cc18ea20c94c","added_by":"auto","created_at":"2025-06-27 14:35:30","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":743583,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExpression analysis of small RNA and chimeric transcripts based on the \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e locus allele. \u003c/strong\u003eSmall RNA sequencing reads were mapped to CHS genes, with each sample analyzed in triplicate. (A) Small RNA sequencing reads mapped to cluster B of the \u003cem\u003eI\u003c/em\u003elocus for alleles \u003cem\u003ei, i\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e\u003cem\u003e, i\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e\u003cem\u003e, L\u003c/em\u003e and \u003cem\u003eI\u003c/em\u003e. The scale represents 150 reads (height of each lane). The genomic structure of \u003cem\u003eCHS1, CHS3, CHS4\u003c/em\u003e, and the chimeric transcript were depicted below the lanes. (B) Quantification of chimeric transcript based on the \u003cem\u003eI\u003c/em\u003e locus allele in seed coats at R6.5 stage. Data were normalized to the YL12 value and presented as mean \u0026nbsp;± \u0026nbsp;s.d. ANOVA Tukey’s test were conducted with three biological and three experimental replicates. n.s. indicates ‘not significant’. (C) Small RNA sequencing reads mapped to intron regions of\u003cem\u003e CHS1, CHS3\u003c/em\u003e, and \u003cem\u003eCHS4\u003c/em\u003e with a scale corresponding to 10 reads (height of each lane). (D) Classification of small RNA sequencing reads mapped to exon and intron regions of the CHS gene family into 22 and 21 bp categories. CHS: chalcone synthase, BL: black-colored soybean, YL: yellow-colored soybean, BR: brown-colored soybean with a saddle pattern, and R6.5: reproductive stage with maturing full seeds.\u003c/p\u003e","description":"","filename":"Figure3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/b817c25982a4e3e5114184df.jpg"},{"id":85565718,"identity":"2370a8e0-8f23-48fd-88d7-79d6271c8c9c","added_by":"auto","created_at":"2025-06-27 14:27:30","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":843437,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExpression patterns of \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eCHS7/CHS8\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e small RNA and mRNA. \u003c/strong\u003e(A) Total Small RNA sequencing reads mapped to \u003cem\u003eGmCHS\u003c/em\u003ein different cultivars: Daeheuk (BL12), Hwangkeum (YL12), PI475822C (BR01), and Williams82 (YL07), corresponding to the \u003cem\u003ei\u003c/em\u003e, \u003cem\u003eI\u003c/em\u003e, i\u003csup\u003ek\u003c/sup\u003e, and i\u003csup\u003ei\u003c/sup\u003e \u0026nbsp;alleles, respectively. (B) Quantification of \u003cem\u003eGmCHS\u003c/em\u003e transcripts in seed coats at the R6.5 stage based on the \u003cem\u003eI\u003c/em\u003e locus allele. Data are normalized to the BL12 \u003cem\u003eCHS1/3\u003c/em\u003e expression value and are presented as mean \u0026nbsp;± \u0026nbsp;s.d., with three biological and three experimental replicates. (C) Small RNA sequencing read mapping plots for \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003e, categorized by the \u003cem\u003eI\u003c/em\u003elocus allele and read length. The scale represents 250 reads (height of each lane). The features under the lanes indicate \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003etranscripts. CHS: chalcone synthase, BL: black-colored soybean, YL: yellow-colored soybean, and BR: brown-colored soybean with a saddle pattern.\u003c/p\u003e","description":"","filename":"Figure4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/f7105a68122e2196ffa220d0.jpg"},{"id":85564891,"identity":"29bdb03c-fcb6-40e2-841a-49c394e6ad61","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":473350,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCHS gene expression in soybean across different tissues and developmental stages. \u003c/strong\u003e(A) Quantification of \u003cem\u003eCHS\u003c/em\u003e expression in various cultivars with different alleles of the \u003cem\u003eI\u003c/em\u003e locus in leaves at R6.5 stage. Expression values are compared to the BL12 \u003cem\u003eCHS1/3\u003c/em\u003e expression level. (B) Quantification of \u003cem\u003eCHS\u003c/em\u003e expression across developmental stages in soybean cultivars with different \u003cem\u003eI\u003c/em\u003e locus alleles. Expression values are compared to the BL12 VC \u003cem\u003eCHS1/3\u003c/em\u003e expression level. The data are presented as mean ± s.d with three biological and three experimental replicates. CHS: chalcone synthase, BL: black-colored soybean, YL: yellow-colored soybean, BR: brown-colored soybean with a saddle pattern, VC: unifoliate leaf emergence, V2–V3: second and third trifoliate leaf emergence, and R6.5: full seed stage with the onset of coloration.\u003c/p\u003e","description":"","filename":"Figure51.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/08f1cd7ead2ad95b85e1ad8a.jpg"},{"id":85564894,"identity":"42f5833b-3891-49f2-b7d2-64d64e8c8dd6","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":571710,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSequence similarity analysis and small RNA sequencing read mapping patterns of GmCHS genes.\u003c/strong\u003e Sequence alignment was performed among the \u003cem\u003eCHS\u003c/em\u003e genes, with alignment coverages represented by pie charts where the brightness corresponds to the alignment value. (A) Sequence alignment within the first exon. (B) Sequence alignment within the second exon. (C) Sequence alignment within the intron regions. (D) Mapping pattern of small RNA sequencing reads across GmCHS genes.\u003c/p\u003e","description":"","filename":"Figure61.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/254651810d4139837a929428.jpg"},{"id":85565719,"identity":"34ce8488-c67f-4ea5-bf0a-885416fd3910","added_by":"auto","created_at":"2025-06-27 14:27:30","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":553491,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSchematic illustration of the mirtron-triggered gene silencing (MTGS) mechanism in soybeans. \u0026nbsp;\u003c/strong\u003eThe diagram depicts the biogenesis of MT-miRNA and secondary miRNA, highlighting the sequential events of the MTGS mechanism. This process begins with the transcription of a chimeric transcript from the \u003cem\u003eI\u003c/em\u003e locus cluster B region of yellow soybean, which differs from that in black soybean. This transcript facilitates the generation of mirtrons, leading to the production of a large quantity of MT-miRNAs (primary miRNA) and their subsequent amplification into secondary-miRNAs. The genomic structures of black soybean (\u003cem\u003eG. soja,\u003c/em\u003e c.v. W05) and yellow soybean (\u003cem\u003eG. max\u003c/em\u003e, c.v. Williams82) are depicted, with exons and introns of each \u003cem\u003eCHS\u003c/em\u003e gene and the chimeric transcript indicated by dark and light segments, respectively. CHS: chalcone synthase, MT-miRNA: mirtron-triggered miRNA.\u003c/p\u003e","description":"","filename":"Figure7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/67bfa896cb9d6a998ae5253b.jpg"},{"id":87035745,"identity":"02c78b33-fb09-490a-88dc-0840733f8e75","added_by":"auto","created_at":"2025-07-18 13:17:09","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":4835461,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/45cbe0b7-8946-4108-90f5-94532808795b.pdf"},{"id":85564878,"identity":"1dbfdf6a-deaa-4739-bc54-efed88ca080a","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":442837,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigure1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/6ca9d07a8f4f048fa2c9fea8.pdf"},{"id":85566005,"identity":"72f957f8-1941-446c-b2e6-4ec9e435dba7","added_by":"auto","created_at":"2025-06-27 14:35:30","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":290350,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigure2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/b3892376832d2367e9d9a5dc.pdf"},{"id":85566004,"identity":"4c47d2ff-f027-42e1-9b57-709e6ab2cad0","added_by":"auto","created_at":"2025-06-27 14:35:30","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":242962,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigure3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/54b5f89a2b2318a8cc2042ec.pdf"},{"id":85564882,"identity":"7ede0494-d7a0-40cb-a805-fb9dd738a896","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":82818,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigure4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/2161c9914087b0c309e1dd60.pdf"},{"id":85564895,"identity":"e1f3028e-9353-4b34-8ab5-b8e1429f4584","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":169955,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigure5.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/08a16c0168660bedd6c9aa7a.pdf"},{"id":85564896,"identity":"0bbbfc58-26f3-4c39-8f4a-2ec4d27216ee","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":142487,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigure6.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/9ee707c5a4515675b0a3ae28.pdf"},{"id":85564892,"identity":"06e7f2fb-a6ea-4208-8923-f25cdcdcf291","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"pdf","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":438256,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigure7.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/0634f5f17bbe369958c0875d.pdf"},{"id":85565720,"identity":"8dbd7099-e90f-4a51-a0a9-85bbb9da70aa","added_by":"auto","created_at":"2025-06-27 14:27:30","extension":"xlsx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":14547,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/3693e41bfadf31c2e6659166.xlsx"},{"id":85564900,"identity":"f128daed-56af-46fd-b5cf-8176776d9d92","added_by":"auto","created_at":"2025-06-27 14:19:31","extension":"xlsx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":9154,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/74459fb1311c6f1c6c3be8d3.xlsx"},{"id":85565717,"identity":"49052c30-c6d2-458c-9388-95039a3ffe2a","added_by":"auto","created_at":"2025-06-27 14:27:30","extension":"xlsx","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":12102,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable3.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/d62a8768f3816e6534efee1e.xlsx"},{"id":85564914,"identity":"52bcc3f2-7976-4226-8cbc-338ec56bf2dd","added_by":"auto","created_at":"2025-06-27 14:19:31","extension":"xlsx","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":14772,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable4.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/74c59842c9678427e4d1d26d.xlsx"},{"id":85564899,"identity":"01cde5f2-1a17-4bb9-8641-b3dbcb8cb8f2","added_by":"auto","created_at":"2025-06-27 14:19:30","extension":"xlsx","order_by":11,"title":"","display":"","copyAsset":false,"role":"supplement","size":12424,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable5.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/ffaf520b8a3ea7c3f9b25ba5.xlsx"},{"id":85566648,"identity":"f6fbd071-0f1a-4e7e-990a-b3017bd5f8f1","added_by":"auto","created_at":"2025-06-27 14:43:31","extension":"xlsx","order_by":12,"title":"","display":"","copyAsset":false,"role":"supplement","size":12966,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable6.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6817785/v1/d1ddd2c5f6ee2dc99fc07097.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Unveiling mirtron-triggered gene silencing mechanisms in soybean seed coat pigmentation","fulltext":[{"header":"Background","content":"\u003cp\u003eSoybeans contain a range of secondary metabolites, such as anthocyanins, lignin, phytoalexins, and flavones, which provide various advantages in interacting with the environment and other organisms, influencing their phenotypes (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). Among these compounds, anthocyanins and other polyphenolic compounds produced via the phenylpropanoid pathway contribute to stress resistance and seed storage capacity (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). They particularly impact seed phenotypes by depositing pigments in the seed coat (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). The accumulation of anthocyanin in the seed coat through the phenylpropanoid pathway results in a darker coloration, facilitating anthocyanin deposition in the outer layer of soybean seeds (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn the flavonoid biosynthesis pathway, chalcone synthase (CHS) is responsible for producing naringenin chalcone from 4-coumaroyl-CoA and 3 malonyl-CoA (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e). Subsequently, downstream enzymes including chalcone isomerase (CHI), flavanone 3-hydroxylase (F3H), flavonoid 3\u0026rsquo;-hydoxylase (F3\u0026rsquo;H), dihydroflavonol-4-reductase (DFR), anthocyanidin synthase (ANS), flavonoid 2-O-glycosyltransferase (UF3GT), and others facilitate the deposition of pigments like anthocyanins and proanthocyanins in the epidermal layer (\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eUntil now, eight genetic loci (\u003cem\u003eI\u003c/em\u003e, \u003cem\u003eR\u003c/em\u003e, \u003cem\u003eT\u003c/em\u003e, \u003cem\u003eW1\u003c/em\u003e, \u003cem\u003eO\u003c/em\u003e, \u003cem\u003eD1\u003c/em\u003e, \u003cem\u003eD2\u003c/em\u003e, and \u003cem\u003eqSC1\u003c/em\u003e) involved in the seed coat pigmentation of soybean have been reported (\u003cspan additionalcitationids=\"CR11 CR12 CR13\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e). Of these, the \u003cem\u003eI\u003c/em\u003e, \u003cem\u003eT\u003c/em\u003e and \u003cem\u003eW1\u003c/em\u003e loci, which function within the phenylpropanoid pathway, are associated with CHS, F3\u0026rsquo;H, and F3\u0026rsquo;5\u0026rsquo;H, respectively (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). The \u003cem\u003eI\u003c/em\u003e locus, which operates at the highest level of the pathway, exhibits dominance in the following order: \u003cem\u003eI\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;\u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e\u0026gt;\u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e\u0026gt;\u003cem\u003ei\u003c/em\u003e. Each allele is associated with specific seed coat and hilum phenotypes: \u003cem\u003eI\u003c/em\u003e results in a yellow coat and a yellow or grey hilum; \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e leads to a yellow coat and a pigmented hilum; \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e produces a partially pigmented coat and a pigmented hilum; and \u003cem\u003ei\u003c/em\u003e results in a fully pigmented coat and hilum (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e). The \u003cem\u003eI\u003c/em\u003e locus consists of two clusters of CHS genes: cluster A includes \u0026lsquo;\u003cem\u003eCHS4\u003c/em\u003e-\u003cem\u003eCHS3\u003c/em\u003e-\u003cem\u003eCHS1\u003c/em\u003e\u0026rsquo;, while cluster B includes \u0026lsquo;\u003cem\u003eCHS1-CHS3\u003c/em\u003e-\u003cem\u003eCHS4\u003c/em\u003e\u0026rsquo; (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eResearch on the \u003cem\u003eI\u003c/em\u003e locus has revealed that the formation of a large hairpin loop by inverted repeat sequences in cluster A and cluster B triggers miRNA production, leading to decreased CHS gene expression in yellow soybeans compared to black soybeans (\u003cspan additionalcitationids=\"CR19\" citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). Analysis of soybean Expressed Sequence Tag (EST) suggested the presence of a small exon between \u003cem\u003eCHS3\u003c/em\u003e and \u003cem\u003eCHS4\u003c/em\u003e in cluster B (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). Subsequent mRNA read mapping and \u003cem\u003ek\u003c/em\u003e-mer analysis of the BAC77G7-a BAC clone indicated the production of chimeric transcripts within cluster B of the \u003cem\u003eI\u003c/em\u003e, \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e, and \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e alleles (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). Further investigations, using mutations in the DICER-LIKE2 (DCL2) responsible for miRNA generation, substantiated the role of RNA interference (RNAi) in seed coat pigmentation. This research infers that the seed coat phenotype in soybean results from RNAi (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). However, there is still a lack of experimental evidence confirming the existence of chimeric transcripts originating from the \u003cem\u003eI\u003c/em\u003e locus and verifying that RNAi in \u003cem\u003eCHS\u003c/em\u003e expression is initiated by these chimeric transcripts.\u003c/p\u003e \u003cp\u003eIn this study, our primary objective was to uncover and provide experimental evidence for the molecular mechanism previously suggested (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e) to regulate the pigmentation of soybean seed coat. We began by validating the precise positioning of the CHS genes and chimeric transcript using the latest \u003cem\u003eGlycine max\u003c/em\u003e reference genome sequence (version 4). Notably, the chimeric transcript exhibited an inverted orientation of \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e within its pre-mRNA intron region. The remarkable sequence similarity of 98.54% between these two genes likely facilitate the formation of a hairpin structure during the splicing process of the chimeric transcript. This structure induces the production of intron-derived pre-miRNA, where the primary miRNA targets the entire CHS gene family with high sequence similarity, leading to the generation of secondary miRNAs. Consequently, this process reduces the expression of the CHS gene family, thereby suppressing seed coat pigmentation. We experimentally validated each step of the \u003cem\u003eCHS\u003c/em\u003e transcription inhibition process mediated by miRNAs derived from the chimeric transcript at the \u003cem\u003eI\u003c/em\u003e locus. To confirm the origin of small RNAs (sRNAs) and their targets, we conducted sRNA sequencing (sRNA-seq) on samples with \u003cem\u003eI, i\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e, \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e, and \u003cem\u003ei\u003c/em\u003e alleles. Through comparative analyses of \u003cem\u003eCHS\u003c/em\u003e expression levels, we provide comprehensive and robust evidence elucidating the process of seed coat pigmentation based on the allele type at the \u003cem\u003eI\u003c/em\u003e locus.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003ePlant materials and growth conditions\u003c/p\u003e \u003cp\u003eA total of 25 soybean varieties were examined in this study: twelve black soybean varieties carrying the \u003cem\u003eI\u003c/em\u003e allele, eleven yellow soybean varieties with the \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e allele, one yellow soybean variety with the \u003cem\u003eI\u003c/em\u003e allele, and one saddle-patterned brown soybean variety with the \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e allele (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Supplementary Table\u0026nbsp;1). All specimens until the seedling stage were cultivated in a controlled plant growth chamber with a photoperiod of 16 hours of light and 8 hours of darkness, at 28℃ during the day and 26℃ at night, maintaining a relative humidity of 60%. Cultivation was carried out from the V1 stage (first trifoliate leaf emergence) to the R6.5 stage (full seed with the onset of coloration) in a glass greenhouse under the same conditions, except for temperature, which was maintained at 32℃ during the day and 28℃ at night. All plants were grown in a mixture of universal potting soil and vermiculite at a 20:1 ratio.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eDNA/ RNA extraction and cDNA synthesis\u003c/p\u003e \u003cp\u003ePlant tissue samples were collected at different developmental stages \u0026mdash; VC (unifoliate leaf emergence), V2\u0026ndash;V3 (second and third trifoliate leaf emergence), and R6.5 (full seed with the onset of coloration) \u0026mdash; across all cultivars, and all samples were collected between 3 PM and 6 PM. Leaf samples were collected during the VC stage, while both leaf and root samples were collected at the V2-V3 stage, and leaf and seed coat samples were obtained at the R6.5 stage. All samples were rapidly frozen in liquid nitrogen and ground into a fine powder. DNA and total RNA were then extracted from all samples, except for seed coat samples, using the yesG\u0026trade; plant DNA extraction kit (GenesGen, Busan, ROK), and the yesR\u0026trade; Total RNA plus kit (GenesGen, Busan, ROK), respectively. For seed coat samples with high polysaccharide content, the cetyltrimethylammonium bromide (CTAB) method(\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e) was employed for DNA extraction, and total RNA was isolated using Plant RNA Purification reagent (Invitrogen, MA, USA). Polysaccharides were removed by treating the samples with 10 M LiCl just before the washing step during the nucleic acid extraction. The quality and integrity of the extracted nucleic acids were assessed by evaluating the 18S/25S rRNA ratio (1:1) via gel electrophoresis and measuring the A260/A280 (1.7\u0026ndash;2.1) using a spectrophotometer. Total RNA was subsequently used for cDNA synthesis, following the supplier\u0026rsquo;s protocol, using SuPrimeScript RT Premix (with oligo dT, 2X) (GenetBio, Nonsan, ROK).\u003c/p\u003e \u003cp\u003ePCR amplification\u003c/p\u003e \u003cp\u003ePolymerase Chain Reaction (PCR) was conducted to amplify DNA using 10ng of extracted DNA and DNA polymerase (Coregen, Busan, ROK). The PCR protocol began with an initial pre-denaturation at 95\u0026deg;C for 2 minutes, followed by 30 cycles of denaturation at 95\u0026deg;C for 20 seconds, annealing at 58\u0026deg;C for 40 seconds, and elongation at 72\u0026deg;C for 1 minute. A final extension step was carried out at 72\u0026deg;C for 5 minutes. All primers used in this study were designed using a Python script based on the nearest-neighbor thermodynamics approach proposed by SantaLucia (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e) for melting temperature (Tm) calculation. (Supplementary Tables\u0026nbsp;2 and 3).\u003c/p\u003e \u003cp\u003eSmall RNA sequencing and mapping\u003c/p\u003e \u003cp\u003eSmall RNA (sRNA) sequencing was performed on twelve samples, each with three biological replicates, from four distinct soybean varieties: Daeheuk (BL12, allele \u003cem\u003ei\u003c/em\u003e), Hwangkeum (YL12, allele \u003cem\u003eI\u003c/em\u003e), PI475822C (BR01, allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e), and Williams82 (YL07, allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e). For each sample, 3 \u0026micro;g of total RNA was extracted from seed coats, followed by small RNA selection, library preparation and small RNA sequencing, conducted by Macrogen, Seoul, ROK. The quality of the single-end reads generated was first assessed using FastQC (v0.11.7; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.bioinformatics.babraham.ac.uk/progects/fastqc\u003c/span\u003e\u003cspan address=\"https://www.bioinformatics.babraham.ac.uk/progects/fastqc\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Adapter-contaminated and low-quality reads were removed using the Trimmomatic (\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e), with the options LEADING:3, TRAILING:3, SLIDINGWINDOW:10, resulting in high-quality reads. These reads were then aligned to the soybean reference genome sequence (Wm82.a4.v1) using the Bowtie aligner (\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e) with the options \u0026ndash;a, --best, and --strata. sRNA read counts for each cultivar were determined using the HTSeq program (\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e), with the options \u0026ndash;m union, -s no, -f bam, -r pos, and \u0026ndash;t transcript, and the read counts were normalized using the DESeq2 R package (v1.42.1)(\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e) with the DESeq2 normalization method for comprehensive sample-to-sample comparisons.\u003c/p\u003e \u003cp\u003eQuantitative real-time PCR analysis\u003c/p\u003e \u003cp\u003eQuantitative real-time PCR (qRT-PCR) was conducted using 10 ng of cDNA as the template for each reaction, with TB Green Premix Ex Taq II (TaKaRa, Shiga, Japan). The PCR procedure included the following steps: Step1, 30 seconds at 95\u0026deg;C; step 2, 5 seconds at 95\u0026deg;C and 20 seconds at 60\u0026deg;C for 40 cycles; step3, 60\u0026deg;C to 95\u0026deg;C with a 0.5\u0026deg;C increment every 5 seconds for melt curve analysis. The delta-delta Ct method (\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e) was employed to analyze expression levels, normalizing the expression values with Tubulin beta 3 (TUBB3) as a house keeping gene (\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e). All \u003cem\u003eCHS\u003c/em\u003e and chimeric transcript-targeting primers used in this study were designed using a Python script based on the nearest-neighbor thermodynamics approach proposed by SantaLucia (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e) for Tm calculation (Supplementary Table\u0026nbsp;6).\u003c/p\u003e \u003cp\u003ePhylogenetic analysis\u003c/p\u003e \u003cp\u003eA phylogenetic analysis was conducted to elucidate the relationships among CHS family genes. CHS gene sequences were extracted from the \u003cem\u003eG. max\u003c/em\u003e genome sequence and aligned using the MUSCLE algorithm (\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e). A neighbor-joining tree was then constructed with 1000 bootstrap replications using MEGA11 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.megasoftware.net/)(44)\u003c/span\u003e\u003cspan address=\"https://www.megasoftware.net/)(44)\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Additionally, to assess sequence similarity within the \u003cem\u003eCHS\u003c/em\u003e family, pairwise sequence alignments were performed using BLASTN (v2.12.0+)(\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eSequence similarity analysis of exons and introns in CHS family genes\u003c/p\u003e \u003cp\u003eTo assess sequence similarity within the CHS gene family, we initially extracted the exon and intron sequences from the soybean reference genome sequence (Wm82.a4.v1). We focused on the sequences corresponding to the first exon, second exon, and intron, which were analyzed using a self-BLASTN search with the blastn-short algorithm to calculate sequence similarity (\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e). However, the results from BLASTN, a local alignment tool, were fragmented across the sequences. To address this limitation and obtain an overall measure of sequence similarity, we developed an in-house Python program to concatenate BLAST (available at GitHub: wntjr3364/concatenateBLAST), to concatenate the fragmented alignment results, providing a comprehensive assessment of sequence similarity across the entire length of the sequences.\u003c/p\u003e \u003cp\u003eAnalysis of sequence variations\u003c/p\u003e \u003cp\u003eTo investigate the genomic structure of soybean (c.v Hwangkem), we followed a comprehensive analysis pipeline. Initial quality assessment was performed using FastQC for the four resequencing accessions from a public database (ENA, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ebi.ac.uk/ena/browser/home\u003c/span\u003e\u003cspan address=\"https://www.ebi.ac.uk/ena/browser/home\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), followed by the removal of low-quality and adapter sequences with the Trimmomatic program. The paired-end reads were then aligned to the \u003cem\u003eG. max\u003c/em\u003e cv. Williams82 assembly version 4.0 (Gmax.v.4.0) using the BWA-MEM algorithm within Burrows-Wheeler Aligner (BWA) with default parameters (\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e). To improve mapping accuracy, duplicated reads were removed from the aligned data using Picard's MarkDuplicates tool (v2.22.1) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://broadinstitute.github.io/picard\u003c/span\u003e\u003cspan address=\"https://broadinstitute.github.io/picard\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analysis\u003c/h2\u003e \u003cp\u003eAll statistical analyses were conducted using the R software (version 4.2.3). One-way and Two-way Analysis of Variance (ANOVA) (α\u0026thinsp;=\u0026thinsp;0.05) were utilized to identify genes with statistically significant changes in gene expression. Subsequently, the significance of difference between sample means was evaluated using Tukey\u0026rsquo;s honestly significant difference (HSD) post hoc test (α\u0026thinsp;=\u0026thinsp;0.05).\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eDivergence in genomic structure of the I locus between black and yellow soybean\u003c/p\u003e \u003cp\u003eAnalysis of black soybean (allele \u003cem\u003ei\u003c/em\u003e) and yellow soybean (allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e) varieties (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Supplementary Table\u0026nbsp;1) revealed structural differences in the genomic region of the \u003cem\u003eI\u003c/em\u003e locus cluster B on chromosome 8, specifically at coordinates 8,499,335-8,499,069 (on the \u0026ndash; strand) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). Sequences upstream of cluster B encompassed the subtilisin promoter and the coding sequences for four exons of subtilisin. PCR amplification of this region showed the absence of expected band in black soybeans with the \u003cem\u003ei\u003c/em\u003e allele. In contrast, this band was present in the yellow soybeans with the \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e allele and in BR01 (PI475822C), which exhibits the saddle-patterned brown phenotype with the \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e allele. However, YL12 (Hwangkeum), despite being a yellow soybean with the \u003cem\u003eI\u003c/em\u003e allele, did not exhibit amplification (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB). Sequencing analysis of the amplified region confirmed a perfect conservation of the translocated segment (Supplementary Fig.\u0026nbsp;1A).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eGeneration of chimeric transcript through genomic structural variation\u003c/p\u003e \u003cp\u003eTo gain insights into the origin of the mirtron, a crucial intermediate molecule in CHS gene silencing, the genomic structure at the \u003cem\u003eI\u003c/em\u003e locus was investigated. The translocation of the subtilisin promoter and its four exons results in the creation of a new open reading frame (ORF) downstream. This leads to the encoding of the fifth exon located 9,077 base pairs (bp) downstream from the fourth exon, resulting in a novel 10,079 bp chimeric pre-mRNA. Subsequent splicing removes a long intron in this transcript, producing a final 723 bp mature mRNA (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003eQuantification of mature chimeric transcripts was performed using qRT-PCR analysis on leaves at the VC stage for ten black soybean varieties (allele \u003cem\u003ei\u003c/em\u003e), six yellow soybean varieties (allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e), and one saddle type soybean variety (allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC). All black soybeans (allele \u003cem\u003ei\u003c/em\u003e) exhibited expression values below the threshold (i.e., not expressed). In contrast, yellow soybean varieties (allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e) and saddle type variety (allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e) showed detectable expression. Notably, YL08 (Haman) exhibited the lowest expression, while YL04 (Paldal) displayed the highest expression.\u003c/p\u003e \u003cp\u003eTo confirm the elimination of long intron between the fourth and fifth exons of the chimeric transcript, PCR amplification was performed on cDNA from leaves and roots at the VC, V2-V3, and R6.5 stages. Results demonstrated that the chimeric transcript was amplified in both leaves and roots at all developmental stages in the \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e and \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e allele types (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD). Sanger sequencing validated the connection between the fourth and fifth exons (Supplementary Fig.\u0026nbsp;1B).\u003c/p\u003e \u003cp\u003eInvestigation of total small RNAs using next-generation sequencing analysis\u003c/p\u003e \u003cp\u003eTo explore the expression patterns of miRNAs targeting CHS genes based on the \u003cem\u003eI\u003c/em\u003e locus and to identify miRNAs originating from \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e within the intron region of the chimeric transcript, small RNA sequencing was conducted on the following samples: Daeheuk (BL12, allele \u003cem\u003ei\u003c/em\u003e), PI475822C (BR01, allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e), Williams82 (YL07, allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e), and Hwangkeum (YL12, allele \u003cem\u003eI\u003c/em\u003e). The small RNA sequencing results revealed that all samples exhibited GC contents within the range of 50\u0026ndash;60%, Q20, Q30 scores of 90% or higher, and generated approximately 10\u0026nbsp;million reads or more (Supplementary Figs.\u0026nbsp;2A and 2B). The reads mapped to the intergenic region in Daeheuk (BL12, allele \u003cem\u003ei\u003c/em\u003e), Hwangkeum (YL12, allele \u003cem\u003eI\u003c/em\u003e), PI475822C (BR01, allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e), and Williams82 (YL07, allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e) at percentages of 88.66%, 89.90%, 90.93%, and 90%, respectively, and to the genic region at percentages of 11.34%, 10.1%, 9.07%, and 10% (Supplementary Fig.\u0026nbsp;2C and Supplementary Table\u0026nbsp;4). PCA analysis resulted in a coverage of 82% in PC1 and PC2, with the samples clustering into three distinct groups (Supplementary Fig.\u0026nbsp;2D).\u003c/p\u003e \u003cp\u003eConsidering that small RNAs in plants are predominantly within the range of 20\u0026thinsp;~\u0026thinsp;24 bp(\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e), our investigation showed that the size of small RNAs fell within the range of 18\u0026thinsp;~\u0026thinsp;25 bp for all cultivars. Additionally, GC contents ranged from 50\u0026ndash;60%, and all reads exhibited Q20 and Q30 scores surpassing 90% (Supplementary Fig.\u0026nbsp;3A). Sequencing data were classified separately for reads in the range of 18\u0026ndash;25 bp for subsequent analysis. The number of produced reads varied, with a maximum of 2.1\u0026nbsp;million, a minimum of 0.3\u0026nbsp;million, and an average of 1.1\u0026nbsp;million reads (Supplementary Fig.\u0026nbsp;3B). Among the produced reads, 21 bp reads were the most abundant, accounting for 15.75 % i Daeheuk, 25 bp reads accounted for 18.17 % i Hwangkeum, 24 bp reads accounted for 18.34 % i PI475822C, and 22 bp reads accounted for 15.25 % i Williams82 (Supplementary Fig.\u0026nbsp;3C).\u003c/p\u003e \u003cp\u003eVerification of CHS-targeting small RNA originating from the chimeric transcript\u003c/p\u003e \u003cp\u003eThe production of small RNAs derived from mirtron and their targeting of CHS genes indicate the mirtron-triggered gene silencing (MTGS) mechanism. To substantiate this mechanism, we mapped small RNAs at the \u003cem\u003eI\u003c/em\u003e locus. Cluster B within the \u003cem\u003eI\u003c/em\u003e locus, where the chimeric transcript originates, comprises \u003cem\u003eCHS1\u003c/em\u003e, \u003cem\u003eCHS3\u003c/em\u003e, \u003cem\u003eCHS4\u003c/em\u003e and a translocated subtilisin genomic segment. We aimed to confirm the generation of \u003cem\u003eCHS\u003c/em\u003e-targeting small RNA, particularly in \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e, which share intron sequences within the chimeric transcript, by analyzing small RNA patterns in both exon and intron regions of \u003cem\u003eCHS1\u003c/em\u003e, \u003cem\u003eCHS3\u003c/em\u003e, and \u003cem\u003eCHS4\u003c/em\u003e. To differentiate between initially generated small RNAs and those subsequently amplified during the small RNA generation process, we categorized them into 22 bp small RNAs (primary) and 21 bp small RNAs (secondary).\u003c/p\u003e \u003cp\u003eUpon examining the overall small RNA expression pattern in the cluster B region of the \u003cem\u003eI\u003c/em\u003e locus, we found that small RNAs containing sequences from \u003cem\u003eCHS1\u003c/em\u003e, \u003cem\u003eCHS3\u003c/em\u003e, and \u003cem\u003eCHS4\u003c/em\u003e were expressed in all soybean varieties except for black soybean (Daeheuk) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). The expression levels of chimeric transcript were measured in the seed coats of each allele, revealing expression in YL12 (allele \u003cem\u003eI\u003c/em\u003e), BR01 (allele \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{k}\\)\u003c/span\u003e\u003c/span\u003e), and YL07 (allele \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{i}\\)\u003c/span\u003e\u003c/span\u003e), but not in BL12 (allele \u003cem\u003ei\u003c/em\u003e) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). Notably, sRNA reads were confirmed in the intron regions of \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e, whereas there were very few reads in the intron region of \u003cem\u003eCHS4\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo differentiate between primary and secondary small RNAs, we categorized them into 22 bp and 21 bp and counted the sequencing reads mapped to exon and intron regions (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD). Small RNAs mapping to exon regions were generally expressed in all cultivars except BL12, with significantly higher expression levels for 21 bp small RNAs compared to 22 bp small RNAs. In contrast, small RNA read mapping to the intron regions was significantly lower in quantity. Interestingly, while there were no specific patterns in 21 bp small RNAs, 22 bp small RNAs exhibited a distinctive count of reads specifically in \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eValidation of the influence of small RNAs on CHS gene expression\u003c/p\u003e \u003cp\u003eThe impact of small RNAs on CHS gene expression was investigated using real-time PCR quantitative analysis. Assessment of the expression levels of small RNAs targeting the entire CHS gene family revealed that BL12 (allele \u003cem\u003ei\u003c/em\u003e) exhibited overall lower expression levels, while YL12 (allele \u003cem\u003eI\u003c/em\u003e), BR01 (allele \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{k}\\)\u003c/span\u003e\u003c/span\u003e) and YL07 (allele \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{i}\\)\u003c/span\u003e\u003c/span\u003e) displayed higher expression levels. Specifically, \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003e showed notably high expression (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA). These small RNAs, functioning as miRNAs, have the potential to repress CHS gene expression due to shared sequence composition.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo elucidate this process, \u003cem\u003eCHS\u003c/em\u003e transcripts were quantified in soybean seed coats at the R6.5 stage. BL12, BL10, and BL11 with the \u003cem\u003ei\u003c/em\u003e allele displayed overall high expression levels, whereas YL12 with the \u003cem\u003eI\u003c/em\u003e allele and YL07 with the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{i}\\)\u003c/span\u003e\u003c/span\u003e allele exhibited very low expression. High expression levels of \u003cem\u003eCHS7/CHS8\u003c/em\u003e were observed in most black soybeans harboring the \u003cem\u003ei\u003c/em\u003e allele (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB).\u003c/p\u003e \u003cp\u003eCloser examination of the sRNA-seq reads mapped to \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003e, which showed high expression at both the mRNA and miRNA levels, revealed that they were concentrated in the 2nd exon rather than the 1st exon, with a preference for the 21 bp size over the 22 bp size (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC). At the R6.5 stage, the expression of the CHS gene in leaves mirrored the pattern observed in the seed coat (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA). In the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cem\u003eI\u003c/em\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{k}\\)\u003c/span\u003e\u003c/span\u003e allele cultivars, the expression levels were low, while in \u003cem\u003ei\u003c/em\u003e allele cultivars, they showed statistically higher expression. When \u003cem\u003eCHS\u003c/em\u003e expression levels were examined in the leaves of BL12 (allele \u003cem\u003ei\u003c/em\u003e), YL07 (allele \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{i}\\)\u003c/span\u003e\u003c/span\u003e), YL12 (allele \u003cem\u003eI\u003c/em\u003e), and BR01 (allele \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{k}\\)\u003c/span\u003e\u003c/span\u003e) at the VC, V2\u0026ndash;V3, and R6.5 stages, they exhibited similarly low expression in the VC and V2\u0026ndash;V3 stages but statistically high expression in the R6.5 stage (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB)\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAnalysis of genomic structure of Hwangkeum using whole genome resequencing data\u003c/p\u003e \u003cp\u003eTo explore the reasons behind the challenge of confirming the genomic structure and chimeric transcript in Hwangkeum, despite its similar sRNA and \u003cem\u003eCHS\u003c/em\u003e transcript expression patterns to Williams82 and PI475822C, we conducted a structural variation analysis using whole genome resequencing data. Four whole-genome resequencing datasets for the Hwangkeum (YL12) variety were obtained from the European Nucleotide Archive (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ebi.ac.uk/ena/browser/home\u003c/span\u003e\u003cspan address=\"https://www.ebi.ac.uk/ena/browser/home\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and aligned to the Williams82 reference genome sequence (Supplementary Fig.\u0026nbsp;7). The resequencing reads showed good alignment with the reference genome sequence. In the \u003cem\u003eI\u003c/em\u003e locus region (chromosome 8; 8,498,281\u0026ndash;8,527,803), the Hwangkeum sequence closely resembled that of Williams82, though several single mismatches were identified in specific sequences. These mismatches were primarily concentrated in the region between cluster A and B.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eOur study aimed to uncover the mechanisms contributing to phenotypic variations in soybean seed coat color, focusing on the differences between black (allele \u003cem\u003ei\u003c/em\u003e) and yellow (allele \u003cem\u003eI, i\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e, \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e) soybeans. Inspired by a comparative analysis between \u003cem\u003eGlycine soja\u003c/em\u003e (W05, black wild soybean) and \u003cem\u003eG. max\u003c/em\u003e (Williams82, yellow soybean) (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e), we investigated the genomic variation in the \u003cem\u003eI\u003c/em\u003e locus cluster B region. Our research confirmed the expression of chimeric transcripts and identified primary miRNAs derived from these transcripts. Our findings demonstrated that the repression of \u003cem\u003eCHS\u003c/em\u003e gene expression accounts for the differences in seed coat color between black and yellow soybeans (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e7\u003c/span\u003e). To validate these results, we analyzed twelve cultivars carrying the \u003cem\u003ei\u003c/em\u003e allele, eleven with \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e allele, one with \u003cem\u003eI\u003c/em\u003e allele, and one with \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e allele (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Supplementary Table\u0026nbsp;1). This comprehensive analysis provided strong evidence linking the observed phenotypic variations to specific allelic differences.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eInitially, the \u003cem\u003eI\u003c/em\u003e locus was identified to harbor an inverted repeat of cluster A and B through a genomic comparison of \u003cem\u003eG. soja\u003c/em\u003e W05 and \u003cem\u003eG. max\u003c/em\u003e Williams82. Upstream of the cluster B region, sequences include the subtilisin promoter and the coding sequences for four exons of subtilisin. Previous studies (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e) reported genomic structural changes leading to the translocation of this region. Through PCR amplification and Sanger sequencing, we confirmed that this region remains intact in the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{i}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{k}\\)\u003c/span\u003e\u003c/span\u003e alleles (Figs.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA, \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB, and Supplementary Fig.\u0026nbsp;1A). The translocation of the subtilisin promoter and its four exons caused a rearrangement, leading to the creation of a new open reading frame (ORF). Consequently, a previously nonexistent 10,079 bp chimeric pre-mRNA is formed, with the final fifth exon encoded 9,077 bp downstream from the fourth exon.\u003c/p\u003e \u003cp\u003eSecondly, the long 9,077 bp intron within this chimeric transcript is spliced to produce a fully mature 723 bp mRNA (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). This transcript was verified through PCR amplification spanning the first and fifth exons (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD). Expression of of the chimeric transcript was detected in VC stage leaves, V2\u0026ndash;V3 stage leaves and roots, and R6.5 stage leaves. Notably, in VC stage leaves, it was expressed in cultivars carrying both the \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e and \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e alleles (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC), supporting the transcription of a 10,079 bp pre-mRNA driven by the relocated subtilisin promoter on chromosome 8. Sequencing confirmed the excision of the 9,077 bp intron during splicing (Supplementary Fig.\u0026nbsp;1B).\u003c/p\u003e \u003cp\u003eThe splicing mechanism involves the 2\u0026rsquo; OH group of an adenine within the branch point sequence (BPS) attacking the phosphate backbone at the 5\u0026rsquo; splice site, leading to the formation of a lariat structure (24\u0026ndash;26). This process brings \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e, which are embedded in an inverted orientation within the intron of chimeric transcript and share 98.54% sequence similarity, into close proximity. Their alignment likely promotes the formation of a complementary base-pairing region, resulting in a hairpin structure known as a mirtron (a short intron-derived pre-miRNA)(\u003cspan additionalcitationids=\"CR28\" citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e). This structural feature suggests the potential for mirtron formation, serving as a precursor to miRNA through subsequent processing.\u003c/p\u003e \u003cp\u003eTo investigate miRNA derived from the intron of chimeric transcript, we conducted small RNA (sRNA) sequencing across different alleles at the \u003cem\u003eI\u003c/em\u003e locus, including Daeheuk (BL12, allele \u003cem\u003ei\u003c/em\u003e), Hwangkeum (YL12, allele \u003cem\u003eI\u003c/em\u003e), Williams82 (YL07, allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e), PI475822C (BR01, allele \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA and Supplementary Fig.\u0026nbsp;2). While the CHS gene family exhibits high sequence similarity in exon regions, except for \u003cem\u003eCHS13\u003c/em\u003e and \u003cem\u003eCHS14\u003c/em\u003e (Supplementary Fig.\u0026nbsp;4), their introns are more divergent. Notably, only \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e share similar intronic sequences (Figs.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e6\u003c/span\u003eA, \u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e6\u003c/span\u003eB, and \u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e6\u003c/span\u003eC). This unique feature enables \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e, which are positioned in an inverted orientation within the intron of chimeric transcripts from cluster B of the \u003cem\u003eI\u003c/em\u003e locus, to potentially form hairpin structures during splicing. These hairpins are processed into primary miRNAs containing \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e sequences. As a result, the primary miRNAs can target other CHS genes with high sequence similarity, triggering the production of secondary miRNAs.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAnalysis of the overall sRNA expression pattern in the cluster B region of the \u003cem\u003eI\u003c/em\u003e locus, along with the expression of chimeric transcripts responsible for primary miRNA plroduction, revealed that \u003cem\u003eCHS1\u003c/em\u003e, \u003cem\u003eCHS3\u003c/em\u003e, and \u003cem\u003eCHS4\u003c/em\u003e \u0026mdash;including their mapped sRNAs and associated mRNA transcripts\u0026mdash; are expressed in all soybean varieties except the black soybean (Daeheuk, BL12) (Figs.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA and \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). Notably, miRNAs expressed in soybean seed coats predominantly exist in a 22-nt primary form, while the secondary form is 21-nt in length (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTo confirm that \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e are the sources of primary miRNAs targeting \u003cem\u003eCHS\u003c/em\u003e genes, we closely examined their intron regions, along with that of \u003cem\u003eCHS4.\u003c/em\u003e This analysis revealed that 22-nt sRNA reads specifically mapped to the introns of \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e, but not in \u003cem\u003eCHS4\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). These 22-nt sRNAs (primary miRNAs) are clearly mapped to \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e, whereas other CHS genes show much lower mapping (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD). Interestingly, 21-nt sRNAs (secondary miRNAs) are distributed more evenly across all CHS genes, suggesting that the amplification of primary miRNAs into secondary miRNAs is limited by sequence divergence (Figs.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD and \u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e6\u003c/span\u003eC).\u003c/p\u003e \u003cp\u003eIn contrast, the exon regions of \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e (YL07) and \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e (BR01) alleles exhibited higher levels of 21-nt sRNAs compared to 22-nt sRNAs, indicating that extensive amplification of secondary miRNAs occurs due to high exon sequence similarity among \u003cem\u003eCHS\u003c/em\u003e genes. Notably, the amplification of secondary miRNAs from primary miRNAs varies across CHS genes, ranging from a 2.37-fold increase (\u003cem\u003eCHS1\u003c/em\u003e) to a 10.48-fold increase (\u003cem\u003eCHS10\u003c/em\u003e), with Williams82 used as the reference (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003eThis observed pattern is attributable to differences in the intronic sequences of \u003cem\u003eCHS1/3\u003c/em\u003e and \u003cem\u003eCHS4\u003c/em\u003e, despite substantial sequence similarity in the exon regions. sRNA read mapping further supports this observation, as reads aligned specifically to regions with high sequence similarity to \u003cem\u003eCHS1/3\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e6\u003c/span\u003e). For \u003cem\u003eCHS2\u003c/em\u003e, \u003cem\u003eCHS4\u003c/em\u003e, \u003cem\u003eCHS5\u003c/em\u003e, \u003cem\u003eCHS6\u003c/em\u003e, \u003cem\u003eCHS9\u003c/em\u003e, \u003cem\u003eCHS10\u003c/em\u003e, and \u003cem\u003eCHS11\u003c/em\u003e, which share homology with \u003cem\u003eCHS1\u003c/em\u003e/\u003cem\u003e3\u003c/em\u003e primarily in the first and second exons, sRNA reads were mapped exclusively to those exonic regions. Similarly, for \u003cem\u003eCHS7\u003c/em\u003e/\u003cem\u003e8\u003c/em\u003e, which shares high similarity only in the second exon, reads were confined to that region. These findings support the conclusion that \u003cem\u003eCHS\u003c/em\u003e-targeting primary miRNAs originate within the intron of the chimeric transcript, explaining the widespread presence of both primary miRNAs and secondary miRNAs in all examined varieties except black soybeans (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003eThe amplification of primary miRNAs into secondary miRNAs is particularly notable in \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003e. sRNA-seq reads mapping to \u003cem\u003eCHS\u003c/em\u003e genes show that these reads are nearly absent in black soybeans, whereas \u003cem\u003eI\u003c/em\u003e, \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e, and \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e genotypes display similar expression patterns (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA), with \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003e exhibiting significantly higher sRNA expression levels. In the seed coat of black soybeans, where CHS gene silencing does not occur, mRNA levels of \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003e are notably elevated (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB). Further analysis of sRNA-seq reads mapped to \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003e, categorized by length, reveals that both 22-nt and 21-nt sRNAs are concentrated in the second exon, which shares a high sequence similarity with \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC). Notably, 21-nt reads are more abundant than 22-nt reads, indicating that \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003e transcripts are initially targeted by 22-nt primary miRNAs, which are subsequently amplified into 21-nt secondary miRNAs. This provides clear evidence of mirtron-triggered miRNA amplification process. This pattern of 22-nt primary miRNAs giving rise to 21-nt secondary miRNAs is consistently observed across the exons of highly sequence-similar \u003cem\u003eCHS1-11\u003c/em\u003e genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e6\u003c/span\u003e and Supplementary Fig.\u0026nbsp;4).\u003c/p\u003e \u003cp\u003eUltimately, primary miRNAs derived from \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e initiate a robust amplification of secondary miRNAs by targeting other CHS transcripts with high sequence similarity. This mechanism affects a broad spectrum of CHS genes, including three \u003cem\u003eCHS1\u003c/em\u003es, one \u003cem\u003eCHS2\u003c/em\u003e, four \u003cem\u003eCHS3\u003c/em\u003es, and one each of \u003cem\u003eCHS5\u003c/em\u003e, \u003cem\u003eCHS6\u003c/em\u003e, \u003cem\u003eCHS7\u003c/em\u003e, \u003cem\u003eCHS8\u003c/em\u003e, \u003cem\u003eCHS9\u003c/em\u003e, \u003cem\u003eCHS10\u003c/em\u003e, and \u003cem\u003eCHS11\u003c/em\u003e, extending the silencing effect beyond chromosome 8 to the entire genome (Supplementary Fig.\u0026nbsp;5 and Supplementary Table\u0026nbsp;5). The abundant production of these primary and secondary miRNAs constitutes a potent post-transcriptional regulatory mechanism that controls CHS gene expression across the genome (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA). This regulatory silencing is not restricted to the seed coat but also occurs in leaves and roots. Notably, the suppression of CHS gene is most pronounced during the reproductive growth stage \u0026mdash; characterized by rapid increase in CHS gene \u0026mdash; and continues through seed maturation. In contrast, minimal silencing is observed during early developmental or vegetative stages (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB).\u003c/p\u003e \u003cp\u003eThe robust production of primary and secondary miRNAs via this mechanism broadly suppresses the expression of CHS genes across diverse soybean cultivars, including Hwangkeum (\u003cem\u003eI\u003c/em\u003e), PI475822C (\u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e), and Williams82 (\u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e) (Figs.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB and \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA). This suppression effectively inhibits \u003cem\u003eCHS\u003c/em\u003e expression, thereby preventing the accumulation of anthocyanin and proanthocyanin in the seed coat (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eWhile specific data on \u003cem\u003eCHS13\u003c/em\u003e and \u003cem\u003eCHS14\u003c/em\u003e was not presented, these genes were excluded from all experimental analyses due to their consistently low expression levels observed across multiple developmental stages, including VC stage leaves, V2\u0026ndash;V3 stage leaves and roots, and R6.5 stage leaves and seed coats of Williams82. Their exclusion was further supported by their low sequence similarity to the broader CHS gene family (Supplementary Fig.\u0026nbsp;4). sRNA-seq read mapping to \u003cem\u003eCHS13\u003c/em\u003e and \u003cem\u003eCHS14\u003c/em\u003e revealed only minimal read counts, irrespective of the \u003cem\u003eI\u003c/em\u003e locus allele or sRNA size. Notably, sequences mapped to the intron region of \u003cem\u003eCHS13\u003c/em\u003e were identified as simple \u0026lsquo;TA\u0026rsquo; sequence repeats (Supplementary Fig.\u0026nbsp;6), suggesting that these reads likely represent background sRNA commonly found throughout the soybean genome, coincidently aligning to a repetitive region rather than originating from the \u003cem\u003eI\u003c/em\u003e locus.\u003c/p\u003e \u003cp\u003eThe \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e allele shares structural features at the \u003cem\u003eI\u003c/em\u003e locus with the \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e and \u003cem\u003eI\u003c/em\u003e alleles but is associated with the \u003cem\u003ek1\u003c/em\u003e mutant, which disrupts the function of Argonaute5, which is a key component of RNA interference (RNAi) pathway. This deficiency in Argonaute5 impairs the silencing of \u003cem\u003eCHS\u003c/em\u003e genes specifically around the hilum, leading to a characteristic saddle-shaped pigmentation pattern on the seed coat. Notably, in the hilum-adjacent region where pigment accumulates, the expression levels of \u003cem\u003eCHS7\u003c/em\u003e and \u003cem\u003eCHS8\u003c/em\u003e are approximately 26-fold higher than in more distal regions (\u003cspan additionalcitationids=\"CR33\" citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e). However, in areas where Argonaut5 remains functional, CHS gene silencing proceeds similarly to that observed in the \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e allele background (Figs.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFor the \u003cem\u003eI\u003c/em\u003e allele (Hwangkeum, YL12), PCR amplification targeting the chimeric transcript region and translocated sequences failed to yield detectable products (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). This discrepancy is likely due to the primers being designed based solely on the Williams82 reference genome, which may not align effectively with the genome sequence of Hwangkeum (YL12, allele \u003cem\u003eI\u003c/em\u003e) (Supplementary Tables\u0026nbsp;2 and 3). As a result, no PCR amplification was observed for regions with low-abundance transcripts (Figs.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB and \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003eDespite the PCR failure, sRNA-seq analysis revealed read alignment patterns consistent with those observed in Williams82 (\u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e), including \u003cem\u003eCHS\u003c/em\u003e targeting and the production of 22-nt primary miRNAs and 21-nt secondary miRNAs. These findings suggest that sequence mismatches, rather than major structural differences, are responsible for the failed PCR amplification and altered read alignment. Resequencing data from Hwangkeum (\u003cem\u003eI\u003c/em\u003e) confirmed that reads predominantly mapped to the \u003cem\u003eI\u003c/em\u003e locus (Supplementary Fig.\u0026nbsp;7), although single nucleotide mismatches were frequent, underscoring the limited specificity of the William82-based primers. Nevertheless, the presence of both primary and secondary miRNAs (Figs.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e and \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e) indicates that the core silencing mechanism is conserved. Consequently, suppression of seed coat pigmentation in Hwangkeum resembles the results observed in Williams82 (\u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sup\u003e) and PI475822C (\u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003ek\u003c/em\u003e\u003c/sup\u003e) is observed.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eThis study presents a comprehensive analysis of genomic variations associated with alleles at the \u003cem\u003eI\u003c/em\u003e locus, confirming the expression of the chimeric transcript arising from these structural differences. Using small RNA sequencing, we identified the generation of primary miRNAs originating from the intronic regions of \u003cem\u003eCHS1\u003c/em\u003e and \u003cem\u003eCHS3\u003c/em\u003e and demonstrated the extensive production of secondary miRNAs targeting the broader \u003cem\u003eCHS\u003c/em\u003e gene family. Our findings elucidate the functional role of mirtrons derived from chimeric transcript, resolving a longstanding question in the field regarding the RNA interference (RNAi) mechanism at the \u003cem\u003eI\u003c/em\u003e locus. Furthermore, by quantifying CHS gene expression across multiple developmental stages and tissues in cultivars carrying different \u003cem\u003eI\u003c/em\u003e locus alleles, we provide strong evidence supporting the mechanism underlying seed coat pigmentation suppression in soybeans. This step-by-step experimental validation integrates previously fragmented insights, leading to the establishment of the mirtron-triggered gene silencing (MTGS) model (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthics approval and consent to participate\u003c/h2\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003ch2\u003eConsent for publication\u003c/h2\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003ch2\u003eAvailability of data and materials\u003c/h2\u003e\n\u003cp\u003eThe datasets generated in this study are available in the NCBI BioProject repository under accession number PRJNA1112103 (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1112103).\u003c/p\u003e\n\u003ch2\u003eCompeting interests\u003c/h2\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThis research was supported by Basic Science Research Program through the National Research Foundation of Korea (NRF) funded by the Ministry of Education (grant number RS-2020-NR049596), Biomaterials Specialized Graduate Program through the Korea Environmental Industry \u0026amp; Technology Institute (KEITI) funded by the Ministry of Environment (MOE).\u003c/p\u003e\n\u003ch2\u003eAuthors\u0026rsquo; contributions\u003c/h2\u003e\n\u003cp\u003eYC: conceptualization, investigation, experiment, and writing \u0026ndash; original draft; J-SP: data curation and analysis; J-HK: conceptualization and analysis; M-GJ: data curation and validation; Y-IJ: conceptualization and data curation; H-KC: conceptualization, funding acquisition, supervision, project administration, writing \u0026ndash; original draft, and writing \u0026ndash; review \u0026amp; editing. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003ch2\u003eAcknowledgments\u003c/h2\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eClough SJ, Tuteja JH, Li M, Marek LF, Shoemaker RC, Vodkin LO. Features of a 103-kb gene-rich region in soybean include an inverted perfect repeat cluster of CHS genes comprising the I locus. Genome. 2004;47(5):819\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiang L, Yang X, Gao X, Yang H, Ma S, Huang S, et al. Multiomics Analyses Reveal the Dual Role of Flavonoids in Pigmentation and Abiotic Stress Tolerance of Soybean Seeds. J Agric Food Chem. 2024;72(6):3231\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu K, Xiao S, Chen Q, Wang Q, Zhang Y, Li K, et al. Changes in the Activity and Transcription of Antioxidant Enzymes in Response to Al Stress in Black Soybeans. Plant Mol Biol Rep. 2013;31(1):141\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZHANG T, KAWABATA K, KITANO R. Preventive Effects of Black Soybean Seed Coat Polyphenols against DNA Damage in Salmonella typhimurium. Food Sci Technol Res. 2013;19(4):685\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu K, Xiao S, Chen Q, Wang Q, Zhang Y, Li K, et al. Changes in the Activity and Transcription of Antioxidant Enzymes in Response to Al Stress in Black Soybeans. Plant Mol Biol Rep. 2013;31(1):141\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZabala G, Zou J, Tuteja J, Gonzalez DO, Clough SJ, Vodkin LO. Transcriptome changes in the phenylpropanoid pathway of Glycine max in response to Pseudomonas syringaeinfection. BMC Plant Biol. 2006;6(1):26.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBRUGLIERA TANAKAY, KALC F, SENIOR G, NAKAMURA MDYSONB. Flower Color Modification by Engineering of the Flavonoid Biosynthetic Pathway: Practical Perspectives. Biosci Biotechnol Biochem. 2010;74(9):1760\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim JH, Park JS, Lee CY, Jeong MG, Xu JL, Choi Y et al. Dissecting seed pigmentation-associated genomic loci and genes by employing dual approaches of reference-based and k-mer-based GWAS with 438 Glycine accessions. Rajcan I, editor. PLoS One. 2020;15(12):e0243085.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim JM, Lee JW, Seo JS, Ha BK, Kwon SJ. Differentially Expressed Genes Related to Isoflavone Biosynthesis in a Soybean Mutant Revealed by a Comparative Transcriptomic Analysis. Plants. 2024;13(5):584.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKohzuma K, Sato Y, Ito H, Okuzaki A, Watanabe M, Kobayashi H, et al. The non-mendelian green cotyledon gene in soybean encodes a small subunit of photosystem II. Plant Physiol. 2017;173(4):2138\u0026ndash;47.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSenda M, Kasai A, Yumoto S, Akada S, Ishikawa R, Harada T, et al. Sequence divergence at chalcone synthase gene in pigmented seed coat soybean mutants of the Inhibitor locus. Genes Genet Syst. 2002;77(5):341\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSenda M, Jumonji A, Yumoto S, Ishikawa R, Harada T, Niizeki M, et al. Analysis of the duplicated CHS1 gene related to the suppression of the seed coat pigmentation in yellow soybeans. Theor Appl Genet. 2002;104(6\u0026ndash;7):1086\u0026ndash;91.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSenda M, Yamaguchi N, Hiraoka M, Kawada S, Iiyoshi R, Yamashita K, et al. Accumulation of proanthocyanidins and/or lignin deposition in buff-pigmented soybean seed coats may lead to frequent defective cracking. Planta. 2017;245(3):659\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYuan B, Yuan C, Wang Y, Liu X, Qi G, Wang Y, et al. Identification of genetic loci conferring seed coat color based on a high-density map in soybean. Front Plant Sci. 2022;13(August):1\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZabala G, Vodkin LO. A rearrangement resulting in small tandem repeats in the F3\u0026prime;5\u0026prime;H gene of white flower genotypes is associated with the soybean W1 locus. Crop Sci. 2007;47(SUPPL 2):113\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTodd JJ, Vodkin LO. Pigmented Soybean. Plant Physiol. 1993;102:663\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSenda M, Masuta C, Ohnishi S, Goto K, Kasai A, Sano T, et al. Patterning of virus-infected Glycine max seed coat is associated with suppression of endogenous silencing of chalcone synthase genes. Plant Cell. 2004;16(4):807\u0026ndash;18.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTuteja JH, Zabala G, Varala K, Hudson M, Vodkin LO, Endogenous. Tissue-Specific Short Interfering RNAs Silence the Chalcone Synthase Gene Family in Glycine max Seed Coats. Plant Cell. 2009;21(10):3063\u0026ndash;77.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTuteja JH, Clough SJ, Chan WC, Vodkin LO. Tissue-Specific Gene Silencing Mediated by a Naturally Occurring Chalcone Synthase Gene Cluster in Glycine max [W]. Plant Cell. 2004;16(4):819\u0026ndash;35.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCho YB, Jones SI, Vodkin L. The Transition from Primary siRNAs to Amplified Secondary siRNAs That Regulate Chalcone Synthase During Development of Glycine max Seed Coats. Freitag M, editor. PLoS One. 2013;8(10):e76954.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJia J, Ji R, Li Z, Yu Y, Nakano M, Long Y, et al. Soybean DICER-LIKE2 Regulates Seed Coat Color via Production of Primary 22-Nucleotide Small Interfering RNAs from Long Inverted Repeats. Plant Cell. 2020;32(12):3662\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao Z, Nie J, Wang H. MicroRNA biogenesis in plant. Plant Growth Regul. 2021;93(1):1\u0026ndash;12.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXie M, Chung CYL, Li MW, Wong FL, Wang X, Liu A, et al. A reference-grade wild soybean genome. Nat Commun. 2019;10(1):1\u0026ndash;12.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDrift R, Velocity R. Encyclopedia of Astrobiology. Encyclopedia of Astrobiology. 2011.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFica SM, Tuttle N, Novak T, Li NS, Lu J, Koodathingal P, et al. RNA catalyses nuclear pre-mRNA splicing. Nature. 2013;503(7475):229\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHang J, Wan R, Yan C. Structural basis of pre-mRNA splicing. 2015;(August):4\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCurtis HJ, Sibley CR, Wood MJA. Mirtrons, an emerging class of atypical miRNA. Wiley Interdiscip Rev RNA. 2012;3(5):617\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFlynt AS, Greimann JC, Chung WJ, Lima CD, Lai EC. MicroRNA Biogenesis via Splicing and Exosome-Mediated Trimming in Drosophila. Mol Cell. 2010;38(6):900\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOkamura K, Hagen JW, Duan H, Tyler DM, Lai EC. The Mirtron Pathway Generates microRNA-Class Regulatory RNAs in Drosophila. Cell. 2007;130(1):89\u0026ndash;100.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZabala G, Vodkin L. Cloning of the Pleiotropic T Locus in Soybean and Two Recessive Alleles That Differentially Affect Structure and Expression of the Encoded Flavonoid 3\u0026prime; Hydroxylase. Genetics. 2003;163(1):295\u0026ndash;309.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZabala G, Zou J, Tuteja J, Gonzalez DO, Clough SJ, Vodkin LO. Transcriptome changes in the phenylpropanoid pathway of Glycine max in response to Pseudomonas syringae infection. BMC Plant Biol. 2006;6(1):26.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePalmer RG, Pfeiffer TW, Buss GR, Kilen TC. Qualitative Genetics. In: Soybeans: Improvement, Production, and Uses. 2016. pp. 137\u0026ndash;233.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCho YB, Jones SI, Vodkin LO. Mutations in Argonaute5 Illuminate Epistatic Interactions of the K1 and I Loci Leading to Saddle Seed Color Patterns in Glycine max. Plant Cell. 2017;29(4):708\u0026ndash;25.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSong J, Liu Z, Hong H, Ma Y, Tian L, Li X, et al. Identification and validation of loci governing seed coat color by combining association mapping and bulk segregation analysis in soybean. PLoS ONE. 2016;11(7):1\u0026ndash;19.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStefanova P, Taseva M, Georgieva T, Gotcheva V, Angelov A. A Modified CTAB Method for DNA Extraction from Soybean and Meat Products. Biotechnol Biotechnol Equip. 2013;27(3):3803\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSantaLucia J, Allawi HT, Seneviratne PA. Improved Nearest-Neighbor Parameters for Predicting DNA Duplex Stability. Biochemistry. 1996;35(11):3555\u0026ndash;62.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014;30(15):2114\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLangmead B, Trapnell C, Pop M, Salzberg SL. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol. 2009;10(3):R25.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnders S, Pyl PT, Huber W. HTSeq\u0026ndash;a Python framework to work with high-throughput sequencing data. Bioinformatics. 2015;31(2):166\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLove MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15(12):550.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLivak KJ, Schmittgen TD. Analysis of Relative Gene Expression Data Using Real-Time Quantitative PCR and the 2\u0026thinsp;\u0026ndash; ∆∆CT Method. Methods. 2001;25(4):402\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao M, Liu Y, Ma X, Shuai Q, Gai J, Li Y. Evaluation of reference genes for normalization of gene expression using quantitative RT-PCR under aluminum, cadmium, and heat stresses in soybean. PLoS ONE. 2017;12(1):1\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEdgar RC. MUSCLE: Multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004;32(5):1792\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTamura K, Stecher G, Kumar S. MEGA11: Molecular Evolutionary Genetics Analysis Version 11. Battistuzzi FU, editor. Mol Biol Evol. 2021;38(7):3022\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCamacho C, Coulouris G, Avagyan V, Ma N, Papadopoulos J, Bealer K, et al. BLAST+: architecture and applications. BMC Bioinformatics. 2009;10(1):421.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAltschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990;215(3):403\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25(14):1754\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"chalcone synthase, mirtron triggered gene silencing, RNA interference, seed coat color, soybean","lastPublishedDoi":"10.21203/rs.3.rs-6817785/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6817785/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eThe diversity of soybean seed coat colors is largely determined by anthocyanin accumulation, regulated by chalcone synthase (CHS) genes at the \u003cem\u003eI\u003c/em\u003e locus. While RNA interference (RNAi) has been implicated in this process, the underlying mechanisms remain incompletely understood. This study investigates the role of mirtrons (miRNA precursor derived from intronic sequences) in gene silencing related to seed coat pigmentation in black and yellow soybean varieties.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eTo validate the presence and function of these mirtrons, we conducted small RNA sequencing (sRNA-seq) for each allele at the \u003cem\u003eI\u003c/em\u003e locus. Our analysis reveals a complex, multi-layered regulatory mechanism involving the genomic architecture the \u003cem\u003eI\u003c/em\u003e locus, mRNA, and sRNA interactions. We identified mirtron-triggered miRNAs (MT-miRNAs) and their amplification into secondary miRNAs, which collectively mediate genome-wide silencing of CHS genes.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eThis study elucidates a cascade of mirtron-triggered gene silencing (MTGS) that regulates seed coat pigmentation soybeans. These findings provide novel insights into RNAi-mediated control of anthocyanin biosynthesis and highlight the significance of mirtrons in plant gene regulation.\u003c/p\u003e","manuscriptTitle":"Unveiling mirtron-triggered gene silencing mechanisms in soybean seed coat pigmentation","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-27 14:19:25","doi":"10.21203/rs.3.rs-6817785/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2b244b75-7017-44ec-bbb6-d09bb4ef5737","owner":[],"postedDate":"June 27th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-07-18T13:08:54+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-27 14:19:25","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6817785","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6817785","identity":"rs-6817785","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00