Transcriptome analysis revealed stress responsive fusion transcripts in Chickpea (Cicer arietinum)

preprint OA: closed
Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-06 · read from full text

This preprint used in-house RNA-seq data from five chickpea tissues and two abiotic stress conditions (drought and salinity) to identify fusion transcripts (chimeric mRNAs) using three fusion-detection tools (FusionMap, STAR-Fusion, MapSplice), aiming to generate high-confidence calls with stringent filtering. Across analyses, 328 unique fusion transcripts were detected, and 69% of fusion junctions showed canonical splice sites consistent with trans-splicing; functional annotation suggested that fusion partners may expand biological functionality. Ten fusion transcripts were validated by RT-PCR followed by Sanger sequencing as the first described fusion transcripts in chickpea, and expression profiling under drought and salinity identified differentially expressed fusion transcripts and their parental genes. A key caveat is that the study emphasizes the remaining challenge of distinguishing true biological events from artifacts, motivating its need for strict filtering and limited experimental validation. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Understanding the transcriptome complexity is essential for deciphering the genome regulation at transcriptional level. High-throughput sequencing technologies facilitated the detection of Fusion transcripts (FTs) which are chimeric mRNA molecules derived from gene fusion due to chromosomal rearrangements or via splicing machinery at the RNA level. In this study, we explored the transcriptome complexity in Cicer arietinum due to fusion events by using high-throughput RNA-Seq datasets in five tissues, and under two abiotic stress conditions. Using three different fusion detection tools, a total of 328 unique FTs were detected. Sequence analysis revealed that 69% of FTs showed the presence of canonical splice site at the junction, which indicates their generation via trans-splicing. Functional annotation and enrichment analyses of fusion partners suggested the expansion of biological functionality. A total of 10 FTs were validated via RT-PCR followed by Sanger sequencing, which are the first FTs described in important legume chickpea. Expression analysis of validated FTs and their parental genes under drought and salinity stress conditions identified differentially expressed fusions. This study offers detailed insight into the fusion landscape of Cicer arietinum and inferred a potential role of FTs during stress responses.
Full text 135,386 characters · extracted from preprint-html · click to expand
Transcriptome analysis revealed stress responsive fusion transcripts in Chickpea (Cicer arietinum) | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Transcriptome analysis revealed stress responsive fusion transcripts in Chickpea (Cicer arietinum) Fiza Hamid, Shafaque Zahra, Shailesh Kumar This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6540588/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Understanding the transcriptome complexity is essential for deciphering the genome regulation at transcriptional level. High-throughput sequencing technologies facilitated the detection of Fusion transcripts (FTs) which are chimeric mRNA molecules derived from gene fusion due to chromosomal rearrangements or via splicing machinery at the RNA level. In this study, we explored the transcriptome complexity in Cicer arietinum due to fusion events by using high-throughput RNA-Seq datasets in five tissues, and under two abiotic stress conditions. Using three different fusion detection tools, a total of 328 unique FTs were detected. Sequence analysis revealed that 69% of FTs showed the presence of canonical splice site at the junction, which indicates their generation via trans-splicing. Functional annotation and enrichment analyses of fusion partners suggested the expansion of biological functionality. A total of 10 FTs were validated via RT-PCR followed by Sanger sequencing, which are the first FTs described in important legume chickpea. Expression analysis of validated FTs and their parental genes under drought and salinity stress conditions identified differentially expressed fusions. This study offers detailed insight into the fusion landscape of Cicer arietinum and inferred a potential role of FTs during stress responses. Abiotic stress Cicer arietinum Fusion transcripts trans-splicing transcriptome complexity Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Mature mRNA molecules are conventionally formed through transcription and post-transcriptional modifications, where introns are excised, and stability is enhanced. Traditionally, mRNA was thought to originate solely from the alternative splicing of a single gene. However, advancements in sequencing technologies have uncovered the existence of novel transcripts, such as fusion transcripts (FTs), which arise by the joining of mRNA molecules derived from different genes (Gingeras 2009 ; Gupta et al. 2018 ). FTs can be generated at the DNA level through genomic rearrangements such as translocation, deletion, duplication, or inversion and give rise to a fused gene (Li et al. 2009a ; Annala et al. 2013 ); or at the RNA level through trans-splicing (Sutton and Boothroyd 1986 ) or read-through transcription (Varley et al. 2014 ), resulting in an increase in transcriptome complexity without a corresponding increase in the gene number. While fusion transcripts are found in both unicellular and multicellular organisms, they were previously considered rare in nature and occasionally dismissed as transcriptional artifacts. Up to now, extensive research has characterized the cellular functions of these transcripts in cancer (Latysheva and Babu 2016 ; Dupain et al. 2019 ). Recent studies indicate that in addition to their role in oncogenesis, FTs have also been reported under normal physiological conditions (Babiceanu et al. 2016 ; Chwalenia et al. 2017 ), and may act as potential regulators of their parental mRNA (Mukherjee et al. 2021 ). However, their presence and functional significance in plants remain largely unexplored. In past decades, efficient big data analysis has facilitated the discovery of fusion transcripts in many plants such as Trifolium pratense (Chao et al. 2018 ), Camellia sinensis (Qiao et al. 2019 ), Oryza sativa (Zhang et al. 2010 ), Arabidopsis thaliana (Singh et al. 2019 ), Brassica rapa (Tan et al. 2019 ), and Zea mays (Zhou et al. 2022 ). These plant-specific fusion transcripts add transcriptome complexity and introduce novel functions, which may be related to stress responses (Zhou et al. 2022 ), metabolic pathways (Hagel and Facchini 2017 ), or phenotypic traits (Chen et al. 2017 ). For instance, maize exhibits various fusion events in response to viral infection, involving proteins like nodulin, flavone synthase, and cation-transporting ATPase (Zhou et al. 2022 ). In tomatoes, the fusion transcript PFP-LeT6 is involved in leaf patterning (Kim et al. 2001 ), while the GN2 chimeric gene in rice, controls plant height, heading date, and grain number (Chen et al. 2017 ). In Arabidopsis thaliana , a protein derived from the fusion of glutamine synthase and nodulin domains regulates root morphogenesis and flagellin-triggered signalling (Doskočilová et al. 2011 ). Studies across species also suggest that fusion events like domain shuffling and DNA fusion enhance the catalytic efficiency of an enzyme (Farrow et al. 2015 ; Li et al. 2016 ; Hagel and Facchini 2017 ). These reports suggest that fusion events result in novel sequences with novel functions either as non-coding RNA or fusion proteins (Winzer et al. 2015 ; Qin et al. 2017 ). Chickpea ( Cicer L .) ranks as the third most cultivated legume crop globally, following common bean and pea (Nathawat et al. 2024 ). Cultivated in over 50 countries, the Indian subcontinent leads in production, contributing approximately 70% of the global output (Koul et al. 2022 ). In addition to its economic significance, chickpea is highly esteemed for its exceptional nutritional content, particularly its rich protein and carbohydrate content. Currently, chickpeas are grown on 15 million hectares worldwide, producing 15.9 million tons (FAO, 2023). However, this yield remains significantly below the crop's potential under optimal conditions. This gap is largely attributed to various biotic and abiotic stresses that hinder productivity. Abiotic stresses, such as salinity and drought, are key factors contributing to yield losses. Research has shown that chickpeas are particularly sensitive to salinity compared to other crops, making it a critical factor limiting yields (Flowers et al. 2010 ; Turner et al. 2013 ). The growing availability of high-throughput transcriptomic data offers valuable insights for tackling these stress-related challenges. Despite advancements in sequencing technologies, the exploration of transcriptome diversity in chickpea, particularly due to fusion events, remains limited (Chitkara et al. 2024 ). Recently, our research group (Chitkara et al. 2024 ) conducted a comprehensive profiling of fusion transcripts using an extensive collection of publicly available RNA-Seq datasets, encompassing diverse tissues and a wide range of conditions. Interestingly, this study revealed that the majority of identified fusion transcripts (~ 74%) were detected in only a single transcriptomic sample. One possible explanation could be the inherent variability introduced by such a large and heterogeneous dataset. Alternatively, many of these uniquely detected fusion transcripts might represent random artifacts rather than true biological events. Moreover, the validation rate was reported to be generally low in Arabidopsis thaliana and Oryza sativa , which could be attributed to high false-positive prediction rates or the inherently low expression levels of fusion transcripts. Notably, no experimental validation was performed for the fusion transcripts identified in Cicer arietinum . The current study addresses this gap by utilizing in-house generated RNA-Seq datasets of chickpea derived from specific abiotic stress conditions and tissue types, enabling the identification of high-confidence fusion transcripts by applying stringent filtering criteria. This enhances the reproducibility, which in turn strengthens the confidence in these fusion events as biologically relevant and potentially regulatory molecules, rather than technical noise. Here, we identified stress-responsive and conserved fusion transcripts in Cicer arietinum using a combination of second- and third-generation RNA-Seq datasets, along with experimental validation. Together, our data provide the genome-wide profiling of fusion events in chickpea and reveal their widespread occurrence and potential functions. Additionally, a benchmarking analysis was performed to evaluate the performance of different fusion detection tools, enabling the identification of the most effective methods for accurate fusion detection in Chickpea. This work thus complements our earlier large-scale study by providing deeper biological interpretation and functional insights into fusion transcript dynamics in chickpea. Ultimately, the investigation of these transcripts could uncover novel mechanisms of stress adaptation and enhance the overall understanding of gene expression dynamics in legumes. Materials and Methods Plant material, growth conditions, and stress treatment Chickpea (Genotype ICC4958) seeds were cultivated by using the established protocol (Garg et al. 2010 ). Seedlings were grown in plastic pots filled with a sterilized mix of agro-peat and vermiculite in a 1:1 ratio, maintained at 22°C with a 14-hour light cycle in a controlled growth chamber. Samples were collected at various developmental stages, including stem and leaves from the seedling stage, and buds, flowers, and pods during the flowering stage. To induce drought stress, 21-day-old seedlings were removed from the pots and placed on tissue paper for five hours. For salinity stress, seedlings were immersed in a beaker filled with a 150 mM NaCl solution at 22°C. Whole seedlings were collected after 5 hours of treatment, and from three independent biological replicates. These samples were immediately flash-frozen in liquid nitrogen and stored at -80°C until further analysis. Total RNA isolation, cDNA library preparation, and sequencing Total RNA was extracted from 100 mg of tissue using the RNeasy Plant Mini Kit (Qiagen). The quantity and quality of the RNA were assessed with a NanoDrop Spectrophotometer (NanoDrop Technologies), ensuring that only samples with a 260/280 ratio between 1.9 and 2.1, and a 260/230 ratio between 2.0 and 2.4, were selected for Illumina sequencing. A total of 26 paired-end RNA-Seq libraries were prepared and sequenced. The quality of the reads was assessed with FastQC ( https://www.bioinformatics.babraham.ac.uk/projects/fastqc/ ), and adapter trimming and filtering of low-quality reads were performed using fastp (Chen et al. 2018 ). High-quality reads were then aligned to the reference genome (ASM33114v1) using HISAT2-2.2.1 (Kim et al. 2019 ). Computational prediction of fusion transcripts Fusion transcripts were predicted using three different fusion detection tools viz. FusionMap (Ge et al. 2011 ), STAR-Fusion (Haas et al. 2019 ), and MapSplice (Wang et al. 2010 ). These tools have different fusion detection strategies and filtering criteria, which help in identifying a broad range of fusion events. These tools were selected based on available benchmarking publications where STAR-Fusion was ranked as the best tool in terms of their high sensitivity, accuracy, and execution time (Haas et al. 2019 ); and MapSplice and FusionMap show good sensitivity for fusion detection (Kumar et al. 2016 ). Gene Ontology and KEGG pathway analysis For all identified fusion transcripts, gene ontology (GO) and KEGG pathway analysis were performed using DAVID (Sherman et al. 2022 ) with a significance threshold of p < 0.05. These analyses help to identify biological processes, cellular components, and pathways potentially impacted by fusion events, shedding light on their functional relevance. Expression analysis of fusion transcripts The expression level of genes involved in fusion formation was analysed based on Transcripts Per Million (TPM) using StringTie (Pertea et al. 2016 ) with the default parameters. To study the impact of fusion formation on the expression of their parental genes, we categorized the samples into two groups for each fusion transcript: (i) fusion-present (F/P) - samples in which the corresponding fusion transcript was detected, and (ii) fusion-absent (F/A) - samples lacking the respective fusion. The expression levels of each parental gene were compared between each sample of the two groups. Genes exhibiting ≥ 2-fold change in expression in F/P samples compared to F/A samples were considered upregulated or downregulated. Validation using Long-Read RNA-Seq data For further validation of fusion transcripts, PacBio long-read RNA-Seq data (PRJNA613159, Jain et al. 2022 ) for Cicer arietinum (ICC4958), were downloaded from the NCBI Sequence Read Archive (SRA). The 200 bp junction sequences of the identified fusion transcripts-comprising 100 bp upstream of the 5′ parental gene breakpoint and 100 bp downstream of the 3′ parental gene breakpoint-were searched against the long-read RNA-Seq data using BLASTn. Hits with greater than 80% sequence identity and alignment length exceeding 150 bp were considered significant. This approach confirmed the presence of predicted fusion transcripts within the long-read data, enhancing confidence in the identified fusion events. Additionally, it offered insights into the putative lengths of the fusion transcripts. Identification of conserved fusion events To identify intra-specific conserved fusions, 103 raw transcriptome data files from 17 Cicer arietinum genotypes were downloaded from the NCBI-SRA ( https://www.ncbi.nlm.nih.gov/sra ). List of RNA-Seq samples from different genotypes used in this study are listed in Table S6. Fusion events were then identified across all genotypes using three fusion detection tools. A binary analysis was conducted to determine the presence or absence of each fusion gene pair in the respective genotypes. For inter-specific conserved fusion detection, fusion genes previously reported in Arabidopsis thaliana (via AtFusionDB: http://www.nipgr.ac.in/AtFusionDB ) were utilized to identify their homologous fusion gene pairs in Cicer arietinum using the OrthoFinder tool. Experimental validation of fusion transcripts To validate fusion transcripts, primers were designed for the fusion transcript sequences taken from the genome 200 bp upstream and 200 bp downstream of the breakpoint, using the OligoCalc (Kibbe 2007 ) and primer BLAST. All primers used for validation are listed in Supplementary Table S7. The cDNA synthesis was carried out with 2µg of RNA using the Verso cDNA synthesis kit (ThermoScientific™). Following RT-PCR and gel electrophoresis, DNA bands were extracted and purified using the GenElute™ Gel Extraction kit (Sigma-Aldrich) and sent for Sanger sequencing at the NIPGR DNA sequencing facility. To quantify the expression of fusion transcripts under drought and salt stress, and across different tissues, quantitative Real-Time PCR was performed. EF1-α was used as an endogenous control gene in this experiment. The real-time PCR reaction mix (10 µl) consisted of 5 µl of 2X SYBR Green Master Mix (Applied Biosystems™), 10 µM of each primer, and 100 ng of cDNA template. PCR amplification was performed using an Applied Biosystems™ qPCR system with thermal cycling conditions of initial denaturation at 95°C for 2 minutes, followed by 40 cycles of 95°C for 15 seconds and 60°C for 1 minute. Each PCR reaction included three biological replicates and three technical replicates. The relative expression levels of fusion transcripts were calculated by using the 2^-ΔΔCt method (Livak and Schmittgen 2001 ). The experimental data were presented as the mean with standard deviation (mean ± SD) derived from three independent biological replicates and three technical replicates. Statistical analysis was conducted by comparing the means of control and stressed plants using one-way analysis of variance (ANOVA) followed by Student’s t-test, with a significance level set at P < 0.05. P-values less than 0.05 were considered statistically significant. Benchmark analysis of fusion detection tools To benchmark the performance of various fusion detection tools for Cicer , we evaluated STAR-Fusion (Haas et al. 2019 ), SQUID (Ma et al. 2018 ), MapSplice (Wang et al. 2010 ), Tophat-Fusion (Kim and Salzberg 2011 ), and FusionMap (Ge et al. 2011 ). The performance of each tool was assessed using both public dataset (PRJNA288321, Garg et al. 2016 ), and our inhouse dataset. For both datasets, we then searched the validated fusion transcripts (true fusions) in the results generated by each tool. The accuracy of fusion prediction was calculated following the method described by Kumar (2016) (Kumar et al. 2016 ), as given below. $$\:Sensitivity\:\left(\%\right)=\left(\frac{TP}{TF}\right)*100$$ $$\:Positive\:Predictive\:Value\:\left(PPV\right)=\frac{TP}{TP+FP}*100$$ $$\:F\:measure=2*\left(S*PPV\right)/(S/PPV)$$ S: Sensitivity, TP: True positive, TF: Total fusions, FP: False positive Results Identification of fusion transcripts in Cicer arietinum To comprehensively investigate fusion events within the chickpea transcriptome, high-throughput paired-end RNA-Seq data were generated using the Illumina sequencing platform. The samples were sequenced from poly(A)-enriched RNAs extracted from five distinct organs, viz. buds, leaves, pods, flowers, and stem (Fig. 1 a), along with two abiotic stress conditions (i.e., drought and salinity). After adapter trimming and quality-check, high-quality reads were mapped onto the CDC Frontier Reference Genome (ASM33114v1) (Varshney et al. 2013 ) using HISAT2 (v2.2.1) (Kim et al. 2019 ), and it was found that > 95% of the reads mapped onto the genome from each paired-end RNA-Seq sample (Table S1 ). In total, 122.47 GB of data was obtained from all samples, representing about 230-fold of the chickpea genome size, and around 852.8 million high-quality reads were generated. Three fusion detection tools were employed for genome-wide identification of fusion transcripts in 26 RNA-Seq datasets of different tissues and two abiotic stress conditions (Table S1 ). The number of overlapping fusion transcripts detected between different tissues and stress samples are shown in Fig. 1 a. The total number of fusion transcripts identified by FusionMap (Ge et al. 2011 ), STAR-Fusion (Haas et al. 2019 ), and MapSplice (Wang et al. 2010 ) are 496, 721 and 533 respectively (Table S2). However, the number of unique fusion transcripts identified is 95, 109, and 140, respectively. A total of 1750 fusion transcripts were identified in our datasets, constituting 328 unique fusions (Table S3), which were derived from 269 unique parental gene pairs. In total, 423 genes were involved in fusion formation, accounting for 1.4% (423/30,344) of the total annotated genes in Cicer arietinum . This proportion is similar to the percentage of genes involved in fusion events in other plant species, such as Arabidopsis (1%), soybean (1.7%), and rice (2.7%) (Cong et al. 2024 ). To verify the reliability of the predicted fusion transcripts, we randomly selected and validated 10 fusion transcripts in multiple biological replicates by PCR and Sanger sequencing (Fig. S1 ). The experimentally validated junction sequences precisely matched the in-silico predicted breakpoint sites, confirming the accuracy of the fusion detection approach. Features of fusion transcripts in Cicer arietinum Initial analysis reveals that interchromosomal (76%) fusions are more prevalent than intrachromosomal (24%) fusions (Fig. 1 b). This indicates that, in Cicer arietinum , the formation of fusion transcripts is not strictly governed by the linear proximity of parental genes, in contrast to Arabidopsis , where intrachromosomal fusions between adjacent genes are more common (Cong et al. 2024 ). This suggests that, beyond linear arrangement of genes, spatial organization may play a role in facilitating fusion events in chickpea and thus underscores the value of further investigation into the spatial organization of the genome. A positive correlation was observed between the number of genes participating in fusion events located at a particular chromosome, and the total number of genes on that chromosome, suggesting there is no chromosome-level preference in fusion formation (Fig. 1 c). Out of the 423 genes involved in fusion events, 306 participate in only one unique fusion event, while 70 genes form fusions with two different partner genes. This indicates that fusion events are highly specific, with individual genes preferentially associating with particular partners rather than engaging in multiple fusion events. However, the mechanisms behind partner selection in fusion events remain unknown. From the validated fusion pairs, all genes were found to fuse with only one partner gene, except LOC101495229, which formed fusions with both LOC101508351 and LOC101506473 (Fig. 2 a). The number of fusion isoforms generated by a fusion gene pair was inversely related to the number of such gene pairs detected, highlighting the specificity of junction sites within fusion genes. Among the validated fusions, all gene pairs had a single junction site, except the LOC101509445_LOC101509981 fusion, which exhibited two isoforms. In both isoforms, the breakpoint in the LOC101509445 gene was identical, whereas LOC101509981 contributed two distinct breakpoints, each located at the exon boundaries of two different alternatively spliced transcripts of the LOC101509981 gene. This is a read-through fusion derived from two adjacent genes, where LOC101509445 is present on the reverse strand and LOC101509981 is present on the forward strand of chromosome 7. In this fusion, the 5’ gene belongs to the ABC transporter I family member 1, and the 3’ gene encodes for tRNA pseudouridine synthase. The two fusion events occur at the existing exon boundaries, hence producing in-frame fusions (Fig. 2 b). To determine the fusion site within the parental genes, we analysed the location of the breakpoint, whether it exists on the exon boundaries of both parents or one or none. It was observed that none type (57%) of fusion transcripts, where breakpoints were present within the exons or in the UTR region, were most abundant, whereas fusion transcripts having breakpoints at the exon borders of either one or both parental genes were 25.3% and 17.6%, respectively (Table S3). Trans-splicing is a known mechanism for fusion transcript formation (Li et al. 2009a ). Junction pattern analysis showed that a significant portion of FTs exhibited canonical splice patterns: GT-AG (69%), however, non-canonical splicing patterns were also found such as GT-AT (4%), AT-AC (3%), CT-GC (3%), etc., implying apart from splicing, other unknown mechanisms may also contribute to fusion formation (Fig. 3 a). Another reported mechanism of fusion formation is transcriptional slippage mediated by the presence of short homologous sequences (SHSs) (Li et al. 2009b ) at the junction but less than 5% of total fusions exhibit SHSs at the breakpoints in Arabidopsis , soybean, rice, and maize (Cong et al. 2024 ). Similar results were observed in chickpea; 4% of total fusions showed SHSs at the junction, and 5.7% had both canonical splice sites and SHSs at their junction (Table S4). All the experimentally validated fusions exhibited canonical splice sites at their junctions (e.g., LOC101515613_LOC113787786; Fig. 3 a), except for LOC101494819_LOC101493433, which showed a non-canonical GA-AG splice site. Additionally, three of the fusions also displayed short homologous sequences (SHSs) at the junctions, e.g., LOC101500313_LOC101505021 (Fig. 3 b). Properties of genes involved in fusion formation Analysis of the transcriptome data showed a significant positive correlation (R = 0.53) between the expression levels of gene pairs involved in fusion formation, suggesting that these genes are linked in terms of expression (Fig. 4 a). Further to evaluate the impact of fusion formation on their parental genes, the expression of fusion parental genes was compared between samples with and without fusion transcripts and it was observed that most of the fusion transcripts (~ 75%) have no significant effect on the expression of their parental genes, suggesting that fusion formation may be independently regulated. However, a subset of parental genes exhibited differential expression in the presence of fusion transcripts, with approximately 21% showing upregulation and 4% showing downregulation (Fig. 4 b). Fusion transcripts may retain or be associated with the biological functions of their parental genes. To gain insights into the potential functions of the fusion transcripts, we employed gene ontology (GO) and KEGG pathway analyses of parental genes involved in fusion formation. The most enriched biological GO terms were translation, photosynthesis, response to light stimulus, and rRNA processing (Fig. 4 c). The most enriched genes in the molecular process analysis showed that fusion transcripts parental genes are involved in the structural constituent of the ribosome and in binding such as RNA binding and chlorophyll-binding and in enzymatic activities such as ATP hydrolysis, carboxylase, and dehydrogenase. The most enriched cellular component GO terms were cytosol, ribosome, and chloroplast. Pathway enrichment analysis showed that these genes are related to the Ribosome, carbon metabolism, and photosynthesis. Together, we can conclude that fusion transcripts originate from genes with diverse functions that are distributed across various cellular compartments, and also show important enzymatic and binding activities. Notably, photosynthesis and ribosome-associated genes are majorly enriched among fusion precursor genes (Fig. 4 c). Validation of fusion transcripts from Long-read RNA-Seq data Illumina sequencing generates short reads that require assembly to reconstruct full-length transcripts. However, this process can introduce errors, particularly in regions with complex splicing patterns or novel splice sites. To overcome these challenges, long-read RNA sequencing technologies, such as PacBio, offer a significant advantage by capturing full-length transcripts in a single read, eliminating the need for assembly. To detect the full length of the identified FTs, publicly available long-read RNA-Seq data of Cicer arietinum were utilized. To identify the full-length sequence of FTs, BLASTn search of 200 bp fusion junction sequences was conducted against the PacBio long-read data. Through this approach, we found 95 fusion transcripts in PacBio data, and their full length was predicted (Table S5). The mean length of identified fusions was around 0.8 kb, with 70% of them having a length less than the mean length. We further compared the length of fusion transcripts and their parental genes and observed that the mean length of fusion transcripts was shorter than that of their parental genes (Fig. S2). This indicates that only specific portions of the parental genes are involved in fusion formation. Intra and Inter-specific conserved fusion events To explore intra-specific conserved fusion events, RNA-Seq data of 17 different chickpea genotypes were analysed. The number of fusion genes identified varied significantly across genotypes, ranging from 41 to 1031 (Fig. 5 a). This variation is likely attributed to differences in the number of samples analysed for each genotype and the sequencing depth of the data. A fusion event was considered in a particular genotype if it was detected even in a single sample. After removing redundancy, interchromosomal fusion transcripts were found to be more prevalent than intrachromosomal fusion transcripts. Fusion genes identified from the ICC4958 genotype (269 fusion gene pairs) were searched across different genotypes (Fig. 5 b). Among these, 140 fusion gene pairs were conserved across multiple genotypes, while the rest were specific to the ICC4958. Notably, 53 fusion gene pairs were detected in 10 or more genotypes, suggesting their potential biological significance and stability within the chickpea population. To identify conserved FTs between Chickpea and Arabidopsis , FTs reported in AtFusionDB (Singh et al. 2019 ) were used. Among the 269 fusion genes identified in chickpea, 19 showed homology with 58 fusion events reported in Arabidopsis within AtFusionDB (Fig. S3a). Gene ontology analysis of these homologous genes revealed their involvement in essential biological processes, including metabolism and environmental responses (Fig. S3b-d). The presence of these fusion transcripts may play a critical role in enhancing biological functions linked to these processes. Validation of fusion transcripts by RT-PCR and Sanger sequencing To quantify the expression of validated fusion transcripts under different abiotic stress and tissues, quantitative real-time PCR was performed. The notable findings from expression analysis reveal that fusion transcripts are expressed at a very low level, which leads to their low detection rate. Under different abiotic stress, three of the validated fusions (LOC101506206_LOC101493600, LOC101500131_LOC101505021, and LOC101495229_LOC101508351) showed differential expression which implies that these fusions might play a crucial role during stress response (Fig. 6 ). Expression analysis of stress-responsive fusion precursor genes indicated that the relative fold change in expression of fusion transcripts was higher than that of their parental genes under stress conditions. This suggests that fusion transcript formation is not solely controlled by its precursor gene expression, but may involve additional regulatory mechanisms (Fig. 6 ). Fusion transcripts can originate from parental genes that exhibit diverse responses to stress, including downregulation, upregulation, or no significant change in expression. Expression analysis of fusion transcripts across different tissues revealed their tissue-specific nature. For example, the LOC101500131_LOC101505021 fusion transcript exhibited upregulation in bud and downregulation in stem and flower, relative to its expression in leaves (Fig. S4). The LOC101506206_LOC101493600 is an interchromosomal fusion derived from a gene encoding a folylpolyglutamate synthase-like enzyme located on chromosome 5 and an uncharacterized gene from chromosome 8. This fusion produces an in-frame transcript that arises from an existing splice site of the 5' gene and a new splice site from the 3’ gene. Since folylpolyglutamate synthase is a single-subunit enzyme, this fusion might add new domains, which may enhance its enzymatic activity (Muralla et al. 2008 ; Li et al. 2016 ). Interestingly, under stress conditions, the fusion transcript exhibited a higher relative fold change in expression compared to either of its precursor genes—LOC101506206, which was upregulated, and LOC101493600, which was downregulated—when compared to control conditions. The LOC101500131_LOC101505021 fusion was identified across all RNA-Seq samples with varying levels of expression among different samples. This fusion is formed by the joining of exon 10 of the LOC101500131 gene to exon 1 of the LOC101505021 gene, both genes are present on the forward strand of DNA and are involved in similar functions that are anaphase-promoting complex. This fusion uses a canonical splice site, however, not at the existing exon boundaries of the parental genes, and generates a frameshift fusion. Interestingly, this fusion showed significant upregulation in both drought and salt stresses as compared to the control condition. In contrast, the expression of the fusion precursor genes remains relatively unchanged in response to stress. Therefore, the regulation of the LOC101500131_LOC101505021 fusion appears to be independent of its parental genes. The LOC101495229_LOC101508351 fusion is formed by joining exon 2 from LOC101495229 and exon 1 from LOC101508351. One of the parental genes acts as a ubiquitin ligase, while the other remains uncharacterized. This fusion occurs at the exon boundary of the LOC101508351 gene and uses a new canonical splice site from the LOC101495229 gene. Both fusion transcript and its parental genes are stress-responsive; however, the relative change in fusion transcript expression under stress conditions is higher compared to parental genes. Comprehensive assessment of fusion detection tools for Cicer arietinum While many tools are available for fusion detection in humans, none of them is exclusively designed for plants, except EricScript-Plants (Benelli et al. 2012 ), which is limited to plants in Ensembl and hence cannot be used for chickpea. In this study, we initially employed FusionMap, STAR-Fusion, and MapSplice to enable the comprehensive detection of fusion transcripts. However, to determine the best-performing tool for fusion detection in chickpea, a comparative analysis of five fusion detection algorithms was conducted. This revealed the following order of tools based on sensitivity and F-measure: FusionMap > STAR-Fusion > MapSplice > SQUID > Tophat-Fusion. A similar ranking was observed with public data, except that STAR-Fusion performed better than FusionMap with public data (Fig. S5a-b). Of note, there is a small overlap in the fusions detected by these tools in both our dataset and public data (Fig. S5c-d). This could be due to false discoveries associated with individual software packages, or the fact that none of the tools is inclusive. Fusion gene pairs LOC101506206_LOC101493600 and LOC101503481_LOC101494793 were commonly detected by 3 out of the 5 fusion detection tools. This analysis suggests that FusionMap detected the maximum number of true fusions, but it also identified a large number of false fusions, whereas STAR-Fusion identified fewer false fusions compared to FusionMap. Hence, FusionMap is the best performer in terms of sensitivity, whereas STAR-Fusion is the best performer in terms of precision. Hence, a combination of STAR-Fusion and FusionMap is a better choice for fusion transcript detection from RNA-Seq data for Cicer until more efficient tools become available. Another factor that affects the detection of FTs from RNA-Seq data is the sequencing depth. It was observed that samples with higher sequencing depth showed a larger number of fusions than those with lower sequencing depth. For instance, in the P3 sample, the number of FTs detected is more than that in CS2 because the sequencing depth of P3 is more than CS2 (Table S2). Discussion In recent years, its already established that the fusion transcripts are not exclusive to tumours but also occur in normal human tissues and a wide variety of species (Babiceanu et al. 2016 ; Cong et al. 2024 ). Although significant progress has been made in mammalian fusion transcript research, studies on fusion transcripts in plants remain limited. While some studies have been conducted on model species such as Arabidopsis and rice, a comprehensive investigation of plant fusion transcripts is still lacking. A few plant-specific fusion transcripts databases are also available (Singh et al. 2019 ; Arya et al. 2024 ). Recently, a comprehensive profiling of fusion transcripts was conducted using public RNA-Seq datasets of Arabidopsis , Rice, and Chickpea (Chitkara et al. 2024 ). Most fusion transcripts were detected in only a single sample, raising concerns about their biological reproducibility and potential as artifacts. The current study addresses this gap through the use of in-house generated RNA-Seq datasets under defined abiotic stress conditions and across tissue types. Here, we present the genome-wide identification of fusion transcripts in Cicer arietinum and explore their potential roles. Our study enhances the current understanding of fusion transcripts in legume plants. Notably, we found that interchromosomal fusion transcripts are more prevalent than those of intrachromosomal fusions in Chickpea, a pattern consistent with maize and soybean. In contrast, rice and Arabidopsis exhibit a higher frequency of intrachromosomal fusions (Cong et al. 2024 ). This indicates that the proportion of interchromosomal and intrachromosomal fusion transcripts is variable among different plants and may be influenced by several factors, including the gene density and compactness of the genome. The majority of gene pairs involved in fusion formation have no other partner genes, with only a few showing fusion isoforms, highlighting the specificity of these events. Additionally, we identified fusion transcripts that are tissue-specific or induced under specific stress conditions. Most parental genes involved in these fusions are multi-exonic and protein-coding. Expression correlation analysis of genes involved in fusions revealed a correlation coefficient (r) of 0.5, aligning with the recent report that suggests genes with similar transcriptional activity are more likely to participate in conserved fusion events (Cong et al. 2024 ). Fusion transcript identification in different genotypes of chickpea revealed that some fusions are conserved across all genotypes, while others are specific to individual genotypes, indicating a link between genomic variation and fusion transcript occurrence. Homologous fusion events between Cicer arietinum and Arabidopsis suggest that the parental genes involved play key roles in essential biological processes, underscoring the conservation of fusion events across different plant species. Most fusion breakpoints exhibited canonical splice sites at the fusion junctions, implying their generation via a splicing mechanism. However, a few fusions showed overlapping sequences at the breakpoints, suggesting that SHSs also contribute to the formation of these fusions in Cicer arietinum . Further investigation revealed that SHSs differ among fusion transcripts and are mostly < 10 bp in length. Due to the short reads generated by Illumina sequencing, full-length fusion transcripts could not be reliably detected. To address this limitation, we employed long-read RNA sequencing to identify the full-length structure of fusion transcripts. Of the 328 fusion events initially identified, 95 were confirmed in the long-read RNA-Seq data, likely due to the lower sequencing depth of this dataset. Out of the predicted fusion transcripts, we confirmed the existence of 10 FTs, marking the first experimental validation of fusion transcripts in Cicer arietinum . Notably, three of these fusions were stress-responsive and exhibited upregulation in drought and salt stress, suggesting a potential function in the plant’s stress response mechanisms. However, despite thorough validation efforts, we were unable to identify any fusion transcripts that were exclusively expressed under a specific stress and absent in control condition. This indicates that while fusion transcripts might be involved in the stress response, their expression may not be strictly limited to stress conditions. It is possible that these fusions represent a broader regulatory mechanism that operates under both normal and stress conditions but becomes more pronounced in response to stress. The confirmation of stress-responsive fusion transcripts paves the way for exploring the functional implications of these fusions in stress adaptation and resilience. Future research could focus on dissecting the biological roles of these fusions, including their impact on gene expression and protein function under stress conditions. Additionally, investigating the potential regulatory networks and pathways associated with these fusions could provide profound understanding into how Cicer arietinum adapts to environmental challenges. Declarations Competing interests The authors have no conflicts of interest to declare. Funding This research is supported by the BT/PR40146/BTIS/137/4/2020, BT/PR40169/BTIS/137/71/2023 research grants by the Department of Biotechnology (DBT), Government of India, and EEQ/2019/000231 research grant from Science and Engineering Research Board (SERB), Department of Science and Technology, Government of India; and core research grant of NIPGR, New Delhi Authors' contributions FH: Conceptualization, Methodology, data curation, analysis, writing-original draft. SZ: Suggestions in data analysis and manuscript writing. SK: Conceptualization, Methodology, reviewed & edited the manuscript, conceived, and coordinated the project and provided overall supervision. All authors read and approved the final version of the manuscript. Acknowledgements FH gratefully acknowledges the Department of Biotechnology (DBT) for providing research fellowship. The authors extend their gratitude to the DBT e-Library Consortium (DeLCON) for providing access to e-material and Computational Biology & Bioinformatics Facility (CBBF) of NIPGR for their support. References Annala MJ, Parker BC, Zhang W, Nykter M (2013) Fusion genes and their discovery using high throughput sequencing. Cancer Lett 340:192–200. https://doi.org/10.1016/J.CANLET.2013.01.011 Arya A, Arora S, Hamid F, Kumar S (2024) PFusionDB: a comprehensive database of plant-specific fusion transcripts. 3 Biotech 2024 14:11 14:1–8. https://doi.org/10.1007/S13205-024-04132-1 Babiceanu M, Qin F, Xie Z, et al (2016) Recurrent chimeric fusion RNAs in non-cancer tissues and cells. Nucleic Acids Res 44:2859–2872. https://doi.org/10.1093/NAR/GKW032 Benelli M, Pescucci C, Marseglia G, et al (2012) Discovering chimeric transcripts in paired-end RNA-seq data by using EricScript. Bioinformatics 28:3232–3239. https://doi.org/10.1093/BIOINFORMATICS/BTS617 Chao Y, Yuan J, Li S, et al (2018) Analysis of transcripts and splice isoforms in red clover (Trifolium pratense L.) by single-molecule long-read sequencing. BMC Plant Biol 18:1–12. https://doi.org/10.1186/S12870-018-1534-8/FIGURES/7 Chen H, Tang Y, Liu J, et al (2017) Emergence of a novel chimeric gene underlying grain number in rice. Genetics 205:993–1002. https://doi.org/10.1534/GENETICS.116.188201/-/DC1 Chen S, Zhou Y, Chen Y, Gu J (2018) fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34:i884–i890. https://doi.org/10.1093/BIOINFORMATICS/BTY560 Chitkara P, Singh A, Gangwar R, et al (2024) The landscape of fusion transcripts in plants: a new insight into genome complexity. BMC Plant Biol 24:1–18. https://doi.org/10.1186/S12870-024-05900-0/FIGURES/6 Chwalenia K, Facemire L, Li H (2017) Chimeric RNAs in cancer and normal physiology. Wiley Interdiscip Rev RNA 8:. https://doi.org/10.1002/WRNA.1427 Cong J, Zhang S, Zhang Q, et al (2024) Conserved features and diversity attributes of chimeric RNAs across accessions in four plants. Plant Biotechnol J 22:3151–3163. https://doi.org/10.1111/PBI.14437 Doskočilová A, Plíhal O, Volc J, et al (2011) A nodulin/glutamine synthetase-like fusion protein is implicated in the regulation of root morphogenesis and in signalling triggered by flagellin. Planta 234:459–476. https://doi.org/10.1007/S00425-011-1419-7 Dupain C, Harttrampf AC, Boursin Y, et al (2019) Discovery of New Fusion Transcripts in a Cohort of Pediatric Solid Cancers at Relapse and Relevance for Personalized Medicine. Molecular Therapy 27:200. https://doi.org/10.1016/J.YMTHE.2018.10.022 Farrow SC, Hagel JM, Beaudoin GAW, et al (2015) Stereochemical inversion of (S)-reticuline by a cytochrome P450 fusion in opium poppy. Nat Chem Biol 11:728–732. https://doi.org/10.1038/NCHEMBIO.1879 Flowers TJ, Gaur PM, Gowda CLL, et al (2010) Salt sensitivity in chickpea. Plant Cell Environ 33:490–509. https://doi.org/10.1111/J.1365-3040.2009.02051.X Garg R, Sahoo A, Tyagi AK, Jain M (2010) Validation of internal control genes for quantitative gene expression studies in chickpea (Cicer arietinum L.). Biochem Biophys Res Commun 396:283–288. https://doi.org/10.1016/J.BBRC.2010.04.079 Garg R, Shankar R, Thakkar B, et al (2016) Transcriptome analyses reveal genotype- and developmental stage-specific molecular responses to drought and salinity stresses in chickpea. Scientific Reports 2016 6:1 6:1–15. https://doi.org/10.1038/srep19228 Ge H, Liu K, Juan T, et al (2011) FusionMap: detecting fusion genes from next-generation sequencing data at base-pair resolution. Bioinformatics 27:1922–1928. https://doi.org/10.1093/BIOINFORMATICS/BTR310 Gingeras TR (2009) Implications of chimaeric non-co-linear transcripts. Nature 2009 461:7261 461:206–211. https://doi.org/10.1038/nature08452 Gupta SK, Luo L, Yen L (2018) RNA-mediated gene fusion in mammalian cells. Proc Natl Acad Sci U S A 115:E12295–E12304. https://doi.org/10.1073/PNAS.1814704115/SUPPL_FILE/PNAS.1814704115.SM02.MOV Haas BJ, Dobin A, Li B, et al (2019) Accuracy assessment of fusion transcript detection via read-mapping and de novo fusion transcript assembly-based methods. Genome Biol 20:. https://doi.org/10.1186/S13059-019-1842-9 Hagel JM, Facchini PJ (2017) Tying the knot: occurrence and possible significance of gene fusions in plant metabolism and beyond. J Exp Bot 68:4029–4043. https://doi.org/10.1093/JXB/ERX152 Jain M, Bansal J, Rajkumar MS, Garg R (2022) An integrated transcriptome mapping the regulatory network of coding and long non-coding RNAs provides a genomics resource in chickpea. Communications Biology 2022 5:1 5:1–18. https://doi.org/10.1038/s42003-022-04083-4 Kibbe WA (2007) OligoCalc: an online oligonucleotide properties calculator. Nucleic Acids Res 35:. https://doi.org/10.1093/NAR/GKM234 Kim D, Paggi JM, Park C, et al (2019) Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nature Biotechnology 2019 37:8 37:907–915. https://doi.org/10.1038/s41587-019-0201-4 Kim D, Salzberg SL (2011) TopHat-Fusion: An algorithm for discovery of novel fusion transcripts. Genome Biol 12:1–15. https://doi.org/10.1186/GB-2011-12-8-R72/FIGURES/6 Kim M, Canio W, Kessler S, Sinha N (2001) Developmental changes due to long-distance movement of a homeobox fusion transcript in tomato. Science 293:287–289. https://doi.org/10.1126/SCIENCE.1059805 Koul B, Sharma K, Sehgal V, et al (2022) Chickpea (Cicer arietinum L.) Biology and Biotechnology: From Domestication to Biofortification and Biopharming. Plants 11:. https://doi.org/10.3390/PLANTS11212926 Kumar S, Vo AD, Qin F, Li H (2016) Comparative assessment of methods for the fusion transcripts detection from RNA-Seq data. Sci Rep 6:21597. https://doi.org/10.1038/srep21597 Latysheva NS, Babu MM (2016) Discovering and understanding oncogenic gene fusions through data intensive computational approaches. Nucleic Acids Res 44:4487–4503. https://doi.org/10.1093/NAR/GKW282 Li H, Wang J, Ma X, Sklar J (2009a) Gene fusions and RNA trans-splicing in normal and neoplastic human cells. Cell Cycle 8:218–222. https://doi.org/10.4161/CC.8.2.7358 Li J, Lee EJ, Chang L, Facchini PJ (2016) Genes encoding norcoclaurine synthase occur as tandem fusions in the Papaveraceae. Scientific Reports 2016 6:1 6:1–12. https://doi.org/10.1038/srep39256 Li X, Zhao L, Jiang H, Wang W (2009b) Short homologous sequences are strongly associated with the generation of chimeric RNAs in eukaryotes. J Mol Evol 68:56–65. https://doi.org/10.1007/S00239-008-9187-0 Livak KJ, Schmittgen TD (2001) Analysis of relative gene expression data using real-time quantitative PCR and the 2(-Delta Delta C(T)) Method. Methods 25:402–408. https://doi.org/10.1006/METH.2001.1262 Ma C, Shao M, Kingsford C (2018) SQUID: Transcriptomic structural variation detection from RNA-seq. Genome Biol 19:1–16. https://doi.org/10.1186/S13059-018-1421-5/FIGURES/7 Mukherjee S, Detroja R, Balamurali D, et al (2021) Computational analysis of sense-antisense chimeric transcripts reveals their potential regulatory features and the landscape of expression in human cells. NAR Genom Bioinform 3:. https://doi.org/10.1093/NARGAB/LQAB074 Muralla R, Chen E, Sweeney C, et al (2008) A Bifunctional Locus (BIO3-BIO1) Required for Biotin Biosynthesis in Arabidopsis. Plant Physiol 146:60. https://doi.org/10.1104/PP.107.107409 Nathawat BDS, Sharma OP, Kumari M, Shivran H (2024) Effect of Nutrients on Wilt in Chickpea. Legume Research 47:152–155. https://doi.org/10.18805/LR-4490 Pertea M, Kim D, Pertea GM, et al (2016) Transcript-level expression analysis of RNA-seq experiments with HISAT, StringTie and Ballgown. Nat Protoc 11:1650–1667. https://doi.org/10.1038/NPROT.2016.095 Qiao D, Yang C, Chen J, et al (2019) Comprehensive identification of the full-length transcripts and alternative splicing related to the secondary metabolism pathways in the tea plant (Camellia sinensis). Scientific Reports 2019 9:1 9:1–13. https://doi.org/10.1038/s41598-019-39286-z Qin F, Zhang Y, Liu J, Li H (2017) SLC45A3-ELK4 functions as a long non-coding chimeric RNA. Cancer Lett 404:53–61. https://doi.org/10.1016/J.CANLET.2017.07.007 Sherman BT, Hao M, Qiu J, et al (2022) DAVID: a web server for functional enrichment analysis and functional annotation of gene lists (2021 update). Nucleic Acids Res 50:W216–W221. https://doi.org/10.1093/NAR/GKAC194 Singh A, Zahra S, Das D, Kumar S (2019) AtFusionDB: a database of fusion transcripts in Arabidopsis thaliana. Database (Oxford) 2019:. https://doi.org/10.1093/DATABASE/BAY135 Sutton RE, Boothroyd JC (1986) Evidence for trans splicing in trypanosomes. Cell 47:527–535. https://doi.org/10.1016/0092-8674(86)90617-3 Tan C, Liu H, Ren J, et al (2019) Single-molecule real-time sequencing facilitates the analysis of transcripts and splice isoforms of anthers in Chinese cabbage (Brassica rapa L. ssp. pekinensis). BMC Plant Biol 19:. https://doi.org/10.1186/S12870-019-2133-Z Tang D, Chen M, Huang X, et al (2023) SRplot: A free online platform for data visualization and graphing. PLoS One 18:e0294236. https://doi.org/10.1371/JOURNAL.PONE.0294236 Turner NC, Colmer TD, Quealy J, et al (2013) Salinity tolerance and ion accumulation in chickpea (Cicer arietinum L.) subjected to salt stress. Plant Soil 365:347–361. https://doi.org/10.1007/S11104-012-1387-0/METRICS Varley KE, Gertz J, Roberts BS, et al (2014) Recurrent read-through fusion transcripts in breast cancer. Breast Cancer Res Treat 146:287. https://doi.org/10.1007/S10549-014-3019-2 Varshney RK, Song C, Saxena RK, et al (2013) Draft genome sequence of chickpea (Cicer arietinum) provides a resource for trait improvement. Nature Biotechnology 2013 31:3 31:240–246. https://doi.org/10.1038/nbt.2491 Wang K, Singh D, Zeng Z, et al (2010) MapSplice: Accurate mapping of RNA-seq reads for splice junction discovery. Nucleic Acids Res 38:e178. https://doi.org/10.1093/NAR/GKQ622 Winzer T, Kern M, King AJ, et al (2015) Plant science. Morphinan biosynthesis in opium poppy requires a P450-oxidoreductase fusion protein. Science 349:309–312. https://doi.org/10.1126/SCIENCE.AAB1852 Zhang G, Guo G, Hu X, et al (2010) Deep RNA sequencing at single base-pair resolution reveals high complexity of the rice transcriptome. Genome Res 20:646. https://doi.org/10.1101/GR.100677.109 Zhou Y, Lu Q, Zhang J, et al (2022) Genome-Wide Profiling of Alternative Splicing and Gene Fusion during Rice Black-Streaked Dwarf Virus Stress in Maize ( Zea mays L.). Genes (Basel) 13:. https://doi.org/10.3390/GENES13030456 Supplementary Files Supplementarymaterial.docx Supplementary Figures Fig. S1 Chromatogram of validated fusion transcripts from chickpea, where black lines mark the junction sites between the two fusion genes Fig. S2 Box plot depicting the relation between the predicted length of fusion transcripts and their parental genes Fig. S3 Inter-specific conserved fusion events. a) Venn diagram showing orthologous fusion events shared between Cicer arietinum and Arabidopsis thaliana . b-d) Gene Ontology (GO) enrichment analysis of orthologous fusion gene pairs, highlighting enriched categories in b) Molecular Function, c) Biological Process, and d) Cellular Component Fig. S4 Relative expression of validated fusion transcripts in Cicer arietinum across different tissues Fig. S5 Benchmark analysis of fusion detection tools using in-house generated RNA-seq data and publicly available RNA-seq data from chickpea. (a) Sensitivity of fusion detection tools evaluated for both datasets, calculated as Sensitivity (%) = (TP/TF) * 100, where TP represents true positives and TF is the total number of fusions. (b) F-measure for each tool, calculated as F measure = 2 * (Sensitivity * PPV) / (Sensitivity + PPV), where PPV (positive predictive value) = TP / (TP + FP). (c-d) Venn diagrams showing the overlapping fusion transcripts identified by different fusion detection tools in the in-house generated RNA-seq data (c) and the publicly available RNA-seq data, (d) highlighting common fusions across the tools. The performance metrics (sensitivity, specificity, and F-measure) were calculated based on the identification of validated fusion transcripts (true fusions) across all tools. Supplementary Tables Table S1. Summary of RNA-Seq reads and mapping statistics. Table S2. Summary of the number of fusion transcripts detected in each sample by three fusion detection methods and sequencing depth of each sample. Table S3. List of fusion transcripts detected in RNA-Seq samples. Table S4. Information regarding Short homologous sequences (SHSs) detected at the junction of fusion transcript. Table S5. List of fusion transcripts detected in Long-read RNA-Seq data. Table S6. List of RNA-Seq samples used to study intra-specific diversity of fusion transcripts. Table S7. List of primers used for qRT-PCR of fusion transcripts and their parental genes. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6540588","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":451013253,"identity":"b977cb7a-8fda-43f5-993f-abb223e5a57e","order_by":0,"name":"Fiza Hamid","email":"","orcid":"","institution":"NIPGR: National Institute of Plant Genome Research","correspondingAuthor":false,"prefix":"","firstName":"Fiza","middleName":"","lastName":"Hamid","suffix":""},{"id":451013254,"identity":"711b8bc7-af67-4866-b10e-ebd69dda6d76","order_by":1,"name":"Shafaque Zahra","email":"","orcid":"","institution":"University of Virginia Health System: UVA Health","correspondingAuthor":false,"prefix":"","firstName":"Shafaque","middleName":"","lastName":"Zahra","suffix":""},{"id":451013255,"identity":"fb4ee77d-1791-4d94-aca4-dc7f6b8a5068","order_by":2,"name":"Shailesh Kumar","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA6UlEQVRIiWNgGAWjYDACCQaGAw8YGBL4GRgboEJwBh4tCUAtkg2kaGEAaTE4QKy75Gc3PzyQUFGXZ3y7ue3h1z0M8vwNzG0P8GkxuHPM4EDCmcPFZncOthvLPGMwnHGAsd0ArxYJoJMS2w4kbruR2CYtcYCBcQMDY5sEXofNSP9wIPFfXeLmGRAt9gS1MNzIAdrSwJy4QSKxTfLDAYZEgloMbuQUHEg4djhxxo3EdmOGAxLJMw4TdtjmDx9q6hL7Z6Q/e/jjgI1tf3v7M/wOQwJszDygaGImVj1IC+MPElSPglEwCkbByAEAYPNSPdhmvIAAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-1872-9903","institution":"NIPGR: National Institute of Plant Genome Research","correspondingAuthor":true,"prefix":"","firstName":"Shailesh","middleName":"","lastName":"Kumar","suffix":""}],"badges":[],"createdAt":"2025-04-27 13:09:47","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6540588/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6540588/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":82180514,"identity":"965551ea-d9a0-4054-82fa-c8d0c2b6c005","added_by":"auto","created_at":"2025-05-07 11:46:48","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1434659,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of fusion transcripts in \u003cem\u003eCicer arietinum\u003c/em\u003e. a) The tissues included in this study and the number of fusions detected are indicated within parentheses following the tissue name. Venn diagram showing the overlap in fusion transcripts across different tissues and stress conditions. b) Circos plot depicting the chromosomal distribution of parental genes involved in fusion events, with each connecting link representing a fusion event. c) Correlation between the total number of genes per chromosome and the number of genes involved in fusion formation on that chromosome\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-6540588/v1/8ee7634094d6ba1bfcc58cf5.png"},{"id":82179946,"identity":"4a4cf8e8-a9e3-47a3-bb72-d43594cc567e","added_by":"auto","created_at":"2025-05-07 11:38:48","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":3671201,"visible":true,"origin":"","legend":"\u003cp\u003eSpecificity of fusion events. a) Bar plot showing the frequency of involvement of a gene in multiple fusion events. An example of a gene involved in multiple fusion events is highlighted. b) Bar plot showing the number of isoforms generated from a fusion gene pair. An example of a fusion gene pair producing two isoforms is presented\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-6540588/v1/1bd2524327ac300864b27796.png"},{"id":82179948,"identity":"6b141ad5-05de-408e-9ed0-e625619646dc","added_by":"auto","created_at":"2025-05-07 11:38:48","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":4977101,"visible":true,"origin":"","legend":"\u003cp\u003eJunction site sequence analysis of fusion transcript. a) The proportion of splice site patterns observed at fusion breakpoints is displayed, with an example fusion transcript demonstrating the canonical splice pattern (GT-AG) at the junction.\u003cbr\u003e\nb) The contribution of different mechanisms to fusion generation is depicted, including trans-splicing and short homologous sequence (SHS) mediated fusion formation. An example fusion transcript displaying both canonical splice sites and SHS at the fusion junction is highlighted in yellow\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-6540588/v1/97092fc48d0fca51d79dd2e5.png"},{"id":82181281,"identity":"8b5cac6d-3eee-4a38-8178-70f829a95ab1","added_by":"auto","created_at":"2025-05-07 11:54:48","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":965893,"visible":true,"origin":"","legend":"\u003cp\u003eFunctional categorization of parental genes involved in fusion. a) Correlation between expression levels of 5’ and 3’ parental genes for each fusion event. Expression levels are presented as log (TPM + 1), with each point representing a fusion event. The R-value represents the Pearson correlation coefficient. b) Scatterplot showing expression levels [log 2(TPM + 1)] of parental genes in samples with (F/P) and without the corresponding fusions (F/A). Colored dots indicate significant changes in parental gene expression (fold change ≥ 2): grey for no change, blue for downregulation, and red for upregulation. c) Gene Ontology and KEGG pathway analysis of parental genes involved in fusion transcript formation. These plots were generated using SRplot (Tang et al. 2023)\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-6540588/v1/9b5aea06abd0a3c3aea13358.png"},{"id":82179952,"identity":"06b336be-74c3-45de-a6fe-5de76b009d13","added_by":"auto","created_at":"2025-05-07 11:38:49","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":301725,"visible":true,"origin":"","legend":"\u003cp\u003eFusion Event Profiles Across Different Chickpea Genotypes. a) The number of fusion events identified in various chickpea genotypes. b) Hierarchical clustering of 17 chickpea genotypes based on the presence or absence of fusion events identified in the ICC4958 genotype, with presence marked in blue and absence in yellow\u003c/p\u003e","description":"","filename":"Figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-6540588/v1/0d860b34a4e1b7c8824640a2.png"},{"id":82179949,"identity":"d2ca9cc3-847e-45f3-9a8a-7aabfdbd8147","added_by":"auto","created_at":"2025-05-07 11:38:48","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":1361864,"visible":true,"origin":"","legend":"\u003cp\u003eRelative expression of validated fusion transcripts and its precursors in \u003cem\u003eCicer arietinum\u003c/em\u003e under drought and salt stress conditions\u003c/p\u003e","description":"","filename":"Figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-6540588/v1/70eb75576f479fa413fa1e5e.png"},{"id":82663699,"identity":"54f3a99e-3f03-46ec-99cd-8737a8fb795e","added_by":"auto","created_at":"2025-05-13 23:29:44","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":11421678,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6540588/v1/927217d3-657a-4e82-9341-4da7878abb6f.pdf"},{"id":82180515,"identity":"2bf03de2-c4fa-4d24-8500-1fd5b37450d7","added_by":"auto","created_at":"2025-05-07 11:46:48","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1103254,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupplementary Figures\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFig. S1\u003c/strong\u003e Chromatogram of validated fusion transcripts from chickpea, where black lines mark the junction sites between the two fusion genes\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFig. S2\u003c/strong\u003e Box plot depicting the relation between the predicted length of fusion transcripts and their parental genes\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFig. S3 \u003c/strong\u003eInter-specific conserved fusion events. a) Venn diagram showing orthologous fusion events shared between \u003cem\u003eCicer arietinum\u003c/em\u003e and \u003cem\u003eArabidopsis thaliana\u003c/em\u003e. b-d) Gene Ontology (GO) enrichment analysis of orthologous fusion gene pairs, highlighting enriched categories in b) Molecular Function, c) Biological Process, and d) Cellular Component\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFig. S4\u003c/strong\u003e Relative expression of validated fusion transcripts in \u003cem\u003eCicer arietinum\u003c/em\u003e across different tissues\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFig. S5\u003c/strong\u003e Benchmark analysis of fusion detection tools using in-house generated RNA-seq data and publicly available RNA-seq data from chickpea. (a) Sensitivity of fusion detection tools evaluated for both datasets, calculated as Sensitivity (%) = (TP/TF) * 100, where TP represents true positives and TF is the total number of fusions. (b) F-measure for each tool, calculated as F measure = 2 * (Sensitivity * PPV) / (Sensitivity + PPV), where PPV (positive predictive value) = TP / (TP + FP). (c-d) Venn diagrams showing the overlapping fusion transcripts identified by different fusion detection tools in the in-house generated RNA-seq data (c) and the publicly available RNA-seq data, (d) highlighting common fusions across the tools. The performance metrics (sensitivity, specificity, and F-measure) were calculated based on the identification of validated fusion transcripts (true fusions) across all tools.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Tables\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable S1.\u003c/strong\u003e Summary of RNA-Seq reads and mapping statistics.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable S2.\u003c/strong\u003e Summary of the number of fusion transcripts detected in each sample by three fusion detection methods and sequencing depth of each sample.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable S3.\u003c/strong\u003e List of fusion transcripts detected in RNA-Seq samples.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable S4.\u003c/strong\u003e Information regarding Short homologous sequences (SHSs) detected at the junction of fusion transcript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable S5.\u003c/strong\u003e List of fusion transcripts detected in Long-read RNA-Seq data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable S6.\u003c/strong\u003e List of RNA-Seq samples used to study intra-specific diversity of fusion transcripts.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable S7.\u003c/strong\u003e List of primers used for qRT-PCR of fusion transcripts and their parental genes.\u003c/p\u003e","description":"","filename":"Supplementarymaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-6540588/v1/edf3265f33b14741cdf8d44d.docx"}],"financialInterests":"","formattedTitle":"Transcriptome analysis revealed stress responsive fusion transcripts in Chickpea (Cicer arietinum)","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMature mRNA molecules are conventionally formed through transcription and post-transcriptional modifications, where introns are excised, and stability is enhanced. Traditionally, mRNA was thought to originate solely from the alternative splicing of a single gene. However, advancements in sequencing technologies have uncovered the existence of novel transcripts, such as fusion transcripts (FTs), which arise by the joining of mRNA molecules derived from different genes (Gingeras \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2009\u003c/span\u003e; Gupta et al. \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). FTs can be generated at the DNA level through genomic rearrangements such as translocation, deletion, duplication, or inversion and give rise to a fused gene (Li et al. \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2009a\u003c/span\u003e; Annala et al. \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2013\u003c/span\u003e); or at the RNA level through trans-splicing (Sutton and Boothroyd \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e1986\u003c/span\u003e) or read-through transcription (Varley et al. \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2014\u003c/span\u003e), resulting in an increase in transcriptome complexity without a corresponding increase in the gene number. While fusion transcripts are found in both unicellular and multicellular organisms, they were previously considered rare in nature and occasionally dismissed as transcriptional artifacts. Up to now, extensive research has characterized the cellular functions of these transcripts in cancer (Latysheva and Babu \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Dupain et al. \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Recent studies indicate that in addition to their role in oncogenesis, FTs have also been reported under normal physiological conditions (Babiceanu et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Chwalenia et al. \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2017\u003c/span\u003e), and may act as potential regulators of their parental mRNA (Mukherjee et al. \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). However, their presence and functional significance in plants remain largely unexplored.\u003c/p\u003e \u003cp\u003eIn past decades, efficient big data analysis has facilitated the discovery of fusion transcripts in many plants such as \u003cem\u003eTrifolium pratense\u003c/em\u003e (Chao et al. \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2018\u003c/span\u003e), \u003cem\u003eCamellia sinensis\u003c/em\u003e (Qiao et al. \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), \u003cem\u003eOryza sativa\u003c/em\u003e (Zhang et al. \u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e2010\u003c/span\u003e), \u003cem\u003eArabidopsis thaliana\u003c/em\u003e (Singh et al. \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), \u003cem\u003eBrassica rapa\u003c/em\u003e (Tan et al. \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), and \u003cem\u003eZea mays\u003c/em\u003e (Zhou et al. \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). These plant-specific fusion transcripts add transcriptome complexity and introduce novel functions, which may be related to stress responses (Zhou et al. \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), metabolic pathways (Hagel and Facchini \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2017\u003c/span\u003e), or phenotypic traits (Chen et al. \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). For instance, maize exhibits various fusion events in response to viral infection, involving proteins like nodulin, flavone synthase, and cation-transporting ATPase (Zhou et al. \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). In tomatoes, the fusion transcript \u003cem\u003ePFP-LeT6\u003c/em\u003e is involved in leaf patterning (Kim et al. \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2001\u003c/span\u003e), while the \u003cem\u003eGN2\u003c/em\u003e chimeric gene in rice, controls plant height, heading date, and grain number (Chen et al. \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). In \u003cem\u003eArabidopsis thaliana\u003c/em\u003e, a protein derived from the fusion of glutamine synthase and nodulin domains regulates root morphogenesis and flagellin-triggered signalling (Doskočilov\u0026aacute; et al. \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). Studies across species also suggest that fusion events like domain shuffling and DNA fusion enhance the catalytic efficiency of an enzyme (Farrow et al. \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Li et al. \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Hagel and Facchini \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). These reports suggest that fusion events result in novel sequences with novel functions either as non-coding RNA or fusion proteins (Winzer et al. \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Qin et al. \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2017\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eChickpea (\u003cem\u003eCicer L\u003c/em\u003e.) ranks as the third most cultivated legume crop globally, following common bean and pea (Nathawat et al. \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Cultivated in over 50 countries, the Indian subcontinent leads in production, contributing approximately 70% of the global output (Koul et al. \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). In addition to its economic significance, chickpea is highly esteemed for its exceptional nutritional content, particularly its rich protein and carbohydrate content. Currently, chickpeas are grown on 15\u0026nbsp;million hectares worldwide, producing 15.9\u0026nbsp;million tons (FAO, 2023). However, this yield remains significantly below the crop's potential under optimal conditions. This gap is largely attributed to various biotic and abiotic stresses that hinder productivity. Abiotic stresses, such as salinity and drought, are key factors contributing to yield losses. Research has shown that chickpeas are particularly sensitive to salinity compared to other crops, making it a critical factor limiting yields (Flowers et al. \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2010\u003c/span\u003e; Turner et al. \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2013\u003c/span\u003e). The growing availability of high-throughput transcriptomic data offers valuable insights for tackling these stress-related challenges.\u003c/p\u003e \u003cp\u003eDespite advancements in sequencing technologies, the exploration of transcriptome diversity in chickpea, particularly due to fusion events, remains limited (Chitkara et al. \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Recently, our research group (Chitkara et al. \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) conducted a comprehensive profiling of fusion transcripts using an extensive collection of publicly available RNA-Seq datasets, encompassing diverse tissues and a wide range of conditions. Interestingly, this study revealed that the majority of identified fusion transcripts (~\u0026thinsp;74%) were detected in only a single transcriptomic sample. One possible explanation could be the inherent variability introduced by such a large and heterogeneous dataset. Alternatively, many of these uniquely detected fusion transcripts might represent random artifacts rather than true biological events. Moreover, the validation rate was reported to be generally low in \u003cem\u003eArabidopsis thaliana\u003c/em\u003e and \u003cem\u003eOryza sativa\u003c/em\u003e, which could be attributed to high false-positive prediction rates or the inherently low expression levels of fusion transcripts. Notably, no experimental validation was performed for the fusion transcripts identified in \u003cem\u003eCicer arietinum\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eThe current study addresses this gap by utilizing in-house generated RNA-Seq datasets of chickpea derived from specific abiotic stress conditions and tissue types, enabling the identification of high-confidence fusion transcripts by applying stringent filtering criteria. This enhances the reproducibility, which in turn strengthens the confidence in these fusion events as biologically relevant and potentially regulatory molecules, rather than technical noise. Here, we identified stress-responsive and conserved fusion transcripts in \u003cem\u003eCicer arietinum\u003c/em\u003e using a combination of second- and third-generation RNA-Seq datasets, along with experimental validation. Together, our data provide the genome-wide profiling of fusion events in chickpea and reveal their widespread occurrence and potential functions. Additionally, a benchmarking analysis was performed to evaluate the performance of different fusion detection tools, enabling the identification of the most effective methods for accurate fusion detection in Chickpea. This work thus complements our earlier large-scale study by providing deeper biological interpretation and functional insights into fusion transcript dynamics in chickpea. Ultimately, the investigation of these transcripts could uncover novel mechanisms of stress adaptation and enhance the overall understanding of gene expression dynamics in legumes.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003ePlant material, growth conditions, and stress treatment\u003c/h2\u003e \u003cp\u003eChickpea (Genotype ICC4958) seeds were cultivated by using the established protocol (Garg et al. \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2010\u003c/span\u003e). Seedlings were grown in plastic pots filled with a sterilized mix of agro-peat and vermiculite in a 1:1 ratio, maintained at 22\u0026deg;C with a 14-hour light cycle in a controlled growth chamber. Samples were collected at various developmental stages, including stem and leaves from the seedling stage, and buds, flowers, and pods during the flowering stage. To induce drought stress, 21-day-old seedlings were removed from the pots and placed on tissue paper for five hours. For salinity stress, seedlings were immersed in a beaker filled with a 150 mM NaCl solution at 22\u0026deg;C. Whole seedlings were collected after 5 hours of treatment, and from three independent biological replicates. These samples were immediately flash-frozen in liquid nitrogen and stored at -80\u0026deg;C until further analysis.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eTotal RNA isolation, cDNA library preparation, and sequencing\u003c/h3\u003e\n\u003cp\u003eTotal RNA was extracted from 100 mg of tissue using the RNeasy Plant Mini Kit (Qiagen). The quantity and quality of the RNA were assessed with a NanoDrop Spectrophotometer (NanoDrop Technologies), ensuring that only samples with a 260/280 ratio between 1.9 and 2.1, and a 260/230 ratio between 2.0 and 2.4, were selected for Illumina sequencing. A total of 26 paired-end RNA-Seq libraries were prepared and sequenced. The quality of the reads was assessed with FastQC (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.bioinformatics.babraham.ac.uk/projects/fastqc/\u003c/span\u003e\u003cspan address=\"https://www.bioinformatics.babraham.ac.uk/projects/fastqc/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), and adapter trimming and filtering of low-quality reads were performed using fastp (Chen et al. \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). High-quality reads were then aligned to the reference genome (ASM33114v1) using HISAT2-2.2.1 (Kim et al. \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2019\u003c/span\u003e).\u003c/p\u003e\n\u003ch3\u003eComputational prediction of fusion transcripts\u003c/h3\u003e\n\u003cp\u003eFusion transcripts were predicted using three different fusion detection tools viz. FusionMap (Ge et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2011\u003c/span\u003e), STAR-Fusion (Haas et al. \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), and MapSplice (Wang et al. \u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e2010\u003c/span\u003e). These tools have different fusion detection strategies and filtering criteria, which help in identifying a broad range of fusion events. These tools were selected based on available benchmarking publications where STAR-Fusion was ranked as the best tool in terms of their high sensitivity, accuracy, and execution time (Haas et al. \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2019\u003c/span\u003e); and MapSplice and FusionMap show good sensitivity for fusion detection (Kumar et al. \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2016\u003c/span\u003e).\u003c/p\u003e\n\u003ch3\u003eGene Ontology and KEGG pathway analysis\u003c/h3\u003e\n\u003cp\u003eFor all identified fusion transcripts, gene ontology (GO) and KEGG pathway analysis were performed using DAVID (Sherman et al. \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) with a significance threshold of p\u0026thinsp;\u0026lt;\u0026thinsp;0.05. These analyses help to identify biological processes, cellular components, and pathways potentially impacted by fusion events, shedding light on their functional relevance.\u003c/p\u003e\n\u003ch3\u003eExpression analysis of fusion transcripts\u003c/h3\u003e\n\u003cp\u003eThe expression level of genes involved in fusion formation was analysed based on Transcripts Per Million (TPM) using StringTie (Pertea et al. \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2016\u003c/span\u003e) with the default parameters. To study the impact of fusion formation on the expression of their parental genes, we categorized the samples into two groups for each fusion transcript: (i) fusion-present (F/P) - samples in which the corresponding fusion transcript was detected, and (ii) fusion-absent (F/A) - samples lacking the respective fusion. The expression levels of each parental gene were compared between each sample of the two groups. Genes exhibiting\u0026thinsp;\u0026ge;\u0026thinsp;2-fold change in expression in F/P samples compared to F/A samples were considered upregulated or downregulated.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eValidation using Long-Read RNA-Seq data\u003c/h2\u003e \u003cp\u003eFor further validation of fusion transcripts, PacBio long-read RNA-Seq data (PRJNA613159, Jain et al. \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) for \u003cem\u003eCicer arietinum\u003c/em\u003e (ICC4958), were downloaded from the NCBI Sequence Read Archive (SRA). The 200 bp junction sequences of the identified fusion transcripts-comprising 100 bp upstream of the 5\u0026prime; parental gene breakpoint and 100 bp downstream of the 3\u0026prime; parental gene breakpoint-were searched against the long-read RNA-Seq data using BLASTn. Hits with greater than 80% sequence identity and alignment length exceeding 150 bp were considered significant. This approach confirmed the presence of predicted fusion transcripts within the long-read data, enhancing confidence in the identified fusion events. Additionally, it offered insights into the putative lengths of the fusion transcripts.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eIdentification of conserved fusion events\u003c/h3\u003e\n\u003cp\u003eTo identify intra-specific conserved fusions, 103 raw transcriptome data files from 17 \u003cem\u003eCicer arietinum\u003c/em\u003e genotypes were downloaded from the NCBI-SRA (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/sra\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/sra\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). List of RNA-Seq samples from different genotypes used in this study are listed in Table S6. Fusion events were then identified across all genotypes using three fusion detection tools. A binary analysis was conducted to determine the presence or absence of each fusion gene pair in the respective genotypes. For inter-specific conserved fusion detection, fusion genes previously reported in \u003cem\u003eArabidopsis thaliana\u003c/em\u003e (via AtFusionDB: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.nipgr.ac.in/AtFusionDB\u003c/span\u003e\u003cspan address=\"http://www.nipgr.ac.in/AtFusionDB\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) were utilized to identify their homologous fusion gene pairs in \u003cem\u003eCicer arietinum\u003c/em\u003e using the OrthoFinder tool.\u003c/p\u003e\n\u003ch3\u003eExperimental validation of fusion transcripts\u003c/h3\u003e\n\u003cp\u003eTo validate fusion transcripts, primers were designed for the fusion transcript sequences taken from the genome 200 bp upstream and 200 bp downstream of the breakpoint, using the OligoCalc (Kibbe \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2007\u003c/span\u003e) and primer BLAST. All primers used for validation are listed in Supplementary Table S7. The cDNA synthesis was carried out with 2\u0026micro;g of RNA using the Verso cDNA synthesis kit (ThermoScientific\u0026trade;). Following RT-PCR and gel electrophoresis, DNA bands were extracted and purified using the GenElute\u0026trade; Gel Extraction kit (Sigma-Aldrich) and sent for Sanger sequencing at the NIPGR DNA sequencing facility.\u003c/p\u003e \u003cp\u003eTo quantify the expression of fusion transcripts under drought and salt stress, and across different tissues, quantitative Real-Time PCR was performed. EF1-α was used as an endogenous control gene in this experiment. The real-time PCR reaction mix (10 \u0026micro;l) consisted of 5 \u0026micro;l of 2X SYBR Green Master Mix (Applied Biosystems\u0026trade;), 10 \u0026micro;M of each primer, and 100 ng of cDNA template. PCR amplification was performed using an Applied Biosystems\u0026trade; qPCR system with thermal cycling conditions of initial denaturation at 95\u0026deg;C for 2 minutes, followed by 40 cycles of 95\u0026deg;C for 15 seconds and 60\u0026deg;C for 1 minute. Each PCR reaction included three biological replicates and three technical replicates. The relative expression levels of fusion transcripts were calculated by using the 2^-ΔΔCt method (Livak and Schmittgen \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2001\u003c/span\u003e). The experimental data were presented as the mean with standard deviation (mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD) derived from three independent biological replicates and three technical replicates. Statistical analysis was conducted by comparing the means of control and stressed plants using one-way analysis of variance (ANOVA) followed by Student\u0026rsquo;s t-test, with a significance level set at P\u0026thinsp;\u0026lt;\u0026thinsp;0.05. P-values less than 0.05 were considered statistically significant.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eBenchmark analysis of fusion detection tools\u003c/h2\u003e \u003cp\u003eTo benchmark the performance of various fusion detection tools for \u003cem\u003eCicer\u003c/em\u003e, we evaluated STAR-Fusion (Haas et al. \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), SQUID (Ma et al. \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2018\u003c/span\u003e), MapSplice (Wang et al. \u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e2010\u003c/span\u003e), Tophat-Fusion (Kim and Salzberg \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2011\u003c/span\u003e), and FusionMap (Ge et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). The performance of each tool was assessed using both public dataset (PRJNA288321, Garg et al. \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2016\u003c/span\u003e), and our inhouse dataset. For both datasets, we then searched the validated fusion transcripts (true fusions) in the results generated by each tool. The accuracy of fusion prediction was calculated following the method described by Kumar (2016) (Kumar et al. \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2016\u003c/span\u003e), as given below.\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:Sensitivity\\:\\left(\\%\\right)=\\left(\\frac{TP}{TF}\\right)*100$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:Positive\\:Predictive\\:Value\\:\\left(PPV\\right)=\\frac{TP}{TP+FP}*100$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:F\\:measure=2*\\left(S*PPV\\right)/(S/PPV)$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eS: Sensitivity, TP: True positive, TF: Total fusions, FP: False positive\u003c/h2\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003e \u003cb\u003eIdentification of fusion transcripts in\u003c/b\u003e \u003cb\u003eCicer arietinum\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo comprehensively investigate fusion events within the chickpea transcriptome, high-throughput paired-end RNA-Seq data were generated using the Illumina sequencing platform. The samples were sequenced from poly(A)-enriched RNAs extracted from five distinct organs, viz. buds, leaves, pods, flowers, and stem (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea), along with two abiotic stress conditions (i.e., drought and salinity). After adapter trimming and quality-check, high-quality reads were mapped onto the CDC Frontier Reference Genome (ASM33114v1) (Varshney et al. \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2013\u003c/span\u003e) using HISAT2 (v2.2.1) (Kim et al. \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), and it was found that \u0026gt;\u0026thinsp;95% of the reads mapped onto the genome from each paired-end RNA-Seq sample (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). In total, 122.47 GB of data was obtained from all samples, representing about 230-fold of the chickpea genome size, and around 852.8\u0026nbsp;million high-quality reads were generated.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThree fusion detection tools were employed for genome-wide identification of fusion transcripts in 26 RNA-Seq datasets of different tissues and two abiotic stress conditions (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). The number of overlapping fusion transcripts detected between different tissues and stress samples are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea. The total number of fusion transcripts identified by FusionMap (Ge et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2011\u003c/span\u003e), STAR-Fusion (Haas et al. \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), and MapSplice (Wang et al. \u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e2010\u003c/span\u003e) are 496, 721 and 533 respectively (Table S2). However, the number of unique fusion transcripts identified is 95, 109, and 140, respectively. A total of 1750 fusion transcripts were identified in our datasets, constituting 328 unique fusions (Table S3), which were derived from 269 unique parental gene pairs. In total, 423 genes were involved in fusion formation, accounting for 1.4% (423/30,344) of the total annotated genes in \u003cem\u003eCicer arietinum\u003c/em\u003e. This proportion is similar to the percentage of genes involved in fusion events in other plant species, such as \u003cem\u003eArabidopsis\u003c/em\u003e (1%), soybean (1.7%), and rice (2.7%) (Cong et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). To verify the reliability of the predicted fusion transcripts, we randomly selected and validated 10 fusion transcripts in multiple biological replicates by PCR and Sanger sequencing (Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). The experimentally validated junction sequences precisely matched the in-silico predicted breakpoint sites, confirming the accuracy of the fusion detection approach.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eFeatures of fusion transcripts in\u003c/b\u003e \u003cb\u003eCicer arietinum\u003c/b\u003e\u003c/p\u003e \u003cp\u003eInitial analysis reveals that interchromosomal (76%) fusions are more prevalent than intrachromosomal (24%) fusions (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). This indicates that, in \u003cem\u003eCicer arietinum\u003c/em\u003e, the formation of fusion transcripts is not strictly governed by the linear proximity of parental genes, in contrast to \u003cem\u003eArabidopsis\u003c/em\u003e, where intrachromosomal fusions between adjacent genes are more common (Cong et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). This suggests that, beyond linear arrangement of genes, spatial organization may play a role in facilitating fusion events in chickpea and thus underscores the value of further investigation into the spatial organization of the genome. A positive correlation was observed between the number of genes participating in fusion events located at a particular chromosome, and the total number of genes on that chromosome, suggesting there is no chromosome-level preference in fusion formation (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ec). Out of the 423 genes involved in fusion events, 306 participate in only one unique fusion event, while 70 genes form fusions with two different partner genes. This indicates that fusion events are highly specific, with individual genes preferentially associating with particular partners rather than engaging in multiple fusion events. However, the mechanisms behind partner selection in fusion events remain unknown.\u003c/p\u003e \u003cp\u003eFrom the validated fusion pairs, all genes were found to fuse with only one partner gene, except LOC101495229, which formed fusions with both LOC101508351 and LOC101506473 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003ea). The number of fusion isoforms generated by a fusion gene pair was inversely related to the number of such gene pairs detected, highlighting the specificity of junction sites within fusion genes. Among the validated fusions, all gene pairs had a single junction site, except the LOC101509445_LOC101509981 fusion, which exhibited two isoforms. In both isoforms, the breakpoint in the LOC101509445 gene was identical, whereas LOC101509981 contributed two distinct breakpoints, each located at the exon boundaries of two different alternatively spliced transcripts of the LOC101509981 gene. This is a read-through fusion derived from two adjacent genes, where LOC101509445 is present on the reverse strand and LOC101509981 is present on the forward strand of chromosome 7. In this fusion, the 5\u0026rsquo; gene belongs to the ABC transporter I family member 1, and the 3\u0026rsquo; gene encodes for tRNA pseudouridine synthase. The two fusion events occur at the existing exon boundaries, hence producing in-frame fusions (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo determine the fusion site within the parental genes, we analysed the location of the breakpoint, whether it exists on the exon boundaries of both parents or one or none. It was observed that none type (57%) of fusion transcripts, where breakpoints were present within the exons or in the UTR region, were most abundant, whereas fusion transcripts having breakpoints at the exon borders of either one or both parental genes were 25.3% and 17.6%, respectively (Table S3).\u003c/p\u003e \u003cp\u003eTrans-splicing is a known mechanism for fusion transcript formation (Li et al. \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2009a\u003c/span\u003e). Junction pattern analysis showed that a significant portion of FTs exhibited canonical splice patterns: GT-AG (69%), however, non-canonical splicing patterns were also found such as GT-AT (4%), AT-AC (3%), CT-GC (3%), etc., implying apart from splicing, other unknown mechanisms may also contribute to fusion formation (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e3\u003c/span\u003ea). Another reported mechanism of fusion formation is transcriptional slippage mediated by the presence of short homologous sequences (SHSs) (Li et al. \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2009b\u003c/span\u003e) at the junction but less than 5% of total fusions exhibit SHSs at the breakpoints in \u003cem\u003eArabidopsis\u003c/em\u003e, soybean, rice, and maize (Cong et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Similar results were observed in chickpea; 4% of total fusions showed SHSs at the junction, and 5.7% had both canonical splice sites and SHSs at their junction (Table S4). All the experimentally validated fusions exhibited canonical splice sites at their junctions (e.g., LOC101515613_LOC113787786; Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e3\u003c/span\u003ea), except for LOC101494819_LOC101493433, which showed a non-canonical GA-AG splice site. Additionally, three of the fusions also displayed short homologous sequences (SHSs) at the junctions, e.g., LOC101500313_LOC101505021 (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e3\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eProperties of genes involved in fusion formation\u003c/h2\u003e \u003cp\u003eAnalysis of the transcriptome data showed a significant positive correlation (R\u0026thinsp;=\u0026thinsp;0.53) between the expression levels of gene pairs involved in fusion formation, suggesting that these genes are linked in terms of expression (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003ea). Further to evaluate the impact of fusion formation on their parental genes, the expression of fusion parental genes was compared between samples with and without fusion transcripts and it was observed that most of the fusion transcripts (~\u0026thinsp;75%) have no significant effect on the expression of their parental genes, suggesting that fusion formation may be independently regulated. However, a subset of parental genes exhibited differential expression in the presence of fusion transcripts, with approximately 21% showing upregulation and 4% showing downregulation (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFusion transcripts may retain or be associated with the biological functions of their parental genes. To gain insights into the potential functions of the fusion transcripts, we employed gene ontology (GO) and KEGG pathway analyses of parental genes involved in fusion formation. The most enriched biological GO terms were translation, photosynthesis, response to light stimulus, and rRNA processing (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003ec). The most enriched genes in the molecular process analysis showed that fusion transcripts parental genes are involved in the structural constituent of the ribosome and in binding such as RNA binding and chlorophyll-binding and in enzymatic activities such as ATP hydrolysis, carboxylase, and dehydrogenase. The most enriched cellular component GO terms were cytosol, ribosome, and chloroplast. Pathway enrichment analysis showed that these genes are related to the Ribosome, carbon metabolism, and photosynthesis. Together, we can conclude that fusion transcripts originate from genes with diverse functions that are distributed across various cellular compartments, and also show important enzymatic and binding activities. Notably, photosynthesis and ribosome-associated genes are majorly enriched among fusion precursor genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003ec).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eValidation of fusion transcripts from Long-read RNA-Seq data\u003c/h2\u003e \u003cp\u003eIllumina sequencing generates short reads that require assembly to reconstruct full-length transcripts. However, this process can introduce errors, particularly in regions with complex splicing patterns or novel splice sites. To overcome these challenges, long-read RNA sequencing technologies, such as PacBio, offer a significant advantage by capturing full-length transcripts in a single read, eliminating the need for assembly. To detect the full length of the identified FTs, publicly available long-read RNA-Seq data of \u003cem\u003eCicer arietinum\u003c/em\u003e were utilized. To identify the full-length sequence of FTs, BLASTn search of 200 bp fusion junction sequences was conducted against the PacBio long-read data. Through this approach, we found 95 fusion transcripts in PacBio data, and their full length was predicted (Table S5). The mean length of identified fusions was around 0.8 kb, with 70% of them having a length less than the mean length. We further compared the length of fusion transcripts and their parental genes and observed that the mean length of fusion transcripts was shorter than that of their parental genes (Fig. S2). This indicates that only specific portions of the parental genes are involved in fusion formation.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eIntra and Inter-specific conserved fusion events\u003c/h2\u003e \u003cp\u003eTo explore intra-specific conserved fusion events, RNA-Seq data of 17 different chickpea genotypes were analysed. The number of fusion genes identified varied significantly across genotypes, ranging from 41 to 1031 (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e5\u003c/span\u003ea). This variation is likely attributed to differences in the number of samples analysed for each genotype and the sequencing depth of the data. A fusion event was considered in a particular genotype if it was detected even in a single sample. After removing redundancy, interchromosomal fusion transcripts were found to be more prevalent than intrachromosomal fusion transcripts. Fusion genes identified from the ICC4958 genotype (269 fusion gene pairs) were searched across different genotypes (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e5\u003c/span\u003eb). Among these, 140 fusion gene pairs were conserved across multiple genotypes, while the rest were specific to the ICC4958. Notably, 53 fusion gene pairs were detected in 10 or more genotypes, suggesting their potential biological significance and stability within the chickpea population.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo identify conserved FTs between Chickpea and \u003cem\u003eArabidopsis\u003c/em\u003e, FTs reported in AtFusionDB (Singh et al. \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) were used. Among the 269 fusion genes identified in chickpea, 19 showed homology with 58 fusion events reported in \u003cem\u003eArabidopsis\u003c/em\u003e within AtFusionDB (Fig. S3a). Gene ontology analysis of these homologous genes revealed their involvement in essential biological processes, including metabolism and environmental responses (Fig. S3b-d). The presence of these fusion transcripts may play a critical role in enhancing biological functions linked to these processes.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eValidation of fusion transcripts by RT-PCR and Sanger sequencing\u003c/h2\u003e \u003cp\u003eTo quantify the expression of validated fusion transcripts under different abiotic stress and tissues, quantitative real-time PCR was performed. The notable findings from expression analysis reveal that fusion transcripts are expressed at a very low level, which leads to their low detection rate. Under different abiotic stress, three of the validated fusions (LOC101506206_LOC101493600, LOC101500131_LOC101505021, and LOC101495229_LOC101508351) showed differential expression which implies that these fusions might play a crucial role during stress response (Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e6\u003c/span\u003e). Expression analysis of stress-responsive fusion precursor genes indicated that the relative fold change in expression of fusion transcripts was higher than that of their parental genes under stress conditions. This suggests that fusion transcript formation is not solely controlled by its precursor gene expression, but may involve additional regulatory mechanisms (Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e6\u003c/span\u003e). Fusion transcripts can originate from parental genes that exhibit diverse responses to stress, including downregulation, upregulation, or no significant change in expression. Expression analysis of fusion transcripts across different tissues revealed their tissue-specific nature. For example, the LOC101500131_LOC101505021 fusion transcript exhibited upregulation in bud and downregulation in stem and flower, relative to its expression in leaves (Fig. S4). The LOC101506206_LOC101493600 is an interchromosomal fusion derived from a gene encoding a folylpolyglutamate synthase-like enzyme located on chromosome 5 and an uncharacterized gene from chromosome 8. This fusion produces an in-frame transcript that arises from an existing splice site of the 5' gene and a new splice site from the 3\u0026rsquo; gene. Since folylpolyglutamate synthase is a single-subunit enzyme, this fusion might add new domains, which may enhance its enzymatic activity (Muralla et al. \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2008\u003c/span\u003e; Li et al. \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). Interestingly, under stress conditions, the fusion transcript exhibited a higher relative fold change in expression compared to either of its precursor genes\u0026mdash;LOC101506206, which was upregulated, and LOC101493600, which was downregulated\u0026mdash;when compared to control conditions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe LOC101500131_LOC101505021 fusion was identified across all RNA-Seq samples with varying levels of expression among different samples. This fusion is formed by the joining of exon 10 of the LOC101500131 gene to exon 1 of the LOC101505021 gene, both genes are present on the forward strand of DNA and are involved in similar functions that are anaphase-promoting complex. This fusion uses a canonical splice site, however, not at the existing exon boundaries of the parental genes, and generates a frameshift fusion. Interestingly, this fusion showed significant upregulation in both drought and salt stresses as compared to the control condition. In contrast, the expression of the fusion precursor genes remains relatively unchanged in response to stress. Therefore, the regulation of the LOC101500131_LOC101505021 fusion appears to be independent of its parental genes. The LOC101495229_LOC101508351 fusion is formed by joining exon 2 from LOC101495229 and exon 1 from LOC101508351. One of the parental genes acts as a ubiquitin ligase, while the other remains uncharacterized. This fusion occurs at the exon boundary of the LOC101508351 gene and uses a new canonical splice site from the LOC101495229 gene. Both fusion transcript and its parental genes are stress-responsive; however, the relative change in fusion transcript expression under stress conditions is higher compared to parental genes.\u003c/p\u003e \u003cp\u003e \u003cb\u003eComprehensive assessment of fusion detection tools for\u003c/b\u003e \u003cb\u003eCicer arietinum\u003c/b\u003e\u003c/p\u003e \u003cp\u003eWhile many tools are available for fusion detection in humans, none of them is exclusively designed for plants, except EricScript-Plants (Benelli et al. \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2012\u003c/span\u003e), which is limited to plants in Ensembl and hence cannot be used for chickpea. In this study, we initially employed FusionMap, STAR-Fusion, and MapSplice to enable the comprehensive detection of fusion transcripts. However, to determine the best-performing tool for fusion detection in chickpea, a comparative analysis of five fusion detection algorithms was conducted. This revealed the following order of tools based on sensitivity and F-measure: FusionMap\u0026thinsp;\u0026gt;\u0026thinsp;STAR-Fusion\u0026thinsp;\u0026gt;\u0026thinsp;MapSplice\u0026thinsp;\u0026gt;\u0026thinsp;SQUID\u0026thinsp;\u0026gt;\u0026thinsp;Tophat-Fusion. A similar ranking was observed with public data, except that STAR-Fusion performed better than FusionMap with public data (Fig. S5a-b). Of note, there is a small overlap in the fusions detected by these tools in both our dataset and public data (Fig. S5c-d). This could be due to false discoveries associated with individual software packages, or the fact that none of the tools is inclusive. Fusion gene pairs LOC101506206_LOC101493600 and LOC101503481_LOC101494793 were commonly detected by 3 out of the 5 fusion detection tools.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThis analysis suggests that FusionMap detected the maximum number of true fusions, but it also identified a large number of false fusions, whereas STAR-Fusion identified fewer false fusions compared to FusionMap. Hence, FusionMap is the best performer in terms of sensitivity, whereas STAR-Fusion is the best performer in terms of precision. Hence, a combination of STAR-Fusion and FusionMap is a better choice for fusion transcript detection from RNA-Seq data for \u003cem\u003eCicer\u003c/em\u003e until more efficient tools become available.\u003c/p\u003e \u003cp\u003eAnother factor that affects the detection of FTs from RNA-Seq data is the sequencing depth. It was observed that samples with higher sequencing depth showed a larger number of fusions than those with lower sequencing depth. For instance, in the P3 sample, the number of FTs detected is more than that in CS2 because the sequencing depth of P3 is more than CS2 (Table S2).\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn recent years, its already established that the fusion transcripts are not exclusive to tumours but also occur in normal human tissues and a wide variety of species (Babiceanu et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Cong et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Although significant progress has been made in mammalian fusion transcript research, studies on fusion transcripts in plants remain limited. While some studies have been conducted on model species such as \u003cem\u003eArabidopsis\u003c/em\u003e and rice, a comprehensive investigation of plant fusion transcripts is still lacking. A few plant-specific fusion transcripts databases are also available (Singh et al. \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Arya et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Recently, a comprehensive profiling of fusion transcripts was conducted using public RNA-Seq datasets of \u003cem\u003eArabidopsis\u003c/em\u003e, Rice, and Chickpea (Chitkara et al. \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Most fusion transcripts were detected in only a single sample, raising concerns about their biological reproducibility and potential as artifacts. The current study addresses this gap through the use of in-house generated RNA-Seq datasets under defined abiotic stress conditions and across tissue types.\u003c/p\u003e \u003cp\u003eHere, we present the genome-wide identification of fusion transcripts in \u003cem\u003eCicer arietinum\u003c/em\u003e and explore their potential roles. Our study enhances the current understanding of fusion transcripts in legume plants. Notably, we found that interchromosomal fusion transcripts are more prevalent than those of intrachromosomal fusions in Chickpea, a pattern consistent with maize and soybean. In contrast, rice and \u003cem\u003eArabidopsis\u003c/em\u003e exhibit a higher frequency of intrachromosomal fusions (Cong et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). This indicates that the proportion of interchromosomal and intrachromosomal fusion transcripts is variable among different plants and may be influenced by several factors, including the gene density and compactness of the genome. The majority of gene pairs involved in fusion formation have no other partner genes, with only a few showing fusion isoforms, highlighting the specificity of these events. Additionally, we identified fusion transcripts that are tissue-specific or induced under specific stress conditions. Most parental genes involved in these fusions are multi-exonic and protein-coding. Expression correlation analysis of genes involved in fusions revealed a correlation coefficient (r) of 0.5, aligning with the recent report that suggests genes with similar transcriptional activity are more likely to participate in conserved fusion events (Cong et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFusion transcript identification in different genotypes of chickpea revealed that some fusions are conserved across all genotypes, while others are specific to individual genotypes, indicating a link between genomic variation and fusion transcript occurrence. Homologous fusion events between \u003cem\u003eCicer arietinum\u003c/em\u003e and \u003cem\u003eArabidopsis\u003c/em\u003e suggest that the parental genes involved play key roles in essential biological processes, underscoring the conservation of fusion events across different plant species. Most fusion breakpoints exhibited canonical splice sites at the fusion junctions, implying their generation via a splicing mechanism. However, a few fusions showed overlapping sequences at the breakpoints, suggesting that SHSs also contribute to the formation of these fusions in \u003cem\u003eCicer arietinum\u003c/em\u003e. Further investigation revealed that SHSs differ among fusion transcripts and are mostly\u0026thinsp;\u0026lt;\u0026thinsp;10 bp in length. Due to the short reads generated by Illumina sequencing, full-length fusion transcripts could not be reliably detected. To address this limitation, we employed long-read RNA sequencing to identify the full-length structure of fusion transcripts. Of the 328 fusion events initially identified, 95 were confirmed in the long-read RNA-Seq data, likely due to the lower sequencing depth of this dataset.\u003c/p\u003e \u003cp\u003eOut of the predicted fusion transcripts, we confirmed the existence of 10 FTs, marking the first experimental validation of fusion transcripts in \u003cem\u003eCicer arietinum\u003c/em\u003e. Notably, three of these fusions were stress-responsive and exhibited upregulation in drought and salt stress, suggesting a potential function in the plant\u0026rsquo;s stress response mechanisms. However, despite thorough validation efforts, we were unable to identify any fusion transcripts that were exclusively expressed under a specific stress and absent in control condition. This indicates that while fusion transcripts might be involved in the stress response, their expression may not be strictly limited to stress conditions. It is possible that these fusions represent a broader regulatory mechanism that operates under both normal and stress conditions but becomes more pronounced in response to stress. The confirmation of stress-responsive fusion transcripts paves the way for exploring the functional implications of these fusions in stress adaptation and resilience. Future research could focus on dissecting the biological roles of these fusions, including their impact on gene expression and protein function under stress conditions. Additionally, investigating the potential regulatory networks and pathways associated with these fusions could provide profound understanding into how \u003cem\u003eCicer arietinum\u003c/em\u003e adapts to environmental challenges.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors have no conflicts of interest to declare.\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research is supported by the BT/PR40146/BTIS/137/4/2020, BT/PR40169/BTIS/137/71/2023 research grants by the Department of Biotechnology (DBT), Government of India, and EEQ/2019/000231 research grant from Science and Engineering Research Board (SERB), Department of Science and Technology, Government of India; and core research grant of NIPGR, New Delhi \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFH: Conceptualization, Methodology, data curation, analysis, writing-original draft. SZ: Suggestions in data analysis and manuscript writing. SK: Conceptualization, Methodology, reviewed \u0026amp; edited the manuscript, conceived, and coordinated the project and provided overall supervision. All authors read and approved the final version of the manuscript.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFH gratefully acknowledges the Department of Biotechnology (DBT) for providing research fellowship. The authors extend their gratitude to the DBT e-Library Consortium (DeLCON) for providing access to e-material and Computational Biology \u0026amp; Bioinformatics Facility (CBBF) of NIPGR for their support.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eAnnala MJ, Parker BC, Zhang W, Nykter M (2013) Fusion genes and their discovery using high throughput sequencing. Cancer Lett 340:192\u0026ndash;200. https://doi.org/10.1016/J.CANLET.2013.01.011\u003c/li\u003e\n \u003cli\u003eArya A, Arora S, Hamid F, Kumar S (2024) PFusionDB: a comprehensive database of plant-specific fusion transcripts. 3 Biotech 2024 14:11 14:1\u0026ndash;8. https://doi.org/10.1007/S13205-024-04132-1\u003c/li\u003e\n \u003cli\u003eBabiceanu M, Qin F, Xie Z, et al (2016) Recurrent chimeric fusion RNAs in non-cancer tissues and cells. Nucleic Acids Res 44:2859\u0026ndash;2872. https://doi.org/10.1093/NAR/GKW032\u003c/li\u003e\n \u003cli\u003eBenelli M, Pescucci C, Marseglia G, et al (2012) Discovering chimeric transcripts in paired-end RNA-seq data by using EricScript. Bioinformatics 28:3232\u0026ndash;3239. https://doi.org/10.1093/BIOINFORMATICS/BTS617\u003c/li\u003e\n \u003cli\u003eChao Y, Yuan J, Li S, et al (2018) Analysis of transcripts and splice isoforms in red clover (Trifolium pratense L.) by single-molecule long-read sequencing. BMC Plant Biol 18:1\u0026ndash;12. https://doi.org/10.1186/S12870-018-1534-8/FIGURES/7\u003c/li\u003e\n \u003cli\u003eChen H, Tang Y, Liu J, et al (2017) Emergence of a novel chimeric gene underlying grain number in rice. Genetics 205:993\u0026ndash;1002. https://doi.org/10.1534/GENETICS.116.188201/-/DC1\u003c/li\u003e\n \u003cli\u003eChen S, Zhou Y, Chen Y, Gu J (2018) fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34:i884\u0026ndash;i890. https://doi.org/10.1093/BIOINFORMATICS/BTY560\u003c/li\u003e\n \u003cli\u003eChitkara P, Singh A, Gangwar R, et al (2024) The landscape of fusion transcripts in plants: a new insight into genome complexity. BMC Plant Biol 24:1\u0026ndash;18. https://doi.org/10.1186/S12870-024-05900-0/FIGURES/6\u003c/li\u003e\n \u003cli\u003eChwalenia K, Facemire L, Li H (2017) Chimeric RNAs in cancer and normal physiology. Wiley Interdiscip Rev RNA 8:. https://doi.org/10.1002/WRNA.1427\u003c/li\u003e\n \u003cli\u003eCong J, Zhang S, Zhang Q, et al (2024) Conserved features and diversity attributes of chimeric RNAs across accessions in four plants. Plant Biotechnol J 22:3151\u0026ndash;3163. https://doi.org/10.1111/PBI.14437\u003c/li\u003e\n \u003cli\u003eDoskočilov\u0026aacute; A, Pl\u0026iacute;hal O, Volc J, et al (2011) A nodulin/glutamine synthetase-like fusion protein is implicated in the regulation of root morphogenesis and in signalling triggered by flagellin. Planta 234:459\u0026ndash;476. https://doi.org/10.1007/S00425-011-1419-7\u003c/li\u003e\n \u003cli\u003eDupain C, Harttrampf AC, Boursin Y, et al (2019) Discovery of New Fusion Transcripts in a Cohort of Pediatric Solid Cancers at Relapse and Relevance for Personalized Medicine. Molecular Therapy 27:200. https://doi.org/10.1016/J.YMTHE.2018.10.022\u003c/li\u003e\n \u003cli\u003eFarrow SC, Hagel JM, Beaudoin GAW, et al (2015) Stereochemical inversion of (S)-reticuline by a cytochrome P450 fusion in opium poppy. Nat Chem Biol 11:728\u0026ndash;732. https://doi.org/10.1038/NCHEMBIO.1879\u003c/li\u003e\n \u003cli\u003eFlowers TJ, Gaur PM, Gowda CLL, et al (2010) Salt sensitivity in chickpea. Plant Cell Environ 33:490\u0026ndash;509. https://doi.org/10.1111/J.1365-3040.2009.02051.X\u003c/li\u003e\n \u003cli\u003eGarg R, Sahoo A, Tyagi AK, Jain M (2010) Validation of internal control genes for quantitative gene expression studies in chickpea (Cicer arietinum L.). Biochem Biophys Res Commun 396:283\u0026ndash;288. https://doi.org/10.1016/J.BBRC.2010.04.079\u003c/li\u003e\n \u003cli\u003eGarg R, Shankar R, Thakkar B, et al (2016) Transcriptome analyses reveal genotype- and developmental stage-specific molecular responses to drought and salinity stresses in chickpea. Scientific Reports 2016 6:1 6:1\u0026ndash;15. https://doi.org/10.1038/srep19228\u003c/li\u003e\n \u003cli\u003eGe H, Liu K, Juan T, et al (2011) FusionMap: detecting fusion genes from next-generation sequencing data at base-pair resolution. Bioinformatics 27:1922\u0026ndash;1928. https://doi.org/10.1093/BIOINFORMATICS/BTR310\u003c/li\u003e\n \u003cli\u003eGingeras TR (2009) Implications of chimaeric non-co-linear transcripts. Nature 2009 461:7261 461:206\u0026ndash;211. https://doi.org/10.1038/nature08452\u003c/li\u003e\n \u003cli\u003eGupta SK, Luo L, Yen L (2018) RNA-mediated gene fusion in mammalian cells. Proc Natl Acad Sci U S A 115:E12295\u0026ndash;E12304. https://doi.org/10.1073/PNAS.1814704115/SUPPL_FILE/PNAS.1814704115.SM02.MOV\u003c/li\u003e\n \u003cli\u003eHaas BJ, Dobin A, Li B, et al (2019) Accuracy assessment of fusion transcript detection via read-mapping and de novo fusion transcript assembly-based methods. Genome Biol 20:. https://doi.org/10.1186/S13059-019-1842-9\u003c/li\u003e\n \u003cli\u003eHagel JM, Facchini PJ (2017) Tying the knot: occurrence and possible significance of gene fusions in plant metabolism and beyond. J Exp Bot 68:4029\u0026ndash;4043. https://doi.org/10.1093/JXB/ERX152\u003c/li\u003e\n \u003cli\u003eJain M, Bansal J, Rajkumar MS, Garg R (2022) An integrated transcriptome mapping the regulatory network of coding and long non-coding RNAs provides a genomics resource in chickpea. Communications Biology 2022 5:1 5:1\u0026ndash;18. https://doi.org/10.1038/s42003-022-04083-4\u003c/li\u003e\n \u003cli\u003eKibbe WA (2007) OligoCalc: an online oligonucleotide properties calculator. Nucleic Acids Res 35:. https://doi.org/10.1093/NAR/GKM234\u003c/li\u003e\n \u003cli\u003eKim D, Paggi JM, Park C, et al (2019) Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nature Biotechnology 2019 37:8 37:907\u0026ndash;915. https://doi.org/10.1038/s41587-019-0201-4\u003c/li\u003e\n \u003cli\u003eKim D, Salzberg SL (2011) TopHat-Fusion: An algorithm for discovery of novel fusion transcripts. Genome Biol 12:1\u0026ndash;15. https://doi.org/10.1186/GB-2011-12-8-R72/FIGURES/6\u003c/li\u003e\n \u003cli\u003eKim M, Canio W, Kessler S, Sinha N (2001) Developmental changes due to long-distance movement of a homeobox fusion transcript in tomato. Science 293:287\u0026ndash;289. https://doi.org/10.1126/SCIENCE.1059805\u003c/li\u003e\n \u003cli\u003eKoul B, Sharma K, Sehgal V, et al (2022) Chickpea (Cicer arietinum L.) Biology and Biotechnology: From Domestication to Biofortification and Biopharming. Plants 11:. https://doi.org/10.3390/PLANTS11212926\u003c/li\u003e\n \u003cli\u003eKumar S, Vo AD, Qin F, Li H (2016) Comparative assessment of methods for the fusion transcripts detection from RNA-Seq data. Sci Rep 6:21597. https://doi.org/10.1038/srep21597\u003c/li\u003e\n \u003cli\u003eLatysheva NS, Babu MM (2016) Discovering and understanding oncogenic gene fusions through data intensive computational approaches. Nucleic Acids Res 44:4487\u0026ndash;4503. https://doi.org/10.1093/NAR/GKW282\u003c/li\u003e\n \u003cli\u003eLi H, Wang J, Ma X, Sklar J (2009a) Gene fusions and RNA trans-splicing in normal and neoplastic human cells. Cell Cycle 8:218\u0026ndash;222. https://doi.org/10.4161/CC.8.2.7358\u003c/li\u003e\n \u003cli\u003eLi J, Lee EJ, Chang L, Facchini PJ (2016) Genes encoding norcoclaurine synthase occur as tandem fusions in the Papaveraceae. Scientific Reports 2016 6:1 6:1\u0026ndash;12. https://doi.org/10.1038/srep39256\u003c/li\u003e\n \u003cli\u003eLi X, Zhao L, Jiang H, Wang W (2009b) Short homologous sequences are strongly associated with the generation of chimeric RNAs in eukaryotes. J Mol Evol 68:56\u0026ndash;65. https://doi.org/10.1007/S00239-008-9187-0\u003c/li\u003e\n \u003cli\u003eLivak KJ, Schmittgen TD (2001) Analysis of relative gene expression data using real-time quantitative PCR and the 2(-Delta Delta C(T)) Method. Methods 25:402\u0026ndash;408. https://doi.org/10.1006/METH.2001.1262\u003c/li\u003e\n \u003cli\u003eMa C, Shao M, Kingsford C (2018) SQUID: Transcriptomic structural variation detection from RNA-seq. Genome Biol 19:1\u0026ndash;16. https://doi.org/10.1186/S13059-018-1421-5/FIGURES/7\u003c/li\u003e\n \u003cli\u003eMukherjee S, Detroja R, Balamurali D, et al (2021) Computational analysis of sense-antisense chimeric transcripts reveals their potential regulatory features and the landscape of expression in human cells. NAR Genom Bioinform 3:. https://doi.org/10.1093/NARGAB/LQAB074\u003c/li\u003e\n \u003cli\u003eMuralla R, Chen E, Sweeney C, et al (2008) A Bifunctional Locus (BIO3-BIO1) Required for Biotin Biosynthesis in Arabidopsis. Plant Physiol 146:60. https://doi.org/10.1104/PP.107.107409\u003c/li\u003e\n \u003cli\u003eNathawat BDS, Sharma OP, Kumari M, Shivran H (2024) Effect of Nutrients on Wilt in Chickpea. Legume Research 47:152\u0026ndash;155. https://doi.org/10.18805/LR-4490\u003c/li\u003e\n \u003cli\u003ePertea M, Kim D, Pertea GM, et al (2016) Transcript-level expression analysis of RNA-seq experiments with HISAT, StringTie and Ballgown. Nat Protoc 11:1650\u0026ndash;1667. https://doi.org/10.1038/NPROT.2016.095\u003c/li\u003e\n \u003cli\u003eQiao D, Yang C, Chen J, et al (2019) Comprehensive identification of the full-length transcripts and alternative splicing related to the secondary metabolism pathways in the tea plant (Camellia sinensis). Scientific Reports 2019 9:1 9:1\u0026ndash;13. https://doi.org/10.1038/s41598-019-39286-z\u003c/li\u003e\n \u003cli\u003eQin F, Zhang Y, Liu J, Li H (2017) SLC45A3-ELK4 functions as a long non-coding chimeric RNA. Cancer Lett 404:53\u0026ndash;61. https://doi.org/10.1016/J.CANLET.2017.07.007\u003c/li\u003e\n \u003cli\u003eSherman BT, Hao M, Qiu J, et al (2022) DAVID: a web server for functional enrichment analysis and functional annotation of gene lists (2021 update). Nucleic Acids Res 50:W216\u0026ndash;W221. https://doi.org/10.1093/NAR/GKAC194\u003c/li\u003e\n \u003cli\u003eSingh A, Zahra S, Das D, Kumar S (2019) AtFusionDB: a database of fusion transcripts in Arabidopsis thaliana. Database (Oxford) 2019:. https://doi.org/10.1093/DATABASE/BAY135\u003c/li\u003e\n \u003cli\u003eSutton RE, Boothroyd JC (1986) Evidence for trans splicing in trypanosomes. Cell 47:527\u0026ndash;535. https://doi.org/10.1016/0092-8674(86)90617-3\u003c/li\u003e\n \u003cli\u003eTan C, Liu H, Ren J, et al (2019) Single-molecule real-time sequencing facilitates the analysis of transcripts and splice isoforms of anthers in Chinese cabbage (Brassica rapa L. ssp. pekinensis). BMC Plant Biol 19:. https://doi.org/10.1186/S12870-019-2133-Z\u003c/li\u003e\n \u003cli\u003eTang D, Chen M, Huang X, et al (2023) SRplot: A free online platform for data visualization and graphing. PLoS One 18:e0294236. https://doi.org/10.1371/JOURNAL.PONE.0294236\u003c/li\u003e\n \u003cli\u003eTurner NC, Colmer TD, Quealy J, et al (2013) Salinity tolerance and ion accumulation in chickpea (Cicer arietinum L.) subjected to salt stress. Plant Soil 365:347\u0026ndash;361. https://doi.org/10.1007/S11104-012-1387-0/METRICS\u003c/li\u003e\n \u003cli\u003eVarley KE, Gertz J, Roberts BS, et al (2014) Recurrent read-through fusion transcripts in breast cancer. Breast Cancer Res Treat 146:287. https://doi.org/10.1007/S10549-014-3019-2\u003c/li\u003e\n \u003cli\u003eVarshney RK, Song C, Saxena RK, et al (2013) Draft genome sequence of chickpea (Cicer arietinum) provides a resource for trait improvement. Nature Biotechnology 2013 31:3 31:240\u0026ndash;246. https://doi.org/10.1038/nbt.2491\u003c/li\u003e\n \u003cli\u003eWang K, Singh D, Zeng Z, et al (2010) MapSplice: Accurate mapping of RNA-seq reads for splice junction discovery. Nucleic Acids Res 38:e178. https://doi.org/10.1093/NAR/GKQ622\u003c/li\u003e\n \u003cli\u003eWinzer T, Kern M, King AJ, et al (2015) Plant science. Morphinan biosynthesis in opium poppy requires a P450-oxidoreductase fusion protein. Science 349:309\u0026ndash;312. https://doi.org/10.1126/SCIENCE.AAB1852\u003c/li\u003e\n \u003cli\u003eZhang G, Guo G, Hu X, et al (2010) Deep RNA sequencing at single base-pair resolution reveals high complexity of the rice transcriptome. Genome Res 20:646. https://doi.org/10.1101/GR.100677.109\u003c/li\u003e\n \u003cli\u003eZhou Y, Lu Q, Zhang J, et al (2022) Genome-Wide Profiling of Alternative Splicing and Gene Fusion during Rice Black-Streaked Dwarf Virus Stress in Maize ( Zea mays L.). Genes (Basel) 13:. https://doi.org/10.3390/GENES13030456\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Abiotic stress, Cicer arietinum, Fusion transcripts, trans-splicing, transcriptome complexity","lastPublishedDoi":"10.21203/rs.3.rs-6540588/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6540588/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eUnderstanding the transcriptome complexity is essential for deciphering the genome regulation at transcriptional level. High-throughput sequencing technologies facilitated the detection of Fusion transcripts (FTs) which are chimeric mRNA molecules derived from gene fusion due to chromosomal rearrangements or via splicing machinery at the RNA level. In this study, we explored the transcriptome complexity in \u003cem\u003eCicer arietinum\u003c/em\u003e due to fusion events by using high-throughput RNA-Seq datasets in five tissues, and under two abiotic stress conditions. Using three different fusion detection tools, a total of 328 unique FTs were detected. Sequence analysis revealed that 69% of FTs showed the presence of canonical splice site at the junction, which indicates their generation via trans-splicing. Functional annotation and enrichment analyses of fusion partners suggested the expansion of biological functionality. A total of 10 FTs were validated via RT-PCR followed by Sanger sequencing, which are the first FTs described in important legume chickpea. Expression analysis of validated FTs and their parental genes under drought and salinity stress conditions identified differentially expressed fusions. This study offers detailed insight into the fusion landscape of \u003cem\u003eCicer arietinum\u003c/em\u003e and inferred a potential role of FTs during stress responses.\u003c/p\u003e","manuscriptTitle":"Transcriptome analysis revealed stress responsive fusion transcripts in Chickpea (Cicer arietinum)","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-07 11:38:44","doi":"10.21203/rs.3.rs-6540588/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2dd299ed-bdcf-4eec-bcf6-0a6d13ef4939","owner":[],"postedDate":"May 7th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-05-13T23:21:26+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-07 11:38:44","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6540588","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6540588","identity":"rs-6540588","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00