Reanalysis of the Tomato Transcriptome Revealed Host Responses to Southern Tomato Virus Infection | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Reanalysis of the Tomato Transcriptome Revealed Host Responses to Southern Tomato Virus Infection Gholamhossein Badeli, Kami Kaboosi, Saeed Nasrollanejad This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8823240/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract A previous study reported that Southern tomato virus (STV)-infected tomato plants were taller, had higher fruit yields. Here, it was aimed to reanalyze the same dataset with updated genomic resources and a bioinformatics pipeline to validate and expand upon those findings.The raw single-end RNA-Seq reads were remapped to the latest tomato reference genome (SL3.0) using Rsubread for alignment and featureCounts for gene quantification. Differential expression was assessed with DESeq2, applying rigorous filtering (FDR < 0.05, |log₂ fold-change| ≥ 1). The results revealed no genes meeting the criteria for differential expression (FDR < 0.05, fold-change ≥ 2). In contrast to the original report, and the host transcriptome remains largely unchanged by STV infection. The few transcripts with the largest expression differences included those in the ethylene biosynthesis and signaling pathway, which were slightly lower in STV-infected plants, mirroring the trends reported previously. Our reanalysis using a RNA-Seq workflow finds that STV infection causes at changes in tomato leaf gene expression, notably slight downregulation of ethylene-related genes, but no broad transcriptomic reprogramming. These results temper the earlier claims of significant transcriptomic effects of STV, instead supporting the notion that STV exerts a minimal and possibly cultivar-specific influence on the host. This study underlines the importance of updated genomic references and statistical approaches. Future studies with larger sample sizes or different developmental stages will be valuable. Southern tomato virus Solanum lycopersicum RNA sequencing differential gene expression persistent plant virus Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Double-stranded RNA viruses that persistently infect plants have been identified in many crop species [ 1 ]. Unlike acute or chronic viral infections that cause obvious disease symptoms, persistent viruses typically establish long-term, low-titer infections with no apparent harm to the host [ 2 ]. These viruses are efficiently seed-transmitted and can spread through plant populations silently, leading some researchers to consider them as mutualistic or commensal symbionts rather than pathogens [ 3 ]. Most persistent plant viruses characterized to date belong to the families Partitiviridae or Endornaviridae , and they have evolved unique lifestyles distinct from those of lytic plant viruses. Southern tomato virus (STV) is a persistent virus of tomato that was first characterized in 2009 by Sabanadzovic et al. [ 4 ]. STV contains a monopartite dsRNA genome (3.5 kb) encoding two overlapping open reading frames and has been classified in the genus Amalgavirus [ 5 ]. Initial reports detected STV in tomatoes in plants showing stunted, bushy growth and fruit discoloration. However, STV is also frequently found in asymptomatic tomato plants or in mixed infections with acute viruses like Pepino mosaic virus. The prevalence of STV in diverse tomato germplasm without obvious disease outbreaks suggested that STV might have negligible or even beneficial effects on its host. To directly evaluate STV’s impact, researchers compared STV-infected vs. STV-free lines of the tomato cultivar M82 under controlled conditions [ 6 ]. Intriguingly, STV-infected plants were taller and produced more fruits and seeds than their STV-free counterparts. The germination rate of seeds from infected plants was also higher, leading the authors to propose that STV confers a mutualistic advantage to the host. Importantly, transcriptome analysis of leaves from these plants indicated that STV infection induces changes in host gene expression. In particular, genes involved in the ethylene biosynthesis and signaling pathway were found to be downregulated in STV-infected tomatoes. Ethylene is a plant hormone that can restrict vegetative growth and promote fruit ripening [ 7 ]; thus, lower ethylene activity in STV-infected plants could mechanistically explain the taller, higher-yield phenotype of infected M82 plants [ 6 ]. The original study highlighted three genes, 1-aminocyclopropane-1-carboxylate oxidase 1 (ACO1), a NAC-domain transcription factor (NAC1), and ethylene-responsive element binding protein (EREB), as key ethylene-related genes whose expression was reduced by STV infection. These findings suggested that STV might modulate the host’s hormonal signaling network to produce beneficial developmental outcomes. While provocative, the transcriptomic evidence for STV’s effects was based on a relatively small RNA-Seq dataset analyzed with then-standard methods. The original analysis used the tomato reference genome SL2.50 (released in 2012) and a TopHat/Cufflinks pipeline to compute fragments per kilobase per million reads (FPKM). Differential expression was inferred from differences in FPKM and validated for a few genes by qPCR [ 6 ], but no genome-wide statistical test like false discovery rate correction was reported for the RNA-Seq results. In the years since, the tomato reference genome has been substantially improved (SL3.0 was released with more complete gene models), and RNA-Seq analysis methods have shifted toward read count-based workflows that offer more robust statistics [ 8 ]. Reanalysis of legacy RNA-Seq data with up-to-date pipelines can either confirm the original findings with greater confidence or reveal potential false positives/negatives introduced by older methods. This is particularly relevant for STV, given the subtle nature of persistent virus effects and the small sample size, the risk of both false negatives and false positives is high. In this study, a comprehensive reanalysis of the RNA-Seq dataset GSE137303 from [ 6 ] is presented, which examines STV-infected vs. STV-free tomato M82 plants. By leveraging the latest S. lycopersicum reference (SL3.0), it was aimed to provide a more rigorous assessment of STV’s transcriptomic impact. Additional exploratory analyses performed, including unsupervised clustering and Gene Ontology enrichment, to situate any observed gene expression changes in a broader biological context. By comparing the reanalysis with the original findings, we seek to clarify whether STV infection reprograms host gene expression, particularly the ethylene pathway, or if the host response is minimal, a question fundamental to understanding the persistent virus–plant interaction paradigm. Our results shed light on the molecular mutualism hypothesis and underscore best practices for transcriptomic studies of host-symbiont interactions in plants. Materials and methods Plant material, dataset, and experimental design Publicly available transcriptome data from S. lycopersicum differing in Southern tomato virus infection status analyzed. The RNA-Seq dataset, originally published by Fukuhara et al. (2019), is accessible in NCBI Gene Expression Omnibus under accession GSE137303. It comprises six leaf mRNA libraries obtained from one-month-old greenhouse-grown tomato plants: three biological replicates of STV-infected plants (STV + ) and three replicates of STV-free control plants (STV − ). All plants were verified for presence or absence of STV by dsRNA extraction and RT-PCR in the original study. The STV-free line was a cured sub-line of M82, and the STV-infected line was an M82 sub-line naturally harboring STV [ 6 ]. This paired design allowed direct comparison of transcriptomes to infer the effect of asymptomatic STV infection. For this study, raw sequencing data were downloaded in FASTQ format. The reference genome assembly was obtained from Ensembl Plants. This reference includes improved assembly continuity and gene annotation for S. lycopersicum . Corresponding reference files, including the genomic DNA FASTA and the comprehensive gene annotation GTF was obtained. Additionally, to account for viral reads, we included the STV genome sequence (GenBank KX949574) as an extra “chromosome” in the reference during read alignment. This enabled simultaneous mapping of reads to both host and virus, although as a persistent virus STV was not expected to produce abundant mRNA. The second major update was switching from an FPKM-based pipeline (TopHat to Cuffnorm) to a count-based differential expression approach. Count-based methods use the discrete read counts per gene and apply established statistical models for significance testing, offering greater sensitivity and control of false positives. The Bioconductor DESeq2 package v1.42.0 was used for this analysis, which utilizes a negative binomial generalized linear model to identify differentially expressed genes with robust estimates of biological variance [ 9 ]. This approach allowed us to apply Benjamini–Hochberg false discovery rate (FDR) correction across all genes, ensuring that any significant hits are not likely due to chance. Read preprocessing and alignment Raw FASTQ files from the six libraries were first subjected to quality control. Adapter sequences and low-quality bases were trimmed using the Trimmomatic tool (v0.39) with default settings, and read quality was verified with FastQC. All reads were confirmed to be 75 nt in length after trimming, as expected from the library protocol. Post-filtering, each library retained between ~ 14 and 18 million reads. High-quality reads were then aligned to the combined reference genome (tomato SL3.0 plus STV) using the Rsubread aligner (R package Rsubread, v2.13.3) [ 10 ]. Since the data derive from mRNA, we built a gapped index to accommodate spliced alignments from the SL3.0 reference FASTA. Alignment was performed in base-space with up to 3 mismatches allowed per read. Rsubread’s align function was set to allow soft-clipping and gapped alignment for introns. We also permitted a maximum insertion or deletion length of 5 bp to account for small indels. Reads mapping to multiple locations in the genome were handled by reporting a single best alignment to avoid inflating counts. After alignment, we used SAMtools through Rsubread to sort and index the resulting BAM files. Gene quantification and normalization Gene-level counting of aligned reads was carried out with featureCounts, as a part of the Rsubread package to produce a matrix of integer read counts per gene. FeatureCounts were provided with the Ensembl SL3.0 gene annotation GTF (Solanum_lycopersicum.SL3.0.61.gff3), which corresponds to ITAG3.0 gene models. FeatureCounts [ 11 ] was run in meta-feature mode, meaning that for each gene or locus it summed reads across all exons, yielding one total count per gene. We allowed featureCounts to count a read if it overlapped any exon of a gene, capturing both fully spliced and pre-mRNA fragments. Reads mapping to multiple genes were excluded by featureCounts to maintain accurate counts. Multi-mapped reads that aligned to multiple loci were also not counted toward any single gene, per the default settings, since their origin is uncertain. After counting, a raw count matrix for 34,075 tomato gene features obtained, including protein-coding genes and non-coding transcripts annotated in Ensembl. Prior to statistical analysis, the count data filtered and normalized. Genes with consistently low read counts were filtered out to reduce noise. Specifically, any gene with < 10 total counts across all six samples was removed from the dataset. Library size normalization was performed using the median ratio method implemented in DESeq2. This method computes size factors for each sample such that the median of the ratio of observed counts to expected counts (geometric mean across samples) is 1. Effectively, it accounts for differences in sequencing depth between libraries. After normalization, count data were transformed to variance-stabilized counts using DESeq2’s variance stabilizing transformation ( vst) for use in visualization and clustering. The vst converts count data to a log 2 -scale while accounting for the mean-variance relationship, producing values that are homoscedastic and suitable for principal component analysis. Differential expression analysis Differential gene expression between the STV-infected and STV-free groups tested using the DESeq2 statistical framework. The model design was simple, with a single factor condition (STV + vs STV − ). DESeq2 fits a negative binomial model for each gene’s counts and uses Wald tests to assess whether the log 2 fold-change (log 2 FC) between conditions is significantly different from zero. We included the three replicates per group to estimate within-group dispersion. Dispersions were empirically estimated and shrunk toward a trend to improve stability for genes with low counts, following DESeq2’s standard procedure. For each gene, an estimated log 2 fold-change (STV + / STV − ) obtained and an associated p -value. Next, a Benjamini–Hochberg [ 12 ] adjustment applied to control the false discovery rate (FDR) across the ~ 25,000 tested genes. Genes were declared significantly differentially expressed if they met two criteria: (1) FDR (adjusted p -value) < 0.05, and (2) absolute log₂FC ≥ 1, which corresponds to a ≥ 2-fold change in expression. This fold-change threshold was chosen to focus on changes likely to be biologically meaningful in the context of the relatively small sample size. Additionally, shrunken log₂ fold-changes was measured using the apeglm method in DESeq2. Shrinkage of effect sizes is recommended to avoid overestimation of fold-changes for genes with low counts or high dispersion. For exploratory purposes and generating lists of candidate genes for enrichment, we also identified genes that would be considered differentially expressed under a more lenient cutoff (FDR < 0.10 or < 0.20). This relaxed threshold is not used to claim significance, but rather to ensure we did not miss any borderline effects that might be biologically relevant in enrichment analysis or in comparisons with the original study’s findings. This was due to no genes achieved the primary significance threshold (FDR < 0.05 and |log₂FC| ≥ 1), highlighting the need for such exploratory follow-ups. Functional annotation of genes Tomato gene identifiers in the Ensembl database are given as Solyc numbers (like Solyc07g007770) which are not immediately intuitive biologically. To aid interpretation, we mapped each gene ID to a gene name or description where possible. The Ensembl Plants ( https://plants.ensembl.org/index.html ) GTF annotation provides a “gene_name” attribute for many genes, which often corresponds to a common gene symbol or a descriptive name. Therefore, these gene symbols extracted, and when unavailable, alternate annotation sources like UniProt and NCBI RefSeq were searched for gene descriptions or putative functions. A custom mapping table was created so that figures and tables could display both the Solyc ID and a gene name or product description. In cases where a gene had no clear symbolic name, the Solyc identifier is retained but noted if it was predicted to encode a certain type of protein. Exploratory data analysis and visualization Exploratory analyses were performed to visualize the overall transcriptomic patterns. In this regard, a principal component analysis (PCA) conducted on the normalized gene expression data. PCA was run on the variance-stabilized expression matrix for the top 500 genes with the highest variance across the six samples. This unsupervised approach reduces the data to principal components (PCs) that capture the largest sources of variation. We examined whether the samples separated along any PC according to STV infection status. A scatter plot of the first two PCs was generated with each sample labeled and colored by condition (STV + vs STV − ). Also, a sample distance heatmap was generated. Using the variance-stabilized data, the Euclidean distance between each pair of samples was evaluated based on the expression of all genes. The goal was to see if STV-infected samples cluster together, indicating a consistent transcriptomic signature of infection, or if samples cluster by other factors, suggesting high within-group variability. Functional enrichment analysis To interpret any biological themes in the data, a Gene Ontology (GO) [ 13 ] enrichment analysis was carried out for differentially expressed genes. Given the small number of genes showing any expression change, we relaxed the significance threshold to include genes with adjusted p < 0.1 (or even < 0.2) for enrichment purposes. We focused on the Biological Process (BP) branch of GO, as our interest was in physiological pathways or processes affected by STV. Accordingly, enrichment was performed using the clusterProfiler package [ 14 ]. GO annotations obtained for tomato genes via the biomaRt [ 15 – 17 ] interface to Ensembl Plants, retrieving all GO terms associated with each Ensembl gene ID. A custom GO term to gene mapping was constructed for S. lycopersicum . For the gene list of interest, clusterProfiler was used to run a hypergeometric test to identify GO terms that are over-represented. The background universe was set as all expressed genes in our dataset. Enrichment p -values were adjusted for multiple testing. We limited our analysis to GO-BP terms to avoid dilution by overly broad categories. Reproducibility and computational environment All analysis steps were performed in the R environment (v4.5.1) on a Windows 10 64-bit system. Bioconductor packages were leveraged for most tasks, ensuring compatibility and version control. Key packages were including, Rsubread v2.13.3 (for alignment and counting), DESeq2 v1.42.0 (for differential expression), tximport v1.20 (for any potential transcript-level import, though not used here), clusterProfiler v4.10.0 (for GO analysis), ggplot2 v3.4.0, pheatmap v1.0.12, and ggrepel v0.9.2 (for labeling plots). Results Sequencing and mapping overview Each of the six RNA-Seq libraries yielded between 14 and 20 million raw single-end reads, consistent with the sequencing depth reported in the original dataset. After quality trimming, over 99% of reads were retained as high-quality. These reads were mapped to the tomato reference genome SL3.0 (Solanum_lycopersicum.SL3.0.dna.toplevel.fa) with high alignment rate. On average, 85–90% of reads per sample aligned to a unique position in the tomato genome and were assigned to an annotated gene feature (Table 1 ). The remaining reads either mapped to intergenic regions or could not be confidently assigned due to multi-mapping. The inclusion of the STV genome sequence in the alignment index did not affect mapping rates to tomato; only a low numbers of reads aligned to the STV dsRNA (< 0.01%), indicating extremely low expression or absence of poly(A)-transcripts from the virus, as expected. Figure 1 (a) Gene-wise dispersion estimates (black), the fitted mean–dispersion trend (red), and the final MAP-moderated dispersions used in modeling (blue) as a function of mean normalized counts. This diagnostic from DESeq2 assesses how dispersion shrinks toward the fitted trend, improving stability for low-count genes in the STV + vs STV− contrast. (b) Total assigned read counts per sample (featureCounts → DESeq2). shown as bars. Samples are S. lycopersicum leaves processed through Rsubread alignment, featureCounts quantification, and DESeq2; counts are used downstream for normalization and vst. Similar library sizes across replicates support comparability prior to differential analysis Table 1 Read alignment and assignment statistics for RNA-Seq libraries from STV + and STV − tomato plants. Sample ID Condition Total reads (M) Assigned to gene ¥ (M) % Assigned Unassigned – No feature (M) Unassigned ¥ – ambiguity (M) STV − _rep1 STV-free 17.65 14.23 80.6% 1.43 1.98 STV − _rep2 STV-free 15.87 13.48 84.9% 0.99 1.46 STV − _rep3 STV-free 13.96 11.61 83.2% 0.87 1.50 STV + _rep1 STV-infected 17.92 15.29 85.4% 1.37 2.11 STV + _rep2 STV-infected 15.31 13.76 89.9% 0.99 1.74 STV + _rep3 STV-infected 19.81 16.16 81.6% 1.20 2.08 “Assigned to gene” indicates reads that mapped to a tomato gene feature by featureCounts . “Unassigned” categories include reads mapping outside annotated exons or those ambiguously mapped. Values in table are in millions of reads (M) All samples showed comparable sequencing depth and mapping success, with no obvious outlier in library preparation or sequencing quality. Thus, any differences observed in gene expression are unlikely to stem from technical artifacts. The raw counts per gene were normalized to account for the small differences in library size (Table 1 shows STV + _rep3 had the highest read count at ~ 19.8M, while STV − _rep3 had ~ 14.0M; DESeq2’s size factor normalization corrected for these differences). Differential gene expression analysis signed no significant hits Applying the DESeq2 differential expression analysis between STV-infected and uninfected tomato samples, no genes found that met the conventional criteria for significant differential expression. In other words, none of the ~ 25,000 expressed genes showed a statistically significant change in expression due to STV infection in this dataset. This is a outstanding result, as it suggests that STV elicits little to no transcriptomic perturbation in M82 tomato leaves, within the sensitivity limits of this experiment. Table 2 summarizes the differential expression results, listing the genes with the largest observed fold-changes, regardless of significance, in STV-infected vs. STV-free plants. Even the top-ranking candidates did not pass the FDR cutoff after multiple-testing correction. For completeness, the top up- and down-regulated genes ranked by smallest nominal p -value, with their expression fold-changes and significance values are presented. Table 2 Top candidates for differential expression between STV-infected and STV-free tomato plants (sorted by nominal p-value). Gene ID (Solyc) Gene Name / Description log₂FC (STV + vs STV − ) Adjusted p-value (FDR) Notes Solyc12g009240.1 Pathogenesis-related (PR)-protein –1.12 0.0831 Encodes PR protein (defense-related) – highest fold change down, n.s. FDR ≈ 0.08 Solyc07g049530.2 ACC oxidase 1 (ACO1) –1.47 0.225 Ethylene biosynthesis enzyme (down in STV + ) Solyc08g067160.3 Uncharacterized protein (DUF247) –1.26 0.736 Expressed protein of unknown function Solyc04g008330.1 beta-Galactosidase 4-like –1.31 1.000 Cell wall degradation enzyme family Solyc02g086990.2 PGIP (Polygalacturonase inhibitor) –0.78 1.000 Defense-related cell wall protein Solyc07g007770 .2 ERF017 (Ethylene-responsive TF) + 3.97 0.157 Ethylene-responsive factor (up in STV + ) nominal p = 0.0157 (highest fold up) Solyc02g031950.1 Glyoxalase II (GLO2) + 1.49 1.000 Stress-related metabolism enzyme, slight up Solyc01g109010.3 GMP synthase [glutamine-hydrolyzing] + 1.09 1.000 Nucleotide biosynthesis enzyme, slight up ACO1 = 1-aminocyclopropane-1-carboxylate oxidase 1; ERF017 is an ethylene-responsive transcription factor. log₂FC is shown as shrunken DESeq2 estimate. A negative log₂FC indicates lower expression in STV-infected plants As shown in Table 2 , the largest fold-change observed was a − 1.12 log 2 decrease for a pathogenesis-related protein gene (Solyc12g009240), which corresponds to roughly a halving of expression in STV-infected plants compared to STV-free. This gene had a p -value of ~ 4.6×10 –6 and an FDR of 0.083; it narrowly missed the 5% FDR threshold. ACO1 (Solyc07g049530), ‘ the key ethylene biosynthetic gene identified in the original study ’, showed a log₂FC of − 1.47, but its adjusted p -value was 0.225, reflecting high variance among replicates. None of the other ethylene pathway genes reached nominal significance; their expression differences were in the expected direction (lower in infected plants) but small relative to variability. For instance, Solyc01g095080 (NAC1) had ~ 1.3-fold lower expression in STV + (log₂FC ≈ − 0.4, FDR > 0.5, see Online Resource 1). ethylene-responsive element binding protein (EREB) was expressed at very low levels in these leaf samples, and any difference was within noise. It’s worth noting that the original study reported a large fold decrease for EREB (on the order of − 5 log 2 ), but that was based on FPKM values near the detection limit. In this count-based analysis, such low-abundance transcripts were filtered or not deemed significant. On the up-regulated side, higher in STV-infected plants, the top candidate was an ethylene-responsive transcription factor (ERF017, Solyc07g007770) which showed a positive log 2 FC of + 3.97, suggesting it was expressed ~ 16× higher in infected plants (p ≈ 0.015, FDR ≈ 0.16), but again not meeting significance. Interestingly, ERF017 is part of the same gene family (AP2/ERF family) involved in ethylene signaling; however, we interpret this result cautiously, as it could be a false positive given the lack of replication support. Other genes with slight up-regulation in STV + included metabolic enzymes like glyoxalase (Solyc02g031950) and GMP synthase (Solyc01g109010), but their changes were minor and not significant. The lack of FDR-significant genes in this reanalysis suggests that either the effect of STV on host gene expression is genuinely minimal, or that the small sample size limits the power to detect subtle changes. Expression trends in ethylene pathway genes Although no genes met significance, it is informative to examine the expression patterns of those ethylene-related genes to see if the preseted data align in direction and magnitude. Figure 2a shows the normalized expression for three key genes, ACO1, NAC1, and ERF (EREB), across the STV + and STV − samples. In all cases, the mean expression in STV-infected plants was lower than in uninfected controls, matching the original finding of downregulation. However, the replication variance was large. For example, ACO1 expression in STV-free samples ranged nearly 4-fold between replicates, and one control plant had much higher ACO1 levels than the others, whereas STV-infected samples were more consistently low. NAC1 showed a similar pattern. Accordingly, two of the STV-free plants had higher NAC1 expression than any of the infected plants, but one STV-free plant had levels comparable to infected. ERF/EREB genes were expressed at low levels, making it difficult to draw conclusions; at least one ERF family gene (ERF017) was paradoxically higher in infected plants, though that may be a separate stress-related factor, whereas the specific EREB gene reported originally was barely detected here. In addition to these targets, it was observed that a pathogenesis-related protein gene (Solyc12g009240) was among the top down-regulated genes in STV-infected plants (Table 2 ). PR proteins are often associated with defense responses. A decrease in a PR transcript might indicate a slightly lower basal defense activation in STV-infected plants. Another gene of interest was WRKY6 (a defense/stress-responsive transcription factor), which also showed a modest decrease in STV-infected plants (log₂FC ≈ − 0.8, nominal p ~ 0.1, FDR ~ 0.9). These slight downward trends in defense-related transcripts could imply that STV-infected plants experience a marginally reduced defense signaling environment, potentially due to lower ethylene. To visualize these patterns more globally, Fig. 2b confirms that virtually all genes cluster around log 2 FC = 0, with very few having low p -values. No points cross the significance threshold lines. The genes with the lowest p -values are labeled for reference. Figure 2b encapsulates that there is no clear differential expression signal attributable to STV infection under the conditions tested. This is in contrast to a typical biotic stress like an acute virus or pathogen infection, where hundreds of plant genes might show significant changes. In the case of STV, the host’s transcriptomic homeostasis appears largely maintained. Figure 2 Expression of selected ethylene pathway genes in STV-free vs. STV-infected tomato leaves. a) Genes not present in the analysis (low/filtered in this experiment): ACS2, ACS4, ACO5, ACO7, E4. No ethylene genes passed FDR 0.05 in this contrast. The most notable trend is ACO1 with shrunk LFC ≈ − 1.47 (raw p ≈ 3.7e − 05, padj ≈ 0.225), but not significant at FDR 5%. All three genes show a trend of lower expression in STV-infected plants, aligning with the original report of ethylene pathway downregulation (b), but the differences are not statistically significant Absence of a strong STV transcriptomic signature in global patterns Given the lack of clear DEGs, we further asked whether unsupervised analyses could detect any subtle but concerted changes in the transcriptome due to STV. The first principal component (PC1) captured ~ 58% of the total variance and interestingly separated one of the STV-free samples (labeled “STV − _rep1” in Fig. 3), away from all other samples. This indicates that this particular replicate had an atypical expression profile compared to the rest. Importantly, the other five samples, including all three STV-infected and the remaining two STV-free clustered relatively closer together. PC2, accounting for ~ 20% variance, separated one STV-infected sample (STV + _rep2) somewhat from the others on the vertical axis. However, there was no clear clustering of samples by STV status. Two of the STV-infected replicates were actually closer to the STV-free cluster center than the outlier STV-free sample was. Figure 3 Principal component analysis of the six transcriptomes. PC1 vs. PC2 plot shows that samples do not separate strongly by infection status To further verify this, Euclidean distances measured between samples based on all genes’ expression. This clustering (Fig. 4 did not group samples strictly by infection status. One STV-infected sample clustered with two STV-free samples, while the other two STV-infected samples formed another cluster with the remaining STV-free sample. Essentially, the samples did not segregate into two distinct clades by treatment, reinforcing that the transcriptomic differences due to STV are subtle due to possible other sources of variation. It is worth noting that all plants were of the same cultivar (M82) and grown simultaneously, so genotype and major growth conditions are controlled. The variability could stem from slight differences in physiological state. Regardless, any STV-induced expression program, if it exists, is not strong enough to dominate these background differences. Figure 4 Between samples measurements. a) Euclidean distances between samples computed from VST expression profiles, displayed as a (b) clustered heatmap with dendrograms for rows and columns. Replicates tend to cluster by condition, providing an additional QC check on sample similarity and potential outliers Next, the top genes examined and ranked by smallest p -value (Fig. 4b). This is a way to show the expression pattern of the most responsive genes, even if they are not significant after FDR correction. Clustering within this finding revealed a minor pattern. In this regard, a subset of genes that were mostly those down-regulated in STV + tend to be co-expressed, and STV-infected samples show slightly lower expression for many of these relative to STV-free. Particularly, several ethylene-related and defense genes like ACO1, PR protein, and WRKY6 appear in this top-50 list and cluster together. However, the consistency was not perfect. For example, one STV-free sample had already low expression of these genes, similar to the infected group. As a result, the infected samples do not form a perfectly distinct cluster separate from all controls on this heatmap either, one control clusters among infected for these genes. Gene Ontology enrichment analysis Consistent with the minimal differential expression detected, Gene Ontology (GO) enrichment analysis did not yield any significant results. We attempted GO Biological Process enrichment on the set of genes with relaxed criteria, the 100 genes with lowest p -values, or those with nominal p < 0.05). No GO terms passed an FDR < 0.05 threshold. The top-scoring GO terms were related to general stress responses and hormone metabolic processes, which is unsurprising given the identity of the genes involved. For example, “ethylene biosynthetic process” (GO:0009693) and “response to fungus” (GO:0009620) appeared in the list due to genes like ACO1 and PR proteins, but with enrichment p ≈ 0.1. With only one or two genes driving each term, these enrichments were not statistically robust. Discussion This comprehensive reanalysis of the tomato transcriptome under STV infection provides a refined picture of host-virus interactions for persistent viruses. The results indicate that STV causes minimal changes in leaf gene expression in S. lycopersicum cv. M82, reinforcing the notion that STV is a truly low-impact symbiont at the molecular level. This finding aligns with STV’s classification as a persistent virus. Such viruses are characterized by stealthy infections that generally do not trigger strong plant defense responses or developmental disturbances. Our findings both confirm and temper the conclusions of the original GEO study (Fukuhara et al., 2019). The previous report noted that STV-infected plants had reduced expression of ethylene biosynthesis/signaling genes and proposed this as a mechanism for the observed taller, higher-yield phenotype [ 6 ]. The presented reanalysis did detect the same trend, ethylene pathway components ACO1, NAC1, and an ethylene-responsive transcription factor were expressed at lower levels in STV + plants, but these differences were not statistically significant after accounting for variance and multiple tests. In other words, we validate the direction of change for ethylene-related genes but find that the magnitude of change is modest and comes with considerable uncertainty. ACO1 transcript levels were on average ~ 2- to 3-fold lower in infected plants, which is consistent with the original fold-change estimate. However, due to one STV-free sample having atypically low ACO1 comparable to infected plants, the difference did not reach significance in our analysis. This highlights how a small sample size can lead to disparate interpretations. The original analysis did not report an FDR or adjust for multiple testing, so their identification of differentially expressed genes was effectively based on nominal p -values or fold-change cutoffs [ 6 ]. Our reanalysis, applying stricter statistics, suggests that none of the expression changes are robust at a genome-wide 5% false discovery rate. It is very possible that some changes like ACO1 downregulation are real but subtle, requiring more replication to distinguish from noise. Assuming STV truly lowers ACO1 (ACC oxidase) expression and possibly other ethylene-related factors in tomato, what does this mean for the plant? Ethylene is known to inhibit stem elongation and promote more bushy growth in many plants. ACO1 is a key enzyme in ethylene biosynthesis, converting ACC to ethylene gas. A reduction in ACO1 could lead to less ethylene production in STV-infected plants. Indeed, Fukuhara et al. (2019) inferred that STV-infected M82 plants might have lower endogenous ethylene, alleviating ethylene’s growth-suppressing effect and thus allowing plants to grow taller [ 6 ]. Our data neither conclusively prove nor refute this mechanism, but they are consistent with it. STV-infected plants do not show any increase in ethylene pathway genes, and if anything, the trend is toward a decrease. The fact that ethylene-related transcripts were among the largest fold-changes suggests that ethylene signaling is indeed the primary pathway impacted by STV. Additionally, a couple of defense-related genes like PR protein and WRKY6 showed slight downregulation in STV-infected plants, which could be a downstream consequence of altered ethylene signaling, since ethylene often modulates pathogen response pathways [ 18 ]. It is also useful to consider what we did not observe. The original study had reported upregulation of a few non-coding RNAs and other genes in STV-infected plants, but in our analysis, most of those either did not appear or were filtered out due to low counts. For example, transcripts denoted as LOC101264228 and LOC101260727 showed dramatic fold increases in the original data. In our pipeline, such transcripts had counts near zero in all samples and were likely excluded. It’s possible that differences in annotation (Ensembl vs. the ITAG2.4 used originally) account for some discrepancies. Nonetheless, none of the highly expressed genes or major metabolic enzymes were differently expressed due to STV. This reinforces that STV’s effect is narrow and subtle, mainly touching a small set of regulatory genes rather than orchestrating a large reprogramming. The results resonate with the idea that many persistent viruses are almost symbiotic, inducing negligible stress on the host. Accordingly, persistent endornaviruses in plants usually do not cause obvious symptoms; transcriptomic studies have found only limited changes. A study on common bean with an endornavirus infection identified on the order of 132 differentially expressed genes, which were enriched in oxidation-reduction processes [ 19 ]. While that number is greater than zero, it’s still relatively small compared to what one sees in acute virus infections. In STV-tomato system, we found zero significant DEGs, suggesting an even milder interaction, or simply reflecting lower statistical power. Either way, the lack of a strong immune response in STV-infected plants is notable. There was no induction of salicylic acid (SA) or jasmonic acid (JA) defense marker genes, no activation of RNA interference genes, which are commonly observed in acute viral infections. This could indicate that defense responses are suppressed or bypassed. Persistent viruses like STV have long evolutionary histories with their hosts and may have evolved mechanisms to avoid triggering host defenses. Alternatively, the mutualistic outcomes might suggest that the host has adapted to tolerate or even integrate the virus’s presence, possibly using it as a sort of internal regulator. The findings also highlight that downregulation of ethylene signaling is not a common feature of all persistent plant viruses. Accordingly, the bean endornavirus study mentioned earlier did not report ethylene pathway changes, but rather changes in redox and stress-related genes [ 19 ]. The authors of the original STV study themselves noted that the ethylene effect seems unique to STV in tomato. Ethylene has broad effects on plant development; by nudging ethylene levels downward, STV might be indirectly promoting a growth-favorable hormone balance. It is conceivable that in the specific genetic background of M82, slightly reduced ethylene could translate to noticeable phenotypic changes like taller vines and more fruit. Our data cannot directly address fruit-specific processes since only leaves were sampled, but one might predict STV could slightly delay ripening. This reanalysis highlights the value of using updated genomic resources and statistical methods. By mapping to the SL3.0 reference, potential alignment errors reduced. Indeed, we achieved > 80% read assignment to genes, indicating a comprehensive capture of transcripts. Although it’s unlikely this would qualitatively change the conclusions given the modest effect sizes. The shift from FPKM to count-based analysis with DESeq2 provided error estimates. It controlled the false discovery rate, which is crucial in a genome-wide study. The finding of ‘none of them meet FDR 0.05’ doesn’t invalidate the primaty observations but suggests caution, the changes are borderline and could be due to noise. For future experiments, an increased sample size would greatly improve power to detect subtle shifts. It is plausible that STV’s effect on gene expression might be more pronounced at a specific stage or in specific tissues. The current data was leaf-only at one month old, which might have been before major ethylene-mediated processes kicked in. It is also worth noting that the tomato cultivar M82 is a lab/reference cultivar often used in research. It is possible that in different tomato genotypes, the response to STV might vary. Some cultivars or wild relatives could mount a stronger response or have different hormone balances. Exploring STV infection in other tomato varieties or species could be informative. Given the seed transmissibility, one could envision breeders intentionally maintaining STV in lines if it indeed boosts yield without negative effects. This reanalysis of the RNA-Seq dataset GSE137303 reveals that Southern tomato virus infection causes no major transcriptional reprogramming in tomato leaves, with the possible exception of dampening the expression of a few ethylene-related genes. Unlike many pathogenic viruses, STV does not trigger broad defense responses or stress pathways in the host, consistent with its asymptomatic nature. The slight downregulation of ethylene biosynthesis and signaling components in STV-infected plants provides a plausible explanation for the previously observed growth and yield benefits. STV appears to be a near-neutral or mildly mutualistic inhabitant of the tomato plant, exerting minimal burden on host cellular processes. This work demonstrates the value of updated genomic resources and methods in reassessing host-virus interactions and emphasizes that biological replicates and proper statistical controls are crucial for distinguishing true biological effects from noise. For future investigations, examining STV’s impact under various conditions is recommend to determine if its mutualistic effects are context-dependent. It would also be worthwhile to directly measure hormone levels and fitness outcomes in presence/absence of STV to solidify causal links. Declarations Funding: The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. Competing Interests: The authors have no relevant financial or non-financial interests to disclose. Acknowledgments: Not applicable. Author Contribution: All authors contributed to the study conception and design. Material preparation, data collection and analysis were performed by Gholamhossein Badeli and Kami Kaboosi. The first draft of the manuscript was written by Gholamhossein Badeli, Kami Kaboosi, and Saeed Nasrollanejad, and all authors commented on previous versions of the manuscript. The manuscript is suppervised by Saeed Nasrollanejad. All authors read and approved the final manuscript. Data Availability: The datasets analysed during the current study are available in the NCBI’s Gene Expression Omnibus (GEO; accession number: GSE137303), BioProject, and Sequence Read Archives (SRA; accession number: PRJNA565137) repositories. Ethics approval: This is an bioinformatics and data analysis study, and no ethical approval is required. Consent to participate: Not applicable. Consent to publish: Not applicable. Clinical trial: Not applicable. References Takahashi H, Fukuhara T, Kitazawa H, Kormelink R (2019) Virus Latency and the Impact on Plants. Front Microbiol 10:2764. https://doi.org/10.3389/FMICB.2019.02764 Harper A, Vijayakumar V, Ouwehand AC et al (2021) Viral Infections, the Microbiome, and Probiotics. Front Cell Infect Microbiol 10:596166. https://doi.org/10.3389/FCIMB.2020.596166/XML Escalante C, Sanz-Saez A, Jacobson A et al (2024) Plant virus transmission during seed development and implications to plant defense system. Front Plant Sci 15:1385456. https://doi.org/10.3389/FPLS.2024.1385456 Sabanadzovic S, Valverde RA, Brown JK et al (2009) Southern tomato virus: The link between the families Totiviridae and Partitiviridae. Virus Res 140:130–137. https://doi.org/10.1016/J.VIRUSRES.2008.11.018 González LE, Peiró R, Rubio L, Galipienso L (2021) Persistent Southern Tomato Virus (STV) Interacts with Cucumber Mosaic and/or Pepino Mosaic Virus in Mixed- Infections Modifying Plant Symptoms, Viral Titer and Small RNA Accumulation. Microorganisms 9:689. https://doi.org/10.3390/MICROORGANISMS9040689 Fukuhara T, Tabara M, Koiwa H, Takahashi H (2020) Effect of asymptomatic infection with southern tomato virus on tomato plants. Arch Virol 165:11–20. https://doi.org/10.1007/S00705-019-04436-1 Liu M, Pirrello J, Chervin C et al (2015) Ethylene Control of Fruit Ripening: Revisiting the Complex Network of Transcriptional Regulation. Plant Physiol 169:2380. https://doi.org/10.1104/PP.15.01361 Xu C, Liu X, Shen G et al (2023) Time-series transcriptome provides insights into the gene regulation network involved in the icariin-flavonoid metabolism during the leaf development of Epimedium pubescens. Front Plant Sci 14:1183481. https://doi.org/10.3389/FPLS.2023.1183481/XML Love MI, Huber W, Anders S (2014) Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biology 2014 15:12 15:550-. https://doi.org/10.1186/S13059-014-0550-8 Liao Y, Smyth GK, Shi W (2019) The R package Rsubread is easier, faster, cheaper and better for alignment and quantification of RNA sequencing reads. Nucleic Acids Res 47:e47–e47. https://doi.org/10.1093/NAR/GKZ114 Liao Y, Smyth GK, Shi W (2014) FeatureCounts: An efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30:923–930. https://doi.org/10.1093/BIOINFORMATICS/BTT656 Benjamini Y, Hochberg Y (1995) Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. J R Stat Soc Ser B Stat Methodol 57:289–300. https://doi.org/10.1111/J.2517-6161.1995.TB02031.X research GOC-N acids 2004 undefined The Gene Ontology (GO) database and informatics resource. academic.oup.com Yu G, Wang LG, Han Y, He QY (2012) clusterProfiler: an R Package for Comparing Biological Themes Among Gene Clusters. OMICS 16:284. https://doi.org/10.1089/OMI.2011.0118 Durinck S, Moreau Y, Kasprzyk A et al (2005) BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis. Bioinformatics 21:3439–3440. https://doi.org/10.1093/BIOINFORMATICS/BTI525 Durinck S, Spellman PT, Birney E, Huber W (2009) Mapping identifiers for the integration of genomic datasets with the R/Bioconductor package biomaRt. Nature Protocols 2009 4:8 4:1184–1191. https://doi.org/10.1038/nprot.2009.97 Smedley D, Haider S, Ballester B et al (2009) BioMart – biological queries made easy. BMC Genomics 2009 10:1 10:22-. https://doi.org/10.1186/1471-2164-10-22 Rosado D, Ackermann A, Spassibojko O et al (2021) WRKY transcription factors and ethylene signaling modify root growth during the shade-avoidance response. Plant Physiol 188:1294. https://doi.org/10.1093/PLPHYS/KIAB493 Khankhum S, Sela N, Osorno JM, Valverde RA (2016) RNAseq analysis of endornavirus-infected vs. endornavirus-free common bean (Phaseolus Vulgaris) cultivar black turtle soup. Front Microbiol 7:1905. https://doi.org/10.3389/FMICB.2016.01905/FULL Additional Declarations No competing interests reported. Supplementary Files OnlineResource1.csv Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8823240","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":591621188,"identity":"d3200f36-3d95-44e7-a663-ee9368081a57","order_by":0,"name":"Gholamhossein Badeli","email":"","orcid":"","institution":"Islamic Azad University","correspondingAuthor":false,"prefix":"","firstName":"Gholamhossein","middleName":"","lastName":"Badeli","suffix":""},{"id":591621189,"identity":"23b58932-d789-4ee9-948f-c679187c6259","order_by":1,"name":"Kami Kaboosi","email":"","orcid":"","institution":"Islamic Azad University","correspondingAuthor":false,"prefix":"","firstName":"Kami","middleName":"","lastName":"Kaboosi","suffix":""},{"id":591621190,"identity":"720aa550-dc05-4e8d-aa92-ea65640e41be","order_by":2,"name":"Saeed Nasrollanejad","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABBElEQVRIiWNgGAWjYBAC+QYwZSEDIg0Y2GwYGCQIaDE4AKYkeKBa0ojQwoCkhYGB7TARWqSbH3+uqJDg4Z/dfKDgR9n5xP7ZzQcfMNTYROPSIj/nmJnkmTMSPBJ3jiUY9py7nTjjzrFkA4ZjabkNuPTcSDBjbGwDOuxGjoEBb9vtxIYbOWYSjA2H8WhJ//wRpEX+Rv4Hw79t5xLnE9aSYyAJ0mJwI4fBmLftQOIGQlqAKsskG4B+MbyRZmAscy7ZeOONtGSDBDx+kZ+RvvljQ4WNnNyN5GeGb8rsZOfdSD744EONDW6HIQE2UCQ5glUmEKEcBJgfAAl7IhWPglEwCkbBCAIAEjteKheyN08AAAAASUVORK5CYII=","orcid":"","institution":"Gorgan University of Agriculture Sciences and Natural Resources","correspondingAuthor":true,"prefix":"","firstName":"Saeed","middleName":"","lastName":"Nasrollanejad","suffix":""}],"badges":[],"createdAt":"2026-02-08 17:08:44","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8823240/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8823240/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":102833409,"identity":"5637b02c-df10-44d7-a973-896a8eca4fcc","added_by":"auto","created_at":"2026-02-17 10:28:28","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":445080,"visible":true,"origin":"","legend":"\u003cp\u003e(a) Gene-wise dispersion estimates (black), the fitted mean–dispersion trend (red), and the final MAP-moderated dispersions used in modeling (blue) as a function of mean normalized counts. This diagnostic from DESeq2 assesses how dispersion shrinks toward the fitted trend, improving stability for low-count genes in the STV+ vs STV− contrast. (b) Total assigned read counts per sample (featureCounts → DESeq2). shown as bars. Samples are S. lycopersicum leaves processed through Rsubread alignment, featureCounts quantification, and DESeq2; counts are used downstream for normalization and vst. Similar library sizes across replicates support comparability prior to differential analysis\u003c/p\u003e","description":"","filename":"fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-8823240/v1/890104dabaad753ddd46dc4e.png"},{"id":102833413,"identity":"cf6870e2-7c74-4a6b-aaa7-c4168fab1e93","added_by":"auto","created_at":"2026-02-17 10:28:28","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":551421,"visible":true,"origin":"","legend":"\u003cp\u003eExpression of selected ethylene pathway genes in STV-free vs. STV-infected tomato leaves. a) Genes not present in the analysis (low/filtered in this experiment): ACS2, ACS4, ACO5, ACO7, E4. No ethylene genes passed FDR 0.05 in this contrast. The most notable trend is ACO1 with shrunk LFC ≈ −1.47 (raw p ≈ 3.7e−05, padj ≈ 0.225), but not significant at FDR 5%. All three genes show a trend of lower expression in STV-infected plants, aligning with the original report of ethylene pathway downregulation (b), but the differences are not statistically significant\u003c/p\u003e","description":"","filename":"fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-8823240/v1/c6350b807cf5e5dad917f6bb.png"},{"id":102833410,"identity":"bf47c019-7019-45ab-92c3-49d45ad0b559","added_by":"auto","created_at":"2026-02-17 10:28:28","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":144029,"visible":true,"origin":"","legend":"\u003cp\u003ePrincipal component analysis of the six transcriptomes. PC1 vs. PC2 plot shows that samples do not separate strongly by infection status\u003c/p\u003e","description":"","filename":"Fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-8823240/v1/c965ec3ab9bd4adabb9289a9.png"},{"id":102833411,"identity":"c0ccce2d-893b-4645-b904-d6dd75643700","added_by":"auto","created_at":"2026-02-17 10:28:28","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1028184,"visible":true,"origin":"","legend":"\u003cp\u003eBetween samples measurements. a) Euclidean distances between samples computed from VST expression profiles, displayed as a (b) clustered heatmap with dendrograms for rows and columns. Replicates tend to cluster by condition, providing an additional QC check on sample similarity and potential outliers\u003c/p\u003e","description":"","filename":"fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-8823240/v1/c29102ac378a8ac1b2708490.png"},{"id":106401612,"identity":"924f36c6-ce39-40ca-b70f-16a6148f84a7","added_by":"auto","created_at":"2026-04-08 09:08:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2970943,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8823240/v1/fa2b4f6c-2665-4d9b-a9d0-927b3bbdb862.pdf"},{"id":102833412,"identity":"d75fad4b-f8f0-4b5a-9185-93355b059033","added_by":"auto","created_at":"2026-02-17 10:28:28","extension":"csv","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":3116138,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineResource1.csv","url":"https://assets-eu.researchsquare.com/files/rs-8823240/v1/4deac2efe6fafd4ecb1c7df2.csv"}],"financialInterests":"No competing interests reported.","formattedTitle":"Reanalysis of the Tomato Transcriptome Revealed Host Responses to Southern Tomato Virus Infection","fulltext":[{"header":"Introduction","content":"\u003cp\u003eDouble-stranded RNA viruses that persistently infect plants have been identified in many crop species [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Unlike acute or chronic viral infections that cause obvious disease symptoms, persistent viruses typically establish long-term, low-titer infections with no apparent harm to the host [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. These viruses are efficiently seed-transmitted and can spread through plant populations silently, leading some researchers to consider them as \u003cem\u003emutualistic\u003c/em\u003e or commensal symbionts rather than pathogens [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Most persistent plant viruses characterized to date belong to the families \u003cem\u003ePartitiviridae\u003c/em\u003e or \u003cem\u003eEndornaviridae\u003c/em\u003e, and they have evolved unique lifestyles distinct from those of lytic plant viruses.\u003c/p\u003e \u003cp\u003eSouthern tomato virus (STV) is a persistent virus of tomato that was first characterized in 2009 by Sabanadzovic et al. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. STV contains a monopartite dsRNA genome (3.5 kb) encoding two overlapping open reading frames and has been classified in the genus \u003cem\u003eAmalgavirus\u003c/em\u003e [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Initial reports detected STV in tomatoes in plants showing stunted, bushy growth and fruit discoloration. However, STV is also frequently found in asymptomatic tomato plants or in mixed infections with acute viruses like Pepino mosaic virus. The prevalence of STV in diverse tomato germplasm without obvious disease outbreaks suggested that STV might have negligible or even beneficial effects on its host.\u003c/p\u003e \u003cp\u003eTo directly evaluate STV\u0026rsquo;s impact, researchers compared STV-infected vs. STV-free lines of the tomato cultivar M82 under controlled conditions [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Intriguingly, STV-infected plants were taller and produced more fruits and seeds than their STV-free counterparts. The germination rate of seeds from infected plants was also higher, leading the authors to propose that STV confers a \u003cem\u003emutualistic advantage\u003c/em\u003e to the host. Importantly, transcriptome analysis of leaves from these plants indicated that STV infection induces changes in host gene expression. In particular, genes involved in the ethylene biosynthesis and signaling pathway were found to be downregulated in STV-infected tomatoes. Ethylene is a plant hormone that can restrict vegetative growth and promote fruit ripening [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]; thus, lower ethylene activity in STV-infected plants could mechanistically explain the taller, higher-yield phenotype of infected M82 plants [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. The original study highlighted three genes, 1-aminocyclopropane-1-carboxylate oxidase 1 (ACO1), a NAC-domain transcription factor (NAC1), and ethylene-responsive element binding protein (EREB), as key ethylene-related genes whose expression was reduced by STV infection. These findings suggested that STV might modulate the host\u0026rsquo;s hormonal signaling network to produce beneficial developmental outcomes.\u003c/p\u003e \u003cp\u003eWhile provocative, the transcriptomic evidence for STV\u0026rsquo;s effects was based on a relatively small RNA-Seq dataset analyzed with then-standard methods. The original analysis used the tomato reference genome SL2.50 (released in 2012) and a TopHat/Cufflinks pipeline to compute fragments per kilobase per million reads (FPKM). Differential expression was inferred from differences in FPKM and validated for a few genes by qPCR [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], but no genome-wide statistical test like false discovery rate correction was reported for the RNA-Seq results. In the years since, the tomato reference genome has been substantially improved (SL3.0 was released with more complete gene models), and RNA-Seq analysis methods have shifted toward read count-based workflows that offer more robust statistics [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Reanalysis of legacy RNA-Seq data with up-to-date pipelines can either confirm the original findings with greater confidence or reveal potential false positives/negatives introduced by older methods. This is particularly relevant for STV, given the subtle nature of persistent virus effects and the small sample size, the risk of both false negatives and false positives is high.\u003c/p\u003e \u003cp\u003eIn this study, a comprehensive reanalysis of the RNA-Seq dataset GSE137303 from [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] is presented, which examines STV-infected vs. STV-free tomato M82 plants. By leveraging the latest \u003cem\u003eS. lycopersicum\u003c/em\u003e reference (SL3.0), it was aimed to provide a more rigorous assessment of STV\u0026rsquo;s transcriptomic impact. Additional exploratory analyses performed, including unsupervised clustering and Gene Ontology enrichment, to situate any observed gene expression changes in a broader biological context. By comparing the reanalysis with the original findings, we seek to clarify whether STV infection reprograms host gene expression, particularly the ethylene pathway, or if the host response is minimal, a question fundamental to understanding the persistent virus\u0026ndash;plant interaction paradigm. Our results shed light on the molecular mutualism hypothesis and underscore best practices for transcriptomic studies of host-symbiont interactions in plants.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003ePlant material, dataset, and experimental design\u003c/h2\u003e \u003cp\u003ePublicly available transcriptome data from \u003cem\u003eS. lycopersicum\u003c/em\u003e differing in Southern tomato virus infection status analyzed. The RNA-Seq dataset, originally published by Fukuhara et al. (2019), is accessible in NCBI Gene Expression Omnibus under accession GSE137303. It comprises six leaf mRNA libraries obtained from one-month-old greenhouse-grown tomato plants: three biological replicates of STV-infected plants (STV\u003csup\u003e+\u003c/sup\u003e) and three replicates of STV-free control plants (STV\u003csup\u003e\u0026minus;\u003c/sup\u003e). All plants were verified for presence or absence of STV by dsRNA extraction and RT-PCR in the original study. The STV-free line was a cured sub-line of M82, and the STV-infected line was an M82 sub-line naturally harboring STV [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. This paired design allowed direct comparison of transcriptomes to infer the effect of asymptomatic STV infection.\u003c/p\u003e \u003cp\u003eFor this study, raw sequencing data were downloaded in FASTQ format. The reference genome assembly was obtained from Ensembl Plants. This reference includes improved assembly continuity and gene annotation for \u003cem\u003eS. lycopersicum\u003c/em\u003e. Corresponding reference files, including the genomic DNA FASTA and the comprehensive gene annotation GTF was obtained. Additionally, to account for viral reads, we included the STV genome sequence (GenBank KX949574) as an extra \u0026ldquo;chromosome\u0026rdquo; in the reference during read alignment. This enabled simultaneous mapping of reads to both host and virus, although as a persistent virus STV was not expected to produce abundant mRNA.\u003c/p\u003e \u003cp\u003eThe second major update was switching from an FPKM-based pipeline (TopHat to Cuffnorm) to a count-based differential expression approach. Count-based methods use the discrete read counts per gene and apply established statistical models for significance testing, offering greater sensitivity and control of false positives. The Bioconductor DESeq2 package v1.42.0 was used for this analysis, which utilizes a negative binomial generalized linear model to identify differentially expressed genes with robust estimates of biological variance [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. This approach allowed us to apply Benjamini\u0026ndash;Hochberg false discovery rate (FDR) correction across all genes, ensuring that any significant hits are not likely due to chance.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eRead preprocessing and alignment\u003c/h3\u003e\n\u003cp\u003eRaw FASTQ files from the six libraries were first subjected to quality control. Adapter sequences and low-quality bases were trimmed using the Trimmomatic tool (v0.39) with default settings, and read quality was verified with FastQC. All reads were confirmed to be 75 nt in length after trimming, as expected from the library protocol. Post-filtering, each library retained between ~\u0026thinsp;14 and 18\u0026nbsp;million reads.\u003c/p\u003e \u003cp\u003eHigh-quality reads were then aligned to the combined reference genome (tomato SL3.0 plus STV) using the Rsubread aligner (R package Rsubread, v2.13.3) [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Since the data derive from mRNA, we built a gapped index to accommodate spliced alignments from the SL3.0 reference FASTA. Alignment was performed in base-space with up to 3 mismatches allowed per read. Rsubread\u0026rsquo;s align function was set to allow soft-clipping and gapped alignment for introns. We also permitted a maximum insertion or deletion length of 5 bp to account for small indels. Reads mapping to multiple locations in the genome were handled by reporting a single best alignment to avoid inflating counts. After alignment, we used SAMtools through Rsubread to sort and index the resulting BAM files.\u003c/p\u003e\n\u003ch3\u003eGene quantification and normalization\u003c/h3\u003e\n\u003cp\u003eGene-level counting of aligned reads was carried out with featureCounts, as a part of the Rsubread package to produce a matrix of integer read counts per gene. FeatureCounts were provided with the Ensembl SL3.0 gene annotation GTF (Solanum_lycopersicum.SL3.0.61.gff3), which corresponds to ITAG3.0 gene models. FeatureCounts [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] was run in \u003cem\u003emeta-feature\u003c/em\u003e mode, meaning that for each gene or locus it summed reads across all exons, yielding one total count per gene. We allowed featureCounts to count a read if it overlapped any exon of a gene, capturing both fully spliced and pre-mRNA fragments. Reads mapping to multiple genes were excluded by featureCounts to maintain accurate counts. Multi-mapped reads that aligned to multiple loci were also not counted toward any single gene, per the default settings, since their origin is uncertain. After counting, a raw count matrix for 34,075 tomato gene features obtained, including protein-coding genes and non-coding transcripts annotated in Ensembl.\u003c/p\u003e \u003cp\u003ePrior to statistical analysis, the count data filtered and normalized. Genes with consistently low read counts were filtered out to reduce noise. Specifically, any gene with \u0026lt;\u0026thinsp;10 total counts across all six samples was removed from the dataset. Library size normalization was performed using the \u003cem\u003emedian ratio\u003c/em\u003e method implemented in DESeq2. This method computes size factors for each sample such that the median of the ratio of observed counts to expected counts (geometric mean across samples) is 1. Effectively, it accounts for differences in sequencing depth between libraries. After normalization, count data were transformed to variance-stabilized counts using DESeq2\u0026rsquo;s variance stabilizing transformation (\u003cem\u003evst)\u003c/em\u003e for use in visualization and clustering. The \u003cem\u003evst\u003c/em\u003e converts count data to a log\u003csub\u003e2\u003c/sub\u003e-scale while accounting for the mean-variance relationship, producing values that are homoscedastic and suitable for principal component analysis.\u003c/p\u003e\n\u003ch3\u003eDifferential expression analysis\u003c/h3\u003e\n\u003cp\u003eDifferential gene expression between the STV-infected and STV-free groups tested using the DESeq2 statistical framework. The model design was simple, with a single factor condition (STV\u003csup\u003e+\u003c/sup\u003e vs STV\u003csup\u003e\u0026minus;\u003c/sup\u003e). DESeq2 fits a negative binomial model for each gene\u0026rsquo;s counts and uses Wald tests to assess whether the log\u003csub\u003e2\u003c/sub\u003e fold-change (log\u003csub\u003e2\u003c/sub\u003eFC) between conditions is significantly different from zero. We included the three replicates per group to estimate within-group dispersion. Dispersions were empirically estimated and shrunk toward a trend to improve stability for genes with low counts, following DESeq2\u0026rsquo;s standard procedure.\u003c/p\u003e \u003cp\u003eFor each gene, an estimated log\u003csub\u003e2\u003c/sub\u003e fold-change (STV\u003csup\u003e+\u003c/sup\u003e / STV\u003csup\u003e\u0026minus;\u003c/sup\u003e) obtained and an associated \u003cem\u003ep\u003c/em\u003e-value. Next, a Benjamini\u0026ndash;Hochberg [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] adjustment applied to control the false discovery rate (FDR) across the ~\u0026thinsp;25,000 tested genes. Genes were declared significantly differentially expressed if they met two criteria: (1) FDR (adjusted \u003cem\u003ep\u003c/em\u003e-value)\u0026thinsp;\u0026lt;\u0026thinsp;0.05, and (2) absolute log₂FC\u0026thinsp;\u0026ge;\u0026thinsp;1, which corresponds to a\u0026thinsp;\u0026ge;\u0026thinsp;2-fold change in expression. This fold-change threshold was chosen to focus on changes likely to be biologically meaningful in the context of the relatively small sample size. Additionally, shrunken log₂ fold-changes was measured using the \u003cem\u003eapeglm\u003c/em\u003e method in DESeq2. Shrinkage of effect sizes is recommended to avoid overestimation of fold-changes for genes with low counts or high dispersion.\u003c/p\u003e \u003cp\u003eFor exploratory purposes and generating lists of candidate genes for enrichment, we also identified genes that would be considered differentially expressed under a more lenient cutoff (FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.10 or \u0026lt;\u0026thinsp;0.20). This relaxed threshold is \u003cem\u003enot\u003c/em\u003e used to claim significance, but rather to ensure we did not miss any borderline effects that might be biologically relevant in enrichment analysis or in comparisons with the original study\u0026rsquo;s findings. This was due to no genes achieved the primary significance threshold (FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05 and |log₂FC| \u0026ge; 1), highlighting the need for such exploratory follow-ups.\u003c/p\u003e\n\u003ch3\u003eFunctional annotation of genes\u003c/h3\u003e\n\u003cp\u003eTomato gene identifiers in the Ensembl database are given as Solyc numbers (like Solyc07g007770) which are not immediately intuitive biologically. To aid interpretation, we mapped each gene ID to a gene name or description where possible. The Ensembl Plants (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://plants.ensembl.org/index.html\u003c/span\u003e\u003cspan address=\"https://plants.ensembl.org/index.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) GTF annotation provides a \u0026ldquo;gene_name\u0026rdquo; attribute for many genes, which often corresponds to a common gene symbol or a descriptive name. Therefore, these gene symbols extracted, and when unavailable, alternate annotation sources like UniProt and NCBI RefSeq were searched for gene descriptions or putative functions. A custom mapping table was created so that figures and tables could display both the Solyc ID and a gene name or product description. In cases where a gene had no clear symbolic name, the Solyc identifier is retained but noted if it was predicted to encode a certain type of protein.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eExploratory data analysis and visualization\u003c/h2\u003e \u003cp\u003eExploratory analyses were performed to visualize the overall transcriptomic patterns. In this regard, a principal component analysis (PCA) conducted on the normalized gene expression data. PCA was run on the variance-stabilized expression matrix for the top 500 genes with the highest variance across the six samples. This unsupervised approach reduces the data to principal components (PCs) that capture the largest sources of variation. We examined whether the samples separated along any PC according to STV infection status. A scatter plot of the first two PCs was generated with each sample labeled and colored by condition (STV\u003csup\u003e+\u003c/sup\u003e vs STV\u003csup\u003e\u0026minus;\u003c/sup\u003e).\u003c/p\u003e \u003cp\u003eAlso, a sample distance heatmap was generated. Using the variance-stabilized data, the Euclidean distance between each pair of samples was evaluated based on the expression of all genes. The goal was to see if STV-infected samples cluster together, indicating a consistent transcriptomic signature of infection, or if samples cluster by other factors, suggesting high within-group variability.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eFunctional enrichment analysis\u003c/h3\u003e\n\u003cp\u003eTo interpret any biological themes in the data, a Gene Ontology (GO) [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] enrichment analysis was carried out for differentially expressed genes. Given the small number of genes showing any expression change, we relaxed the significance threshold to include genes with adjusted \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.1 (or even \u0026lt;\u0026thinsp;0.2) for enrichment purposes. We focused on the Biological Process (BP) branch of GO, as our interest was in physiological pathways or processes affected by STV.\u003c/p\u003e \u003cp\u003eAccordingly, enrichment was performed using the clusterProfiler package [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. GO annotations obtained for tomato genes via the biomaRt [\u003cspan additionalcitationids=\"CR16\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] interface to Ensembl Plants, retrieving all GO terms associated with each Ensembl gene ID. A custom GO term to gene mapping was constructed for \u003cem\u003eS. lycopersicum\u003c/em\u003e. For the gene list of interest, clusterProfiler was used to run a hypergeometric test to identify GO terms that are over-represented. The background universe was set as all expressed genes in our dataset. Enrichment \u003cem\u003ep\u003c/em\u003e-values were adjusted for multiple testing. We limited our analysis to GO-BP terms to avoid dilution by overly broad categories.\u003c/p\u003e\n\u003ch3\u003eReproducibility and computational environment\u003c/h3\u003e\n\u003cp\u003eAll analysis steps were performed in the R environment (v4.5.1) on a Windows 10 64-bit system. Bioconductor packages were leveraged for most tasks, ensuring compatibility and version control. Key packages were including, \u003cem\u003eRsubread\u003c/em\u003e v2.13.3 (for alignment and counting), \u003cem\u003eDESeq2\u003c/em\u003e v1.42.0 (for differential expression), \u003cem\u003etximport\u003c/em\u003e v1.20 (for any potential transcript-level import, though not used here), \u003cem\u003eclusterProfiler\u003c/em\u003e v4.10.0 (for GO analysis), \u003cem\u003eggplot2\u003c/em\u003e v3.4.0, \u003cem\u003epheatmap\u003c/em\u003e v1.0.12, and \u003cem\u003eggrepel\u003c/em\u003e v0.9.2 (for labeling plots).\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eSequencing and mapping overview\u003c/h2\u003e \u003cp\u003eEach of the six RNA-Seq libraries yielded between 14 and 20\u0026nbsp;million raw single-end reads, consistent with the sequencing depth reported in the original dataset. After quality trimming, over 99% of reads were retained as high-quality. These reads were mapped to the tomato reference genome SL3.0 (Solanum_lycopersicum.SL3.0.dna.toplevel.fa) with high alignment rate. On average, 85\u0026ndash;90% of reads per sample aligned to a unique position in the tomato genome and were assigned to an annotated gene feature (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The remaining reads either mapped to intergenic regions or could not be confidently assigned due to multi-mapping. The inclusion of the STV genome sequence in the alignment index did not affect mapping rates to tomato; only a low numbers of reads aligned to the STV dsRNA (\u0026lt;\u0026thinsp;0.01%), indicating extremely low expression or absence of poly(A)-transcripts from the virus, as expected.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;1\u003c/b\u003e \u003cem\u003e(a) Gene-wise dispersion estimates (black), the fitted mean\u0026ndash;dispersion trend (red), and the final MAP-moderated dispersions used in modeling (blue) as a function of mean normalized counts. This diagnostic from DESeq2 assesses how dispersion shrinks toward the fitted trend, improving stability for low-count genes in the STV\u0026thinsp;+\u0026thinsp;vs STV\u0026minus; contrast. (b) Total assigned read counts per sample (featureCounts \u0026rarr; DESeq2). shown as bars. Samples are\u003c/em\u003e S. lycopersicum \u003cem\u003eleaves processed through Rsubread alignment, featureCounts quantification, and DESeq2; counts are used downstream for normalization and vst. Similar library sizes across replicates support comparability prior to differential analysis\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eRead alignment and assignment statistics for RNA-Seq libraries from STV\u003csup\u003e+\u003c/sup\u003e and STV\u003csup\u003e\u0026minus;\u003c/sup\u003e tomato plants.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSample ID\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCondition\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTotal reads (M)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAssigned to gene \u003csup\u003e\u0026yen;\u003c/sup\u003e(M)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e% Assigned\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eUnassigned \u0026ndash; No feature (M)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eUnassigned \u003csup\u003e\u0026yen;\u003c/sup\u003e \u0026ndash; ambiguity (M)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSTV\u003csup\u003e\u0026minus;\u003c/sup\u003e_rep1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSTV-free\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e17.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e14.23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e80.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.43\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.98\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSTV\u003csup\u003e\u0026minus;\u003c/sup\u003e_rep2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSTV-free\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e15.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e13.48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e84.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSTV\u003csup\u003e\u0026minus;\u003c/sup\u003e_rep3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSTV-free\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e13.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e11.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e83.2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.50\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSTV\u003csup\u003e+\u003c/sup\u003e_rep1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSTV-infected\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e17.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e15.29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e85.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.37\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e2.11\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSTV\u003csup\u003e+\u003c/sup\u003e_rep2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSTV-infected\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e15.31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e13.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e89.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.74\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSTV\u003csup\u003e+\u003c/sup\u003e_rep3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSTV-infected\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e19.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e16.16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e81.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e2.08\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"7\" nameend=\"c7\" namest=\"c1\"\u003e \u003cp\u003e\u0026ldquo;Assigned to gene\u0026rdquo; indicates reads that mapped to a tomato gene feature by \u003cem\u003efeatureCounts\u003c/em\u003e. \u0026ldquo;Unassigned\u0026rdquo; categories include reads mapping outside annotated exons or those ambiguously mapped. Values in table are in millions of reads (M)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAll samples showed comparable sequencing depth and mapping success, with no obvious outlier in library preparation or sequencing quality. Thus, any differences observed in gene expression are unlikely to stem from technical artifacts. The raw counts per gene were normalized to account for the small differences in library size (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows STV\u003csup\u003e+\u003c/sup\u003e_rep3 had the highest read count at ~\u0026thinsp;19.8M, while STV\u003csup\u003e\u0026minus;\u003c/sup\u003e_rep3 had\u0026thinsp;~\u0026thinsp;14.0M; DESeq2\u0026rsquo;s size factor normalization corrected for these differences).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eDifferential gene expression analysis signed no significant hits\u003c/h2\u003e \u003cp\u003eApplying the DESeq2 differential expression analysis between STV-infected and uninfected tomato samples, no genes found that met the conventional criteria for significant differential expression. In other words, none of the ~\u0026thinsp;25,000 expressed genes showed a statistically significant change in expression due to STV infection in this dataset. This is a outstanding result, as it suggests that STV elicits little to no transcriptomic perturbation in M82 tomato leaves, within the sensitivity limits of this experiment. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e summarizes the differential expression results, listing the genes with the largest observed fold-changes, regardless of significance, in STV-infected vs. STV-free plants. Even the top-ranking candidates did not pass the FDR cutoff after multiple-testing correction. For completeness, the top up- and down-regulated genes ranked by smallest nominal \u003cem\u003ep\u003c/em\u003e-value, with their expression fold-changes and significance values are presented.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eTop candidates for differential expression between STV-infected and STV-free tomato plants (sorted by nominal p-value).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGene ID (Solyc)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGene Name / Description\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003elog₂FC (STV\u003csup\u003e+\u003c/sup\u003e vs STV\u003csup\u003e\u0026minus;\u003c/sup\u003e)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAdjusted p-value (FDR)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNotes\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSolyc12g009240.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePathogenesis-related (PR)-protein\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u0026ndash;1.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0831\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eEncodes PR protein (defense-related) \u0026ndash; highest fold change down, n.s. FDR\u0026thinsp;\u0026asymp;\u0026thinsp;0.08\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSolyc07g049530.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eACC oxidase 1 (ACO1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u0026ndash;1.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.225\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eEthylene biosynthesis enzyme (down in STV\u003csup\u003e+\u003c/sup\u003e)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSolyc08g067160.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUncharacterized protein (DUF247)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u0026ndash;1.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.736\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eExpressed protein of unknown function\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSolyc04g008330.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ebeta-Galactosidase 4-like\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u0026ndash;1.31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCell wall degradation enzyme family\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSolyc02g086990.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePGIP (Polygalacturonase inhibitor)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u0026ndash;0.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDefense-related cell wall protein\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSolyc07g007770\u003cb\u003e.2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eERF017 (Ethylene-responsive TF)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;3.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.157\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eEthylene-responsive factor (up in STV\u003csup\u003e+\u003c/sup\u003e) nominal p\u0026thinsp;=\u0026thinsp;0.0157 (highest fold up)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSolyc02g031950.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGlyoxalase II (GLO2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;1.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eStress-related metabolism enzyme, slight up\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSolyc01g109010.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGMP synthase [glutamine-hydrolyzing]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;1.09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNucleotide biosynthesis enzyme, slight up\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"5\"\u003e\u003cem\u003eACO1\u0026thinsp;=\u0026thinsp;1-aminocyclopropane-1-carboxylate oxidase 1; ERF017 is an ethylene-responsive transcription factor. log₂FC is shown as shrunken DESeq2 estimate. A negative log₂FC indicates lower expression in STV-infected plants\u003c/em\u003e\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAs shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, the largest fold-change observed was a \u0026minus;\u0026thinsp;1.12 log\u003csub\u003e2\u003c/sub\u003e decrease for a pathogenesis-related protein gene (Solyc12g009240), which corresponds to roughly a halving of expression in STV-infected plants compared to STV-free. This gene had a \u003cem\u003ep\u003c/em\u003e-value of ~\u0026thinsp;4.6\u0026times;10\u003csup\u003e\u0026ndash;6\u003c/sup\u003e and an FDR of 0.083; it narrowly missed the 5% FDR threshold. ACO1 (Solyc07g049530), \u0026lsquo;\u003cem\u003ethe key ethylene biosynthetic gene identified in the original study\u003c/em\u003e\u0026rsquo;, showed a log₂FC of \u0026minus;\u0026thinsp;1.47, but its adjusted \u003cem\u003ep\u003c/em\u003e-value was 0.225, reflecting high variance among replicates. None of the other ethylene pathway genes reached nominal significance; their expression differences were in the expected direction (lower in infected plants) but small relative to variability. For instance, Solyc01g095080 (NAC1) had\u0026thinsp;~\u0026thinsp;1.3-fold lower expression in STV\u003csup\u003e+\u003c/sup\u003e (log₂FC \u0026asymp; \u0026minus;\u0026thinsp;0.4, FDR\u0026thinsp;\u0026gt;\u0026thinsp;0.5, \u003cem\u003esee\u003c/em\u003e Online Resource 1). ethylene-responsive element binding protein (EREB) was expressed at very low levels in these leaf samples, and any difference was within noise. It\u0026rsquo;s worth noting that the original study reported a large fold decrease for EREB (on the order of \u0026minus;\u0026thinsp;5 log\u003csub\u003e2\u003c/sub\u003e), but that was based on FPKM values near the detection limit. In this count-based analysis, such low-abundance transcripts were filtered or not deemed significant.\u003c/p\u003e \u003cp\u003eOn the up-regulated side, higher in STV-infected plants, the top candidate was an ethylene-responsive transcription factor (ERF017, Solyc07g007770) which showed a positive log\u003csub\u003e2\u003c/sub\u003eFC of +\u0026thinsp;3.97, suggesting it was expressed\u0026thinsp;~\u0026thinsp;16\u0026times; higher in infected plants (p\u0026thinsp;\u0026asymp;\u0026thinsp;0.015, FDR\u0026thinsp;\u0026asymp;\u0026thinsp;0.16), but again not meeting significance. Interestingly, ERF017 is part of the same gene family (AP2/ERF family) involved in ethylene signaling; however, we interpret this result cautiously, as it could be a false positive given the lack of replication support. Other genes with slight up-regulation in STV\u003csup\u003e+\u003c/sup\u003e included metabolic enzymes like glyoxalase (Solyc02g031950) and GMP synthase (Solyc01g109010), but their changes were minor and not significant.\u003c/p\u003e \u003cp\u003eThe lack of FDR-significant genes in this reanalysis suggests that either the effect of STV on host gene expression is genuinely minimal, or that the small sample size limits the power to detect subtle changes.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eExpression trends in ethylene pathway genes\u003c/h2\u003e \u003cp\u003eAlthough no genes met significance, it is informative to examine the expression patterns of those ethylene-related genes to see if the preseted data align in direction and magnitude. Figure\u0026nbsp;2a shows the normalized expression for three key genes, ACO1, NAC1, and ERF (EREB), across the STV\u003csup\u003e+\u003c/sup\u003e and STV\u003csup\u003e\u0026minus;\u003c/sup\u003e samples. In all cases, the mean expression in STV-infected plants was lower than in uninfected controls, matching the original finding of downregulation. However, the replication variance was large. For example, ACO1 expression in STV-free samples ranged nearly 4-fold between replicates, and one control plant had much higher ACO1 levels than the others, whereas STV-infected samples were more consistently low. NAC1 showed a similar pattern. Accordingly, two of the STV-free plants had higher NAC1 expression than any of the infected plants, but one STV-free plant had levels comparable to infected. ERF/EREB genes were expressed at low levels, making it difficult to draw conclusions; at least one ERF family gene (ERF017) was paradoxically higher in infected plants, though that may be a separate stress-related factor, whereas the specific EREB gene reported originally was barely detected here.\u003c/p\u003e \u003cp\u003eIn addition to these targets, it was observed that a pathogenesis-related protein gene (Solyc12g009240) was among the top down-regulated genes in STV-infected plants (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). PR proteins are often associated with defense responses. A decrease in a PR transcript might indicate a slightly lower basal defense activation in STV-infected plants. Another gene of interest was WRKY6 (a defense/stress-responsive transcription factor), which also showed a modest decrease in STV-infected plants (log₂FC \u0026asymp; \u0026minus;\u0026thinsp;0.8, nominal p\u0026thinsp;~\u0026thinsp;0.1, FDR\u0026thinsp;~\u0026thinsp;0.9). These slight downward trends in defense-related transcripts could imply that STV-infected plants experience a marginally reduced defense signaling environment, potentially due to lower ethylene.\u003c/p\u003e \u003cp\u003eTo visualize these patterns more globally, Fig.\u0026nbsp;2b confirms that virtually all genes cluster around log\u003csub\u003e2\u003c/sub\u003eFC\u0026thinsp;=\u0026thinsp;0, with very few having low \u003cem\u003ep\u003c/em\u003e-values. No points cross the significance threshold lines. The genes with the lowest \u003cem\u003ep\u003c/em\u003e-values are labeled for reference. Figure\u0026nbsp;2b encapsulates that there is no clear differential expression signal attributable to STV infection under the conditions tested. This is in contrast to a typical biotic stress like an acute virus or pathogen infection, where hundreds of plant genes might show significant changes. In the case of STV, the host\u0026rsquo;s transcriptomic homeostasis appears largely maintained.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;2\u003c/b\u003e \u003cem\u003eExpression of selected ethylene pathway genes in STV-free vs. STV-infected tomato leaves. a) Genes not present in the analysis (low/filtered in this experiment): ACS2, ACS4, ACO5, ACO7, E4. No ethylene genes passed FDR 0.05 in this contrast. The most notable trend is ACO1 with shrunk LFC\u0026thinsp;\u0026asymp;\u0026thinsp;\u0026minus;\u0026thinsp;1.47 (raw p\u0026thinsp;\u0026asymp;\u0026thinsp;3.7e\u0026thinsp;\u0026minus;\u0026thinsp;05, padj\u0026thinsp;\u0026asymp;\u0026thinsp;0.225), but not significant at FDR 5%. All three genes show a trend of lower expression in STV-infected plants, aligning with the original report of ethylene pathway downregulation (b), but the differences are not statistically significant\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eAbsence of a strong STV transcriptomic signature in global patterns\u003c/h2\u003e \u003cp\u003eGiven the lack of clear DEGs, we further asked whether unsupervised analyses could detect any subtle but concerted changes in the transcriptome due to STV. The first principal component (PC1) captured\u0026thinsp;~\u0026thinsp;58% of the total variance and interestingly separated one of the STV-free samples (labeled \u0026ldquo;STV\u003csup\u003e\u0026minus;\u003c/sup\u003e_rep1\u0026rdquo; in Fig.\u0026nbsp;3), away from all other samples. This indicates that this particular replicate had an atypical expression profile compared to the rest. Importantly, the other five samples, including all three STV-infected and the remaining two STV-free clustered relatively closer together. PC2, accounting for ~\u0026thinsp;20% variance, separated one STV-infected sample (STV\u003csup\u003e+\u003c/sup\u003e_rep2) somewhat from the others on the vertical axis. However, there was no clear clustering of samples by STV status. Two of the STV-infected replicates were actually closer to the STV-free cluster center than the outlier STV-free sample was.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;3\u003c/b\u003e \u003cem\u003ePrincipal component analysis of the six transcriptomes. PC1 vs. PC2 plot shows that samples do not separate strongly by infection status\u003c/em\u003e\u003c/p\u003e \u003cp\u003eTo further verify this, Euclidean distances measured between samples based on all genes\u0026rsquo; expression. This clustering (Fig.\u0026nbsp;4 did not group samples strictly by infection status. One STV-infected sample clustered with two STV-free samples, while the other two STV-infected samples formed another cluster with the remaining STV-free sample. Essentially, the samples did not segregate into two distinct clades by treatment, reinforcing that the transcriptomic differences due to STV are subtle due to possible other sources of variation. It is worth noting that all plants were of the same cultivar (M82) and grown simultaneously, so genotype and major growth conditions are controlled. The variability could stem from slight differences in physiological state. Regardless, any STV-induced expression program, if it exists, is not strong enough to dominate these background differences.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;4\u003c/b\u003e \u003cem\u003eBetween samples measurements. a) Euclidean distances between samples computed from VST expression profiles, displayed as a (b) clustered heatmap with dendrograms for rows and columns. Replicates tend to cluster by condition, providing an additional QC check on sample similarity and potential outliers\u003c/em\u003e\u003c/p\u003e \u003cp\u003eNext, the top genes examined and ranked by smallest \u003cem\u003ep\u003c/em\u003e-value (Fig.\u0026nbsp;4b). This is a way to show the expression pattern of the most responsive genes, even if they are not significant after FDR correction. Clustering within this finding revealed a minor pattern. In this regard, a subset of genes that were mostly those down-regulated in STV\u003csup\u003e+\u003c/sup\u003e tend to be co-expressed, and STV-infected samples show slightly lower expression for many of these relative to STV-free. Particularly, several ethylene-related and defense genes like ACO1, PR protein, and WRKY6 appear in this top-50 list and cluster together. However, the consistency was not perfect. For example, one STV-free sample had already low expression of these genes, similar to the infected group. As a result, the infected samples do not form a perfectly distinct cluster separate from all controls on this heatmap either, one control clusters among infected for these genes.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eGene Ontology enrichment analysis\u003c/h2\u003e \u003cp\u003eConsistent with the minimal differential expression detected, Gene Ontology (GO) enrichment analysis did not yield any significant results. We attempted GO Biological Process enrichment on the set of genes with relaxed criteria, the 100 genes with lowest \u003cem\u003ep\u003c/em\u003e-values, or those with nominal p\u0026thinsp;\u0026lt;\u0026thinsp;0.05). No GO terms passed an FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05 threshold. The top-scoring GO terms were related to general stress responses and hormone metabolic processes, which is unsurprising given the identity of the genes involved. For example, \u0026ldquo;ethylene biosynthetic process\u0026rdquo; (GO:0009693) and \u0026ldquo;response to fungus\u0026rdquo; (GO:0009620) appeared in the list due to genes like ACO1 and PR proteins, but with enrichment \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026asymp;\u0026thinsp;0.1. With only one or two genes driving each term, these enrichments were not statistically robust.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis comprehensive reanalysis of the tomato transcriptome under STV infection provides a refined picture of host-virus interactions for persistent viruses. The results indicate that STV causes minimal changes in leaf gene expression in \u003cem\u003eS. lycopersicum\u003c/em\u003e cv. M82, reinforcing the notion that STV is a truly low-impact symbiont at the molecular level. This finding aligns with STV\u0026rsquo;s classification as a persistent virus. Such viruses are characterized by stealthy infections that generally do not trigger strong plant defense responses or developmental disturbances.\u003c/p\u003e \u003cp\u003eOur findings both confirm and temper the conclusions of the original GEO study (Fukuhara et al., 2019). The previous report noted that STV-infected plants had reduced expression of ethylene biosynthesis/signaling genes and proposed this as a mechanism for the observed taller, higher-yield phenotype [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. The presented reanalysis did detect the same trend, ethylene pathway components ACO1, NAC1, and an ethylene-responsive transcription factor were expressed at lower levels in STV\u003csup\u003e+\u003c/sup\u003e plants, but these differences were not statistically significant after accounting for variance and multiple tests. In other words, we validate the direction of change for ethylene-related genes but find that the magnitude of change is modest and comes with considerable uncertainty. ACO1 transcript levels were on average\u0026thinsp;~\u0026thinsp;2- to 3-fold lower in infected plants, which is consistent with the original fold-change estimate. However, due to one STV-free sample having atypically low ACO1 comparable to infected plants, the difference did not reach significance in our analysis. This highlights how a small sample size can lead to disparate interpretations. The original analysis did not report an FDR or adjust for multiple testing, so their identification of differentially expressed genes was effectively based on nominal \u003cem\u003ep\u003c/em\u003e-values or fold-change cutoffs [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Our reanalysis, applying stricter statistics, suggests that none of the expression changes are robust at a genome-wide 5% false discovery rate. It is very possible that some changes like ACO1 downregulation are real but subtle, requiring more replication to distinguish from noise.\u003c/p\u003e \u003cp\u003eAssuming STV truly lowers ACO1 (ACC oxidase) expression and possibly other ethylene-related factors in tomato, what does this mean for the plant? Ethylene is known to inhibit stem elongation and promote more bushy growth in many plants. ACO1 is a key enzyme in ethylene biosynthesis, converting ACC to ethylene gas. A reduction in ACO1 could lead to less ethylene production in STV-infected plants. Indeed, Fukuhara et al. (2019) inferred that STV-infected M82 plants might have lower endogenous ethylene, alleviating ethylene\u0026rsquo;s growth-suppressing effect and thus allowing plants to grow taller [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Our data neither conclusively prove nor refute this mechanism, but they are consistent with it. STV-infected plants do not show any \u003cem\u003eincrease\u003c/em\u003e in ethylene pathway genes, and if anything, the trend is toward a decrease. The fact that ethylene-related transcripts were among the largest fold-changes suggests that ethylene signaling is indeed the primary pathway impacted by STV. Additionally, a couple of defense-related genes like PR protein and WRKY6 showed slight downregulation in STV-infected plants, which could be a downstream consequence of altered ethylene signaling, since ethylene often modulates pathogen response pathways [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIt is also useful to consider what we did not observe. The original study had reported upregulation of a few non-coding RNAs and other genes in STV-infected plants, but in our analysis, most of those either did not appear or were filtered out due to low counts. For example, transcripts denoted as LOC101264228 and LOC101260727 showed dramatic fold increases in the original data. In our pipeline, such transcripts had counts near zero in all samples and were likely excluded. It\u0026rsquo;s possible that differences in annotation (Ensembl vs. the ITAG2.4 used originally) account for some discrepancies. Nonetheless, none of the highly expressed genes or major metabolic enzymes were differently expressed due to STV. This reinforces that STV\u0026rsquo;s effect is narrow and subtle, mainly touching a small set of regulatory genes rather than orchestrating a large reprogramming.\u003c/p\u003e \u003cp\u003eThe results resonate with the idea that many persistent viruses are almost symbiotic, inducing negligible stress on the host. Accordingly, persistent endornaviruses in plants usually do not cause obvious symptoms; transcriptomic studies have found only limited changes. A study on common bean with an endornavirus infection identified on the order of 132 differentially expressed genes, which were enriched in oxidation-reduction processes [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. While that number is greater than zero, it\u0026rsquo;s still relatively small compared to what one sees in acute virus infections. In STV-tomato system, we found zero significant DEGs, suggesting an even milder interaction, or simply reflecting lower statistical power. Either way, the lack of a strong immune response in STV-infected plants is notable. There was no induction of salicylic acid (SA) or jasmonic acid (JA) defense marker genes, no activation of RNA interference genes, which are commonly observed in acute viral infections. This could indicate that defense responses are suppressed or bypassed. Persistent viruses like STV have long evolutionary histories with their hosts and may have evolved mechanisms to avoid triggering host defenses. Alternatively, the mutualistic outcomes might suggest that the host has adapted to tolerate or even integrate the virus\u0026rsquo;s presence, possibly using it as a sort of internal regulator.\u003c/p\u003e \u003cp\u003eThe findings also highlight that downregulation of ethylene signaling is not a common feature of all persistent plant viruses. Accordingly, the bean endornavirus study mentioned earlier did not report ethylene pathway changes, but rather changes in redox and stress-related genes [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. The authors of the original STV study themselves noted that the ethylene effect seems unique to STV in tomato. Ethylene has broad effects on plant development; by nudging ethylene levels downward, STV might be indirectly promoting a growth-favorable hormone balance. It is conceivable that in the specific genetic background of M82, slightly reduced ethylene could translate to noticeable phenotypic changes like taller vines and more fruit. Our data cannot directly address fruit-specific processes since only leaves were sampled, but one might predict STV could slightly delay ripening.\u003c/p\u003e \u003cp\u003eThis reanalysis highlights the value of using updated genomic resources and statistical methods. By mapping to the SL3.0 reference, potential alignment errors reduced. Indeed, we achieved\u0026thinsp;\u0026gt;\u0026thinsp;80% read assignment to genes, indicating a comprehensive capture of transcripts. Although it\u0026rsquo;s unlikely this would qualitatively change the conclusions given the modest effect sizes. The shift from FPKM to count-based analysis with DESeq2 provided error estimates. It controlled the false discovery rate, which is crucial in a genome-wide study. The finding of \u0026lsquo;none of them meet FDR 0.05\u0026rsquo; doesn\u0026rsquo;t invalidate the primaty observations but suggests caution, the changes are borderline and could be due to noise. For future experiments, an increased sample size would greatly improve power to detect subtle shifts. It is plausible that STV\u0026rsquo;s effect on gene expression might be more pronounced at a specific stage or in specific tissues. The current data was leaf-only at one month old, which might have been before major ethylene-mediated processes kicked in.\u003c/p\u003e \u003cp\u003eIt is also worth noting that the tomato cultivar M82 is a lab/reference cultivar often used in research. It is possible that in different tomato genotypes, the response to STV might vary. Some cultivars or wild relatives could mount a stronger response or have different hormone balances. Exploring STV infection in other tomato varieties or species could be informative. Given the seed transmissibility, one could envision breeders intentionally maintaining STV in lines if it indeed boosts yield without negative effects.\u003c/p\u003e \u003cp\u003eThis reanalysis of the RNA-Seq dataset GSE137303 reveals that Southern tomato virus infection causes no major transcriptional reprogramming in tomato leaves, with the possible exception of dampening the expression of a few ethylene-related genes. Unlike many pathogenic viruses, STV does not trigger broad defense responses or stress pathways in the host, consistent with its asymptomatic nature. The slight downregulation of ethylene biosynthesis and signaling components in STV-infected plants provides a plausible explanation for the previously observed growth and yield benefits. STV appears to be a near-neutral or mildly mutualistic inhabitant of the tomato plant, exerting minimal burden on host cellular processes. This work demonstrates the value of updated genomic resources and methods in reassessing host-virus interactions and emphasizes that biological replicates and proper statistical controls are crucial for distinguishing true biological effects from noise. For future investigations, examining STV\u0026rsquo;s impact under various conditions is recommend to determine if its mutualistic effects are context-dependent. It would also be worthwhile to directly measure hormone levels and fitness outcomes in presence/absence of STV to solidify causal links.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e The authors declare that no funds, grants, or other support were received during the preparation of this manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests:\u003c/strong\u003e The authors have no relevant financial or non-financial interests to disclose.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contribution:\u003c/strong\u003e All authors contributed to the study conception and design. Material preparation, data collection and analysis were performed by Gholamhossein Badeli and Kami Kaboosi. The first draft of the manuscript was written by Gholamhossein Badeli, Kami Kaboosi, and Saeed Nasrollanejad, and all authors commented on previous versions of the manuscript. The manuscript is suppervised by Saeed Nasrollanejad. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability:\u0026nbsp;\u003c/strong\u003eThe datasets analysed during the current study are available in the NCBI\u0026rsquo;s Gene Expression Omnibus (GEO; accession number: GSE137303), BioProject, and Sequence Read Archives (SRA; accession number: PRJNA565137) repositories.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval:\u003c/strong\u003e This is an bioinformatics and data analysis study, and no ethical approval is required.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to participate:\u003c/strong\u003e Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to publish:\u003c/strong\u003e Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClinical trial:\u003c/strong\u003e Not applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eTakahashi H, Fukuhara T, Kitazawa H, Kormelink R (2019) Virus Latency and the Impact on Plants. Front Microbiol 10:2764. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/FMICB.2019.02764\u003c/span\u003e\u003cspan address=\"10.3389/FMICB.2019.02764\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarper A, Vijayakumar V, Ouwehand AC et al (2021) Viral Infections, the Microbiome, and Probiotics. Front Cell Infect Microbiol 10:596166. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/FCIMB.2020.596166/XML\u003c/span\u003e\u003cspan address=\"10.3389/FCIMB.2020.596166/XML\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEscalante C, Sanz-Saez A, Jacobson A et al (2024) Plant virus transmission during seed development and implications to plant defense system. Front Plant Sci 15:1385456. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/FPLS.2024.1385456\u003c/span\u003e\u003cspan address=\"10.3389/FPLS.2024.1385456\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSabanadzovic S, Valverde RA, Brown JK et al (2009) Southern tomato virus: The link between the families Totiviridae and Partitiviridae. Virus Res 140:130\u0026ndash;137. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/J.VIRUSRES.2008.11.018\u003c/span\u003e\u003cspan address=\"10.1016/J.VIRUSRES.2008.11.018\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGonz\u0026aacute;lez LE, Peir\u0026oacute; R, Rubio L, Galipienso L (2021) Persistent Southern Tomato Virus (STV) Interacts with Cucumber Mosaic and/or Pepino Mosaic Virus in Mixed- Infections Modifying Plant Symptoms, Viral Titer and Small RNA Accumulation. Microorganisms 9:689. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/MICROORGANISMS9040689\u003c/span\u003e\u003cspan address=\"10.3390/MICROORGANISMS9040689\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFukuhara T, Tabara M, Koiwa H, Takahashi H (2020) Effect of asymptomatic infection with southern tomato virus on tomato plants. Arch Virol 165:11\u0026ndash;20. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/S00705-019-04436-1\u003c/span\u003e\u003cspan address=\"10.1007/S00705-019-04436-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu M, Pirrello J, Chervin C et al (2015) Ethylene Control of Fruit Ripening: Revisiting the Complex Network of Transcriptional Regulation. Plant Physiol 169:2380. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1104/PP.15.01361\u003c/span\u003e\u003cspan address=\"10.1104/PP.15.01361\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu C, Liu X, Shen G et al (2023) Time-series transcriptome provides insights into the gene regulation network involved in the icariin-flavonoid metabolism during the leaf development of Epimedium pubescens. Front Plant Sci 14:1183481. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/FPLS.2023.1183481/XML\u003c/span\u003e\u003cspan address=\"10.3389/FPLS.2023.1183481/XML\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLove MI, Huber W, Anders S (2014) Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biology 2014 15:12 15:550-. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/S13059-014-0550-8\u003c/span\u003e\u003cspan address=\"10.1186/S13059-014-0550-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiao Y, Smyth GK, Shi W (2019) The R package Rsubread is easier, faster, cheaper and better for alignment and quantification of RNA sequencing reads. Nucleic Acids Res 47:e47\u0026ndash;e47. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/NAR/GKZ114\u003c/span\u003e\u003cspan address=\"10.1093/NAR/GKZ114\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiao Y, Smyth GK, Shi W (2014) FeatureCounts: An efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30:923\u0026ndash;930. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/BIOINFORMATICS/BTT656\u003c/span\u003e\u003cspan address=\"10.1093/BIOINFORMATICS/BTT656\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBenjamini Y, Hochberg Y (1995) Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. J R Stat Soc Ser B Stat Methodol 57:289\u0026ndash;300. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/J.2517-6161.1995.TB02031.X\u003c/span\u003e\u003cspan address=\"10.1111/J.2517-6161.1995.TB02031.X\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eresearch GOC-N acids 2004 undefined The Gene Ontology (GO) database and informatics resource. academic.oup.com\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu G, Wang LG, Han Y, He QY (2012) clusterProfiler: an R Package for Comparing Biological Themes Among Gene Clusters. OMICS 16:284. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1089/OMI.2011.0118\u003c/span\u003e\u003cspan address=\"10.1089/OMI.2011.0118\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDurinck S, Moreau Y, Kasprzyk A et al (2005) BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis. Bioinformatics 21:3439\u0026ndash;3440. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/BIOINFORMATICS/BTI525\u003c/span\u003e\u003cspan address=\"10.1093/BIOINFORMATICS/BTI525\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDurinck S, Spellman PT, Birney E, Huber W (2009) Mapping identifiers for the integration of genomic datasets with the R/Bioconductor package biomaRt. Nature Protocols 2009 4:8 4:1184\u0026ndash;1191. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/nprot.2009.97\u003c/span\u003e\u003cspan address=\"10.1038/nprot.2009.97\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmedley D, Haider S, Ballester B et al (2009) BioMart \u0026ndash; biological queries made easy. BMC Genomics 2009 10:1 10:22-. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/1471-2164-10-22\u003c/span\u003e\u003cspan address=\"10.1186/1471-2164-10-22\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRosado D, Ackermann A, Spassibojko O et al (2021) WRKY transcription factors and ethylene signaling modify root growth during the shade-avoidance response. Plant Physiol 188:1294. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/PLPHYS/KIAB493\u003c/span\u003e\u003cspan address=\"10.1093/PLPHYS/KIAB493\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhankhum S, Sela N, Osorno JM, Valverde RA (2016) RNAseq analysis of endornavirus-infected vs. endornavirus-free common bean (Phaseolus Vulgaris) cultivar black turtle soup. Front Microbiol 7:1905. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/FMICB.2016.01905/FULL\u003c/span\u003e\u003cspan address=\"10.3389/FMICB.2016.01905/FULL\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Southern tomato virus, Solanum lycopersicum, RNA sequencing, differential gene expression, persistent plant virus","lastPublishedDoi":"10.21203/rs.3.rs-8823240/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8823240/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eA previous study reported that Southern tomato virus (STV)-infected tomato plants were taller, had higher fruit yields. Here, it was aimed to reanalyze the same dataset with updated genomic resources and a bioinformatics pipeline to validate and expand upon those findings.The raw single-end RNA-Seq reads were remapped to the latest tomato reference genome (SL3.0) using Rsubread for alignment and featureCounts for gene quantification. Differential expression was assessed with DESeq2, applying rigorous filtering (FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05, |log₂ fold-change| \u0026ge; 1). The results revealed no genes meeting the criteria for differential expression (FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05, fold-change\u0026thinsp;\u0026ge;\u0026thinsp;2). In contrast to the original report, and the host transcriptome remains largely unchanged by STV infection. The few transcripts with the largest expression differences included those in the ethylene biosynthesis and signaling pathway, which were slightly lower in STV-infected plants, mirroring the trends reported previously. Our reanalysis using a RNA-Seq workflow finds that STV infection causes at changes in tomato leaf gene expression, notably slight downregulation of ethylene-related genes, but no broad transcriptomic reprogramming. These results temper the earlier claims of significant transcriptomic effects of STV, instead supporting the notion that STV exerts a minimal and possibly cultivar-specific influence on the host. This study underlines the importance of updated genomic references and statistical approaches. Future studies with larger sample sizes or different developmental stages will be valuable.\u003c/p\u003e","manuscriptTitle":"Reanalysis of the Tomato Transcriptome Revealed Host Responses to Southern Tomato Virus Infection","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-17 10:28:23","doi":"10.21203/rs.3.rs-8823240/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d23a00e8-8232-4025-a526-dde7d60a8054","owner":[],"postedDate":"February 17th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-03-21T06:55:01+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-17 10:28:23","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8823240","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8823240","identity":"rs-8823240","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.