Machine Learning Enables Viral Genome-Agnostic Classification of RNA Virus Infections from Host Transcriptomes

preprint OA: closed
📄 Open PDF Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-03 · read from full text

This preprint evaluates whether viral species can be classified from host transcriptome data alone, without using viral genome sequence information, by testing two host-derived feature sets—differentially expressed genes and alignment-free nucleotide k-mer spectra—on publicly available RNA-seq from Huh7 and Calu-3 cells infected with multiple negative-sense RNA viruses. Using hierarchical clustering and Random Forest models across timepoints (12–24 hours post infection), the authors report that Random Forests could accurately distinguish infections by different viruses, with Influenza A showing the strongest, most distinct host signatures while Ebola and Lassa virus responses were subtler yet still classifiable, and they used label permutation/read shuffling as non-randomness controls. They describe that the differential expression signals are driven mainly by pathways related to translation, immune response, and chromatin remodeling, but the analysis is limited to specific cell lines and in vitro timepoints, which the study design explicitly constrains. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

ABSTRACT Targeted PCR diagnosis of RNA viruses is sequence dependent, meaning that the accuracy of the assay depends on the identity of the viral sequence. However, sequence-targeted assays can miss novel or divergent viruses. We test whether the host transcriptome alone can classify RNA virus infections without using viral sequences. Using publicly available data on Huh7 and Calu-3 cells experimentally infected with diverse negative-sense RNA, we evaluate two host-derived feature sets (differential expression dataset of identified genes and alignment-free nucleotide k-mer spectra) and compare hierarchical clustering with Random Forests. Across datasets and timepoints (12 to 24 hpi), Random forests accurately distinguished cells infected with different viruses. Controls with label permutation and read shuffling across experimental conditions established non-random performance. Influenza A infections exhibited the strongest, most distinct signatures, whereas Ebola and Lassa virus responses were subtler yet still classifiable. These results show that host-only transcriptomics encodes virus-specific, complex signals that machine learning can exploit, enabling genome-agnostic classification of infection statuses. This approach could aid early outbreak triage when primer based detection methods fail or viral genomes are unknown and complements sequence-based discovery.
Full text 40,765 characters · extracted from oa-pdf · 7 sections · click to expand

Abstract

Targeted PCR diagnosis of RNA viruses is sequence dependent, meaning that the accuracy of the assay depends on the identity of the viral sequence. However, sequence-targeted assays can miss novel or divergent viruses. We test whether the host transcriptome alone can classify RNA virus infections without using viral sequences. Using publicly available data on Huh7 and Calu-3 cells experimentally infected with diverse negative-sense RNA, we evaluate two host-derived feature sets (differential expression dataset of identified genes and alignment-free nucleotide k-mer spectra) and compare hierarchical clustering with Random Forests. Across datasets and timepoints (12 to 24 hpi), Random forests accurately distinguished cells infected with different viruses. Controls with label permutation and read shuffling across experimental conditions established non-random performance. Influenza A infections exhibited the strongest, most distinct signatures, whereas Ebola and Lassa virus responses were subtler yet still classifiable. These results show that host-only transcriptomics encodes virus-specific, complex signals that machine learning can exploit, enabling genome-agnostic classification of infection statuses. This approach could aid early outbreak triage when primer based detection methods fail or viral genomes are unknown and complements sequence-based discovery.

Introduction

(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint Zoonotic viruses—those that spill over from animal reservoirs into human populations—pose a growing threat to global health. Anthropogenic global change and biodiversity loss are accelerating the frequency of spillover events by altering ecosystems, disrupting host-pathogen dynamics, and increasing contact between humans and wildlife1,2. While only a subset of animal viruses known to be are capable of infecting humans, those that do are disproportionately likely to cause severe disease and, in some cases, global pandemics3–7. Despite this risk, the vast majority of viruses circulating in wildlife remain undescribed8–10. Most have never been isolated, sequenced, or even detected9–11. This knowledge gap severely limits our ability to assess which viruses might pose a threat to public health. In addition, once in humans, many viruses may cause subclinical or misdiagnosed diseases (such as undiagnosed febrile diseases that could resemble relatively common diseases like malaria or influenza), further complicating efforts to identify them before they spread12,13. (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint Targeted PCR-based assays face similar challenges. These methods depend on primers designed to match conserved regions of known viral genomes. If a virus’s genome diverges too much from the expected sequence, primer binding may be inefficient or fail entirely, resulting in false negatives14–17. Even assays designed with degenerate primers—such as pan-family PCRs—can fail to detect divergent strains. For example, a second genotype of Hendra virus (HeV-g2) circulated undetected in a surveilled population of bats due to mismatches with primers designed for the original Hendra virus genome18. Similarly, a recombinant coronavirus (CCoV-HuPn- 2018) that caused febrile illness in children in Malaysia and Haiti escaped detection by coronavirus-specific RT-PCR19,20. This is particularly concerning because Henipaviruses and Coronaviruses are both considered WHO priority pathogens that are flagged for their potential to cause pandemics with high fatality rates, and in these two (and potentially many more) cases they were missed21. (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint Modern virus discovery relies heavily on sequence-based methods (such as genome amplification by targeted PCR for detection and sequencing), particularly high-throughput sequencing and metagenomics22,23. These approaches enable the recovery of viral genome fragments directly from clinical or environmental samples without prior knowledge of the virus’s identity24,25. However, sequence-based discovery typically requires some level of similarity between the unknown virus and existing sequences in reference databases26–28. Highly divergent viruses may go undetected if they lack conserved genes or recognizable sequence motifs 29,30. This is a particularly relevant problem for RNA viruses, which tend to have high mutation rates, genetic variability, and extreme expression differences between genes making it difficult to pick a good PCR target across families—and even within genera31. For example, transcription in Ebola virus (EBOV) follows a 3’ – 5’ gradient with the RNA-dependent RNA polymerase being the least expressed gene32. As such, there is an urgent need for virus detection strategies that do not depend on prior sequence information. One alternative is to detect infection by measuring changes in the host cell, rather than detecting the virus directly. Viral infection can induce characteristic transcriptional responses in host cells32,33, which may serve as indirect markers of infection. Host-based diagnostic methods are already used clinically to distinguish between viral and bacterial infections by analyzing expression patterns of 29 genes that are up- or downregulated during infection 34,35. Some studies have used short k-mers from viral genomes to train machine learning models that can classify viral infections from total RNA-seq data, without relying on alignment to reference genomes to determine the presence or absence of one type of viral infection or to discriminate between bacterial or viral infection36–39. (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint However, few studies have systematically tested whether host transcriptomic responses alone— without any viral sequence—can be used to classify infections across different viral families. It remains unclear whether the host response provides enough resolution to infer the identity or family of the infecting virus, especially across diverse cell types and timepoints. In this study, we test the hypothesis that viral species can be inferred from host transcriptomic responses alone. We evaluate two feature types (short nucleotide k-mer spectra from total RNA and differential gene expression) and assess their utility for classifying viral infections using machine learning and hierarchical clustering. We focus on negative-sense RNA viruses infecting mammalian cell lines, using publicly available transcriptomics datasets40,41. Our panel includes experimental infections of Huh7 cells (a human hepatocyte immortalized cell line) and Calu-3 cells (a human lung epithelial immortalized cell line) with viruses including Influenza H1N1, H5N1, H7N7, H7N9, EBOV , Marburg virus (MARV), Respiratory Syncytial Virus (RSV), Nipah virus (NiV), Lassa virus (LASV), Rift Valley Fever virus (RVFV), and Sandfly Fever Sardinia Virus (SFSV). Across both cell types, we show that different viruses elicit distinct, classifiable host responses—supporting the possibility of sequence-independent viral classification using host transcriptional data alone.

Results

Gene expression patterns elicited by infection vary in magnitude and differentiability To test the hypothesis that each viral infection results in a differentiable response compared to other viruses, we performed hierarchical clustering of differential gene expression profiles from experimentally infected Huh7 cells at 12 and 24 hours post infection. In short, hierarchical clustering is a method for grouping similar data points into clusters by progressively merging the (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint closest pairs to form a tree-like structure called a dendrogram. We used the Euclidean distance, which is a measure of the straight-line distance between two points in space, calculated as the square root of the sum of the squared differences between their corresponding coordinates42. The magnitude and specificity of the host transcriptional response varied across viral infections, resulting in distinct patterns of gene expression. Hierarchical clustering using the top 100 (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint differentially expressed genes effectively separated infected from uninfected samples (Figure 1). While some viral infections elicited sharply divergent transcriptomic profiles, others were more subtle and clustered closer to uninfected controls. Notably, infections with Influenza A viruses (H1N1 and H5N1) induced the most distinct expression patterns, clustering farthest from the uninfected samples and from other viruses. Several other viruses—including NiV , RSV , SFSV , RVFV , and MARV—also produced expression profiles distinguishable from uninfected controls, as their transcriptomes were placed outside of the subtree containing most mock-infected samples. In contrast, infections with EBOV and LASV produced host responses that were less distinguishable, with infected samples Figure 1: Hierarchical clustering of gene expression profiles from Huh7 cells 12 hours post infection (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint clustering within or adjacent to the subtree of uninfected controls (Figure 1). The genes driving these distinctions were primarily involved in translation, immune response, and chromatin remodeling, followed by genes associated transport, cell signaling, and metabolism (Figure 2). Figure 2: Z scores of top 100 differentially expressed genes in Huh7 cells infected with various negative-sense RNA viruses after 12 hours post infection and mock (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint A consistent feature across most infections was the downregulation of genes involved in translation, a known host antiviral response43–45; however, this trend was reversed in H5N1 and H1N1, and to a lesser extent RVFV , where translation-associated genes were upregulated relative to uninfected samples. Furthermore, only H5N1, H1N1, and MARV exhibited clustering patterns so distinct that their replicates grouped into clearly defined, virus-specific subtrees. The patterns of gene expression elicited by closely related Influenza A viruses are specific and differentiable by hierarchical clustering and random forest classification To test whether Influenza A viruses consistently induced differentiable responses in a different cell type and across viral genotypes, we used experimental infection micro-array gene expression data from Calu-3 cells experimentally infected with H1N1, H5N1, H7N7, and H7N9 viruses 41. We found that all of the four viruses were able to be classified by both Hierarchical clustering and random forest classification at 12 and 24 hours post infection (Figure 3). Figure 3: Hierarchical clustering Tree Purity and Random forest Classification F1 score of Differential Gene Expression Profiles originating from Calu-3 cells infected with various Influenza A viruses or mock at 3, 7, 12, and 24 hours post infection (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint The distribution and composition of nucleotides in an infected cell varies in a virus-specific manner To test whether the nucleotide composition of an infected cell can be used to predict which virus was infecting the cell, we compared experimental data to simulated parallel datasets where the nucleotides were randomly shuffled within each sequencing read or a dataset where the infection status label was randomly assigned to each kmer abundance profile. Infection status was differentiable for kmer lengths 8 to 12 above random chance across all of the viruses tested via hierarchical clustering using Euclidean distance. Virally infected transcriptomes were differentiable to random forest classification above randomly assigned data for the entire k-mer range, and was differentiable to the model in contrast with nucleotide- shuffled data when the k-mer length was greater than 4 quantified by purity and F1 score which measure the accuracy of hierarchical clustering or random forests to group samples of the same category together, respectively (Figure 4). (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint The contribution of nucleotides from negative-sense RNA viruses drives the differentiability of k-mer abundance profiles in hierarchical clustering but not random forest classification To test if the viral RNA itself drives the differentiability of the transcriptomes in Random forest classification and Hierarchical Cluster, we created a synthetic dataset where the viral reads were depleted. In transcriptomes where reads originating from each virus were depleted, at no point in the kmer range were the transcriptomes differentiable above the shuffled or permuted data. In Random forest classification, kmers above length 2 were able to inform the model above random chance, but not to consistently high accuracy (Figure 5). Figure 4: Performance of Hierarchical clustering measured by Tree Purity (a), and Random forest F1 Score (b), Recall (c), and Precision (d) on infected and uninfected Huh7 transcriptomes after 12 hours post infection and synthetic datasets where all of the reads in the original transcriptomes were randomly shuffled, or where each k-mer abundance profile was randomly assigned a virus (ie. permuted) (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint Enrichment of viral RNA in transcriptomic data enhances the differentiability to random forest classification and hierarchical clustering When viral reads were doubled in the dataset, the resulting differentiability by hierarchical clustering and random forest classification increased. For hierarchical clustering, the data was differentiable above random chance over the k-mer range 4 – 12 and was differentiable by random forest classification for all k-mer lengths tested (Figure 6). Figure 5: Performance of hierarchical clustering measured by tree purity (a) and random forest F1 score (b) on infected and uninfected Huh7 transcriptomes at 12 and 24 hours post infection that were computationally depleted of viral reads (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint

Discussion

To replicate and spread, all viruses must manipulate host cell biology46. At the cellular level, this often involves subverting host gene expression and regulatory networks to promote viral replication, suppress immune responses, and evade detection46. Viruses systemically manipulate the host’s immune system, inhibit antiviral signaling pathways, rewire host metabolism, or disrupt intercellular communication to maximize their replication fitness47–49. Despite their compact genomes, RNA viruses encode multifunctional proteins that carry out both essential replication tasks and host antagonism46. These dual roles are a product of strong selective pressures: with limited coding space, many RNA viruses have evolved proteins and genome structures that both facilitate replication and interfere with host defenses. Viral replication itself often hijacks or suppresses core host processes at every stage—transcription, translation, signaling, and transport47,50–52. Figure 6: Performance of Hierarchical clustering measured by Tree Purity (a) and Random forest F1 Score (b) on infected and uninfected Huh7 transcriptomes at 12 and 24 hours post infection that were computationally enriched for viral reads (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint In turn, host cells have evolved sophisticated and dynamic genetic mechanisms to detect and restrict viral infection. Mammalian cells encode hundreds of innate immune genes that are activated in response to viral invasion, though the specific pathways engaged can vary depending on the virus and host cell type33,50–53. Effective immune responses must be both rapid and specific: activating the appropriate antiviral programs without triggering unnecessary or energetically costly responses is critical to controlling infection without collateral damage. These diverse, dynamic responses are reflected in the transcriptome, making them an attractive substrate for computational classification. A clear example of this functional specialization is seen in the differential response to DNA versus RNA viruses. For instance, infection with a DNA virus such as mpox (monkeypox) activates cytosolic DNA sensors like cGAS, DNA-PK, and IFI16, which initiate downstream signaling to induce interferon-β and other antiviral genes13,14. In contrast, RNA viruses such as SARS-CoV-2 or Influenza A are recognized primarily through double-stranded RNA intermediates by RIG-I and MDA552. These sensors engage the mitochondrial adapter MA VS to trigger antiviral transcriptional programs52. Transcriptomic studies have shown that even among related viral strains, host responses can differ markedly. For instance, in the 2022 outbreak of mpox (Clade IIb), several genes were significantly upregulated (e.g., SLC2A3, ATP2B1, VEGFA) or downregulated (MAP3K8, IL1A, SGK1), reflecting a distinct host signature53. Similarly, SARS-CoV-2 variants induce variant- specific responses: the Delta variant was associated with upregulation of EGR1 and IFIT, while Alpha triggered genes like SYVN1, CH25H, VIPR1, and others54. In this study, we show that host transcriptomic data—whether represented as normalized gene expression values or alignment-free k-mer frequency spectra—contain sufficient information to (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint classify the infecting virus using machine learning. For viruses like Influenza A, which contribute substantial amounts of viral RNA to the transcriptome and elicit strong transcriptional changes, both hierarchical clustering and Random forest classification achieved high performance. In contrast, even for viruses with more subtle effects on host gene expression and lower transcriptomic viral burden, Random forest classifiers were still able to differentiate viral genotypes with high accuracy. This likely reflects the ability of Random forests to capture non- linear patterns in complex, high-dimensional data55. Where hierarchical clustering failed to separate infected from uninfected samples—especially in cases where the host response was mild or variable—Random forest models consistently identified the correct viral infection. This suggests that while traditional clustering methods detect global patterns of perturbation, machine learning models can exploit finer-grained, distributed signals that are not apparent through unsupervised methods. This study has several important limitations. Our analyses are based on infections in two immortalized human cell lines (Huh7 and Calu-3) under controlled laboratory conditions, which may not capture the complexity of in vivo or clinical infections. Host responses in primary cells, across tissues, or in whole organisms can differ in both timing and magnitude, potentially affecting classification performance. The datasets are further constrained to mid-stage infection (12–24 hpi); earlier or later responses may be weaker, noisier, or biologically distinct. In addition, our work is limited to negative-sense RNA viruses, leaving open how well this approach generalizes to positive-sense RNA or DNA viruses with different replication dynamics. Finally, clinical samples often present additional challenges—mixed infections, inter-individual variability, and technical artifacts such as batch effects or sequencing depth—that may reduce robustness. Although our depletion and enrichment experiments begin to address the role of viral (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint read abundance, the minimal data requirements and cross-platform reproducibility remain unresolved. Several directions could strengthen the translational relevance of this framework. Extending benchmarking to additional cell types, primary cultures, and eventually clinical samples will be essential for testing whether host-only signatures generalize beyond controlled in vitro settings. Longitudinal sampling across broader post-infection windows could help determine the earliest time points at which responses become diagnostic and the duration for which they remain informative.. Despite these caveats, our findings provide a clear proof-of-concept. We show that host transcriptomes alone are sufficient to distinguish between infections with diverse RNA viruses, even when viral reads are scarce and when unsupervised clustering fails. This underscores that host responses are not merely generic markers of infection but encode virus-specific, nonlinear signatures that can be captured by machine learning. In practice, such genome-agnostic signals could support early diagnostic triage when primers are unreliable or viral genomes are unknown, complementing sequence-based approaches. These results align with a growing body of work on transcriptomic diagnostics in bacterial sepsis and respiratory infections, suggesting a broader paradigm in which host responses serve as sensitive biosensors of pathogen class. By demonstrating consistent classification across multiple viral families, we add to evidence that systems-level host responses can bridge key gaps between pathogen discovery, outbreak surveillance, and clinical diagnostics. (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint DATA A V AILABILITY The data used in this study were obtained from the Sequencing Read Archive (PRJNA1074963) and Gene Expression Omnibus (GSE49840). From PRJNA1074963, we created synthetic permuted, shuffled, enriched, and depleted derivative datasets.

Materials and methods

Data sources. We analyzed publicly available RNA-seq datasets from human hepatocellular carcinoma (Huh7) and lung adenocarcinoma (Calu-3) cell lines infected with a panel of negative-sense RNA viruses, including influenza A subtypes (H1N1, H5N1, H7N7, H7N9), Ebola virus (EBOV), Marburg virus (MARV), Nipah virus (NiV), Rift Valley fever virus (RVFV), respiratory syncytial virus (RSV), Sandfly fever Sicilian virus (SFSV), and Lassa virus (LASV). Metadata on infection time point (3, 6, 7, 12 or 24 hpi), replicate, and sequencing depth were extracted from SRA and GEO project accessions (PRJNA1074963, GSE49840). We used the first 5,000 reads from each .fastq file, because we saw that the nucleotide k-mer spectra converged around 1000- 5000 reads (Suppl. Fig. S1). Feature generation. We derived two complementary feature sets from host transcriptomes: (i) normalized differential gene expression (TPM) values and (ii) alignment-free k-mer frequency spectra (k = 1–15). Viral reads were excluded or enriched as controls by re-mapping and sub-sampling. Classification and clustering. (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint To evaluate separability of host responses, we applied hierarchical clustering (Euclidean distance, Ward linkage) and Random Forest classifiers. For Random Forests, we used 200 trees, class-balanced stratified 4-fold cross-validation (smallest class n = 4), and default scikit-learn parameters unless otherwise specified. Controls. We implemented strict controls, including label permutation, read shuffling, viral read depletion, and viral read enrichment, to confirm that observed signal reflected biologically meaningful host responses. Reproducibility. All analyses were performed in Python (v3.11) using pandas, scikit-learn, Biopython, and SciPy. Data accessions are provided above; analysis code will be publicly archived on Zenodo before publication. All of the code is hosted on the following GitHub repository: https://github.com/blatuscaspot/multivirus-classification Supplemental Material

Acknowledgements

We acknowledge the input of Dr. John Parker at Cornell University in suggesting to use shuffled nucleotides. This research was enabled in part by support provided by Calcul Québec (calculquebec.ca) and the Digital Research Alliance of Canada (alliancecan.ca). (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint Funding BK is supported by a Canadian Institutes for Health Research Doctoral Foreign Study Award: Machine Learning to Identify Viral Pathogens through Host Cell Gene Expression Patterns, and was supported by a gift from the Griffin Foundation, as well as a Fellows-In-Residence Award from Verena (an NSF Biology Integration Institute). TP was funded through award 223764/Z/21/Z from the Wellcome Trust. TP, BK, and SNS are supported by the U.S. National Science Foundation (NSF DBI 2213854).

References

1. Carlson, C. J. et al. Pathogens and planetary change. Nat. Rev. Biodivers. 1, 32–49 (2025). 2. Carlson, C. J. et al. Climate change increases cross-species viral transmission risk. Nature 607, 555–562 (2022). 3. Plowright, R. K. et al. Ecological countermeasures to prevent pathogen spillover and subsequent pandemics. Nat. Commun. 15, 2577 (2024). 4. Gurley, E. S. & Plowright, R. K. A Roadmap of Primary Pandemic Prevention Through Spillover Investigation - V olume 31, Number 8—August 2025 - Emerging Infectious Diseases journal - CDC. doi:10.3201/eid3108.250442. 5. Ruiz-Aravena, M. et al. Ecology, evolution and spillover of coronaviruses from bats. Nat. Rev. Microbiol. 20, 299–314 (2022). 6. Guth, S., Visher, E., Boots, M. & Brook, C. E. Host phylogenetic distance drives trends in virus virulence and transmissibility across the animal-human interface. Philos. Trans. R. Soc. Lond. B. Biol. Sci. 374, 20190296 (2019). 7. Guth, S. et al. Bats host the most virulent-but not the most dangerous-zoonotic viruses. Proc. Natl. Acad. Sci. U. S. A. 119, e2113628119 (2022). (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint 8. Kawasaki, J., Kojima, S., Tomonaga, K. & Horie, M. Hidden Viral Sequences in Public Sequencing Data and Warning for Future Emerging Diseases. mBio 12, e01638 (2021). 9. Anthony, S. J. et al. A Strategy To Estimate Unknown Viral Diversity in Mammals. mBio 4, e00598-13 (2013). 10. Kiselev, D. et al. Current Trends in Diagnostics of Viral Infections of Unknown Etiology. Viruses 12, 211 (2020). 11. Lipkin, W. I. The changing face of pathogen discovery and surveillance. Nat. Rev. Microbiol. 11, 133–141 (2013). 12. A Zoonotic Henipavirus in Febrile Patients in China | New England Journal of Medicine. https://www.nejm.org/doi/full/10.1056/NEJMc2202705. 13. Bartekwa, J. W., Abene, E. E., Luka, P. D., Yilgwan, C. S. & Shehu, N. Y . COVID-19 Subclinical Infection and Immunity: A Review. Niger. J. Med. J. Natl. Assoc. Resid. Dr. Niger. 30, 631–636 (2021). 14. Chow, C.-K., Qin, K., Lau, L.-T. & Cheung-Hoi Yu, A. Significance of a Single- Nucleotide Primer Mismatch in Hepatitis B Virus Real-Time PCR Diagnostic Assays. J. Clin. Microbiol. 49, 4418–4419 (2011). 15. Holbrook, M. G. et al. Updated and Validated Pan-Coronavirus PCR Assay to Detect All Coronavirus Genera. Viruses 13, 599 (2021). 16. Whiley, D. M. & Sloots, T. P. Sequence variation in primer targets affects the accuracy of viral quantitative PCR. J. Clin. Virol. Off. Publ. Pan Am. Soc. Clin. Virol. 34, 104–107 (2005). 17. Khan, K. A. & Cheung, P. Evaluation of the Sequence Variability within the PCR Primer/Probe Target Regions of the SARS-CoV-2 Genome. Bio-Protoc. 10, e3871 (2020). (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint 18. Taylor, J. et al. Novel variant Hendra virus genotype 2 infection in a horse in the greater Newcastle region, New South Wales, Australia. One Health 15, 100423 (2022). 19. Lednicky, J. A. et al. Isolation of a Novel Recombinant Canine Coronavirus From a Visitor to Haiti: Further Evidence of Transmission of Coronaviruses of Zoonotic Origin to Humans. Clin. Infect. Dis. 75, e1184–e1187 (2022). 20. Zehr, J. D. et al. Recent Zoonotic Spillover and Tropism Shift of a Canine Coronavirus Is Associated with Relaxed Selection and Putative Loss of Function in NTD Subdomain of Spike Protein. Viruses 14, 853 (2022). 21. Pathogens prioritization: a scientific framework for epidemic and pandemic research preparedness. https://www.who.int/publications/m/item/pathogens-prioritization-a-scientific- framework-for-epidemic-and-pandemic-research-preparedness. 22. Ho, T. & Tzanetakis, I. E. Development of a virus detection and discovery pipeline using next generation sequencing. Virology 471–473, 54–60 (2014). 23. Woolhouse, M., Scott, F., Hudson, Z., Howey, R. & Chase-Topping, M. Human viruses: discovery and emergence. Philos. Trans. R. Soc. B Biol. Sci. 367, 2864–2871 (2012). 24. Chin, P.-J. et al. Virus detection by short read high throughput sequencing in a high virus low cellular background. Npj Vaccines 10, 61 (2025). 25. Minicka, J., Zarzyńska-Nowak, A., Budzyńska, D., Borodynko-Filas, N. & Hasiów- Jaroszewska, B. High-Throughput Sequencing Facilitates Discovery of New Plant Viruses in Poland. Plants 9, 820 (2020). 26. Edgar, R. C. et al. Petabase-scale sequence alignment catalyses viral discovery. Nature 602, 142–147 (2022). (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint 27. Ulhuq, F. R. et al. Analysis of the ARTIC V4 and V4.1 SARS-CoV-2 primers and their impact on the detection of Omicron BA.1 and BA.2 lineage-defining mutations. Microb. Genomics 9, mgen000991 (2023). 28. Nesti, D. R. et al. Development of a semicomprehensive detection method for paramyxoviruses and its validation using Indonesian bats. Sci. Rep. 15, 19154 (2025). 29. Phan, M. V . T. et al. Identification of missed viruses by metagenomic sequencing of clinical respiratory samples from Kenya. Sci. Rep. 12, 202 (2022). 30. Kamau, E. et al. Recent sequence variation in probe binding site affected detection of respiratory syncytial virus group B by real-time RT-PCR. J. Clin. Virol. 88, 21–25 (2017). 31. Duffy, S. Why are RNA virus mutation rates so damn high? PLoS Biol. 16, e3000003 (2018). 32. Brauburger, K., Boehmann, Y ., Krähling, V . & Mühlberger, E. Transcriptional Regulation in Ebola Virus: Effects of Gene Border Structure and Regulatory Elements on Gene Expression and Polymerase Scanning Behavior. J. Virol. 90, 1898–1909 (2016). 33. Merchant, M. et al. SARS-CoV-2 variants induce increased inflammatory gene expression but reduced interferon responses and heme synthesis as compared with wild type strains. Sci. Rep. 14, 25734 (2024). 34. Tong-Minh, K. et al. A 29-mRNA host response test to identify bacterial and viral infections and to predict 30-day mortality in emergency department patients with suspected infections: A prospective observational cohort study. Diagn. Microbiol. Infect. Dis. 111, 116599 (2025). (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint 35. Sampson, D. et al. Blood transcriptomic discrimination of bacterial and viral infections in the emergency department: a multi-cohort observational validation study. BMC Med. 18, 185 (2020). 36. Riquier, S. et al. Kmerator Suite: design of specific k-mer signatures and automatic metadata discovery in large RNA-seq datasets. NAR Genomics Bioinforma. 3, lqab058 (2021). 37. Alam, Md. N. U. & Chowdhury, U. F. Short k-mer abundance profiles yield robust machine learning features and accurate classifiers for RNA viruses. PLoS ONE 15, e0239381 (2020). 38. Allesøe, R. L. et al. Automated download and clean-up of family-specific databases for kmer-based virus identification. Bioinformatics 37, 705–710 (2020). 39. Mouratidis, I. et al. kmerDB: A database encompassing the set of genomic and proteomic sequence information for each species. Comput. Struct. Biotechnol. J. 23, 1919–1928 (2024). 40. Hofmann, N. et al. Distinct negative-sense RNA viruses induce a common set of transcripts encoding proteins forming an extensive network. J. Virol. 98, e0093524 (2024). 41. Josset, L., Zeng, H., Kelly, S. M., Tumpey, T. M. & Katze, M. G. Transcriptomic characterization of the novel avian-origin influenza A (H7N9) virus: specific host response and responses intermediate between avian (H5N1 and H7N7) and human (H3N2) viruses and implications for treatment options. mBio 5, e01102-01113 (2014). 42. Ultsch, A. & Lötsch, J. Euclidean distance-optimized data transformation for cluster analysis in biomedical data (EDOtrans). BMC Bioinformatics 23, 233 (2022). 43. Livingstone, M. et al. Assessment of mTOR-Dependent Translational Regulation of Interferon Stimulated Genes. PLoS ONE 10, e0133482 (2015). (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint 44. Karousis, E. D., Schubert, K. & Ban, N. Coronavirus takeover of host cell translation and intracellular antiviral response: a molecular perspective. EMBO J. 43, 151–167 (2024). 45. Kroczynska, B., Mehrotra, S., Arslan, A. D., Kaur, S. & Platanias, L. C. Regulation of Interferon-Dependent mRNA Translation of Target Genes. J. Interferon Cytokine Res. 34, 289–296 (2014). 46. Kaza, B. & Aguilar, H. C. Pathogenicity and virulence of henipaviruses. Virulence 14, 2273684 (2023). 47. Rosenke, K. et al. UK B.1.1.7 variant exhibits increased respiratory replication and shedding in nonhuman primates. BioRxiv Prepr. Serv. Biol. 2021.06.11.448134 (2021) doi:10.1101/2021.06.11.448134. 48. Hansen, F. et al. SARS-CoV-2 reinfection prevents acute respiratory disease in Syrian hamsters but not replication in the upper respiratory tract. Cell Rep. 38, 110515 (2022). 49. Ithinji, D. G. et al. Multivalent viral particles elicit safe and efficient immunoprotection against Nipah Hendra and Ebola viruses. NPJ Vaccines 7, 166 (2022). 50. Welcome to Interferome. https://interferome.org/interferome/home.jspx. 51. Lu, Y . & Zhang, L. DNA-Sensing Antiviral Innate Immunity in Poxvirus Infection. Front. Immunol. 11, 1637 (2020). 52. Rehwinkel, J. & Gack, M. U. RIG-I-like receptors: their regulation and roles in RNA sensing. Nat. Rev. Immunol. 20, 537–551 (2020). 53. Debnath, J. P. et al. Identification of potential biomarkers for 2022 Mpox virus infection: a transcriptomic network analysis and machine learning approach. Sci. Rep. 15, 2922 (2025). (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint 54. SARS-CoV-2 variants induce increased inflammatory gene expression but reduced interferon responses and heme synthesis as compared with wild type strains | Scientific Reports. https://www.nature.com/articles/s41598-024-76401-1. 55. Ryo, M. & Rillig, M. C. Statistically reinforced machine learning for nonlinear patterns and variable interactions. Ecosphere 8, e01976 (2017). (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. The copyright holder for this preprintthis version posted September 11, 2025. ; https://doi.org/10.1101/2025.09.11.675678doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-08-09T06:42:26.407065+00:00