Hardware-Aware RNA-seq Diagnostics: Plant Virus Detection via Cloud and AI | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Hardware-Aware RNA-seq Diagnostics: Plant Virus Detection via Cloud and AI Elisson Silva, Paolo Margaria, Rosana Blawid, Eder J. Oliveira, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6667739/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: High-throughput sequencing (HTS) has transformed plant pathogen diagnostics by enabling comprehensive analysis of RNA-seq data. However, the size and complexity of HTS datasets often exceed the computational resources available in phytopathology laboratories, particularly in low- and middle-income regions. Efficient processing of HTS data remains a major bottleneck for timely and accessible virus detection. Results: To address these challenges, we introduce here two complementary strategies that mitigate hardware constraints in virus diagnostics, considering as study case the detection and discovery of viruses in cassava, a major food crop in tropical and subtropical regions worldwide. First, we deployed PhytoPipe, a state-of-the-art phytosanitary pipeline, as a cloud-based service on Amazon Web Services (AWS), allowing laboratories to run sophisticated workflows without high-performance computing infrastructure. PhytoPipe integrates unguided virus detection methods—including read classification, assembly-based annotation, and reference-based mapping—and was validated in this work using RNA-seq data from virus-infected cassava samples. Second, we developed an artificial neural network (ANN) classifier based on multi-head attention mechanisms , trained to distinguish viral reads from plant and bacterial sequences in unassembled HTS data (raw reads). By encoding k-mers at the protein level, the model achieves accurate classification while significantly reducing computational overhead. The ANN’s filtering capability improves the efficiency of downstream assembly and annotation tools, enabling virus detection on resource-constrained systems. Conclusion: Our findings demonstrate that cloud-based bioinformatics pipelines and machine learning classifiers can substantially lower the hardware demands of HTS-based pathogen diagnostics. These strategies expand access to advanced viral detection workflows, making them more scalable and applicable in diverse research and diagnostic settings, including laboratories with limited computational infrastructure. Virus detection RNA-seq data Sequencing read classification Cloud service Deep learning Full Text Additional Declarations No competing interests reported. Supplementary Files BMCBioInf2025Supl.pdf Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6667739","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":490446120,"identity":"44906514-1ad4-473a-9627-c27b69b0bad7","order_by":0,"name":"Elisson Silva","email":"","orcid":"","institution":"Universidade Federal de Pernambuco","correspondingAuthor":false,"prefix":"","firstName":"Elisson","middleName":"","lastName":"Silva","suffix":""},{"id":490446121,"identity":"1906c086-ab19-4234-bf7e-a9478358ec9d","order_by":1,"name":"Paolo Margaria","email":"","orcid":"","institution":"Leibniz Institute DSMZ, German Collection of Microorganisms and Cell Cultures GmbH","correspondingAuthor":false,"prefix":"","firstName":"Paolo","middleName":"","lastName":"Margaria","suffix":""},{"id":490446122,"identity":"d0a67a38-393f-4f73-9304-fb9ea2a9b4cc","order_by":2,"name":"Rosana Blawid","email":"","orcid":"","institution":"Universidade Federal Rural de Pernambuco","correspondingAuthor":false,"prefix":"","firstName":"Rosana","middleName":"","lastName":"Blawid","suffix":""},{"id":490446123,"identity":"a3b2a178-fafb-4116-9c7a-efed1dd80c88","order_by":3,"name":"Eder J. Oliveira","email":"","orcid":"","institution":"Embrapa Mandioca e Fruticultura","correspondingAuthor":false,"prefix":"","firstName":"Eder","middleName":"J.","lastName":"Oliveira","suffix":""},{"id":490446124,"identity":"787e533d-6ef9-473f-8ad5-8e84bda70a21","order_by":4,"name":"Stephan Winter","email":"","orcid":"","institution":"Leibniz Institute DSMZ, German Collection of Microorganisms and Cell Cultures GmbH","correspondingAuthor":false,"prefix":"","firstName":"Stephan","middleName":"","lastName":"Winter","suffix":""},{"id":490446125,"identity":"ae05d7e9-ef71-4ed9-bce7-3dfac9bfca1b","order_by":5,"name":"Stefan Blawid","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA2UlEQVRIiWNgGAWjYDACCeYGBgYDZjkGZhDPBiRCUAsjWIsxREsa0VoYmBMbGIjVIj+7sfEzT4F1en877wFmnoQ7cgzSvQ/wajG4c7BZmscgPXfGYb4EoJZnxgwyxw3wa5FIbJDOMTicu4GZx4CZ98fhxAaJNAIOm5HY/BuoJd0ApIUn4XA9QS0MNxLbQLYkwLQkMBDSAvRLm/Ufg3TDGYd5DA7OSXhm2CZzjIDDZjcfvjnjj7U8f/8ZwwdvEu7I80u3EXAYMjgAQmwkaIDpGgWjYBSMglGABgArs0CeB5+k+wAAAABJRU5ErkJggg==","orcid":"","institution":"Universidade Federal de Pernambuco","correspondingAuthor":true,"prefix":"","firstName":"Stefan","middleName":"","lastName":"Blawid","suffix":""}],"badges":[],"createdAt":"2025-05-15 01:08:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6667739/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6667739/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":92179163,"identity":"06bd63d8-6d20-4f5b-aa19-31f04fc04f14","added_by":"auto","created_at":"2025-09-25 13:23:23","extension":"json","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8223,"visible":true,"origin":"","legend":"","description":"","filename":"d40ec53fd3d741e083d618aac7e167d8.json","url":"https://assets-eu.researchsquare.com/files/rs-6667739/v1/7e0503b2e328858795038e3d.json"},{"id":103081487,"identity":"7af5823e-fdc5-43d7-b144-823afd9035a0","added_by":"auto","created_at":"2026-02-20 14:41:34","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":664953,"visible":true,"origin":"","legend":"","description":"","filename":"BMCBioInf2025.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6667739/v1_covered_7abf2025-b39b-4d92-b29d-cadd2cd2b7bc.pdf"},{"id":92179169,"identity":"5d316f4e-1946-4b76-9733-6b0830742dcd","added_by":"auto","created_at":"2025-09-25 13:23:24","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":524455,"visible":true,"origin":"","legend":"","description":"","filename":"BMCBioInf2025Supl.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6667739/v1/728ccfac49879da9abacb726.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Hardware-Aware RNA-seq Diagnostics: Plant Virus Detection via Cloud and AI","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":true,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Virus detection, RNA-seq data, Sequencing read classification, Cloud service, Deep learning","lastPublishedDoi":"10.21203/rs.3.rs-6667739/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6667739/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Background: High-throughput sequencing (HTS) has transformed plant pathogen diagnostics by enabling comprehensive analysis of RNA-seq data. However, the size and complexity of HTS datasets often exceed the computational resources available in phytopathology laboratories, particularly in low- and middle-income regions. Efficient processing of HTS data remains a major bottleneck for timely and accessible virus detection.\n\nResults: To address these challenges, we introduce here two complementary strategies that mitigate hardware constraints in virus diagnostics, considering as study case the detection and discovery of viruses in cassava, a major food crop in tropical and subtropical regions worldwide. First, we deployed PhytoPipe, a state-of-the-art phytosanitary pipeline, as a cloud-based service on Amazon Web Services (AWS), allowing laboratories to run sophisticated workflows without high-performance computing infrastructure. PhytoPipe integrates unguided virus detection methods—including read classification, assembly-based annotation, and reference-based mapping—and was validated in this work using RNA-seq data from virus-infected cassava samples. Second, we developed an artificial neural network (ANN) classifier based on multi-head attention mechanisms , trained to distinguish viral reads from plant and bacterial sequences in unassembled HTS data (raw reads). By encoding k-mers at the protein level, the model achieves accurate classification while significantly reducing computational overhead. The ANN’s filtering capability improves the efficiency of downstream assembly and annotation tools, enabling virus detection on resource-constrained systems.\n\nConclusion: Our findings demonstrate that cloud-based bioinformatics pipelines and machine learning classifiers can substantially lower the hardware demands of HTS-based pathogen diagnostics. These strategies expand access to advanced viral detection workflows, making them more scalable and applicable in diverse research and diagnostic settings, including laboratories with limited computational infrastructure.","manuscriptTitle":"Hardware-Aware RNA-seq Diagnostics: Plant Virus Detection via Cloud and AI","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-25 13:23:15","doi":"10.21203/rs.3.rs-6667739/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7fbcf27d-1ef1-4c32-8b26-654ff390bff1","owner":[],"postedDate":"September 25th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-02-20T14:41:01+00:00","versionOfRecord":[],"versionCreatedAt":"2025-09-25 13:23:15","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6667739","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6667739","identity":"rs-6667739","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.