{"paper_id":"00a0478b-8071-4726-8ae7-f46cbe15f076","body_text":"1 \nNanopore- and AI-empowered metagenomic viability inference 1 \nHarika Urel1,2,3, Sabrina Benassou 4, Tim Reska 1,2,3, Hanna Marti 5, Enrique Rayo 5, Edward J. 2 \nMartin6, Michael Schloter 3,7, James M. Ferguson 8, Stefan Kesselheim 4, Nicole Borel 5, Lara 3 \nUrban1,2,3* 4 \n 5 \n1 Helmholtz AI, Helmholtz Zentrum Muenchen, Neuherberg, Germany 6 \n2 Helmholtz Pioneer Campus, Helmholtz Zentrum Muenchen, Neuherberg, Germany 7 \n3 Technical University of Munich, School of Life Sciences, Freising, Germany 8 \n4 Jülich Supercomputing Centre, Forschungszentrum Jülich, Germany 9 \n5 Institute of Veterinary Pathology, University of Zürich, Switzerland 10 \n6 Institute of Ecology and Evolution, School of Biological Sciences, University of Edinburgh, 11 \nUnited Kingdom 12 \n7 Comparative Microbiome Analysis, Helmholtz Zentrum Muenchen, Neuherberg, Germany  13 \n8 Garvan Institute of Medical Research, Sydney, Australia 14 \n 15 \n*Corresponding author: lara.h.urban@gmail.com 16 \n 17 \nAbstract 18 \nThe ability to differentiate between viable and dead microorganisms in metagenomic samples 19 \nis crucial for various microbial inferences, ranging from assessing ecosystem functions of 20 \nenvironmental microbiomes to inferring the virulence of potential pathogens. While established 21 \nviability-resolved metagenomic approaches are labor-intensive as well as biased and lacking 22 \nin sensitivity, we here introduce a new fully computational framework that leverages nanopore 23 \nsequencing technology to assess microbial viability directly from freely available nanopore 24 \nsignal data. Our approach utilizes deep neural networks to learn features from such raw 25 \nnanopore signal data that can distinguish DNA from viable and dead microorganisms in a 26 \ncontrolled experimental setting. The application of explainable AI tools then allows us to 27 \nrobustly pinpoint the signal patterns in the nanopore raw data that allow the model to make 28 \nviability predictions at high accuracy. Using the model predictions as well as efficient 29 \nexplainable AI -based rules, we show that our framework can be leveraged in a real -world 30 \napplication to estimate the viability of pathogenic Chlamydia, where traditional culture-based 31 \nmethods suffer from inherently high false negative rates. This application shows that our 32 \nviability model captures predictive patterns in the nanopore signal that can in principle be 33 \nutilized to predict viability across taxonomic boundaries and indendent of the killing method 34 \nused to induce bacterial cell death. While the generalizability of our computational framework 35 \nneeds to be assessed in more detail, we here demonstrate for the first time the potential of 36 \nanalyzing freely available nanopore signal data to infer the viability of microorganisms, with 37 \nmany applications in environmental, veterinary, and clinical settings. 38 \n 39 \nAuthor summary  40 \nMetagenomics investigates the entirety of DNA isolated from an environment or a sample to 41 \nholistically understand microbial diversity in terms of known and newly discovered 42 \nmicroorganisms and their ecosystem functions. Unlike traditional culturing of microorganisms, 43 \nmetagenomics is not able to differentiate between viable and dead microorganisms since DNA 44 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n2 \nmight readily persist under different environmental circumstances. The viability of 45 \nmicroorganisms is, however, of importance when making inferences about a microorganism’s 46 \nmetabolic potential, a pathogen’s virulence, or an entire microbiome’s impact on its 47 \nenvironment. As existing viability -resolved metagenomic approaches are labor -intensive, 48 \nexpensive, and lack sensitivity, we here investigate our hypothesis if freely available nanopore 49 \nsequencing signal data, which captures DNA molecule information beyond the DNA sequence, 50 \nmight be leveraged to infer such viability. This hypothesis assumes that DNA from dead 51 \nmicroorganisms accumulates certain damage signatures that reflect microbial viability and can 52 \nbe read from nanopore signal data using fully computational frameworks. We here show first 53 \nevidence that such a computational framework might be feasible by training a deep model on 54 \ncontrolled experimental data to predict viability at high accuracy, exploring what the model has 55 \nlearned, and applying it to an independent real-world dataset of an infectious pathogen. While 56 \nthe generalizability of this computational framework needs to be assessed in much more detail, 57 \nwe demonstrate that freely available data might be usable for relevant viability inferences in 58 \nenvironmental, veterinary, and clinical settings. 59 \n 60 \nIntroduction 61 \nWhile microbial cultivation remains a foundational technique in microbiology to assess the 62 \ntaxonomic composition of microbial communities and to properly understand their  physiology 63 \nand ecosystem functions [ 1], only a small fraction of microbial diversity has been isolated in 64 \npure culture [2]. This limitation has led to undiscovered functions and biased representations 65 \nof the phylogenetic diversity of  microbial communities in nearly all of Earth’s environments [2]. 66 \nWhile medically relevant microorganisms of the human microbiome constitute an exemption 67 \nsince they have been disproportionately well studied through microbial cultures [3], the clinical 68 \napplication of microbial cultivation for pathogen profiling is also limited due to its time -69 \nconsuming and labor-intensive nature [4]. 70 \nThe first studies of the so -called “microbial dark matter” have been enabled by advances in 71 \nculture-independent molecular methodology [ 5], and have been based on amplifications of 72 \nconserved marker regions such as ribosomal RNA genes [ 6]. Such targeted metabarcoding 73 \napproaches, however, suffer from several limitations: They can often not provide strain - or 74 \neven species -level taxonomic resolution, are highly dependent on genomic database 75 \ncompleteness, do not allow for any functional inferences or virulence annotations, and often 76 \nintroduce amplification bias due to differential amplification efficiency and primer mismatches, 77 \nwhich can significantly distort the representation of microbial community compositions [7].  78 \nMetagenomics, on the other hand, is a shotgun sequencing -based molecular methodology 79 \nthat can assess the entirety of DNA isolated from an environment or a sample and de novo 80 \nassemblies of potentially complete microbial genomes of all present microorganisms; such 81 \ngenome-based approaches provide a variety of phylogenetically informative sequences for 82 \ntaxonomic classification, information about the metabolic and virulence potential of 83 \nmicroorganisms, and the potential to identify completely novel genes [8, 9].  84 \nEspecially long-read metagenomic approaches have shown great promise in achieving highly 85 \ncontiguous de novo assemblies through the recovery of high-quality metagenome-assembled 86 \ngenomes (MAGs) from complex environments; specifically, the latest advances in nanopore 87 \nsequencing technologies have resulted in high sequencing accuracies of very long sequencing 88 \nreads of up to millions of bases, which allowed for the generation of hundreds of MAGs from 89 \nmetagenomic data, including the generation of closed circularized genomes [ 10, 11]. 90 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n3 \nNanopore sequencing technology is based on the interpretation of the disruption of an ionic 91 \ncurrent due to a motor protein guiding individual nucleotide strands through nanopores 92 \nembedded in an electrically resistant polymer membrane at a consistent translocation speed 93 \n[12]. This raw nanopore signal, or “squiggle” data, can then be translated into nucleotide 94 \nsequence using bespoke neural network -based basecalling algorithms [ 13], which – when 95 \nefficiently embedded on powerful GPUs – can generate genomic data in real -time. The 96 \nportable character and straightforward implementation of nanopore sequencing at very upfront 97 \ninvestment costs further make this technology accessible for fast microbial and pathogens 98 \nassessments at point of interest all around the world, including in low - and middle-income 99 \ncountries [14]. 100 \nIn contrast to cultivation -based approaches, molecular methods suffer from their inherent 101 \ndeficiency of not being able to differentiate between viable and dead microorganisms [ 2, 15]. 102 \nWhile cultivation -based approaches only detect viable microorganisms, DNA might remain 103 \nintact and therefore accessible by molecular methods despite the respective microorganisms 104 \nbeing dead [ 15]. This might be especially relevant in the context of infection prevention and 105 \ncontrol and pathogen monitoring in the clinical setting, where certain disinfection methods or 106 \nthe use of systemic antibiotics often kill the bacteria before the DNA is destroyed [15, 16], but 107 \nalso for understanding the ecosystem functions of thus far understudied microbiomes [2]: For 108 \nexample, the air microbiome has been shown to be remarkable diverse and variable when 109 \nassessed through nanopore metagenomics [17], but given the low biomass of this environment 110 \nit is expected that many microorganisms might be dead and stem from adjacent environments 111 \nsuch as soil or water. The persistence of the DNA of dead microorganisms in the environment 112 \nmight hereby depend on many factors, including external conditions such as temperature, pH, 113 \nand microbial activity, and internal, taxon -specific parameters such as microbial cell wall 114 \ncomposition. Viability-resolved metagenomics would, however, be crucial for the interpretation 115 \nof metagenomic data, ranging from outbreak source detection [18], food safety [19] and public 116 \nhealth investigations [20], to ecosystem function inferences [21]. 117 \nTo assess microbial viability from genomic data, several approaches have been developed: 118 \nCulture-dependent viability methods combine the advantages of cultivation -based and 119 \nmolecular approaches by growing certain microorganisms of interest on selective media; this 120 \napproach, however, remains time -consuming and labor-intensive and suffers from the same 121 \nselectivity of growth media and culturable microorganisms as purely cultivation -based 122 \napproaches [ 22], especially for fastidious or obligate intracellular microorganisms [ 23]. 123 \nMicrobial viability can further be assessed through metabolic activity, where microbial cells are 124 \nincubated with specific substrates whose metabolization produces detectable signals such as 125 \nATP production, tetrazolium salt reduction, and radiolabeled substrate incorporation; this 126 \napproach may, however, not be applicable to all microbial taxa and environmental conditions, 127 \nand the results might be heavily influenced by the physiological state of the microbial cells [24]. 128 \nWhile RNA has been used as a viable/dead marker due to its intrinsic instability [ 25, 26, 27], 129 \nthe metatranscriptome has to be stable enough to be detectable in certain environments, but 130 \nunstable enough to reflect viability; if only one gene is targeted, the gene analyzed further has 131 \nto be continuously expressed [ 15]. Further issues might arise from the relatively challenging 132 \nextraction protocols due to the RNA’s instability, and from the relatively evolutionary 133 \nconservation of gene sequences, which might hamper taxonomic resolution.  134 \nFinally, an aspect that can be used for viability -resolved metagenomics is the physical 135 \ndifference between viable and dead cells: Viability PCR (vPCR) uses DNA -intercalating dyes 136 \nsuch as ethidium monoazide (EMA) or propidium monoazide (PMA) to differentiate between 137 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n4 \nviable and dead cells. These dyes penetrate only dead cells with compromised membranes 138 \nand bind to their DNA via covalent bonds upon photoactivation, preventing it from being 139 \namplified during subsequent PCR [15, 16]; this approach has been applied to a diverse array 140 \nof Gram-negative and -positive bacteria, and to assess the effectiveness of disinfection and 141 \nheat treatment [24]. It, however, relies on the assumption that membrane integrity is a reliable 142 \nindicator of viability, which can lead to overestimation of viability if cells lose viability without 143 \nimmediate membrane compromise [28] and can be biased by the dye’s variable permeability 144 \nacross different microbial cell wall structures [ 29, 30]. The dependence of the approach on 145 \nphotoactivation further means that turbid material might hamper the efficiency of the dye [31].  146 \nAll these established viability -resolved metagenomic approaches are labor -intensive, require 147 \nadditional reagents and sample processing, and are often biased and lack sensitivity. We here 148 \nhypothesized that the raw, freely available nanopore signal from metagenomic datasets might 149 \nbe leveraged to infer microbial viability, assuming that the native DNA from dead 150 \nmicroorganisms accumulates detectable squiggle signatures due to, e.g., external damage, 151 \nthe lack of DNA repair mechanisms, or the enzymatic activity [ 32, 33, 34]. Such an analysis 152 \nframework would be fully computational and utilize squiggle data that is automatically obtained 153 \nwith nanopore sequencing. While raw nanopore data is known to contain information about 154 \nepigenetic modifications [35, 36, 37] and oxidative stress at specific human telomere sites [38], 155 \nthe applicability to assess microbial viability has not yet been tested.  156 \nIn this study, we produced experimental nanopore sequencing data from viable and dead 157 \nEscherichia coli cultures to optimize deep neural networks to predict viability just from the 158 \nnanopore squiggle signal. We then applied explainable AI (XAI) tools, which allow us to identify 159 \nthe specific nanopore signal patterns in the input data that allow the model to deliver high -160 \naccuracy predictions as an output. We finally showed that our computational framework can 161 \nbe leveraged in a real -world application to estimate the viability of pathogenic Chlamydia 162 \nabortus, an obligate intracellular bacterial species with a complex biphasic lifecycle causing 163 \nenzootic abortion in sheep and goats [ 39]. Together, we show that our framework can in 164 \nprinciple predict viability across taxonomic boundaries and independent of the killing method 165 \nused to induce bacterial cell death. While we have not yet tested the generalizability of our 166 \ncomputational framework, we here demonstrate for the first time the potential of analyzing 167 \nfreely available nanopore signal data to infer the viability of microorganisms, with many 168 \napplications in environmental, veterinary, and clinical settings. 169 \n 170 \nResults & Discussion 171 \nNeural Network Training 172 \nWe generated controlled training data by nanopore sequencing native DNA of viable and dead 173 \nE. coli (Materials and Methods). We killed E. coli cultures using different stressors to then 174 \nisolate the extracellular DNA and expose it to natural degradation. We only obtained enough 175 \nDNA for subsequent shotgun sequencing from the viable culture and from the culture killed 176 \nthrough rapid UV exposure (viable: 212 ng/μL; UV: 5.46 ng/μL; heat shock: 0.03 ng/μL; bead 177 \nbeating: 0.67 ng/μL; Materials and Methods). We repeated this experiment and confirmed that 178 \nrapid heat shock as well as bead beating exposure again resulted in very low DNA 179 \nconcentrations, suggesting quick and complete DNA degradation. We hypothesize that UV 180 \nexposure constitutes the only stressor that simultaneously destroys bacterial cell walls as well 181 \nas inactivates DNA-degrading enzymes. In the case of killing by heat shock and bead beating, 182 \nsuch enzymes might have remained active and would have been carried forward by our 183 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n5 \nextracellular DNA isolation approaches (Materials and Methods) to then degrade all genomic 184 \nmaterial during natural exposure. We therefore created nanopore shotgun sequencing of the 185 \nviable and the UV-exposed culture, which resulted in 2.92 Gbases (Gb; median read length of 186 \n2,476 b) and 2.69 Gb (median read length of 1,606 b) of sequencing output, respectively 187 \n(Materials and Methods).  188 \nWe then tested the implementation of different neural network architectures to predict the 189 \nbinary viability state from the raw nanopore data (0=viable; 1=dead; Materials and Methods). 190 \nWe processed the E. coli nanopore signal, or “squiggle”, data, cut it into altogether 3,181,600 191 \nsignal chunks of 10k signals, and separated the chunks into balanced training (60%), validation 192 \n(20%), and test (20%) set along each original sequencing read to avoid that signal chunks 193 \nfrom the same read would end up in the same dataset (Materials and Methods). These signal 194 \nchunks can be treated as 1D time series signal data of consistent length. We trained the 195 \ndifferent model architectures using different learning rates (LRs) up to 1,000 epochs, 196 \nassessing the models’ performance based on training and validation loss after each epoch 197 \n(Materials and Methods; Table S1; Fig S1). The loss plot of our best -performing model, a 198 \nresidual neural network with convolutional input layers  (configuration ResNet1; LR=1e -4; 199 \nTable S1 ; Fig S1 ; Materials and Methods) shows minimal overfitting when the minimum 200 \nvalidation loss is reached at epoch 667 ( Fig 1A ). The other residual neural network 201 \narchitectures (ResNet2, ResNet3), on the other hand, resulted in overfitting to the training data 202 \nat any LR, and the transformer architecture did not reach the minimum validation loss of 203 \nResNet1 (Fig S1). We next only focused on ResNet1 and optimized its probability threshold 204 \nusing the validation set; in order to obtain a high accuracy and F1 score, we maintained the 205 \nprobability threshold at the default value of 0.5 ( Fig 1B ), which resulted in a good final 206 \nperformance on the test data with an accuracy of 0.83 and a F1 score of 0.81 ( Fig 1B, inlet) 207 \nas well as Area Under the Curve (AUC) values of 0.90 (Area Under the Receiver Operating 208 \nCharacteristic curve; AUROC) and 0.92 (Area Under the Precision -Recall curve; AUPR), 209 \nrespectively (Fig 1C).  210 \n 211 \n 212 \nFig 1. Training of the Viability Residual Neural Network (ResNet) 213 \n(A) Model loss for validation and test sets across 1,000 epochs; the minimum validation loss of the ResNet was 214 \nreached at epoch 677. (B) Prediction probability threshold optimization on the validation dataset resulted in a 215 \nprobability threshold of 0.5 for obtaining maximum accuracy and F1 score. Inlet: Performance of the final ResNet 216 \nmodel ResNet1 on the test data (Materials and Methods). (C) Test performance of the ResNet1 in terms of 217 \nPrecision-Recall (PR; pink) and Receiver Operating Characteristic (ROC; green) curves and their respective Areas 218 \nUnder the Curve (AUPR, AUROC).  219 \n 220 \nWe also trained the same residual neural network architecture ResNet1 on the basecalled 221 \nnanopore data of viable and dead E. coli at a standardized chunk size of 800 b, which roughly 222 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n6 \ncorresponds to a signal chunk size of 10k signals (Materials and Methods). Independently of 223 \nif we only basecalled the canonical bases or used a N6 -methyladenine (6mA) modification -224 \naware basecalling model (Materials and Methods), the model could not be trained to 225 \ndistinguish viable from dead data just from DNA sequence data ( Table S1). This shows that 226 \nour model captures patterns in the squiggle data that goes beyond the encoding of nucleotides 227 \nand their known epigenetic modifications. While this was expected since we used the same E. 228 \ncoli culture with the same reference genome to create the viable and dead datasets, we hereby 229 \nruled out that our squiggle -based model captured any random differences in DNA sequence 230 \ncontext between the two datasets that might have occurred by chance.  231 \n 232 \nWe additionally obtained the performance of ResNet1 for different signal chunk sizes (Fig S2; 233 \nTable S1), and we found that viability prediction performance was possible from a minimum 234 \nchunk size of approximately 5k, but can be further improved with increasing chunk size. This 235 \nshows that larger signal chunks contain more information that can be used by our model to 236 \nmake accurate per-chunk predictions while the resulting reduced size of the dataset, especially 237 \nof the training dataset, did not influence the learning ability of the model. While this might be 238 \nvaluable information for future applications where potentially the full length of a read could be 239 \nleveraged, we here decided to stick to a signal chunk size of 10k signals, which had already 240 \nresulted in good performance ( Fig 1; Table S1) and which can be applied to relatively small 241 \nsequencing reads as is often the case for metagenomic datasets (e.g., [17]). 242 \n 243 \nExplainable AI 244 \nWe implemented Class Activation Maps (CAM) as an XAI method [ 40] to identify the most 245 \nimportant regions in the nanopore signal data that inform the model’s viability classifications 246 \n(Materials and Methods; Fig 2A ). We found that “dead” signal chunks exhibited discrete 247 \nregions of increased CAM values (“CAM regions” defined at CAM values>0.8; Fig 2B  for 248 \nseveral true positive classifications of the test dataset). To confirm the importance of these 249 \nCAM regions for the model’s final predictions, we applied consecutive masking of the regions 250 \nwith the highest CAM values within each nanopore signal chunk (Fig S3 for several examples); 251 \nwe observed that the prediction probability for being classified as “dead” decreased with 252 \nincreased masking of CAM-relevant regions, either by consecutively masking regions using a 253 \nconsistent mask size or by increasing the mask size (from 100 to 2k signals; Fig 2C). This 254 \nshows that the CAM application reliably pinpoints patterns in the nanopore signal that are 255 \npredictive for our viability model.  256 \nWe used the CAM regions to manually investigate the squiggle signals, and found that many 257 \nCAM regions of “dead” signal chunks included a sudden substantial drop in the nanopore 258 \nsignal. We therefore developed a simple algorithm that identifies such sudden drops, and 259 \napplied this XAI rule to our test dataset (Materials and Methods; Fig 2D). We here classified 260 \nany signal chunk with at least one sudden drop as “dead”, and all others as “viable”. While this 261 \nsimplified algorithm led to a drop in overall performance, we could still reach a relatively good 262 \noverall accuracy of 0.68 (in comparison to 0.83 of the full model; Fig 2E). While the XAI rule 263 \nmaintained performance of specificity and precision, we observed a substantial drop in recall 264 \nin comparison to the full model (now 0.39 instead of 0.74). This shows that while the absence 265 \nof a sudden drop in the nanopore signal data seems to reliably predict viability, not all “dead” 266 \nsignal chunks seem to contain such a sudden drop. While this sudden -drop detection still 267 \nseems to be at the core of our model’s interpretability (when focusing on high-confidence true 268 \npositive chunks at p>0.99, the recall increased to 0.68), the model seems to additionally detect 269 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n7 \nmore subtle patterns in the nanopore signal data which allow it to increase recall while 270 \nmaintaining specificity and precision.  271 \nBased on our previous experience with squiggle data analysis [ 41, 42], we hypothesize that 272 \nthe substantial sudden drops in nanopore signal might be caused by a twist or kink in the DNA 273 \nbackbone, for example from 6-4pp pyrimidine dimers. The drop would then mark the event of 274 \na pore getting blocked due to such damage. Such a twist could also lead to a stalling signal if 275 \nit impairs the motor protein from processing the DNA strand, which we indeed partially 276 \nobserved in our data (e.g., top signal chunk of Fig 2B/D). While more detailed future squiggle 277 \nanalyses as well as application of our model to other viability studies will hopefully shed more 278 \nlight on the biological, chemical, and physical features detected by our model, we conclude 279 \nthat in our current study both, UV exposure ( E. coli) and heat shock (Chlamydia) might have 280 \ncaused such twists in the DNA backbone. 281 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n8 \n 282 \nFig 2. XAI for Interpretability of the Viability ResNet. 283 \n(A) The Class Activation Maps (CAMs) leverage the global average pooling (GAP) layer right before the fully 284 \nconnected (FC) layers of the residual neural network to map model interpretability onto the input features; they are 285 \ngenerated by aggregating the final convolutional layer's feature maps through a weighted sum, highlighting 286 \nnanopore signal regions that allow the neural network to make accurate predictions. (B) Exemplary nanopore signal 287 \nchunks that were classified as “dead” at a prediction probability of p>0.99, and their CAM values. Higher CAM 288 \nvalues indicate stronger feature map activations. (C) Impact of consecutive masking (n=5) of the signal region with 289 \nthe highest CAM value per signal chunk (x -axis) on the model’s prediction probability (y -axis); five different mask 290 \nsizes (from 100 to 2k signals) were used. (D) Application of a simplified XAI rule that classifies signal chunks 291 \naccording to the presence of a “sudden drop” (green: mean signal per chunk; red: threshold for sudden drop 292 \ndefinition; yellow: identification of sudden drops in the exemplary signal chunks). (E) Comparison of the 293 \nperformance of the full model (ResNet) with the simplified XAI rule based on the entire test dataset.  294 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n9 \nApplication to pathogenic Chlamydia 295 \nWe next applied our computational viability framework to distinguish viable from dead 296 \nChlamydia abortus, one of the most common causes of infectious abortion in small ruminants 297 \nworldwide [43, 44], and a zoonotic pathogen causing pneumonia and miscarriage in humans 298 \n[45]. C. abortus are obligate intracellular bacterial species with a complex biphasic life cycle 299 \nthat form intracellular vacuoles termed inclusions. These unique properties render both 300 \ncultivation- and vPCR-based approaches for viability estimations complicated [46]. In order to 301 \napply our computational viability framework to this pathogen, we decided to use two differently 302 \ntreated samples of the same strain of C. abortus samples for which both cultivation- and vPCR-303 \nbased approaches had predicted viability with high certainty. We chose a “viable” and a heat-304 \ntreated (“dead”) sample of the same culture for which we could ascertain viability and non -305 \nviability, respectively (Materials and Methods); heat treatment was used since it constitutes 306 \nthe standard approach for killing C. abortus  [46]. Briefly, for ascertaining viability and non -307 \nviability of the two samples, respectively, we applied cultivation and propidium monoazide 308 \n(PMA)-based vPCR. In the case of cultivation, heat -treated C. abortus were unable to form 309 \nviable inclusions in cell culture, whereas the untreated sample showed high infectivity with 310 \n2.6e6 inclusion forming units per mL (IFU/ml). In the case of vPCR, both samples were treated 311 \nwith and without PMA enhancer with or without PMA [ 46] to then quantify the respective 312 \namounts of chlamydial DNA with a sensitive Chlamydia-specific qPCR ( 47; Materials and 313 \nMethods). The difference in quantity between PMA-treated and untreated DNA was expressed 314 \nas Δlog10 Chlamydia per mL, resulting in 0.42 and 4.2 Δlog10 Chlamydia per mL for the control 315 \nand the heat-treated sample, respectively (Table 1). These data are comparable to a previous 316 \nstudy in which fresh C. trachomatis culture was heat -killed and a viability ratio determined 317 \nranging from 0% to 100% resulting in a 3.01 and 0.37 Δlog10 Chlamydia per mL for 0% and 318 \n100% viable, respectively [ 46]. Both cultivation and vPCR therefore confirmed that heat -319 \ntreatment had completely inactivated previously viable C. abortus for subsequent culture and 320 \nhad strongly reduced the amount of “viable” DNA using vPCR. 321 \n 322 \nTable 1. Viability PCR (vPCR) results of a “viable” and a “dead” C. abortus sample. 323 \nA viable and dead (heat-treated at 95°C for 10 min) sample of the same Chlamydia (C. abortus S26/3) stock were 324 \ntested for viability using vPCR: PMA-untreated vPCR reflects total Chlamydia content; PMA-treated vPCR reflects 325 \nviable Chlamydia content; IFU describes the number of Inclusion Forming Units. 326 \n 327 \nCondition Titer [IFU/mL] PMA-untreated \nvPCR [copy number \nper mL] \nPMA-treated vPCR \n[copy number per \nmL] \nViability ratio [%] \nviable 2.67e7 \n \n4.59e7 \n \n \n1.73e7 \n \n37.6 \ndead 0 9.03e7 5.52e7 \n \n0.01 \n 328 \nWe next created nanopore shotgun sequencing of these “viable” and “dead” C. abortus 329 \nsamples, which resulted in 37.50 Mb (median read length of 1,458 b) and 10.38 Mb (median 330 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n10 \nread length of 496 b) of sequencing output, respectively (Materials and Methods). We 331 \nsubsequently confirmed that the sequencing data was indeed created from the low -332 \nconcentration C. abortus samples, and not from any contaminating microorganisms: The  de 333 \nnovo assembly resulted in 18 and 39 contigs, respectively, which all mapped to C. abortus 334 \n(NCBI nt taxonomy id; 83555; Materials and Methods).  335 \nWe processed this nanopore sequencing data into 42,335 and 13,312 nanopore signal chunks 336 \nof 10k signal length for the viable and dead sample, respectively. We used our model 337 \n(ResNet1) to make computational viability predictions on this dataset, which resulted in an 338 \naccuracy of 0.85 and a F1 score of 0.68 at the previously optimized prediction probability 339 \nthreshold of 0.5. These performance values for the C. abortus  application are therefore 340 \ncomparable to the performance on the original test E. coli data, while the microorganism, killing 341 \nmethod, and DNA extractions methods were substantially different ( Fig 3A; top row). While 342 \nour model’s prediction probability distributions across signal chunks (Fig 3A; bottom row) were 343 \ncomparable between viable E. coli and C. abortus, dead C. abortus resulted in a more bimodal 344 \ndistribution than dead E. coli, with several “dead” chunks receiving close -to-zero prediction 345 \nprobabilities, suggesting viability (Fig 3A; bottom row; right column). The probability threshold-346 \nindependent application of our model to C. abortus consequently resulted in decreased AUC 347 \nvalues (AUROC of 0.80 and AUPR of 0.64; Fig 3B) in comparison to its application to E. coli 348 \n(AUROC of 0.90 and AUPR of 0.92; Fig 1C). While this difference in performance might just 349 \nmean that ResNet1 is less sensitive at detecting “dead” signal chunks in the  C. abortus data 350 \nthan in the E. coli data that it was trained on, it might also reflect true variability: The C. abortus 351 \nsample that was nanopore -sequenced to create “dead” nanopore signal data had strongly 352 \nreduced amounts of “viable” DNA according to the vPCR results but might have still contained 353 \nDNA from viable C. abortus bacterial cells – different from our dead E. coli sample which was 354 \nstringently filtered to only contain extracellular DNA and subsequently subjected to natural 355 \ndegradation (Materials and Methods). The latter hypothesis might be supported by the 356 \nobservation that the application of our CAM-based XAI to the “dead” C. abortus signal chunks 357 \ndetected similar patterns as in the “dead” E. coli signal chunks (Fig 3C). An application of our 358 \nXAI rule (Materials and Methods) further resulted in similar performance as its application to 359 \nthe E. coli signal chunks with good accuracy (0.78), high specificity (0.91), and low recall 360 \n(0.36).  361 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n11 \n 362 \nFig 3. Application of the Viability ResNet to the infectious pathogen Chlamydia. 363 \n(A) Model performance comparisons between E. coli (left) and C. abortus (right) datasets at the optimized prediction 364 \nprobability threshold of 0.5. Top row : Binary model predictions for viable and dead E. coli  and C. abortus , 365 \nrespectively; bottom row : Normalized distribution of model prediction probabilities across all signal chunks, 366 \nrespectively. For E. coli, all test signal chunks are visualized (388k “viable” and “dead” chunks, respectively), and 367 \nfor C. abortus, all signal chunks are visualized  (42,335 “viable” and 13,312 “dead” chunks). (B) Performance of the 368 \nmodel on the C. abortus dataset in terms of Precision -Recall (PR; pink) and Receiver Operating Characteristic 369 \n(ROC; green) curves and their respective Areas Under the Curve (AUPR, AUROC). (C) Exemplary Chlamydia 370 \nnanopore signal chunks that were correctly classified as “dead”, and their CAM values. Higher CAM values indicate 371 \nstronger feature map activations. 372 \n 373 \nAI- and nanopore-empowered viability-resolved metagenomics 374 \nWhile metagenomic approaches provide the unique opportunity of generating de novo 375 \nassemblies and potentially complete microbial genomes to explore the “microbial dark matter” 376 \nas well as to infer potential functions such as metabolic and virulence potential [ 8, 9], they 377 \nhave suffered from their inability to differentiate between viable and dead microorganism [ 2, 378 \n15]. Such viability inferences can, however, distort any microbial inference, ranging from 379 \nassessing ecosystem functions of environmental microbiomes to inferring the virulence of 380 \npotential pathogens. As established viability -resolved metagenomic approaches are labor -381 \nintensive as well as biased and lack sensitivity (e.g., 22), we here show first evidence that a 382 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n12 \nfully computational framework based on residual neural networks with convolutional data 383 \nprocessing layers can leverage raw nanopore signal data, also known as squiggle data, to 384 \nmake accurate inferences about microbial viability (test set accuracy of 0.83).   385 \nOur subsequent XAI analyses point to the potential role of DNA backbone damage for 386 \nachieving accurate model predictions; however, more work is needed to fully understand the 387 \nbiological, physical, or chemical underpinnings of our viability model predictions. First, the 388 \nsimplified XAI rule is not sufficient to correctly classify the majority of “dead” signal chunks 389 \n(recall of 0.39), which means that the residual neural network has apparently picked up on 390 \nadditional, more ambiguous signal patterns that allow the full model to make more sensitive 391 \npredictions (recall of 0.74). Besides exploring the underlying rules of such additional signal 392 \npatterns, we will also apply the viability model to more datasets to assess its generalizability. 393 \nEspecially the probing of different taxonomic groups and killing methods should help us tease 394 \napart the origins of our current XAI results. Our study, however, already provides first evidence 395 \nthat our viability model captures predictive patterns in the nanopore signal that can in principle 396 \nbe utilized to predict viability across taxonomic boundaries and independent of the killing 397 \nmethod. The application to estimate the viability of pathogenic Chlamydia (prediction accuracy 398 \nof 0.85) is hereby of potentially immediate interest to veterinary scientists since traditional 399 \nmethods for assessing the pathogen’s viability have been labor -intensive and suffered from 400 \ninherently high false negative rates.  401 \nWhile the generalizability of our model needs to be assessed in much more detail, including 402 \nfor other microbial taxa such as spore -forming bacteria or fungi and for real -world 403 \nmetagenomics applications, the potential benefits of such a widely applicable computational 404 \nframework could be immense, with many potential applications in environmental, veterinary, 405 \nand clinical settings. As is the case for epigenetic inferences [35, 36, 37], the viability inference-406 \nenabling squiggle data is a complementary output of any nanopore sequencing experiment of 407 \nnative DNA, that is then usually basecalled and archived for future re -basecalling after 408 \npotential basecalling model improvements. This means that any future nanopore -based 409 \nmetagenomic study could make viability predictions for free without additional costs and 410 \nlaboratory work, and that any existing archived nanopore data could be assessed in terms of 411 \nits microorganisms’ viability – which would allow us to quantify the impact of dead 412 \nmicroorganisms on metagenomics in general, and to further explore factors such as species - 413 \nand environment -specificity. We finally anticipate that quantitative AI modeling has the 414 \npotential to inform more differentiated viability assessments, which might help quantify or even 415 \ntime degradation events and decipher the impact of dormancy on metagenomic studies [ 48, 416 \n49].  417 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n13 \nMaterials and Methods 418 \nTraining Data Generation 419 \nWe cultured E. coli K12 in 200 mL Luria -Bertani (LB) medium for 24 hours at 37°C to reach 420 \nthe log phase of the growth curve. The culture was then used to inoculate four 200 mL LB 421 \nmedia in 1L Erlenmeyer flasks, which were again incubated for 24 hours to reach the growth 422 \nlog phase. One of the media was used as viable control, i.e. DNA was extracted from 750 uL 423 \nof the living culture using the spin-column based QIAGEN PowerSoil Pro Kit (QIAGEN, 2018, 424 \nHilden, Germany), following the manufacturers’ instructions. The remaining three cultures 425 \nwere killed by one of the following stressors: UV irradiation at 254 nm for 15 min, heat shock 426 \nat 120°C for 5 min, or bead beating for 30 min. To then separate extracellular DNA from cell 427 \ndebris and intact bacterial cells, we centrifuged the media for 10 min at 4,000 x g and filtered 428 \nthe supernatant through 0.2 µm filters. The resulting extracellular DNA was subsequently kept 429 \nat room temperature for 5 days to simulate the natural accumulation of DNA degradation. DNA 430 \nfrom dead bacteria was extracted from these samples using the same extraction approach 431 \nfollowing the QIAGEN PowerSoil Pro Kit protocol, but the first lysis buffer step was omitted 432 \nsince cell lysis had already happened.  433 \nWe then used Oxford Nanopore Technologies’ Rapid Barcoding library preparation kit 434 \n(RBK114-24 V14), R10.4.1 MinION flow cells, and MinKNOW software v23.04.5 for shotgun 435 \nnanopore sequencing of the “viable” and “dead” DNA. We used four barcodes for each sample, 436 \nresulting in DNA input of 800 ng and 218 ng for the preparation of the “viable” and “dead” 437 \nlibrary, respectively. We ran each library for 24 h, using two different flow cells to avoid any 438 \ncross contamination, and filtered the resulting nanopore data at a minimum read length of 20 439 \nb. Raw nanopore data was created using the standard translocation speed of 400 b/s, and a 440 \nsampling frequency of 5 kHz. 441 \nWe applied Dorado v4.2.0 ( https://github.com/nanoporetech/dorado) SUP -basecalling 442 \n(dna_r10.4.1_e8.2_400bps_sup@v4.2.0) and 6mA -aware SUP-basecalling (6mA@v1) to all 443 \nnanopore reads that had passed internal data quality thresholds to obtain E. coli  DNA 444 \nsequence data. We subsequently removed sequencing adapters and barcodes using 445 \nPorechop v0.2.3 (https://github.com/rrwick/Porechop).  446 \n 447 \nNeural Network Architecture and Training 448 \nWe tested the implementation of different residual neural networks and transformer 449 \narchitectures to predict the binary viability state from the raw nanopore data (0=viable; 450 \n1=dead). The first residual neural network, ResNet1, consists of four layers, each containing 451 \ntwo bottleneck blocks. Each bottleneck block consists of convolutional layers, batch 452 \nnormalization, and a rectified linear unit (ReLu) activation function. Each of the four layers 453 \nconsists of an increasing number of convolutional channels: 20, 30, 45, and 65, respectively, 454 \nfollowed by global average pooling and a fully connected layer, resulting in 66,916 parameters. 455 \nThe model then uses a softmax function to convert logits, the raw outputs from the fully 456 \nconnected layer, into predicted probabilities ranging from 0 to 1. We evaluated the training of 457 \nthe model using the Adam optimizer for mini-batch gradient descent at three different LRs (1e-458 \n3, 1e-4, and 1e-5), training the model up to 1,000 epochs and at a batch size of 1,000 signal 459 \nchunks. We initialized the model using Kaiming initialization. For ResNet2, we increased the 460 \nnumber of convolutional channels to 40, 60, 90, and 135, respectively, resulting in 1,828,777 461 \nparameters. For ResNet3, we increased the number of convolutional channels to 512, 30, 45, 462 \nand 67, respectively, resulting in 2,479,140 parameters. The transformer model was based on 463 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n14 \na positional encoding, convolutional layer with a channel number of 24 and one block of one 464 \nattention head, resulting 219,586 parameters. 465 \nWe processed the E. coli squiggle data by excluding the first 1,500 signal points (potential 466 \nnoise, adapter sequences, or barcodes), then cutting it into signal chunks of 10k signals, and 467 \nseparated the chunks into balanced training (60%), validation (20%), and test (20%) set along 468 \neach original sequencing read to avoid that signal chunks from the same read would end up 469 \nin the same dataset. We pooled the viable and dead signals chunks to obtain exactly balanced 470 \ntraining, validation, and test sets. For normalizing each chunk, we subtracted the median per 471 \nchunk and divided it by the median absolute deviation (MAD) to make the signal data robust 472 \nto outliers. We then scaled the signal by the MAD scaling factor 1.4826, and replaced outliers 473 \nexceeding 3.5 times the scaled MAD by the mean of their two neighboring values. 474 \nWe also trained ResNet1 on the basecalled nanopore data (with or without 6mA basecalling) 475 \nof viable and dead E. coli at a standardized chunk size of 800 b, which roughly corresponds 476 \nto a signal chunk size of 10k signals. For encoding, we used a one -hot encoding method to 477 \nturn DNA sequence into unique binary vectors. We then concatenated and saved these 478 \nencoded sequences as tensors for training and testing. We finally trained ResNet1 on signal 479 \nchunks of different signal lengths, ranging from 1k to 20k signals.  480 \n 481 \nExplainable AI 482 \nWe utilized CAMs to identify and visualize signals regions that influenced the model's decision-483 \nmaking. As feature maps from the final convolutional layer undergo a global average pooling 484 \nlayer where each map is averaged and concatenated, we can calculate the weighted sum of 485 \nthese feature maps using the weights of the fully connected layer and project it back onto the 486 \npreprocessed signal [ 40]. To do so, we implemented CAM in Python/PyTorch. During the 487 \nforward pass, we ensure that the feature maps from the last convolutional layer are captured. 488 \nTo compute the CAM, we use the weights of the model's output layer for the class of interest, 489 \nmultiplying these weights with the corresponding feature maps and then summing them up. 490 \nWe convert the resulting CAM to an array and normalize its values to a range of 0 to 1. We 491 \nthen overlay the CAM on the original input signal to identify the regions most influential in the 492 \nmodel's decision -making process. We additionally used the Remora API to match raw 493 \nnanopore data to the corresponding Dorado -basecalled bases, to then manually investigate 494 \nany obvious sequence abnormalities in the CAM regions. 495 \nFor consecutive masking of the CAM regions with the highest CAM values, we used Python 496 \nto first obtain and normalize the CAM values of all true -positive signal chunks at p>0.5 497 \n(n=286,179), identify the maximum value, mask the signal region (i.e., setting to zero after 498 \nMAD-normlization) at a specified mask size (between 100 and 2k signals) centered around 499 \nthe maximum -CAM signal, and obtain updated prediction probabilities. We repeated this 500 \nmasking step five times and calculated the confidence interval at each masking step: We 501 \nobtained the mean and Standard Error of the Mean (SEM) of the newly calculated prediction 502 \nprobabilities across all signal chunks to calculate the 95% confidence interval at mean ± 1.96 503 \n* SEM. We plotted the results using matplotlib. 504 \nWe next used Python to develop an algorithm to obtain a simplified XAI rule to distinguish 505 \n“dead” from “viable” signals chunks based on our CAM results by identifying the presence of 506 \nat least one sudden drop in the chunk. To identify those sudden drops, we first calculated the 507 \nmean and standard deviation (SD) of each signal chunk, and found that a scaling factor of 3 508 \nidentified most manually detected sudden drops at a vertical threshold of mean - 3 * SD. 509 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n15 \nApplication to pathogenic Chlamydia 510 \nThe Chlamydia abortus strain S26/3 associated with ovine abortion (provided by Dr. G.E. 511 \nJones, Moredun Research Institute, Edinburgh, UK) was cultured as described by Borel et al.  512 \n[50]. Briefly, the strain was first grown in embryonated chicken eggs and stored at -80°C 513 \nfollowing 1:2 dilution in sucrose-phosphate-glutamate buffer (SPG). For propagation, the strain 514 \nwas cultured in HEp -2 cells (ATCC CCL -23). Infectious elementary bodies were separated 515 \nfrom cell debris as well as non -infectious chlamydial reticulate bodies using a renografin 516 \ndensity gradient [51], resuspended in SPG, and stored in aliquots at -80°C. To determine the 517 \nviability of this stock, one aliquot was thawed on ice and separated into two tubes of which one 518 \nwas heat-treated for 10 min at 95°C. Both the control (“viable”) and the heat -treated (“dead”) 519 \nsamples were then divided into subsamples, which were subsequently used for cultivation, 520 \nvPCR, and nanopore sequencing.  521 \nFor viability determination by culture, 100 µl per sample was used to infect two glass coverslips 522 \n(13 mm in diameter, ThermoScientific, Waltham, MA, USA) in 24 -well plates (TPP Techno 523 \nPlastic Product AG, Trasadingen, Switzerland) seeded to confluence with LLC -MK2 cells (524 \nrhesus monkey kidney cell line; provided by IZSLER, Brescia, Italy) [52]. Following inoculation 525 \nof the monolayer, infection was enhanced by centrifugation for 1  h at 25°C (1000 x g). After 526 \n48 h of incubation at 37°C (5% CO ₂), cultures were fixed for 10  min in ice -cold methanol. 527 \nCoverslips were then processed using a well -established immunofluorescence assay [ 53]. 528 \nBriefly, DNA was stained with 1 μg/mL 4', 6-diamidino-2’-phenylindole dihydrochloride (DAPI, 529 \nMolecular Probes, Eugene, OR, USA). In parallel, inclusions were labeled with a 530 \nChlamydiaceae-specific primary antibody ( Chlamydiaceae LPS, Clone ACI -P; Progen, 531 \nGermany), which was diluted 1:200  in blocking solution consisting of 1% bovine serum 532 \nalbumin (BSA, St. Louis, MO, USA) in phosphate -buffered saline (PBS, GIBCO, Invitrogen, 533 \nCarlsbad, CA, USA). Inclusions were then visualized with Alexa Fluor 488 goat anti -mouse 534 \n(Molecular Probes) diluted 1:500 in blocking solution. As a final step, coverslips were washed 535 \nwith PBS, mounted with FluoreGuard (Hard Set; ScyTek Laboratories Inc., Logan, UT, USA) 536 \non glass slides, and inclusions determined using a Leica DMLB fluorescence microscope 537 \n(Leica Microsystems, Wetzlar, Germany) and a 10X ocular objective (Leica L-Plan 10x/ 25 M, 538 \nLeica Microsystems). In parallel, a three -fold dilution series of the sample was performed in 539 \n96-well plates (TPP) and processed as above. The number of IFU/ml was then determined 540 \nusing the Nikon Ti Eclipse epifluorescence microscope (Nikon, Tokyo, Japan) at a 20X 541 \nmagnification [52]. 542 \nFor vPCR, two 100  µl subsamples were taken and mixed with 200 µl SPG and 100 µl PMA 543 \nenhancer for Gram Negative Bacteria (5X, Biotium, Fremont, CA, USA). PMAxx (Biotium) at a 544 \nfinal concentration of 50  µM was added (“PMA -treated”) or not (“untreated”) to the 545 \nsubsamples. Samples were then exposed to a 650 -W light source using a PMA -Lite LED 546 \nPhotolysis Device (Biotium) for 5  min, followed by 2 min on ice and additional light exposure 547 \nfor 5 min [46]. 548 \nFor vPCR as well as nanopore sequencing, DNA was extracted using the DNeasy® Blood and 549 \nTissue Kit (QIAGEN, Hilden, Germany) according to the manufacturer’s instructions. The 550 \namount of chlamydial DNA in all vPCR and nanopore sequencing samples was quantified with 551 \na sensitive Chlamydiaceae qPCR [ 47]. For subsequent nanopore sequencing of the DNA 552 \nextracts (viable: 0.62 ng/μL; dead: 0.17 ng/μL), we followed the same approach as described 553 \nfor E. coli above. We used three barcodes for each sample, resulting in DNA input of 18.66 ng 554 \nand 5.07 ng for the “viable” and “dead” library, respectively. After following the same 555 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n16 \nprocessing steps as established for the E. coli experiment, we filtered reads at a minimum 556 \naverage quality score of 8 and a minimum length of 100 b using Nanofilt v2.8.0 [ 54], and 557 \ncreated de novo assemblies using metaFlye v2.9.1 [55], followed by polishing using minimap2 558 \nv2.17 [56] and three rounds of Racon v1.5 [ 57]. We finally used Kraken2 v2.0.7 [ 58] and the 559 \nNCBI nt database for taxonomic classification of the assembled contigs. 560 \n 561 \nData Availability Statement 562 \nAll raw data has been made publicly available via ENA (study accession number: 563 \nPRJEB76420). All code has been made publicly available via Github: 564 \nhttps://github.com/Genomics4OneHealth/Squiggle4Viability.git. 565 \n 566 \nFinancial Disclosure Statement 567 \nThis study was funded by a Helmholtz Principal Investigator Grant awarded to LU. HU was 568 \nsupported by the Helmholtz Association under the joint research school “HIDSS-006 - Munich 569 \nSchool for Data Science@Helmholtz, TUM&LMU”. EM was supported by an EASTBIO 570 \nstudentship, funded by BBSRC Grant Number BB/M010996/1, and an STFC Food Network+ 571 \nScoping Grant. SB and SK were supported by Helmholtz AI’s Helmholtz Association Initiative 572 \nand Networking Fund. Computational Resources were provided by Helmholtz Munich and by 573 \nthe Helmholtz Association Initiative and Networking Fund HAICORE partition at the 574 \nForschungszentrum Jülich. LU, HU, and JMF have previously received travel and 575 \naccommodation expenses to speak at Oxford Nanopore Technologies’ conferences.  576 \n 577 \nAcknowledgments 578 \nWe thank the laboratory of the Research Unit of Comparative Microbiome Analysis at 579 \nHelmholtz Munich, Germany, especially Cornelia Galonska, for their support in processing the 580 \nEscherichia coli samples. We thank the laboratory of the Institute of Veterinary Pathology, 581 \nSwitzerland, especially Theresa Pesch, for their support in processing the Chlamydia abortus 582 \nsamples. We further thank Valentin Rauscher for his help in training several deep models 583 \nduring his internship in the Urban research group at Helmholtz Munich.  584 \n 585 \nCompeting interests 586 \nThe authors declare no competing interests.   587 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n17 \nReferences 588 \n1. Lewis WH, Tahon G, Geesink P, Sousa DZ, Ettema TJG. Innovations to culturing the uncultured microbial 589 \nmajority. Nat Rev Microbiol. 2021 Apr;19(4):225-40. 590 \n2. Lloyd KG, Steen AD, Ladau J, Yin J, Crosby L. Phylogenetically novel uncultured microbial cells dominate 591 \nEarth microbiomes. mSystems. 2018 Sep 25;3(5). 592 \n3. Human Microbiome Jumpstart Reference Strains Consortium, Nelson KE, Weinstock GM, Highlander SK, 593 \nWorley KC, Creasy HH, et al. A catalog of reference genomes from the human microbiome. Science. 2010 594 \nMay 21;328(5981):994-9. 595 \n4. Sauerborn E, Corredor C, Reska T, Perlas A, Vargas da Fonseca Atum S, Goldman N, Wantia N, Prazeres 596 \nda Costa C, Foster -Nyarko E, Urban L. Detection of hidden antibiotic resistance through real -time 597 \ngenomics. Preprint. Research Square. 2023 Dec 4. Available from: https://doi.org/10.21203/rs.3.rs-598 \n3620416/v1. 599 \n5. Handelsman J. Metagenomics: application of genomics to uncultured microorganisms. Microbiol Mol Biol 600 \nRev. 2004 Dec;68(4):669-85.. 601 \n6. Pace NR. Mapping the tree of life: progress and prospects. Microbiol Mol Biol Rev. 2009 Dec;73(4):565-602 \n76. 603 \n7. Brooks JP, Edwards DJ, Harwich MD, Rivera MC, Fettweis JM, Serrano MG, et al. The truth about 604 \nmetagenomics: quantifying and counteracting bias in 16S rRNA studies. BMC Microbiol. 2015;15:66. 605 \n8. Dick GJ, Andersson AF, Baker BJ, Simmons SL, Thomas BC, Yelton AP, Banfield JF. Community -wide 606 \nanalysis of microbial genome sequence signatures. Genome Biol. 2009;10(8). 607 \n9. Quince C, Walker AW, Simpson JT, Loman NJ, Segata N. Shotgun metagenomics, from sampling to 608 \nanalysis. Nat Biotechnol. 2017 Sep;35(9):833-44. 609 \n10. Sereika M, Kirkegaard RH, Karst SM, Michaelsen TY, Sørensen EA, Wollenberg RD, Albertsen M. Oxford 610 \nNanopore R10.4 long-read sequencing enables the generation of near -finished bacterial genomes from 611 \npure cultures and metagenomes without short -read or reference polishing. Nat Methods. 2022 612 \nJul;19(7):823-6. 613 \n11. Liu L, Yang Y, Deng Y, Zhang T. Nanopore long -read-only metagenomics enables complete and high -614 \nquality genome reconstruction from mock and complex metagenomes. Microbiome. 2022 Dec 615 \n2;10(1):209. 616 \n12. Jain M, Olsen HE, Paten B, Akeson M. The Oxford Nanopore MinION: delivery of nanopore sequencing 617 \nto the genomics community. Genome Biol. 2016;17:239. 618 \n13. Pagès-Gallego M, de Ridder J. Comprehensive benchmark and architectural analysis of deep learning 619 \nmodels for nanopore sequencing basecalling. Genome Biol. 2023 Apr 11;24(1):71.  620 \n14. Urban L, Perlas A, Francino O, Martí -Carreras J, Muga BA, Mwangi JW, et al. Real -time genomics for 621 \nOne Health. Mol Syst Biol. 2023 Aug 8;19(8). 622 \n15. Nogva HK, Drømtorp SM, Nissen H, Rudi K. Ethidium monoazide for DNA -based differentiation of viable 623 \nand dead bacteria by 5'-nuclease PCR. Biotechniques. 2003 Apr;34(4):804-8, 810, 812-3. 624 \n16. Hellmann KT, Tuura CE, Fish J, Patel JM, Robinson DA. Viability -resolved metagenomics reveals 625 \nantagonistic colonization dynamics of Staphylococcus epidermidis strains on preterm infant skin. 626 \nmSphere. 2021 Oct 27;6(5). 627 \n17. Reska T, Pozdniakova S, Borras S, Schloter M, Cañas L, Perlas A, Rodó X, Winkler B, Schnitzler JP, 628 \nUrban L. Air monitoring by nanopore sequencing. bioRxiv. 2023 Dec 19:2023.12.19.572325. 629 \ndoi:10.1101/2023.12.19.572325. 630 \n18. Morré SA, Sillekens PT, Jacobs MV, de Blok S, Ossewaarde JM, van Aarle P, et al. Monitoring of 631 \nChlamydia trachomatis infections after antibiotic treatment using RNA detection by nucleic acid sequence 632 \nbased amplification. Mol Pathol. 1998 Jun;51(3):149-54. 633 \n19. Herman L. Detection of viable and dead Listeria monocytogenes by PCR. Food Microbiol. 1997 634 \nApr;14(2):103-110. 635 \n20. Min J, Baeumner AJ. Highly sensitive and specific detection of viable Escherichia coli in drinking water. 636 \nAnal Biochem. 2002 Apr 15;303(2):186-93. 637 \n21. Urban L, Holzer A, Baronas JJ, Hall MB, Braeuninger-Weimer P, Scherm MJ, et al. Freshwater monitoring 638 \nby nanopore sequencing. eLife. 2021 Jan 19;10. 639 \n22. Kumar SS, Ghosh AR. Assessment of bacterial viability: a comprehensive review on recent advances and 640 \nchallenges. Microbiology (Reading). 2019 Jun;165(6):593-610. 641 \n23. Taylor-Brown A, Madden D, Polkinghorne A. Culture -independent approaches to chlamydial genomics. 642 \nMicrob Genom. 2018 Feb;4(2). 643 \n24. Nocker A, Camper AK. Novel approaches toward preferential detection of viable cells using nucleic acid 644 \namplification techniques. FEMS Microbiol Lett. 2009 Feb;291(2):137-42. 645 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n18 \n25. Sheridan GE, Masters CI, Shallcross JA, Mackey BM. Detection of mRNA by reverse transcription -PCR 646 \nas an indicator of viability in Escherichia coli cells. Appl Environ Microbiol. 1998 Apr;64(4):1313-8. 647 \n26. Wong VY, Duval MX. Inter -laboratory variability in array -based RNA quantification methods. Genomics 648 \nInsights. 2013 May 6;6:13-24. 649 \n27. Hønsvall BK, Robertson LJ. From research lab to standard environmental analysis tool: Will NASBA make 650 \nthe leap? Water Res. 2017 Feb 1;109:389-397.  651 \n28. Janssen KJH, Dirks JAMC, Dukers -Muijrers NHTM, Hoebe CJPA, Wolffs PFG. Review of Chlamydia 652 \ntrachomatis viability methods: assessing the clinical diagnostic impact of NAAT positive results. Expert 653 \nRev Mol Diagn. 2018 Aug;18(8):739-747.  654 \n29. Nocker A, Cheung CY, Camper AK. Comparison of propidium monoazide with ethidium monoazide for 655 \ndifferentiation of live vs. dead bacteria by selective removal of DNA from dead cells. J Microbiol Methods. 656 \n2006 Nov;67(2):310-20. 657 \n30. Emerson JB, Adams RI, Román CMB, Brooks B, Coil DA, Dahlhausen K, et al. Schrödinger's microbes: 658 \nTools for distinguishing the living from the dead in microbial ecosystems. Microbiome. 2017 Aug 659 \n16;5(1):86. 660 \n31. Bae S, Wuertz S. Discrimination of viable and dead fecal Bacteroidales bacteria by quantitative PCR with 661 \npropidium monoazide. Appl Environ Microbiol. 2009 May;75(9):2940-4.  662 \n32. Courcelle J, Donaldson JR, Chow KH, Courcelle CT. DNA damage -induced replication fork regression 663 \nand processing in Escherichia coli. Science. 2003 Feb 14;299(5609):1064-7. 664 \n33. Setlow RB, Swenson PA, Carrier WL. Thymine dimers and inhibition of DNA synthesis by ultraviolet 665 \nirradiation of cells. Science. 1963 Dec 13;142(3598):1464-6. 666 \n34. Shibai A, Takahashi Y, Ishizawa Y, Motooka D, Nakamura S, Ying BW, Tsuru S. Mutation accumulation 667 \nunder UV radiation in Escherichia coli. Sci Rep. 2017 Nov 6;7(1):14531.  668 \n35. Ni P, Huang N, Zhang Z, Wang DP, Liang F, Miao Y, Xiao CL, Luo F, Wang J. DeepSignal: detecting DNA 669 \nmethylation state from Nanopore sequencing reads using deep -learning. Bioinformatics. 2019 Nov 670 \n1;35(22):4586-95. 671 \n36. Stoiber M, Quick J, Egan R, Lee JE, Celniker S, Neely RK, Loman N, Pennacchio LA, Brown J. De novo 672 \nidentification of DNA modifications enabled by genome-guided nanopore signal processing. bioRxiv. 2016 673 \nDec 15;094672. doi:10.1101/094672. 674 \n37. Wang X, Ahsan MU, Zhou Y, Wang K. Transformer -based DNA methylation detection on ionic signals 675 \nfrom Oxford Nanopore sequencing data. Quant Biol. 2023;11(3):287-96. 676 \n38. An N, Fleming AM, White HS, Burrows CJ. Nanopore detection of 8 -oxoguanine in the human telomere 677 \nrepeat sequence. ACS Nano. 2015 Apr 28;9(4):4296-307. 678 \n39. Turin L, Surini S, Wheelhouse N, Rocchi MS. Recent advances and public health implications for 679 \nenvironmental exposure to Chlamydia abortus: from enzootic to zoonotic disease. Vet Res. 2022 May 680 \n31;53(1):37. 681 \n40. Zhou B, Khosla A, Lapedriza A, Oliva A, Torralba A. Learning deep features for discriminative localization. 682 \nIn: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2016 Jun 27 -30; 683 \nLas Vegas, NV, USA. p. 2921-9. 684 \n41. Ferguson JM, Smith MA. SquiggleKit: a toolkit for manipulating nanopore signal data. Bioinformatics. 2019 685 \nDec 15;35(24):5372-3. 686 \n42. Gamaarachchi H, Samarakoon H, Jenner SP, Ferguson JM, Amos TG, Hammond JM, et al. Fast 687 \nnanopore sequencing data analysis with SLOW5. Nat Biotechnol. 2022 Jul;40(7):1026-9.  688 \n43. Borel N, Polkinghorne A, Pospischil A. A review on chlamydial diseases in animals: still a challenge for 689 \npathologists? Vet Pathol. 2018 May;55(3):374-90. 690 \n44. Borel N, Sachse K. Zoonotic transmission of Chlamydia spp.: known for 140 years, but still 691 \nunderestimated. In: Sing A, editor. Zoonoses: infections affecting humans and animals. Cham: Springer 692 \nInternational Publishing; 2023. p. 1-28. 693 \n45. Borel N, Marti H, Pospischil A, Pesch T, Prähauser B, Wunderlin S, et al. Chlamydiae in human intestinal 694 \nbiopsy samples. Pathog Dis. 2018 Nov 1;76(8). 695 \n46. Janssen KJ, Hoebe CJ, Dukers-Muijrers NH, Eppings L, Lucchesi M, Wolffs PF. Viability-PCR shows that 696 \nNAAT detects a high proportion of DNA from non -viable Chlamydia trachomatis. PLoS One. 2016 Nov 697 \n3;11(11). 698 \n47. Loehrer S, Hagenbuch F, Marti H, Pesch T, Hässig M, Borel N. Longitudinal study of Chlamydia pecorum 699 \nin a healthy Swiss cattle population. PLoS One. 2023 Dec 11;18(12). 700 \n48. McDonald MD, Owusu -Ansah C, Ellenbogen JB, Malone ZD, Ricketts MP, Frolking SE, et al. What is 701 \nmicrobial dormancy? Trends Microbiol. 2024 Feb;32(2):142-50. 702 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n19 \n49. Potgieter M, Bester J, Kell DB, Pretorius E. The dormant blood microbiome in chronic, inflammatory 703 \ndiseases. FEMS Microbiol Rev. 2015 Jul;39(4):567-91. 704 \n50. Borel N, Dumrese C, Ziegler U, Schifferli A, Kaiser C, Pospischil A. Mixed infections with Chlamydia and 705 \nporcine epidemic diarrhea virus - a new in vitro model of chlamydial persistence. BMC Microbiol. 2010 Jul 706 \n27;10:201. 707 \n51. Howard L, Orenstein NS, King NW. Purification on renografin density gradients of Chlamydia trachomatis 708 \ngrown in the yolk sac of eggs. Appl Microbiol. 1974 Jan;27(1):102-6. 709 \n52. Marti H, Biggel M, Shima K, Onorini D, Rupp J, Charette SJ, Borel N. Chlamydia suis displays high 710 \ntransformation capacity with complete cloning vector integration into the chromosomal rrn -nqrF plasticity 711 \nzone. Microbiol Spectr. 2023 Dec 12;11(6). 712 \n53. Leonard CA, Schoborg RV, Borel N. Damage/danger associated molecular patterns (DAMPs) modulate 713 \nChlamydia pecorum and C. trachomatis serovar E inclusion development in vitro. PLoS One. 2015 Aug 714 \n6;10(8). 715 \n54. De Coster W, D'Hert S, Schultz DT, Cruts M, Van Broeckhoven C. NanoPack: visualizing and processing 716 \nlong-read sequencing data. Bioinformatics. 2018 Aug 1;34(15):2666-9. 717 \n55. Kolmogorov M, Bickhart DM, Behsaz B, Gurevich A, Rayko M, Shin SB, Kuhn K, Yuan J, Polevikov E, 718 \nSmith TPL, Pevzner PA. metaFlye: scalable long -read metagenome assembly using repeat graphs. Nat 719 \nMethods. 2020 Nov;17(11):1103-10. 720 \n56. Li H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 2018 Sep 15;34(18):3094 -721 \n100. 722 \n57. Vaser R, Sović I, Nagarajan N, Šikić M. Fast and accurate de novo genome assembly from long 723 \nuncorrected reads. Genome Res. 2017 May;27(5):737-46. 724 \n58. Wood DE, Lu J, Langmead B. Improved metagenomic analysis with Kraken 2. Genome Biol. 2019 Nov 725 \n28;20(1):257.  726 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n20 \nSupporting information 727 \n 728 \nTable S1. Performance Metrics of all Deep Neural Network Architectures tested for Viabiliy Inferences. 729 \nTest set performance of residual neural network (ResNet) and transformer architectures, trained on various data 730 \nmodalities (Nanopore “Signal” or DNA sequence aka “Nucleotide”) and various signal chunk sizes (LR=1e-4).  731 \n 732 \nModel \nArchitecture \nData \nModality \nLength \n[Signal/Base] \nAccuracy F1 Precision Sensitivity Specificity AUROC AUPR \nResNet1 Signal 10K 0.83 0.84 0.78 0.91 0.74 0.90 0.87 \nResNet2 Signal 10K 0.83 0.85 0.77 0.94 0.72 0.89 0.86 \nResNet3 Signal 10K 0.81 0.82 0.76 0.89 0.72 0.87 0.83 \nTransformer Signal 10K 0.79 0.82 0.73 0.92 0.67 0.86 0.82 \nResNet1 Nucleotide \n(A,C,G,T) \n800 0.51 0.51 0.51 0.52 0.50 0.52 0.53 \nResNet1 Nucleotide \n(A,C,G,T,M) \n800 0.51 0.50 0.51 0.50 0.52 0.51 0.51 \nResNet1 Signal 1K 0.58 0.65 0.55 0.80 0.35 0.61 0.58 \nResNet1 Signal 5K 0.71 0.76 0.66 0.89 0.54 0.78 0.73 \nResNet1 Signal 7K 0.77 0.79 0.72 0.88 0.65 0.83 0.79 \nResNet1 Signal 12K 0.84 0.85 0.77 0.97 0.71 0.91 0.88 \nResNet1 Signal 20K 0.89 0.90 0.85 0.95 0.83 0.94 0.91 \n 733 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n21 \n 734 \nFig S1. Training and Validation Loss across Deep Neural Network Architectures tested for Nanopore-Signal 735 \nbased Viability Inference, and across different Learning Rates. 736 \n(A-C) Model loss of ResNet1 at LRs of 1e-3, 1e-4, and 1e-5; (D-F) Model loss of ResNet2 at LRs of 1e-3, 1e-4, and 737 \n1e-5; (G-I) Model loss of ResNet3 at LRs of 1e -3, 1e-4, and 1e-5; and (J-L) Model loss of the transformer models 738 \nat LRs of 1e-3, 1e-4, and 1e-5. The solid blue line indicates the training loss, the solid red line indicates the validation 739 \nloss, and the dashed line indicates the minimum validation loss from the final ResNet1, LR=1e-4, model. 740 \n 741 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n22 \n 742 \nFig S2. Training and Validation Loss of ResNet1 at various signal chunk sizes. 743 \nThe signal chunk size varies from (A) 1k, (B) 5k, (C) 7k, (D) 10k, to (E) 12k and (F) 20k. The solid blue line indicates 744 \nthe training loss, the solid red line indicates the validation loss, and the dashed line indicates the minimum validation 745 \nloss from the final (D) ResNet1model using a signal chunk size of 10k. 746 \n 747 \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint \n\n \n23 \n \nFig S3. Exemplary drops in ResNet1 prediction probabilities in nanopore signal chunks after consecutive masking of the signal region with the respectively highest \nCAM value. Figure headers: signal chunk ID / total number of masked signal values / prediction probability per signal chunk “prob”. Left to right: Five exemplary nanopore signal \nchunks (length of 10k signals). Top to bottom : Consecutive masking of 200 signal values per masking event (Materials and Methods). Legends: Red-colored CAM value \nvizualisations.  \n \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}