Abstract
18
The ability to differentiate between viable and dead microorganisms in metagenomic samples 19
is crucial for various microbial inferences, ranging from assessing ecosystem functions of 20
environmental microbiomes to inferring the virulence of potential pathogens. While established 21
viability-resolved metagenomic approaches are labor-intensive as well as biased and lacking 22
in sensitivity, we here introduce a new fully computational framework that leverages nanopore 23
sequencing technology to assess microbial viability directly from freely available nanopore 24
signal data. Our approach utilizes deep neural networks to learn features from such raw 25
nanopore signal data that can distinguish DNA from viable and dead microorganisms in a 26
controlled experimental setting. The application of explainable AI tools then allows us to 27
robustly pinpoint the signal patterns in the nanopore raw data that allow the model to make 28
viability predictions at high accuracy. Using the model predictions as well as efficient 29
explainable AI -based rules, we show that our framework can be leveraged in a real -world 30
application to estimate the viability of pathogenic Chlamydia, where traditional culture-based 31
Methods
suffer from inherently high false negative rates. This application shows that our 32
viability model captures predictive patterns in the nanopore signal that can in principle be 33
utilized to predict viability across taxonomic boundaries and indendent of the killing method 34
used to induce bacterial cell death. While the generalizability of our computational framework 35
needs to be assessed in more detail, we here demonstrate for the first time the potential of 36
analyzing freely available nanopore signal data to infer the viability of microorganisms, with 37
many applications in environmental, veterinary, and clinical settings. 38
39
Author summary 40
Metagenomics investigates the entirety of DNA isolated from an environment or a sample to 41
holistically understand microbial diversity in terms of known and newly discovered 42
microorganisms and their ecosystem functions. Unlike traditional culturing of microorganisms, 43
metagenomics is not able to differentiate between viable and dead microorganisms since DNA 44
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
2
might readily persist under different environmental circumstances. The viability of 45
microorganisms is, however, of importance when making inferences about a microorganism’s 46
metabolic potential, a pathogen’s virulence, or an entire microbiome’s impact on its 47
environment. As existing viability -resolved metagenomic approaches are labor -intensive, 48
expensive, and lack sensitivity, we here investigate our hypothesis if freely available nanopore 49
sequencing signal data, which captures DNA molecule information beyond the DNA sequence, 50
might be leveraged to infer such viability. This hypothesis assumes that DNA from dead 51
microorganisms accumulates certain damage signatures that reflect microbial viability and can 52
be read from nanopore signal data using fully computational frameworks. We here show first 53
evidence that such a computational framework might be feasible by training a deep model on 54
controlled experimental data to predict viability at high accuracy, exploring what the model has 55
learned, and applying it to an independent real-world dataset of an infectious pathogen. While 56
the generalizability of this computational framework needs to be assessed in much more detail, 57
we demonstrate that freely available data might be usable for relevant viability inferences in 58
environmental, veterinary, and clinical settings. 59
60
Introduction
61
While microbial cultivation remains a foundational technique in microbiology to assess the 62
taxonomic composition of microbial communities and to properly understand their physiology 63
and ecosystem functions [ 1], only a small fraction of microbial diversity has been isolated in 64
pure culture [2]. This limitation has led to undiscovered functions and biased representations 65
of the phylogenetic diversity of microbial communities in nearly all of Earth’s environments [2]. 66
While medically relevant microorganisms of the human microbiome constitute an exemption 67
since they have been disproportionately well studied through microbial cultures [3], the clinical 68
application of microbial cultivation for pathogen profiling is also limited due to its time -69
consuming and labor-intensive nature [4]. 70
The first studies of the so -called “microbial dark matter” have been enabled by advances in 71
culture-independent molecular methodology [ 5], and have been based on amplifications of 72
conserved marker regions such as ribosomal RNA genes [ 6]. Such targeted metabarcoding 73
approaches, however, suffer from several limitations: They can often not provide strain - or 74
even species -level taxonomic resolution, are highly dependent on genomic database 75
completeness, do not allow for any functional inferences or virulence annotations, and often 76
introduce amplification bias due to differential amplification efficiency and primer mismatches, 77
which can significantly distort the representation of microbial community compositions [7]. 78
Metagenomics, on the other hand, is a shotgun sequencing -based molecular methodology 79
that can assess the entirety of DNA isolated from an environment or a sample and de novo 80
assemblies of potentially complete microbial genomes of all present microorganisms; such 81
genome-based approaches provide a variety of phylogenetically informative sequences for 82
taxonomic classification, information about the metabolic and virulence potential of 83
microorganisms, and the potential to identify completely novel genes [8, 9]. 84
Especially long-read metagenomic approaches have shown great promise in achieving highly 85
contiguous de novo assemblies through the recovery of high-quality metagenome-assembled 86
genomes (MAGs) from complex environments; specifically, the latest advances in nanopore 87
sequencing technologies have resulted in high sequencing accuracies of very long sequencing 88
reads of up to millions of bases, which allowed for the generation of hundreds of MAGs from 89
metagenomic data, including the generation of closed circularized genomes [ 10, 11]. 90
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
3
Nanopore sequencing technology is based on the interpretation of the disruption of an ionic 91
current due to a motor protein guiding individual nucleotide strands through nanopores 92
embedded in an electrically resistant polymer membrane at a consistent translocation speed 93
[12]. This raw nanopore signal, or “squiggle” data, can then be translated into nucleotide 94
sequence using bespoke neural network -based basecalling algorithms [ 13], which – when 95
efficiently embedded on powerful GPUs – can generate genomic data in real -time. The 96
portable character and straightforward implementation of nanopore sequencing at very upfront 97
investment costs further make this technology accessible for fast microbial and pathogens 98
assessments at point of interest all around the world, including in low - and middle-income 99
countries [14]. 100
In contrast to cultivation -based approaches, molecular methods suffer from their inherent 101
deficiency of not being able to differentiate between viable and dead microorganisms [ 2, 15]. 102
While cultivation -based approaches only detect viable microorganisms, DNA might remain 103
intact and therefore accessible by molecular methods despite the respective microorganisms 104
being dead [ 15]. This might be especially relevant in the context of infection prevention and 105
control and pathogen monitoring in the clinical setting, where certain disinfection methods or 106
the use of systemic antibiotics often kill the bacteria before the DNA is destroyed [15, 16], but 107
also for understanding the ecosystem functions of thus far understudied microbiomes [2]: For 108
example, the air microbiome has been shown to be remarkable diverse and variable when 109
assessed through nanopore metagenomics [17], but given the low biomass of this environment 110
it is expected that many microorganisms might be dead and stem from adjacent environments 111
such as soil or water. The persistence of the DNA of dead microorganisms in the environment 112
might hereby depend on many factors, including external conditions such as temperature, pH, 113
and microbial activity, and internal, taxon -specific parameters such as microbial cell wall 114
composition. Viability-resolved metagenomics would, however, be crucial for the interpretation 115
of metagenomic data, ranging from outbreak source detection [18], food safety [19] and public 116
health investigations [20], to ecosystem function inferences [21]. 117
To assess microbial viability from genomic data, several approaches have been developed: 118
Culture-dependent viability methods combine the advantages of cultivation -based and 119
molecular approaches by growing certain microorganisms of interest on selective media; this 120
approach, however, remains time -consuming and labor-intensive and suffers from the same 121
selectivity of growth media and culturable microorganisms as purely cultivation -based 122
approaches [ 22], especially for fastidious or obligate intracellular microorganisms [ 23]. 123
Microbial viability can further be assessed through metabolic activity, where microbial cells are 124
incubated with specific substrates whose metabolization produces detectable signals such as 125
ATP production, tetrazolium salt reduction, and radiolabeled substrate incorporation; this 126
approach may, however, not be applicable to all microbial taxa and environmental conditions, 127
and the results might be heavily influenced by the physiological state of the microbial cells [24]. 128
While RNA has been used as a viable/dead marker due to its intrinsic instability [ 25, 26, 27], 129
the metatranscriptome has to be stable enough to be detectable in certain environments, but 130
unstable enough to reflect viability; if only one gene is targeted, the gene analyzed further has 131
to be continuously expressed [ 15]. Further issues might arise from the relatively challenging 132
extraction protocols due to the RNA’s instability, and from the relatively evolutionary 133
conservation of gene sequences, which might hamper taxonomic resolution. 134
Finally, an aspect that can be used for viability -resolved metagenomics is the physical 135
difference between viable and dead cells: Viability PCR (vPCR) uses DNA -intercalating dyes 136
such as ethidium monoazide (EMA) or propidium monoazide (PMA) to differentiate between 137
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
4
viable and dead cells. These dyes penetrate only dead cells with compromised membranes 138
and bind to their DNA via covalent bonds upon photoactivation, preventing it from being 139
amplified during subsequent PCR [15, 16]; this approach has been applied to a diverse array 140
of Gram-negative and -positive bacteria, and to assess the effectiveness of disinfection and 141
heat treatment [24]. It, however, relies on the assumption that membrane integrity is a reliable 142
indicator of viability, which can lead to overestimation of viability if cells lose viability without 143
immediate membrane compromise [28] and can be biased by the dye’s variable permeability 144
across different microbial cell wall structures [ 29, 30]. The dependence of the approach on 145
photoactivation further means that turbid material might hamper the efficiency of the dye [31]. 146
All these established viability -resolved metagenomic approaches are labor -intensive, require 147
additional reagents and sample processing, and are often biased and lack sensitivity. We here 148
hypothesized that the raw, freely available nanopore signal from metagenomic datasets might 149
be leveraged to infer microbial viability, assuming that the native DNA from dead 150
microorganisms accumulates detectable squiggle signatures due to, e.g., external damage, 151
the lack of DNA repair mechanisms, or the enzymatic activity [ 32, 33, 34]. Such an analysis 152
framework would be fully computational and utilize squiggle data that is automatically obtained 153
with nanopore sequencing. While raw nanopore data is known to contain information about 154
epigenetic modifications [35, 36, 37] and oxidative stress at specific human telomere sites [38], 155
the applicability to assess microbial viability has not yet been tested. 156
In this study, we produced experimental nanopore sequencing data from viable and dead 157
Escherichia coli cultures to optimize deep neural networks to predict viability just from the 158
nanopore squiggle signal. We then applied explainable AI (XAI) tools, which allow us to identify 159
the specific nanopore signal patterns in the input data that allow the model to deliver high -160
accuracy predictions as an output. We finally showed that our computational framework can 161
be leveraged in a real -world application to estimate the viability of pathogenic Chlamydia 162
abortus, an obligate intracellular bacterial species with a complex biphasic lifecycle causing 163
enzootic abortion in sheep and goats [ 39]. Together, we show that our framework can in 164
principle predict viability across taxonomic boundaries and independent of the killing method 165
used to induce bacterial cell death. While we have not yet tested the generalizability of our 166
computational framework, we here demonstrate for the first time the potential of analyzing 167
freely available nanopore signal data to infer the viability of microorganisms, with many 168
applications in environmental, veterinary, and clinical settings. 169
170
Results
& Discussion 171
Neural Network Training 172
We generated controlled training data by nanopore sequencing native DNA of viable and dead 173
E. coli (Materials and Methods). We killed E. coli cultures using different stressors to then 174
isolate the extracellular DNA and expose it to natural degradation. We only obtained enough 175
DNA for subsequent shotgun sequencing from the viable culture and from the culture killed 176
through rapid UV exposure (viable: 212 ng/μL; UV: 5.46 ng/μL; heat shock: 0.03 ng/μL; bead 177
beating: 0.67 ng/μL; Materials and Methods). We repeated this experiment and confirmed that 178
rapid heat shock as well as bead beating exposure again resulted in very low DNA 179
concentrations, suggesting quick and complete DNA degradation. We hypothesize that UV 180
exposure constitutes the only stressor that simultaneously destroys bacterial cell walls as well 181
as inactivates DNA-degrading enzymes. In the case of killing by heat shock and bead beating, 182
such enzymes might have remained active and would have been carried forward by our 183
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
5
extracellular DNA isolation approaches (Materials and Methods) to then degrade all genomic 184
Material
during natural exposure. We therefore created nanopore shotgun sequencing of the 185
viable and the UV-exposed culture, which resulted in 2.92 Gbases (Gb; median read length of 186
2,476 b) and 2.69 Gb (median read length of 1,606 b) of sequencing output, respectively 187
(Materials and Methods). 188
We then tested the implementation of different neural network architectures to predict the 189
binary viability state from the raw nanopore data (0=viable; 1=dead; Materials and Methods). 190
We processed the E. coli nanopore signal, or “squiggle”, data, cut it into altogether 3,181,600 191
signal chunks of 10k signals, and separated the chunks into balanced training (60%), validation 192
(20%), and test (20%) set along each original sequencing read to avoid that signal chunks 193
from the same read would end up in the same dataset (Materials and Methods). These signal 194
chunks can be treated as 1D time series signal data of consistent length. We trained the 195
different model architectures using different learning rates (LRs) up to 1,000 epochs, 196
assessing the models’ performance based on training and validation loss after each epoch 197
(Materials and Methods; Table S1; Fig S1). The loss plot of our best -performing model, a 198
residual neural network with convolutional input layers (configuration ResNet1; LR=1e -4; 199
Table S1 ; Fig S1 ; Materials and Methods) shows minimal overfitting when the minimum 200
validation loss is reached at epoch 667 ( Fig 1A ). The other residual neural network 201
architectures (ResNet2, ResNet3), on the other hand, resulted in overfitting to the training data 202
at any LR, and the transformer architecture did not reach the minimum validation loss of 203
ResNet1 (Fig S1). We next only focused on ResNet1 and optimized its probability threshold 204
using the validation set; in order to obtain a high accuracy and F1 score, we maintained the 205
probability threshold at the default value of 0.5 ( Fig 1B ), which resulted in a good final 206
performance on the test data with an accuracy of 0.83 and a F1 score of 0.81 ( Fig 1B, inlet) 207
as well as Area Under the Curve (AUC) values of 0.90 (Area Under the Receiver Operating 208
Characteristic curve; AUROC) and 0.92 (Area Under the Precision -Recall curve; AUPR), 209
respectively (Fig 1C). 210
211
212
Fig 1. Training of the Viability Residual Neural Network (ResNet) 213
(A) Model loss for validation and test sets across 1,000 epochs; the minimum validation loss of the ResNet was 214
reached at epoch 677. (B) Prediction probability threshold optimization on the validation dataset resulted in a 215
probability threshold of 0.5 for obtaining maximum accuracy and F1 score. Inlet: Performance of the final ResNet 216
model ResNet1 on the test data (Materials and Methods). (C) Test performance of the ResNet1 in terms of 217
Precision-Recall (PR; pink) and Receiver Operating Characteristic (ROC; green) curves and their respective Areas 218
Under the Curve (AUPR, AUROC). 219
220
We also trained the same residual neural network architecture ResNet1 on the basecalled 221
nanopore data of viable and dead E. coli at a standardized chunk size of 800 b, which roughly 222
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
6
corresponds to a signal chunk size of 10k signals (Materials and Methods). Independently of 223
if we only basecalled the canonical bases or used a N6 -methyladenine (6mA) modification -224
aware basecalling model (Materials and Methods), the model could not be trained to 225
distinguish viable from dead data just from DNA sequence data ( Table S1). This shows that 226
our model captures patterns in the squiggle data that goes beyond the encoding of nucleotides 227
and their known epigenetic modifications. While this was expected since we used the same E. 228
coli culture with the same reference genome to create the viable and dead datasets, we hereby 229
ruled out that our squiggle -based model captured any random differences in DNA sequence 230
context between the two datasets that might have occurred by chance. 231
232
We additionally obtained the performance of ResNet1 for different signal chunk sizes (Fig S2; 233
Table S1), and we found that viability prediction performance was possible from a minimum 234
chunk size of approximately 5k, but can be further improved with increasing chunk size. This 235
shows that larger signal chunks contain more information that can be used by our model to 236
make accurate per-chunk predictions while the resulting reduced size of the dataset, especially 237
of the training dataset, did not influence the learning ability of the model. While this might be 238
valuable information for future applications where potentially the full length of a read could be 239
leveraged, we here decided to stick to a signal chunk size of 10k signals, which had already 240
resulted in good performance ( Fig 1; Table S1) and which can be applied to relatively small 241
sequencing reads as is often the case for metagenomic datasets (e.g., [17]). 242
243
Explainable AI 244
We implemented Class Activation Maps (CAM) as an XAI method [ 40] to identify the most 245
important regions in the nanopore signal data that inform the model’s viability classifications 246
(Materials and Methods; Fig 2A ). We found that “dead” signal chunks exhibited discrete 247
regions of increased CAM values (“CAM regions” defined at CAM values>0.8; Fig 2B for 248
several true positive classifications of the test dataset). To confirm the importance of these 249
CAM regions for the model’s final predictions, we applied consecutive masking of the regions 250
with the highest CAM values within each nanopore signal chunk (Fig S3 for several examples); 251
we observed that the prediction probability for being classified as “dead” decreased with 252
increased masking of CAM-relevant regions, either by consecutively masking regions using a 253
consistent mask size or by increasing the mask size (from 100 to 2k signals; Fig 2C). This 254
shows that the CAM application reliably pinpoints patterns in the nanopore signal that are 255
predictive for our viability model. 256
We used the CAM regions to manually investigate the squiggle signals, and found that many 257
CAM regions of “dead” signal chunks included a sudden substantial drop in the nanopore 258
signal. We therefore developed a simple algorithm that identifies such sudden drops, and 259
applied this XAI rule to our test dataset (Materials and Methods; Fig 2D). We here classified 260
any signal chunk with at least one sudden drop as “dead”, and all others as “viable”. While this 261
simplified algorithm led to a drop in overall performance, we could still reach a relatively good 262
overall accuracy of 0.68 (in comparison to 0.83 of the full model; Fig 2E). While the XAI rule 263
maintained performance of specificity and precision, we observed a substantial drop in recall 264
in comparison to the full model (now 0.39 instead of 0.74). This shows that while the absence 265
of a sudden drop in the nanopore signal data seems to reliably predict viability, not all “dead” 266
signal chunks seem to contain such a sudden drop. While this sudden -drop detection still 267
seems to be at the core of our model’s interpretability (when focusing on high-confidence true 268
positive chunks at p>0.99, the recall increased to 0.68), the model seems to additionally detect 269
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
7
more subtle patterns in the nanopore signal data which allow it to increase recall while 270
maintaining specificity and precision. 271
Based on our previous experience with squiggle data analysis [ 41, 42], we hypothesize that 272
the substantial sudden drops in nanopore signal might be caused by a twist or kink in the DNA 273
backbone, for example from 6-4pp pyrimidine dimers. The drop would then mark the event of 274
a pore getting blocked due to such damage. Such a twist could also lead to a stalling signal if 275
it impairs the motor protein from processing the DNA strand, which we indeed partially 276
observed in our data (e.g., top signal chunk of Fig 2B/D). While more detailed future squiggle 277
analyses as well as application of our model to other viability studies will hopefully shed more 278
light on the biological, chemical, and physical features detected by our model, we conclude 279
that in our current study both, UV exposure ( E. coli) and heat shock (Chlamydia) might have 280
caused such twists in the DNA backbone. 281
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
8
282
Fig 2. XAI for Interpretability of the Viability ResNet. 283
(A) The Class Activation Maps (CAMs) leverage the global average pooling (GAP) layer right before the fully 284
connected (FC) layers of the residual neural network to map model interpretability onto the input features; they are 285
generated by aggregating the final convolutional layer's feature maps through a weighted sum, highlighting 286
nanopore signal regions that allow the neural network to make accurate predictions. (B) Exemplary nanopore signal 287
chunks that were classified as “dead” at a prediction probability of p>0.99, and their CAM values. Higher CAM 288
values indicate stronger feature map activations. (C) Impact of consecutive masking (n=5) of the signal region with 289
the highest CAM value per signal chunk (x -axis) on the model’s prediction probability (y -axis); five different mask 290
sizes (from 100 to 2k signals) were used. (D) Application of a simplified XAI rule that classifies signal chunks 291
according to the presence of a “sudden drop” (green: mean signal per chunk; red: threshold for sudden drop 292
definition; yellow: identification of sudden drops in the exemplary signal chunks). (E) Comparison of the 293
performance of the full model (ResNet) with the simplified XAI rule based on the entire test dataset. 294
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
9
Application to pathogenic Chlamydia 295
We next applied our computational viability framework to distinguish viable from dead 296
Chlamydia abortus, one of the most common causes of infectious abortion in small ruminants 297
worldwide [43, 44], and a zoonotic pathogen causing pneumonia and miscarriage in humans 298
[45]. C. abortus are obligate intracellular bacterial species with a complex biphasic life cycle 299
that form intracellular vacuoles termed inclusions. These unique properties render both 300
cultivation- and vPCR-based approaches for viability estimations complicated [46]. In order to 301
apply our computational viability framework to this pathogen, we decided to use two differently 302
treated samples of the same strain of C. abortus samples for which both cultivation- and vPCR-303
based approaches had predicted viability with high certainty. We chose a “viable” and a heat-304
treated (“dead”) sample of the same culture for which we could ascertain viability and non -305
viability, respectively (Materials and Methods); heat treatment was used since it constitutes 306
the standard approach for killing C. abortus [46]. Briefly, for ascertaining viability and non -307
viability of the two samples, respectively, we applied cultivation and propidium monoazide 308
(PMA)-based vPCR. In the case of cultivation, heat -treated C. abortus were unable to form 309
viable inclusions in cell culture, whereas the untreated sample showed high infectivity with 310
2.6e6 inclusion forming units per mL (IFU/ml). In the case of vPCR, both samples were treated 311
with and without PMA enhancer with or without PMA [ 46] to then quantify the respective 312
amounts of chlamydial DNA with a sensitive Chlamydia-specific qPCR ( 47; Materials and 313
Methods). The difference in quantity between PMA-treated and untreated DNA was expressed 314
as Δlog10 Chlamydia per mL, resulting in 0.42 and 4.2 Δlog10 Chlamydia per mL for the control 315
and the heat-treated sample, respectively (Table 1). These data are comparable to a previous 316
study in which fresh C. trachomatis culture was heat -killed and a viability ratio determined 317
ranging from 0% to 100% resulting in a 3.01 and 0.37 Δlog10 Chlamydia per mL for 0% and 318
100% viable, respectively [ 46]. Both cultivation and vPCR therefore confirmed that heat -319
treatment had completely inactivated previously viable C. abortus for subsequent culture and 320
had strongly reduced the amount of “viable” DNA using vPCR. 321
322
Table 1. Viability PCR (vPCR) results of a “viable” and a “dead” C. abortus sample. 323
A viable and dead (heat-treated at 95°C for 10 min) sample of the same Chlamydia (C. abortus S26/3) stock were 324
tested for viability using vPCR: PMA-untreated vPCR reflects total Chlamydia content; PMA-treated vPCR reflects 325
viable Chlamydia content; IFU describes the number of Inclusion Forming Units. 326
327
Condition Titer [IFU/mL] PMA-untreated
vPCR [copy number
per mL]
PMA-treated vPCR
[copy number per
mL]
Viability ratio [%]
viable 2.67e7
4.59e7
1.73e7
37.6
dead 0 9.03e7 5.52e7
0.01
328
We next created nanopore shotgun sequencing of these “viable” and “dead” C. abortus 329
samples, which resulted in 37.50 Mb (median read length of 1,458 b) and 10.38 Mb (median 330
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
10
read length of 496 b) of sequencing output, respectively (Materials and Methods). We 331
subsequently confirmed that the sequencing data was indeed created from the low -332
concentration C. abortus samples, and not from any contaminating microorganisms: The de 333
novo assembly resulted in 18 and 39 contigs, respectively, which all mapped to C. abortus 334
(NCBI nt taxonomy id; 83555; Materials and Methods). 335
We processed this nanopore sequencing data into 42,335 and 13,312 nanopore signal chunks 336
of 10k signal length for the viable and dead sample, respectively. We used our model 337
(ResNet1) to make computational viability predictions on this dataset, which resulted in an 338
accuracy of 0.85 and a F1 score of 0.68 at the previously optimized prediction probability 339
threshold of 0.5. These performance values for the C. abortus application are therefore 340
comparable to the performance on the original test E. coli data, while the microorganism, killing 341
method, and DNA extractions methods were substantially different ( Fig 3A; top row). While 342
our model’s prediction probability distributions across signal chunks (Fig 3A; bottom row) were 343
comparable between viable E. coli and C. abortus, dead C. abortus resulted in a more bimodal 344
distribution than dead E. coli, with several “dead” chunks receiving close -to-zero prediction 345
probabilities, suggesting viability (Fig 3A; bottom row; right column). The probability threshold-346
independent application of our model to C. abortus consequently resulted in decreased AUC 347
values (AUROC of 0.80 and AUPR of 0.64; Fig 3B) in comparison to its application to E. coli 348
(AUROC of 0.90 and AUPR of 0.92; Fig 1C). While this difference in performance might just 349
mean that ResNet1 is less sensitive at detecting “dead” signal chunks in the C. abortus data 350
than in the E. coli data that it was trained on, it might also reflect true variability: The C. abortus 351
sample that was nanopore -sequenced to create “dead” nanopore signal data had strongly 352
reduced amounts of “viable” DNA according to the vPCR results but might have still contained 353
DNA from viable C. abortus bacterial cells – different from our dead E. coli sample which was 354
stringently filtered to only contain extracellular DNA and subsequently subjected to natural 355
degradation (Materials and Methods). The latter hypothesis might be supported by the 356
observation that the application of our CAM-based XAI to the “dead” C. abortus signal chunks 357
detected similar patterns as in the “dead” E. coli signal chunks (Fig 3C). An application of our 358
XAI rule (Materials and Methods) further resulted in similar performance as its application to 359
the E. coli signal chunks with good accuracy (0.78), high specificity (0.91), and low recall 360
(0.36). 361
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
11
362
Fig 3. Application of the Viability ResNet to the infectious pathogen Chlamydia. 363
(A) Model performance comparisons between E. coli (left) and C. abortus (right) datasets at the optimized prediction 364
probability threshold of 0.5. Top row : Binary model predictions for viable and dead E. coli and C. abortus , 365
respectively; bottom row : Normalized distribution of model prediction probabilities across all signal chunks, 366
respectively. For E. coli, all test signal chunks are visualized (388k “viable” and “dead” chunks, respectively), and 367
for C. abortus, all signal chunks are visualized (42,335 “viable” and 13,312 “dead” chunks). (B) Performance of the 368
model on the C. abortus dataset in terms of Precision -Recall (PR; pink) and Receiver Operating Characteristic 369
(ROC; green) curves and their respective Areas Under the Curve (AUPR, AUROC). (C) Exemplary Chlamydia 370
nanopore signal chunks that were correctly classified as “dead”, and their CAM values. Higher CAM values indicate 371
stronger feature map activations. 372
373
AI- and nanopore-empowered viability-resolved metagenomics 374
While metagenomic approaches provide the unique opportunity of generating de novo 375
assemblies and potentially complete microbial genomes to explore the “microbial dark matter” 376
as well as to infer potential functions such as metabolic and virulence potential [ 8, 9], they 377
have suffered from their inability to differentiate between viable and dead microorganism [ 2, 378
15]. Such viability inferences can, however, distort any microbial inference, ranging from 379
assessing ecosystem functions of environmental microbiomes to inferring the virulence of 380
potential pathogens. As established viability -resolved metagenomic approaches are labor -381
intensive as well as biased and lack sensitivity (e.g., 22), we here show first evidence that a 382
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
12
fully computational framework based on residual neural networks with convolutional data 383
processing layers can leverage raw nanopore signal data, also known as squiggle data, to 384
make accurate inferences about microbial viability (test set accuracy of 0.83). 385
Our subsequent XAI analyses point to the potential role of DNA backbone damage for 386
achieving accurate model predictions; however, more work is needed to fully understand the 387
biological, physical, or chemical underpinnings of our viability model predictions. First, the 388
simplified XAI rule is not sufficient to correctly classify the majority of “dead” signal chunks 389
(recall of 0.39), which means that the residual neural network has apparently picked up on 390
additional, more ambiguous signal patterns that allow the full model to make more sensitive 391
predictions (recall of 0.74). Besides exploring the underlying rules of such additional signal 392
patterns, we will also apply the viability model to more datasets to assess its generalizability. 393
Especially the probing of different taxonomic groups and killing methods should help us tease 394
apart the origins of our current XAI results. Our study, however, already provides first evidence 395
that our viability model captures predictive patterns in the nanopore signal that can in principle 396
be utilized to predict viability across taxonomic boundaries and independent of the killing 397
method. The application to estimate the viability of pathogenic Chlamydia (prediction accuracy 398
of 0.85) is hereby of potentially immediate interest to veterinary scientists since traditional 399
Methods
for assessing the pathogen’s viability have been labor -intensive and suffered from 400
inherently high false negative rates. 401
While the generalizability of our model needs to be assessed in much more detail, including 402
for other microbial taxa such as spore -forming bacteria or fungi and for real -world 403
metagenomics applications, the potential benefits of such a widely applicable computational 404
framework could be immense, with many potential applications in environmental, veterinary, 405
and clinical settings. As is the case for epigenetic inferences [35, 36, 37], the viability inference-406
enabling squiggle data is a complementary output of any nanopore sequencing experiment of 407
native DNA, that is then usually basecalled and archived for future re -basecalling after 408
potential basecalling model improvements. This means that any future nanopore -based 409
metagenomic study could make viability predictions for free without additional costs and 410
laboratory work, and that any existing archived nanopore data could be assessed in terms of 411
its microorganisms’ viability – which would allow us to quantify the impact of dead 412
microorganisms on metagenomics in general, and to further explore factors such as species - 413
and environment -specificity. We finally anticipate that quantitative AI modeling has the 414
potential to inform more differentiated viability assessments, which might help quantify or even 415
time degradation events and decipher the impact of dormancy on metagenomic studies [ 48, 416
49]. 417
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
13
Materials and methods
418
Training Data Generation 419
We cultured E. coli K12 in 200 mL Luria -Bertani (LB) medium for 24 hours at 37°C to reach 420
the log phase of the growth curve. The culture was then used to inoculate four 200 mL LB 421
media in 1L Erlenmeyer flasks, which were again incubated for 24 hours to reach the growth 422
log phase. One of the media was used as viable control, i.e. DNA was extracted from 750 uL 423
of the living culture using the spin-column based QIAGEN PowerSoil Pro Kit (QIAGEN, 2018, 424
Hilden, Germany), following the manufacturers’ instructions. The remaining three cultures 425
were killed by one of the following stressors: UV irradiation at 254 nm for 15 min, heat shock 426
at 120°C for 5 min, or bead beating for 30 min. To then separate extracellular DNA from cell 427
debris and intact bacterial cells, we centrifuged the media for 10 min at 4,000 x g and filtered 428
the supernatant through 0.2 µm filters. The resulting extracellular DNA was subsequently kept 429
at room temperature for 5 days to simulate the natural accumulation of DNA degradation. DNA 430
from dead bacteria was extracted from these samples using the same extraction approach 431
following the QIAGEN PowerSoil Pro Kit protocol, but the first lysis buffer step was omitted 432
since cell lysis had already happened. 433
We then used Oxford Nanopore Technologies’ Rapid Barcoding library preparation kit 434
(RBK114-24 V14), R10.4.1 MinION flow cells, and MinKNOW software v23.04.5 for shotgun 435
nanopore sequencing of the “viable” and “dead” DNA. We used four barcodes for each sample, 436
resulting in DNA input of 800 ng and 218 ng for the preparation of the “viable” and “dead” 437
library, respectively. We ran each library for 24 h, using two different flow cells to avoid any 438
cross contamination, and filtered the resulting nanopore data at a minimum read length of 20 439
b. Raw nanopore data was created using the standard translocation speed of 400 b/s, and a 440
sampling frequency of 5 kHz. 441
We applied Dorado v4.2.0 ( https://github.com/nanoporetech/dorado) SUP -basecalling 442
(
[email protected]) and 6mA -aware SUP-basecalling (6mA@v1) to all 443
nanopore reads that had passed internal data quality thresholds to obtain E. coli DNA 444
sequence data. We subsequently removed sequencing adapters and barcodes using 445
Porechop v0.2.3 (https://github.com/rrwick/Porechop). 446
447
Neural Network Architecture and Training 448
We tested the implementation of different residual neural networks and transformer 449
architectures to predict the binary viability state from the raw nanopore data (0=viable; 450
1=dead). The first residual neural network, ResNet1, consists of four layers, each containing 451
two bottleneck blocks. Each bottleneck block consists of convolutional layers, batch 452
normalization, and a rectified linear unit (ReLu) activation function. Each of the four layers 453
consists of an increasing number of convolutional channels: 20, 30, 45, and 65, respectively, 454
followed by global average pooling and a fully connected layer, resulting in 66,916 parameters. 455
The model then uses a softmax function to convert logits, the raw outputs from the fully 456
connected layer, into predicted probabilities ranging from 0 to 1. We evaluated the training of 457
the model using the Adam optimizer for mini-batch gradient descent at three different LRs (1e-458
3, 1e-4, and 1e-5), training the model up to 1,000 epochs and at a batch size of 1,000 signal 459
chunks. We initialized the model using Kaiming initialization. For ResNet2, we increased the 460
number of convolutional channels to 40, 60, 90, and 135, respectively, resulting in 1,828,777 461
parameters. For ResNet3, we increased the number of convolutional channels to 512, 30, 45, 462
and 67, respectively, resulting in 2,479,140 parameters. The transformer model was based on 463
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
14
a positional encoding, convolutional layer with a channel number of 24 and one block of one 464
attention head, resulting 219,586 parameters. 465
We processed the E. coli squiggle data by excluding the first 1,500 signal points (potential 466
noise, adapter sequences, or barcodes), then cutting it into signal chunks of 10k signals, and 467
separated the chunks into balanced training (60%), validation (20%), and test (20%) set along 468
each original sequencing read to avoid that signal chunks from the same read would end up 469
in the same dataset. We pooled the viable and dead signals chunks to obtain exactly balanced 470
training, validation, and test sets. For normalizing each chunk, we subtracted the median per 471
chunk and divided it by the median absolute deviation (MAD) to make the signal data robust 472
to outliers. We then scaled the signal by the MAD scaling factor 1.4826, and replaced outliers 473
exceeding 3.5 times the scaled MAD by the mean of their two neighboring values. 474
We also trained ResNet1 on the basecalled nanopore data (with or without 6mA basecalling) 475
of viable and dead E. coli at a standardized chunk size of 800 b, which roughly corresponds 476
to a signal chunk size of 10k signals. For encoding, we used a one -hot encoding method to 477
turn DNA sequence into unique binary vectors. We then concatenated and saved these 478
encoded sequences as tensors for training and testing. We finally trained ResNet1 on signal 479
chunks of different signal lengths, ranging from 1k to 20k signals. 480
481
Explainable AI 482
We utilized CAMs to identify and visualize signals regions that influenced the model's decision-483
making. As feature maps from the final convolutional layer undergo a global average pooling 484
layer where each map is averaged and concatenated, we can calculate the weighted sum of 485
these feature maps using the weights of the fully connected layer and project it back onto the 486
preprocessed signal [ 40]. To do so, we implemented CAM in Python/PyTorch. During the 487
forward pass, we ensure that the feature maps from the last convolutional layer are captured. 488
To compute the CAM, we use the weights of the model's output layer for the class of interest, 489
multiplying these weights with the corresponding feature maps and then summing them up. 490
We convert the resulting CAM to an array and normalize its values to a range of 0 to 1. We 491
then overlay the CAM on the original input signal to identify the regions most influential in the 492
model's decision -making process. We additionally used the Remora API to match raw 493
nanopore data to the corresponding Dorado -basecalled bases, to then manually investigate 494
any obvious sequence abnormalities in the CAM regions. 495
For consecutive masking of the CAM regions with the highest CAM values, we used Python 496
to first obtain and normalize the CAM values of all true -positive signal chunks at p>0.5 497
(n=286,179), identify the maximum value, mask the signal region (i.e., setting to zero after 498
MAD-normlization) at a specified mask size (between 100 and 2k signals) centered around 499
the maximum -CAM signal, and obtain updated prediction probabilities. We repeated this 500
masking step five times and calculated the confidence interval at each masking step: We 501
obtained the mean and Standard Error of the Mean (SEM) of the newly calculated prediction 502
probabilities across all signal chunks to calculate the 95% confidence interval at mean ± 1.96 503
* SEM. We plotted the results using matplotlib. 504
We next used Python to develop an algorithm to obtain a simplified XAI rule to distinguish 505
“dead” from “viable” signals chunks based on our CAM results by identifying the presence of 506
at least one sudden drop in the chunk. To identify those sudden drops, we first calculated the 507
mean and standard deviation (SD) of each signal chunk, and found that a scaling factor of 3 508
identified most manually detected sudden drops at a vertical threshold of mean - 3 * SD. 509
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
15
Application to pathogenic Chlamydia 510
The Chlamydia abortus strain S26/3 associated with ovine abortion (provided by Dr. G.E. 511
Jones, Moredun Research Institute, Edinburgh, UK) was cultured as described by Borel et al. 512
[50]. Briefly, the strain was first grown in embryonated chicken eggs and stored at -80°C 513
following 1:2 dilution in sucrose-phosphate-glutamate buffer (SPG). For propagation, the strain 514
was cultured in HEp -2 cells (ATCC CCL -23). Infectious elementary bodies were separated 515
from cell debris as well as non -infectious chlamydial reticulate bodies using a renografin 516
density gradient [51], resuspended in SPG, and stored in aliquots at -80°C. To determine the 517
viability of this stock, one aliquot was thawed on ice and separated into two tubes of which one 518
was heat-treated for 10 min at 95°C. Both the control (“viable”) and the heat -treated (“dead”) 519
samples were then divided into subsamples, which were subsequently used for cultivation, 520
vPCR, and nanopore sequencing. 521
For viability determination by culture, 100 µl per sample was used to infect two glass coverslips 522
(13 mm in diameter, ThermoScientific, Waltham, MA, USA) in 24 -well plates (TPP Techno 523
Plastic Product AG, Trasadingen, Switzerland) seeded to confluence with LLC -MK2 cells (524
rhesus monkey kidney cell line; provided by IZSLER, Brescia, Italy) [52]. Following inoculation 525
of the monolayer, infection was enhanced by centrifugation for 1 h at 25°C (1000 x g). After 526
48 h of incubation at 37°C (5% CO ₂), cultures were fixed for 10 min in ice -cold methanol. 527
Coverslips were then processed using a well -established immunofluorescence assay [ 53]. 528
Briefly, DNA was stained with 1 μg/mL 4', 6-diamidino-2’-phenylindole dihydrochloride (DAPI, 529
Molecular Probes, Eugene, OR, USA). In parallel, inclusions were labeled with a 530
Chlamydiaceae-specific primary antibody ( Chlamydiaceae LPS, Clone ACI -P; Progen, 531
Germany), which was diluted 1:200 in blocking solution consisting of 1% bovine serum 532
albumin (BSA, St. Louis, MO, USA) in phosphate -buffered saline (PBS, GIBCO, Invitrogen, 533
Carlsbad, CA, USA). Inclusions were then visualized with Alexa Fluor 488 goat anti -mouse 534
(Molecular Probes) diluted 1:500 in blocking solution. As a final step, coverslips were washed 535
with PBS, mounted with FluoreGuard (Hard Set; ScyTek Laboratories Inc., Logan, UT, USA) 536
on glass slides, and inclusions determined using a Leica DMLB fluorescence microscope 537
(Leica Microsystems, Wetzlar, Germany) and a 10X ocular objective (Leica L-Plan 10x/ 25 M, 538
Leica Microsystems). In parallel, a three -fold dilution series of the sample was performed in 539
96-well plates (TPP) and processed as above. The number of IFU/ml was then determined 540
using the Nikon Ti Eclipse epifluorescence microscope (Nikon, Tokyo, Japan) at a 20X 541
magnification [52]. 542
For vPCR, two 100 µl subsamples were taken and mixed with 200 µl SPG and 100 µl PMA 543
enhancer for Gram Negative Bacteria (5X, Biotium, Fremont, CA, USA). PMAxx (Biotium) at a 544
final concentration of 50 µM was added (“PMA -treated”) or not (“untreated”) to the 545
subsamples. Samples were then exposed to a 650 -W light source using a PMA -Lite LED 546
Photolysis Device (Biotium) for 5 min, followed by 2 min on ice and additional light exposure 547
for 5 min [46]. 548
For vPCR as well as nanopore sequencing, DNA was extracted using the DNeasy® Blood and 549
Tissue Kit (QIAGEN, Hilden, Germany) according to the manufacturer’s instructions. The 550
amount of chlamydial DNA in all vPCR and nanopore sequencing samples was quantified with 551
a sensitive Chlamydiaceae qPCR [ 47]. For subsequent nanopore sequencing of the DNA 552
extracts (viable: 0.62 ng/μL; dead: 0.17 ng/μL), we followed the same approach as described 553
for E. coli above. We used three barcodes for each sample, resulting in DNA input of 18.66 ng 554
and 5.07 ng for the “viable” and “dead” library, respectively. After following the same 555
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
16
processing steps as established for the E. coli experiment, we filtered reads at a minimum 556
average quality score of 8 and a minimum length of 100 b using Nanofilt v2.8.0 [ 54], and 557
created de novo assemblies using metaFlye v2.9.1 [55], followed by polishing using minimap2 558
v2.17 [56] and three rounds of Racon v1.5 [ 57]. We finally used Kraken2 v2.0.7 [ 58] and the 559
NCBI nt database for taxonomic classification of the assembled contigs. 560
561
Data Availability Statement 562
All raw data has been made publicly available via ENA (study accession number: 563
PRJEB76420). All code has been made publicly available via Github: 564
https://github.com/Genomics4OneHealth/Squiggle4Viability.git. 565
566
Financial Disclosure Statement 567
This study was funded by a Helmholtz Principal Investigator Grant awarded to LU. HU was 568
supported by the Helmholtz Association under the joint research school “HIDSS-006 - Munich 569
School for Data Science@Helmholtz, TUM&LMU”. EM was supported by an EASTBIO 570
studentship, funded by BBSRC Grant Number BB/M010996/1, and an STFC Food Network+ 571
Scoping Grant. SB and SK were supported by Helmholtz AI’s Helmholtz Association Initiative 572
and Networking Fund. Computational Resources were provided by Helmholtz Munich and by 573
the Helmholtz Association Initiative and Networking Fund HAICORE partition at the 574
Forschungszentrum Jülich. LU, HU, and JMF have previously received travel and 575
accommodation expenses to speak at Oxford Nanopore Technologies’ conferences. 576
577
Acknowledgments 578
We thank the laboratory of the Research Unit of Comparative Microbiome Analysis at 579
Helmholtz Munich, Germany, especially Cornelia Galonska, for their support in processing the 580
Escherichia coli samples. We thank the laboratory of the Institute of Veterinary Pathology, 581
Switzerland, especially Theresa Pesch, for their support in processing the Chlamydia abortus 582
samples. We further thank Valentin Rauscher for his help in training several deep models 583
during his internship in the Urban research group at Helmholtz Munich. 584
585
Competing interests 586
The authors declare no competing interests. 587
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
17
References
588
1. Lewis WH, Tahon G, Geesink P, Sousa DZ, Ettema TJG. Innovations to culturing the uncultured microbial 589
majority. Nat Rev Microbiol. 2021 Apr;19(4):225-40. 590
2. Lloyd KG, Steen AD, Ladau J, Yin J, Crosby L. Phylogenetically novel uncultured microbial cells dominate 591
Earth microbiomes. mSystems. 2018 Sep 25;3(5). 592
3. Human Microbiome Jumpstart Reference Strains Consortium, Nelson KE, Weinstock GM, Highlander SK, 593
Worley KC, Creasy HH, et al. A catalog of reference genomes from the human microbiome. Science. 2010 594
May 21;328(5981):994-9. 595
4. Sauerborn E, Corredor C, Reska T, Perlas A, Vargas da Fonseca Atum S, Goldman N, Wantia N, Prazeres 596
da Costa C, Foster -Nyarko E, Urban L. Detection of hidden antibiotic resistance through real -time 597
genomics. Preprint. Research Square. 2023 Dec 4. Available from: https://doi.org/10.21203/rs.3.rs-598
3620416/v1. 599
5. Handelsman J. Metagenomics: application of genomics to uncultured microorganisms. Microbiol Mol Biol 600
Rev. 2004 Dec;68(4):669-85.. 601
6. Pace NR. Mapping the tree of life: progress and prospects. Microbiol Mol Biol Rev. 2009 Dec;73(4):565-602
76. 603
7. Brooks JP, Edwards DJ, Harwich MD, Rivera MC, Fettweis JM, Serrano MG, et al. The truth about 604
metagenomics: quantifying and counteracting bias in 16S rRNA studies. BMC Microbiol. 2015;15:66. 605
8. Dick GJ, Andersson AF, Baker BJ, Simmons SL, Thomas BC, Yelton AP, Banfield JF. Community -wide 606
analysis of microbial genome sequence signatures. Genome Biol. 2009;10(8). 607
9. Quince C, Walker AW, Simpson JT, Loman NJ, Segata N. Shotgun metagenomics, from sampling to 608
analysis. Nat Biotechnol. 2017 Sep;35(9):833-44. 609
10. Sereika M, Kirkegaard RH, Karst SM, Michaelsen TY, Sørensen EA, Wollenberg RD, Albertsen M. Oxford 610
Nanopore R10.4 long-read sequencing enables the generation of near -finished bacterial genomes from 611
pure cultures and metagenomes without short -read or reference polishing. Nat Methods. 2022 612
Jul;19(7):823-6. 613
11. Liu L, Yang Y, Deng Y, Zhang T. Nanopore long -read-only metagenomics enables complete and high -614
quality genome reconstruction from mock and complex metagenomes. Microbiome. 2022 Dec 615
2;10(1):209. 616
12. Jain M, Olsen HE, Paten B, Akeson M. The Oxford Nanopore MinION: delivery of nanopore sequencing 617
to the genomics community. Genome Biol. 2016;17:239. 618
13. Pagès-Gallego M, de Ridder J. Comprehensive benchmark and architectural analysis of deep learning 619
models for nanopore sequencing basecalling. Genome Biol. 2023 Apr 11;24(1):71. 620
14. Urban L, Perlas A, Francino O, Martí -Carreras J, Muga BA, Mwangi JW, et al. Real -time genomics for 621
One Health. Mol Syst Biol. 2023 Aug 8;19(8). 622
15. Nogva HK, Drømtorp SM, Nissen H, Rudi K. Ethidium monoazide for DNA -based differentiation of viable 623
and dead bacteria by 5'-nuclease PCR. Biotechniques. 2003 Apr;34(4):804-8, 810, 812-3. 624
16. Hellmann KT, Tuura CE, Fish J, Patel JM, Robinson DA. Viability -resolved metagenomics reveals 625
antagonistic colonization dynamics of Staphylococcus epidermidis strains on preterm infant skin. 626
mSphere. 2021 Oct 27;6(5). 627
17. Reska T, Pozdniakova S, Borras S, Schloter M, Cañas L, Perlas A, Rodó X, Winkler B, Schnitzler JP, 628
Urban L. Air monitoring by nanopore sequencing. bioRxiv. 2023 Dec 19:2023.12.19.572325. 629
doi:10.1101/2023.12.19.572325. 630
18. Morré SA, Sillekens PT, Jacobs MV, de Blok S, Ossewaarde JM, van Aarle P, et al. Monitoring of 631
Chlamydia trachomatis infections after antibiotic treatment using RNA detection by nucleic acid sequence 632
based amplification. Mol Pathol. 1998 Jun;51(3):149-54. 633
19. Herman L. Detection of viable and dead Listeria monocytogenes by PCR. Food Microbiol. 1997 634
Apr;14(2):103-110. 635
20. Min J, Baeumner AJ. Highly sensitive and specific detection of viable Escherichia coli in drinking water. 636
Anal Biochem. 2002 Apr 15;303(2):186-93. 637
21. Urban L, Holzer A, Baronas JJ, Hall MB, Braeuninger-Weimer P, Scherm MJ, et al. Freshwater monitoring 638
by nanopore sequencing. eLife. 2021 Jan 19;10. 639
22. Kumar SS, Ghosh AR. Assessment of bacterial viability: a comprehensive review on recent advances and 640
challenges. Microbiology (Reading). 2019 Jun;165(6):593-610. 641
23. Taylor-Brown A, Madden D, Polkinghorne A. Culture -independent approaches to chlamydial genomics. 642
Microb Genom. 2018 Feb;4(2). 643
24. Nocker A, Camper AK. Novel approaches toward preferential detection of viable cells using nucleic acid 644
amplification techniques. FEMS Microbiol Lett. 2009 Feb;291(2):137-42. 645
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
18
25. Sheridan GE, Masters CI, Shallcross JA, Mackey BM. Detection of mRNA by reverse transcription -PCR 646
as an indicator of viability in Escherichia coli cells. Appl Environ Microbiol. 1998 Apr;64(4):1313-8. 647
26. Wong VY, Duval MX. Inter -laboratory variability in array -based RNA quantification methods. Genomics 648
Insights. 2013 May 6;6:13-24. 649
27. Hønsvall BK, Robertson LJ. From research lab to standard environmental analysis tool: Will NASBA make 650
the leap? Water Res. 2017 Feb 1;109:389-397. 651
28. Janssen KJH, Dirks JAMC, Dukers -Muijrers NHTM, Hoebe CJPA, Wolffs PFG. Review of Chlamydia 652
trachomatis viability methods: assessing the clinical diagnostic impact of NAAT positive results. Expert 653
Rev Mol Diagn. 2018 Aug;18(8):739-747. 654
29. Nocker A, Cheung CY, Camper AK. Comparison of propidium monoazide with ethidium monoazide for 655
differentiation of live vs. dead bacteria by selective removal of DNA from dead cells. J Microbiol Methods. 656
2006 Nov;67(2):310-20. 657
30. Emerson JB, Adams RI, Román CMB, Brooks B, Coil DA, Dahlhausen K, et al. Schrödinger's microbes: 658
Tools for distinguishing the living from the dead in microbial ecosystems. Microbiome. 2017 Aug 659
16;5(1):86. 660
31. Bae S, Wuertz S. Discrimination of viable and dead fecal Bacteroidales bacteria by quantitative PCR with 661
propidium monoazide. Appl Environ Microbiol. 2009 May;75(9):2940-4. 662
32. Courcelle J, Donaldson JR, Chow KH, Courcelle CT. DNA damage -induced replication fork regression 663
and processing in Escherichia coli. Science. 2003 Feb 14;299(5609):1064-7. 664
33. Setlow RB, Swenson PA, Carrier WL. Thymine dimers and inhibition of DNA synthesis by ultraviolet 665
irradiation of cells. Science. 1963 Dec 13;142(3598):1464-6. 666
34. Shibai A, Takahashi Y, Ishizawa Y, Motooka D, Nakamura S, Ying BW, Tsuru S. Mutation accumulation 667
under UV radiation in Escherichia coli. Sci Rep. 2017 Nov 6;7(1):14531. 668
35. Ni P, Huang N, Zhang Z, Wang DP, Liang F, Miao Y, Xiao CL, Luo F, Wang J. DeepSignal: detecting DNA 669
methylation state from Nanopore sequencing reads using deep -learning. Bioinformatics. 2019 Nov 670
1;35(22):4586-95. 671
36. Stoiber M, Quick J, Egan R, Lee JE, Celniker S, Neely RK, Loman N, Pennacchio LA, Brown J. De novo 672
identification of DNA modifications enabled by genome-guided nanopore signal processing. bioRxiv. 2016 673
Dec 15;094672. doi:10.1101/094672. 674
37. Wang X, Ahsan MU, Zhou Y, Wang K. Transformer -based DNA methylation detection on ionic signals 675
from Oxford Nanopore sequencing data. Quant Biol. 2023;11(3):287-96. 676
38. An N, Fleming AM, White HS, Burrows CJ. Nanopore detection of 8 -oxoguanine in the human telomere 677
repeat sequence. ACS Nano. 2015 Apr 28;9(4):4296-307. 678
39. Turin L, Surini S, Wheelhouse N, Rocchi MS. Recent advances and public health implications for 679
environmental exposure to Chlamydia abortus: from enzootic to zoonotic disease. Vet Res. 2022 May 680
31;53(1):37. 681
40. Zhou B, Khosla A, Lapedriza A, Oliva A, Torralba A. Learning deep features for discriminative localization. 682
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2016 Jun 27 -30; 683
Las Vegas, NV, USA. p. 2921-9. 684
41. Ferguson JM, Smith MA. SquiggleKit: a toolkit for manipulating nanopore signal data. Bioinformatics. 2019 685
Dec 15;35(24):5372-3. 686
42. Gamaarachchi H, Samarakoon H, Jenner SP, Ferguson JM, Amos TG, Hammond JM, et al. Fast 687
nanopore sequencing data analysis with SLOW5. Nat Biotechnol. 2022 Jul;40(7):1026-9. 688
43. Borel N, Polkinghorne A, Pospischil A. A review on chlamydial diseases in animals: still a challenge for 689
pathologists? Vet Pathol. 2018 May;55(3):374-90. 690
44. Borel N, Sachse K. Zoonotic transmission of Chlamydia spp.: known for 140 years, but still 691
underestimated. In: Sing A, editor. Zoonoses: infections affecting humans and animals. Cham: Springer 692
International Publishing; 2023. p. 1-28. 693
45. Borel N, Marti H, Pospischil A, Pesch T, Prähauser B, Wunderlin S, et al. Chlamydiae in human intestinal 694
biopsy samples. Pathog Dis. 2018 Nov 1;76(8). 695
46. Janssen KJ, Hoebe CJ, Dukers-Muijrers NH, Eppings L, Lucchesi M, Wolffs PF. Viability-PCR shows that 696
NAAT detects a high proportion of DNA from non -viable Chlamydia trachomatis. PLoS One. 2016 Nov 697
3;11(11). 698
47. Loehrer S, Hagenbuch F, Marti H, Pesch T, Hässig M, Borel N. Longitudinal study of Chlamydia pecorum 699
in a healthy Swiss cattle population. PLoS One. 2023 Dec 11;18(12). 700
48. McDonald MD, Owusu -Ansah C, Ellenbogen JB, Malone ZD, Ricketts MP, Frolking SE, et al. What is 701
microbial dormancy? Trends Microbiol. 2024 Feb;32(2):142-50. 702
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
19
49. Potgieter M, Bester J, Kell DB, Pretorius E. The dormant blood microbiome in chronic, inflammatory 703
diseases. FEMS Microbiol Rev. 2015 Jul;39(4):567-91. 704
50. Borel N, Dumrese C, Ziegler U, Schifferli A, Kaiser C, Pospischil A. Mixed infections with Chlamydia and 705
porcine epidemic diarrhea virus - a new in vitro model of chlamydial persistence. BMC Microbiol. 2010 Jul 706
27;10:201. 707
51. Howard L, Orenstein NS, King NW. Purification on renografin density gradients of Chlamydia trachomatis 708
grown in the yolk sac of eggs. Appl Microbiol. 1974 Jan;27(1):102-6. 709
52. Marti H, Biggel M, Shima K, Onorini D, Rupp J, Charette SJ, Borel N. Chlamydia suis displays high 710
transformation capacity with complete cloning vector integration into the chromosomal rrn -nqrF plasticity 711
zone. Microbiol Spectr. 2023 Dec 12;11(6). 712
53. Leonard CA, Schoborg RV, Borel N. Damage/danger associated molecular patterns (DAMPs) modulate 713
Chlamydia pecorum and C. trachomatis serovar E inclusion development in vitro. PLoS One. 2015 Aug 714
6;10(8). 715
54. De Coster W, D'Hert S, Schultz DT, Cruts M, Van Broeckhoven C. NanoPack: visualizing and processing 716
long-read sequencing data. Bioinformatics. 2018 Aug 1;34(15):2666-9. 717
55. Kolmogorov M, Bickhart DM, Behsaz B, Gurevich A, Rayko M, Shin SB, Kuhn K, Yuan J, Polevikov E, 718
Smith TPL, Pevzner PA. metaFlye: scalable long -read metagenome assembly using repeat graphs. Nat 719
Methods. 2020 Nov;17(11):1103-10. 720
56. Li H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 2018 Sep 15;34(18):3094 -721
100. 722
57. Vaser R, Sović I, Nagarajan N, Šikić M. Fast and accurate de novo genome assembly from long 723
uncorrected reads. Genome Res. 2017 May;27(5):737-46. 724
58. Wood DE, Lu J, Langmead B. Improved metagenomic analysis with Kraken 2. Genome Biol. 2019 Nov 725
28;20(1):257. 726
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
20
Supporting information 727
728
Table S1. Performance Metrics of all Deep Neural Network Architectures tested for Viabiliy Inferences. 729
Test set performance of residual neural network (ResNet) and transformer architectures, trained on various data 730
modalities (Nanopore “Signal” or DNA sequence aka “Nucleotide”) and various signal chunk sizes (LR=1e-4). 731
732
Model
Architecture
Data
Modality
Length
[Signal/Base]
Accuracy F1 Precision Sensitivity Specificity AUROC AUPR
ResNet1 Signal 10K 0.83 0.84 0.78 0.91 0.74 0.90 0.87
ResNet2 Signal 10K 0.83 0.85 0.77 0.94 0.72 0.89 0.86
ResNet3 Signal 10K 0.81 0.82 0.76 0.89 0.72 0.87 0.83
Transformer Signal 10K 0.79 0.82 0.73 0.92 0.67 0.86 0.82
ResNet1 Nucleotide
(A,C,G,T)
800 0.51 0.51 0.51 0.52 0.50 0.52 0.53
ResNet1 Nucleotide
(A,C,G,T,M)
800 0.51 0.50 0.51 0.50 0.52 0.51 0.51
ResNet1 Signal 1K 0.58 0.65 0.55 0.80 0.35 0.61 0.58
ResNet1 Signal 5K 0.71 0.76 0.66 0.89 0.54 0.78 0.73
ResNet1 Signal 7K 0.77 0.79 0.72 0.88 0.65 0.83 0.79
ResNet1 Signal 12K 0.84 0.85 0.77 0.97 0.71 0.91 0.88
ResNet1 Signal 20K 0.89 0.90 0.85 0.95 0.83 0.94 0.91
733
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
21
734
Fig S1. Training and Validation Loss across Deep Neural Network Architectures tested for Nanopore-Signal 735
based Viability Inference, and across different Learning Rates. 736
(A-C) Model loss of ResNet1 at LRs of 1e-3, 1e-4, and 1e-5; (D-F) Model loss of ResNet2 at LRs of 1e-3, 1e-4, and 737
1e-5; (G-I) Model loss of ResNet3 at LRs of 1e -3, 1e-4, and 1e-5; and (J-L) Model loss of the transformer models 738
at LRs of 1e-3, 1e-4, and 1e-5. The solid blue line indicates the training loss, the solid red line indicates the validation 739
loss, and the dashed line indicates the minimum validation loss from the final ResNet1, LR=1e-4, model. 740
741
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
22
742
Fig S2. Training and Validation Loss of ResNet1 at various signal chunk sizes. 743
The signal chunk size varies from (A) 1k, (B) 5k, (C) 7k, (D) 10k, to (E) 12k and (F) 20k. The solid blue line indicates 744
the training loss, the solid red line indicates the validation loss, and the dashed line indicates the minimum validation 745
loss from the final (D) ResNet1model using a signal chunk size of 10k. 746
747
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
23
Fig S3. Exemplary drops in ResNet1 prediction probabilities in nanopore signal chunks after consecutive masking of the signal region with the respectively highest
CAM value. Figure headers: signal chunk ID / total number of masked signal values / prediction probability per signal chunk “prob”. Left to right: Five exemplary nanopore signal
chunks (length of 10k signals). Top to bottom : Consecutive masking of 200 signal values per masking event (Materials and Methods). Legends: Red-colored CAM value
vizualisations.
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted June 11, 2024. ; https://doi.org/10.1101/2024.06.10.598221doi: bioRxiv preprint
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.