Efilter: A Tool for Identifying Unexpected and Erroneous Taxids in Sequencing Data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Efilter: A Tool for Identifying Unexpected and Erroneous Taxids in Sequencing Data Sharon Bewick, Xianghui Dong, David Karig, William F. Fagan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-29648/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: Multiple types of error can enter metagenomics sample analysis, including contamination during sample collection, library preparation and sequencing, as well as incorrect taxonomic assignment by bioinformatics packages. Often, such errors either go unidentified, or else are removed, ad hoc , based on user knowledge of microbial ecology. However, becausedifferent researchers are more or less familiar with the ecologies of the various organisms in their systems, filtering is applied non-uniformly at best, with differences in the degree of filtering between studies and between taxonomic groups within a single study. Results: In this paper, we present EFILTER – a tool that capitalizes on decades of research in microbial ecology to identify suspicious or spurious taxa in metagenomics samples based on habitat information. Conclusions: EFILTER allows all microbiome researchers, regardless of background, to examine taxon lists for unusual entries, and to do so in a manner that is systematic and without bias. Biotechnology and Bioengineering taxonomic identification microbial habitat association online tool metagenomics datasets contamination bioinformatics errors novel niches taxon discovery Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Background The ability to extract, clone and analyze DNA directly from environmental samples has revolutionized disciplines ranging from biomedical science to environmental microbiology[1-3]. Compared to historic, culture-based approaches,sequencing methods offer unprecedented advancements in terms of cost, effort and speed. They also improve detection of full community membership by exposing unculturable organisms [4]. However, sequencing studies are not without their own set of difficulties.One of the greatest challenges lies in interpreting sequencingresults, which can be complicated by issues of contamination [5-7], sequencing error [8, 9], and bioinformatics misidentification[10]. Although improved laboratory and bioinformatics techniques canhelp to alleviate these difficulties, increased focus on detecting rare taxa, along with more specific classification down to species and even strains, means that interpretation of sequencing data will remain a significant challenge for the foreseeable future. Contamination is perhaps the largest and most difficult to address source of confusion in sequencing data. Whereasculture-based methods only detect organisms that can be reliably and repeatedly grown from an environmental sample, sequencing techniques identify even small pieces of non-viable DNA. While this is one of the strengths of the sequencing approach, and the main reason that it can uncover unculturable organisms, thisis also why sequencing methods arefar more susceptible to issues with contamination[5-7]. Unfortunately, contamination can occur at any step along the processing pipeline, ranging from issues during sample collection and storage [11], though reagent contamination[12-15] to room source contamination andcontaminating bacteria from the mouth and skin of the researcher during sample preparation [16, 17].While it is relatively straightforward to detect contaminants in samples that should contain a single organism, for example microbial DNA in the cow genome [5], or contamination of pure bacterial cultures with other bacterial taxa [18], identifying contamination in complex,mixed microbial communities presents a more difficult, and still open challenge. Another major issue with sequencing data is imperfect taxonomic assignment. This challenge is also not easy to overcome, since there is an inherent trade-off between finding all taxa within a sample, and confidence/accuracy in taxon prediction [10]. Thus, bioinformatics methods that are highly sensitive are also more likely to report organisms not actually present. Meanwhile, more conservative classification methods can reduce the number of false positives, but do so at a cost of increasing false negatives. Misclassificationcan be more or less problematic, depending on the microbial community involved, the type of sequencing method employed[19-21], the bioinformatics toolkit used[10, 22, 23], the reference databases available [24],and the quality (e.g., sequencing platform statistics, read depth, and read length) of the sequencing data[25-27].As with contamination, even in simple systems with appropriately chosen sequencing and bioinformatics methods, classification errors can have a significant impact on conclusions, for example by massively inflating diversity or the relative importance of rare organisms. One method for identifying potentially problematic taxIDs in sequencing data is to rely on knowledge of microbial ecology. Finding a common plant pathogen in a human gut microbiome dataset, for example, might raise a red flag. Whether consciously or not, researchers use this approach to clean their data, for instance by deciding whether or not to trust all of the output taxa from a particular experiment [11]. Unfortunately, use of such information requires extensive knowledge of microbial behavior and characteristics – knowledge that is lacking among many, if not most researchers performing metagenomics analysis. Indeed, even for researchers with strong backgrounds in microbial ecology, the vast diversity of microbes from different environments means that it is impossible for any single researcher to be familiar with all common microbial habitats. Consequently, even with extensive knowledge of the ecologies of individual bacterial taxa, data verification based on microbial ecology is likely to be biased and ad hoc . Nevertheless, if such verification can be standardized and automated, it remains a promising method for addressing the challenges of spurious taxIDs in microbiome datasets. In this paper, we present EFILTER ( https://efilter.shinyapps.io/EFilter-app/ )– a computational tool that performs E nvironmental FILTER ing by leveraging known relationships between bacterial taxa and specific habitats in order to identify taxIDs that are surprising or unlikely given the source of a metagenomics sample. EFILTER incorporates all habitat information available in Bergey’s Manual of Systematic Bacteriology and National Center for Biotechnology Information (NCBI) records from the International Journal of Systematic and Evolutionary Microbiology (IJSEM). As such, EFILTER allows users to assess the likelihood of taxIDs being in their datasets based on known microbial ecology. Further, it does so without requiring an extensive background in microbial ecology, and also without introducing bias that might otherwise arise based on user experience and knowledge of specific taxonomic groups. To demonstrate the strength and performance of EFILTER, we test our tool on 4 datasets,which include both 16S rRNA and shotgun sequencing methods, analyzed with different bioinformatics pipelines and studied at different taxonomic ranks. Our analysis shows that EFILTER can reliably provide lists of unexpected/spurious organisms in metagenomics samples, offering a new,ecologically motivated tool for validation of microbiome datasets. Materials And Methods The EFilter Database We constructed a database of microbe-environment associations using information from Bergey’s Manual of Systematic Bacteriology and the National Center for Biotechnology Information (NCBI) Nucleotide database. For Bergey’s , we converted pdf files of Volumes IIB, IIC, III, IV, and V as well as additional chapters on Aquificae , Chlorobi , Chloroflexi , Chrysiogenetes , Crenarchaeota , Cyanobacteria , Deferribacteres , Deinococcus-Thermus , Euryarchaeota , Nitrospirae , Thermodesulfobacteria , Thermomicrobia , and Thermotogae to .txt files using the pdfminer functions in Python. The text was then separated into segments for each taxonomic group. Text describing genera and species was mined for environmental information. We did not use text describing higher taxonomic groups, since this information tends to be too broad to be useful. To identify sentences containing environmental information within genus and species descriptions, we used the keywords in Table A.1 (see Appendix A). For the NCBI Nucleotide database, we only used entries with IJSEM[Journal] in the journal field. Our assumption was that all of these entries represent newly defined taxa. Although taxa can be defined in other journals, the International Journal of Systematic and Evolutionary Microbiology (IJSEM) represents one of the largest publishers of new taxon descriptions. Most other journals do not specialize in taxonomy, and thus may contain other types of studies where taxonomic specification is not precise/accurate. Our goal was to avoid including habitat information from inaccurate sources, since this is the very problem that EFILTER seeks to correct. Habitat information was taken from the ‘isolation source/’ field of the NCBI entries. NCBI/IJSEM entries were only used for taxa that did not exist in Bergey’s (i.e., taxa that were defined since publication of the most recent Bergey’s volume for that taxonomic group). Habitat descriptions from Bergey’s Manual of Systematic Bacteriology and NCBI were stored in tab-delimited .txt files (see Appendix B), along with relevant taxonomic information, forming the basis of our environmental database. For species, all searches are performed on species-level descriptions harvested from either Bergey’s or NCBI/IJSEM. For genera, searches are performed on genus-level descriptions from Bergey’s , as well as all species-level descriptions for species within the genus. For higher taxonomic ranks, searches are performed on all genus-level and species-level descriptions within the particular taxonomic group. User-Defined Search Terms The entire environmental database can be queried for keywords. In order to minimize search time, EFILTER contains a look-up table linking each unique keyword from the database (see above) to the species/genera in whose descriptions that keyword appears. That is, we built a table of keywords wherein each keyword (column) maps to many rows, with each row being a specific species/genus in the database. All keywords in our look-up table are full, stand-alone words. However, in many cases, users may be interested in related words (e.g., ‘dog’ and ‘dogs’) as well. For this reason, whenever a keyword is typed into EFILTER, EFILTER automatically generates a list of all larger, stand-alone words (‘ dog s, dog -bite’) containing the inputted keyword. This enables users to select the precise keyword combinations that they want to include. Thus, for the ‘dog’ example, the user may want to include ‘dogs’, ‘dog-bite’, and ’dogbite’ , but not ‘peptidoglycan’, ‘endogenous’, etc. Finally, the environmental database can be queried for phrases (e.g., ‘dogs and cats’). In this case, however, there is no look-up table, and the phrase is searched for directly through the database. Consequently, searches that involve querying for phrases are somewhat slower. Pre-defined Environments Although possible, it may be hard and time-consuming for individual users to think of all potential keywords associated with a particular microbial environment. For this reason, we pre-defined keyword combinations for 46 common habitats falling within six broad classes (animal, body sites, food, plant, environment and specialized). To generate keyword combinations, we manually curated all keywords appearing in ³5 species/genus descriptions and then went through these keywords individually to decide whether they or any of their word derivatives fell into our common environmental categories. We believe that our pre-defined environment feature will be the most useful for the majority of users. Lists of keyword combinations used for each pre-defined environment (ranging from 6 words for ‘sand’ to 199 words + derivatives for ‘other mammals’) can be found in Appendix C. Results General Usage The EFILTER database contains 11633 species and 34 phyla (see Table B.1, Appendix B). 72% of species have habitat information specified, while 45% fall into at least one pre-defined environment (see Methods). 100% of phyla have habitat information specified, while 97.1% fall into at least one pre-defined environment. The pre-defined environment with the largest number of taxa is ‘soil’, which includes 17.5% of species. The next largest is ‘human’, which includes 8.4% of species. The pre-defined environments with the fewest number of taxa are ‘amphibian’, ‘reptile’ and ‘symbiont’, with 0.06%, 0.3% and 0.3% of species respectively (see Appendix B, Table B.2). EFILTER inputs .txt files (including batch upload) in the form of lists of either Latin names or taxIDs. Inputs in the form of taxIDs are mapped onto Latin names in order to infer habitat associations. Only well defined taxa (i.e., taxa with entries in either Bergeys or NCBI/IJSEM, see Methods)areincluded in the database. EFILTER allows Latin names or taxIDs referencing any taxonomic rank from phylum to species, including lists with mixed ranks. Taxonomic rank is automatically determined based on name or taxID. Each taxon in an input file is queried against the EFILTER database using a defined environmental filter. Environmental filters can be constructed as user-defined keywords, user-defined phrases, or pre-defined environments (see Methods). In addition, users can construct any logical combination (‘AND’, ‘OR’, ‘NOT’, for example ‘eye AND human NOT other mammal’) of single filters. A video tutorial covering filter definition can be found at ( https://www.youtube.com/watch?v=UU1rPTUZPTE ). Upon file upload, EFILTER outputs the number of taxa at each taxonomic rank as well as the number of taxa that were not found in the EFILTER database. Missing taxa include taxa that are not well-defined, taxa that were defined more recently than the most recent EFILTER update, or taxa that are not bacterial (for example fungal taxa from the output of shotgun sequencing). These taxa are ignored, meaning that there is no need to curate shotgun sequencing data to include strictly bacteria. EFILTER also automatically reports the pre-defined environmental classes for each unique taxon in an uploaded dataset, starting from the input taxonomic rankup to the taxonomic rank of family. The primary goal of EFILTER is to identify organisms that are consistent/inconsistent with the sampled environment (i.e., the user selected filter). Thus, when a taxon list is queried against a chosen filter, EFILTER outputs the number of taxa with at least one habitat in the database that matches (passes) the filter.Taxa without at least one matching habitat are identified as failing the filter. In addition, for each input taxon, EFILTER reports whether the taxon passed or failed at taxonomic ranks higher than the inputted rank. Finally, EFILTER outputs the names of taxa that passed the filter at taxonomic ranks lower than the inputted rank. This latter feature can be useful, for example, to make an educated guess as to which species might be present given a list of genera from a 16S rRNA sequencing run. Samples versus Controls As a first demonstration of EFILTER performance, we used our own data, which consisted of 16S rRNA sequencing of 375skin microbiome samples from the foreheads of 50 individuals, along with 100paired controls (blank swabs collected before and/or after each person was sampled). Figure 1 shows the percentage of genera passing four different filters averaged over the samples and paired controls for each person. For the ‘skin’ filter (A), more organisms passed in the samples relative to the controls for the majority (48/50 or 96%) of people. By contrast, for the ‘built’ filter (C) more organisms passed in the controls relative to the samples for the majority (44/50 or 88%) of people. As expected, for logical OR combinations of filters, the percentages of passing taxa were greater overall (compare, for example Figures 1A and 1B or Figures 1C and 1D). However, once again, for the ‘human/mammal/body-site’ filter (C) more taxa passed in the samples relative to the controls for the majority (40/50 or 80%) of people, whereas the reverse was true for the ‘built/water’ filter (D) for the majority (30/50 or 60%) of people. In general, one would expect that a skin/human/mammal/body-site filter would perform better on human samples than on non-human (e.g., control) samples, which is what we see. Likewise, assuming that many contaminants come from room air or sequencing preparation steps, one would assume that a built/water filter would perform better on control samples than on human samples, which is again what we see. Samples from Different Sources As a second demonstration of EFILTER performance, we comparedvarious EFILTER pre-defined environmental filters across a variety of environmental samples. Datasets were downloaded as lists of genera directly from the MG-RAST website ( https://www.mg-rast.org/ ). Specifically, we considered hot springs (mgp5265, mgp6907, mgp5356, mgp5355, mgp5270, mgp5269, mgp5268, mgp5266), tropical soil (mgp4362), a freshwater lake (mgp19525), oceans (mgp20413), cheese (mgp14606), human nares (mgp385), and horses guts (mgp7746). All studies involved shotgun sequencing. We used the same number samples from each environment (8,limited by the availability of hot springs samples; in all cases we selected the 8 samples with the lowest MG-RAST ID numbers), and restricted our analysis to the top 25 most abundant taxa in each sample. Figure 2 shows confusion matrices for samples versus environmental filters.Focusing on the averaged confusion matrix (right panel), we see strong EFILTER performance, with red/orange (high percentage of passing taxa) concentrated along the diagonal and blue/green (low percentage of passing taxa) off-diagonal. Overall, across the seven paired environments and samples, an average of 72% of taxa passed their expected filters. Passing rates were highest (88.5%) for the anterior nares, and lowest (62.5%) for the lake. Filter fidelity to sample was quite high. For five out of seven samples (hot springs, soil, oceans, cheese and anterior nares), the expected filter (extreme, soil, salt water, dairy, human) performed best. Even when this was not the case, the expected filter still performed well. For the lake study, for instance, both the salt water and soil filters outperformed the fresh water filter (note that all three are in the‘environment’ class, see Methods), though only marginally (63.5% and 62.8% of taxa passing for the salt water and soil filters, relative to 62.5% for the fresh water filter). Likewise, for the horse gut study, the human filter outperformed the other mammal filter (note that both are in the ‘animal’ class), although once again, the other mammal filter still performed well (79.4% passing for the human filter relative to 75.9% for the other mammal filter). Sample fidelity to filter was not as good. For three out of seven filters (extreme, soil and human) the highest percentages of passing taxa were identified in matched samples (hot springs, soil, and anterior nares respectively). For the remaining four filters, however, the highest percentages of passing taxa were identifiedin non-matched samples (the fresh and salt water filters identified the highest percentage of passing taxa in the soil sample, while the dairy and other mammal filtersidentified the highest percentage of passing taxa in the anterior nares). Samples from Different Body Sites As a third demonstration of EFILTER performance, we considered a shot-gun sequencing dataset from the Human Microbiome Project (HMP), including 690 samples from 15 body-sites. Importantly, in this dataset all samples fall under a single filter class (‘body-site’, see Methods). As in the previous section, we only considered the top 25 most abundant taxa in each sample. Figure 3 shows confusion matrices for species- and genus-level data for each sample individually (A), and as sample averages over each body region (B; notice that, unlike Figure 2, there are different numbers of samples contributing to each body region). EFILTER performance on the HMP dataset was even higher than on the environmental dataset, with an average of 40.9% of species and 76.0% of genera passing the appropriately matched filter. At the same time, however, filter fidelity to sample was lower (compare the right panel in Figure 3B to Figure 2). That is, at the rank of both genus and species, the expected filter (gut/digestive system, oral) performed best for only two (St, To+Kg+Bm+Sa+Sb+Sp+Td) out of five body regions. For the remaining three regions, the gut/digestive system and/or oral filters performed better. The gut/digestive system filter likely performed well because it has the largest number of taxa (see Appendix B). It is less clear why the oral filter out-performed the other filters, since it has fewer taxa than the ear/nose/throat filter at the ranks of both genus and species and has approximately the same number of taxa as the skin filter, atthe rank of genus. HMP sample fidelity to filter, on the other hand, was strong, at least at the rank of species. In particular, four out of five filters (skin, oral, gut/digestive system, female reproductive system) identified the highest percentage of passing taxa in their matched samples (Rc, To+Kg+Bm+Sa+Sb+Sp+Td, St, Va). Sample fidelity to filter was slightly weaker at the rank of genus, where only two out of five filters (oral, gut/digestive system) identified the highest percentage of passing taxa in their matched samples. This fidelity, was similar to the fidelity observed in our environmental dataset (see Figure 2).One particularly striking feature in Figure 3 is the poor performance of all but the gut/digestive system filter on stool samples. Whereas >56% of genera from every other body region passed every body-site filter, only 17-35% of stool genera passednon-gut/digestive system filters. At the same time, however, 95% of stool genera passed the gut/digestive system filter. This suggests that stool taxa are unique to the gut environment, whereas taxa from other body sites exhibit broad body distributions. Samples from Different Bioinformatics Pipelines As a final demonstration of EFILTER performance, we compared EFILTER output for different bioinformatics pipelines. A common problem with microbiome studies is that different pipelines can give different taxon lists, including substitutions among distantly related organisms. EFILTER can be used to gauge performance of different bioinformatics methods and to identify potentially spurious taxa introduced through taxonomic assignment steps. To illustrate this, we used a shotgun sequencing dataset of the human skin microbiome [28], processed using both Kraken [29] and MetaPhlAn [30], and thresholded at various different read percentages. Results are shown in Figure 4, which also serves as an example of the EFILTER graphical output. With larger thresholds, the pass:fail ratio increases for both bioinformatics pipelines. In other words, a larger fraction of the rare tail fails, regardless of the bioinformatics method used. This should not be surprising. Everything from contaminants and transient taxa to inherently rare/understudied organisms and bioinformatics errors are expected to contribute to the rare tail. In keeping with the consensus that Kraken has a tendency to over classify [10, 31], the pass:fail ratio for Kraken is significantly lower than it is for MetaPhlAn. This is particularly true when there is no threshold, with Kraken givingpass rates of 42% and35% for genera and species respectively relative to MetaPhlAn’s 67% and 61% at the same ranks. With a 1% threshold, both pipelines givesimilar results for species (72% passing), while Kraken actually shows a higher pass:fail ratio for genera (92% versus 88%). Surprisingly, though, despite Kraken’s lower pass:fail ratio for most scenarios, MetaPhlAn usuallyidentifies a greater absolute number of passing taxa. Thus, while it is likely that some of the failing taxa identified by Kraken (and MetaPhlAn) are truly present on skin,based on overall performance as judged by environmental consistency of assigned taxa, MetaPhlAn appears to be the better pipeline for this dataset. Discussion A wealth of microbiome research has emerged over the past decade. Most of it, however, has been plagued by contamination, sequencing errors and taxonomic misidentification[6-10]. This has led to a lack of reproducibility across pipelines and amongst labs. Recently, there has been a call for improved standards and validation of microbiome datasets, including efforts like the MicroBiome Quality Control (MBQC) project [15] and the Critical Assessment of Metagenome Interpretation (CAMI) initiative ( http://microbiome-cosi.org/cami ), as well as standards development by the National Institute of Standards and Technology (NIST) [32]. In this paper, we introduce EFILTER as a new tool for helping to assess the quality of microbiome datasets and for identifying suspicious taxIDs within them. EFILTER is unique amongst microbiome informatics approaches in that it leverages ecological information to target organisms that are unexpected based on sample source. Ultimately, EFILTER provides lists of suspicious taxIDs. The question, of course, is what to do with this information. In general, we advocate against indiscriminant dismissal of taxa that fail a chosen filter. Rather, two approaches are possible. First, depending on the analysis being performed, analysis can be run with and without discarding the problematic taxa in order to determine whether any conclusions change and, if they do, which and how many. In Figure 5, for example, we show how conclusions about relative sample diversity differ depending on whether analysis includes all taxa or only those taxa passing the salt water OR fresh water OR general water filter for the tara ocean dataset from MG-RAST (mgp20413). Notably, 93.6% of diversity orderings remain unchanged when failing taxa are ignored, suggesting that suspicious taxa are not overly problematic for the conclusions of thisparticular analysis with this particular dataset. A second approach for using EFILTER data is to carefully examine taxa that fail and to make decisions about whether to include them based on additional knowledge or further experimentation. Depending on the system, the list of failing taxa can still be large. Thus, we recommend using broad filters and focusing on particularly egregious failures. That is, either failures that involve very abundant taxa, or else failures that extend to higher taxonomic ranks. We illustrate such an approach in Table C.1 of Appendix C for our own skin microbiome dataset. Notably, although we cannot be certain about the source of any taxon failures, we can make educated guesses about a sizeable fraction. Some taxa, for example Geodermatophilus and Methylobacterium , are known contaminants of blanks from other sequencing studies, suggesting a contamination origin. Others, for example Rhodothermus and Rheinheimera are present in samples and controls taken at the same time or are ubiquitous in controls, again suggesting a contamination origin. Our lab has cultured certain suspicious taxa, for example Enhydrobacter , directly from skin, suggesting an unreported habitat for this organism. Finally, some taxa, including Modestobacter and Hymenobacter , are found in a range of samples and in no controls, hinting that they may be real, as yet undiscovered taxa from human skin. Supporting this claim, both Modestobacter and Hymenobacter have been found in other human microbiome datasets. Modestobacter , in particular, would be interesting to explore further. Although it is currently known from extreme (desert), rock/stone and soil environments it showed up in both 16S amplicon sequencing from our skin samples and in shotgun sequencing from a separate skin study performed in a different lab using an entirely different sequencing and bioinformatics pipeline[28]. This suggests that the source could be an as yet unknown species of Modestobacter that resides on human skin. Although the goal of EFILTER is to identify suspicious taxa in metagenomics samples, it is worth pointing out that EFILTER also works reasonably well for determining sample source origin, at least when the possible sources are quite distinct. In Figure 2, for instance, it is possible to discriminate soil/water sources, animal sources and extreme sources based on the percentage of abundant (top 25) genera passing the associated filters. Such discrimination is more difficult for closely related sources, for example for sources from different mammals, sources from different types of water, or sources from different human body regions (see Figure 3). A future goal would be to extend EFILTER such that source discrimination is improved, possibly by leveraging additional habitat information from sequencing studies. On a similar note, one of the downsides of the existing EFILTER database is that it only uses habitat information from validly published taxa descriptions. Although this ensures that contaminated or otherwise compromised metagenomics samples do not influence the EFILTER database, it means that EFILTER does not capitalize on the widely available, though imperfect, sequencing data currently in the public domain. A future direction would be to extend EFILTER to include these data, but to do so in a manner that incorporates confidence in the data source. This could result in higher percentages of taxa passing the appropriate filters, and also improved source discrimination. Conclusions Contamination, sequencing error and bioinformatics misidentifications will continue to plague microbiome research for the foreseeable future. Historically, many questionable results from sequencing studies have been addressed ad hoc by researchers who realize that particular organisms should not be present in particular samples. With EFILTER, we extend this capability to anyone involved in microbiome research, and do so in a way that allows rigorous, systematic, and repeatabledataset cleaning and validation based on ecologically relevant habitat information. Declarations Ethics approval and consent to participate: Not applicable Consent for publication: Not applicable Availability of data and material: All data generated or analysed during this study are included in this published article [and its supplementary information files]. Competing Interests: There are no competing interests. Funding: This material is based upon work supported by, or in part by, the U. S. Army Research Laboratory and the U. S. Army Research Office under contract/grant number #W911NF-14-1-0490. Authors' contributions: SB conceived of the idea, built the database, and ran analyses on test datasets; XD developed the online platform in RShiny; SB, XD, DK and WFF contributed to idea development, results interpretation, and manuscript writing. Acknowledgements: Not applicable References Stahl DA, Lane DJ, Olsen G, Pace NR: Analysis of hydrothermal vent-associated symbionts by ribosomal RNA sequences. Science 1984, 224: 409-412. Amann RI, Ludwig W, Schleifer K-H: Phylogenetic identification and in situ detection of individual microbial cells without cultivation. Microbiological reviews 1995, 59: 143-169. Venter JC, Remington K, Heidelberg JF, Halpern AL, Rusch D, Eisen JA, Wu D, Paulsen I, Nelson KE, Nelson W: Environmental genome shotgun sequencing of the Sargasso Sea. science 2004, 304: 66-74. Riesenfeld CS, Schloss PD, Handelsman J: Metagenomics: genomic analysis of microbial communities. Annu Rev Genet 2004, 38: 525-552. Merchant S, Wood DE, Salzberg SL: Unexpected cross-species contamination in genome sequencing projects. PeerJ 2014, 2: e675. Strong MJ, Xu G, Morici L, Bon-Durant SS, Baddoo M, Lin Z, Fewell C, Taylor CM, Flemington EK: Microbial contamination in next generation sequencing: implications for sequence-based analysis of clinical samples. PLoS Pathog 2014, 10: e1004437. Lusk RW: Diverse and widespread contamination evident in the unmapped depths of high throughput sequencing data. PloS one 2014, 9: e110808. Kunin V, Engelbrektson A, Ochman H, Hugenholtz P: Wrinkles in the rare biosphere: pyrosequencing errors can lead to artificial inflation of diversity estimates. Environmental microbiology 2010, 12: 118-123. Lee CK, Herbold CW, Polson SW, Wommack KE, Williamson SJ, McDonald IR, Cary SC: Groundtruthing next-gen sequencing for microbial ecology–biases and errors in community structure estimates from PCR amplicon pyrosequencing. PloS one 2012, 7: e44224. Peabody MA, Van Rossum T, Lo R, Brinkman FS: Evaluation of shotgun metagenomics sequence classification methods using in silico and in vitro simulated communities. BMC bioinformatics 2015, 16: 362. DeLong EF: Microbial community genomics in the ocean. Nature Reviews Microbiology 2005, 3: 459-469. Jervis-Bardy J, Leong LE, Marri S, Smith RJ, Choo JM, Smith-Vaughan HC, Nosworthy E, Morris PS, O’Leary S, Rogers GB: Deriving accurate microbiota profiles from human samples with low bacterial content through post-sequencing processing of Illumina MiSeq data. Microbiome 2015, 3: 19. Woyke T, Sczyrba A, Lee J, Rinke C, Tighe D, Clingenpeel S, Malmstrom R, Stepanauskas R, Cheng J-F: Decontamination of MDA reagents for single cell whole genome amplification. PloS one 2011, 6: e26161. Motley ST, Picuri JM, Crowder CD, Minich JJ, Hofstadler SA, Eshoo MW: Improved multiple displacement amplification (iMDA) and ultraclean reagents. BMC genomics 2014, 15: 443. Sinha R, Abnet CC, White O, Knight R, Huttenhower C: The microbiome quality control project: baseline study design and future directions. Genome biology 2015, 16: 276. Salter SJ, Cox MJ, Turek EM, Calus ST, Cookson WO, Moffatt MF, Turner P, Parkhill J, Loman NJ, Walker AW: Reagent and laboratory contamination can critically impact sequence-based microbiome analyses. BMC biology 2014, 12: 87. Weiss S, Amir A, Hyde ER, Metcalf JL, Song SJ, Knight R: Tracking down the sources of experimental contamination in microbiome studies. Genome biology 2014, 15: 564. Olson ND, Zook JM, Morrow JB, Lin NJ: Using metagenomic methods to detect organismal contaminants in microbial materials. PeerJ Preprints; 2017. Ranjan R, Rani A, Metwally A, McGee HS, Perkins DL: Analysis of the microbiome: advantages of whole genome shotgun versus 16S amplicon sequencing. Biochemical and biophysical research communications 2016, 469: 967-977. Qunfeng D, Claudia V: Evaluation of the RDP classifier accuracy using 16S rRNA gene variable regions. Metagenomics 2012, 2012 . Claesson MJ, Wang Q, O'sullivan O, Greene-Diniz R, Cole JR, Ross RP, O'toole PW: Comparison of two next-generation sequencing technologies for resolving highly complex microbiota composition using tandem variable 16S rRNA gene regions. Nucleic acids research 2010 : gkq873. Ricke DO, Shcherbina A, Chiu N: Evaluating performance of metagenomic characterization algorithms using in silico datasets generated with FASTQSim. bioRxiv 2016 : 046532. Hall RJ, Draper JL, Nielsen FG, Dutilh BE: Beyond research: a primer for considerations on using viral metagenomics in the field and clinic. Frontiers in microbiology 2015, 6 . Golob JL, Margolis E, Hoffman NG, Fredricks DN: Evaluating the accuracy of amplicon-based microbiome computational pipelines on simulated human gut microbial communities. BMC bioinformatics 2017, 18: 283. Schloss PD: The effects of alignment quality, distance calculation method, sequence filtering, and region on the analysis of 16S rRNA gene-based studies. PLoS Comput Biol 2010, 6: e1000844. Wommack KE, Bhavsar J, Ravel J: Metagenomics: read length matters. Applied and environmental microbiology 2008, 74: 1453-1463. Mosher JJ, Bowman B, Bernberg EL, Shevchenko O, Kan J, Korlach J, Kaplan LA: Improved performance of the PacBio SMRT technology for 16S rDNA sequencing. Journal of microbiological methods 2014, 104: 59-60. Oh J, Byrd AL, Deming C, Conlan S, Barnabas B, Blakesley R, Bouffard G, Brooks S, Coleman H, Dekhtyar M: Biogeography and individuality shape function in the human skin metagenome. Nature 2014, 514: 59. Wood DE, Salzberg SL: Kraken: ultrafast metagenomic sequence classification using exact alignments. Genome biology 2014, 15: R46. Segata N, Waldron L, Ballarini A, Narasimhan V, Jousson O, Huttenhower C: Metagenomic microbial community profiling using unique clade-specific marker genes. Nature methods 2012, 9: 811. Clooney AG, Fouhy F, Sleator RD, O’Driscoll A, Stanton C, Cotter PD, Claesson MJ: Comparing apples and oranges?: next generation sequencing and its impact on microbiome analysis. PLoS One 2016, 11: e0148028. Stulberg E, Fravel D, Proctor LM, Murray DM, LoTempio J, Chrisey L, Garland J, Goodwin K, Graber J, Harris MC: An assessment of US microbiome research. Nature microbiology 2016, 1: 15015. Supplementary Files BewicketalCL.docx AdditionalFile1.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-29648","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research","associatedPublications":[],"authors":[{"id":593691,"identity":"263d0552-da64-49f8-8205-9cd6fd9d92f7","order_by":1,"name":"Sharon Bewick","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+0lEQVRIiWNgGAWjYDACCQiVwMDA2MDYUHEAzDvwgHgtZw4w8IC0JBCnBaipsQ2ihQGfFvnZzc8+fNzBkMcvkdz4cea8O3L2YocfAm2xk9NtwK7F4M4x45kzzzAUS85IbJbcuO2ZMY90mgFQS7Kx2QEcWiQSjJl52xgSN9xObGN8uO1wYo90AkjLgcRtOLTIz0j/DNayH6xlDkhL+ge8Whhu5EBtkQZq2dgA0pKD3xaDGznFjDPbGIol7j9slpxx7LAxz+2cggMJBrj9AnTYZoaPbcAQ6zn+8GNPzWE59tnpmz98qLCTw6UFCv5j2I5X+SgYBaNgFIwCAgAA1/lmPtHy2DUAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-2563-5761","institution":"University of Maryland","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Sharon","middleName":"","lastName":"Bewick","suffix":""},{"id":593692,"identity":"dd23bd0d-76fc-4673-825f-58b63bc20b6f","order_by":2,"name":"Xianghui Dong","email":"","orcid":"","institution":"University of Maryland","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xianghui","middleName":"","lastName":"Dong","suffix":""},{"id":593693,"identity":"5f863848-1381-4d4e-8c0a-a570453fabb3","order_by":3,"name":"David Karig","email":"","orcid":"","institution":"Clemson University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"David","middleName":"","lastName":"Karig","suffix":""},{"id":593694,"identity":"0c1bb480-b481-4f2e-87f3-97b45c935042","order_by":4,"name":"William F. Fagan","email":"","orcid":"","institution":"University of Maryland","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"William","middleName":"F.","lastName":"Fagan","suffix":""}],"badges":[],"createdAt":"2020-05-19 04:52:52","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-29648/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-29648/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":1153918,"identity":"f8a1809e-465d-4876-a278-342351fbd5df","added_by":"auto","created_at":"2020-05-21 21:07:53","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":36056,"visible":true,"origin":"","legend":"Average percentage of taxa passing the pre-defined (A) skin, (B) skin OR human OR other mammal OR eye OR oral OR ear/nose/throat OR female reproductive systems OR male reproductive systems OR general reproduction OR lymph circulatory system OR nervous system OR bone/muscle OR liver/urinary tract/pancreas OR gut/digestive system (C) built and (D) built OR fresh water OR salt water OR general water filters for each of our 50 sets of samples (x-axis) and paired controls (y-axis). Samples consisted of 5-11 forensic swabs rubbed against the skin at 2 cm intervals across the entire width of each person’s forehead. Paired controls consisted of forensic swabs exposed to room air before and after sampling from each person. Samples were analyzed using 16S rRNA sequencing of the V3-V4 region followed by QIIME based on 97% sequence similarity. Lists of taxa were generated by including any genus that was present at \u003e0.1% of reads in any sample/control. A greater fraction of sample (control) taxa pass the filter for any point that lies below (above) the dashed grey line in each panel.","description":"","filename":"Fig1.JPG","url":"https://assets-eu.researchsquare.com/files/rs-29648/v1/Fig1.JPG"},{"id":1153920,"identity":"a347d0e3-d68f-4bee-8034-9cc5178984cf","added_by":"auto","created_at":"2020-05-21 21:07:53","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":36662,"visible":true,"origin":"","legend":"Confusion matrices showing the percentage of taxa passing each pre-defined environmental filter (x-axis) against each sample (y-axis) considering samples individually (left panel), or averaged over studies/sources (right panel). Red indicates that the majority of taxa passed the filter; blue indicates that the majority of taxa did not pass the filter. Filters are ordered such that those coming from the same broad class (e.g., environment, animal) are next to each other. Strong EFILTER performance is indicated by warm colors along the diagonal and cool colors in off-diagonal regions.","description":"","filename":"Fig2.JPG","url":"https://assets-eu.researchsquare.com/files/rs-29648/v1/Fig2.JPG"},{"id":1153921,"identity":"8ec97381-3a5d-41d7-84ea-34e7d4399430","added_by":"auto","created_at":"2020-05-21 21:07:54","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":42596,"visible":true,"origin":"","legend":"Confusion matrices showing the percentage of species (left panels) and genera (right panels) passing each pre-defined environmental filter (x-axis) against each sample (y-axis) considering samples individually (A), or averaged over body regions (B). Red indicates that the majority of taxa passed the filter; blue indicates that the majority of taxa did not pass the filter. For this dataset, FASTA files were downloaded directly from the HMP ftp server on 07/10/2016 (ftp://public-ftp.ihmpdcc.org/HMGI/). At the time of download, 690 samples were available from 15 body sites. Because of the small number of samples from mid-vagina and vaginal introitus, we did not consider vaginal sites separately. Likewise, we pooled left and right retroauricular creases, leaving 12 distinct body sites. FASTA files were analyzed for taxonomy using the default settings in MetaPhlAn2.","description":"","filename":"Fig3.JPG","url":"https://assets-eu.researchsquare.com/files/rs-29648/v1/Fig3.JPG"},{"id":1153922,"identity":"d26df273-6b50-4d76-bd1a-8e822c3c1f9a","added_by":"auto","created_at":"2020-05-21 21:07:54","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":42464,"visible":true,"origin":"","legend":"Comparison of Kraken and MetaPhlAn taxon lists generated from the FASTA files in [28] for both genera (A) and species (B). For this analysis, we only considered taxa that were present in at least one sample at levels above the defined threshold percentage of reads (in brackets). Full taxon lists were generated by pooling the taxa in all samples. For Kraken, a reference database was constructed using the complete genomes in RefSeq for the bacterial (2,199 taxonomic IDs), archaeal (165 taxonomic IDs), and viral (4,011 taxonomic IDs) domains, as well as eight representative fungal taxonomic IDs, the Plasmodium falciparum 3D7 genome, the human genome, and the UniVec Core database (ftp://ftp.ncbi.nlm.nih.gov/pub/UniVec). Low complexity regions of the microbial reference sequences were masked using the dustmasker program with a DUST level of 20 [http://www.ncbi.nlm.nih.gov/pubmed/16796549]. After masking, every 31-mer nucleotide sequence present in the collection of reference FASTA sequences is stored at the taxonomic ID of the lowest common ancestor among the leaf nodes that share that 31-mer. MetaPhlAn was run using the default settings in metaphlan2. Results are shown for the EFILTER pre-defined logical combination human OR other mammal OR eye OR oral OR ear/nose/throat OR female reproductive systems OR male reproductive systems OR general reproduction OR lymph circulatory system OR nervous system OR bone/muscle OR liver/urinary tract/pancreas OR gut/digestive system.","description":"","filename":"Fig4.JPG","url":"https://assets-eu.researchsquare.com/files/rs-29648/v1/Fig4.JPG"},{"id":1153923,"identity":"610a1357-07b8-4df4-a2d9-d019049f2ebb","added_by":"auto","created_at":"2020-05-21 21:07:54","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":27381,"visible":true,"origin":"","legend":"Pairwise combinations of samples from the tara oceans dataset, where black indicates that the sample with the highest diversity depends on whether taxa failing the salt water/fresh water/general water filter are included, and white indicates that it does not. Data were downloaded as lists of genera directly from the MG-RAST website.","description":"","filename":"Fig5.JPG","url":"https://assets-eu.researchsquare.com/files/rs-29648/v1/Fig5.JPG"},{"id":13532129,"identity":"748b132d-3253-47ab-aa4e-d15d586f3a07","added_by":"auto","created_at":"2021-09-17 01:15:48","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1130401,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-29648/v1/72d41448-89c9-4601-a813-e282216214f5.pdf"},{"id":1153919,"identity":"68189b52-6b27-436a-beb3-39303600833d","added_by":"auto","created_at":"2020-05-21 21:07:53","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":261517,"visible":true,"origin":"","legend":"","description":"","filename":"BewicketalCL.docx","url":"https://assets-eu.researchsquare.com/files/rs-29648/v1/BewicketalCL.docx"},{"id":1153917,"identity":"d4cd99a8-021b-4c2f-ad9f-2cef6b597457","added_by":"auto","created_at":"2020-05-21 21:07:53","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":66933,"visible":true,"origin":"","legend":"","description":"","filename":"AdditionalFile1.docx","url":"https://assets-eu.researchsquare.com/files/rs-29648/v1/AdditionalFile1.docx"}],"financialInterests":"","formattedTitle":"\u003cp\u003eEfilter: A Tool for Identifying Unexpected and Erroneous Taxids in Sequencing Data\u003c/p\u003e","fulltext":[{"header":"Background","content":"\u003cp\u003eThe ability to extract, clone and analyze DNA directly from environmental samples has revolutionized disciplines ranging from biomedical science to environmental microbiology[1-3].\u0026nbsp; Compared to historic, culture-based approaches,sequencing methods offer unprecedented advancements in terms of cost, effort and speed.\u0026nbsp; They also improve detection of full community membership by exposing unculturable organisms [4].\u0026nbsp; However, sequencing studies are not without their own set of difficulties.One of the greatest challenges lies in interpreting sequencingresults, which can be complicated by issues of contamination [5-7], sequencing error [8, 9], and bioinformatics misidentification[10].\u0026nbsp; Although improved laboratory and bioinformatics techniques canhelp to alleviate these difficulties, increased focus on detecting rare taxa, along with more specific classification down to species and even strains, means that interpretation of sequencing data will remain a significant challenge for the foreseeable future.\u003c/p\u003e\n\u003cp\u003eContamination is perhaps the largest and most difficult to address source of confusion in sequencing data.\u0026nbsp; Whereasculture-based methods only detect organisms that can be reliably and repeatedly grown from an environmental sample, sequencing techniques identify even small pieces of non-viable DNA.\u0026nbsp; While this is one of the strengths of the sequencing approach, and the main reason that it can uncover unculturable organisms, thisis also why sequencing methods arefar more susceptible to issues with contamination[5-7].\u0026nbsp; Unfortunately, contamination can occur at any step along the processing pipeline, ranging from issues during sample collection and storage [11], though reagent contamination[12-15] to room source contamination andcontaminating bacteria from the mouth and skin of the researcher during sample preparation [16, 17].While it is relatively straightforward to detect contaminants in samples that should contain a single organism, for example microbial DNA in the cow genome [5], or contamination of pure bacterial cultures with other bacterial taxa [18], identifying contamination in complex,mixed microbial communities presents a more difficult, and still open challenge.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; Another major issue with sequencing data is imperfect taxonomic assignment.\u0026nbsp; This challenge is also not easy to overcome, since there is an inherent trade-off between finding all taxa within a sample, and confidence/accuracy in taxon prediction [10].\u0026nbsp; Thus, bioinformatics methods that are highly sensitive are also more likely to report organisms not actually present.\u0026nbsp; Meanwhile, more conservative classification methods can reduce the number of false positives, but do so at a cost of increasing false negatives. \u0026nbsp;Misclassificationcan be more or less problematic, depending on the microbial community involved, the type of sequencing method employed[19-21], the bioinformatics toolkit used[10, 22, 23], the reference databases available [24],and the quality (e.g., sequencing platform statistics, read depth, and read length) of the sequencing data[25-27].As with contamination, even in simple systems with appropriately chosen sequencing and bioinformatics methods, classification errors can have a significant impact on conclusions, for example by massively inflating diversity or the relative importance of rare organisms.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; One method for identifying potentially problematic taxIDs in sequencing data is to rely on knowledge of microbial ecology.\u0026nbsp; Finding a common plant pathogen in a human gut microbiome dataset, for example, might raise a red flag.\u0026nbsp; Whether consciously or not, researchers use this approach to clean their data, for instance by deciding whether or not to trust all of the output taxa from a particular experiment [11].\u0026nbsp; Unfortunately, use of such information requires extensive knowledge of microbial behavior and characteristics \u0026ndash; knowledge that is lacking among many, if not most researchers performing metagenomics analysis.\u0026nbsp; Indeed, even for researchers with strong backgrounds in microbial ecology, the vast diversity of microbes from different environments means that it is impossible for any single researcher to be familiar with all common microbial habitats.\u0026nbsp; Consequently, even with extensive knowledge of the ecologies of individual bacterial taxa, data verification based on microbial ecology is likely to be biased and \u003cem\u003ead hoc\u003c/em\u003e.\u0026nbsp; Nevertheless, if such verification can be standardized and automated, it remains a promising method for addressing the challenges of spurious taxIDs in microbiome datasets.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; In this paper, we present EFILTER (\u003ca href=\"https://efilter.shinyapps.io/EFilter-app/\"\u003ehttps://efilter.shinyapps.io/EFilter-app/\u003c/a\u003e)\u0026ndash; a computational tool that performs\u003cstrong\u003e\u003cem\u003eE\u003c/em\u003e\u003c/strong\u003e\u003cem\u003environmental \u003cstrong\u003eFILTER\u003c/strong\u003eing \u003c/em\u003eby leveraging known relationships between bacterial taxa and specific habitats in order to identify taxIDs that are surprising or unlikely given the source of a metagenomics sample.\u0026nbsp; EFILTER incorporates all habitat information available in\u003cem\u003eBergey\u0026rsquo;s Manual of Systematic Bacteriology\u003c/em\u003e and \u003cem\u003eNational Center for Biotechnology Information \u003c/em\u003e(NCBI) records from the \u003cem\u003eInternational Journal of Systematic and Evolutionary Microbiology\u003c/em\u003e(IJSEM).\u0026nbsp; As such, EFILTER allows users to assess the likelihood of taxIDs being in their datasets based on known microbial ecology.\u0026nbsp; Further, it does so without requiring an extensive background in microbial ecology, and also without introducing bias that might otherwise arise based on user experience and knowledge of specific taxonomic groups.\u0026nbsp; To demonstrate the strength and performance of EFILTER, we test our tool on 4 datasets,which include both 16S rRNA and shotgun sequencing methods, analyzed with different bioinformatics pipelines and studied at different taxonomic ranks.\u0026nbsp; Our analysis shows that EFILTER can reliably provide lists of unexpected/spurious organisms in metagenomics samples, offering a new,ecologically motivated tool for validation of microbiome datasets.\u0026nbsp;\u003c/p\u003e"},{"header":"Materials And Methods","content":"\u003ch2\u003eThe EFilter Database\u003c/h2\u003e\n\u003cp\u003eWe constructed a database of microbe-environment associations using information from \u003cem\u003eBergey\u0026rsquo;s Manual of Systematic Bacteriology\u003c/em\u003e and the National Center for Biotechnology Information (NCBI) Nucleotide database.\u0026nbsp; For \u003cem\u003eBergey\u0026rsquo;s\u003c/em\u003e, we converted pdf files of Volumes IIB, IIC, III, IV, and V as well as additional chapters on \u003cem\u003eAquificae\u003c/em\u003e, \u003cem\u003eChlorobi\u003c/em\u003e, \u003cem\u003eChloroflexi\u003c/em\u003e, \u003cem\u003eChrysiogenetes\u003c/em\u003e,\u003cem\u003e Crenarchaeota\u003c/em\u003e, \u003cem\u003eCyanobacteria\u003c/em\u003e,\u003cem\u003e Deferribacteres\u003c/em\u003e, \u003cem\u003eDeinococcus-Thermus\u003c/em\u003e, \u003cem\u003eEuryarchaeota\u003c/em\u003e, \u003cem\u003eNitrospirae\u003c/em\u003e, \u003cem\u003eThermodesulfobacteria\u003c/em\u003e,\u003cem\u003e Thermomicrobia\u003c/em\u003e, and \u003cem\u003eThermotogae \u003c/em\u003eto .txt files using the pdfminer functions in Python.\u0026nbsp; The text was then separated into segments for each taxonomic group.\u0026nbsp; Text describing genera and species was mined for environmental information.\u0026nbsp; We did not use text describing higher taxonomic groups, since this information tends to be too broad to be useful.\u0026nbsp; To identify sentences containing environmental information within genus and species descriptions, we used the keywords in Table A.1 (see Appendix A).\u0026nbsp; For the NCBI Nucleotide database, we only used entries with IJSEM[Journal] in the journal field.\u0026nbsp; Our assumption was that all of these entries represent newly defined taxa.\u0026nbsp; Although taxa can be defined in other journals, the \u003cem\u003eInternational Journal of Systematic and Evolutionary Microbiology\u003c/em\u003e (IJSEM) represents one of the largest publishers of new taxon descriptions.\u0026nbsp; Most other journals do not specialize in taxonomy, and thus may contain other types of studies where taxonomic specification is not precise/accurate.\u0026nbsp; Our goal was to avoid including habitat information from inaccurate sources, since this is the very problem that EFILTER seeks to correct.\u0026nbsp; Habitat information was taken from the \u0026lsquo;isolation source/\u0026rsquo; field of the NCBI entries.\u0026nbsp; NCBI/IJSEM entries were only used for taxa that did not exist in \u003cem\u003eBergey\u0026rsquo;s \u003c/em\u003e(i.e., taxa that were defined since publication of the most recent \u003cem\u003eBergey\u0026rsquo;s \u003c/em\u003evolume for that taxonomic group).\u003c/p\u003e\n\u003cp\u003eHabitat descriptions from \u003cem\u003eBergey\u0026rsquo;s Manual of Systematic Bacteriology\u003c/em\u003e and NCBI were stored in tab-delimited .txt files (see Appendix B), along with relevant taxonomic information, forming the basis of our environmental database.\u0026nbsp; For species, all searches are performed on species-level descriptions harvested from either \u003cem\u003eBergey\u0026rsquo;s\u003c/em\u003e or NCBI/IJSEM.\u0026nbsp; For genera, searches are performed on genus-level descriptions from \u003cem\u003eBergey\u0026rsquo;s\u003c/em\u003e, as well as all species-level descriptions for species within the genus.\u0026nbsp; For higher taxonomic ranks, searches are performed on all genus-level and species-level descriptions within the particular taxonomic group.\u003c/p\u003e\n\u003ch2\u003eUser-Defined Search Terms\u003c/h2\u003e\n\u003cp\u003eThe entire environmental database can be queried for keywords.\u0026nbsp; In order to minimize search time, EFILTER contains a look-up table linking each unique keyword from the database (see above) to the species/genera in whose descriptions that keyword appears.\u0026nbsp; That is, we built a table of keywords wherein each keyword (column) maps to many rows, with each row being a specific species/genus in the database.\u0026nbsp; All keywords in our look-up table are full, stand-alone words.\u0026nbsp; However, in many cases, users may be interested in related words (e.g., \u0026lsquo;dog\u0026rsquo; and \u0026lsquo;dogs\u0026rsquo;) as well.\u0026nbsp; For this reason, whenever a keyword is typed into EFILTER, EFILTER automatically generates a list of all larger, stand-alone words (\u0026lsquo;\u003cstrong\u003edog\u003c/strong\u003es, \u003cstrong\u003edog\u003c/strong\u003e-bite\u0026rsquo;) containing the inputted keyword.\u0026nbsp; This enables users to select the precise keyword combinations that they want to include.\u0026nbsp; Thus, for the \u0026lsquo;dog\u0026rsquo; example, the user may want to include \u0026lsquo;dogs\u0026rsquo;, \u0026lsquo;dog-bite\u0026rsquo;, and \u0026rsquo;dogbite\u0026rsquo; , but not \u0026lsquo;peptidoglycan\u0026rsquo;, \u0026lsquo;endogenous\u0026rsquo;, etc.\u0026nbsp; Finally, the environmental database can be queried for phrases (e.g., \u0026lsquo;dogs and cats\u0026rsquo;).\u0026nbsp; In this case, however, there is no look-up table, and the phrase is searched for directly through the database.\u0026nbsp; Consequently, searches that involve querying for phrases are somewhat slower.\u003c/p\u003e\n\u003ch2\u003ePre-defined Environments\u003c/h2\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; Although possible, it may be hard and time-consuming for individual users to think of all potential keywords associated with a particular microbial environment.\u0026nbsp; For this reason, we pre-defined keyword combinations for 46 common habitats falling within six broad classes (animal, body sites, food, plant, environment and specialized).\u0026nbsp; To generate keyword combinations, we manually curated all keywords appearing in \u0026sup3;5 species/genus descriptions and then went through these keywords individually to decide whether they or any of their word derivatives fell into our common environmental categories.\u0026nbsp; We believe that our pre-defined environment feature will be the most useful for the majority of users.\u0026nbsp; Lists of keyword combinations used for each pre-defined environment (ranging from 6 words for \u0026lsquo;sand\u0026rsquo; to 199 words + derivatives for \u0026lsquo;other mammals\u0026rsquo;) can be found in Appendix C.\u0026nbsp;\u003c/p\u003e"},{"header":"Results","content":"\u003ch2\u003eGeneral Usage\u003c/h2\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; The EFILTER database contains 11633 species and 34 phyla (see Table B.1, Appendix B).\u0026nbsp; 72% of species have habitat information specified, while 45% fall into at least one pre-defined environment (see Methods). 100% of phyla have habitat information specified, while 97.1% fall into at least one pre-defined environment.\u0026nbsp; The pre-defined environment with the largest number of taxa is \u0026lsquo;soil\u0026rsquo;, which includes 17.5% of species.\u0026nbsp; The next largest is \u0026lsquo;human\u0026rsquo;, which includes 8.4% of species. The pre-defined environments with the fewest number of taxa are \u0026lsquo;amphibian\u0026rsquo;, \u0026lsquo;reptile\u0026rsquo; and \u0026lsquo;symbiont\u0026rsquo;, with 0.06%, 0.3% and 0.3% of species respectively (see Appendix B, Table B.2).\u003c/p\u003e\n\u003cp\u003eEFILTER inputs .txt files (including batch upload) in the form of lists of either Latin names or taxIDs. Inputs in the form of taxIDs are mapped onto Latin names in order to infer habitat associations.\u0026nbsp; Only well defined taxa (i.e., taxa with entries in either \u003cem\u003eBergeys\u003c/em\u003e or NCBI/IJSEM, see Methods)areincluded in the database.\u0026nbsp; EFILTER allows Latin names or taxIDs referencing any taxonomic rank from phylum to species, including lists with mixed ranks.\u0026nbsp; Taxonomic rank is automatically determined based on name or taxID.\u0026nbsp; Each taxon in an input file is queried against the EFILTER database using a defined environmental filter.\u0026nbsp; Environmental filters can be constructed as user-defined keywords, user-defined phrases, or pre-defined environments (see Methods).\u0026nbsp; In addition, users can construct any logical combination (\u0026lsquo;AND\u0026rsquo;, \u0026lsquo;OR\u0026rsquo;, \u0026lsquo;NOT\u0026rsquo;, for example \u0026lsquo;eye AND human NOT other mammal\u0026rsquo;) of single filters.\u0026nbsp; A video tutorial covering filter definition can be found at (\u003ca href=\"https://www.youtube.com/watch?v=UU1rPTUZPTE\"\u003ehttps://www.youtube.com/watch?v=UU1rPTUZPTE\u003c/a\u003e).\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; Upon file upload, EFILTER outputs the number of taxa at each taxonomic rank as well as the number of taxa that were not found in the EFILTER database. Missing taxa include taxa that are not well-defined, taxa that were defined more recently than the most recent EFILTER update, or taxa that are not bacterial (for example fungal taxa from the output of shotgun sequencing).\u0026nbsp; These taxa are ignored, meaning that there is no need to curate shotgun sequencing data to include strictly bacteria.\u0026nbsp; EFILTER also automatically reports the pre-defined environmental classes for each unique taxon in an uploaded dataset, starting from the input taxonomic rankup to the taxonomic rank of family.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; The primary goal of EFILTER is to identify organisms that are consistent/inconsistent with the sampled environment (i.e., the user selected filter).\u0026nbsp; Thus, when a taxon list is queried against a chosen filter, EFILTER outputs the number of taxa with at least one habitat in the database that matches (passes) the filter.Taxa without at least one matching habitat are identified as failing the filter.\u0026nbsp; In addition, for each input taxon, EFILTER reports whether the taxon passed or failed at taxonomic ranks higher than the inputted rank.\u0026nbsp; Finally, EFILTER outputs the names of taxa that passed the filter at taxonomic ranks lower than the inputted rank.\u0026nbsp; This latter feature can be useful, for example, to make an educated guess as to which species might be present given a list of genera from a 16S rRNA sequencing run.\u003c/p\u003e\n\u003ch2\u003eSamples versus Controls\u003c/h2\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; As a first demonstration of EFILTER performance, we used our own data, which consisted of 16S rRNA sequencing of 375skin microbiome samples from the foreheads of 50 individuals, along with 100paired controls (blank swabs collected before and/or after each person was sampled).\u0026nbsp; Figure 1 shows the percentage of genera passing four different filters averaged over the samples and paired controls for each person.\u0026nbsp; For the \u0026lsquo;skin\u0026rsquo; filter (A), more organisms passed in the samples relative to the controls for the majority (48/50 or 96%) of people.\u0026nbsp; By contrast, for the \u0026lsquo;built\u0026rsquo; filter (C) more organisms passed in the controls relative to the samples for the majority (44/50 or 88%) of people.\u0026nbsp; As expected, for logical OR combinations of filters, the percentages of passing taxa were greater overall (compare, for example Figures 1A and 1B or Figures 1C and 1D).\u0026nbsp; However, once again, for the \u0026lsquo;human/mammal/body-site\u0026rsquo; filter (C) more taxa passed in the samples relative to the controls for the majority (40/50 or 80%) of people, whereas the reverse was true for the \u0026lsquo;built/water\u0026rsquo; filter (D) for the majority (30/50 or 60%) of people.\u0026nbsp; In general, one would expect that a skin/human/mammal/body-site filter would perform better on human samples than on non-human (e.g., control) samples, which is what we see.\u0026nbsp; Likewise, assuming that many contaminants come from room air or sequencing preparation steps, one would assume that a built/water filter would perform better on control samples than on human samples, which is again what we see.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eSamples from Different Sources\u003c/h2\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; As a second demonstration of EFILTER performance, we comparedvarious EFILTER pre-defined environmental filters across a variety of environmental samples.\u0026nbsp; Datasets were downloaded as lists of genera directly from the MG-RAST website (\u003ca href=\"https://www.mg-rast.org/\"\u003ehttps://www.mg-rast.org/\u003c/a\u003e).\u0026nbsp; Specifically, we considered hot springs (mgp5265, mgp6907, mgp5356, mgp5355, mgp5270, mgp5269, mgp5268, mgp5266), tropical soil (mgp4362), a freshwater lake (mgp19525), oceans (mgp20413), cheese (mgp14606), human nares (mgp385), and horses guts (mgp7746).\u0026nbsp; All studies involved shotgun sequencing.\u0026nbsp; We used the same number samples from each environment (8,limited by the availability of hot springs samples; in all cases we selected the 8 samples with the lowest MG-RAST ID numbers), and restricted our analysis to the top 25 most abundant taxa in each sample.\u0026nbsp; Figure 2 shows confusion matrices for samples versus environmental filters.Focusing on the averaged confusion matrix (right panel), we see strong EFILTER performance, with red/orange (high percentage of passing taxa) concentrated along the diagonal and blue/green (low percentage of passing taxa) off-diagonal.\u0026nbsp; Overall, across the seven paired environments and samples, an average of 72% of taxa passed their expected filters.\u0026nbsp; Passing rates were highest (88.5%) for the anterior nares, and lowest (62.5%) for the lake.\u003c/p\u003e\n\u003cp\u003eFilter fidelity to sample was quite high.\u0026nbsp; For five out of seven samples (hot springs, soil, oceans, cheese and anterior nares), the expected filter (extreme, soil, salt water, dairy, human) performed best.\u0026nbsp; Even when this was not the case, the expected filter still performed well.\u0026nbsp; For the lake study, for instance, both the salt water and soil filters outperformed the fresh water filter (note that all three are in the\u0026lsquo;environment\u0026rsquo; class, see Methods), though only marginally (63.5% and 62.8% of taxa passing for the salt water and soil filters, relative to 62.5% for the fresh water filter).\u0026nbsp; Likewise, for the horse gut study, the human filter outperformed the other mammal filter (note that both are in the \u0026lsquo;animal\u0026rsquo; class), although once again, the other mammal filter still performed well (79.4% passing for the human filter relative to 75.9% for the other mammal filter).\u0026nbsp; Sample fidelity to filter was not as good.\u0026nbsp; For three out of seven filters (extreme, soil and human) the highest percentages of passing taxa were identified in matched samples (hot springs, soil, and anterior nares respectively).\u0026nbsp; For the remaining four filters, however, the highest percentages of passing taxa were identifiedin non-matched samples (the fresh and salt water filters identified the highest percentage of passing taxa in the soil sample, while the dairy and other mammal filtersidentified the highest percentage of passing taxa in the anterior nares).\u003c/p\u003e\n\u003ch2\u003eSamples from Different Body Sites\u003c/h2\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; As a third demonstration of EFILTER performance, we considered a shot-gun sequencing dataset from the Human Microbiome Project (HMP), including 690 samples from 15 body-sites.\u0026nbsp; Importantly, in this dataset all samples fall under a single filter class (\u0026lsquo;body-site\u0026rsquo;, see Methods).\u0026nbsp; As in the previous section, we only considered the top 25 most abundant taxa in each sample. Figure 3 shows confusion matrices for species- and genus-level data for each sample individually (A), and as sample averages over each body region (B; notice that, unlike Figure 2, there are different numbers of samples contributing to each body region).\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; EFILTER performance on the HMP dataset was even higher than on the environmental dataset, with an average of 40.9% of species and 76.0% of genera passing the appropriately matched filter.\u0026nbsp; At the same time, however, filter fidelity to sample was lower (compare the right panel in Figure 3B to Figure 2).\u0026nbsp; That is, at the rank of both genus and species, the expected filter (gut/digestive system, oral) performed best for only two (St, To+Kg+Bm+Sa+Sb+Sp+Td) out of five body regions.\u0026nbsp; For the remaining three regions, the gut/digestive system and/or oral filters performed better.\u0026nbsp; The gut/digestive system filter likely performed well because it has the largest number of taxa (see Appendix B).\u0026nbsp; It is less clear why the oral filter out-performed the other filters, since it has fewer taxa than the ear/nose/throat filter at the ranks of both genus and species and has approximately the same number of taxa as the skin filter, atthe rank of genus.\u003c/p\u003e\n\u003cp\u003eHMP sample fidelity to filter, on the other hand, was strong, at least at the rank of species.\u0026nbsp; In particular, four out of five filters (skin, oral, gut/digestive system, female reproductive system) identified the highest percentage of passing taxa in their matched samples (Rc, To+Kg+Bm+Sa+Sb+Sp+Td, St, Va).\u0026nbsp; Sample fidelity to filter was slightly weaker at the rank of genus, where only two out of five filters (oral, gut/digestive system) identified the highest percentage of passing taxa in their matched samples.\u0026nbsp; This fidelity, was similar to the fidelity observed in our environmental dataset (see Figure 2).One particularly striking feature in Figure 3 is the poor performance of all but the gut/digestive system filter on stool samples.\u0026nbsp; Whereas \u0026gt;56% of genera from every other body region passed every body-site filter, only 17-35% of stool genera passednon-gut/digestive system filters.\u0026nbsp; At the same time, however, 95% of stool genera passed the gut/digestive system filter.\u0026nbsp; This suggests that stool taxa are unique to the gut environment, whereas taxa from other body sites exhibit broad body distributions.\u003c/p\u003e\n\u003ch2\u003eSamples from Different Bioinformatics Pipelines\u003c/h2\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; As a final demonstration of EFILTER performance, we compared EFILTER output for different bioinformatics pipelines.\u0026nbsp; A common problem with microbiome studies is that different pipelines can give different taxon lists, including substitutions among distantly related organisms.\u0026nbsp; EFILTER can be used to gauge performance of different bioinformatics methods and to identify potentially spurious taxa introduced through taxonomic assignment steps.\u0026nbsp; To illustrate this, we used a shotgun sequencing dataset of the human skin microbiome [28], processed using both Kraken [29] and MetaPhlAn [30], and thresholded at various different read percentages.\u0026nbsp; Results are shown in Figure 4, which also serves as an example of the EFILTER graphical output. \u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWith larger thresholds, the pass:fail ratio increases for both bioinformatics pipelines.\u0026nbsp; In other words, a larger fraction of the rare tail fails, regardless of the bioinformatics method used.\u0026nbsp; This should not be surprising.\u0026nbsp; Everything from contaminants and transient taxa to inherently rare/understudied organisms and bioinformatics errors are expected to contribute to the rare tail.\u0026nbsp; In keeping with the consensus that Kraken has a tendency to over classify [10, 31], the pass:fail ratio for Kraken is significantly lower than it is for MetaPhlAn.\u0026nbsp; This is particularly true when there is no threshold, with Kraken givingpass rates of 42% and35% for genera and species respectively relative to MetaPhlAn\u0026rsquo;s 67% and 61% at the same ranks. With a 1% threshold, both pipelines givesimilar results for species (72% passing), while Kraken actually shows a higher pass:fail ratio for genera (92% versus 88%).\u0026nbsp; Surprisingly, though, despite Kraken\u0026rsquo;s lower pass:fail ratio for most scenarios, MetaPhlAn usuallyidentifies a greater absolute number of passing taxa.\u0026nbsp; Thus, while it is likely that some of the failing taxa identified by Kraken (and MetaPhlAn) are truly present on skin,based on overall performance as judged by environmental consistency of assigned taxa, MetaPhlAn appears to be the better pipeline for this dataset.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eA wealth of microbiome research has emerged over the past decade.\u0026nbsp; Most of it, however, has been plagued by contamination, sequencing errors and taxonomic misidentification[6-10].\u0026nbsp; This has led to a lack of reproducibility across pipelines and amongst labs.\u0026nbsp; Recently, there has been a call for improved standards and validation of microbiome datasets, including efforts like the MicroBiome Quality Control (MBQC) project [15] and the Critical Assessment of Metagenome Interpretation (CAMI) initiative (\u003ca href=\"http://microbiome-cosi.org/cami\"\u003ehttp://microbiome-cosi.org/cami\u003c/a\u003e), as well as standards development by the National Institute of Standards and Technology (NIST) [32].\u0026nbsp; In this paper, we introduce EFILTER as a new tool for helping to assess the quality of microbiome datasets and for identifying suspicious taxIDs within them.\u0026nbsp; EFILTER is unique amongst microbiome informatics approaches in that it leverages ecological information to target organisms that are unexpected based on sample source.\u003c/p\u003e\n\u003cp\u003eUltimately, EFILTER provides lists of suspicious taxIDs.\u0026nbsp; The question, of course, is what to do with this information.\u0026nbsp; In general, we advocate against indiscriminant dismissal of taxa that fail a chosen filter.\u0026nbsp; Rather, two approaches are possible.\u0026nbsp; First, depending on the analysis being performed, analysis can be run with and without discarding the problematic taxa in order to determine whether any conclusions change and, if they do, which and how many.\u0026nbsp; In Figure 5, for example, we show how conclusions about relative sample diversity differ depending on whether analysis includes all taxa or only those taxa passing the salt water OR fresh water OR general water filter for the tara ocean dataset from MG-RAST (mgp20413).\u0026nbsp; Notably, 93.6% of diversity orderings remain unchanged when failing taxa are ignored, suggesting that suspicious taxa are not overly problematic for the conclusions of thisparticular analysis with this particular dataset.\u003c/p\u003e\n\u003cp\u003eA second approach for using EFILTER data is to carefully examine taxa that fail and to make decisions about whether to include them based on additional knowledge or further experimentation.\u0026nbsp; Depending on the system, the list of failing taxa can still be large.\u0026nbsp; Thus, we recommend using broad filters and focusing on particularly egregious failures.\u0026nbsp; That is, either failures that involve very abundant taxa, or else failures that extend to higher taxonomic ranks.\u0026nbsp; We illustrate such an approach in Table C.1 of Appendix C for our own skin microbiome dataset.\u0026nbsp; Notably, although we cannot be certain about the source of any taxon failures, we can make educated guesses about a sizeable fraction.\u0026nbsp; Some taxa, for example \u003cem\u003eGeodermatophilus\u003c/em\u003eand \u003cem\u003eMethylobacterium\u003c/em\u003e, are known contaminants of blanks from other sequencing studies, suggesting a contamination origin.\u0026nbsp; Others, for example \u003cem\u003eRhodothermus\u003c/em\u003e and \u003cem\u003eRheinheimera \u003c/em\u003eare present in samples and controls taken at the same time or are ubiquitous in controls, again suggesting a contamination origin.\u0026nbsp; Our lab has cultured certain suspicious taxa, for example \u003cem\u003eEnhydrobacter\u003c/em\u003e, directly from skin, suggesting an unreported habitat for this organism.\u0026nbsp; Finally, some taxa, including\u003cem\u003eModestobacter\u003c/em\u003e and \u003cem\u003eHymenobacter\u003c/em\u003e, are found in a range of samples and in no controls, hinting that they may be real, as yet undiscovered taxa from human skin.\u0026nbsp; Supporting this claim, both \u003cem\u003eModestobacter\u003c/em\u003e and \u003cem\u003eHymenobacter\u003c/em\u003e have been found in other human microbiome datasets.\u0026nbsp; \u003cem\u003eModestobacter\u003c/em\u003e, in particular, would be interesting to explore further.\u0026nbsp; Although it is currently known from extreme (desert), rock/stone and soil environments it showed up in both 16S amplicon sequencing from our skin samples and in shotgun sequencing from a separate skin study performed in a different lab using an entirely different sequencing and bioinformatics pipeline[28].\u0026nbsp; This suggests that the source could be an as yet unknown species of \u003cem\u003eModestobacter \u003c/em\u003ethat resides on human skin.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; Although the goal of EFILTER is to identify suspicious taxa in metagenomics samples, it is worth pointing out that EFILTER also works reasonably well for determining sample source origin, at least when the possible sources are quite distinct.\u0026nbsp; In Figure 2, for instance, it is possible to discriminate soil/water sources, animal sources and extreme sources based on the percentage of abundant (top 25) genera passing the associated filters.\u0026nbsp; Such discrimination is more difficult for closely related sources, for example for sources from different mammals, sources from different types of water, or sources from different human body regions (see Figure 3).\u0026nbsp; A future goal would be to extend EFILTER such that source discrimination is improved, possibly by leveraging additional habitat information from sequencing studies.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; On a similar note, one of the downsides of the existing EFILTER database is that it only uses habitat information from validly published taxa descriptions.\u0026nbsp; Although this ensures that contaminated or otherwise compromised metagenomics samples do not influence the EFILTER database, it means that EFILTER does not capitalize on the widely available, though imperfect, sequencing data currently in the public domain.\u0026nbsp; A future direction would be to extend EFILTER to include these data, but to do so in a manner that incorporates confidence in the data source.\u0026nbsp; This could result in higher percentages of taxa passing the appropriate filters, and also improved source discrimination.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eContamination, sequencing error and bioinformatics misidentifications will continue to plague microbiome research for the foreseeable future.\u0026nbsp; Historically, many questionable results from sequencing studies have been addressed \u003cem\u003ead hoc\u003c/em\u003e by researchers who realize that particular organisms should not be present in particular samples.\u0026nbsp; With EFILTER, we extend this capability to anyone involved in microbiome research, and do so in a way that allows rigorous, systematic, and repeatabledataset cleaning and validation based on ecologically relevant habitat information.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthics approval and consent to participate:\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003ch2\u003eConsent for publication:\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003ch2\u003eAvailability of data and material:\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003eAll data generated or analysed during this study are included in this published article [and its supplementary information files].\u003c/p\u003e\n\u003ch2\u003eCompeting Interests:\u003c/h2\u003e\n\u003cp\u003eThere are no competing interests.\u003c/p\u003e\n\u003ch2\u003eFunding:\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003eThis material is based upon work supported by, or in part by, the U. S. Army Research Laboratory and the U. S. Army Research Office under contract/grant number #W911NF-14-1-0490.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAuthors' contributions:\u0026nbsp; SB conceived of the idea, built the database, and ran analyses on test datasets;\u0026nbsp; XD developed the online platform in RShiny; SB, XD, DK and WFF contributed to idea development, results interpretation, and manuscript writing.\u003c/p\u003e\n\u003ch2\u003eAcknowledgements:\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eStahl DA, Lane DJ, Olsen G, Pace NR: \u003cstrong\u003eAnalysis of hydrothermal vent-associated symbionts by ribosomal RNA sequences.\u003c/strong\u003e\u003cem\u003eScience \u003c/em\u003e1984, \u003cstrong\u003e224:\u003c/strong\u003e409-412.\u003c/li\u003e\n\u003cli\u003eAmann RI, Ludwig W, Schleifer K-H: \u003cstrong\u003ePhylogenetic identification and in situ detection of individual microbial cells without cultivation.\u003c/strong\u003e\u003cem\u003eMicrobiological reviews \u003c/em\u003e1995, \u003cstrong\u003e59:\u003c/strong\u003e143-169.\u003c/li\u003e\n\u003cli\u003eVenter JC, Remington K, Heidelberg JF, Halpern AL, Rusch D, Eisen JA, Wu D, Paulsen I, Nelson KE, Nelson W: \u003cstrong\u003eEnvironmental genome shotgun sequencing of the Sargasso Sea.\u003c/strong\u003e\u003cem\u003escience \u003c/em\u003e2004, \u003cstrong\u003e304:\u003c/strong\u003e66-74.\u003c/li\u003e\n\u003cli\u003eRiesenfeld CS, Schloss PD, Handelsman J: \u003cstrong\u003eMetagenomics: genomic analysis of microbial communities.\u003c/strong\u003e\u003cem\u003eAnnu Rev Genet \u003c/em\u003e2004, \u003cstrong\u003e38:\u003c/strong\u003e525-552.\u003c/li\u003e\n\u003cli\u003eMerchant S, Wood DE, Salzberg SL: \u003cstrong\u003eUnexpected cross-species contamination in genome sequencing projects.\u003c/strong\u003e\u003cem\u003ePeerJ \u003c/em\u003e2014, \u003cstrong\u003e2:\u003c/strong\u003ee675.\u003c/li\u003e\n\u003cli\u003eStrong MJ, Xu G, Morici L, Bon-Durant SS, Baddoo M, Lin Z, Fewell C, Taylor CM, Flemington EK: \u003cstrong\u003eMicrobial contamination in next generation sequencing: implications for sequence-based analysis of clinical samples.\u003c/strong\u003e\u003cem\u003ePLoS Pathog \u003c/em\u003e2014, \u003cstrong\u003e10:\u003c/strong\u003ee1004437.\u003c/li\u003e\n\u003cli\u003eLusk RW: \u003cstrong\u003eDiverse and widespread contamination evident in the unmapped depths of high throughput sequencing data.\u003c/strong\u003e\u003cem\u003ePloS one \u003c/em\u003e2014, \u003cstrong\u003e9:\u003c/strong\u003ee110808.\u003c/li\u003e\n\u003cli\u003eKunin V, Engelbrektson A, Ochman H, Hugenholtz P: \u003cstrong\u003eWrinkles in the rare biosphere: pyrosequencing errors can lead to artificial inflation of diversity estimates.\u003c/strong\u003e\u003cem\u003eEnvironmental microbiology \u003c/em\u003e2010, \u003cstrong\u003e12:\u003c/strong\u003e118-123.\u003c/li\u003e\n\u003cli\u003eLee CK, Herbold CW, Polson SW, Wommack KE, Williamson SJ, McDonald IR, Cary SC: \u003cstrong\u003eGroundtruthing next-gen sequencing for microbial ecology\u0026ndash;biases and errors in community structure estimates from PCR amplicon pyrosequencing.\u003c/strong\u003e\u003cem\u003ePloS one \u003c/em\u003e2012, \u003cstrong\u003e7:\u003c/strong\u003ee44224.\u003c/li\u003e\n\u003cli\u003ePeabody MA, Van Rossum T, Lo R, Brinkman FS: \u003cstrong\u003eEvaluation of shotgun metagenomics sequence classification methods using in silico and in vitro simulated communities.\u003c/strong\u003e\u003cem\u003eBMC bioinformatics \u003c/em\u003e2015, \u003cstrong\u003e16:\u003c/strong\u003e362.\u003c/li\u003e\n\u003cli\u003eDeLong EF: \u003cstrong\u003eMicrobial community genomics in the ocean.\u003c/strong\u003e\u003cem\u003eNature Reviews Microbiology \u003c/em\u003e2005, \u003cstrong\u003e3:\u003c/strong\u003e459-469.\u003c/li\u003e\n\u003cli\u003eJervis-Bardy J, Leong LE, Marri S, Smith RJ, Choo JM, Smith-Vaughan HC, Nosworthy E, Morris PS, O\u0026rsquo;Leary S, Rogers GB: \u003cstrong\u003eDeriving accurate microbiota profiles from human samples with low bacterial content through post-sequencing processing of Illumina MiSeq data.\u003c/strong\u003e\u003cem\u003eMicrobiome \u003c/em\u003e2015, \u003cstrong\u003e3:\u003c/strong\u003e19.\u003c/li\u003e\n\u003cli\u003eWoyke T, Sczyrba A, Lee J, Rinke C, Tighe D, Clingenpeel S, Malmstrom R, Stepanauskas R, Cheng J-F: \u003cstrong\u003eDecontamination of MDA reagents for single cell whole genome amplification.\u003c/strong\u003e\u003cem\u003ePloS one \u003c/em\u003e2011, \u003cstrong\u003e6:\u003c/strong\u003ee26161.\u003c/li\u003e\n\u003cli\u003eMotley ST, Picuri JM, Crowder CD, Minich JJ, Hofstadler SA, Eshoo MW: \u003cstrong\u003eImproved multiple displacement amplification (iMDA) and ultraclean reagents.\u003c/strong\u003e\u003cem\u003eBMC genomics \u003c/em\u003e2014, \u003cstrong\u003e15:\u003c/strong\u003e443.\u003c/li\u003e\n\u003cli\u003eSinha R, Abnet CC, White O, Knight R, Huttenhower C: \u003cstrong\u003eThe microbiome quality control project: baseline study design and future directions.\u003c/strong\u003e\u003cem\u003eGenome biology \u003c/em\u003e2015, \u003cstrong\u003e16:\u003c/strong\u003e276.\u003c/li\u003e\n\u003cli\u003eSalter SJ, Cox MJ, Turek EM, Calus ST, Cookson WO, Moffatt MF, Turner P, Parkhill J, Loman NJ, Walker AW: \u003cstrong\u003eReagent and laboratory contamination can critically impact sequence-based microbiome analyses.\u003c/strong\u003e\u003cem\u003eBMC biology \u003c/em\u003e2014, \u003cstrong\u003e12:\u003c/strong\u003e87.\u003c/li\u003e\n\u003cli\u003eWeiss S, Amir A, Hyde ER, Metcalf JL, Song SJ, Knight R: \u003cstrong\u003eTracking down the sources of experimental contamination in microbiome studies.\u003c/strong\u003e\u003cem\u003eGenome biology \u003c/em\u003e2014, \u003cstrong\u003e15:\u003c/strong\u003e564.\u003c/li\u003e\n\u003cli\u003eOlson ND, Zook JM, Morrow JB, Lin NJ: \u003cstrong\u003eUsing metagenomic methods to detect organismal contaminants in microbial materials.\u003c/strong\u003e PeerJ Preprints; 2017.\u003c/li\u003e\n\u003cli\u003eRanjan R, Rani A, Metwally A, McGee HS, Perkins DL: \u003cstrong\u003eAnalysis of the microbiome: advantages of whole genome shotgun versus 16S amplicon sequencing.\u003c/strong\u003e\u003cem\u003eBiochemical and biophysical research communications \u003c/em\u003e2016, \u003cstrong\u003e469:\u003c/strong\u003e967-977.\u003c/li\u003e\n\u003cli\u003eQunfeng D, Claudia V: \u003cstrong\u003eEvaluation of the RDP classifier accuracy using 16S rRNA gene variable regions.\u003c/strong\u003e\u003cem\u003eMetagenomics \u003c/em\u003e2012, \u003cstrong\u003e2012\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003eClaesson MJ, Wang Q, O'sullivan O, Greene-Diniz R, Cole JR, Ross RP, O'toole PW: \u003cstrong\u003eComparison of two next-generation sequencing technologies for resolving highly complex microbiota composition using tandem variable 16S rRNA gene regions.\u003c/strong\u003e\u003cem\u003eNucleic acids research \u003c/em\u003e2010\u003cstrong\u003e:\u003c/strong\u003egkq873.\u003c/li\u003e\n\u003cli\u003eRicke DO, Shcherbina A, Chiu N: \u003cstrong\u003eEvaluating performance of metagenomic characterization algorithms using in silico datasets generated with FASTQSim.\u003c/strong\u003e\u003cem\u003ebioRxiv \u003c/em\u003e2016\u003cstrong\u003e:\u003c/strong\u003e046532.\u003c/li\u003e\n\u003cli\u003eHall RJ, Draper JL, Nielsen FG, Dutilh BE: \u003cstrong\u003eBeyond research: a primer for considerations on using viral metagenomics in the field and clinic.\u003c/strong\u003e\u003cem\u003eFrontiers in microbiology \u003c/em\u003e2015, \u003cstrong\u003e6\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003eGolob JL, Margolis E, Hoffman NG, Fredricks DN: \u003cstrong\u003eEvaluating the accuracy of amplicon-based microbiome computational pipelines on simulated human gut microbial communities.\u003c/strong\u003e\u003cem\u003eBMC bioinformatics \u003c/em\u003e2017, \u003cstrong\u003e18:\u003c/strong\u003e283.\u003c/li\u003e\n\u003cli\u003eSchloss PD: \u003cstrong\u003eThe effects of alignment quality, distance calculation method, sequence filtering, and region on the analysis of 16S rRNA gene-based studies.\u003c/strong\u003e\u003cem\u003ePLoS Comput Biol \u003c/em\u003e2010, \u003cstrong\u003e6:\u003c/strong\u003ee1000844.\u003c/li\u003e\n\u003cli\u003eWommack KE, Bhavsar J, Ravel J: \u003cstrong\u003eMetagenomics: read length matters.\u003c/strong\u003e\u003cem\u003eApplied and environmental microbiology \u003c/em\u003e2008, \u003cstrong\u003e74:\u003c/strong\u003e1453-1463.\u003c/li\u003e\n\u003cli\u003eMosher JJ, Bowman B, Bernberg EL, Shevchenko O, Kan J, Korlach J, Kaplan LA: \u003cstrong\u003eImproved performance of the PacBio SMRT technology for 16S rDNA sequencing.\u003c/strong\u003e\u003cem\u003eJournal of microbiological methods \u003c/em\u003e2014, \u003cstrong\u003e104:\u003c/strong\u003e59-60.\u003c/li\u003e\n\u003cli\u003eOh J, Byrd AL, Deming C, Conlan S, Barnabas B, Blakesley R, Bouffard G, Brooks S, Coleman H, Dekhtyar M: \u003cstrong\u003eBiogeography and individuality shape function in the human skin metagenome.\u003c/strong\u003e\u003cem\u003eNature \u003c/em\u003e2014, \u003cstrong\u003e514:\u003c/strong\u003e59.\u003c/li\u003e\n\u003cli\u003eWood DE, Salzberg SL: \u003cstrong\u003eKraken: ultrafast metagenomic sequence classification using exact alignments.\u003c/strong\u003e\u003cem\u003eGenome biology \u003c/em\u003e2014, \u003cstrong\u003e15:\u003c/strong\u003eR46.\u003c/li\u003e\n\u003cli\u003eSegata N, Waldron L, Ballarini A, Narasimhan V, Jousson O, Huttenhower C: \u003cstrong\u003eMetagenomic microbial community profiling using unique clade-specific marker genes.\u003c/strong\u003e\u003cem\u003eNature methods \u003c/em\u003e2012, \u003cstrong\u003e9:\u003c/strong\u003e811.\u003c/li\u003e\n\u003cli\u003eClooney AG, Fouhy F, Sleator RD, O\u0026rsquo;Driscoll A, Stanton C, Cotter PD, Claesson MJ: \u003cstrong\u003eComparing apples and oranges?: next generation sequencing and its impact on microbiome analysis.\u003c/strong\u003e\u003cem\u003ePLoS One \u003c/em\u003e2016, \u003cstrong\u003e11:\u003c/strong\u003ee0148028.\u003c/li\u003e\n\u003cli\u003eStulberg E, Fravel D, Proctor LM, Murray DM, LoTempio J, Chrisey L, Garland J, Goodwin K, Graber J, Harris MC: \u003cstrong\u003eAn assessment of US microbiome research.\u003c/strong\u003e\u003cem\u003eNature microbiology \u003c/em\u003e2016, \u003cstrong\u003e1:\u003c/strong\u003e15015.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"taxonomic identification, microbial habitat association, online tool, metagenomics datasets, contamination, bioinformatics errors, novel niches, taxon discovery","lastPublishedDoi":"10.21203/rs.3.rs-29648/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-29648/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u0026nbsp;\u003c/strong\u003eMultiple types of error can enter metagenomics sample analysis, including contamination during sample collection, library preparation and sequencing, as well as incorrect taxonomic assignment by bioinformatics packages.\u0026nbsp;Often, such errors either go unidentified, or else are removed, \u003cem\u003ead hoc\u003c/em\u003e, based on user knowledge of microbial ecology.\u0026nbsp;However, becausedifferent researchers are more or less familiar with the ecologies of the various organisms in their systems, filtering is applied non-uniformly at best, with differences in the degree of filtering between studies and between taxonomic groups within a single study.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults:\u0026nbsp;\u003c/strong\u003eIn this paper, we present EFILTER – a tool that capitalizes on decades of research in microbial ecology to identify suspicious or spurious taxa in metagenomics samples based on habitat information.\u0026nbsp;\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions:\u0026nbsp;\u003c/strong\u003eEFILTER allows all microbiome researchers, regardless of background, to examine taxon lists for unusual entries, and to do so in a manner that is systematic and without bias.\u003c/p\u003e","manuscriptTitle":"Efilter: A Tool for Identifying Unexpected and Erroneous Taxids in Sequencing Data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2020-05-21 21:07:53","doi":"10.21203/rs.3.rs-29648/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e3675c1f-6a4b-48b1-98d2-edbf80676cee","owner":[],"postedDate":"May 21st, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":105030,"name":"Biotechnology and Bioengineering"}],"tags":[],"updatedAt":"2020-06-03T15:48:10+00:00","versionOfRecord":[],"versionCreatedAt":"2020-05-21 21:07:53","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-29648","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-29648","identity":"rs-29648","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.