CoMIT: A Bioinformatic Pipeline for Risk-based Prediction of COVID-19 Test Inclusivity | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article CoMIT: A Bioinformatic Pipeline for Risk-based Prediction of COVID-19 Test Inclusivity Diane M. Walker, Wendy A. Smith, Lia Gale, Jacob T. Wolff, Connor P. Healy, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5783329/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 12 Feb, 2025 Read the published version in BMC Bioinformatics → Version 1 posted 4 You are reading this latest preprint version Abstract Background The global Coronavirus Disease 2019 (COVID-19) pandemic highlighted the need to quickly diagnose infections to identify and prevent viral spread in the population. In response to the pandemic, BioFire Defense leveraged its PCR-based “lab-in-a-pouch” technology for expedited development of the BioFire® COVID-19 Test, a novel in vitro diagnostic detecting SARS-CoV-2 nucleic acid in human samples. Following clearance of an in vitro diagnostic device, regulatory bodies such as the U.S. Food and Drug Administration (FDA) require regular post market surveillance to monitor test performance against viral lineages circulating in the field, using predictive in silico inclusivity evaluations. Exponential increases in the number of sequences deposited in bioinformatic repositories such as GISAID, during the pandemic, impeded progress in meeting these post market requirements. In response, BioFire Defense developed a new bioinformatic tool to overcome scalability problems and the loss of accuracy encountered with the standard inclusivity method. Results The Coronavirus Monitoring for Inclusivity Tool (CoMIT) uses the Variant Sorter Algorithm to sidestep multiple sequence alignments, a significant barrier inherent in the standard inclusivity method. The implementation of CoMIT and its Variant Sorter Algorithm are described. Automated summary tables and visualizations from a typical inclusivity evaluation are presented. We report our approach to filter and display relevant information in the pipeline outputs using risk factors tied to test performance. Conclusions BioFire Defense has developed CoMIT, an automated bioinformatic pipeline for efficient processing and reporting of variant inclusivity from the GISAID EpiCoV™ repository. This tool ensures continuous and comprehensive post market evaluations of BioFire COVID-19 Test performance even from datasets large enough to impede standard inclusivity analyses. CoMIT’s low computational space complexity and modular code allow this tool to be generalized for inclusivity monitoring of multianalyte or single analyte tests with complex assay designs and/or highly variable targets. CoMIT’s databasing capabilities and metadata handling hold the potential for new investigations to improve readiness for future outbreaks. COVID-19 SARS-CoV-2 GISAID EpiCoV™ PCR genomic variants detection in silico inclusivity BioFire FilmArray System in vitro diagnostic device bioinformatics tool Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Background The Coronavirus Disease 2019 (COVID-19) pandemic brought challenges never faced in the modern era, necessitating accelerated timelines and processes to address urgent public health needs. The global outbreak spawned the rapid development of diagnostic tests capable of detecting the presence of SARS-CoV-2 (the etiological agent of COVID-19) in human samples ( 1 – 6 ). In response, BioFire Defense leveraged its existing BioFire® FilmArray® PCR-based technology to develop a test specific for the identification of SARS-CoV-2 nucleic acid from patient samples. On March 24, 2020, BioFire Defense received initial Emergency Use Authorization (EUA) by the U.S. Food and Drug Administration (FDA) for the BioFire® COVID-19 Test. A 510k clearance was granted for the BioFire® COVID-19 Test 2 on November 1, 2021, which has identical chemistry to the EUA version, becoming the first single-analyte, PCR-based COVID-19 in vitro diagnostic (IVD) device to receive FDA clearance. Over the course of the pandemic, the SARS-CoV-2 genome evolved rapidly, resulting in a burgeoning population of genomic variants. The Global Initiative on Sharing All Influenza Data (GISAID) EpiCoV™ database was quickly organized as the public sequence repository for compiling viral genomes from human cases ( 7 , 8 ) and many countries undertook large scale sequencing efforts. The number of deposited sequences was beginning to accelerate when the BioFire® COVID-19 Test EUA was granted, just 13 days after the World Health Organization (WHO) declared the COVID-19 outbreak a global pandemic. As more sequence data became available, the need to frequently assess SARS-CoV-2 variants and their potential impacts to test performance was quickly apparent ( 9 ). Such evaluations of test inclusivity must consider a combination of genomic factors and lineage-associated clinical phenotypes (such as increased transmissibility) in order to make informed decisions regarding when corrective actions or mitigations may be needed. For example, mutations falling within primer binding regions can reduce template-primer affinity and extensibility to retard or prevent PCR amplification ( 9 – 12 ). This risk may increase in the case of highly transmissible variants. Lineages harboring these mutations may quickly gain prominence in the population and escape detection in clinical specimens, especially with low viral titers. FDA and other regulatory bodies require regular viral sequence monitoring of authorized and cleared products to ensure these tests continue to identify positive cases, including emergent strains circulating both in the US and globally ( 13 ). Challenges in the in-silico Inclusivity Evaluation Process In silico inclusivity evaluations use publicly available sequence data to approximate risks to detection in the field. The process includes building a multiple sequence alignment (MSA) of intended sequence targets (inclusive sequences) and comparing nucleotide changes across primer binding regions against a reference sequence. Figure 1 reports high-level steps for standard in silico inclusivity evaluations of SARS-CoV-2 variants, similar to inclusivity processes reported in the literature ( 9 ). Online tools such as the Basic Local Alignment Search Tool (BLAST) can also be used, where amplicons or primers are queried against the National Center for Biotechnology Information (NCBI) sequence library ( 14 ), although limitations exist with this method ( 15 ). At BioFire Defense, datasets for inclusive sequences are generally small and their evaluations managed using the MSA-based approach outlined in Fig. 1 . As global outbreaks and subsequent sequence submissions caused a surge in publicly available SARS-CoV-2 sequence data, concerns grew for monitoring test performance. The combination of increasingly large data sets and the accelerated demand for new analyses exposed limitations in the standard in silico inclusivity evaluation process. Computational bottlenecks in the MSA step resulted in increased failures and reduced accuracy of large alignments, slowing the evaluation progress and complicating data interpretation. Figure 2 shows the number of complete, high coverage, human-host SARS-CoV-2 sequences collected from January 2020 through August 2022 and submitted to the GISAID EpiCoV™ database through 2 November 2022. The exponential growth of sequence datasets combined with the demands of responding to FDA monitoring requirements and customer inquiries necessitated development of a more scalable and reliable evaluation approach. Here, we describe a fully automated and adaptive bioinformatic pipeline for comprehensive monitoring and reporting BioFire COVID-19 Test performance using a predictive, risk-based strategy. Implementation The Variant Sorter Algorithm CoMIT is written as an R package. The code executes a pipeline that initially builds an empty database housing data for an evaluation (an existing database can also be updated). The pipeline takes GISAID EpiCoV™ sequence submissions (FASTA) and associated metadata (TSV) files as the input. Sequences are processed by the Variant Sorter Algorithm and the resulting data are added to the database. The BioFire COVID-19 Test uses seven nested and multiplexed SARS-COV-2 target regions, or assays, requiring inclusivity surveillance of 30 individual primers. A diagram of the Variant Sorter Algorithm – the main portion of the CoMIT pipeline – with its inputs, high-level operations, and database file output is shown in Fig. 3 . The Variant Sorter Algorithm uses an iterative string-matching comparison to identify primer variants (i.e., mutations exclusively found within primer binding regions of the test). For each sequence in the submission set, a small search space is defined around the presumptive location of the primer binding region. Novel primer variants identified by the Variant Sorter Algorithm are assigned a unique identification number and their mutation characteristics (e.g., primer affected, position, type) are captured in the database. These steps are repeated for each primer binding region. This process ensures that only a small proportion of sequences require alignment for classification. A recent inclusivity run showed 0.2% of sequences required alignment (50/21947 sequences). Pairwise alignments are performed using the DECIPHER R package ( 16 ). Sequences processed by the Variant Sorter Algorithm are stored in a relational database generated for the run (alternatively, an existing database can be specified to which new data is appended during the run). The database holds 11 tables storing sequence, assay, and primer variant data. A database schema is provided as an additional file ( Additional File 1 ). Structured Query Language Database Processing and Visualization Code After algorithm processing, other CoMIT package functions can be run on a database to generate summary tables and visualizations. Risk criteria based on primer variant prevalence, mutation severity, co-occurrence, and variant lineage type help identify primer variants predicted to be the highest risk to inclusivity. These criteria are applied in different ways to filter, highlight, and stratify data and can be modified, as needed. Figure 4 shows a flowchart for a typical in silico inclusivity analysis using the CoMIT pipeline. Risk-based Reporting Key factors for evaluating risk are described in this section, including mutational severity, co-occurrence, prevalence, and variant lineage. Figure 5 provides a summary of these considerations. Sequences harboring mutations to any test primers are identified in the evaluation, and characteristics of mutations (such as mutation position along the primer-spanning region) are leveraged to predict impacts at the individual assay level. Mismatches falling within the last five bases of the 3’ end of a primer binding region are more likely to interfere with amplification ( 11 , 19 – 22 ). Therefore, sequences carrying 3’ end mutations are labeled as a severity risk in the evaluation (Fig. 5, orange circle). When considering the risk of complete test failure, inclusivity evaluations may be further complicated by complex test designs. The BioFire COVID-19 Test 2 leverages a nested, multiplex PCR approach, targeting five independent regions of the SARS-CoV-2 genome. Detection of the expected amplicon from only one region is required to successfully elicit a SARS-CoV-2 detected result. Any sequences with mismatches to all or multiple assay primers are identified as a co-occurrence risk in the assessments (Fig. 5, grey circle) Genetic evolution of SARS-CoV-2 variants resulting in increased pathogenicity of the virus in human hosts can have significant public health impacts ( 23 ). The Centers for Disease Control and Prevention and WHO evaluate and classify emerging variants based on potential or known impacts to effectiveness of medical treatments, severity of disease, and transmissibility ( 18 , 24 ). Variant lineages associated with official designations given by US and global health organizations (e.g., Variants of Concern) are considered a prevalence risk (Fig. 5, yellow circle). The prevalence risk is also assessed for unclassified variant lineages when represented at a significant frequency in the sequence dataset ( 13 ). Sequences characterized by an overlap of any two risk factors (Fig. 5, regions indicated by 1–3) would be considered high risk, whereas sequences characterized by all risk factors (Fig. 5, area indicated by 4) are of the greatest concern due to the potential negative impacts on diagnostic accuracy. Sequences carrying primer spanning mutations flagged as high risk in these predictive evaluations are escalated for wet benchtop testing and/or thermodynamic modeling analysis ( 15 , 25 , 26 ). Results Five automated visualizations were developed to summarize processed sequence data and enable clear and concise reporting of results. An example of visualization outputs for a candidate evaluation are shown as figures and tables (Tables 1 – 3 , Figs. 6 and 7) and as additional files ( Additional Files 2–5 ). Each output in the pipeline features two or more risk indicators (i.e., co-occurrence, prevalence, lineage, and growth). All outputs (except Fig. 7 ) can be filtered based on mutational severity (i.e., when a primer-spanning mutation is positioned within 10 base pairs of the 3’ end). Database Breakdown Table The Database Breakdown Table provides a summary of collection date, variant identity (Pangolin lineage and WHO label), sequence frequencies and frequency changes of variants included in the analysis. These data are taken from the GISAID metadata associated with each sequence analyzed and can be used to clearly summarize the dataset included in the evaluation. Table 1 shows an example Database Breakdown table for a typical in silico inclusivity evaluation. Date columns refer to sample collection dates, which should include sequences from patient samples collected in the most recent three-month period. Sequence frequencies are reported for the entire dataset ( All Sequences ) and stratified by Pangolin lineage ( 27 ) and WHO label (i.e., Variants of Concern, Variants of Interest, Variants Under Monitoring); these annotations are updated for every evaluation using a variant mapping file sourced from publicly available information ( 17 , 18 ). Frequency changes are compared between the one-month sequence data (newest) and a superset of the most recent three months. Growth is represented as yellow shading when delta frequencies increase or decrease between three- and one-month sequence datasets. Variant lineages with notable frequency changes in the most recent month (i.e., greater than or equal to five percent change) are shaded in this example. Delta frequency thresholds can be modified, as needed. Identifiable Mutations in Each Assay Region As shown in Fig. 3 , the CoMIT tool first bins data according to previously identified mutations in the assay primer regions. These mutations are summarized in a table like that shown in Table 2 . The number and frequency of each lineage is recorded as an indicator of prevalence within the dataset, in this example over a 3-month period. Within these lineages, the frequencies of sequences with observed mutations are recorded in each assay column (e.g. assay 2a, 2c etc.). This gives visibility to assays which may have reduced sensitivity with emerging lineage variation. In the case of the COVID-19 Test, the Test is comprised of seven assays (2a, 2c-2g); the co-occurring mutated sequences column indicates the frequency of sequences within each lineage that contains, in this example, mutations in 5 or 6 of the assays on the Test, and that could be potentially at risk of missed or late detection. Table 2 provides an example summary table for sequences containing identifiable primer spanning mutations across COVID-19 Test assays. The Sequences by Lineage column shows the most recent three-month period sequence frequencies (count and rate) stratified by lineage exactly matching the Database Breakdown table. The remaining columns report lineage-stratified frequencies of sequences harboring an assay-specific primer variant (i.e., a primer-spanning mutation or set of mutations) ( Mutated Sequences by Assay ) and sequences with co-occurring primer variants across multiple assays ( Co-occurring Mutated Sequences) . The bottom row shows sequence frequencies of primer variants by assay ( Summary: All Sequences by Assays ). Sequence frequencies below one percent are shaded in blue; frequencies equal to or greater than five percent are shaded yellow. A summary table filtering for 3’ end mutations can be generated to represent high-risk mutations ( Additional File 2 ). A version of this table showing an expanded section for Co-occurring Mutated Sequences is also available as an additional file ( Additional File 3 ). Table 3 provides a detailed breakdown of sequences with mutations under multiple assay primers. Specifically, it shows the number of assays affected by a mutation under those primers, organized by lineage. The columns show the number of assays affected, increasing from left to right. Sequence frequencies (counts and rates) are reported by lineage ( Sequences by Lineage ) and by increasing co-occurrence risk based on the total number of assays impacted ( # Assays Affected, columns 0 through ≥6 ). The Summary: All Sequences by Assays section shows sequence frequencies based on co-occurrence risk. Blue shading indicates sequence frequencies below one percent; frequencies equal to or greater than five percent are shaded yellow. A version of this table with filtering for 3’ end mutations is provided ( Additional File 4 ). Figure 6 visualizes the lineages with mutations under multiple sets of assay primers, along with their frequencies (counts and rate). All primer-spanning mutations are reported. Purple shading indicates the impact of primer variant combinations on each individual assay and combinations when they are compounded across assays (indicating co-occurrence risks). Sequence counts and percent frequency for each combination are shown at the bottom of the figure. This figure can be filtered for 3’ end mutations only ( Additional File 5 ). Trending of Primer Variants Post-market surveillance not only tracks newly emerging sequence variants but also trends their frequencies. This monitoring helps assess the risk of missed detection based on prevalence. Variant prevalence becomes one of the risk criteria used to assess whether the diagnostic test is still functional in an evolving outbreak. Figure 7 summarizes characteristics of individual primer variants at or above 0.1% frequency in the sequence dataset compared with the previous 3-month period (note: these datasets represent nonoverlapping time periods totaling six months). The figure consists of three sections: a histogram showing assay location of primer variants and their percent frequencies in the current dataset (top section), a table detailing the primer variant characteristics and trending based on a comparison with the previous 3-month period (middle section), and a stacked bar graph displaying the primer variant distribution across lineages (bottom section). Trending symbols indicate a 0.1% frequency change in the sequence dataset compared with the previous 3-month period. An equivalent symbol represents a delta frequency less than 0.1%. The delta frequency thresholds defining inclusion criteria and trending symbols can be adjusted, as needed. Taken together, the automated outputs of the CoMIT pipeline provide summary tables and visualizations with risk-based features and modifiable thresholds for added flexibility in reporting evaluation results. Discussion CoMIT was developed specifically for in silico inclusivity evaluations of the BioFire COVID-19 Test, a single analyte, PCR-based IVD designed for use with BioFire® FilmArray® Systems. Evaluating SARS-CoV-2 genomes as they evolve through human infection is required by regulatory bodies to ensure reliable detection of COVID-19 cases in the US and globally ( 28 ). The standard inclusivity approach includes a sequence alignment step, which presented a computational bottleneck with the increasing volume and rate of sequence data needing to be analyzed. The Variant Sorter Algorithm identifies and catalogues primer variants using iterative string matching and binning functions, an efficient process to sidestep the predominance of MSAs in the standard approach. The bioinformatic analysis and visualization pipeline handles large volumes of sequence data with automated results reporting and databasing capabilities for regular comprehensive post market in silico inclusivity monitoring. CoMIT’s low computational space complexity requires minimal memory, allowing it to be run on a personal computer. In silico inclusivity monitoring serves many purposes and its results inform different audiences: online to customers in the BioFire COVID-19 Test Reactivity Technical Note ( 29 ), to regulators in FDA submissions, and companywide as required for internal trending purposes. For added flexibility and clarity in reporting, the pipeline applies risk-based parameters to summary tables and visualizations, as primer variants with these characteristics pose the greatest risks to overall the test performance. For example, figures and tables can be filtered to reporting only primer variants likely to disrupt the PCR reaction (i.e., 3’ end mutations). Co-occurrence, lineage, frequency, and delta frequency (i.e., growth) are also featured prominently in the outputs. The visualizations leverage auto-generated shading, data stratification, and symbols to identify prevalence risks in currently circulating variants using both lineage associations and growth characteristics of unclassified sequence populations. These risk criteria can be adjusted as needed to align with FDA or other post market requirements. High risk sequences can be flagged for wet benchtop testing to empirically confirm any predicted performance impacts. Despite being built for in silico inclusivity testing of COVID-19 tests, CoMIT has been developed as an accessible and user-friendly R package. Researchers can easily download and utilize CoMIT to query and test the inclusivity of their own primer sequences against any organism in the GISAID database. Details on accessing and downloading the CoMIT R package can be found in the Availability section. CoMIT’s Variant Sorter Algorithm has been adapted at BioFire Defense for inclusivity monitoring of different pathogens. We developed a modified version of CoMIT to evaluate in silico inclusivity of the Lassa virus for the BioFire® Global Fever Special Pathogens Panel (an IVD cleared by the FDA). The Lassa assays have complex designs because of the genetic diversity of the Lassa virus species which can be as high as 24.6% between lineages ( 30 ). The pipeline and algorithm are currently being expanded for processing sequences for multianalyte panels. CoMIT is primarily designed to assess the potential impact of sequence variants on the efficacy of a diagnostic test which can be used to issue a warning when a new variant is likely to escape detection. However, CoMIT could also be leveraged in epidemiology to monitor viral evolution, provide early detection of emerging variants, and inform outbreak response. CoMIT’s databasing capability allows for analyses of sequence datasets to gain new information. GISAID metadata contains details on sequence submissions, including variant information (i.e., Nextstrain clade, variant and Pangolin lineage), case data (i.e., location of exposure, demographics, reporting hospital/laboratory), and amino acid change constellation summaries. Tracking changes in viral properties such as rate and location of spread, disease severity, and variant lineage could aid in forecasting where new outbreaks may occur and when appropriate countermeasures may be needed. The CoMIT pipeline is currently limited to evaluate only primer binding sites. Successful detection on the BioFire® FilmArray® Systems depends on post run melting temperature (Tm) analyses, and variants with mutations (such as large indels in the amplicon region) that fundamentally change characteristics of the amplified region and could result in a missed detection. We are developing an amplicon tracking feature for monitoring changes to the inner amplicon region. Our inner amplicon tracker will leverage thermodynamic models to predict when sequence changes, such as indels or accumulation of single nucleotide polymorphisms, impact Tm. This new feature highlights the adaptability of CoMIT for improved predictions. Conclusion CoMIT leverages publicly available SARS-CoV-2 sequence data and metadata in GISAID EpiCoV™ repository to predict BioFire COVID-19 Test performance in the field. The pipeline can process large datasets with low computational space complexity and leverages adjustable, risk-based summarization features for easily digestible reports of highly complex testing targets. Its flexible database design and improved metadata handling provide opportunities for new epidemiological investigations of both emerging and archived case data, with the potential to improve readiness for future outbreaks as an early warning system for new variants. Abbreviations COVID-19 : Coronavirus Disease 2019 CoMIT : Coronavirus Monitoring for Inclusivity Tool EUA : Emergency Use Authorization FDA : U.S. Food and Drug Administration IVD : in vitro diagnostic GISAID : Global Initiative on Sharing All Influenza Data WHO : World Health Organization MSA : multiple sequence alignment Tm : melting temperature Declarations Availability and Requirements Project name : The Coronavirus Monitoring for Inclusivity (CoMIT) Pipeline Project Project home page : https://bitbucket.org/biofiredefense/comit/src/main/ Operating system : Platform independent Programming language : R Other requirements : Package dependencies License : CC-BY-4.0 Any restrictions to use by non-academics : Yes Ethics approval and consent to participate : Not applicable. Consent for publication : Not applicable. Availability of data and materials : The in-silico inclusivity evaluation presented in this report is based on 24,911 SARS-CoV-2 sequences and associated metadata available from January 1, 2023, up to March 31, 2023, via gisaid.org/EPI_SET_230509yp. A Supplemental Table describing the sequence dataset is provided as an additional file ( Additional File 6 ). The CoMIT R Package and instructions for its use are available at https://bitbucket.org/biofiredefense/comit/src/main/. Competing interests : W.A. Smith, J.T. Wolff, C.P. Healy, H.F. Van Hollebeke, and M. Kim are current employees of Biofire Defense, LLC; D. Walker, L. Gale, and A. Stephenson are former employees of Biofire Defense, LLC; D. Walker, W.A. Smith, and M. Kim are shareholders of Biofire Defense, LLC. Funding : This project was funded internally by BioFire Defense, LLC. Authors' contributions : Tool conceptualization: D.W., L.G., H.F.VH., W.A.S., and M.K.; software development: D.W., L.G., A.S, J.W., and C.H.; data curation: D.W., LG., and A.S.; evaluation methodology: D.W., L.G., H.F.VH., W.A.S., and M.K.; visualizations: L.G., A.S., D.W., H.F.VH., and W.A.S.; project administration: D.W. and M.K.; database validation: L.G. and H.F.VH.; benchmarking: A.S., code reviews: D.W., L.G., A.S., J.W., and C.H.; literature review: L.G., A.S., D.W., H.F.VH.; writing – original draft: D.W.; writing – review and editing: D.W., L.G., H.F.VH., A.S., W.A.S., J.W., C.H., and M.K; supervision: M.K. All authors have read and agreed to the published version of the manuscript. Acknowledgements : We gratefully acknowledge all data contributors, i.e., the Authors and their Originating laboratories responsible for obtaining the specimens, and their Submitting laboratories for generating the genetic sequence and metadata and sharing via the GISAID Initiative, on which this research is based. We thank members of the BioFire Defense Regulatory Affairs and Research and Development departments, including Kristin Casper, Dave Rabiger, and Jason Nielson for their thoughtful reviews of the manuscript. Thanks to Scott Glaittli for his technical expertise on the CoMIT software package. References Gao J, Quan L. Current Status of Diagnostic Testing for SARS-CoV-2 Infection and Future Developments: A Review. Med Sci Monit Int Med J Exp Clin Res. 2020 Dec 17;26:e928552-1-e928552-7. Nguyen NNT, McCarthy C, Lantigua D, Camci-Unal G. Development of Diagnostic Tests for Detection of SARS-CoV-2. Diagnostics. 2020 Nov;10(11):905. Ravi N, Cortade DL, Ng E, Wang SX. Diagnostics for SARS-CoV-2 detection: A comprehensive review of the FDA-EUA COVID-19 testing landscape. Biosens Bioelectron. 2020 Oct 1;165:112454. Jayamohan H, Lambert CJ, Sant HJ, Jafek A, Patel D, Feng H, et al. SARS-CoV-2 pandemic: a review of molecular diagnostic tools including sample collection and commercial response with associated advantages and limitations. Anal Bioanal Chem. 2021 Jan 1;413(1):49–71. Jalandra R, Yadav AK, Verma D, Dalal N, Sharma M, Singh R, et al. Strategies and perspectives to develop SARS-CoV-2 detection methods and diagnostics. Biomed Pharmacother. 2020 Sep 1;129:110446. Mitchell SL, St. George K, Rhoads DD, Butler-Wu SM, Dharmarha V, McNult P, et al. Understanding, Verifying, and Implementing Emergency Use Authorization Molecular Diagnostics for the Detection of SARS-CoV-2 RNA. J Clin Microbiol. 2020 Jul 23;58(8):e00796-20. GISAID Initiative [Internet]. [cited 2022 Dec 20]. Available from: https://www.epicov.org/epi3/frontend#56093 Khare S, Gurry C, Freitas L, Schultz MB, Bach G, Diallo A, et al. GISAID’s Role in Pandemic Response. China CDC Wkly. 2021 Dec 3;3(49):1049–51. Khan KA, Cheung P. Presence of mismatches between diagnostic PCR assays and coronavirus SARS-CoV-2 genome. R Soc Open Sci. 2022 Aug 3;7(6):200636. Cha RS, Thilly WG. Specificity, efficiency, and fidelity of PCR. Genome Res. 1993 Dec 1;3(3):S18–29. Bru D, Martin-Laurent F, Philippot L. Quantification of the Detrimental Effect of a Single Primer-Template Mismatch by Real-Time PCR Using the 16S rRNA Gene as an Example. Appl Environ Microbiol. 2008 Mar;74(5):1660–3. Rejali NA, Moric E, Wittwer CT. The Effect of Single Mismatches on Primer Extension. Clin Chem. 2018 May 1;64(5):801–9. Policy for Evaluating Impact of Viral Mutations on COVID-19 Tests (Revised) - Guidance for Test Developers and Food and Drug Administration Staff. BLAST: Basic Local Alignment Search Tool [Internet]. [cited 2023 Feb 24]. Available from: https://blast.ncbi.nlm.nih.gov/Blast.cgi SantaLucia J. Physical Principles and Visual-OMP Software for Optimal PCR Design. In: Yuryev A, editor. PCR Primer Design [Internet]. Totowa, NJ: Humana Press; 2007 [cited 2022 Sep 20]. p. 3–33. (Walker JM, editor. Methods in Molecular Biology TM ; vol. 402). Available from: http://link.springer.com/10.1007/978-1-59745-528-2_1 Wright ES. DECIPHER: harnessing local sequence context to improve protein multiple sequence alignment. BMC Bioinformatics. 2015 Oct 6;16(1):322. Pango designation [Internet]. CoV-lineages; 2023 [cited 2023 Feb 7]. Available from: https://github.com/cov-lineages/pango-designation/blob/106720cbb83f1cd10a55ab537f84967d8b6c2e7a/lineage_notes.txt Tracking SARS-CoV-2 variants [Internet]. [cited 2022 Nov 8]. Available from: https://www.who.int/activities/tracking-SARS-CoV-2-variants Rychlik W. Priming efficiency in PCR. BioTechniques. 1995 Jan 1;18(1):84–6, 88–90. Wu JH, Hong PY, Liu WT. Quantitative effects of position and type of single mismatch on single base primer extension. J Microbiol Methods. 2009 Jun 1;77(3):267–75. Stadhouders R, Pas SD, Anber J, Voermans J, Mes THM, Schutten M. The Effect of Primer-Template Mismatches on the Detection and Quantification of Nucleic Acids Using the 5′ Nuclease Assay. J Mol Diagn. 2010 Jan 1;12(1):109–17. Kim M, Smith WA, Van Hollebeke H. Personal communication. Aleem A, Akbar Samad AB, Slenker AK. Emerging Variants of SARS-CoV-2 And Novel Therapeutics Against Coronavirus (COVID-19). In: StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing; 2022 [cited 2023 Feb 22]. Available from: http://www.ncbi.nlm.nih.gov/books/NBK570580/ CDC. Centers for Disease Control and Prevention. 2020 [cited 2023 Feb 22]. Coronavirus Disease 2019 (COVID-19). Available from: https://www.cdc.gov/coronavirus/2019-ncov/variants/variant-classifications.html Mann T, Humbert R, Dorschner M, Stamatoyannopoulos J, Noble WS. A thermodynamic approach to PCR primer design. Nucleic Acids Res. 2009 Jul 1;37(13):e95–e95. Howson ELA, Orton RJ, Mioulet V, Lembo T, King DP, Fowler VL. GoPrime: Development of an In Silico Framework to Predict the Performance of Real-Time PCR Primers and Probes Using Foot-and-Mouth Disease Virus as a Model. Pathogens. 2020 Apr 20;9(4):303. Rambaut A, Holmes EC, O’Toole Á, Hill V, McCrone JT, Ruis C, et al. A dynamic nomenclature proposal for SARS-CoV-2 lineages to assist genomic epidemiology. Nat Microbiol. 2020 Nov;5(11):1403–7. Health C for D and R. U.S. Food and Drug Administration. FDA; 2021 [cited 2022 Nov 14]. Policy for Evaluating Impact of Viral Mutations on COVID-19 Tests. Available from: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/policy-evaluating-impact-viral-mutations-covid-19-tests BioFire® COVID-19 Test [Internet]. BioFire Defense. [cited 2022 Nov 13]. Available from: https://www.biofiredefense.com/covid-19test/ Bowen MD, Rollin PE, Ksiazek TG, Hustad HL, Bausch DG, Demby AH, et al. Genetic Diversity among Lassa Virus Strains. J Virol. 2000 Aug;74(15):6992–7004. Tables Tables are available in the Supplementary Files section. Additional Declarations Competing interest reported. W.A. Smith, J.T. Wolff, C.P. Healy, H.F. Van Hollebeke, and M. Kim are current employees of Biofire Defense, LLC; D. Walker, L. Gale, and A. Stephenson are former employees of Biofire Defense, LLC; D. Walker, W.A. Smith, and M. Kim are shareholders of Biofire Defense, LLC. Supplementary Files AdditionalMaterials.docx Tables.docx Cite Share Download PDF Status: Published Journal Publication published 12 Feb, 2025 Read the published version in BMC Bioinformatics → Version 1 posted Editorial decision: Accepted 10 Jan, 2025 Editor assigned by journal 08 Jan, 2025 Submission checks completed at journal 08 Jan, 2025 First submitted to journal 07 Jan, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5783329","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":399437764,"identity":"306bdf60-7895-4430-b647-526570241675","order_by":0,"name":"Diane M. Walker","email":"","orcid":"","institution":"BioFire Defense, LLC","correspondingAuthor":false,"prefix":"","firstName":"Diane","middleName":"M.","lastName":"Walker","suffix":""},{"id":399437765,"identity":"3a8f12db-d9b2-4779-869a-a6ad6a6168ae","order_by":1,"name":"Wendy A. Smith","email":"","orcid":"","institution":"BioFire Defense, LLC","correspondingAuthor":false,"prefix":"","firstName":"Wendy","middleName":"A.","lastName":"Smith","suffix":""},{"id":399437766,"identity":"fd8292a4-27b3-4b9d-a03d-e1931c8f2e9b","order_by":2,"name":"Lia Gale","email":"","orcid":"","institution":"BioFire Defense, LLC","correspondingAuthor":false,"prefix":"","firstName":"Lia","middleName":"","lastName":"Gale","suffix":""},{"id":399437767,"identity":"037a1dd5-9202-4c61-b14f-81477065a6bc","order_by":3,"name":"Jacob T. Wolff","email":"","orcid":"","institution":"BioFire Defense, LLC","correspondingAuthor":false,"prefix":"","firstName":"Jacob","middleName":"T.","lastName":"Wolff","suffix":""},{"id":399437768,"identity":"b9dad51f-ef20-481b-b8ef-9c887bf4323c","order_by":4,"name":"Connor P. Healy","email":"","orcid":"","institution":"BioFire Defense, LLC","correspondingAuthor":false,"prefix":"","firstName":"Connor","middleName":"P.","lastName":"Healy","suffix":""},{"id":399437769,"identity":"259eb628-0fe7-4672-8af5-d63f182480c5","order_by":5,"name":"Hannah F. Hollebeke","email":"","orcid":"","institution":"BioFire Defense, LLC","correspondingAuthor":false,"prefix":"","firstName":"Hannah","middleName":"F.","lastName":"Hollebeke","suffix":""},{"id":399437770,"identity":"b53b7371-2f58-48ec-b368-d9b5562f20b2","order_by":6,"name":"Ashlie Stephenson","email":"","orcid":"","institution":"BioFire Defense, LLC","correspondingAuthor":false,"prefix":"","firstName":"Ashlie","middleName":"","lastName":"Stephenson","suffix":""},{"id":399437771,"identity":"61182e50-fae5-4ca9-8e23-0818d2807f2e","order_by":7,"name":"Marianne Kim","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABOElEQVRIie3QPUvDQBjA8SccJEtC15Rg+hWeElAziF8lWdJF9w5iT8Sb2l2hHyIgSDYvHNQlkjVSh3bJVKHdWnzBI6KCSXV1uP90bz+OOwCV6h+GHIBXIzuksKI2GOdyYoK7lfjfJKXalSSmqIi3lRzyz1GbArGoPM3/ILt393O+SR5dcMjF9UGy75qGNZ8t+ujutQSHdf+mRrIepqOs9MDV2PQ4sz2TGF53nKHnX0aBNsymNcIj4BYTIa0Is8NbouuOxV7CuDCRyMWfBPMS0lcmBtSRu74kQ6Ibz9YbDuI8XzaSIgIhbwmgfcam2gfR5T9ggPwImkkJYoeJLrNT9jCq3qKT9niC3biIMG14C+YRWT0x0WnZvbLYJKeu2Zpoy8UJdjAX89m6XyNf6Y2rfOt5lUqlUv3WO7zgd8ygoAovAAAAAElFTkSuQmCC","orcid":"","institution":"BioFire Defense, LLC","correspondingAuthor":true,"prefix":"","firstName":"Marianne","middleName":"","lastName":"Kim","suffix":""}],"badges":[],"createdAt":"2025-01-07 17:23:07","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5783329/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5783329/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12859-025-06046-y","type":"published","date":"2025-02-12T15:57:09+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":73459782,"identity":"a4775f79-5c0b-44a7-b26f-167a6c5b5445","added_by":"auto","created_at":"2025-01-10 07:44:38","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":107144,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eHigh-level Steps for Conducting Standard in silico Inclusivity Evaluations for SARS-CoV-2 Variants.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/48ecf420701b216cea2ab477.png"},{"id":73459781,"identity":"fca24a8b-a211-4861-824c-70326186690d","added_by":"auto","created_at":"2025-01-10 07:44:38","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":89653,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGlobal Initiative for Sharing All Influenza Data (GISAID) EpiCoV™ SARSCoV-2 Sequence Submissions by Collection Month. \u003c/strong\u003eComplete, high coverage, human-host viral sequences with submission dates through 2 November 2022 and collection dates from January through August 2022 are included. Vertical lines indicate the dates when emergency use authorization (EUA) and 510(K) clearance was granted to BioFire Defense.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/5708a32ab8197d593af10623.png"},{"id":73460058,"identity":"ab6dab2e-6b41-4e46-9a67-53f9b3a91427","added_by":"auto","created_at":"2025-01-10 07:52:38","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":118863,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe Variant Sorter Algorithm in the Coronavirus Monitoring for Inclusivity (CoMIT) Pipeline. \u003c/strong\u003eSequences containing a primer-spanning match to the reference sequence are identified in the first classification step (Match bin). The red symbol indicates the portion of the algorithm where a successful classification process stops. When no match is found, the algorithm compares the sequence region against previously characterized primer variants in the database, and sequences with exact matches to existing variants are assigned the same primer variant ID (Known Primer Variant bin). When no classification has been made, the algorithm generates a library of all possible single mismatch variations and compares them to the unknown primer variant to find a match. If still no match is found, the algorithm generates all possible double mismatches for the primer and searches the expected primer region for a match. After these steps if a match has not been found, a pairwise sequence alignment against the reference sequence is used to classify the mutation (New Primer Variant bin). The novel primer variant is given a new primer variant ID and logged in the database.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/787cf93619e21c92ca76c2ff.png"},{"id":73459790,"identity":"be3f80b1-b06b-4d29-90fd-721592c048d1","added_by":"auto","created_at":"2025-01-10 07:44:38","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":86917,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFlowchart for the Typical in silico Inclusivity Evaluation Using the Coronavirus Monitoring for Inclusivity (CoMIT) Pipeline. \u003c/strong\u003eThe CoMIT pipeline requires two inputs: GISAID sequence submissions restricted by complete, high coverage and human-host filters and GISAID sequence metadata. The dotted red line highlights the portion of the pipeline involved in a Variant Sorter Run, including processing by the Variant Sorter Algorithm. Results are captured in a SQLite database where SQL queries are leveraged as inputs for the visualization and summary code (in R). Outputs of the pipeline are the visualizations. A variant mapping file is updated for each evaluation using two website sources (17,18). DBMS: database management system; GISAID: Global Initiative on Sharing Influenza Data; SQL: structured query language\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/c5af14aecc80ce6fbf39246a.png"},{"id":73459798,"identity":"c7b9e548-e335-40e1-a85b-82c82bf4882a","added_by":"auto","created_at":"2025-01-10 07:44:38","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":192983,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eRisk Considerations for Evaluating SARS-CoV-2 Variant Detection.\u003c/strong\u003e The diagram outlines factors with increased risk for affecting overall test performance. All mutations are tracked and reported based on risk criteria and elevated risks are identified when overlapping risk characteristics are present (labels 1-4). Highest risk to overall test performance is predicted when sequences carrying primer-spanning mutations both negatively impact performance of all or several individual assays and occur at a significant prevalence in the sequence dataset (e.g., 5% or higher) (label 4).\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/47d85649a84ddbce60f9f803.png"},{"id":73459788,"identity":"b2bd873f-c709-4454-aa64-9b12af12dca7","added_by":"auto","created_at":"2025-01-10 07:44:38","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":41784,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003ePrimer Variant Combinations in a Typical in silico Inclusivity Evaluation. \u003c/strong\u003e\u003c/em\u003e\u003cem\u003eThe figure is divided into two sections (a heatmap and histogram) to summarize co-occurring primer variant combinations in the dataset. The number of assays affected is shown across the top with each column relating to a specific primer variant combination group. The top heatmap shows abbreviated assay names along the left y-axis (2a, 2c, 2d, 2e, 2f, 2g, and 2h). Purple shading indicates the specific assays impacted in the various primer variant combination groups. The number of assays affected increases from left to right. The bottom histogram reports the percent frequency of sequences in the dataset represented in each primer variant combination group.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/e3291baff25a3baad6d1af7d.png"},{"id":73460063,"identity":"fa3bb1f4-c728-4b96-8e9d-c3b5fa8a3a51","added_by":"auto","created_at":"2025-01-10 07:52:38","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":65020,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003ePrimer Variant Characteristics and Trending. \u003c/strong\u003e\u003c/em\u003e\u003cem\u003eAssay names of affected primers are shown along the top with shared primer regions indicated by both assay names (A, C, C/D, D, E, E/F, G, H). The column indicated by All combines all assay data. Each row relates to the specific primer variant indicated by the assay name column. The top part of the figure shows a histogram with raw count frequencies of each primer variant\u003c/em\u003e\u003cem\u003e\u003cstrong\u003e. \u003c/strong\u003e\u003c/em\u003e\u003cem\u003eThe middle section shows characteristics of each primer variant including prevalence, growth trend compared with the previous three-month dataset\u003c/em\u003e\u003csup\u003e\u003cem\u003e*\u003c/em\u003e\u003c/sup\u003e\u003cem\u003e, primer affected, mutation position(s), and primer/template mismatch (or DL## to indicate a deletion and number of base pairs spanning the deletion). The bottom section of the figure provides a bar chart showing the lineage distribution of the primer variant.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e\u003cem\u003e*\u003c/em\u003e\u003c/sup\u003e\u003cem\u003eUp arrow indicates a 0.1 or greater delta increase, downward arrow indicates 0.1 or greater delta decrease, equivalent arrows indicate a less than 0.1 growth change.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/f5203d1a0eecb763aa5569b4.png"},{"id":76487462,"identity":"72628472-0dbd-46bb-bda2-7bd9147d4e0b","added_by":"auto","created_at":"2025-02-17 16:07:38","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1697129,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/c2647a83-4455-4e46-9a13-b9c9192092cd.pdf"},{"id":73459787,"identity":"1e01b70a-db8c-4fce-a65a-b2e47ed6cd40","added_by":"auto","created_at":"2025-01-10 07:44:38","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1165858,"visible":true,"origin":"","legend":"","description":"","filename":"AdditionalMaterials.docx","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/62ef649f619fe5c514d9f9bc.docx"},{"id":73459789,"identity":"f45a815e-30ae-4831-beb8-702596a79706","added_by":"auto","created_at":"2025-01-10 07:44:38","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":864539,"visible":true,"origin":"","legend":"","description":"","filename":"Tables.docx","url":"https://assets-eu.researchsquare.com/files/rs-5783329/v1/1384d1bfc18baedae0c7f00d.docx"}],"financialInterests":"Competing interest reported. W.A. Smith, J.T. Wolff, C.P. Healy, H.F. Van Hollebeke, and M. Kim are current employees of Biofire Defense, LLC; D. Walker, L. Gale, and A. Stephenson are former employees of Biofire Defense, LLC; D. Walker, W.A. Smith, and M. Kim are shareholders of Biofire Defense, LLC.","formattedTitle":"CoMIT: A Bioinformatic Pipeline for Risk-based Prediction of COVID-19 Test Inclusivity","fulltext":[{"header":"Background","content":"\u003cp\u003eThe Coronavirus Disease 2019 (COVID-19) pandemic brought challenges never faced in the modern era, necessitating accelerated timelines and processes to address urgent public health needs. The global outbreak spawned the rapid development of diagnostic tests capable of detecting the presence of SARS-CoV-2 (the etiological agent of COVID-19) in human samples (\u003cspan class=\"CitationRef\"\u003e1\u003c/span\u003e–\u003cspan class=\"CitationRef\"\u003e6\u003c/span\u003e). In response, BioFire Defense leveraged its existing BioFire® FilmArray® PCR-based technology to develop a test specific for the identification of SARS-CoV-2 nucleic acid from patient samples. On March 24, 2020, BioFire Defense received initial Emergency Use Authorization (EUA) by the U.S. Food and Drug Administration (FDA) for the BioFire® COVID-19 Test. A 510k clearance was granted for the BioFire® COVID-19 Test 2 on November 1, 2021, which has identical chemistry to the EUA version, becoming the first single-analyte, PCR-based COVID-19 \u003cem\u003ein vitro\u003c/em\u003e diagnostic (IVD) device to receive FDA clearance.\u003c/p\u003e\n\u003cp\u003eOver the course of the pandemic, the SARS-CoV-2 genome evolved rapidly, resulting in a burgeoning population of genomic variants. The Global Initiative on Sharing All Influenza Data (GISAID) EpiCoV™ database was quickly organized as the public sequence repository for compiling viral genomes from human cases (\u003cspan class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e8\u003c/span\u003e) and many countries undertook large scale sequencing efforts. The number of deposited sequences was beginning to accelerate when the BioFire® COVID-19 Test EUA was granted, just 13 days after the World Health Organization (WHO) declared the COVID-19 outbreak a global pandemic.\u003c/p\u003e\n\u003cp\u003eAs more sequence data became available, the need to frequently assess SARS-CoV-2 variants and their potential impacts to test performance was quickly apparent (\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e). Such evaluations of test inclusivity must consider a combination of genomic factors and lineage-associated clinical phenotypes (such as increased transmissibility) in order to make informed decisions regarding when corrective actions or mitigations may be needed. For example, mutations falling within primer binding regions can reduce template-primer affinity and extensibility to retard or prevent PCR amplification (\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e–\u003cspan class=\"CitationRef\"\u003e12\u003c/span\u003e). This risk may increase in the case of highly transmissible variants. Lineages harboring these mutations may quickly gain prominence in the population and escape detection in clinical specimens, especially with low viral titers. FDA and other regulatory bodies require regular viral sequence monitoring of authorized and cleared products to ensure these tests continue to identify positive cases, including emergent strains circulating both in the US and globally (\u003cspan class=\"CitationRef\"\u003e13\u003c/span\u003e).\u003c/p\u003e\n\u003cdiv\u003e\n \u003ch2\u003eChallenges in the \u003cem\u003ein-silico\u003c/em\u003e Inclusivity Evaluation Process\u003c/h2\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cem\u003eIn silico\u003c/em\u003e inclusivity evaluations use publicly available sequence data to approximate risks to detection in the field. The process includes building a multiple sequence alignment (MSA) of intended sequence targets (inclusive sequences) and comparing nucleotide changes across primer binding regions against a reference sequence. Figure 1 reports high-level steps for standard \u003cem\u003ein silico\u003c/em\u003e inclusivity evaluations of SARS-CoV-2 variants, similar to inclusivity processes reported in the literature (\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e). Online tools such as the Basic Local Alignment Search Tool (BLAST) can also be used, where amplicons or primers are queried against the National Center for Biotechnology Information (NCBI) sequence library (\u003cspan class=\"CitationRef\"\u003e14\u003c/span\u003e), although limitations exist with this method (\u003cspan class=\"CitationRef\"\u003e15\u003c/span\u003e). At BioFire Defense, datasets for inclusive sequences are generally small and their evaluations managed using the MSA-based approach outlined in \u003cstrong\u003eFig.\u0026nbsp;1\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eAs global outbreaks and subsequent sequence submissions caused a surge in publicly available SARS-CoV-2 sequence data, concerns grew for monitoring test performance. The combination of increasingly large data sets and the accelerated demand for new analyses exposed limitations in the standard \u003cem\u003ein silico\u003c/em\u003e inclusivity evaluation process. Computational bottlenecks in the MSA step resulted in increased failures and reduced accuracy of large alignments, slowing the evaluation progress and complicating data interpretation. Figure 2 shows the number of complete, high coverage, human-host SARS-CoV-2 sequences collected from January 2020 through August 2022 and submitted to the GISAID EpiCoV™ database through 2 November 2022. The exponential growth of sequence datasets combined with the demands of responding to FDA monitoring requirements and customer inquiries necessitated development of a more scalable and reliable evaluation approach. Here, we describe a fully automated and adaptive bioinformatic pipeline for comprehensive monitoring and reporting BioFire COVID-19 Test performance using a predictive, risk-based strategy.\u003c/p\u003e\n"},{"header":"Implementation","content":"\u003cp\u003eThe Variant Sorter Algorithm\u003c/p\u003e\u003cp\u003eCoMIT is written as an R package. The code executes a pipeline that initially builds an empty database housing data for an evaluation (an existing database can also be updated). The pipeline takes GISAID EpiCoV™ sequence submissions (FASTA) and associated metadata (TSV) files as the input. Sequences are processed by the Variant Sorter Algorithm and the resulting data are added to the database. The BioFire COVID-19 Test uses seven nested and multiplexed SARS-COV-2 target regions, or assays, requiring inclusivity surveillance of 30 individual primers. A diagram of the Variant Sorter Algorithm – the main portion of the CoMIT pipeline – with its inputs, high-level operations, and database file output is shown in \u003cstrong\u003eFig.\u0026nbsp;3\u003c/strong\u003e.\u003c/p\u003e\u003cp\u003eThe Variant Sorter Algorithm uses an iterative string-matching comparison to identify \u003cem\u003eprimer variants\u003c/em\u003e (i.e., mutations exclusively found within primer binding regions of the test). For each sequence in the submission set, a small search space is defined around the presumptive location of the primer binding region. Novel primer variants identified by the Variant Sorter Algorithm are assigned a unique identification number and their mutation characteristics (e.g., primer affected, position, type) are captured in the database. These steps are repeated for each primer binding region. This process ensures that only a small proportion of sequences require alignment for classification. A recent inclusivity run showed 0.2% of sequences required alignment (50/21947 sequences). Pairwise alignments are performed using the DECIPHER R package (\u003cspan class=\"CitationRef\"\u003e16\u003c/span\u003e). Sequences processed by the Variant Sorter Algorithm are stored in a relational database generated for the run (alternatively, an existing database can be specified to which new data is appended during the run). The database holds 11 tables storing sequence, assay, and primer variant data. A database schema is provided as an additional file (\u003cstrong\u003eAdditional File 1\u003c/strong\u003e).\u003c/p\u003e\u003ch2\u003eStructured Query Language Database Processing and Visualization Code\u003c/h2\u003e\u003cp\u003eAfter algorithm processing, other CoMIT package functions can be run on a database to generate summary tables and visualizations. Risk criteria based on primer variant prevalence, mutation severity, co-occurrence, and variant lineage type help identify primer variants predicted to be the highest risk to inclusivity. These criteria are applied in different ways to filter, highlight, and stratify data and can be modified, as needed. Figure\u0026nbsp;4 shows a flowchart for a typical \u003cem\u003ein silico\u003c/em\u003e inclusivity analysis using the CoMIT pipeline.\u003c/p\u003e\u003ch3\u003eRisk-based Reporting\u003c/h3\u003e\u003cp\u003eKey factors for evaluating risk are described in this section, including mutational severity, co-occurrence, prevalence, and variant lineage. Figure\u0026nbsp;5 provides a summary of these considerations. Sequences harboring mutations to any test primers are identified in the evaluation, and characteristics of mutations (such as mutation position along the primer-spanning region) are leveraged to predict impacts at the individual assay level. Mismatches falling within the last five bases of the 3’ end of a primer binding region are more likely to interfere with amplification (\u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e19\u003c/span\u003e–\u003cspan class=\"CitationRef\"\u003e22\u003c/span\u003e). Therefore, sequences carrying 3’ end mutations are labeled as a \u003cem\u003eseverity risk\u003c/em\u003e in the evaluation (Fig.\u0026nbsp;5, orange circle).\u003c/p\u003e\u003cp\u003eWhen considering the risk of complete test failure, inclusivity evaluations may be further complicated by complex test designs. The BioFire COVID-19 Test 2 leverages a nested, multiplex PCR approach, targeting five independent regions of the SARS-CoV-2 genome. Detection of the expected amplicon from only one region is required to successfully elicit a SARS-CoV-2 detected result. Any sequences with mismatches to all or multiple assay primers are identified as a \u003cem\u003eco-occurrence risk\u003c/em\u003e in the assessments (Fig.\u0026nbsp;5, grey circle)\u003c/p\u003e\u003cp\u003eGenetic evolution of SARS-CoV-2 variants resulting in increased pathogenicity of the virus in human hosts can have significant public health impacts (\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e). The Centers for Disease Control and Prevention and WHO evaluate and classify emerging variants based on potential or known impacts to effectiveness of medical treatments, severity of disease, and transmissibility (\u003cspan class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e24\u003c/span\u003e). Variant lineages associated with official designations given by US and global health organizations (e.g., Variants of Concern) are considered a \u003cem\u003eprevalence risk\u003c/em\u003e (Fig.\u0026nbsp;5, yellow circle). The prevalence risk is also assessed for unclassified variant lineages when represented at a significant frequency in the sequence dataset (\u003cspan class=\"CitationRef\"\u003e13\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eSequences characterized by an overlap of any two risk factors (Fig.\u0026nbsp;5, regions indicated by 1–3) would be considered high risk, whereas sequences characterized by all risk factors (Fig.\u0026nbsp;5, area indicated by 4) are of the greatest concern due to the potential negative impacts on diagnostic accuracy. Sequences carrying primer spanning mutations flagged as high risk in these predictive evaluations are escalated for wet benchtop testing and/or thermodynamic modeling analysis (\u003cspan class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e).\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eFive automated visualizations were developed to summarize processed sequence data and enable clear and concise reporting of results. An example of visualization outputs for a candidate evaluation are shown as figures and tables (Tables \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e, \u003cstrong\u003eFigs.\u0026nbsp;6 and 7)\u003c/strong\u003e and as additional files (\u003cstrong\u003eAdditional Files 2\u0026ndash;5\u003c/strong\u003e). Each output in the pipeline features two or more risk indicators (i.e., co-occurrence, prevalence, lineage, and growth). All outputs (except \u003cstrong\u003eFig.\u0026nbsp;7\u003c/strong\u003e) can be filtered based on mutational severity (i.e., when a primer-spanning mutation is positioned within 10 base pairs of the 3\u0026rsquo; end).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003ctable id=\"Tab1\" border=\"1\"\u003e\u003c/table\u003e\n\u003c/div\u003e\n\u003ch3\u003eDatabase Breakdown Table\u003c/h3\u003e\n\u003cp\u003eThe Database Breakdown Table provides a summary of collection date, variant identity (Pangolin lineage and WHO label), sequence frequencies and frequency changes of variants included in the analysis. These data are taken from the GISAID metadata associated with each sequence analyzed and can be used to clearly summarize the dataset included in the evaluation.\u003c/p\u003e\n\u003cp\u003eTable \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e shows an example \u003cem\u003eDatabase Breakdown\u003c/em\u003e table for a typical \u003cem\u003ein silico\u003c/em\u003e inclusivity evaluation. Date columns refer to sample collection dates, which should include sequences from patient samples collected in the most recent three-month period. Sequence frequencies are reported for the entire dataset (\u003cstrong\u003eAll Sequences\u003c/strong\u003e) and stratified by Pangolin lineage (\u003cspan class=\"CitationRef\"\u003e27\u003c/span\u003e) and WHO label (i.e., Variants of Concern, Variants of Interest, Variants Under Monitoring); these annotations are updated for every evaluation using a variant mapping file sourced from publicly available information (\u003cspan class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e18\u003c/span\u003e). Frequency changes are compared between the one-month sequence data (newest) and a superset of the most recent three months. Growth is represented as yellow shading when delta frequencies increase or decrease between three- and one-month sequence datasets. Variant lineages with notable frequency changes in the most recent month (i.e., greater than or equal to five percent change) are shaded in this example. Delta frequency thresholds can be modified, as needed.\u003c/p\u003e\n\u003ch3\u003eIdentifiable Mutations in Each Assay Region\u003c/h3\u003e\n\u003cp\u003eAs shown in \u003cstrong\u003eFig.\u0026nbsp;3\u003c/strong\u003e, the CoMIT tool first bins data according to previously identified mutations in the assay primer regions. These mutations are summarized in a table like that shown in Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e. The number and frequency of each lineage is recorded as an indicator of prevalence within the dataset, in this example over a 3-month period. Within these lineages, the frequencies of sequences with observed mutations are recorded in each assay column (e.g. assay 2a, 2c etc.). This gives visibility to assays which may have reduced sensitivity with emerging lineage variation. In the case of the COVID-19 Test, the Test is comprised of seven assays (2a, 2c-2g); the co-occurring mutated sequences column indicates the frequency of sequences within each lineage that contains, in this example, mutations in 5 or 6 of the assays on the Test, and that could be potentially at risk of missed or late detection.\u003c/p\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\u003c/table\u003e\n\u003cp\u003e\u003c/p\u003e\n\u003cp\u003eTable \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e provides an example summary table for sequences containing identifiable primer spanning mutations across COVID-19 Test assays. The \u003cem\u003eSequences by Lineage\u003c/em\u003e column shows the most recent three-month period sequence frequencies (count and rate) stratified by lineage exactly matching the \u003cem\u003eDatabase Breakdown\u003c/em\u003e table. The remaining columns report lineage-stratified frequencies of sequences harboring an assay-specific primer variant (i.e., a primer-spanning mutation or set of mutations) (\u003cem\u003eMutated Sequences by Assay\u003c/em\u003e) and sequences with co-occurring primer variants across multiple assays (\u003cem\u003eCo-occurring Mutated Sequences)\u003c/em\u003e. The bottom row shows sequence frequencies of primer variants by assay (\u003cem\u003eSummary: All Sequences by Assays\u003c/em\u003e). Sequence frequencies below one percent are shaded in blue; frequencies equal to or greater than five percent are shaded yellow. A summary table filtering for 3\u0026rsquo; end mutations can be generated to represent high-risk mutations (\u003cstrong\u003eAdditional File 2\u003c/strong\u003e). A version of this table showing an expanded section for \u003cem\u003eCo-occurring Mutated Sequences\u003c/em\u003e is also available as an additional file (\u003cstrong\u003eAdditional File 3\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3\u003c/strong\u003e provides a detailed breakdown of sequences with mutations under multiple assay primers. Specifically, it shows the number of assays affected by a mutation under those primers, organized by lineage. The columns show the number of assays affected, increasing from left to right. Sequence frequencies (counts and rates) are reported by lineage (\u003cem\u003eSequences by Lineage\u003c/em\u003e) and by increasing co-occurrence risk based on the total number of assays impacted (\u003cem\u003e# Assays Affected,\u0026nbsp;\u003c/em\u003ecolumns\u003cem\u003e\u0026nbsp;0\u003c/em\u003e through \u003cem\u003e\u0026ge;6\u003c/em\u003e). The \u003cem\u003eSummary: All Sequences by Assays\u003c/em\u003e section shows sequence frequencies based on co-occurrence risk. Blue shading indicates sequence frequencies below one percent; frequencies equal to or greater than five percent are shaded yellow. A version of this table with filtering for 3\u0026rsquo; end mutations is provided (\u003cstrong\u003eAdditional File 4\u003c/strong\u003e).\u003c/p\u003e\n\u003ctable id=\"Tab3\" border=\"1\"\u003e\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eFigure 6\u003c/strong\u003e visualizes the lineages with mutations under multiple sets of assay primers, along with their frequencies (counts and rate). All primer-spanning mutations are reported. Purple shading indicates the impact of primer variant combinations on each individual assay and combinations when they are compounded across assays (indicating co-occurrence risks). Sequence counts and percent frequency for each combination are shown at the bottom of the figure. This figure can be filtered for 3\u0026rsquo; end mutations only (\u003cstrong\u003eAdditional File 5\u003c/strong\u003e).\u003c/p\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n \u003ch2\u003eTrending of Primer Variants\u003c/h2\u003e\n \u003cp\u003ePost-market surveillance not only tracks newly emerging sequence variants but also trends their frequencies. This monitoring helps assess the risk of missed detection based on prevalence. Variant prevalence becomes one of the risk criteria used to assess whether the diagnostic test is still functional in an evolving outbreak.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eFigure 7\u003c/strong\u003e summarizes characteristics of individual primer variants at or above 0.1% frequency in the sequence dataset compared with the previous 3-month period (note: these datasets represent nonoverlapping time periods totaling six months). The figure consists of three sections: a histogram showing assay location of primer variants and their percent frequencies in the current dataset (top section), a table detailing the primer variant characteristics and trending based on a comparison with the previous 3-month period (middle section), and a stacked bar graph displaying the primer variant distribution across lineages (bottom section). Trending symbols indicate a 0.1% frequency change in the sequence dataset compared with the previous 3-month period. An equivalent symbol represents a delta frequency less than 0.1%. The delta frequency thresholds defining inclusion criteria and trending symbols can be adjusted, as needed.\u003c/p\u003e\n \u003cp\u003eTaken together, the automated outputs of the CoMIT pipeline provide summary tables and visualizations with risk-based features and modifiable thresholds for added flexibility in reporting evaluation results.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eCoMIT was developed specifically for \u003cem\u003ein silico\u003c/em\u003e inclusivity evaluations of the BioFire COVID-19 Test, a single analyte, PCR-based IVD designed for use with BioFire\u0026reg; FilmArray\u0026reg; Systems. Evaluating SARS-CoV-2 genomes as they evolve through human infection is required by regulatory bodies to ensure reliable detection of COVID-19 cases in the US and globally (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e). The standard inclusivity approach includes a sequence alignment step, which presented a computational bottleneck with the increasing volume and rate of sequence data needing to be analyzed. The Variant Sorter Algorithm identifies and catalogues primer variants using iterative string matching and binning functions, an efficient process to sidestep the predominance of MSAs in the standard approach. The bioinformatic analysis and visualization pipeline handles large volumes of sequence data with automated results reporting and databasing capabilities for regular comprehensive post market \u003cem\u003ein silico\u003c/em\u003e inclusivity monitoring. CoMIT\u0026rsquo;s low computational space complexity requires minimal memory, allowing it to be run on a personal computer.\u003c/p\u003e \u003cp\u003e \u003cem\u003eIn silico\u003c/em\u003e inclusivity monitoring serves many purposes and its results inform different audiences: online to customers in the BioFire COVID-19 Test Reactivity Technical Note (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e), to regulators in FDA submissions, and companywide as required for internal trending purposes. For added flexibility and clarity in reporting, the pipeline applies risk-based parameters to summary tables and visualizations, as primer variants with these characteristics pose the greatest risks to overall the test performance. For example, figures and tables can be filtered to reporting only primer variants likely to disrupt the PCR reaction (i.e., 3\u0026rsquo; end mutations). Co-occurrence, lineage, frequency, and delta frequency (i.e., growth) are also featured prominently in the outputs. The visualizations leverage auto-generated shading, data stratification, and symbols to identify prevalence risks in currently circulating variants using both lineage associations and growth characteristics of unclassified sequence populations. These risk criteria can be adjusted as needed to align with FDA or other post market requirements. High risk sequences can be flagged for wet benchtop testing to empirically confirm any predicted performance impacts.\u003c/p\u003e \u003cp\u003eDespite being built for \u003cem\u003ein silico\u003c/em\u003e inclusivity testing of COVID-19 tests, CoMIT has been developed as an accessible and user-friendly R package. Researchers can easily download and utilize CoMIT to query and test the inclusivity of their own primer sequences against any organism in the GISAID database. Details on accessing and downloading the CoMIT R package can be found in the Availability section.\u003c/p\u003e \u003cp\u003eCoMIT\u0026rsquo;s Variant Sorter Algorithm has been adapted at BioFire Defense for inclusivity monitoring of different pathogens. We developed a modified version of CoMIT to evaluate \u003cem\u003ein silico\u003c/em\u003e inclusivity of the Lassa virus for the BioFire\u0026reg; Global Fever Special Pathogens Panel (an IVD cleared by the FDA). The Lassa assays have complex designs because of the genetic diversity of the Lassa virus species which can be as high as 24.6% between lineages (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e). The pipeline and algorithm are currently being expanded for processing sequences for multianalyte panels.\u003c/p\u003e \u003cp\u003eCoMIT is primarily designed to assess the potential impact of sequence variants on the efficacy of a diagnostic test which can be used to issue a warning when a new variant is likely to escape detection. However, CoMIT could also be leveraged in epidemiology to monitor viral evolution, provide early detection of emerging variants, and inform outbreak response. CoMIT\u0026rsquo;s databasing capability allows for analyses of sequence datasets to gain new information. GISAID metadata contains details on sequence submissions, including variant information (i.e., Nextstrain clade, variant and Pangolin lineage), case data (i.e., location of exposure, demographics, reporting hospital/laboratory), and amino acid change constellation summaries. Tracking changes in viral properties such as rate and location of spread, disease severity, and variant lineage could aid in forecasting where new outbreaks may occur and when appropriate countermeasures may be needed.\u003c/p\u003e \u003cp\u003eThe CoMIT pipeline is currently limited to evaluate only primer binding sites. Successful detection on the BioFire\u0026reg; FilmArray\u0026reg; Systems depends on post run melting temperature (Tm) analyses, and variants with mutations (such as large indels in the amplicon region) that fundamentally change characteristics of the amplified region and could result in a missed detection. We are developing an amplicon tracking feature for monitoring changes to the inner amplicon region. Our inner amplicon tracker will leverage thermodynamic models to predict when sequence changes, such as indels or accumulation of single nucleotide polymorphisms, impact Tm. This new feature highlights the adaptability of CoMIT for improved predictions.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eCoMIT leverages publicly available SARS-CoV-2 sequence data and metadata in GISAID EpiCoV\u0026trade; repository to predict BioFire COVID-19 Test performance in the field. The pipeline can process large datasets with low computational space complexity and leverages adjustable, risk-based summarization features for easily digestible reports of highly complex testing targets. Its flexible database design and improved metadata handling provide opportunities for new epidemiological investigations of both emerging and archived case data, with the potential to improve readiness for future outbreaks as an early warning system for new variants.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003e\u003cstrong\u003eCOVID-19\u003c/strong\u003e: Coronavirus Disease 2019\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCoMIT\u003c/strong\u003e: Coronavirus Monitoring for Inclusivity Tool\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEUA\u003c/strong\u003e: Emergency Use Authorization\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFDA\u003c/strong\u003e: U.S. Food and Drug Administration\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eIVD\u003c/strong\u003e: \u003cem\u003ein vitro\u003c/em\u003e diagnostic\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eGISAID\u003c/strong\u003e:\u0026nbsp;Global Initiative on Sharing All Influenza Data\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eWHO\u003c/strong\u003e:\u0026nbsp;World Health Organization\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMSA\u003c/strong\u003e: multiple sequence alignment\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTm\u003c/strong\u003e: melting temperature\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eAvailability and Requirements \u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProject name\u003c/strong\u003e: The Coronavirus Monitoring for Inclusivity (CoMIT) Pipeline Project\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProject home page\u003c/strong\u003e: \u0026nbsp;https://bitbucket.org/biofiredefense/comit/src/main/\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eOperating system\u003c/strong\u003e: Platform independent\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProgramming language\u003c/strong\u003e: R\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eOther requirements\u003c/strong\u003e: Package dependencies\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLicense\u003c/strong\u003e: CC-BY-4.0\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAny restrictions to use by non-academics\u003c/strong\u003e: Yes\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e: Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e: Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e: The \u003cem\u003ein-silico inclusivity\u0026nbsp;\u003c/em\u003eevaluation presented in this report is based on 24,911 SARS-CoV-2 sequences and associated metadata available from January 1, 2023, up to March 31, 2023, via gisaid.org/EPI_SET_230509yp. A Supplemental Table describing the sequence dataset is provided as an additional file (\u003cstrong\u003eAdditional File 6\u003c/strong\u003e). The CoMIT R Package and instructions for its use are available at https://bitbucket.org/biofiredefense/comit/src/main/.\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e: W.A. Smith, J.T. Wolff, C.P. Healy, H.F. Van Hollebeke, and M. Kim are current employees of Biofire Defense, LLC; D. Walker, L. Gale, and A. Stephenson are former employees of Biofire Defense, LLC; D. Walker, W.A. Smith, and M. Kim are shareholders of Biofire Defense, LLC.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e: This project was funded internally by BioFire Defense, LLC.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e: Tool conceptualization: D.W., L.G., H.F.VH., W.A.S., and M.K.; software development: D.W., L.G., A.S, J.W., and C.H.; data curation: D.W., LG., and A.S.; evaluation methodology: D.W., L.G., H.F.VH., W.A.S., and M.K.; visualizations: L.G., A.S., D.W., H.F.VH., and W.A.S.; project administration: D.W. and M.K.; database validation: L.G. and H.F.VH.; benchmarking: A.S., code reviews: D.W., L.G., A.S., J.W., and C.H.; literature review: L.G., A.S., D.W., H.F.VH.; writing \u0026ndash; original draft: D.W.; writing \u0026ndash; review and editing: D.W., L.G., H.F.VH., A.S., W.A.S., J.W., C.H., and M.K; supervision: M.K. All authors have read and agreed to the published version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e: We gratefully acknowledge all data contributors, i.e., the Authors and their Originating laboratories responsible for obtaining the specimens, and their Submitting laboratories for generating the genetic sequence and metadata and sharing via the GISAID Initiative, on which this research is based. We thank members of the BioFire Defense Regulatory Affairs and Research and Development departments, including Kristin Casper, Dave Rabiger, and Jason Nielson for their thoughtful reviews of the manuscript. Thanks to Scott Glaittli for his technical expertise on the CoMIT software package.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eGao J, Quan L. Current Status of Diagnostic Testing for SARS-CoV-2 Infection and Future Developments: A Review. Med Sci Monit Int Med J Exp Clin Res. 2020 Dec 17;26:e928552-1-e928552-7. \u003c/li\u003e\n\u003cli\u003eNguyen NNT, McCarthy C, Lantigua D, Camci-Unal G. Development of Diagnostic Tests for Detection of SARS-CoV-2. Diagnostics. 2020 Nov;10(11):905. \u003c/li\u003e\n\u003cli\u003eRavi N, Cortade DL, Ng E, Wang SX. Diagnostics for SARS-CoV-2 detection: A comprehensive review of the FDA-EUA COVID-19 testing landscape. Biosens Bioelectron. 2020 Oct 1;165:112454. \u003c/li\u003e\n\u003cli\u003eJayamohan H, Lambert CJ, Sant HJ, Jafek A, Patel D, Feng H, et al. SARS-CoV-2 pandemic: a review of molecular diagnostic tools including sample collection and commercial response with associated advantages and limitations. Anal Bioanal Chem. 2021 Jan 1;413(1):49\u0026ndash;71. \u003c/li\u003e\n\u003cli\u003eJalandra R, Yadav AK, Verma D, Dalal N, Sharma M, Singh R, et al. Strategies and perspectives to develop SARS-CoV-2 detection methods and diagnostics. Biomed Pharmacother. 2020 Sep 1;129:110446. \u003c/li\u003e\n\u003cli\u003eMitchell SL, St. George K, Rhoads DD, Butler-Wu SM, Dharmarha V, McNult P, et al. Understanding, Verifying, and Implementing Emergency Use Authorization Molecular Diagnostics for the Detection of SARS-CoV-2 RNA. J Clin Microbiol. 2020 Jul 23;58(8):e00796-20. \u003c/li\u003e\n\u003cli\u003eGISAID Initiative [Internet]. [cited 2022 Dec 20]. Available from: https://www.epicov.org/epi3/frontend#56093\u003c/li\u003e\n\u003cli\u003eKhare S, Gurry C, Freitas L, Schultz MB, Bach G, Diallo A, et al. GISAID\u0026rsquo;s Role in Pandemic Response. China CDC Wkly. 2021 Dec 3;3(49):1049\u0026ndash;51. \u003c/li\u003e\n\u003cli\u003eKhan KA, Cheung P. Presence of mismatches between diagnostic PCR assays and coronavirus SARS-CoV-2 genome. R Soc Open Sci. 2022 Aug 3;7(6):200636. \u003c/li\u003e\n\u003cli\u003eCha RS, Thilly WG. Specificity, efficiency, and fidelity of PCR. Genome Res. 1993 Dec 1;3(3):S18\u0026ndash;29. \u003c/li\u003e\n\u003cli\u003eBru D, Martin-Laurent F, Philippot L. Quantification of the Detrimental Effect of a Single Primer-Template Mismatch by Real-Time PCR Using the 16S rRNA Gene as an Example. Appl Environ Microbiol. 2008 Mar;74(5):1660\u0026ndash;3. \u003c/li\u003e\n\u003cli\u003eRejali NA, Moric E, Wittwer CT. The Effect of Single Mismatches on Primer Extension. Clin Chem. 2018 May 1;64(5):801\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003ePolicy for Evaluating Impact of Viral Mutations on COVID-19 Tests (Revised) - Guidance for Test Developers and Food and Drug Administration Staff. \u003c/li\u003e\n\u003cli\u003eBLAST: Basic Local Alignment Search Tool [Internet]. [cited 2023 Feb 24]. Available from: https://blast.ncbi.nlm.nih.gov/Blast.cgi\u003c/li\u003e\n\u003cli\u003eSantaLucia J. Physical Principles and Visual-OMP Software for Optimal PCR Design. In: Yuryev A, editor. PCR Primer Design [Internet]. Totowa, NJ: Humana Press; 2007 [cited 2022 Sep 20]. p. 3\u0026ndash;33. (Walker JM, editor. Methods in Molecular Biology\u003csup\u003eTM\u003c/sup\u003e; vol. 402). Available from: http://link.springer.com/10.1007/978-1-59745-528-2_1\u003c/li\u003e\n\u003cli\u003eWright ES. DECIPHER: harnessing local sequence context to improve protein multiple sequence alignment. BMC Bioinformatics. 2015 Oct 6;16(1):322. \u003c/li\u003e\n\u003cli\u003ePango designation [Internet]. CoV-lineages; 2023 [cited 2023 Feb 7]. Available from: https://github.com/cov-lineages/pango-designation/blob/106720cbb83f1cd10a55ab537f84967d8b6c2e7a/lineage_notes.txt\u003c/li\u003e\n\u003cli\u003eTracking SARS-CoV-2 variants [Internet]. [cited 2022 Nov 8]. Available from: https://www.who.int/activities/tracking-SARS-CoV-2-variants\u003c/li\u003e\n\u003cli\u003eRychlik W. Priming efficiency in PCR. BioTechniques. 1995 Jan 1;18(1):84\u0026ndash;6, 88\u0026ndash;90. \u003c/li\u003e\n\u003cli\u003eWu JH, Hong PY, Liu WT. Quantitative effects of position and type of single mismatch on single base primer extension. J Microbiol Methods. 2009 Jun 1;77(3):267\u0026ndash;75. \u003c/li\u003e\n\u003cli\u003eStadhouders R, Pas SD, Anber J, Voermans J, Mes THM, Schutten M. The Effect of Primer-Template Mismatches on the Detection and Quantification of Nucleic Acids Using the 5\u0026prime; Nuclease Assay. J Mol Diagn. 2010 Jan 1;12(1):109\u0026ndash;17. \u003c/li\u003e\n\u003cli\u003eKim M, Smith WA, Van Hollebeke H. Personal communication. \u003c/li\u003e\n\u003cli\u003eAleem A, Akbar Samad AB, Slenker AK. Emerging Variants of SARS-CoV-2 And Novel Therapeutics Against Coronavirus (COVID-19). In: StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing; 2022 [cited 2023 Feb 22]. Available from: http://www.ncbi.nlm.nih.gov/books/NBK570580/\u003c/li\u003e\n\u003cli\u003eCDC. Centers for Disease Control and Prevention. 2020 [cited 2023 Feb 22]. Coronavirus Disease 2019 (COVID-19). Available from: https://www.cdc.gov/coronavirus/2019-ncov/variants/variant-classifications.html\u003c/li\u003e\n\u003cli\u003eMann T, Humbert R, Dorschner M, Stamatoyannopoulos J, Noble WS. A thermodynamic approach to PCR primer design. Nucleic Acids Res. 2009 Jul 1;37(13):e95\u0026ndash;e95. \u003c/li\u003e\n\u003cli\u003eHowson ELA, Orton RJ, Mioulet V, Lembo T, King DP, Fowler VL. GoPrime: Development of an In Silico Framework to Predict the Performance of Real-Time PCR Primers and Probes Using Foot-and-Mouth Disease Virus as a Model. Pathogens. 2020 Apr 20;9(4):303. \u003c/li\u003e\n\u003cli\u003eRambaut A, Holmes EC, O\u0026rsquo;Toole \u0026Aacute;, Hill V, McCrone JT, Ruis C, et al. A dynamic nomenclature proposal for SARS-CoV-2 lineages to assist genomic epidemiology. Nat Microbiol. 2020 Nov;5(11):1403\u0026ndash;7. \u003c/li\u003e\n\u003cli\u003eHealth C for D and R. U.S. Food and Drug Administration. FDA; 2021 [cited 2022 Nov 14]. Policy for Evaluating Impact of Viral Mutations on COVID-19 Tests. Available from: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/policy-evaluating-impact-viral-mutations-covid-19-tests\u003c/li\u003e\n\u003cli\u003eBioFire\u0026reg; COVID-19 Test [Internet]. BioFire Defense. [cited 2022 Nov 13]. Available from: https://www.biofiredefense.com/covid-19test/\u003c/li\u003e\n\u003cli\u003eBowen MD, Rollin PE, Ksiazek TG, Hustad HL, Bausch DG, Demby AH, et al. Genetic Diversity among Lassa Virus Strains. J Virol. 2000 Aug;74(15):6992\u0026ndash;7004. \u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTables are available in the Supplementary Files section.\u003c/p\u003e\n"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-bioinformatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"binf","sideBox":"Learn more about [BMC Bioinformatics](http://bmcbioinformatics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/binf","title":"BMC Bioinformatics","twitterHandle":"@BMC_Bioinformatics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"COVID-19, SARS-CoV-2, GISAID EpiCoV™, PCR, genomic variants, detection, in silico inclusivity, BioFire FilmArray System, in vitro diagnostic device, bioinformatics tool","lastPublishedDoi":"10.21203/rs.3.rs-5783329/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5783329/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eThe global Coronavirus Disease 2019 (COVID-19) pandemic highlighted the need to quickly diagnose infections to identify and prevent viral spread in the population. In response to the pandemic, BioFire Defense leveraged its PCR-based \u0026ldquo;lab-in-a-pouch\u0026rdquo; technology for expedited development of the BioFire\u0026reg; COVID-19 Test, a novel \u003cem\u003ein vitro\u003c/em\u003e diagnostic detecting SARS-CoV-2 nucleic acid in human samples. Following clearance of an \u003cem\u003ein vitro\u003c/em\u003e diagnostic device, regulatory bodies such as the U.S. Food and Drug Administration (FDA) require regular post market surveillance to monitor test performance against viral lineages circulating in the field, using predictive \u003cem\u003ein silico\u003c/em\u003e inclusivity evaluations. Exponential increases in the number of sequences deposited in bioinformatic repositories such as GISAID, during the pandemic, impeded progress in meeting these post market requirements. In response, BioFire Defense developed a new bioinformatic tool to overcome scalability problems and the loss of accuracy encountered with the standard inclusivity method.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eThe Coronavirus Monitoring for Inclusivity Tool (CoMIT) uses the Variant Sorter Algorithm to sidestep multiple sequence alignments, a significant barrier inherent in the standard inclusivity method. The implementation of CoMIT and its Variant Sorter Algorithm are described. Automated summary tables and visualizations from a typical inclusivity evaluation are presented. We report our approach to filter and display relevant information in the pipeline outputs using risk factors tied to test performance.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eBioFire Defense has developed CoMIT, an automated bioinformatic pipeline for efficient processing and reporting of variant inclusivity from the GISAID EpiCoV\u0026trade; repository. This tool ensures continuous and comprehensive post market evaluations of BioFire COVID-19 Test performance even from datasets large enough to impede standard inclusivity analyses. CoMIT\u0026rsquo;s low computational space complexity and modular code allow this tool to be generalized for inclusivity monitoring of multianalyte or single analyte tests with complex assay designs and/or highly variable targets. CoMIT\u0026rsquo;s databasing capabilities and metadata handling hold the potential for new investigations to improve readiness for future outbreaks.\u003c/p\u003e","manuscriptTitle":"CoMIT: A Bioinformatic Pipeline for Risk-based Prediction of COVID-19 Test Inclusivity","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-01-10 07:44:33","doi":"10.21203/rs.3.rs-5783329/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Accepted","date":"2025-01-10T16:18:20+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-01-08T12:57:02+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-01-08T12:54:20+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Bioinformatics","date":"2025-01-07T17:10:19+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-bioinformatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"binf","sideBox":"Learn more about [BMC Bioinformatics](http://bmcbioinformatics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/binf","title":"BMC Bioinformatics","twitterHandle":"@BMC_Bioinformatics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"3526ab0f-e210-4487-b0cc-fd92e63caae2","owner":[],"postedDate":"January 10th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-02-17T15:59:53+00:00","versionOfRecord":{"articleIdentity":"rs-5783329","link":"https://doi.org/10.1186/s12859-025-06046-y","journal":{"identity":"bmc-bioinformatics","isVorOnly":false,"title":"BMC Bioinformatics"},"publishedOn":"2025-02-12 15:57:09","publishedOnDateReadable":"February 12th, 2025"},"versionCreatedAt":"2025-01-10 07:44:33","video":"","vorDoi":"10.1186/s12859-025-06046-y","vorDoiUrl":"https://doi.org/10.1186/s12859-025-06046-y","workflowStages":[]},"version":"v1","identity":"rs-5783329","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5783329","identity":"rs-5783329","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.