Life Identification Numbers: A bacterial strain nomenclature approach
preprint
OA: gold
CC-BY-4.0
Abstract
Unified strain taxonomies are needed for the epidemiological surveillance of bacterial pathogens and international communication in microbiological research. Core genome multilocus sequence typing (cgMLST) holds great promise for standardized high-resolution strain genotyping. However, this approach faces challenges including classification instability and disconnection of new nomenclature from widely adopted classical MLST identifiers. This essay discusses the cgMLST-based Life Identification Number (LIN) method, recently proposed as a stable multilevel strain taxonomy system applicable to most bacterial pathogens. We describe how LIN codes are implemented and used in practice for precise strain definitions and epidemiological tracking. Glossary Multilocus sequence typing (MLST) A genotyping method applied mostly to microbial strains to study population structure and epidemiology, based on comparing the nucleotide sequences of a small number (typically seven) of housekeeping protein-coding genes. In MLST, allele numbers are assigned to each sequence variant (allele) of a given gene. The MLST genotype of a bacterial strain is defined by the combination of the allele numbers observed at the genes that are included in the genotyping scheme. A sequence type (ST) is assigned to each unique combination of alleles, called an MLST profile. MLST was invented in 1998 and became a de-facto standard taxonomy of bacterial strains, albeit at low resolution. Core genome MLST An extension of MLST that analyzes sequence variation across hundreds to thousands of conserved (core) genes, shared by all strains of a species, providing higher resolution typing for genomic epidemiology and evolutionary studies. cgMLST schemes typically comprise 2000 to 4000 genes, depending on the genome size and genetic variation (in terms of presence/absence of genes) within bacterial species. A core genome sequence type (cgST) can be assigned to unique cgMLST profiles, i.e., a unique combination of cgMLST allelic numbers. Whole Genome Sequencing (WGS) A method that determines the complete DNA sequence of an organism’s genome in a single process, providing comprehensive information for comparative genetic analyses based on cgMLST or other analytic methods. Single Nucleotide Polymorphisms (SNPs) Variations at a single base position in the DNA sequence among individuals isolates, strains or species, used as genetic markers for studying for example, evolutionary relationships or strain identity. Average nucleotide identity (ANI) A measure of genomic similarity between two organisms, calculated as the average percentage of identical nucleotides in orthologous genomic regions; commonly used to assess species-level relatedness in prokaryotes. Taxonomy Here, we apply the word taxonomy to bacterial strains as a system of classifying, naming and identifying strains based on shared genetic characteristics as defined by e.g., cgMLST.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-4.0