Results
Characteristics of Haemophilus influenzae clinical isolates collection.
Of the ~1,600 H. influenzae genomes used in this study, ~ 1/3 were obtained from online
databases with the remainder sequenced in our laboratory, the majority of which were
obtained from the biorepository at the Center for Infectious Diseases and Immunology at
Rochester General Hospital Research Institute, maintained by two us (MEP and RK). The clinical
phenotypes explored in this study were: health state (sick or healthy), anatomical infection site,
and patient age. Our isolates are primarily from patients (70%) vs. healthy subjects (30%) (Fig
1A); infection type was by anatomic site to probe for tissue-specific adaptations (Fig 1B). We
also parsed for age: senior > 65 years; adult 18-65 years; child 2-18 years; and baby < 2 years of
age (Fig 1C). Gene groups (GG) were defined from the pan-genomic analyses as protein
sequences containing 75% or greater homology (Hogg et al 2007) resulting in 4,610 GG for
analysis.
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
7
Figure 1. Summary of Haemophilus influenzae clinical isolates and their genes. Percentage of all
isolates by: health state (A), infection type (B), and age (C). Gene groups of protein sequences
containing >75% similarity were analyzed. Shown are total number of sequence variants per gene
group (D), and total number of protein sequences (redundant sequences included) per gene group
(E), sorted by ascending. These values were then plotted together as unique variants vs total protein
sequences (F).
Figure 1
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
8
Within the se gene groups are variants; represented by each unique protein sequence. Total
variants for each gene group were plotted in ascending order ( Fig 1D ). Many gene groups
contained a single variant, and about half of all gene groups have less than 10 variants. At the
other end of the spectrum, a few gene groups contained many variants. Low variant total i s
consistent with such genes being highly conserved, as expected with core genes. A single
sequence variant can represent any number of strains that contain that sequence . Therefore,
total proteins represented were enumerated for each gene group and also plotted in ascending
order (Fig 1E). A large number of gene groups contain roughly 1,600 total sequences, indicating
roughly one per strain (1,600 total strains). Any gene groups with a value greater than 1,600
contain, on average, more than one copy per genome. However, when the two data sets were
combined as total protein sequences represented versus number of variants, the gene groups
with roughly 1,600 total sequences showed a larg e amount of sequence diversity ( Fig 1F ).
Surprisingly, there were no gene groups with only one variant that were present less than 600
times, suggesting that there is a large amount of genotypic diversity amongst all isolates, even in
core genes.
Process overview of machine learning model application and analysis
A pipeline was constructed to explore intragen ic diversity of H. influenzae for hidden properties
that correlate with various clinical phenotypes (Fig 2A). The crux of this algorithm is a previously
developed unsupervised machine learning language model ( Rives et al 2021 ). This model uses
only amino acid sequence information which is then converted to a numerical vector that reflects
biological aspects of the protein. This is done by inferring meaning of the amino acid from its
context within the protein, similar to words in a sentence.
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
9
B)
C)
E)
D)
Figure 2
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
10
Figure 2. Process overview and sample input and outputs. A) Flowchart of the application and
analysis of an unsupervised language model on protein sequences to identify clusters of variants
that correlate with clinical metadata. B) Modeling: protein sequences are processed by the model
and converted to numerica l vectors. C) Visualization: numerical vectors are shown in reduced
dimensions to visualize proximity. This step was not included in the main pipeline ; it is shown
here for visualization purposes only. D) Clustering: HDBScan was used to generate clusters from
multidimensional positioning E) Correlation with metadata: Chi squared analysis was performed
to check for correlations between cluster and clinical metadata
An example of this transformation is shown in Fig 2B. These numerical vectors are bulky and can
be simplified for visualization by dimension reduction, however this step was only done in this
study for visualization and was not included in the main pipeline to maintain reproduceable
results. Clustering occurred at the multidimensional level, prior to dimension reduction. We used
uniform manifold approximation and projection (UMAP) to reduce the vector to two dimensions
so that it can be plotted for visualization ( Fig 2C). Each point represents a unique protein
sequence within the gene group. Clustering was performed using HDBScan with a minimum
threshold of six variants for a cluster ( Fig 2D). Variants identified as not belonging to a cluster
(noise) were dropped. B ecause we were interested in comparing multiple clusters within one
gene group , gene groups with only one cluster were discontinued. A gene group with two or
more clusters then had additional data for total protein counts appended. This is because a single
sequence variant often represents many occurrences of that sequence. Clinical metadata was
also retrieved and appended to the protein data table, corresponding to the strain harboring
each protein. A contingency table of cluster and the category of interest (ex: infection type) was
created and a C hi squared test of independence was calculated. A p- value of 4x10
-5 was
determined to be statistically significant based on a Bonferroni correction corresponding to the
number of gene groups that were analyzed (groups contain ing more than one cluster). The
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
11
resulting significant genes are reported herein, and further analysis was conducted depending
on the category of interest.
Correlations of cluster and health state
To evaluate the machine learning model’s ability to find meaningful differences amongst variants,
we looked for correlations between variant cluster and health state. Health state here refers to
isolates from healthy or sick patients. Isolates with an unknown health state were removed from
the analysis. For gene groups with two or more clusters, a contingency table was generated for
cluster and health state to enumerate counts for each combination of cluster number and health
state. To determine the relationship between cluster and health state, a C hi squared test of
independence was performed. If the p-value was less than 4x10
-5 the gene group was considered
significant, as discussed in the previous section. 79 gene groups were found to be significant for
dependence of cluster and heath state (Table 3.1). Of those significant gene groups, nine had at
least one cluster with greater than 95% isolates from sick patients, and five of those had a cluster
that was 99% or higher (Fig 3A). Gene group hcat had less than 1% sick . This analysis identified
clusters enriched for a particular health state. To identify which cellular processes were
represented by the gene groups that were significantly associated with one of the test states ,
COG categories were predicted by functional domain analysis using eggNOG- mapper
(Cantalapiedra et al 2021) to predict. The most prevalent amino acid sequence for each gene
group was used for this analysis. Unknown function was the most represented COG category (Fig
3B), although it was also the most numerous category amongst all of the gene groups. When
normalized for change in prevalence, defense mechanisms and inorganic ion transport categories
showed the largest change (Fig 3C). A contingency table for gene group
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
12
tbp1 shows clusters containing predominantly sick isolates (Fig 4A). tbp1 was no longer
significant when strain duplicates were removed from each cluster, yet it remains of interest
because of its high percentage of sick isolates in many clusters (Fig 4B).
Figure 3. Correlations of cluster and health
state. A) Significant gene groups that had at
least one cluster with >99% sick or healthy
isolates. B) COG category analysis of all
significant gene groups. C) Change in
prevalence of each COG category.
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
13
Figure 4. A) Contingency table of cluster and health state for gene group tbpA . B Plotting tbpA
clusters by percentage of isolates that are sick.
Correlations between clusters and infection type
All identified clusters were then overlaid with the tissue/organ (ear, lung, carriage, eye, invasive
and other) from which the corresponding isolate was recovered, and a contingency table was
created of number of clusters versus infection type. This was used to conduct a Chi squared test
of independence. Using the same p -value as above , 285 gene groups were found to be
significantly associated with the tissue of origin (Table 2). Four of the gene groups with significant
associations contained a cluster with ≥ 99% prevalence for one or more origin sites, ten from the
lung and two NP carriage (Figure 5A). Three contained a cluster with ≥ 99% prevalence in the
lung, and the hcat gene group was predominantly associated with carriage. COG category analysis
of significant gene groups showed many with: unknown function; cell wall synthesis; and amino
acid transport and metabolism ( Figure 5B). After normalization to the greater gene group
collection, only the latter two showed an increased in prevalence (Figure 5C), both of which are
common targets of antibiotics. Tbp1 was a top hit again; two clusters had a 100% prevalence in
the lung, and a third had 93%. Analysis of the contingency table for tbpA clusters versus infection
B A
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
14
type or health state, it was observed that four clusters from tbpA are predominantly associated
with isolates from the diseased lung, while the remaining cluster resembled the greater collection
(Fig 6A&B)
Figure 5. Correlations of cluster and
infection type. A) Significant gene
groups that had at least one cluster with
>99% one infection type. B) COG
category analysis of all significant gene
groups. C) Change in prevalence of each
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
15
Figure 6. Correlations of cluster and infection type. A) Contingency table of tbpA cluster and infection
type. B) Plotting tbpA clusters by percentage found in the lung.
Analysis of tbpA
tbpA was selected for further exploration because four of its five clusters contained a higher
percentage of isolates that were from lung infections. tbpA encodes a n outer membrane
transferrin binding protein, Tbp1 ; that serves as a virulence factor involved in iron uptake from
the host. It is essential for H. influenzae growth on media with human transferrin as the sole iron
source (Gray-Owen et al 1995). Our analysis showed that gene group tpbA consists of 421 unique
variants from 1,782 total protein sequences. Modeling and clustering of the unique amino acid
sequences resulted in five clusters. Cluster 3 includes 368 protein variants and 1596 total
proteins, and in terms of infection type prevalence, resembles the greater collection. Clusters 1
and 2 represent 21 and 30 strains respectively, all of which were isolated from patients with lung
illnesses. Clusters 4 and 0 are also primarily found in isolat es from the lung , but each of these
clusters consist of < 10 strains. Reducing dimensions of the numerical vectors shows the spatial
proximity of the clusters in 2 dimensions, down from 1,280 (Fig 7A). Some differences among the
clusters show low resolution due to clustering prior to dimension reduction. However, the overall
image shows a general separation of clusters 0, 1, 2, and 4 from cluster 3. Clusters 0, 1, 2, and 4
A B
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
16
displayed a reduced average protein length compared to cluster 3 (Fig 7B) suggesting a truncation
of the gene. A pfam domain analysis of the variants in each cluster revealed a significant decrease
in the prevalence of either or both of the receptor and plug domains, from the isolates recovered
from lung patients (Fig 7C). We noted that all strains present in clusters 0, 1, 2, and 4 were also
present in cluster 3. Therefore, it is likely that these clusters are truncated duplications that
occurred in strains that have a full-length copy of tbpA. Due to the small number of isolates found
in these clusters we sought to determine if this phenomenon was confined to (possibly related)
strains obtained from a single collaborator, but this was not the case as these clusters consisted
of isolates from two unrelated cohorts: cystic fibrosis isolates from Seattle, Washington USA, and
COPD isolates from Madrid Spain. This observation suggests that this phenomenon of probable
tbpA partial duplication has occurred repeatedly within patients with different lung diseases
Figure 7. Analysis of Tbp1. A) clustering of T bp1 sequence variants as classified by the model. B)
Average amino acid length of Tbp1 by cluster. C) Prevalence of Pfam domains by cluster
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
17
Correlations of cluster and patient age
Finally, we correlated variant clusters w ith patient age. 286 gene groups were found to be
significant by a Chi squared analysis between cluster number and age group (Table 3). Of the
significant gene groups, only three had at least one cluster with greater than 99% prevalence of
isolates belonging to a single age group (Fig 8A). These were annotated as: hgpB, a hemoglobin-
haptoglobin binding protein; g roup_287, a predicted glycosyltransferase family 8 protein ; and
msfA_1, a hypothetical related to a previously characterized virulence factor by our group (Kress-
Bennett et al 2016) . A COG category analysis revealed that the significant gene groups are
enriched in unknown function, amino acid transport and metabolism, translation, and cell wall
synthesis ( Fig 8B). After normalization, amino acid transport and metabolism, inorganic ion
transport and metabolism showed the largest increase in prevalence (Fig 8C).
Figure 8. Correlations of clusters and patient age. A) Significant gene groups that had at least
one cluster with >99% in one age group. B) COG category analysis of all significant gene groups.
C) Change in prevalence of each COG category.
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
18
References
Ehrlich GD, Hiller NL, Hu FZ. 2008. What makes Pathogens Pathogenic. Genome Biology 9: 225
Duell BL, Su YC, Riesbeck K. 2016. Host-pathogen interactions of nontypeable Haemophilus influenzae:
from commensal to pathogen. FEBS Lett 590: 3840-53
Khattak ZE, Anjum F. 2023. Haemophilus influenzae infection. In: StatPearls [Internet]. Treasure Island
(FL): StatPearls Publishing; 2025
Wen S, Feng D, Chen D, Yang L, Xu Z. 2020. Molecular epidemiology and evolution of Haemophilus
influenzae. Infect Genet Evol 80: 104205.
Tsang RS. Serotyping and population genetics of invasive Haemophilus influenzae. 2008. J Clin Microbiol
46: 1159.
Zhang H, Patenaude B, Zhang H, Jit M, Fang H. 2024. Global vaccine coverage and childhood survival
estimates: 1990-2019. Bull World Health Organ.
102: 276-287.
Soeters HM, Blain A, Pondo T, Doman B, Farley MM, Harrison LH, Lynfield R, Miller L, Petit S, Reingold A, et al.
2018. Current Epidemiology and Trends in Invasive Haemophilus influenzae Disease-United States,
2009-2015. Clin Infect Dis 67: 881-9.
Soeters HM, Oliver SE, Plumb ID, Blain AE, Zulz T, Simons BC, Barnes M, Farley MM, Harrison LH,
Lynfield R, et al. 2021. Epidemiology of Invasive
Haemophilus influenzae Serotype a Disease-
United States, 2008-2017. Clin Infect Dis 73: e371-e9
Shuel M, Knox N, Tsang RSW. 2021. Global population structure of Haemophilus influenzae serotype a
(Hia) and emergence of invasive Hia disease: capsule switching or capsule replacement? Can J
Microbiol. 2021 67: 875-884.
Topaz N, Tsang R, Deghmane AE, Claus H, Lâm TT, Litt D, Bajanca-Lavado MP, Pérez-Vázquez M, Vestrheim
D, Giufrè M, et al. 2022. Phylogenetic Structure and Comparative Genomics of Multi-National
Invasive Haemophilus influenzae Serotype a Isolates. Front Microbiol 13: 856884.
Ulanova M, Tsang RS, Goldfarb DM, Smieja M, Huska B, Luinstra K, Le Saux N. 2024. Prevalence
of Haemophilus influenzae in the nasopharynx of children from regions with varying incidence of
invasive H. influenzae serotype a disease: Canadian Immunization Research Network (CIRN) study. Int J
Circumpolar Health
83: 2371111.
Whyte KE, Hoang L, Sekirov I, Shuel ML, Hoang W, Tsang RS. 2020. Emergence of a clone of
invasive fucK-negative serotype e Haemophilus influenzae in British Columbia. J Assoc Med
Microbiol Infect Dis Can 5: 29-34.
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
29
Takla A, Schonfeld V, Claus H, Krone M, An der Heiden M, Koch J, Vogel U, Wichmann O, Lâm TT. 2020.
Invasive Haemophilus influenzae infections in Germany after the introduction of routine childhood
immunization, 2001-2016. Open Forum Infect Dis 7: ofaa444.
Agrawal A, Murphy TF. 2011. Haemophilus influenzae infections in the H. influenzae type b conjugate
vaccine era. J Clin Microbiol 49: 3728-32.
Hu YL, Lee PI, Hsueh PR, Lu CY, Chang LY, Huang LM, Chang TH, Chen JM. 2021. Predominant role of
Haemophilus influenzae in the association of conjunctivitis, acute otitis media and acute bacterial
paranasal sinusitis in children. Sci Rep 11: 11.
Post JC, Preston RA, Aul JJ, Larkins-Pettigrew M, Rydquist-White J, Anderson KW, Wadowsky RM, Reagan DR,
Walker ES, Kingsley LA, et al. 1995. Molecular analysis of bacterial pathogens in otitis media with
effusion. JAMA 273: 1598-604.
Rayner MG Zhang Y Gorry MC Chen Y Post JC Ehrlich GD. 1998. Evidence of bacterial metabolic activity in
culture-negative otitis media with effusion. JAMA 279: 296-299, 1998.
Ehrlich GD, Veeh R, Wang X, Costerton JW, Hayes JD, Hu FZ, Daigle BJ, Ehrlich MD, Post JC. 2002. Mucosal
Biofilm Formation on Middle-ear Mucosa in the Chinchilla Model of Otitis Media. JAMA 287: 1710-
1715.
Shen K, Wang X, Post JC, Ehrlich GD. 2003. Molecular and Translational Research Approaches for
the study of Bacterial Pathogenesis in Otitis Media. In: Evidence-based Otitis Media, 2nd ed. (ed
Rosenfeld R and Bluestone CD) pp91-119. B.C. Decker Inc, Hamilton.
Hall-Stoodley L, Hu FZ, Stoodley P, Nistico L, Link TR, Burrows A ,Post JC, Ehrlich GD, Kerschner KE.
2006. Direct Detection of Bacterial Biofilms on the Middle-Ear Mucosa of Children with Chronic Otitis
Media. JAMA 296: 202-211.
Nistico L, Gieseke A, Kreft R, Coticchia JM, Burrows A, Khampang P, Liu Y, Kerschner JE, Post JC. et al.
2
011. Pathogenic Biofilms in Adenoids: a Reservoir for Persistent Bacteria. J Clin Micro 490: 1411-20.
Short B, Carson S, Devlin AC, Reihill JA, Crilly A, MacKay W, Ramage G, Williams C, Lundy FT,
McGarvey LP, et al. 2021. Non-typeable Haemophilus influenzae chronic colonization in chronic
obstructive pulmonary disease (COPD). Crit Rev Microbiol 47: 192-205.
Murphy TF, Brauer AL, Schiffmacher AT, Sethi S. 2004. Persistent colonization by
Haemophilus influenzae in chronic obstructive pulmonary disease. Am J Respir Crit Care Med
170: 266- 72.
Moleres J, Fernández-Calvet A, Ehrlich RL, Martí S, Pérez-Regidor L, Euba B, Rodríguez-Arce I,
Balashov S, Cuevas E, Liñares J, et al. 2018. Antagonistic pleiotropy in a bifunctional fatty
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
30
acid transporter during adaptation of Haemophilus influenzae to chronic lung infection
associated with COPD. mBio 9: e01176-18.
Starner TD, Zhang N, Kim G, Apicella MA, McCray PB, Jr. Haemophilus influenzae forms
biofilms on airway epithelia: implications in cystic fibrosis. Am J Respir Crit Care Med 174:
213-20.
Martin D, Dbouk RH, Deleon-Carnes M, del Rio C, Guarner J. 2013. Haemophilus influenzae
acute endometritis with bacteremia: case report and literature review. Diagn Microbiol
Infect Dis. 76: 235-6.
Heise T, Langereis JD, Rossing E, de Jonge MI, Adema GJ, Büll C, Boltje TJ. 2018. Selective
inhibition of sialic acid-based molecular mimicry in Haemophilus influenzae abrogates serum
resistance. Cell Chem Biol 25: 1279-1285.
Kress-Bennett J, Hiller NL, Eutsey R, Powell E, Longwell MJ, Hillman T, Blackwell T, Byers B,
Post JC, Hu F. 2016. Identification and characterization of msf, a novel virulence factor in
Haemophilus influenzae. PLoS ONE 11: e0149891.
Ehrlich GD, Hu ZF, Post JC. 2004. Role for Biofilms in Infectious Disease. In: Microbial
Biofilms. Eds. Ghannoum M, and O’Toole GA) pp332-358. ASM Press, Washington, DC.
Post JC, Stoodley P, Hall-Stoodley L, Ehrlich GD. 2004. The Role of Biofilms in
Otolaryngologic Infections. Current Opinion Otolaryngology Head Neck Surgery. 12: 185-190.
Schaudinn C, Stoodley P, Kainović A, O’Keeffe T, Costerton JW, Robinson D, Baum M, Ehrlich
GD, Webster P. 2007. Bacterial Biofilms, Other Structures Seen as Mainstream Concepts Get
used to it: bacterial microcolonies form regular shapes, such as nanowires or honeycomb-
like structures. Microbe 2: 231-237.
Kerschner JE, Erdos G, Hu FZ, Burrows A, Cioffi J, Hayes J, Keefe R, Janto B, Post JC, Ehrlich
GD. 2010. Characterization of normal and Haemophilus influenzae-infected mucosal cDNA
libraries in chinchilla middle ear mucosa. Ann Otol Rhinol Laryngol 119: 270-278.
Devaraj A, Buzzo J, Rocco CJ, Bakaletz LO, Goodman SD. 2018. The DNABII family of proteins
is comprised of the only nucleoid associated proteins required for nontypeable Haemophilus
influenzae biofilm structure. Microbiologyopen 7: e00563.
Buzzo JR, Devaraj A, Gloag ES, Jurcisek JA, Robledo-Avila F, Kesler T, Wilbanks K, Mashburn-
Warren L, Balu S, Wickham J, et al. 2021. Z-form extracellular DNA is a structural
component of the bacterial biofilm matrix. Cell 184: 5740-5758.e17.
Asbell PA, DeCory HH. 2018. Antibiotic resistance among bacterial conjunctival pathogens
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
31
collected in the Antibiotic Resistance Monitoring in Ocular Microorganisms (ARMOR)
surveillance study. PLoS One 13: e0205814.
Jansen WT, Verel A, Beitsma M, Verhoef J, Milatovic D. 2006. Longitudinal European
surveillance study of antibiotic resistance of Haemophilus influenzae. J Antimicrob
Chemother 58: 873-7.
Atack JM, Day CJ, Poole J, Brockman KL, Timms JRL, Winter LE, Haselhorst T, Bakaletz LO,
Barenkamp SJ, Jennings MP. 2020. The Nontypeable Haemophilus influenzae Major Adhesin
Hia Is a Dual-Function Lectin That Binds to Human- Specific Respiratory Tract Sialic Acid
Glycan Receptors. mBio. 11: e02714-20
Fernández-Calvet A, Euba B, Gil-Campillo C, Catalan-Moreno A, Moleres J, Martí S, Merlos A,
Langereis JD, García-Del Portillo F, Bakaletz LO, Ehrlich GD, et al. 2021. Phase variation in
HMW1a controls a phenotypic switch in Haemophilus influenzae associated with
pathoadaptation during persistent infection. mBio 12: e0078921.
Murphy TF, Kirkham C, D'Mello A, Sethi S, Pettigrew MM, Tettelin H. 2023. Adaptation of
nontypeable Haemophilus influenzae in human airways in COPD: genome rearrangements
and modulation of expression of HMW1 and HMW2. mBio. 14: e0014023
Morton DJ, Seale TW, Bakaletz LO, Jurcisek JA, Smith A, VanWagoner TM, Whitby PW, Stull
TL. 2009. The heme- binding protein (HbpA) of Haemophilus influenzae as a virulence
determinant. Int J Med Microbiol 299: 479-88.
Gray-Owen SD, Loosmore S, Schryvers AB. 1995. Identification and characterization of
genes encoding the human transferrin-binding proteins from Haemophilus influenzae. Infect
Immun 64: 1201-10.
Ren D, Walker AN, Daines DA. 2012. Toxin-antitoxin loci vapBC-1 and vapXD contribute to
survival and virulence in nontypeable Haemophilus influenzae. BMC Microbiol 12: 263.
Daines DA, Jarisch J, Smith AL. 2004. Identification and characterization of a nontypeable
Haemophilus influenzae putative toxin-antitoxin locus. BMC Microbiol 4: 30.
Ishak N, Tikhomirova A, Bent SJ, Ehrlich GD, Hu FZ, Kidd SP. 2014. There is a specific
response to pH by isolates of Haemophilus influenzae and this has a direct influence on
biofilm formation. BMC Microbiol 14: 47.
Fleischmann RD, Adams MD, White O, Clayton RA, Kirkness EF, Kerlavage AR, Bult CJ, Tomb
JF, Dougherty BA, Merrick JM, et al. 1995. Whole-genome random sequencing and
assembly of Haemophilus influenzae Rd. Science 269: 496-512.
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
32
Harrison A, Dyer DW, Gillaspy A, Ray WC, Mungur R, Carson MB, Zhong H, Gipson J, Gipson
M, Johnson LS, et al. 2005. Genomic sequence of an otitis media isolate of nontypeable
Haemophilus influenzae: comparative study with H. influenzae serotype d, strain KW20. J
Bacteriol 187: 4627-36.
Ehrlich GD, Hu FZ, Shen K, Stoodley P, Post JC. 2005. Bacterial plurality as a general
mechanism driving persistence in chronic infections. Clin Orthop Relat Res 437: 20-24.
Buchinsky FJ, Forbes M, Hayes J, Hu FZ, Greenberg P, Post JC, Ehrlich GD. 2007. Phenotypic
plurality among clinical strains of nontypeable Haemophilus influenzae determined by
symptom severity in the chinchilla model of otitis media. BMC Micro 7: 56.
Ehrlich GD, Ahmed A, Earl J, Hiller NL, Costerton JW, Stoodley P, Post JC, DeMeo P, Hu FZ.
2010. The Distributed Genome Hypothesis as a Rubric for Understanding Evolution in situ
During Chronic Infectious Processes. FEMS Immunol Med Micro 59: 269-79.
Nistico L, Earl J, Hiller L, Ahmed A, Retchless A, Janto B, Costerton JC, Hu FZ, Ehrlich GD.
2014. Using the Core and Supra Genomes to Determine Diversity and Natural Proclivities
among Bacterial Strains. 2014. In: Application of Molecular Microbiological Methods eds.
Skovhus TL, Caffrey S, Hubert C) pp Caister Academic Press, Norfolk, UK.
Hammond JA, Gordon EA, Socarras KM, Mell JC, Ehrlich GD. 2020. Beyond the Pangenome:
Current Perspectives on the Functional and Practical Outcomes of the Distributed Genome
Hypothesis. Biochemical Society Transactions 48: 2437-2455.
Innamorati KA, Earl JP, Aggarwal SD, Ehrlich GD, Hiller NL. 2020. The bacterial guide to
designing a diversified portfolio. In: The Pangenome (eds. Tettelin H and Medinie D) pp 51-
88. Springer, Cham, Switzerland.
Shen, K., Antalis, P., Gladitz, J., Sayeed, S., Ahmed, A., Yu, S., Hayes, J., Johnson, S., Dice, B.,
Dopico, et al. 2005. Identification, distribution, and expression of novel (nonRd) genes in
ten clinical isolates of nontypeable Haemophilus influenzae. Infect Immun 73: 3479-3491.
Gladitz J, Antalis P, Hu FZ Post JC Ehrlich GD. 2005. Codon usage comparison of novel genes
in clinical isolates of Haemophilus influenzae. Nucleic Acids Res 33: 3644-58.
Hogg JS, Hu FZ, Janto B, Boissy R, Hayes J, Keefe R, Post JC, Ehrlich GD. 2007.
Characterization and modeling of the Haemophilus influenzae core and supragenomes
based on the complete genomic sequences of Rd and 12 clinical nontypeable strains.
Genome Biol 8: R103.
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
33
Hall HG, Ehrlich GD, Hu FZ. 2010. Pan-genome analysis provides much higher strain typing
resolution than does MLST. Microbiology 156: 1060-1068.
Boissy R, Ahmed A, Janto B, Earl J, Hall BJ, Hogg J, Pusch GD, Hiller NL, Powell E, Hayes J.
2011. Comparative supragenomic analyses among the pathogens Staphylococcus aureus,
Streptococcus pneumoniae, and Haemophilus influenzae using a modification of the finite
supragenome model. BMC Genomics 12: 187.
Eutsey R, Hiller NL, Earl J, Janto B, Dahlgren M, Ahmed A, Powell A, Schultz M, Gilsdorf J,
Zhang L. 2013. Design and validation of a supragenome hybridization array for
determination of the genomic content of Haemophilus influenzae isolates. BMC Genomics
14: 484.
Kosar K. 2023. Using accessory gene content to predict the clinical isolation source of
Haemophilus influenzae. Master’s Thesis Drexel University
Lopatkin AJ, Bening SC, Manson AL, Stokes JM, Kohanski MA, Badran AH, Earl AM, Cheney
NJ, Yang JH, Collins JJ. 2021. Clinically relevant mutations in core metabolic genes confer
antibiotic resistance. Science. 371: 6531.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I.
2023. Attention is all you need [Preprint]. arXiv:1706.03762v7
Zhang TH, Hasib MM, Chiu YC, Han ZF, Jin YF, Flores M, Chen Y, Huang Y. 2022. Transformer
for Gene Expression Modeling (T-GEM): An Interpretable Deep Learning Model for Gene
Expression-Based Phenotype Predictions. Cancers (Basel) 14: 4763.
Du, W., Zhao, L., Diao, K. Zheng Y, Yang Q, Zhu Z, Zhu X, Tang D. 2025. A versatile
CRISPR/Cas9 system off-target prediction tool using language model. Commun Biol 8: 882.
Dampier W, Link RW, Earl JP, Collins M, De Souza DR, Koser K, Nonnemacher MR, Wigdahl B.
2022. HIV-Bidirectional Encoder Representations from Transformers: a set of pretrained
transformers for accelerating HIV deep learning tasks. Front Virol 17: 880618
Wiatrak M, Viñas Torné R, Ntemourtsidou M, Dinan AM, Abelson DC, Arora D, Brbić M,
Weimann A, Floto RA. 2025. A contextualised protein language model reveals the functional
syntax of bacterial evolution [Preprint]. bioRxiv. 20: 665723.
Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, Guo D, Ott M, Zitnick CL, Ma J, Fergus R. 2021.
Biological structure and function emerge from scaling unsupervised learning to 250 million
protein sequences. Proc Natl Acad Sci U S A 118: e2016239118.
Cantalapiedra CP, Hernandez-Plaza A, Letunic I, Bork P, Huerta-Cepas J. 2021. eggNOG-
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
34
mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the
Metagenomic Scale. Mol Biol Evol 38: 5825-9.
Wolf, Thomas, et al. 2020. Transformers: State-of-the-Art Natural Language Processing.
arXiv:1910.03771. https://doi.org/10.48550/arXiv.1910.03771
Traxler MF, Summers SM, Nguyen HT, Zacharia VM, Hightower GA, Smith JT, et al. 2008. The
global, ppGpp-mediated stringent response to amino acid starvation in Escherichia coli. Mol
Microbiol 68: 1128-48.
Yahara K, Didelot X, Jolley KA, Kobayashi I, Maiden MC, Sheppard SK, et al. The Landscape of
Realized Homologous Recombination in Pathogenic Bacteria. Mol Biol Evol 33: 456-71.
Poje G, Redfield RJ. 2003. Transformation of Haemophilus influenzae. In: Methods in
Molecular Medicine (Vol. 71).
Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, Lesin VM, Nikolenko SI,
Pham S, Prjibelski AD, et al. 2012. SPAdes: A new genome assembly algorithm and its
applications to single-cell sequencing. J Comp Bio, 19 (5).
Chin C S, Alexander DH, Marks P, Klammer AA, Drake J, Heiner C, Clum A, Copeland A,
Huddleston J, Eichler EE, et al. 2013. Nonhybrid, finished microbial genome assemblies from
long-read SMRT sequencing data. Nature Methods, 10 (6).
Hunt M, Silva ND, Otto TD, Parkhill J, Keane JA, Harris SR. 2015. Circlator: automated
circularization of genome assemblies using long sequencing reads. Genome Biol. 29: 294.
De Chiara M, Hood D, Muzzi A, Pickard DJ, Perkins T, Pizza, M, Dougan G, Rappuoli R, Moxon
ER, Soriani M, Donati C. 2014. Genome sequencing of disease and carriage isolates of
nontypeable Haemophilus influenzae identifies discrete population structure. PNAS USA
111: 5439–5444.
Dröge J, Gregor I, McHardy AC. 2015. Taxator-tk: Precise taxonomic assignment of
metagenomes by fast approximation of evolutionary neighborhoods. Bioinformatics 31:
817–824.
Seemann T. 2014. Prokka: Rapid prokaryotic genome annotation. Bioinformatics, 30: 2068–
2069.
Watts SC, Holta KE 2019. HICAP: In silico serotyping of the haemophilus influenzae capsule
locus. J Clin Micro 57(6).
Page AJ, Cummins CA, Hunt M, Wong VK, Reuter S, Holden MTG, Fookes M, Falush D, Keane
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint
35
JA, Parkhill J. 2015. Roary: Rapid large-scale prokaryote pan genome analysis.
Bioinformatics 31: 3691–3693
Tonkin-Hill G, MacAlasdair N, Ruis C, Weimann A, Horesh G, Lees JA, Gladstone RA, Lo S,
Beaudoin C, Floto RA, et al. 2020. Producing polished prokaryotic pangenomes with the
Panaroo pipeline. Genome Biology 21(1).
Kolde R. 2019. Package ‘pheatmap’: Pretty heatmaps. Version 1.0.12.
Löytynoja A. 2014. Phylogeny-aware alignment with PRANK. Methods in Molecular Biology,
1079.
Stamatakis A. 2014. RAxML version 8: A tool for phylogenetic analysis and post-analysis of
large phylogenies. Bioinformatics 30(9).
Paradis E, Schliep K. 2019. Ape 5.0: An environment for modern phylogenetics and
evolutionary analyses in R. Bioinformatics 35: 526–528.
Xu S, Li L, Luo X, Chen, M, Tang W, Zhan L, Dai Z, Lam TT, Guan Y, Yu G. 2022. Ggtree: A
serialized data object for visualization of a phylogenetic tree and annotation data. IMeta,
1(4).
Jolley KA, Chan MS, & Maiden MCJ. 2004. mlstdbNet - Distributed multi-locus sequence
typing (MLST) databases. BMC Bioinformatics 5.
Lees JA, Harris SR, Tonkin-Hill G, Gladstone RA, Lo SW, Weiser JN, Corander J, Bentley SD,
Croucher NJ. 2019. Fast and flexible bacterial genomic epidemiology with PopPUNK.
Genome Research 29: 304–316.
(which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for this preprintthis version posted March 15, 2026. ; https://doi.org/10.64898/2026.03.12.711436doi: bioRxiv preprint