Abstract
In this work, we analyzed 184 metagenomes from sub-Saharan Africa to characterize the functional landscape of ARGs using a combination of homology-based annotation and protein language model embeddings. We obtained 5,066 high-confidence ARG protein sequences from the African metagenomes, which we compared with 6,052 reference ARGs from CARD using embeddings generated with the ESM-2 protein language model. Additionally, we used a random forest classifier to determine the role of amino acid sequence and physiochemical features in discriminating ARGs and non-ARG sequences. The curated dataset revealed a predominance of ESKAPE pathogens in the resistome. β-lactam resistance was the most prevalent functional class, accounting for 4,046 ARG assignments (79.87%). At the country level, Burkina Faso, Malawi, and Benin exhibited the highest ARG hits per sample, thus demonstrating geographic heterogeneity in ARG burden. Protein language modeling demonstrated that African ARG sequences largely occupied the same functional subspace as globally curated CARD proteins. Supervised machine learning based on protein compositional and physicochemical features achieved high discriminatory performance between ARGs and non-ARGs (accuracy = 0.94, ROC-AUC = 0.99). Feature importance analysis identified amino acid composition, protein length, molecular weight, and isoelectric point as key discriminators, with statistically significant differences between ARG and non-ARG proteins (Mann-Whitney U test, p < 0.001). These results suggest that ARGs in African metagenomes are shaped primarily by ecological filtering and antibiotic selection pressure rather than by the emergence of novel resistance functions. This work provides a functional baseline for AMR surveillance in Africa and highlights the value of protein language models for resistome-scale analyses.
Full text
2,387 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
In this work, we analyzed 184 metagenomes from sub-Saharan Africa to characterize the functional landscape of ARGs using a combination of homology-based annotation and protein language model embeddings. We obtained 5,066 high-confidence ARG protein sequences from the African metagenomes, which we compared with 6,052 reference ARGs from CARD using embeddings generated with the ESM-2 protein language model. Additionally, we used a random forest classifier to determine the role of amino acid sequence and physiochemical features in discriminating ARGs and non-ARG sequences. The curated dataset revealed a predominance of ESKAPE pathogens in the resistome. β-lactam resistance was the most prevalent functional class, accounting for 4,046 ARG assignments (79.87%). At the country level, Burkina Faso, Malawi, and Benin exhibited the highest ARG hits per sample, thus demonstrating geographic heterogeneity in ARG burden. Protein language modeling demonstrated that African ARG sequences largely occupied the same functional subspace as globally curated CARD proteins. Supervised machine learning based on protein compositional and physicochemical features achieved high discriminatory performance between ARGs and non-ARGs (accuracy = 0.94, ROC-AUC = 0.99). Feature importance analysis identified amino acid composition, protein length, molecular weight, and isoelectric point as key discriminators, with statistically significant differences between ARG and non-ARG proteins (Mann-Whitney U test, p < 0.001). These results suggest that ARGs in African metagenomes are shaped primarily by ecological filtering and antibiotic selection pressure rather than by the emergence of novel resistance functions. This work provides a functional baseline for AMR surveillance in Africa and highlights the value of protein language models for resistome-scale analyses.
Competing Interest Statement
The authors have declared no competing interest.
Footnotes
Data Availability
All sequencing accessions analyzed in this study are publicly available from their respective ENA and SRA repositories. Processed antimicrobial resistance gene (ARG) hit data are available for download at https://www.bio-africa.org/downloads. All scripts used for data processing, feature extraction, embedding generation, and machine-learning analyses are openly available at https://github.com/ngangao/AfriARM.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.