{"paper_id":"274e1978-6e08-46a5-b5e6-d5d175a19e22","body_text":"Abstract\nPredicting gene function is a pivotal and challenging step in genomic and metagenomic data analysis. Current automatic annotation tools typically rely on the single most similar sequence from the query database. The sparsity of data per annotation makes it challenging to confidently assign gene function for underrepresented genes. Here, we present a contrastive learning framework for functional annotation. FAMUS (Functional Annotation Method Using Supervised contrastive learning) compares query sequences to profile Hidden Markov Model databases and transforms the similarity scores into a condensed vector space that minimizes the distance of proteins from the same family. The similarity scores of a query to all profiles are used for its representation instead of considering only the top-ranking hit. In a protein family assignment task, FAMUS outperformed KEGG’s native KofamScan for KEGG Orthology annotation and InterPro’s InterProScan for PANTHER family annotation. We thus created four protein annotation models using protein families from the KEGG Orthology, InterPro family, OrthoDB, and EggNOG databases. All four models are available as a conda package and via our user-friendly web server, allowing users to annotate large-scale datasets. FAMUS is the first comprehensive and modular annotation framework based on contrastive learning. It supports both pre-defined and user-specific databases for tailored annotation, and can be easily integrated into any genomic and metagenomic analysis pipeline to facilitate accurate, large-scale functional annotation.\nCompeting Interest Statement\nThe authors have declared no competing interest.","source_license":"CC-BY-4.0","license_restricted":false}