Benchmarking datasets for machine learning in protein function prediction

preprint OA: closed CC-BY-NC-4.0

Abstract

ABSTRACT Remarkable progress has been achieved by machine learning, particularly in accurate prediction of protein tertiary structures. Despite these advances, accurately annotating protein functions through machine learning approaches remains challenging, primarily due to the limited availability of large-scale benchmarking data. In this study, we addressed this gap by systematically screening proteins from the UniProt database for functional annotations, resulting in the creation of a benchmarking dataset that includes protein sequences and their corresponding annotations. The Protein Annotation Dataset (PAD) is a resource available to train a wide range of machine learning models for assignment of function annotations to previously unlabeled proteins. We curated a comprehensive dataset comprising four categories of functional annotations using enzyme commission (EC) numbers and gene ontology (GO) terms. The dataset was subsequently partitioned into training, validation, and test subsets. Furthermore, we incorporated an independent set from 12 diverse species, enabling the development and evaluation of innovative machine learning models.
Full text 1,289 characters · extracted from oa-html · click to expand
ABSTRACT Remarkable progress has been achieved by machine learning, particularly in accurate prediction of protein tertiary structures. Despite these advances, accurately annotating protein functions through machine learning approaches remains challenging, primarily due to the limited availability of large-scale benchmarking data. In this study, we addressed this gap by systematically screening proteins from the UniProt database for functional annotations, resulting in the creation of a benchmarking dataset that includes protein sequences and their corresponding annotations. The Protein Annotation Dataset (PAD) is a resource available to train a wide range of machine learning models for assignment of function annotations to previously unlabeled proteins. We curated a comprehensive dataset comprising four categories of functional annotations using enzyme commission (EC) numbers and gene ontology (GO) terms. The dataset was subsequently partitioned into training, validation, and test subsets. Furthermore, we incorporated an independent set from 12 diverse species, enabling the development and evaluation of innovative machine learning models. Competing Interest Statement The authors have declared no competing interest. Footnotes ↵† email: benoit.kornmann{at}bioch.ox.ac.uk

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

References (14)

Source provenance

crossref
last seen: 2026-06-03T07:23:58.345424+00:00
europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-NC-4.0