Enhanced Binding Affinity Prediction via Transformer‑Based Protein Embeddings and Ensemble Learning
preprint
OA: closed
Abstract
Binding Affinity Prediction is pivotal in determining the efficacy of interactions between proteins and ligands. It plays a critical role in drug discovery and therapeutic agent development. This study proposes a new hybrid ensemble learning model for protein-ligand affinity prediction that utilizes ProtT5 protein embeddings and Morgan molecular fingerprints for ligand representation. We used ProtT5 for protein sequence embedding and 2,048-bit Morgan Fingerprints to capture ligand structures. The ensemble approach performed best when a Random Forest regressor was paired with data encoded by the T5 Transformer model, achieving MSE = 0.5678, R 2 = 0.6062, MAE = 0.3246, and Pearson correlation = 0.7787 on the benchmark Devis Dataset. Our approach is distinctive in that it combines protein sequence embeddings with offline frozen Random Forest regression, enhancing accuracy due to the data’s intricate interdependence and highly multivariate nature. Hyperparameter tuning was conducted with Random Search, and ProtT5 embeddings coupled with Morgan Fingerprints served to depict ligands and proteins respectively. Compared with earlier models that utilized basic Random Forest and SVR, this approach demonstrates improved accuracy, interpretive capability, and less complexity in required inputs—merely protein sequences and SMILES strings. Here, the interplay between advanced embeddings and classical machine learning methods are harnessed to improve the accuracy and interpretability of the model.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00