In-Silico Stability Predictors: Investigation of Performance Towards balanced Experimental Data

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

Change in protein stability, quantified as the change in Gibbs free energy of folding (ΔΔG) in kcal/mol, plays a crucial role in functional alterations of proteins, with misfolding and destabilization commonly associated with pathogenicity. The past two decades have brought the development of bioinformatics tools leveraging evolutionary knowledge, empirical force fields, and machine learning to predict stability alterations. However, existing tools are often optimized towards or trained on limited experimental data, leading to unbalanced datasets and potential overfitting. The research objective is to benchmark selected stability predictors using an unbiased and balanced dataset, with AlphaFold structures as input. We demonstrate a performance decline when balancing data across amino acids, stabilizing and destabilizing mutations, and protein representatives, highlighting that redundancy alone is insufficient for benchmarking correction. Additionally, we illustrate that a protein structure ensemble from molecular dynamics acts as a superior input compared to a single static structure. At the same time, coarse-grained methodologies tend to decrease the output quality.
Full text 1,272 characters · extracted from oa-doi-fallback · click to expand
Abstract Change in protein stability, quantified as the change in Gibbs free energy of folding (ΔΔG) in kcal/mol, plays a crucial role in functional alterations of proteins, with misfolding and destabilization commonly associated with pathogenicity. The past two decades have brought the development of bioinformatics tools leveraging evolutionary knowledge, empirical force fields, and machine learning to predict stability alterations. However, existing tools are often optimized towards or trained on limited experimental data, leading to unbalanced datasets and potential overfitting. The research objective is to benchmark selected stability predictors using an unbiased and balanced dataset, with AlphaFold structures as input. We demonstrate a performance decline when balancing data across amino acids, stabilizing and destabilizing mutations, and protein representatives, highlighting that redundancy alone is insufficient for benchmarking correction. Additionally, we illustrate that a protein structure ensemble from molecular dynamics acts as a superior input compared to a single static structure. At the same time, coarse-grained methodologies tend to decrease the output quality. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00