Developing machine-learning-based amyloid predictors with Cross-Beta DB

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-14

This study developed and benchmarked machine learning predictors for amyloid formation using the novel Cross-Beta DB dataset, with the random-forest-based Cross-Beta RF Predictor showing improved performance over existing methods.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-14 · read from full text

The paper describes the creation of Cross-Beta DB, a database of naturally formed cross-β amyloids, motivated by the role of amyloid aggregation in disease and function and the need for dedicated datasets for benchmarking computational predictors. Using the Cross-Beta DB dataset, the authors trained and benchmarked multiple machine-learning amyloidogenicity predictors and report that a random-forest-based model, Cross-Beta RF Predictor, outperformed existing methods. The main limitation stated is the historical lack of datasets specifically dedicated to naturally occurring cross-β amyloids, which the new database is intended to address, while the work focuses on cross-β structures (typically regions longer than ~15 residues) rather than other aggregation forms. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Due to shifts in environmental conditions, mutations, or interactions with other biomolecules, some proteins that would normally be soluble can undergo aggregation, resulting in the formation of clumps of amyloid fibrils. Understanding of this phenomenon is of paramount importance due not only to its association with various diseases (including Alzheimer’s disease), but also due to increasingly abundant evidence for its functional roles. Numerous studies have demonstrated that the propensity to form amyloids is coded by the amino acid sequence and this finding has paved the way for the development of several computational predictors of amyloidogenicity. The ultimate objective of computational methods is to accurately predict the formation of disease-related and functionally relevant amyloids that occur in vivo . These amyloid fibrils are known to form very specific “cross-β” structures of protein regions longer than about 15 residues. Remarkably, despite the significance of the naturally occurring amyloids, there has been a lack of datasets specifically dedicated to them. Hence, we built Cross-Beta DB, a database composed of cross-β amyloids formed in natural conditions. This database is expected to be indispensable for benchmarking amyloid predictors. We used the Cross-Beta DB to train and benchmark several such algorithms, using machine learning. The best-performing of these, the random-forest-based Cross-Beta RF Predictor, demonstrated superior performance over the other existing methods, fostering high expectations for an improved prediction of naturally occurring amyloids.
Full text 1,690 characters · extracted from oa-doi-fallback · click to expand
Abstract Due to shifts in environmental conditions, mutations, or interactions with other biomolecules, some proteins that would normally be soluble can undergo aggregation, resulting in the formation of clumps of amyloid fibrils. Understanding of this phenomenon is of paramount importance due not only to its association with various diseases (including Alzheimer’s disease), but also due to increasingly abundant evidence for its functional roles. Numerous studies have demonstrated that the propensity to form amyloids is coded by the amino acid sequence and this finding has paved the way for the development of several computational predictors of amyloidogenicity. The ultimate objective of computational methods is to accurately predict the formation of disease-related and functionally relevant amyloids that occur in vivo. These amyloid fibrils are known to form very specific “cross-β” structures of protein regions longer than about 15 residues. Remarkably, despite the significance of the naturally occurring amyloids, there has been a lack of datasets specifically dedicated to them. Hence, we built Cross-Beta DB, a database composed of cross-β amyloids formed in natural conditions. This database is expected to be indispensable for benchmarking amyloid predictors. We used the Cross-Beta DB to train and benchmark several such algorithms, using machine learning. The best-performing of these, the random-forest-based Cross-Beta RF Predictor, demonstrated superior performance over the other existing methods, fostering high expectations for an improved prediction of naturally occurring amyloids. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-NC-ND-4.0