PrePCI: A structure- and chemical similarity-informed database of predicted protein compound interactions

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

The PrePCI database predicts over 5 billion protein-compound interactions by combining structural modeling and chemical similarity to identify potential binding partners for proteins and compounds.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

We describe the Predicting Protein Compound Interactions (PrePCI) database which comprises over 5 billion predicted interactions between nearly 7 million chemical compounds and 19,797 human proteins. PrePCI relies on a proteome-wide database of structural models based on both traditional modeling techniques and the AlphaFold Protein Structure Database. Sequence and structural similarity-based metrics are established between template proteins in the Protein Data Bank, T, that bind small molecules, C, and proteins in the models database, Q. When these metrics pass a sequence threshold value, it is assumed that C also binds to Q with a probability derived from machine learning. If the relationship is based on structure, this probability is based on a scoring function that measures the extent to which C is compatible with the binding site of Q as described in the LT-scanner algorithm. For every predicted complex derived in this way, chemical similarity based on the Tanimoto Coefficient identifies other small molecules that may bind to Q. A likelihood ratio for the binding of C to Q is obtained from naïve Bayesian statistics. The PrePCI algorithm performs well under different validations. It can be queried by entering a UniProt ID for a protein and obtaining a list of compounds predicted to bind to it along with associated probabilities. Alternatively, entering an identifier for the compound outputs a list of proteins it is predicted to bind. Specific applications of the database are described and a strategy is introduced to use PrePCI as a first step in a docking screen.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-NC-ND-4.0