Abstract
Creating a proteomewide atlas of reversible
and covalent protein-ligand binding sites
would transform our understanding of pro-
tein functions and accelerate therapeutic dis-
covery; however, current approaches face
significant challenges. Chemical proteomics
is constrained by limited proteome coverage
and data heterogeneity, while existing ma-
chine learning models have limited predictive
power due to structural dependencies and
inadequate handling of heterogeneous ex-
perimental labels. Here we developed AiPP ,
an AI protein profiling platform that provides
comprehensive annotations of ligand inter-
action sites directly from sequence. AiPP
is powered by the evolutionary scale protein
large language models (ESM3 and ESMC)
and two new databases, LigCysABPP (cys-
teine ligandability from activity-based pro-
tein profiling or ABPP) and LigBind3D (re-
versible binding sites in co-crystal structures).
LigCysABPP comprises >700,000 cysteine-
site records from 15 ABPP studies covering
>10,000 human proteins. We developed an
evolutionary representation-based clustering
framework followed by consensus analysis
to reconcile and augment heterogeneous ex-
perimental annotations. Two complementary
iterative data expansion protocols were im-
plemented to enhance model performance
and generalization. AiPP recovers 80% (Top-
1) of all cysteine liganding events from the
Protein Data Bank with AUPRC of 84% and
AUROC of 89%. Furthermore, it recapitu-
lates the consistently and heterogeneously
liganded cysteines determined by a recent
ABPP study using over 400 cancer cell lines.
AiPP provides a starting point for creating a
proteome-wide atlas of protein-ligand interac-
tions and systematic discovery of druggable
sites. The integrative approach that leverages
evolutionary information to harmonize large-
scale proteomic data can be broadly applied
to protein research.
Introduction
The human proteome comprises ∼20,300
distinct genes that encode over 200,000 pro-
teoforms1 (proteins and their variants aris-
ing from splicing, post-translational modifica-
tions, single-nucleotide polymorphisms and
mutations); however, fewer than 900 proteins
have been therapeutically targeted by the
FDA-approved drugs (https://go.drugbank.
com/).2 This gap underscores both the urgent
need and vast potential to expand the drug-
gable proteome. First developed as a chem-
ical proteomic approach that utilizes small-
molecule covalent probes (known as the
activity-based probes) to report on enzyme
1
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
activities in complex biological systems, 3,4
activity-based protein profiling (ABPP) has
evolved into a powerful and versatile tech-
nology. By combining activity- or reactivity-
based probes with mass spectrometry (MS)
based quantitative proteomics, ABPP has en-
abled several major advances, 5,6 including
the discovery of reactive cysteines7 and char-
acterization of small-molecule–protein inter-
actions (ligandability) in cells, 8 as well as the
development of covalent inhibitors. 9 In par-
ticular, the isotopic tandem orthogonal pro-
teolysis (isoTOP) ABPP experiments utilizing
competition ratios (CRs) between covalent
probes and fragments first introduced by Cra-
vatt and Weerapana 7,10,11 have led to quan-
titative assessment of cysteine ligandability
across over 10,000 proteins using different
human cell lines 8,12–24 Recently, the ABPP
technology has been further developed for
high-throughput screening through minimiz-
ing the amount of starting protein materials
and instrument time 15 as well as using label-
free quantification.18,25
Despite the enormous potential of ABPP
and rapid progress made in recent years, sev-
eral challenges remain, hindering efforts of
employing ABPP to create a comprehensive,
accurate map of cysteine-directed ligandable
sites across the proteome for covalent drug
design. One major challenge is limited pro-
teome coverage due to the dependence on
protein abundance, probe chemistry, and cell
type (e.g., membrane-bound proteins are un-
derrepresented in lysates due to insolubil-
ity).5,25–30 A second major challenge is the
variability in detected proteins, cysteines and
their assessed ligandability across the pub-
lished ABPP datasets (see refs 25,30–32 and
later discussions). This variability may at least
in part be attributed to the diverse cellular con-
texts25,31 and to specific experimental condi-
tions,30 including probe chemistry and con-
centration, fragment concentration, treatment
duration, cell type (e.g., cell lysates vs. live
cells), cellular state (e.g., mitotic vs. asyn-
chronous conditions; protein mutations; post-
translational modifications; redox state),33 cell
line (e.g., different cancers 31), specific elec-
trophilic fragments used, 25 and quantification
protocols (e.g., MS protocol and CR threshold
for defining a liganding event). An additional
challenge arises from the inherent limitations
of shotgun proteomics, 34 which prevents the
unambiguous identification of a subset of lig-
andable cysteines. 18,25 To address these lim-
itations, recent studies implemented label-
free and data-independent acquisition proto-
cols,18,25 demonstrating improved proteome
coverage depth and enhanced data complete-
ness compared to traditional approaches.
In contrast to proteomics approaches which
offer broad coverage across the proteome
but with significant data variability, X-ray crys-
tallography delivers high-resolution structural
information on protein-ligand interactions for
a limited number of complexes. Drawing
from the cysteine-liganded co-crystal struc-
tures deposited in the Protein Data Bank
(PDB), several databases (e.g., CovPDB, 35
CovBinderInPDB,36 and LigCys3D 37) have
been developed, which annotate nearly 800
proteins with chemically modified cysteines.
Based on these databases and experimen-
tal or AlphaFold 38 structure models, ma-
chine learning (ML) models, including support
vector models (SVMs), 39 boosted decision
trees,37,40 graph neural networks (GNNs), 41
and convolutional neural networks (CNNs), 37
have been developed for structure-based pre-
diction of ligandable cysteines, achieving re-
call and precision as high as 95%. How-
ever, these models are not yet deployable
for proteome-wide mapping of ligandable cys-
teine sites due to three critical limitations:
(i) the extremely small training dataset and
lack of protein diversity; (ii) the absence of
reliable controls for unligandable cysteines;
and (iii) the requirement for protein structures,
which often cannot be met even with state-of-
the-art structure prediction tools such as Al-
phaFold3.42
Recently, the ABPP-derived databases
(CysDB43 and TopCysteineDB 32) have been
developed and used to train structure-based
ML models for predicting cysteine reactiv-
ity or ligandability. 32,44 Based on CysDB 43
and other ABPP data (∼800 reactive cys-
2
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
teines) and associated crystal structures,
Forli, Backus, and coworkers developed the
random forest model CIAA to predict reactive
cysteines with an accuracy of 68% in exter-
nal validation.44 Using TopCysteineDB, which
combines PDB- (nearly 800 liganded cys-
teines) with ABPP-derived cysteine ligand-
ability data (over 9,000 ligandable cysteines
compiled from three publications 13,15,31 ) and
experimental or AlphaFold structures, Gohlke
and coworkers 32 developed the XGBoost
model TopCySPAL to predict cysteine ligand-
ability. A unique feature of TopCySPAL is that
it is trained with negative cysteines defined as
those undetected in liganded proteins quan-
tified by ABPP. 32 In the 20% hold-out test,
TopCySPAL achieved AUROC and AUPRC
of 96.4% and 91.4%, respectively. Another
structure-based ML model was the CNN 31
developed using the DrugMap database (an
ABPP dataset obtained with 400 cancer cell
lines), which achieved AUROC of 69.3% in
external evaluation.
Despite recent advances, existing ML mod-
els for predicting cysteine ligandability remain
severely constrained in both performance and
applicability. First, current approaches for
developing ML models inadequately leverage
ABPP datasets and lack rigorous frameworks
for assignment of ligandability labels. These
models fail to reconcile the inherent variabil-
ity in experimental labels, leading to incon-
sistent training data and unreliable predic-
tions. Second, this reconciliation challenge is
compounded by the absence of standardized
or rigorous evaluation metrics and bench-
marking protocols, making it difficult to rig-
orously assess model performance or com-
pare different approaches in a meaningful
way. Finally, and perhaps most critically, exist-
ing models require high-quality protein struc-
tures as input, rendering them inapplicable to
substantial portions of the proteome that re-
main structurally unresolved even with state-
of-the-art structure prediction tools such as
AlphaFold3,42 creating a significant coverage
gap in ligandability assessment.
To overcome the ML challenges and un-
lock the full potential of ABPP , we developed
a sequence-based AI Protein Profiling (AiPP)
platform for annotating reversible ligand bind-
ing (LigBind) and (cysteine-directed) cova-
lently ligandable residues (LigCys) across the
proteome. AiPP leverages the state-of-the-
art protein large language models (pLLM)
ESM Cambrian (ESMC) 45 as the foundation
model. We curated a comprehensive ABPP
database LigCysABPP comprising 15 inde-
pendent datasets published since 2016. We
developed a representation-based clustering
strategy to reconcile and augment the lig-
andability labels. A similar strategy is ap-
plied to refine the labels in LigBind3D. We
further developed two iterative learning pro-
tocols to systematically expand the ML train-
ing data, achieving enhanced model accu-
racy and generalization. Beyond predict-
ing ligandable sites, AiPP incorporates RIDA,
a disorder annotation platform that identifies
short disordered molecular recognition fea-
tures (MoRFs) with a propensity for binding-
induced folding Cysteine-containing MoRFs
are particularly interesting, as they may serve
as anchor points for protein-protein interac-
tion (PPI) drug design. AiPP provides a
unified platform, paving the way for creating
proteome-wide annotation of druggable pock-
ets and their adjacent covalently ligandable
residues.
Results
and Discussion
Curation of an ML-ready cysteine ligand-
ability database LigCysABPP. From 15
independent ABPP datasets (also referred
to as sources) published between 2016 and
2025,8,12–25 we extracted 683,192 peptide-
derived records of cysteine ligandability for
58,704 unique cysteine residues across
14,417 proteins in human proteomes. Fol-
lowing data clean-up steps, including protein
sequence reconstruction, record validation,
and consolidation of UniProt IDs (UIDs) refer-
ring to identical protein sequences, the raw
entries were converted to 669,908 structured,
site-specific cysteine ligandability records for
53,867 cysteine residues across 11,017 pro-
3
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
a
Other liganded
215 (203)
ABPP quantified proteins
10,649 (15,559)
Drug targets
70 (87)
LC3D
285 (290)
unliganded proteins: 4,147(30,663Cys)
db
c
Figure 1.
Figure 1: Overview of the ABPP-quantified proteins and cysteines in LigCysABPP database. a. The Lig-
CysABPP database is manually curated from a collection of 15 published cysteine-directed ABPP studies. 8,12–25
Peptide-derived records are extracted from each publication, validated against UniProt, consolidated across iden-
tical sequences and subsequences, augmented with unquantified (unseen) cysteine-sites on otherwise quantified
proteins, and filtered for compatibility with downstream processing steps. b. LigCysABPP contains the cysteine
ligandability records of 10,649 ABPP-quantified proteins, of which 6,502 are liganded and 4,147 are unliganded. A
liganded cysteine is defined as having at least 1 pos ABPP record while a liganded protein is defined as having at
least 1 liganded cysteine. Among the liganded proteins, 931 are known drug targets (with 2,293 liganded cysteines),
while the rest of 5,571 other liganded proteins contain 13,266 liganded cysteines. The LC3D database contains
285 unique proteins (290 cysteines liganded in the co-crystal structures), among which only 70 proteins (87 drug
targets) are also identified as liganded by ABPP .c. Functional classes of drug targets, other liganded proteins, and
unliganded proteins according to The Human Protein Atlas classifications46(https://www.proteinatlas.org/). d.
Histogram of the number of ABPP liganded cysteines per protein shows that most liganded proteins contain 1-3
liganded cysteines.
teins. Each record contains the ligandabil-
ity label (pos for liganded and neg for un-
liganded) given by the study authors for a
specific cysteine site in a protein under a set
of experimental conditions and cellular con-
text. To fully characterize the cysteines in
ABPP-quantified proteins, we added unquan-
tified (provisionally neg) records for cysteines
belonging to the ABPP-quantified proteins
but undetected by the probes. To accommo-
date the ESM-based ML, protein sequences
shorter than 30 or longer than 2046 amino
acids were removed. This procedure results
in 703,135 total curated records describing
ligandability (pos, neg, or unquantified) of
140,459 unique cysteine sites across 10,649
distinct proteins (Fig. 1a). Detailed extrac-
tion, validation, and consolidation logs are
provided in Supplemental Data (SD6–SD11).
Functional analysis of ABPP-quantified
proteins. We define a liganded cysteine as
one with at least one positive ABPP record.
A supporting source for a liganded cysteine
is one that contains at least one pos record
for that cysteine. Among the 10,649 ABPP-
quantified proteins, 6,502 are liganded (con-
taining at least one liganded cysteine) and
4,147 are unliganded (Fig. 1b). Among the
liganded proteins, 931 (with 2,293 pos cys-
teines) are known drug targets, while the
rest of 5,571 proteins (with 13,266 pos cys-
teines) represent unexplored potential ther-
apeutic opportunities (Fig. 1b). Among the
liganded drug targets and other proteins,
enzymes constitute the largest functional
4
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
family, followed by transporters, nuclear re-
ceptors, and others (Fig. 1c). It is note-
worthy that transporters and nuclear recep-
tors are substantially more prevalent among
other ligandable proteins, each represent-
ing over one-third the abundance of en-
zymes (Fig. 1c). Interestingly, GPCRs are
sparsely represented, likely because most
cysteine residues are not positioned near
ligand-binding pockets. It is intriguing that
the majority of unliganded proteins (2,593
out of 4,147) do not belong to any functional
class identified by The Human Protein Atlas46
(https://www.proteinatlas.org/), suggest-
ing that a significant fraction of the human
proteome remains to be functionally charac-
terized.
The overlap between the ABPP-liganded
proteins and those evidenced by co-crystal
structures in PDB is extremely small, only 60
proteins (42 are known drug targets, Fig. 1b).
This limited overlap may be attributed to the
historical focus of structural studies on bacte-
rial and other non-human proteins, which are
overrepresented in the PDB. The limited over-
lap also demonstrates the potential of ABPP
to greatly expand the known ligandable pro-
teome. The vast ABPP data also allows us
to ask a pertinent question: how many lig-
andable cysteines are in a single protein?
Interestingly, while the majority of proteins
contain only one ligandable cysteine, a siz-
able number of proteins harbor 2–4 ligandable
cysteines (Fig. 1d), suggesting that cysteine-
directed probes or inhibitors may be designed
to target allosteric and other ligandable pock-
ets.
Analysis of ABPP sources and records
to assess consensus in cysteine ligand-
ability. In order to devise an ML label-
ing scheme based on LigCysABPP records,
we first analyzed the consensus among 15
sources for liganded cysteines, which are
grouped by the number of sources that quan-
tified them (n =1,...,12). For instance, 2,996
cysteines were quantified by a single source,
while 2,644 and 2,154 were quantified by 2
and 3 sources, respectively. We then exam-
ined the consensus levels by determining how
many liganded cysteines are supported by m
number of sources (m =1,...,8). This analysis
highlights the challenge of reconciling source-
specific labels, as considerable variability ex-
ists. For example, among cysteines quanti-
fied by 2 sources, only 310 (11%) are sup-
ported by both sources, while 2,334 (89%)
showed conflicting results – pos in one source
and neg in the other. Among cysteines quan-
tified by 3 sources, only 82 (4%) are unani-
mously supported, while 449 (21%) are pos in
2 sources and 1,623 (75%) are pos in only 1
source. This pattern of decreasing consensus
becomes more pronounced as the number of
sources increases (Fig. 2a and Supplemental
Fig. S1).
In the above analysis, a supporting source
can have any number of pos records. To fur-
ther dissect the ABPP noise, we also exam-
ined the number of supporting records for lig-
anded cysteines (Fig. 2b). For liganded cys-
teines supported by 1 source, 7,787 have only
1 pos record, 1,374 have 2 pos records, and
816 have 3–4 pos records. For liganded cys-
teines supported by 2 sources, 814 have 3–4
positive records. As the number of support-
ing sources increases, the number of liganded
cysteines with positive records decreases –
with 5 sources, the number of liganded cys-
teines with at least 5 pos records is below 30.
The steep decline in well-supported liganded
cysteines suggests that requiring agreement
among more than 4 sources may be too strin-
gent for practical ML applications, as it would
severely limit the training data size.
The significant variation in cysteine ligand-
ability likely reflects differences in experimen-
tal conditions (e.g., probe and fragment con-
centrations) and cellular context (e.g., cell
line-specific factors) as shown in SI Data Ta-
ble S1. The influence of cellular context on
cysteine ligandability was recently 31 demon-
strated by an ABPP study of over 400 can-
cer proteomes, which identified nearly 400
cysteines exhibiting varying degrees of en-
gagement by an identical fragment (“hetero-
geneous ligandable”). While cellular con-
5
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
c
a
2,996 2,334 1,623
449
310
275
106
72
235 95
60
92
82
146
164 1168787
162
75
100
94
101
179241
105
115
50
53
82
37
436
1,010
754353
141
n=10 (479) n=11 (355) n=12 (234)
119
24
36
n=1 (2,996) n=2 (2,644) n=3 (2,154) n=4 (1,589)
n=9 (572)
n=5 (1,285) n=6 (1,002) n=7 (877) n=8 (680)
549 416 306
b Cys quantified in n sources (publications) and liganded in:
1 source 2 sources 3 sources ≥4 sources
In Pos Cluster
In Neg Cluster
Unlabeled
94,128
Unquantified Cys
74,190
11,665
8,273
Supporting sources
Positive records
1 3 5 7 92 4 6 8 10
5
10
15
20
Figure 2: Analysis of cysteine ligandability across ABPP sources and records. a. Each pie chart displays
the number of liganded cysteines quantified by n sources. A liganded cysteine is defined as one that has at least
one pos ABPP record. Each pie segment represents the number of cysteines supported by 1 (blue), 2 (green), 3
(yellow), and ≥4 (purple) sources. b. The number of liganded cysteines with the number of supporting sources
and positive records. A supporting source for a specific cysteine is defined as one that contains at least one pos
record for that cysteine. A total of 7,787 cysteines are labeled pos by only 1 record (black). No liganded cysteine
is supported by >10 sources. Liganded cysteines with >25 records are negligible and therefore excluded from the
plot. c. Following representation-based clustering, 8,273 unquantified cysteines are labeled pos, 11,665 labeled
neg, while 74,190 cysteines remain unlabeled.
texts will be addressed in our future stud-
ies, the present work focuses on cysteines
that demonstrate consistent labeling patterns
across different experiment conditions and
cell lines.
Using representation-based clustering to
derive cysteine ligandability labels for ML.
The variability in cysteine ligandability across
the ABPP datasets presents a significant
challenge for ML. Encouraged by our recent
work47 demonstrating that residue-level rep-
resentations (also known as per-token em-
beddings) from ultra-large pLLMs such as
ESMC45,48 encode not only evolutionary infor-
mation but also structural and local biochemi-
cal environments, we leveraged the residue-
specific representations to reconcile the la-
bels for liganded and unliganded cysteines.
Specifically, we performed clustering analysis
on all cysteines in ABPP-quantified proteins
using pairwise similarity scores computed
from ESMC-derived residue-specific embed-
ding vectors. For each cluster, we identified a
representative cysteine as the node exhibiting
the highest average similarity to all other clus-
ter members. This approach yielded compact,
non-overlapping clusters of cysteine sites with
high embedding similarity, each anchored by
a representative node used for consensus la-
beling and subsequent model training. To pre-
vent leakage, clusters were never split across
data partitions and all clusters sharing a UID
were assigned to the same partition.
To derive high-confidence ligandability la-
bels from the ABPP data, we performed
cluster-based label assignment on the full
set of representation-based clusters accord-
ing to varying levels of consensus, defined
as a combination of supporting sources and
6
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
records (denoted nS–mR). For example, un-
der the 4S–4R criteria, a cluster is labeled pos
only if it contains at least four positive records
from four different sources. Conversely, a
cluster is labeled neg if it contains no posi-
tive records and at least four negative records
from unique sources. Clusters that did not
meet either pos or neg labeling condition were
excluded from the training set. To augment
coverage, unquantified cysteines in quantified
proteins cast default negative votes and in-
herit a cluster-based label.
Applying the 4S–4R criteria yielded 936
pos and 4,590 neg labeled clusters, cov-
ering 143,914 site-level records describ-
ing 11,869 distinct cysteine residues across
11,541 unique proteins. Importantly, this ap-
proach enabled us to assign labels to 21%
of ABPP-unquantified cysteines, classifying
8,273 as pos and 11,665 as neg (Fig. 2c).
The cluster representatives were then used
for downstream model training. Additional de-
tails are provided in the Supplemental Meth-
ods, including data validation, representation-
level deduplication, and procedures with
representation-based criteria to prevent data
leakage between training, validation, and ex-
ternal evaluation.
Note, while conventional sequence identity
filtering methods, such as CD-HIT, 49 focus
on eliminating sequence similarity, they fail to
account for functionally convergent or struc-
turally similar ligandable sites that may exist
within otherwise divergent proteins. In con-
trast, clustering based on pLLM representa-
tion goes beyond simple sequence compar-
ison by integrating local biochemical, struc-
tural, and evolutionary information. Thus,
representation-based criteria are more strin-
gent and effective in reducing the risk of data
leakage.
Support from 4 sources may be an opti-
mal tradeoff between label confidence and
data coverage. To examine how label con-
fidence depends on the number of quantify-
ing sources, we measured variability using
Shannon entropy (a measure of labeling un-
certainty) of the binary pos/neg distribution for
cysteine clusters quantified byn sources. The
number of clusters falls sharply as the thresh-
old increases; for example, requiring at least 4
quantifying sources reduces the set to 11,913
clusters compared to 27,410 with a 2-source
threshold (Fig. 3a, top). Entropy is high for
clusters supported by 2 or 3 sources, reaches
its minimum at 4 sources (Fig. 3a, bottom),
and then rises again at thresholds ≥ 7, peak-
ing at 12 sources where only∼150 highly het-
erogeneous clusters remain.
Fig. 3b complements this analysis by show-
ing how positive cysteines accumulate as ad-
ditional sources are added under different
consensus rules (1S–1R, 2S–2R, . . ., 5S–
5R). At the 4-source threshold, the num-
ber of positives normalized by the na ¨ıve to-
tal (all cysteines reported as liganded in any
record) qualitatively plateaus: each additional
source contributes relatively few new positives
(Fig. 3b, top). Together, these results sug-
gest that requiring evidence from at least 4
supporting sources (defined here as quanti-
fying sources that reported a cysteine as lig-
anded at least once; the 4S–4R rule) provides
a practical balance between label confidence
and data coverage for annotating ligandable
cysteines with the currently available data.
Development of the sequence-only LigCys
models. Inspired by our recent demonstra-
tion that ESMC-based, sequence-only mod-
els can accurately predict residue-specific
pKa values, particularly for cysteines despite
limited training data, 47 we developed the Lig-
Cys model for predicting cysteine ligandabil-
ity. The model employs ESMC as the founda-
tional pLLM, using frozen 2,560-dimensional
per-residue embeddings passed through a
three-layer multilayer perceptron (MLP) clas-
sifier trained on the ABPP-derived data de-
scribed below. Embeddings from layer 76
of ESMC were used (Methods), as prelimi-
nary testing showed this layer provided opti-
mal performance for LigCys predictions. Con-
sistent with prior observations, 47,50 different
transformer layers capture distinct evolution-
7
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
a b
4S-4R
Figure 3: Analysis of how source threshold affects label variability and data coverage. a. Top: Bar plot
showing the number of cysteine clusters supported by at least n independent sources, regardless of the number of
underlying records. For clarity, bins corresponding to odd values ofn are unlabeled. Bottom: Label variability within
each cluster, stratified by the number of supporting sources. Variability is quantified using the Shannon entropy of
the binary label distribution ( p, the fraction of sources labeling a cysteine as positive), scaled to a 0–100% range.
Boxes denote the interquartile range; medians are shown as horizontal dashed lines; whiskers span the 10th–90th
percentiles. A minimum in entropy is observed at n ≥ 4, suggesting this threshold maximizes labeling consistency
across sources. b. Top: Incremental gain in the number of positive cysteines when increasing from n−1 to n
sources. Bottom: For each consensus threshold (e.g., 1S–1R through 5S–5R), we randomly sample n sources
from a pool of 15, repeat 200 times, and compute the average number of positive cysteines. This is normalized by
the na¨ıve total (i.e., all cysteines reported as positive in any single record). The 95% confidence interval is shown
for the na¨ıve curve. The 4S–4R consensus curve (highlighted in purple) plateaus after four sources, supporting its
use as a data-driven threshold for robust labeling.
ary and structural features at different rates;
layer 76 embeddings were therefore also
used for LigBind models. We also evaluated
pLLM fine-tuning approaches using ESM2, 51
but the resulting models did not outperform
the representation-only ESMC approach for
cysteine ligandability prediction.
LigCys models trained on the 4S–4R ABPP
dataset show the strongest generaliza-
tion. Due to the highly variable nature of
ABPP data, both model training and validation
are extremely challenging. To address val-
idation challenges, we used crystallography-
validated ligandable cysteines from the LC3D
database as an external test set. LC3D
cysteines and any representation-based clus-
ters containing these cysteines were removed
from the training set to prevent data leak-
age (Supplemental Methods). Therefore, the
LC3D test set is not only orthogonal to the
ABPP training set but also provides the most
stringent validation for model’s ability to gen-
eralize.
We evaluated how ABPP data consensus
requirement affects the model performance
by training LigCys models with ligandability la-
bels derived from increasingly stringent con-
sensus criteria (Table 1). These LigCys mod-
els are compared using the standard metrics,
including the most pertinent metric Top-1 re-
covery, which is defined as the probability that
the cysteine with the highest predicted classi-
fication score in a protein is a true pos given
that the protein has at least one pos cys-
8
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
Table 1: External evaluation of LigCys models trained using ABPP datasets with varying con-
sensus thresholdsa
ABPP training data External LC3D evaluation (%)
Criteria Pos Neg Prot PRC ROC Prec Rec Top-1 Top-1*
1S–1R 13430 55441 7051 60.0 72.2 38.4 37.8 40.5 44.5
1S–2R 6078 16791 4216 70.2 80.9 55.9 55.1 56.8 59.6
1S–3R 3396 8600 2829 76.6 85.8 62.3 61.9 65.3 68.2
1S–4R 2094 5048 2004 78.1 86.3 65.3 64.9 68.3 72.9
1S–5R 1418 3354 1495 80.1 87.4 70.0 69.6 71.7 75.9
2S–2R 3467 6838 2617 77.0 84.2 65.3 64.7 66.2 70.3
2S–3R 2198 4765 1883 77.3 85.5 64.3 64.1 66.6 70.3
2S–4R 1450 3326 1404 78.8 86.4 69.6 68.8 70.9 75.1
2S–5R 988 2258 1041 79.5 86.6 71.7 70.9 72.2 76.8
3S–3R 1115 1920 997 78.8 86.6 68.8 68.1 70.0 74.2
3S–4R 867 1583 837 80.2 88.3 68.3 67.5 71.3 75.5
3S–5R 649 1244 669 80.8 87.7 72.6 71.7 73.5 77.6
4S–4R 384 642 389 81.2 87.3 74.3 73.2 75.2 79.4
4S–5R 347 590 358 80.3 86.1 74.3 73.2 74.3 78.9
5S–5R 142 236 150 79.8 85.6 71.3 70.3 72.6 76.8
aAt a consensus threshold nS–mR, the label assignment requires at least m records from n different sources (see
main text). Each model is an ensemble of 200 models (see Supplemental Methods). Performance is measured on
the orthogonal LC3D testset on the per-protein basis. Cysteines without covalent ligation are provisionally labeled
neg. Precision and recall of predicting ligandable cysteines are calculated with a classification threshold of 0.5.
Top-1 recovery refers to the probability that the cysteine with the highest classification score is a true pos given that
the protein has at least one pos cysteine. Top-1* refers to the evaluation using LC3D augmented with labels from
representation-based clustering (see Supplemental data).
teine. Since LC3D negatives reflect the ab-
sence of observed liganding event rather than
validated inactivity, AUROC/AUPRC should
be interpreted with caution; we therefore em-
phasize ranked recovery metrics (e.g., Top-
1). Confirming the substantial heterogeneity
across ABPP data, the 1S–1R model, despite
being trained on the largest dataset (13,430
pos and 55,441 neg cysteines), delivers the
worst performance, with AUROC, AUPRC,
and Top-1 recovery values of 72.2%, 60.0%,
and 40.5%, respectively. At the single-source
level, performance improves dramatically as
the number of required pos records increases,
reaching AUROC, AUPRC, and Top-1 recov-
ery of 87.3%, 81.2%, and 71.7%, respectively,
despite a ten-fold reduction in training data. A
similar trend of significant performance gain
with increasing number of pos records is ob-
served at the 2- and 3-source levels. How-
ever, at the 4-source level, performance with
5 pos records (4S–5R) is very similar to that
with 4 records (4S–4R), which has the AU-
ROC, AUPRC, and Top-1 of 87.3%, 81.2%,
and 75.2%, suggesting that model perfor-
mance may be plateaued. Indeed, the 5-
source model (5S–5R) shows decreased per-
formance across all three metrics by 2–5%
compared to the 4S–5R model, which may be
attributed to the substantial (more than 50%)
reduction in training data (only 142 pos and
236 neg cysteines).
This comparison shows that the 4S–4R
training set yields the best model performance
(Table 1, highlighted row) for generalization
to the completely orthogonal external vali-
dation data. This observation is consistent
with the analyses in Fig. 2 and 3, suggest-
9
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
ing that the 4S–4R training set balances la-
bel confidence with dataset size for effective
generalization. To further assess model per-
formance, we computed Top-1* recovery for
the LC3D test set that received additional
label assignments through representation-
based clustering among themselves (referred
to augmented LC3D or LC3D*). The 4S–
4R model remains the top-performing one,
achieving AUROC, AUPRC, and Top-1* of
87.3%, 81.2%, and 79.4%, respectively.
Enhanced model performance through it-
erative training data expansion guided by
LC3D-based evaluation. Although the 4S–
4R dataset achieves superior performance
relative to others, its small size may limit
the model’s generalization capability. To ad-
dress this limitation, we hypothesized that us-
ing LC3D as a validation benchmark would
allow us to identify additional cysteines (and
proteins) with high-confidence labels from the
broader ABPP dataset to expand the training
data. To test this hypothesis, we developed
an iterative procedure guided by LC3D Top-1
recovery, starting from a truncated version of
the high-confidence 4S–4R dataset contain-
ing cluster representatives from proteins with
both positive and negative cysteines. In this
truncated version (denoted 4S–4R t), 23 pro-
teins quantified by 10 or more sources with
at least one pos and one neg cysteine are
withheld for model evaluations. The candidate
training pool included cysteines from all 4S–
4Rt clusters, extending beyond the baseline
to encompass non-representative members
as well as proteins supported exclusively by
positive or exclusively by negative evidence.
In iterations 1–3, we generated 100
batches, each consisting of 80, 100, and 150
randomly selected cysteines from the can-
didate pool, respectively. From iteration 4
onward, we generated 100 batches of 175
randomly selected cysteines. In each itera-
tion, batches were evaluated by incorporating
them into the training set and training ensem-
bles of 24 models (6 data splits with 4 random
seeds per split). The best-performing batch
was selected for further validation using 10
new random splits with 20 models per split,
with Top-1 recovery serving as the evaluation
metric. If Top-1 recovery exceeded the pre-
vious iteration’s baseline, the best batch was
accepted.
Through six iterations, Top-1 recovery in-
creased from 70.9% to 79.9% on LC3D and
from 75.1% to 84.5% on LC3D*, accompa-
nied by steady improvements in AUROC from
85.5% to 88.9% and AUPRC from 78.1% to
84.2% (Table 2). In iteration 7, no further gain
was observed, indicating that performance
had plateaued under this random-batch ex-
pansion scheme. Hereafter we will call the
Iter6 model LigCys-S. Relative to the base-
line 4S–4R t dataset, the expanded training
set nearly tripled in size to 1,099 proteins,
driven by over twofold increase in negative
cysteines (1,345) and a nearly 30% increase
in positives (441). Using the same data, we
trained a structure-aware model (6-SA) to test
if adding additional structural features would
enhance predictive power (Table 2, row 6-SA).
Interestingly, along with other calculated met-
rics, the Top-1 recovery of LC3D liganded cys-
teines is decreased by 1.7 %. This suggests
that the evolutionary driven pLLM-based rep-
resentations capture the majority of informa-
tion pertinent to cysteine ligandability.
Enhanced model performance through it-
erative training data expansion guided by
ABPP-based evaluation. As random sam-
pling in the LC3D-guided procedure may
leave parts of the candidate pool unexplored,
we hypothesized that a systematic, cross-
validation (CV) driven approach could more
efficiently identify additional high-confidence
data. To test this hypothesis, we developed
an expansion strategy guided by CV perfor-
mance on the ABPP data, starting from the
same baseline dataset (4S–4R t) as in the
LC3D-guided expansion. The same candi-
date data pool was also used.
At each iteration, candidate batches of fixed
composition (25 positives and the remain-
der negatives) were assembled, with batch
10
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
Table 2: Performance metrics for iterative training data expansion guided by LC3D-based evalu-
ation (top block) and ABPP-based cross validation (bottom block).
ABPP training dataa ABPP valid. (%)b External LC3D evaluationc (%)
LC3D-guided expansion
Iter Pos Neg Prot PRCE ROC PRC ROC Prec Rec Top-1 Top-1*
0 354 574 366 74 74.6 78.1 85.5 70.5 69.6 70.9 75.1
0-SA 354 574 366 79 76.5 80.1 86.9 71.3 70.5 72.6 76.8
1 359 649 442 92 78.2 79.5 85.8 71.3 70.3 72.2 76.8
2 369 739 534 94 76.1 80.4 86.8 73.2 73.4 73.5 78.1
3 384 874 674 92 74.2 81.8 87.8 73.9 72.8 74.8 79.8
4 401 1032 827 115 77.5 82.9 88.0 78.2 76.9 78.2 82.8
5 425 1186 966 139 79.4 83.0 87.8 76.9 75.6 78.2 82.8
6 441 1345 1099 141 78.2 84.2 88.9 77.3 77.7 79.9 84.5
6-SA 441 1345 1099 128 76.7 83.8 88.7 76.9 75.6 78.2 83.2
ABPP-guided expansion
Tourd Pos Neg Prot PRCE ROC PRC ROC Prec Rec Top-1 Top-1*
0 354 574 366 74 74.6 78.1 85.5 70.5 69.6 70.9 75.1
1 378 643 452 93 80.2 80.1 86.8 72.6 71.7 73.1 77.3
2 468 967 817 122 82.1 81.5 88.3 78.1 72.1 74.4 79.4
3 555 1432 1258 148 82.6 80.1 87.5 71.3 70.3 71.8 76.0
4 645 1998 1744 187 83.5 79.3 87.1 69.8 69.2 70.5 76.0
4-SA 645 1998 1744 178 80.9 80.5 87.7 69.8 69.2 71.7 76.8
Blendede - - - - - 83.2 88.4 75.4 74.5 76.9 85.5
a In the LC3D-guided data expansion, Top-1 recovery of LC3D liganded cysteines was used to select candidate
data batches (complete data in Supplemental Data). The total number of pos, neg cysteines, and distinct proteins
in the training set is given after each iteration. The start of the expansion (Iter 0) used the 4S–4R t dataset, which
excludes 23 proteins quantified by 10 sources from the original 4S–4R dataset. b Model validation metrics are
AUROC (ROC) and AUPRC enrichment (PRCE), defined as(AUPRC−random)/random. Here random represents
the AUPRC of a random classifier (% positive samples in the dataset). c For external evaluation against the LC3D
dataset, precision (Prec) and recall (Rec) were calculated using the default classification threshold of 0.5. Top-1
and Top-1* were calculated using LC3D and LC3D* (label augmented by clustering) datasets, respectively. d In
the ABPP-guided data expansion, AUPRC enrichment (PRCE) was used to select candidate data batches. Tour
denotes the number of successive tournaments (effectively iterations 1, 5, 9, and 13; see Supplemental Data). All
other details follow the notes above. e Blended denotes ensemble prediction using both Iter 6 and Tour 4 models.
size set to 100 cysteines in early iterations
and increased to 175 as the pool expanded.
Batches were prioritized by ensemble uncer-
tainty, and a batch was admitted only if it sig-
nificantly exceeded the AUPRC enrichment
(PRCE) of the previous iteration baseline in
the CV. PRCE (enrichment of AUPRC over
random guess) was used due to the vari-
able pos:neg Cys ratio in each iteration. To
broaden the search, advancing batches were
further refined by simple genetic algorithm
steps, preserving the best performers while
generating crossover and mutation variants.
Expansion proceeded through four tourna-
ments, each comprising four iterations. Here,
a “tournament” refers to a series of eval-
uations in which candidate batches com-
pete against the baseline: top performers
11
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
are retained and recombined, and only the
strongest batch is accepted at each itera-
tion of the tournament. Expansion terminated
when the final tournament ended without any
accepted batch. This acceptance scheme
yielded steady performance gains: Top-1 re-
covery stayed constant at 70.9% to 70.5%
on LC3D and at 75.1% to 76.0% on LC3D*,
with corresponding increases in AUROC and
AUPRC (Table 2, under ABPP-guided expan-
sion). Hereafter we will refer to the Tour4
model as LigCys-A. The final expanded set
comprised 1,744 proteins and 2,633 cysteine
residues, representing a 1.8-fold increase in
positives and a 3.4-fold increase in negatives
relative to the 4S–4Rt baseline. Performance
gains plateaued after the fourth tournament,
indicating that the available candidate pool
had been exhausted. As in the LC3D-guided
expansion, we also trained a structure-aware
model using the final expanded data. The
resulting model (4-SA) has a small increase
of 1.2% for AUPRC and Top-1 recovery com-
pared to the sequence-only model.
Beyond the LC3D test, we evaluated the
LigCys-S and LigCys-A models against the
ABPP hold-out set comprised of 23 proteins
quantified by 10 or more sources, each con-
taining at least one positive and one negative
cysteine. From this hold-out set, we created 9
test sets using the nS-nR consensus thresh-
olds (n =1,...,9). As expected, the number
of labeled cysteines decreases from 1S-1R
(∼300 cysteines) to just under 100 at 4S-4R
and 7 at 9S-9R (Supplemental Fig. S2). The
AUPRC of both models increases with the
consensus level of the test set, from ∼68%
(67 for LigCys-A model) at 1S-1R to 73%
(75 for LigCys-A model) at 4S-4R and nearly
100% at 9S-9R (Supplemental Fig. S2). Sim-
ilarly, the model recall steadily increases from
below 40% at 1S-1R to 100% at 9S-9R (Sup-
plemental Fig. S2). These improvements
stem from increased TPs and decreased FNs
with growing consensus level (Supplemental
Fig. S2). The analysis demonstrates that
the models perform better on consistently lig-
anded cysteines, corroborating our initial find-
ings (Table 1). Consistent with the LC3D test,
the ABPP hold-out test suggests that LigCys-
S and LigCys-A models perform similarly, with
LigCys-S showing slightly fewer false posi-
tives when predicting highly ligandable cys-
teines.
Development of an ML-ready LigBind3D
database and the sequence-only LigBind
models for predicting reversible ligand
binding residues. Since truly ligandable
cysteines are located at or near a reversible
binding pocket, we developed sequence-only
LigBind models (LigBind-seq) to predict re-
versible ligand-binding residues complemen-
tary to cysteine ligandability predictions. To
develop the LigBind models, we created an
ML-ready database called LigBind3D based
on BioLiP2, 52 an updated manually curated
dataset of experimental structures from bio-
logically relevant protein–ligand complexes in
the PDB. BioLiP252 is significantly larger than
commonly used datasets for protein-ligand
affinity predictions (e.g., BindingDB 53) which
contain only protein–ligand interactions with
experimentally determined binding affinities.
We selected BioLiP2 entries of monomeric
proteins containing small molecules with
molecular weights between 150–600 Da
and applied a custom structural annotation
pipeline (PickPocket, unpublished) to vali-
date biologically relevant protein-ligand inter-
actions. Following the various filtering steps,
we reconstructed the full protein sequence di-
rectly from the atomic coordinate file. Ligand-
binding residues were labeled based on a
4.5 ˚A heavy-atom distance cutoff. These
residues were converted into structured, site-
level records. Details are given in Supplemen-
tal Methods. The final LigBind3D database
comprises 687,712 site-level records across
1,998 unique proteins. This database offers
a direct structural perspective on reversible
protein-ligand interactions, complementing
the LigCysABPP and LC3D databases. In
contrast to LigCys3D, which restricts to cys-
teines, LigBind3D labels every residue in a
protein, reflecting the broader diversity of
reversible protein-ligand interactions. Sim-
12
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
ilar to the treatment of LigCysABPP data,
representation-based clustering was applied
to the LigBind3D data for creating the training,
validation, and test datasets.
The LigBind-seq model employs a simpler
single-layer feedforward network architecture
that processes fixed 2,560-dimensional per-
token representations extracted from ESMC
layer 76. Preliminary screening showed that
the simpler architecture was sufficient to cap-
ture ligand-binding patterns from sequence
alone. In addition to the sequence-only mod-
els, we also developed structure-aware Lig-
Bind models (LigBind-SA) by incorporating
3D geometric features through a pretrained
transformer model derived from the PeSTo ar-
chitecture54 alongside the pLLM embeddings.
Details are given in Supplemental Methods.
In the hold-out test, LigBind-seq achieves
AUROC, AUPRC, and Top-10 scores of
93.9%, 74.5%, and 78.1%, respectively, while
LigBind-SA attains 95.1%, 77.5%, and 80.1%
(Table 3). The modest performance improve-
ments of LigBind-SA relative to LigBind-seq
indicate that the ESMC-derived representa-
tions provide the majority of the model’s pre-
dictive power.
Table 3: Performance of the sequence-only
and structure-aware LigBind models in pre-
dicting reversible ligand-binding residuesa
Metrics (%) LigBind-seq LigBind-SA
AUROC* 93.9 95.1
AUPRC* 74.5 77.5
Prec* 72.1 71.1
Rec* 66.4 71.5
Top-5* 79.2 81.2
Top-10* 78.1 80.1
Top-20* 76.6 79.5
a All models were trained and evaluated* using cluster-
propagated labels derived via representation-based
clustering. Performance metrics reflect a post hoc
ensemble of 200 models (Supplemental Methods).
Per-residue ligandability probabilities were averaged
across all 200 models and thresholded at 0.5. All met-
rics were computed on a per-protein basis using the
held-out test set of 129 proteins.
LigBind predictions offer complementary
validation for LigCys predictions. Cova-
lent ligandability of cysteines requires prox-
imity to a reversible binding pocket. There-
fore, accurate LigBind predictions not only
provide reversible pocket information but may
also serve as complementary validation for
LigCys predictions. To test this complemen-
tary relationship, we applied both LigBind and
LigCys models to the LC3D test set with pro-
teins in the LigBind training set excluded. We
then calculated the radial distribution function
(RDF) of predicted ligand-binding residues
surrounding TP , TN, FP , and FN cysteines
(Fig. 4). Both LigBind-seq and LigBind-SA
models were used. Except for TNs, all RDFs
exhibit a peak around 4–6 ˚A. Notably, both
LigBind-seq and LigBind-SA RDFs for TPs
show the highest peak, consistent with the ex-
pectation that ligandable cysteines are proxi-
mal to reversible binding pockets. The peak of
the LigBind-SA RDF (dashed) is a bit higher
compared to the LigBind-seq RDF (solid),
suggesting that incorporating structural fea-
tures enhances pocket prominence. Reas-
suringly, both LigBind-seq and LigBind-SA
RDFs for TNs lack peaks, consistent with the
expectation that unligandable cysteines are
not located near ligand-binding pockets. Re-
markably, the LigBind-seq RDF for FNs ex-
hibits the second highest peak, suggesting
that LigBind-seq predictions could help iden-
tify overlooked ligandable cysteines based on
pocket proximity. Conversely, the LigBind-
seq and LigBind-SA RDFs for FPs display
very low peaks, suggesting that incorrectly
predicted ligandable cysteines can be filtered
out based on their lack of proximal binding
pockets. These results demonstrate that inte-
grating LigCys and LigBind predictions within
a unified framework would further enhance
the reliability of covalent ligandability assess-
ment.
Development of the AiPP platform for
sequence-based predictions of reversible
and covalent ligand binding sites across
the proteome. Building on the LigCys and
13
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
𝑔 𝑟
Distance 𝑟 from Cys of interest (Å)
0 20 255 10 158
100
200
(tail normalized) TP (n=15)
TN (n=167)
FP (n=8)
FN (n=122)
LigBind-seq
LigBind-SA
300
400
500
600
700
Figure 4: Radial distribution function of LigBind-
predicted ligand-binding residues around cysteine
sites, stratified by the LigCys prediction outcome
(TP, TN, FP, FN).Distance is the minimum all-atom Eu-
clidean distance between the cysteine of interest and
each predicted ligand-binding residue. Radial distribu-
tion function g(r) is normalized by the mean shell den-
sity over the last 20% of the distance range. Solid and
dashed lines represent the LigBind-seq and LigBind-
SA predictions, respectively, on the pruned LC3D test
set that excludes proteins in the LigBind training set.
LigBind models, we developed the Artificial
Intelligence Protein Profiling (AiPP) platform
as a unified, sequence-based framework
for proteome-wide ligandable site identifica-
tion (Fig. 5). AiPP integrates reversible and
covalent ligand-binding predictions as well
as a disorder predictor to support rational
drug discovery. The LigBind module detects
residues that spatially stabilize ligand binding;
the LigCys module identifies cysteines sus-
ceptible to modification by a covalent ligand;
and the RIDA module (based on the RIDAO
server56) predicts per-residue disorder and
molecular recognition features (MoRFs) that
can fold upon binding. Together, these pre-
dictions map ligandable binding sites capa-
ble of retaining cysteine-reactive compounds
sufficiently long for covalent bond forma-
tion—guiding targeted covalent drug design,
while simultaneously identifying dynamic or
induced-fit binding regions amenable to PPI
inhibitor development. Additionally, the re-
cently developed KaML-ESM 47 and KaML-
CBtree57 models are incorporated to offer
reactivity information for cysteines of interest
based on sequence only or structure.
Starting from a protein sequence, AiPP gen-
erates a predicted 3D structure (via a local
version of ESM3 48), a structural and func-
tional disorder annotation (via a local ver-
sion of RIDAO server 56), and the residue-
specific representations (from the pLLM
ESMC45). The ESMC-based representa-
tions are processed by LigBind and LigCys
modules to generate predictions of reversible
and cysteine-directed covalent ligand binding
sites, respectively. LigCys generates predic-
tions from the LigCys-S and LigCys-A mod-
els, as well as a blended prediction that en-
sembles results from both models. Option-
ally, structure-aware versions of the predic-
tions are provided. Results are visualized on
the ESM3 predicted structure with potential
MoRFs. Optionally, the predicted structure is
processed by the pretrained transformers de-
rived from the PeSTo architecture54 to gener-
ate geometric representations, which are then
integrated with the ESMC-based representa-
tions as well as hand-engineered features to
enhance predictions by LigBind and LigCys.
Additionally, cysteine reactivity (using pKa as
a proxy) and solvent accessibility (buried frac-
tion) are annotated to inform ligand design.
Evaluation on the consistently and hetero-
geneously liganded cysteines from a re-
cent ABPP study. 31 A recent ABPP study
using over 400 cancer cell lines demonstrated
that while a majority of the quantified cys-
teines were consistently liganded across cell
lines, a subset of cysteines showed varied
ligandability depending on the cell line (re-
ferred to as heterogeneous cysteine ligand-
ability).31 We made predictions for 14 pro-
teins containing the consistently and hetero-
geneously liganded cysteines featured in the
publication31 (DrugMap dataset) but absent
in the LigCys training set. Note, although
LigCysABPP database does not contain the
DrugMap dataset, there is a significant over-
lap in proteins/cysteines. LigCys predicted
14
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
UniProtID: P05141
ADP/ATP translocase 2
Reversable Binding ResidueCovalently Ligandable Cysteine
C129
(Unligandable)
C257 (Top-3)
pKa=8.2;
exposure=0%
C57 (Top-1)
pKa=7.7;
exposure=0%
C347 (Top-1)
pKa=6.5;
exposure=25%
C160 (Top-2)
pKa=8.1;
exposure=0%
UniProtID: P0DPH7
Tubulin alpha-3C
b c
a
Figure 5: Architecture of the AI protein profiler (AiPP) for annotating covalent and reversible ligand
binding sites given a protein sequence. a. Starting from a protein sequence, AiPP generates a predicted 3D
structure (ESM3), structural and functional disorder annotations (RIDA), and ESMC-based representations. These
representations are processed by three specialized modules: LigCys, LigBind, and KaML-ESM to predict covalent,
reversible ligand binding residues, cysteine reactivities, respectively. Results are mapped onto the ESM3 predicted
structure. Optionally, the predicted structure is processed by a transformer model (PeSTo) to generate geometric
representations, which are then integrated with the pLLM-based embeddings as well as hand-engineered features
to generate LigCys-SA and LigBind-SA models. b. For TUBA3C, LigCys predicted C347 as Top-1 ligandable,
supporting the finding that it is consistently liganded across cell lines. 31 C347 is at the binding interface with β-
tubulin in the adjacent dimer of the protofilaments that comprise microtubule structures. For visualization, we added
a surface rendering of β-tubulin based on the electron diffraction model (PDB ID: 1TUB). 55 c. For SLC25A5,
LigCys predicted C57, C160, C257, and C129 as Top-1, Top-2, and Top-3 ligandable, respectively, while C129 was
predicted as unligandable, consistent with the order of engagement frequency across cell lines.31
tubulin alpha-3C (TUBA3C, UID P0DPH7)
Cys347 and SOX10 (UID P56693) Cys71 as
Top-1 ligandable (Fig. 5b), corroborating the
consistent engagement across different can-
cer cell lines. 31 Although this comparison is
limited, the data suggest that LigCys models
can confidently predict consistently ligandable
cysteines.
Based on the electron microscopy model of
α/β-tubulin dimer, 55 Cys347 is positioned at
the binding interface with β-tubulin from the
adjacent dimer in the protofilaments that com-
prise the microtubule structures. Given the
disease relevance of tubulin proteins, the con-
vergence of computational predictions and
experimental evidence for Cys347 ligandabil-
ity suggests potential drug design opportu-
nity for disrupting or stabilizing this protein-
15
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
protein interaction (PPI). SOX10 transcription
factor is a major dependency of melanoma
cells.31 Notably, the compelling ligandability of
Cys71, which is located at the dimerization
domain, facilitated a proof-of-concept ligand
design that stabilizes the PPI, leading to the
disruption of melanoma transcriptional signal-
ing.31
Among the heterogeneously liganded cys-
teines31 not in the LigCys training set, Lig-
Cys predictions vary. While PCBP1 (Q15365)
Cys109, LARS1 (Q9P2J5) Cys70, and
RPS3A (P61247) Cys201, are predicted as
non-ligandable, SLC25A5 (P05141) Cys257
and ACAT1 (A0A5F9ZI66) Cys413 were each
predicted as the Top-2 ligandable cysteine in
their respective proteins. Interestingly, closer
examination of our prediction and the pub-
lished ABPP data 31 for SLC25A5 revealed
that the LigCys ranking of the 3 predicted lig-
andable cysteines (Fig. 5c) is identical to the
order of the positive engagement frequency
across cell lines. Note, a positive engage-
ment is defined by ≥60% engagement, cor-
responding to the commonly used CR thresh-
old of 4. 31 While this level of agreement may
be anecdotal given our limited comparison,
the data suggest that LigCys models may al-
ready capture some information about hetero-
geneous ligandability through the prediction
ranking. While our LC3D-based evaluations
focused solely on the Top-1 predictions, users
can explore additional candidates by exam-
ining cysteines ranked 2nd, 3rd, or lower in
their ligandability.
Concluding Discussion
We present the first unified AI platform to en-
able sequence-based, scalable annotation of
reversible and covalent ligand-binding sites
across the proteome. This platform is en-
abled through the new ML-ready databases
and methodological advances that address
several critical challenges in proteome-wide
ML.
By incorporating positive, negative, and un-
quantified ABPP records from 15 publica-
tions, the LigCysABPP database substan-
tially exceeds the scope and comprehensive-
ness of existing ABPP databases, including
CysDB,43 DrugMap, 31 and TopCysteineDB.32
LigCysABPP comprises 608,898 quantified
and 94,237 unquantified cysteine records of
10,649 unique proteins. Of these proteins,
6,502 contain at least 1 liganded cysteine
(a total of 15,559 liganded cysteines) while
4,147 contain only unliganded and unquan-
tified cysteines. Building upon BioLip2 52 the
LigBind3D database offers more comprehen-
sive coverage of protein-ligand interactions
than commonly used binding databases such
as BindingDB. 53 LigBind3D contains anno-
tations for reversible ligand-binding residues
across 1,998 unique proteins derived from
co-crystal structures. Together with LC3D
(a refined database of covalently liganded
cysteines evidenced by co-crystal structures),
LigCysABPP and LigBind3D provide solid
foundations for developing ML models aimed
at proteome-wide discovery of covalent (Lig-
Cys) and reversible (LigBind) ligand-binding
sites.
A major bottleneck in developing ML mod-
els for predicting ligand binding residues is
experimental label variability (e.g., conflicting
annotations for ligandable cysteines) and in-
completeness (lack of high-confidence nega-
tives for covalent or reversible ligand-binding
residues). We addressed these issues by de-
veloping a pLLM representation-based clus-
tering approach to reconcile and augment la-
bels. To develop robust LigCys models from
ABPP data, we established a consensus anal-
ysis framework evaluating models trained on
data at varying consensus thresholds, using
LC3D as an orthogonal validation dataset to
identify the most reliable baseline. Building
from this optimized baseline, we implemented
two iterative approaches – LC3D- (orthogonal
validation) and ABPP-guided (inherent valida-
tion) – to systematically increase the train-
ing set size and diversity, achieving substan-
tial improvements in model performance and
generalization.
The structural requirement of ML mod-
els significantly limits practical applications.
16
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
We addressed this constraint by leverag-
ing state-of-the-art pLLM ESMC 45 to de-
velop LigCys and LigBind models capable
of rapid, sequence-based prediction of co-
valent and reversible ligand-binding sites
for proteome-scale exploration. The LigCys
models achieved 79.9% Top-1 recovery of
all known liganded cysteines from co-crystal
structures. The LigBind models demonstrated
strong performance on the hold-out test set of
129 proteins, achieving an AUROC of 93.9%,
AUPRC of 74.5%, and Top-10 recovery rate
of 78.1%.
LigCys and LigBind models are integrated
into AiPP , a unified AI platform designed
for proteome-scale annotation of covalent
and reversible ligand-binding sites. Addition-
ally, AiPP incorporates several complemen-
tary predictive tools: a meta disorder pre-
dictor RIDA (based on RIDAO 56) for identi-
fying MoRFs critical to PPI drug design, a
cysteine reactivity predictor 47,57 for additional
functional studies, and ESM3 48 for protein
structure prediction. The integration of struc-
tural information (LigBind-SA) enhances bind-
ing site predictions and provides users with vi-
sualization capabilities. Importantly, AiPP can
be systematically improved (e.g., by incorpo-
rating continuously growing data from ABPP
experiments) and expanded (e.g., by adding
prediction capabilities for lysine and other co-
valently ligandable residues).
Current work has several caveats that can
be addressed in the future. Although limited
comparison to the recent DrugMap work 31
suggest that LigCys models may already
capture some information about heteroge-
neous ligandability through the prediction
ranking system, current LigCys models are
currently agnostic to cellular context, e.g.,
redox conditions, post-translational modifi-
cations, and mutational states. Two recent
ABPP studies revealed that phosphoryla-
tion33 and specific cancer-cell environment 31
can significantly modulate cysteine reactiv-
ity and/or ligandability in a subset of cys-
teines. The underlying mechanisms may in-
volve phosphorylation-induced protein con-
formational changes, shifts in cellular redox
state, and accumulated cysteine-proximal
mutations in malignant cells. 31,33 Molecular
dynamics (MD) simulations 58,59 have demon-
strated that the nucleophilicity of cysteines
(and other residues such as lysines) can be
conformation-dependent, with kinases pro-
viding clear examples of how reactivity varies
between functionally relevant conformational
states (e.g., DFG-in vs. DFG-out forms). We
envision that the contextual information can
be added to the training data and at inference
stage to provide proteoform-specific predic-
tions in the future. This enhancement may re-
solve the divergent performance observed be-
tween models trained on datasets expanded
using LC3D- vs. ABPP-guided strategies. An-
other limitation of current AiPP is the LigBind
model which was trained using a database
consisting of monomer proteins. In addition to
addressing these limitations, we also envision
coupling AiPP with ligand docking and gener-
ation models to advance proteome-wide drug
design. Finally, the representation-based la-
bel analysis and augmentation approaches
developed in this work are broadly applica-
ble for ML-experiment integration, establish-
ing a generalizable framework for develop-
ing proteomics-derived ML models across di-
verse applications.
Supplemental Materials
Supplemental Methods contains detailed pro-
cedures for constructing the LigCysABPP ,
LC3D, and LigBind3D databases; pLLM-
based data clustering, label assignment and
reconciliation; LigCys and LigBind model ar-
chitectures, training, and validation; and Lig-
Cys training data expansion. Supplemental
Figures contains a more detailed analysis ac-
companying Figure 1; model performance
analysis using the ABPP hold-out set.
Supplemental Data
A Supplemental Data Excel file contains the
following 20 sheets:
17
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
Benchmarking studies: ESMC representa-
tion (SD2), consensus (SD3), source abla-
tion (SD4), source prioritization (SD5), option-
ablation (SD20) and split-level noise (SD21).
Data extraction logs: ABPP extraction log
(SD6), log of ambiguous and multivalue entry
splitting (SD7), ABPP statistical report (SD8),
ABPP orphaned records (dropped due to lack
of unambiguous support) (SD9).
Sequence and clustering records: validated
UID sequences in FASTA format (SD10),
UID deduplication log (SD11), precluster
records (SD12), cysteine (Cys) clusters re-
port (SD13).
Production distillation: 4S–4R distillation re-
port (SD14).
Support reports: UID-level (SD15), ROI-level
(SD16), and representative-level (SD17).
ABPP hold-out set details: UID-level (SD18)
and ROI-level (SD19).
Note: Production runs refer to combined
ABPP and LC3D data processed in a single
distillation workflow, as illustrated in Fig. 1a.
Data Availability
Databases and models are available at
https://github.com/JanaShenLab/AiPP/,
which also contains a link to the cor-
responding web server https://aipp.
computchem.org/. Two data files accom-
pany this manuscript: the Supplemental Data
spreadsheet (SD2–SD21) and the LigCys-
ABPP
v1.xlsx spreadsheet; both are mirrored
with code, models, and data in the repository.
Acknowledgments
Research reported in this publication was
supported by the National Institute Of Gen-
eral Medical Sciences (R35GM148261) and
National Cancer Institute (R01CA256557) of
the National Institutes of Health. The con-
tent is solely the responsibility of the authors
and does not necessarily represent the official
views of the National Institutes of Health. We
thank Benjamin Cravatt from the Scripps Re-
search Institute for insightful comments and
discussion. We thank Chetan Mishra and
Marius Wiggert from Evolutionary Scale for
facilitating the usage of ESM Cambrian and
ESM3. This work was supported by an Evolu-
tionaryScale compute grant.
References
(1) Aebersold, R. et al. How Many Human
Proteoforms Are There? Nat. Chem.
Biol. 2018, 14, 206–214.
(2) Wishart, D. S. et al. DrugBank 5.0: A
Major Update to the DrugBank Database
for 2018. Nucleic Acids Res. 2018, 46,
D1074–D1082.
(3) Liu, Y .; Patricelli, M. P .; Cravatt, B. F .
Activity-Based Protein Profiling: The
Serine Hydrolases. Proc. Natl. Acad.
Sci. U.S.A. 1999, 96, 14694–14699.
(4) Greenbaum, D.; Medzihradszky, K. F .;
Burlingame, A.; Bogyo, M. Epoxide Elec-
trophiles as Activity-Dependent Cysteine
Protease Pro¢ling and Discovery Tools.
Chem. Biol. 2000, 7, 569–581.
(5) Chan, W. C.; Sharifzadeh, S.;
Buhrlage, S. J.; Marto, J. A. Chemo-
proteomic Methods for Covalent Drug
Discovery. Chem. Soc. Rev. 2021, 50,
8361–8381.
(6) Cravatt, B. F . Activity-Based Protein Pro-
filing – Finding General Solutions to Spe-
cific Problems. Isr. J. Chem. 2023,
(7) Weerapana, E.; Wang, C.; Simon, G. M.;
Richter, F .; Khare, S.; Dillon, M. B. D.;
Bachovchin, D. A.; Mowen, K.; Baker, D.;
Cravatt, B. F . Quantitative Reactivity Pro-
filing Predicts Functional Cysteines in
Proteomes. Nature 2010, 468, 790–795.
(8) Backus, K. M.; Correia, B. E.;
Lum, K. M.; Forli, S.; Horning, B. D.;
Gonz´alez-P´aez, G. E.; Chatterjee, S.;
18
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
Lanning, B. R.; Teijaro, J. R.; Ol-
son, A. J.; Wolan, D. W.; Cravatt, B. F .
Proteome-Wide Covalent Ligand Dis-
covery in Native Biological Systems.
Nature 2016, 534, 570–574.
(9) Roberts, A. M.; Miyamoto, D. K.; Huff-
man, T. R.; Bateman, L. A.; Ives, A. N.;
Akopian, D.; Heslin, M. J.; Contr-
eras, C. M.; Rape, M.; Skibola, C. F .; No-
mura, D. K. Chemoproteomic Screening
of Covalent Ligands Reveals UBA5 As
a Novel Pancreatic Cancer Target. ACS
Chem. Biol. 2017, 12, 899–904.
(10) Weerapana, E.; Speers, A. E.;
Cravatt, B. F . Tandem Orthogonal
Proteolysis-Activity-Based Protein Pro-
filing (TOP-ABPP)—a General Method
for Mapping Sites of Probe Modification
in Proteomes. Nat. Protoc. 2007, 2,
1414–1425.
(11) Wang, C.; Weerapana, E.;
Blewett, M. M.; Cravatt, B. F . A Chemo-
proteomic Platform to Quantitatively Map
Targets of Lipid-Derived Electrophiles.
Nat. Methods 2014, 11, 79–85.
(12) Bar-Peled, L. et al. Chemical Proteomics
Identifies Druggable Vulnerabilities in a
Genetically Defined Cancer. Cell 2017,
171, 696–709.e23.
(13) Vinogradova, E. V. et al. An Activity-
Guided Map of Electrophile-Cysteine In-
teractions in Primary Human T Cells.
Cell 2020, 182, 1009–1026.e29.
(14) Cao, J.; Boatner, L. M.; Desai, H. S.;
Burton, N. R.; Armenta, E.; Chan, N. J.;
Castell´on, J. O.; Backus, K. M. Multi-
plexed CuAAC Suzuki–Miyaura Label-
ing for Tandem Activity-Based Chemo-
proteomic Profiling. Anal. Chem. 2021,
93, 2610–2618.
(15) Kuljanin, M.; Mitchell, D. C.;
Schweppe, D. K.; Gikandi, A. S.;
Nusinow, D. P .; Bulloch, N. J.; Vino-
gradova, E. V.; Wilson, D. L.; Kool, E. T.;
Mancias, J. D.; Cravatt, B. F .; Gygi, S. P .
Reimagining High-Throughput Profiling
of Reactive Cysteines for Cell-Based
Screening of Large Electrophile Li-
braries. Nat. Biotechnol. 2021, 39,
630–641.
(16) Y an, T.; Desai, H. S.; Boatner, L. M.;
Y en, S. L.; Cao, J.; Palafox, M. F .; Jami-
Alahmadi, Y .; Backus, K. M. SP3-FAIMS
Chemoproteomics for High-Coverage
Profiling of the Human Cysteinome**.
ChemBioChem 2021, 22, 1841–1851.
(17) Tao, Y .; Remillard, D.; Vino-
gradova, E. V.; Y okoyama, M.;
Banchenko, S.; Schwefel, D.; Melillo, B.;
Schreiber, S. L.; Zhang, X.; Cra-
vatt, B. F . Targeted Protein Degradation
by Electrophilic PROTACs That Stereos-
electively and Site-Specifically Engage
DCAF1. J. Am. Chem. Soc. 2022, 144,
18688–18699.
(18) Y ang, F .; Jia, G.; Guo, J.; Liu, Y .;
Wang, C. Quantitative Chemopro-
teomic Profiling with Data-Independent
Acquisition-Based Mass Spectrometry.
J. Am. Chem. Soc. 2022, 144, 901–911.
(19) Y an, T.; Boatner, L. M.; Cui, L.;
Tontonoz, P . J.; Backus, K. M. Defining
the Cell Surface Cysteinome Using Two-
Step Enrichment Proteomics. JACS Au
2023, 3, 3506–3523.
(20) Burton, N. R.; Polasky, D. A.; Shik-
wana, F .; Ofori, S.; Y an, T.; Geis-
zler, D. J.; Veiga Leprevost, F . D.;
Nesvizhskii, A. I.; Backus, K. M. Solid-
Phase Compatible Silane-Based Cleav-
able Linker Enables Custom Isobaric
Quantitative Chemoproteomics. J. Am.
Chem. Soc. 2023, 145, 21303–21318.
(21) Koo, T.-Y .; Lai, H.; Nomura, D. K.;
Chung, C. Y .-S. N-Acryloylindole-alkyne
(NAIA) Enables Imaging and Profiling
New Ligandable Cysteines and Oxidized
Thiols by Chemoproteomics. Nat. Com-
mun. 2023, 14, 3564.
19
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
(22) Burton, N. R.; Backus, K. M. Functional-
izing Tandem Mass Tags for Streamlin-
ing Click-Based Quantitative Chemopro-
teomics. Commun. Chem. 2024, 7, 80.
(23) Njomen, E.; Hayward, R. E.; De-
Meester, K. E.; Ogasawara, D.;
Dix, M. M.; Nguyen, T.; Ashby, P .; Si-
mon, G. M.; Schreiber, S. L.; Melillo, B.;
Cravatt, B. F . Multi-Tiered Chemical Pro-
teomic Maps of Tryptoline Acrylamide–
Protein Interactions in Cancer Cells.
Nat. Chem. 2024, 16, 1592–1604.
(24) Tian, C.; Sun, L.; Liu, K.; Fu, L.;
Zhang, Y .; Chen, W.; He, F .; Y ang, J.
Proteome-Wide Ligandability Maps of
Drugs with Diverse Cysteine-Reactive
Chemotypes. Nat. Commun. 2025, 16,
4863.
(25) Biggs, G. S. et al. Robust Proteome
Profiling of Cysteine-Reactive Frag-
ments Using Label-Free Chemopro-
teomics. Nat. Commun. 2025, 16, 73.
(26) Abegg, D.; Frei, R.; Cerato, L.;
Prasad Hari, D.; Wang, C.; Waser, J.;
Adibekian, A. Proteome-Wide Profiling
of Targets of Cysteine Reactive Small
Molecules by Using Ethynyl Benziodox-
olone Reagents. Angew. Chem. Int. Ed.
2015, 54, 10852–10857.
(27) Shannon, D. A.; Weerapana, E. Covalent
Protein Modification: The Current Land-
scape of Residue-Specific Electrophiles.
Curr. Opin. Chem. Biol.2015, 24, 18–26.
(28) Long, M. J. C.; Aye, Y . Privileged Elec-
trophile Sensors: A Resource for Cova-
lent Drug Development. Cell Chem. Biol.
2017, 24, 787–800.
(29) Wang, S.; Tian, Y .; Wang, M.; Wang, M.;
Sun, G.-b.; Sun, X.-b. Advanced
Activity-Based Protein Profiling Applica-
tion Strategies for Drug Development.
Front. Pharmacol. 2018, 9.
(30) White, M. E.; Gil, J.; Tate, E. W.
Proteome-Wide Structural Analysis
Identifies Warhead- and Coverage-
Specific Biases in Cysteine-Focused
Chemoproteomics. Cell Chem. Biol.
2023, 30, 828–838.e4.
(31) Takahashi, M. et al. DrugMap: A Quanti-
tative Pan-Cancer Analysis of Cysteine
Ligandability. Cell 2024, 187, 2536–
2556.e30.
(32) Bonus, M.; Greb, J.; Majmudar, J. D.;
Boehm, M.; Korczynska, M.; Nazemi, A.;
Mathiowetz, A. M.; Gohlke, H. TopCys-
teineDB: A Cysteinome-wide Database
Integrating Structural and Chemopro-
teomics Data for Cysteine Ligandabil-
ity Prediction. J. Mol. Biol. 2025, 437,
169196.
(33) Kemper, E. K.; Zhang, Y .; Dix, M. M.;
Cravatt, B. F . Global Profiling of
Phosphorylation-Dependent Changes in
Cysteine Reactivity. Nat. Methods 2022,
19, 341–352.
(34) Mann, M.; Kulak, N. A.; Nagaraj, N.;
Cox, J. The Coming Age of Complete,
Accurate, and Ubiquitous Proteomes.
Mol. Cell 2013, 49, 583–590.
(35) Gao, M.; Moumbock, A. F . A.;
Qaseem, A.; Xu, Q.; G ¨unther, S.
CovPDB: A High-Resolution Coverage
of the Covalent Protein–Ligand Inter-
actome. Nucleic Acids Res. 2022, 50,
D445–D450.
(36) Guo, X.-K.; Zhang, Y . CovBinder-
InPDB: A Structure-Based Covalent
Binder Database. J. Chem. Inf. Model.
2022, 62, 6057–6068.
(37) Liu, R.; Clayton, J.; Shen, M.; Bhatna-
gar, S.; Shen, J. Machine Learning Mod-
els to Interrogate Proteome-Wide Co-
valent Ligandabilities Directed at Cys-
teines. JACS Au 2024, 4, 1374–1384.
(38) Jumper, J. et al. Highly Accurate Protein
Structure Prediction with AlphaFold. Na-
ture 2021, 596, 583–589.
20
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
(39) Zhang, W.; Pei, J.; Lai, L. Statisti-
cal Analysis and Prediction of Covalent
Ligand Targeted Cysteine Residues. J.
Chem. Inf. Model. 2017, 57, 1453–1460.
(40) Reimer, B. M.; Awoonor-Williams, E.;
Golosov, A. A.; Hornak, V. Cov-
CysPredictor: Predicting Selective Co-
valently Modifiable Cysteines Using Pro-
tein Structure and Interpretable Machine
Learning. J. Chem. Inf. Model. 2025, 65,
544–553.
(41) Du, H.; Jiang, D.; Gao, J.; Zhang, X.;
Jiang, L.; Zeng, Y .; Wu, Z.; Shen, C.;
Xu, L.; Cao, D.; Hou, T.; Pan, P .
Proteome-Wide Profiling of the
Covalent-Druggable Cysteines with
a Structure-Based Deep Graph Learn-
ing Network. AAAS Research 2022,
2022, 9873564.
(42) Abramson, J. et al. Accurate Structure
Prediction of Biomolecular Interactions
with AlphaFold 3. Nature 2024, 630,
493–500.
(43) Boatner, L. M.; Palafox, M. F .;
Schweppe, D. K.; Backus, K. M.
CysDB: A Human Cysteine Database
Based on Experimental Quantitative
Chemoproteomics. Cell Chem. Biol.
2023, 30, 683–698.e3.
(44) Boatner, L. M.; Eberhardt, J.; Shik-
wana, F .; Holcomb, M.; Lee, P .;
Houk, K. N.; Forli, S.; Backus, K. M.
CIAA: Integrated Proteomics and Struc-
tural Modeling for Understanding Cys-
teine Reactivity with Iodoacetamide
Alkyne. ACS Chem. Biol. 2025, 20,
1669–1682.
(45) ESM Team. ESM Cambrian: Re-
vealing the Mysteries of Pro-
teins with Unsupervised Learning.
https://www.evolutionaryscale.ai/blog/esm-
cambrian.
(46) Uhlen, M.; Oksvold, P .; Fagerberg, L.;
Lundberg, E.; Jonasson, K.; Fors-
berg, M.; Zwahlen, M.; Kampf, C.;
Wester, K.; Hober, S.; Wernerus, H.;
Bj¨orling, L.; Ponten, F . Towards a
Knowledge-Based Human Protein Atlas.
Nat. Biotechnol. 2010, 28, 1248–1250.
(47) Shen, M.; Dayhoff, G. W.; Shen, J. Pro-
tein Electrostatic Properties Are Fine-
Tuned Through Evolution. 2025.
(48) Hayes, T. et al. Simulating 500 Mil-
lion Y ears of Evolution with a Language
Model. Science 2025, 387, 850–858.
(49) Fu, L.; Niu, B.; Zhu, Z.; Wu, S.;
Li, W. CD-HIT: Accelerated for Clus-
tering the next-Generation Sequencing
Data. Bioinformatics 2012, 28, 3150–
3152.
(50) Vig, J.; Madani, A.; Varshney, L. R.;
Xiong, C.; Socher, R.; Rajani, N. F .
BERTology Meets Biology: Interpreting
Attention in Protein Language Models.
ICLR 2021. 2021.
(51) Lin, Z.; Akin, H.; Rao, R.; Hie, B.; Zhu, Z.;
Lu, W.; Smetanin, N.; Verkuil, R.; Ka-
beli, O.; Shmueli, Y .; Fazel-Zarandi, M.;
Sercu, T.; Candido, S.; Rives, A.
Evolutionary-Scale Prediction of Atomic-
Level Protein Structure with a Language
Model. Science 2023, 379, 1123–1130.
(52) Zhang, C.; Zhang, X.; Freddolino, L.;
Zhang, Y . BioLiP2: An Updated Struc-
ture Database for Biologically Relevant
Ligand–Protein Interactions. Nucl. Acids
Res. 2024, 52, D404–D412.
(53) Gilson, M. K.; Liu, T.; Baitaluk, M.;
Nicola, G.; Hwang, L.; Chong, J. Bind-
ingDB in 2015: A Public Database
for Medicinal Chemistry, Computational
Chemistry and Systems Pharmacology.
Nucl. Acids Res. 2016, 44, D1045–
D1053.
(54) Krapp, L. F .; Abriata, L. A.; Cort ´es Ro-
driguez, F .; Dal Peraro, M. PeSTo:
Parameter-Free Geometric Deep Learn-
ing for Accurate Prediction of Protein
21
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
Binding Interfaces. Nat. Commun. 2023,
14, 2175.
(55) L ¨owe, J.; Li, H.; Downing, K.; No-
gales, E. Refined Structure ofαβ-Tubulin
at 3.5 ˚A Resolution 1 1Edited by I. A. Wil-
son. J. Mol. Biol. 2001, 313, 1045–1057.
(56) Dayhoff II, G. W.; Uversky, V. N. Rapid
Prediction and Analysis of Protein In-
trinsic Disorder. Protein Sci. 2022, 31,
e4496.
(57) Shen, M.; Kortzak, D.; Ambrozak, S.;
Bhatnagar, S.; Buchanan, I.; Liu, R.;
Shen, J. KaMLs for Predicting Protein
p K a Values and Ionization States: Are
Trees All Y ou Need? J. Chem. Theory
Comput. 2025, 21, 1446–1458.
(58) Liu, R.; Yue, Z.; Tsai, C.-C.; Shen, J. As-
sessing Lysine and Cysteine Reactivities
for Designing Targeted Covalent Kinase
Inhibitors. J. Am. Chem. Soc. 2019, 141,
6553–6560.
(59) Liu, R.; Verma, N.; Henderson, J. A.;
Zhan, S.; Shen, J. Profiling MAP Kinase
Cysteines for Targeted Covalent Inhibitor
Design. RSC Med. Chem. 2022, 13, 54–
63.
22
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 8, 2025. ; https://doi.org/10.1101/2025.09.07.670677doi: bioRxiv preprint
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.