SAIR: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset

preprint OA: closed
📄 Open PDF Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-06 · read from full text

SAIR introduces the Structurally Augmented IC50 Repository (SAIR), a large dataset of 5,244,285 protein–ligand 3D structures spanning 1,048,857 unique protein–ligand systems, curated from ChEMBL and BindingDB and computationally folded using the Boltz-1x model. The authors characterize protein and ligand distributions and assess structural fidelity with PoseBusters, finding that about 3% of structures show physical anomalies, mainly internal energy violations. As a benchmark, they evaluate several binding affinity prediction approaches (Vina, Vinardo, Onionnet-2, and AEV-PLIG) and report that ML models outperform traditional scoring functions but do not yield high correlation with ground-truth affinities, attributing this to the need for models fine-tuned to synthetic structure distributions; a key limitation is that correlation remains low despite improved structural fidelity. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

A bstract Accurate prediction of protein-ligand binding affinities remains a cornerstone problem in drug discovery. While binding affinity is inherently dictated by the 3D structure and dynamics of protein-ligand complexes, current deep learning approaches are limited by the lack of high-quality experimental structures with annotated binding affinities. To address this limitation, we introduce the Struc-turally Augmented IC50 Repository ( SAIR ), the largest publicly available dataset of protein-ligand 3D structures with associated activity data. The dataset com-prises 5, 244, 285 structures across 1, 048, 857 unique protein-ligand systems, cu-rated from the ChEMBL and BindingDB databases, which were then computa-tionally folded using the Boltz-1x model. We provide a comprehensive charac-terization of the dataset, including distributional statistics of proteins and ligands, and evaluate the structural fidelity of the folded complexes using PoseBusters. Our analysis reveals that approximately 3% of structures exhibit physical anoma-lies, predominantly related to internal energy violations. As an initial demon-stration, we benchmark several binding affinity prediction methods, including empirical scoring functions (Vina, Vinardo), a 3D convolutional neural network (Onionnet-2), and a graph neural network (AEV-PLIG). While machine learning-based models consistently outperform traditional scoring function methods, nei-ther exhibit a high correlation with ground truth affinities, highlighting the need for models specifically fine-tuned to synthetic structure distributions. This work provides a foundation for developing and evaluating next-generation structure and binding-affinity prediction models and offers insights into the structural and phys-ical underpinnings of protein-ligand interactions. The dataset can be found at https://www.sandboxaq.com/sair .
Full text 61,487 characters · extracted from preprint-html · click to expand
SAIR: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results SAIR: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset View ORCID Profile Pablo Lemos , View ORCID Profile Zane Beckwith , View ORCID Profile Sasaank Bandi , Maarten van Damme , View ORCID Profile Jordan Crivelli-Decker , Benjamin J. Shields , Thomas Merth , View ORCID Profile Punit K. Jha , View ORCID Profile Nicola De Mitri , View ORCID Profile Tiffany J. Callahan , AJ Nish , Paul Abruzzo , Romelia Salomon-Ferrer , Martin Ganahl doi: https://doi.org/10.1101/2025.06.17.660168 Pablo Lemos 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Pablo Lemos For correspondence: pablo.lemos{at}sandboxaq.com zane.beckwith{at}sandboxaq.com Zane Beckwith 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Zane Beckwith For correspondence: pablo.lemos{at}sandboxaq.com zane.beckwith{at}sandboxaq.com Sasaank Bandi 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Sasaank Bandi Maarten van Damme 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site Jordan Crivelli-Decker 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Jordan Crivelli-Decker Benjamin J. Shields 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site Thomas Merth 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site Punit K. Jha 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Punit K. Jha Nicola De Mitri 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Nicola De Mitri Tiffany J. Callahan 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Tiffany J. Callahan AJ Nish 2 NVIDIA Find this author on Google Scholar Find this author on PubMed Search for this author on this site Paul Abruzzo 2 NVIDIA Find this author on Google Scholar Find this author on PubMed Search for this author on this site Romelia Salomon-Ferrer 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site Martin Ganahl 1 SandboxAQ Find this author on Google Scholar Find this author on PubMed Search for this author on this site Abstract Full Text Info/History Metrics Data/Code Preview PDF A bstract Accurate prediction of protein-ligand binding affinities remains a cornerstone problem in drug discovery. While binding affinity is inherently dictated by the 3D structure and dynamics of protein-ligand complexes, current deep learning approaches are limited by the lack of high-quality experimental structures with annotated binding affinities. To address this limitation, we introduce the Struc-turally Augmented IC50 Repository ( SAIR ), the largest publicly available dataset of protein-ligand 3D structures with associated activity data. The dataset com-prises 5, 244, 285 structures across 1, 048, 857 unique protein-ligand systems, cu-rated from the ChEMBL and BindingDB databases, which were then computa-tionally folded using the Boltz-1x model. We provide a comprehensive charac-terization of the dataset, including distributional statistics of proteins and ligands, and evaluate the structural fidelity of the folded complexes using PoseBusters. Our analysis reveals that approximately 3% of structures exhibit physical anoma-lies, predominantly related to internal energy violations. As an initial demon-stration, we benchmark several binding affinity prediction methods, including empirical scoring functions (Vina, Vinardo), a 3D convolutional neural network (Onionnet-2), and a graph neural network (AEV-PLIG). While machine learning-based models consistently outperform traditional scoring function methods, nei-ther exhibit a high correlation with ground truth affinities, highlighting the need for models specifically fine-tuned to synthetic structure distributions. This work provides a foundation for developing and evaluating next-generation structure and binding-affinity prediction models and offers insights into the structural and phys-ical underpinnings of protein-ligand interactions. The dataset can be found at https://www.sandboxaq.com/sair . 1 I ntroduction Understanding the interaction between proteins and ligands is a fundamental problem in chemical biology and drug discovery. The binding affinity of a ligand to its target protein, as well as to off-target proteins, is a critical parameter when it comes to designing small-molecule drugs. In principle, the binding affinity can be derived from the structural information of the protein-ligand complex, as the three-dimensional structure of the complex describes most of the interaction between the protein and the ligand. In practice, however, there are limitations to predicting affinity from structural data, both from an experimental and a computational perspective. From an experimental perspective, despite tremendous advances in the field, there are significant challenges generating experimental structures for some proteins, limiting the accessibility and resolution of some structures. Addition-ally, in spite of significant advancement in higher throughput methods, the effort needed to produce experimental structural information makes it difficult to efficiently integrate it into the design cy-cle. From the computational perspective, despite their tremendous utility, traditional methods used to calculate binding affinities like MM/GBSA ( Wang et al., 2019 ) and free energy methods (e.g. Cournia et al., 2017 ; Crivelli-Decker et al., 2024 ; York, 2023 ) rely on the use of force fields which limit their accuracy, while quantum mechanical based methods remain prohibitively expensive. Ad-ditionally, traditional methods strongly depend on the quality of the protein-ligand complex, with small inaccuracies producing large errors in the estimation of binding affinities. To address these limitations, one approach is to learn surrogate functions that approximate bind-ing affinity directly from protein sequences and ligand SMILES representations (e.g., Ö ztürk et al., 2019; Jiang et al., 2022 ; Limbu & Dakshanamurthy, 2022 ). However, binding affinity is fundamen-tally determined by the three-dimensional (3D) structure of the protein–ligand complex, which is not fully captured by primary sequence information alone. As a result, deep learning methods that operate on 3D structural inputs—whether using convolutional neural networks (e.g., Jiménez et al., 2018 ; Zheng et al., 2019 ; Wang et al., 2021 ) or graph neural networks (e.g., Son & Kim, 2021 )—are generally more accurate and robust. An important issue for scaling deep learning-based affinity prediction methods utilizing 3D struc-tures is the availability of high-quality crystal structure data and accurate binding affinity values. The number of known protein-ligand structures (both from cryo-EM or from X-ray crystallography) that are paired with measured binding affinity values is fairly limited considering the number of possible combinations that may exist in nature ( Askr et al., 2023 ; Libouban et al., 2023 ; Wang, 2024 ; Zeng et al., 2024 ). Moreover, surrogate models trained on these data are heavily impacted by the quality of crystal structures and accuracy of measured binding affinities used. There has been substantial progress in recent years at improving the coverage of these datasets, however existing datasets still lack sufficient coverage in both protein and ligand space. One possible solution to address the lack of data coverage, is to augment existing datasets with high-confidence computationally folded structures, through a process termed distillation. This ap-proach is frequently used in the field and has been applied to a number of recent works such as in AlphaFold ( Jumper et al., 2021 ; Abramson et al., 2024 ), Chai-1 ( Chai Discovery, 2024 ), or Neu-ralPlexer ( Qiao et al., 2024 ). Other groups have attempted to leverage experimentally determined strucutres by leveraging the PDB and linking them with other external resources like BindingDB. For example, the PLINDER dataset ( Durairaj et al., 2024 ), utilized this approach to curate nearly 450k protein-ligand pairs for both apo and holo structures. However, experimentally derived affinity values were only included if available from BindingDB and only represents a smaller fraction of the full dataset. In another example, the CrossDocked dataset Francoeur et al. (2020) used experimen-tally solved bound structures and docked bound ligands to other similar binding pockets. However, this dataset doesn’t include experimentally verified affinity values and only includes poses with bi-nary classification labels for use in downstream machine learning tasks. Thus, there is a clear need to further enhance the availability of protein-ligand pairs with bioactivity and binding affinity data to enable the training of large-scale supervised machine learning tasks. In this work, we present the Structurally Augmented IC50 Repository ( SAIR ): the largest publicly available dataset of protein-ligand 3D structures 1, with annotated potencies (some example struc-tures are shown in Fig. 1 ). Our data consists of 1, 048, 857 protein-ligand complexes, which were obtained from the ChEMBL ( Gaulton et al., 2012 ; Zdrazil et al., 2024 ) and BindingDB ( Liu et al., 2007 ; 2025 ) datasets. The data were preprocessed and filtered as described in §2, and the pro-tein–ligand complex structures were folded using the Boltz-1x model ( Wohlwend et al., 2024 ), a publicly available implementation inspired by AlphaFold 3. Results derived from our dataset are presented in §3, including the performance of two binding affinity prediction models: one trained on protein sequence and SMILES representations, and another trained on 3D structural data. We conclude with a summary of key findings in §4. Download figure Open in new tab Figure 1: Three example protein-ligand complexes from our SAIR dataset. Protein chains are color coded by pLDDT score with blue, yellow, and red regions corresponding to high, medium, and low confidence regions. From left to right the first complex corresponds to sample 4702 with Uniprot ID C7C422 and Ligand InchiKey BFOYHAUOQCJINB-UHFFFAOYSA-N, the second complex cor-responds to sample 501640 with Uniprot ID P28062 and Ligand InchiKey ODIWGDQSSAWSIL-GKWCIWIWSA-N, and the third complex corresponds to sample 253438 with Uniprot ID P07858 and Ligand InchiKey SQIHYRIDUCIHLT-HKUYNNGSSA-N. Our main contribution is the public release of the SAIR dataset. The paper describes how the dataset was obtained, and presents some analyses using the data. The data is available at https://www.sandboxaq.com/sair . Appendix §A, goes into more detail about the available data. 2 D ataset C onstruction 2.1 D ataset C uration Bioactivity data were obtained from the ChEMBL35 release ( Gaulton et al., 2012 ) and BindingDB (1Q2025) ( Liu et al., 2007 ), and subsequently curated using a minimal set of filters designed to retain a large volume of high-quality data. The specific filters are described below. ChEMBL35 Removed entries missing ligand SMILES or pchembl values. Removed entries which: did not have a UniProt ID for the protein target, referenced multi-ple protein targets, or referenced a protein variant. Removed entries with a data validity comment 1 . Removed entries where standard relation was . This step ensures that any measured values obtained are within the limit of detection for the assay. Only included assays that were flagged by ChEMBL as measuring binding (e.g., K i , IC50, K d ). Removed measurements outside of a reasonable biochemical assay dynamic range (1 pM < x < 100 µ M). BindingDB Removed entries missing molecule SMILES or IC50 values. Removed entries without a UniProt ID for the protein target or referenced multiple protein targets. Removed entries where reported IC50 values contained inequalities (i.e. ). This step ensures that any measured values obtained are within the limit of detection for the assay. Remove measurements outside of a reasonable biochemical assay dynamic range (1 pM < x < 100 µ M). After initial curation, data from both sources were merged into a single table. While this curation strategy can introduce variability in IC50 values by combining data from different assays Landrum & Riniker (2024) , this dataset is still fully compatible with the maximal curation strategy outlined in Landrum & Riniker (2024) for data points from ChEMBL. Because BindingDB does not perform curation at the level of specific assays, the maximal curation strategy is not compatible with that source. For protein-ligand complexes that appear in both ChEMBL and BindingDB, we keep the information from both datasets. All bio-activity values were converted to pIC50 units ( − log 10 ). SMILES strings for the ligand molecular structures were standardized by the removal of salts, protonation at neutral pH (where possible), and canonicalization using RDkit. Note that the choice of neutral pH for the ligand pro-tonation is immaterial for the subsequent computational prediction of the protein-ligand structures, as current cofolding models do not predict the positions of hydrogen atoms. A coarse ligand library filter was applied to exclude likely false positives/false negatives by remov-ing PAINS and molecules with molecular weights exceeding 1250 Da. Protein-ligand complexes containing proteins with more than 2000 amino acid residues were excluded, in order to increase the probability of successful prediction by the cofolding model on current GPU hardware. Next, duplicate entries were removed based on UniProt accession and canonical SMILES. The amino acid sequence for each protein was obtained from its UniProt entry using the accession number provided in the ChEMBL or BindingDB dataset. Note that this canonical sequence from UniProt may differ from the one used in the original bioactivity assay. For instance, the experimental protein may have been a truncated construct, a mutant, or a specific quaternary structure (e.g., a homodimer), whereas our analysis used the monomeric sequence from the database. Finally, to avoid data leakage when using this dataset to train or evaluate models that use structural data from the PDB for training, protein-ligand systems that already have experimentally-solved structures in the PDB were removed. The existence of a corresponding structure in the PDB was determined by finding the Chemical Component Dictionary (CCD) identifier of the ligand (by first computing its InChIKeyHeller et al. (2015) using RDKit) and looking for matches to this (Uniprot ID, CCD ID) pair in the PDB. This search utilized the RCSB GraphQL search APIrcs, and the PDBe REST API provided by EMBL-EBISIF. This results in 1, 048, 857 complexes, with 936, 702 from ChEMBL and 613, 597 from BindingDB (see Table 2 ). Note the number of complexes from each source adds up to more than the total number, because of complexes that appear in both sources. For structure generation, duplicated complexes were only folded once. The distribution of pIC50 values is shown in Fig. 2 . Download figure Open in new tab Figure 2: Histogram showing the distribution of pIC50 values stratified by data source (ChEMBL and BindingDB) and assay type (biochemical, cell-based, and unknown). Note that, as shown in Ta-ble 2, the assay type could not be inferred for the majority of complexes based on the curated assay descriptions. As a result, the bottom panels include only those complexes with known assays. For all histograms except the first one, we also plot the overall distribution, in grey. View this table: View inline View popup Download powerpoint Table 1: Comparison of Protein-Ligand Binding Databases View this table: View inline View popup Download powerpoint Table 2: Data distribution by source and assay. Note that the total number between both assays is larger than the number of generated structures. That is because of protein-ligand pairs that appear in both datasets 2.2 S tructure P rediction We used the Boltz-1x folding model to generate 3D structures for all protein-ligand complexes described in §2. Boltz-1 is a publicly available implementation of AlphaFold 3 2 . Additionally, Boltz-1x extends this model by introducing a guiding potential to the diffusion process to prevent clashes, resulting in more physically realistic binding poses. Due to the lack of quaternary structure information, all complexes were treated as monomeric assemblies. We generated five structure samples per complex, as this represents the maximum number we can compute on a single GPU for the longest protein sequence in the dataset (see §4 below for more details). While it is common practice to increase sample diversity by varying random seeds across multiple runs, we did not apply this technique due to resource constraints 3 . The Boltz-1x model was run using three recycling steps and 200 sampling steps (any other settings are the defaults as of the boltz-1x release). Note that, as all these systems are monomeric, the pairing strategy is irrelevant. Multiple sequence alignments (MSAs) for input to the model were generated using the MMseqs2 tool (Steinegger, 2017) (via the ColabFold Mirdita et al. (2022) project). This used the UniRef30 sequence database version 2302 and the ColabFoldDB metagenomic sequence database version 202108. 3 R esults 3.1 D ata S tatistics 3.1.1 P roteins The 1, 048, 857 protein-ligand systems in the dataset comprise 5, 149 unique protein sequences. Of these 5, 149 proteins present in our dataset, 2, 150 are believed to have no structures deposited in the PDB. The distribution of sequence lengths is shown in Fig. 3 . Most sequences fall within the 300-500 amino acid range. Beyond 500 residues, the frequency decreases steadily, with very long sequences (e.g., > 1500 amino acids) appearing only rarely. Download figure Open in new tab Figure 3: Distribution of protein sequence lengths across all unique entries in the dataset. Sequence clustering with MMseqs2 (Steinegger, 2017) using reasonable values for the minimum sequence identity and minimum coverage (MMseqs2 flags --min-seq-id and -c , respectively) revealed the presence of a large number of singleton clusters. For example, setting --min-seq-id and -c to [0.8, 0.8], [0.5, 0.7], and [0.3, 0.2] resulted in 3793, 2818, and 1862 clusters, respectively. Proteins were assigned to a family by using metadata provided by the UniProt database. We first classified them into enzymes or non-enzymes by looking at the presence of an enzymatic activity number (EC number). Enzymes were further subdivided into kinases (EC = 2.7 .x ), phosphatases (EC = 3.1 .x ) and other enzymes. Non-enzymes were subdivided by looking at their gene ontology codes (GO code). For example, the presence of GO = 0004879 implies that the protein is a nuclear receptor. The distribution of protein families accross different assays is shown in Fig. 4 . We see that the biochemical assay has a larger proportion of phosphatases, kinases and enzymes, while the cell assay has more nuclear receptors and GPCRs. There is also a larger number of proteins in the cell assay, for which we could not parse family information. Download figure Open in new tab Figure 4: Distribution of protein families across all unique entries in the dataset. We could not parse the protein family information for a fraction of the proteins, which are shown under ‘other’. 3.1.2 L igands We used RDKit ( Landrum et al., 2025 ) to compute a range of chemical descriptors, as summarized in Table 3 . The values capture the statistics of the unique ligands in the dataset. View this table: View inline View popup Download powerpoint Table 3: Chemical descriptors across the dataset (aggregation of unique ligand entries) It is worth noting that we did not perform any filtering on the dataset on the basis of things like “drug-likeness” of the ligands, for example filtering samples with ligands below a certain molecular weight. This is to avoid losing useful and chemically-meaningful data points of biologically-relevant species, such as ionic cofactors or small organic fragments that can teach a model protein-small-molecule interaction chemistry. It is our expectation that users will choose to filter samples based on ligand characteristics according to their use case. 3.2 P ose B usters We evaluated all generated protein-ligand structures using PoseBusters ( Buttenschoen et al., 2024 ), with results summarized in Table 4 . Overall, Boltz-1x performs well in generating physically valid structures, with only approximately 3% of structures failing any PoseBusters check. This number is consistent with the performance reported in Wohlwend et al. (2024) . Notably, only 0.53% of protein-ligand complexes had all five generated structures fail. Focusing on different protein families, we find that the model has lower failure rates for kinases, and phosphatases; and fails more often for GPCRs. One interesting case are nuclear receptors, where the overall failure rate is low, but there is a relatively large number of complexes, for which all structures failed (more than for any other family) indicating that some of the nuclear receptor complexes are particularly hard to fold. View this table: View inline View popup Download powerpoint Table 4: Summary of PoseBusters results by family For the assemblies that failed PoseBusters validation, we present a more fine-grained analysis of individual test outcomes in Fig. 5 . Across all assay types 4 the most frequent source of failure is the internal energy check, which accounts for more than half of all PoseBusters failures. Other common failure modes include the number of bonds, abnormal bond lengths, and ligands that could not be loaded by RDKit. Download figure Open in new tab Figure 5: Failure rate for each PoseBusters check, defined as the number of structures failing a given check divided by the total number of failed structures. Bar colors indicate assay type. The first entry in the x-axis corresponds to cases where RDKit failed to load the ligand. 3.3 B oltz C onfidence M etrics Given that we have access to experimental binding potencies for the corresponding complexes, we assessed whether Boltz-1x’s confidence metrics correlate with binding affinity. Prior work has demonstrated that AlphaFold confidence scores correlate with binding affinity in protein–protein interactions ( Zambaldi et al., 2024 ) (PPIs). Here, we explore whether similar correlations exist in the context of protein–ligand interactions. Results are shown in Fig. 6 .Focusing on the blue bars (Spearman correlation averaged across all assay types), we observe a significant correlation be-tween certain confidence metrics, particularly those involving the protein-ligand interface —namely iPTM, complex iPDE, and complex iPLDDT—and experimental potency. These findings suggest that Boltz-1x’s structural confidence metrics provide some predictive signal for protein–ligand bind-ing affinity. Notably, the strength of the correlation varies by assay type: it is highest for biochemical assays and weakest for cell assays. We hypothesize that this is caused by the great specificity and accuracy of biochemical assays, while cellular and homogenate assays may introduce additional confounding factors such as off-target binding, permeability, and intracellular dynamics. To further probe protein–ligand interaction quality, we introduce a new metric — interaction PTM — defined as the average of the off-diagonal values in the pair chains ptm confidence head. This metric captures the confidence of the protein with respect to the ligand, and vice versa, and is analogous to the “interaction PAE” described in Zambaldi et al. (2024) . We find that interaction PTM exhibits strong correlation with binding affinity ( r s = 0.25), ranking second only to iPTM ( r s = 0.27) in predictive power across our dataset. Download figure Open in new tab Figure 6: Comparison of the Spearman correlation r s between different Boltz-1x confidence metrics, and the experimental IC50 activity, by assay type. The signs shown in the title of each panel indicate the expected direction of correlation: negative for PDE and iPDE (as they represent distances), and positive for all other metrics. We can furthermore look at the similarity between generated protein chains and protein chains from the training dataset. We find a high degree of correlation (Spearman correlation of 0.47) between the global PTM confidence and the highest TM-score (normalized by query length) to structures in the training set. This metric is independent of the ligand, which partly explains the poor correlation to the binding affinity. It is better suited for evaluating the global shape of the protein. 3.4 B inding A ffinity M odels We use the SAIR dataset to benchmark the performance of several binding affinity prediction models. The combination of cofolding-generated structures with structure-based affinity predic-tion represents a promising and increasingly adopted approach in the scientific community. By providing high-throughput structural models paired with experimentally observed IC50 values, the SAIR dataset enables rigorous evaluation of this emerging class of predictive methods. The field of protein–ligand binding affinity prediction is supported by a vast and diverse body of literature, and a comprehensive comparison of all available methods is beyond the scope of this work. Instead, we focus on three representative and methodologically distinct approaches to binding affinity prediction: Empirical scoring functions : We employ two different traditional empirical scoring functions: AutoDock Vina (henceforth referred to as Vina) ( Trott & Olson, 2010 ) and Vinardo ( Quiroga & Villarreal, 2016 ), both calculated using the GNINA library ( McNutt et al., 2021 ). We additionally evaluate first minimizing the ligand pose using the Vina scor-ing function before scoring the resulting structure, again using Vina (this is referred to as “Vina minimized”). Convolutional neural network (CNN): As a first method of structure-based machine learning affinity prediction, we employ a three-dimensional CNN, and represent the input by projecting it into 3D voxels. There are various available 3D CNN methods for affinity prediction, but we use Onionnet-2 ( Wang et al., 2021 ), one of the state-of-the-art methods. Graph neural network (GNN): As an alternative structure-based machine learning method, we consider a GNN. GNNs are, theoretically, better suited for the task of affinity prediction, as protein-ligand systems are easily represented as graphs. However, regression from graphs is generally a harder task than regression using voxels, as graph convolutions are non-trivial ( Zhang et al., 2019 ). As our GNN, we use the AEV-PLIG model ( Warren et al., 2024 ; Valsson et al., 2025 ), which recently showed state-of-the-art performance in structure-based binding affinity prediction. In all cases (with the exception of the “Vina minimized” approach, as explained above), the given affinity prediction tool is evaluated on the predicted three-dimensional protein-ligand structure as-is. There are many other methods that could be used for benchmarking binding affinity prediction. For example, the recently developed Boltz-2 5 accomplishes accurate binding affinity prediction via re-gression from intermediate embeddings. However, it is trained on similar data to what we present here, therefore it was not used for comparison. On the side of physics-based methods, Free Energy Perturbation (FEP) methods are considered the most accurate, however they are very computation-ally expensive, making it difficult to use them in a dataset of millions of structures like the one presented in this work. We use four metrics to compare the performance of the various binding affinity methods: Spearman Correlation : The Spearman correlation between the predicted and experimen-tal binding affinity, defined as: where d i is the difference between the ranks of the predicted and experimental binding affinities, and n is the number of samples. Pearson Correlation : The Pearson correlation between the predicted and experimental binding affinity, defined as: where x i and y i are the predicted and experimental binding affinities, respectively, and x̅ and y̅ are the means of the predicted and experimental binding affinities, respectively. Kendall’s Tau : Kendall’s Tau is a measure of the ordinal association between two quanti-ties, defined as: where n c is the number of concordant pairs, and n d is the number of discordant pairs, and n is the number of samples. Area Under the Curve (AUC) : The AUC is a measure of the ability of a model to distin-guish between positive and negative samples. It is defined as the area under the Receiver Operating Characteristic (ROC) curve, which is a plot of the true positive rate against the false positive rate. To calculate the AUC, we first need to define a threshold for the pre-dicted binding affinity, and then compute the true positive rate (TPR) and false positive rate (FPR) for that threshold. We use a threshold of 100 µM , which is a common threshold for binding affinity prediction. We restrict our evaluation to structures derived from ChEMBL, as BindingDB includes experimental protein–ligand complexes that were used in the training of AEV-PLIG. Although the structures in our dataset are synthetically generated and not identical to those used in training, we exclude BindingDB entries to minimize the risk of data leakage and to ensure a fair comparison. As an additional consideration, AEV-PLIG was trained using the BindingNet database Li et al. (2024) , which contains synthetically generated structure-activity relationship data derived from ChEMBL. This could account for AEV-PLIG’s enhanced performance. Nevertheless, it should be noted that the BindingNet training set represents a minor portion of the overall SAIR database ( ≤ 5%), and the structures are likely distinct from those in the Boltz computed set. We present the results of the model comparison in Fig. 7 . Across all assay types, the GNN method achieves the highest performance, followed by the CNN method, with the empirical scoring func-tions performing the worst. However, none of the methods achieve a very high correlation, with Spearman correlations comparable to the ones achieved by some of the interface confidence met-rics (even though those were not specifically tuned for binding affinity prediction), such as iPTM, also shown the figure. The shaded bars in the figure show the results when we only keep structures for which Boltz-1x predicts a high confidence ( > 0.8). We find that keeping only these structures improves performance of almost all models, as the structures are more likely to be correct. Download figure Open in new tab Figure 7: The various metrics we use to compare predicted and experimental binding affinity, for each of the binding affinity prediction methods studied in the paper, and for each of the assays. The shaded lines show results when using only the structures for which Boltz-1x predicted a high confidence. It is important to note that both the GNN and CNN models were originally trained on experimental structures, and our evaluation is conducted on synthetic structures generated via cofolding. Fine-tuning these models on a subset of the synthetic dataset would likely improve their performance and better align them with the structural distribution seen at inference time. 3.5 P ocket diversity We can use our model to gain insight into the effect that changing the input ligand has in the gen-erated protein conformation. If we give Boltz-1x sufficiently distinct ligands, is the model able to detect different appropriate binding sites, or will it re-use the pockets it has seen during training? To address this, we need to define a pocket. First, we define pocket residue as every residue that has a non-hydrogen atom within 6 Å to the closest ligand atom, a cutoff commonly adopted in the field. The set of pocket residues defines a pocket, and two pockets are considered similar if where we defined the threshold to be 0.8. That allows us to cluster the detected pockets per protein into groups, such that no group shares a similar pocket. Fig. 8 shows both the diversity in pockets over the five generated structures per protein/ligand com-plex (left panel), and the diversity for a given protein as we change ligands (right). AlphaFold3-like models are known to generate similar conformations for different diffusion samples, which we also find when looking at the pocket diversity, with most complexes generating ligands in the same pocket for all five generated samples. Download figure Open in new tab Figure 8: Pocket diversity for different samples and different ligands. Pocket similarity was deter-mined using Eq. (4) . Left : The diversity of pockets in the five generated structures per protein/ligand complex. Right : The diversity of pockets for different proteins. However when looking at a protein chain, with different ligands, we find a significant fraction of systems with a variable set of pockets. The most extreme example of this is protein P10636, where we found well over a thousand different potential binding sites for 345 different ligands, two of which are visualized in figure 9 . Download figure Open in new tab Figure 9: Diverse binding sites in protein P10636. This shows a potential reason for the fat tail in the number of distinct pockets per protein. When Boltz-1x is uncertain about the structure or if the protein is very flexible, then we find a very diverse set of protein conformations. The pocket-residues will similarly change a lot, and there is no well defined binding site. We can also study pocket similarity to the training set. Fig. 10 shows the result of performing a similarity search of all generated structures against the Boltz-1x training data, the pocket-LDDT score, as defined in Durairaj et al. (2024) . The pocket-LDDT is defined by structurally aligning predicted structures to ground truth structures, and calculating the average LDDT over the backbone carbon atoms in the aligned pocket residues. In our case, this score is not a measure of correctness, but more a measure of similarity to the training dataset. Download figure Open in new tab Figure 10: The distribution of pocket LDDT’s calculated, comparing the generated structures to the training dataset. We find no correlation between the pocket-LDDT and the interface confidence (Spearman corre-lation of −0.02), but we do find most pockets to be highly similar to the training data. We have previously seen that Boltz is able to generate multiple distinct binding poses, which together implies that Boltz successfully places ligand atoms in plausible looking pockets. 4 C onclusions In this work, we introduce the Structurally Augmented IC50 Repository ( SAIR ), a large-scale dataset of protein–ligand 3D structures paired with annotated binding affinities. Comprising 5, 244, 285 synthetically-generated structures representing 1, 048, 857 protein-ligand complexes, each annotated with experimentally-determined potency, the dataset is designed to significantly ex-pand the volume of data available for training and evaluating structure-based deep learning models in drug discovery. We rigorously evaluated the quality of the generated structures using PoseBusters and observed a low overall failure rate of approximately 3%. To assess the utility of the dataset for predictive modeling, we benchmarked several structure-based binding affinity prediction methods. Graph neural networks performed best, followed by convolutional neural networks and empirical scoring function methods. However, all models achieved only modest correlations, comparable to those obtained from the folding model’s interface confidence metrics. This suggests that models trained on experimental structures may not generalize well to synthetic data, highlighting the potential need for fine-tuning on generated complexes. Looking ahead, this work opens several promising avenues for future research. We plan to evaluate the performance of our binding affinity prediction algorithm AQFEP ( Crivelli-Decker et al., 2024 ) on the SAIR dataset, to directly compare its performance with the methods benchmarked in this paper. We also aim to use SAIR to improve the performance of AQFEP. More broadly, fine-tuning existing affinity prediction models—or developing new architectures specifically optimized for syn-thetic protein–ligand complexes—could lead to significant gains in predictive accuracy. Beyond affinity prediction, the dataset may also support self-distillation strategies for cofolding models, as has been demonstrated in other data modalities. Finally, SAIR represents a valuable resource for inverse design tasks, enabling generative ap-proaches to create new ligands conditioned on a target protein—further expanding the possibilities for structure-based drug discovery. A uthor C ontributions PL co-lead the data generation effort, and led the writing of the manuscript. ZB co-lead the data generation effort, and significantly contributed to the manuscript. SB, MvD, TM and PJ made important contributions to the data generation, and to improv-ing the manuscript. JCD and BS led the data curation effort, and contributed to manuscript writing. NDM contributed to cloud computing efforts. TC and RSF significantly contributed to improving the manuscript. AJN and PA provided support for GPU access. MG proposed the original idea, and supervised the effort. H ardware S pecifications This work was performed on an H100 cluster in DGX Cloud on Google Cloud provisioned by NVIDIA’s AI Accelerator team. This was initially provisioned as a managed 35-Node (280 GPUs) cluster, and was non-disruptively scaled up to 96 nodes (760 GPUs) to allow for scale-out workloads, with support from the NVIDIA AI Accelerator and NVIDIA’s BioNemo team. Along the way, the NVIDIA AI Accelerator team helped SandboxAQ through both infrastructure and workload optimization, leveraging technologies like Parallelstore persistent volumes to speed up data transfer to and between nodes within their managed GKE environment. Highly granular node, operator, scheduler, and GPU metrics were captured at all stages to help NVIDIA AI Accelerator and SandboxAQ engineering teams collect the necessary information to identify bottlenecks and tweak configurations to achieve the highest workload throughput and effective GPU utilization. SandboxAQ and NVIDIA AI Accelerator were able to achieve > 95% GPU compute utilization during these runs. In total, these calculations used approximately 130k GPU-hours of compute time on this cluster. D ata A vailability The SAIR dataset can be found at https://www.sandboxaq.com/sair . This page provides links for downloading all components of the dataset, which is hosted on Google Cloud. Data generated and made available by SandboxAQ is shared under the Creative Commons Attri-bution Non-Commercial-ShareAlike License (CC BY-NC-SA 4.0). Data from other sources cited, such as ChEMBL and BindingDB, are made available under their respective licenses. SandboxAQ intends to grant permission for commercial use of this data to third parties, and does not intend to charge for such use, subject to a third party requesting such additional permissions using this form and receiving written authorization from SandboxAQ. A IC50 D ataset S tructure and C ontents This appendix provides a detailed description of the SAIR dataset, its file organization, and the contents of its primary data files. A.1 D ataset M anifest The IC50 dataset is organized into the following primary components: sair.parquet : This is the central dataframe of the dataset, containing all curated IC50 data, associated original source metadata, results from PoseBusters structural valid-ity checks, Boltz-1x prediction confidence measures, and other relevant metadata for each protein-ligand complex. structures/ : This directory contains all the predicted 3D structures generated by the Boltz-1x co-folding model. For each unique protein-ligand complex, five distinct predicted structures (referred to as “models”) are provided. Each structure is stored as a .cif file, named according to the convention sample _model_.cif . Here, entry id corresponds to the entry id field in sair.parquet , and model identifies the specific sample (0 through 4) from the co-folding model. prediction confidences/ : This directory stores the raw .json confidence files directly produced by the Boltz-1x model. The files follow the naming convention confidence_sample__model_.json , with entry id and model consistent with the structure files. These files contain detailed confidence metrics beyond those summarized in sair.parquet . prediction plddts/ : This directory contains the raw .npz files for Predicted Lo-cal Distance Difference Test (pLDDT) values, also output directly by the Boltz-1x model. These files are named plddt_model_.npz , using the same entry id and model identifiers. pLDDT values provide residue-level confidence esti-mates for the predicted structures. A.2 sair.parquet C olumn D escriptions The sair.parquet file serves as the main tabular data source, consolidating key information for each structure-potency pair. Table 5 , Table 6 , Table 7 and Table 8 ; describe each column in sair.parquet, categorized for clarity. View this table: View inline View popup Download powerpoint Table 5: Description of Columns in sair.parquet : Identifiers and Inputs View this table: View inline View popup Download powerpoint Table 6: Description of Columns in sair.parquet : Potency Data and Metadata View this table: View inline View popup Download powerpoint Table 7: Description of Columns in sair.parquet : Binding Affinity Predictions and Model Confidence View this table: View inline View popup Download powerpoint Table 8: Description of Columns in sair.parquet : PoseBusters Results Acknowledgments This research was supported by the NVIDIA AI Accelerator team on an NVIDIA DGX Cloud Cluster. We extend our sincere gratitude to NVIDIA for their valuable collaboration and support throughout this project. We thank Amit Kadan, Kenneth Goossens, Kevin Ryzcko, Mary Pitman, Takeshi Yamazaki, Andrea Bortolato and Adam Lewis; for their valuable feedback on the manuscript. We also thank ChEMBL and BindingDB for creating and maintaining their open-source datasets, which were fundamental to this work. Furthermore, we are grateful to the MIT Jameel Clinic for making the Boltz model openly available. Footnotes https://www.sandboxaq.com/sair ↵ 1 The data validity comment was introduced in ChEMBL15 and includes information about the quality of the entry and to allow users to make an informed decision on whether to include that value in their analyses ( https://chembl.blogspot.com/2020/10/data-checks.html ). ↵ 2 There are some minor changes between AlphaFold 3 and Boltz-1, such as the strategy used for Multiple Sequence Alignment (MSA) subsampling. ↵ 3 Further, Boltz-1’s MSA subsampling is deterministic with respect to the random seed, unlike other cofold-ing models such as AlphaFold3 and Chai-1 ( Chai Discovery, 2024 )), where seed variation is a primary source of stochasticity. As a result, we do not expect significant diversity gains from seed variation in Boltz-1x. ↵ 4 We do not show results for the homogenate assay, as there are not enough structures, compared with the rest. ↵ 5 while no pre-print is available in any official pre-print server at the point of writing, the Boltz-2 pa-per can be found in https://cdn.prod.website-files.com/68404fd075dba49e58331ad9/6842ee1285b9af247ac5a122_boltz2.pdf R eferences Embl-ebi: Rest calls based on sifts mappings . https://www.ebi.ac.uk/pdbe/api/doc/sifts.html . Accessed: 2025-06-13 . Rcsb pdb data api . https://data.rcsb.org/#data-api . Accessed: 2025-06-13 . ↵ Josh Abramson , Jonas Adler , Jack Dunger , Richard Evans , Tim Green , Alexander Pritzel , Olaf Ronneberger , Lindsay Willmore , Andrew J Ballard , Joshua Bambrick , et al. Accurate structure prediction of biomolecular interactions with alphafold 3 . Nature , 630 ( 8016 ): 493 – 500 , 2024 . OpenUrl CrossRef PubMed ↵ Heba Askr , Enas Elgeldawi , Heba Aboul Ella , Yaseen AMM Elshaier , Mamdouh M Gomaa , and Aboul Ella Hassanien . Deep learning in drug discovery: an integrative review and future chal-lenges . Artificial Intelligence Review , 56 ( 7 ): 5975 – 6037 , 2023 . OpenUrl CrossRef PubMed ↵ Martin Buttenschoen , Garrett M Morris , and Charlotte M Deane . Posebusters: Ai-based docking methods fail to generate physically valid poses or generalise to novel sequences . Chemical Sci-ence , 15 ( 9 ): 3130 – 3139 , 2024 . OpenUrl CrossRef PubMed ↵ Chai Discovery . Chai-1: Decoding the molecular interactions of life . bioRxiv , 2024 . doi: 10.1101/2024.10.10.615955 . URL https://www.biorxiv.org/content/early/2024/10/11/2024.10.10.615955 . OpenUrl Abstract / FREE Full Text ↵ Zoe Cournia , Bryce Allen , and Woody Sherman . Relative binding free energy calculations in drug discovery: recent advances and practical considerations . Journal of chemical information and modeling , 57 ( 12 ): 2911 – 2937 , 2017 . OpenUrl CrossRef PubMed ↵ Jordan E Crivelli-Decker , Zane Beckwith , Gary Tom , Ly Le , Sheenam Khuttan , Romelia Salomon- Ferrer , Jackson Beall , Rafael Gómez-Bombarelli , and Andrea Bortolato . Machine learning guided aqfep: a fast and efficient absolute free energy perturbation solution for virtual screening . Journal of Chemical Theory and Computation , 20 ( 16 ): 7188 – 7198 , 2024 . OpenUrl ↵ Janani Durairaj , Yusuf Adeshina , Zhonglin Cao , Xuejin Zhang , Vladas Oleinikovas , Thomas Duig-nan , Zachary McClure , Xavier Robin , Daniel Kovtun , Emanuele Rossi , et al. Plinder: The protein-ligand interactions dataset and evaluation resource . bioRxiv , pp. 2024 – 07 , 2024 . ↵ Paul G. Francoeur , Tomohide Masuda , Jocelyn Sunseri , Andrew Jia , Richard B. Iovanisci , Ian Sny-der , and David R. Koes . Three-dimensional convolutional neural networks and a cross-docked data set for structure-based drug design . Journal of Chemical Information and Modeling , 60 ( 9 ): 4200 – 4215 , 2020 . doi: 10.1021/acs.jcim.0c00411 . URL https://doi.org/10.1021/acs.jcim.0c00411. OpenUrl CrossRef PubMed ↵ Anna Gaulton , Louisa J Bellis , A Patricia Bento , Jon Chambers , Mark Davies , Anne Hersey , Yvonne Light , Shaun McGlinchey , David Michalovich , Bissan Al-Lazikani , et al. Chembl: a large-scale bioactivity database for drug discovery . Nucleic acids research , 40 ( D1 ): D1100 – D1107 , 2012 . OpenUrl CrossRef PubMed Web of Science Stephen R Heller , Alan McNaught , Igor Pletnev , Stephen Stein , and Dmitrii Tchekhovskoi . Inchi, the iupac international chemical identifier . Journal of cheminformatics , 7 : 1 – 34 , 2015 . OpenUrl CrossRef PubMed ↵ Mingjian Jiang , Shuang Wang , Shugang Zhang , Wei Zhou , Yuanyuan Zhang , and Zhen Li . Sequence-based drug-target affinity prediction using weighted graph neural networks . BMC ge-nomics , 23 ( 1 ): 449 , 2022 . ↵ José Jiménez , Miha Skalic , Gerard Martinez-Rosell , and Gianni De Fabritiis . K deep: protein–ligand absolute binding affinity prediction via 3d-convolutional neural networks . Journal of chemical information and modeling , 58 ( 2 ): 287 – 296 , 2018 . OpenUrl CrossRef PubMed ↵ John Jumper , Richard Evans , Alexander Pritzel , Tim Green , Michael Figurnov , Olaf Ronneberger , Kathryn Tunyasuvunakool , Russ Bates , Augustin ΁ídek , Anna Potapenko , et al. Highly accurate protein structure prediction with alphafold . nature , 596 ( 7873 ): 583 – 589 , 2021 . OpenUrl CrossRef PubMed ↵ Greg Landrum , Paolo Tosco , Brian Kelley , Ricardo Rodriguez , David Cosgrove , Riccardo Vianello , sriniker , Peter Gedeck , Gareth Jones , Eisuke Kawashima , NadineSchneider , Dan Nealschneider , Andrew Dalke , Matt Swain , Brian Cole , tadhurst cdd , Samo Turk , Aleksandr Savelev , Alain Vaucher , Maciej Wójcikowski , Ichiru Take , Rachel Walker , Vincent F. Scalfani , Hussein Faara , Kazuya Ujihara , Daniel Probst , Niels Maeder , Jeremy Monat , Juuso Lehtivarjo , and guillaume godin . rdkit/rdkit: 2025 03 3 (q1 2025) release , June 2025 . URL doi: 10.5281/zenodo.15605628 . OpenUrl CrossRef ↵ Gregory A. Landrum and Sereina Riniker . Combining ic50 or ki values from different sources is a source of significant noise . Journal of Chemical Information and Modeling , 64 ( 5 ): 1560 – 1567 , 2024 . doi: 10.1021/acs.jcim.4c00049 . URL https://doi.org/10.1021/acs.jcim.4c00049. OpenUrl CrossRef PubMed ↵ Xuelian Li , Cheng Shen , Hui Zhu , Yujian Yang , Qing Wang , Jincai Yang , and Niu Huang . A high-quality data set of protein–ligand binding interactions via comparative complex structure modeling . Journal of Chemical Information and Modeling , 64 ( 7 ): 2454 – 2466 , Jan 2024 . doi: 10.1021/acs.jcim.3c01170 . OpenUrl CrossRef PubMed ↵ Pierre-Yves Libouban , Samia Aci-Sèche , Jose Carlos Gómez-Tamayo , Gary Tresadern , and Pascal Bonnet . The impact of data on structure-based binding affinity predictions using deep neural networks . International Journal of Molecular Sciences , 24 ( 22 ): 16120 , 2023 . ↵ Sarita Limbu and Sivanesan Dakshanamurthy . A new hybrid neural network deep learning method for protein–ligand binding affinity prediction and de novo drug design . International Journal of Molecular Sciences , 23 ( 22 ): 13912 , 2022 . ↵ Tiqing Liu , Yuhmei Lin , Xin Wen , Robert N Jorissen , and Michael K Gilson . Bindingdb: a web-accessible database of experimentally determined protein–ligand binding affinities . Nucleic acids research , 35 ( suppl 1 ): D198 – D201 , 2007 . OpenUrl CrossRef PubMed Web of Science ↵ Tiqing Liu , Linda Hwang , Stephen K Burley , Carmen I Nitsche , Christopher Southan , W Patrick Walters , and Michael K Gilson . Bindingdb in 2024: a fair knowledgebase of protein-small molecule binding data . Nucleic acids research , 53 ( D1 ): D1633 – D1644 , 2025 . OpenUrl CrossRef PubMed ↵ Andrew T McNutt , Paul Francoeur , Rishal Aggarwal , Tomohide Masuda , Rocco Meli , Matthew Ragoza , Jocelyn Sunseri , and David Ryan Koes . Gnina 1.0: molecular docking with deep learn-ing . Journal of cheminformatics , 13 ( 1 ): 43 , 2021 . ↵ Milot Mirdita , Konstantin Schütze , Yoshitaka Moriwaki , Lim Heo , Sergey Ovchinnikov , and Martin Steinegger . Colabfold: making protein folding accessible to all . Nature methods , 19 ( 6 ): 679 – 682 , 2022 . OpenUrl CrossRef PubMed Hakime Öztürk , Elif Ozkirimli , and Arzucan Özgür . Widedta: prediction of drug-target binding affinity . arXiv preprint arXiv:1902.04166 , 2019 . ↵ Zhuoran Qiao , Weili Nie , Arash Vahdat , Thomas F Miller III . , and Animashree Anandkumar . State-specific protein–ligand complex structure prediction with a multiscale deep generative model . Nature Machine Intelligence , 6 ( 2 ): 195 – 208 , 2024 . OpenUrl CrossRef ↵ Rodrigo Quiroga and Marcos A Villarreal . Vinardo: A scoring function based on autodock vina improves scoring, docking, and virtual screening . PloS one , 11 ( 5 ): e0155183 , 2016 . OpenUrl CrossRef PubMed ↵ Jeongtae Son and Dongsup Kim . Development of a graph convolutional neural network model for efficient prediction of protein-ligand binding affinities . PloS one , 16 ( 4 ): e0249404 , 2021 . OpenUrl CrossRef PubMed Söding J. Steinegger , M. Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets . Nature Biotechnology , 35 : 1026 – 1028 , 2017 . OpenUrl CrossRef PubMed ↵ Oleg Trott and Arthur J Olson . Autodock vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading . Journal of computational chemistry , 31 ( 2 ): 455 – 461 , 2010 . OpenUrl CrossRef PubMed Web of Science ↵ Ísak Valsson , Matthew T Warren , Charlotte M Deane , Aniket Magarkar , Garrett M Morris , and Philip C Biggin . Narrowing the gap between machine learning scoring functions and free energy perturbation using augmented data . Communications Chemistry , 8 ( 1 ): 41 , 2025 . ↵ Ercheng Wang , Huiyong Sun , Junmei Wang , Zhe Wang , Hui Liu , John ZH Zhang , and Tingjun Hou . End-point binding free energy calculation with mm/pbsa and mm/gbsa: strategies and applications in drug design . Chemical reviews , 119 ( 16 ): 9478 – 9508 , 2019 . OpenUrl CrossRef PubMed ↵ Huiwen Wang . Prediction of protein–ligand binding affinity via deep learning models . Briefings in Bioinformatics , 25 ( 2 ): bbae081 , 2024 . OpenUrl CrossRef PubMed ↵ Zechen Wang , Liangzhen Zheng , Yang Liu , Yuanyuan Qu , Yong-Qiang Li , Mingwen Zhao , Yuguang Mu , and Weifeng Li . Onionnet-2: a convolutional neural network model for predicting protein-ligand binding affinity based on residue-atom contacting shells . Frontiers in chemistry , 9 : 753002 , 2021 . ↵ Matthew Warren , Charlotte Deane , Aniket Magarkar , Garrett Morris , Philip Biggin , et al. How to make machine learning scoring functions competitive with fep . 2024 . ↵ Jeremy Wohlwend , Gabriele Corso , Saro Passaro , Mateo Reveiz , Ken Leidal , Wojtek Swiderski , Tally Portnoi , Itamar Chinn , Jacob Silterra , Tommi Jaakkola , et al. Boltz-1: Democratizing biomolecular interaction modeling . bioRxiv , pp. 2024 – 11 , 2024 . ↵ Darrin M. York . Modern alchemical free energy methods for drug discoveryexplained . ACS Physical Chemistry Au , 3 ( 6 ): 478 – 491 , 2023 . doi: 10.1021/acsphyschemau.3c00033 . OpenUrl CrossRef PubMed ↵ Vinicius Zambaldi , David La , Alexander E Chu , Harshnira Patani , Amy E Danson , Tristan OC Kwan , Thomas Frerix , Rosalia G Schneider , David Saxton , Ashok Thillaisundaram , et al. De novo design of high-affinity protein binders with alphaproteo . arXiv preprint arXiv:2409.08022 , 2024 . ↵ Barbara Zdrazil , Eloy Felix , Fiona Hunter , Emma J Manners , James Blackshaw , Sybilla Corbett , Marleen de Veij , Harris Ioannidis , David Mendez Lopez , Juan F Mosquera , et al. The chembl database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods . Nucleic acids research , 52 ( D1 ): D1180 – D1192 , 2024 . OpenUrl CrossRef PubMed ↵ Xin Zeng , Shu-Juan Li , Shuang-Qing Lv , Meng-Liang Wen , and Yi Li . A comprehensive review of the recent advances on predicting drug-target affinity based on deep learning . Frontiers in Pharmacology , 15 : 1375522 , 2024 . ↵ Si Zhang , Hanghang Tong , Jiejun Xu , and Ross Maciejewski . Graph convolutional networks: a comprehensive review . Computational Social Networks , 6 ( 1 ): 1 – 23 , 2019 . OpenUrl CrossRef ↵ Liangzhen Zheng , Jingrong Fan , and Yuguang Mu . Onionnet: a multiple-layer intermolecular-contact-based convolutional neural network for protein–ligand binding affinity prediction . ACS omega , 4 ( 14 ): 15956 – 15965 , 2019 . OpenUrl CrossRef PubMed View the discussion thread. Back to top Previous Next Posted June 21, 2025. Download PDF Data/Code Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following SAIR: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share SAIR: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset Pablo Lemos , Zane Beckwith , Sasaank Bandi , Maarten van Damme , Jordan Crivelli-Decker , Benjamin J. Shields , Thomas Merth , Punit K. Jha , Nicola De Mitri , Tiffany J. Callahan , AJ Nish , Paul Abruzzo , Romelia Salomon-Ferrer , Martin Ganahl bioRxiv 2025.06.17.660168; doi: https://doi.org/10.1101/2025.06.17.660168 Share This Article: Copy Citation Tools SAIR: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset Pablo Lemos , Zane Beckwith , Sasaank Bandi , Maarten van Damme , Jordan Crivelli-Decker , Benjamin J. Shields , Thomas Merth , Punit K. Jha , Nicola De Mitri , Tiffany J. Callahan , AJ Nish , Paul Abruzzo , Romelia Salomon-Ferrer , Martin Ganahl bioRxiv 2025.06.17.660168; doi: https://doi.org/10.1101/2025.06.17.660168 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Biochemistry Subject Areas All Articles Animal Behavior and Cognition (7617) Biochemistry (17633) Bioengineering (13856) Bioinformatics (41841) Biophysics (21399) Cancer Biology (18529) Cell Biology (25422) Clinical Trials (138) Developmental Biology (13352) Ecology (19860) Epidemiology (2067) Evolutionary Biology (24281) Genetics (15582) Genomics (22461) Immunology (17700) Microbiology (40295) Molecular Biology (17140) Neuroscience (88413) Paleontology (666) Pathology (2823) Pharmacology and Toxicology (4813) Physiology (7632) Plant Biology (15107) Scientific Communication and Education (2042) Synthetic Biology (4284) Systems Biology (9808) Zoology (2267)

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00