Large scale prospective evaluation of co-folding across 557 Mac1-ligand complexes and three virtual screens

preprint OA: closed
Full text JSON View at publisher

Abstract

Accurate prediction of ligand-bound protein complexes and ranking them by affinity are central problems in drug discovery. While deep learning co-folding methods can help address these challenges, their evaluation has been hampered by the difficulties in assessing independence from training data and insufficiently large test sets. Here we test the ability of co-folding methods to predict the structures of 557 ligands bound to the SARS-CoV-2 NSP3 macrodomain (Mac1) that were determined after the training cut-off dates. AlphaFold3 (AF3), Boltz-2, and Chai-1 each reproduced >50% of the Mac1 ligand poses to better than 2 Å RMSD of experiment. Despite the potential for co-folding to describe protein conformational changes that stabilize ligand binding, we did not find that common conformational rearrangements, including peptide flip and a large loop opening, were recapitulated by the co-folding prediction. For AF3 and Chai-1, ligand pose prediction confidence weakly, but significantly, tracked experimental potency, while DOCK3.7 energies were only weakly correlated. Boltz-2 affinity predictions showed the strongest correlation with measured potency and, after calibration, achieved lower mean absolute error than a baseline predictor. We next assessed whether co-folding scores could rescore docking hit-lists to distinguish true ligands from non-binders among hundreds of molecules prospectively experimentally tested against AmpC β-lactamase, the dopamine D4 and the σ₂ receptors. AF3 ligand pose confidence values did not separate true ligands from high-scoring false-positives as effectively as docking scores or Boltz-2 affinity predictions did. Taken together, the modest, but independent correlations of docking score and co-folding confidence or affinity suggests that integrating physics-based and deep-learning and approaches may help with hit prioritization and subsequent optimization in structure-based ligand discovery.
Full text 3,809 characters · extracted from oa-doi-fallback · click to expand
Abstract Accurate prediction of ligand-bound protein complexes and ranking them by affinity are central problems in drug discovery. While deep learning co-folding methods can help address these challenges, their evaluation has been hampered by the difficulties in assessing independence from training data and insufficiently large test sets. Here we test the ability of co-folding methods to predict the structures of 557 ligands bound to the SARS-CoV-2 NSP3 macrodomain (Mac1) that were determined after the training cut-off dates. AlphaFold3 (AF3), Boltz-2, and Chai-1 each reproduced >50% of the Mac1 ligand poses to better than 2 Å RMSD of experiment. Despite the potential for co-folding to describe protein conformational changes that stabilize ligand binding, we did not find that common conformational rearrangements, including peptide flip and a large loop opening, were recapitulated by the co-folding prediction. For AF3 and Chai-1, ligand pose prediction confidence weakly, but significantly, tracked experimental potency, while DOCK3.7 energies were only weakly correlated. Boltz-2 affinity predictions showed the strongest correlation with measured potency and, after calibration, achieved lower mean absolute error than a baseline predictor. We next assessed whether co-folding scores could rescore docking hit-lists to distinguish true ligands from non-binders among hundreds of molecules prospectively experimentally tested against AmpC β-lactamase, the dopamine D4 and the σ₂ receptors. AF3 ligand pose confidence values did not separate true ligands from high-scoring false-positives as effectively as docking scores or Boltz-2 affinity predictions did. Taken together, the modest, but independent correlations of docking score and co-folding confidence or affinity suggests that integrating physics-based and deep-learning and approaches may help with hit prioritization and subsequent optimization in structure-based ligand discovery. Competing Interest Statement N.J.K.: The Krogan laboratory has received research support from Vir Biotechnology, F. Hoffmann-La Roche, and Rezo Therapeutics. N.J.K. has financially compensated consulting agreements with Maze Therapeutics and Interline Therapeutics. He is on the Board of Directors and is President of Rezo Therapeutics and is a shareholder in Tenaya Therapeutics, Maze Therapeutics, Rezo Therapeutics, GEn1E Lifesciences, and Interline Therapeutics. A.R.R.: A.R.R. is a co-founder and holds equity in BridgeBio Oncology Therapeutics, Tatara Therapeutics, and Elgia Therapeutics. A.A.: A.A. is a co-founder of Tango Therapeutics, Azkarra Therapeutics, and Kytarro; serves on boards and advisory boards as listed; receives research support from SPARC; and holds patents on the use of PARP inhibitors held jointly with AstraZeneca. B.K.S.: B.K.S. is co-founder of BlueDolphin LLC, Epiodyne Inc, and Deep Apple Therapeutics, Inc., and serves on the SRB of Genentech and the SAB of Schrodinger LLC and Vilya Therapeutics. J.S.F.: J.S.F. holds equity in Relay Therapeutics, Impossible Foods, Arda Therapeutics, Profluent Bio, Interdict Bio (co-founder), Vilya Therapeutics, and Edison Scientific Inc and is a paid consultant for Relay Therapeutics, Profluent Bio, Vilya Therapeutics, and Monimoi Therapeutics. All other authors declare no competing interests. Footnotes We addressed comments from three peer reviewers and one community reviewer. Changes include: (1) new supplementary figures situating tested compound sets within the broader PDB chemical space; (2) semi-logarithmic ROC plots added to better reflect early enrichment performance; (3) clarified interpretation of co-folding pose accuracy despite incorrect protein conformational modeling; (4) minor text and terminology corrections throughout. Core conclusions are unchanged.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00