PocketGNN: A Cross-Modal Framework Unifying Local 3D Pocket Geometry and Global Sequence Semantics for Enzyme Kinetic Prediction

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

The enzyme turnover number ( k cat ) is a pivotal kinetic parameter for understanding bio-catalytic efficiency, yet its accurate prediction remains a grand challenge due to the complex interplay between local physicochemical constraints and global evolutionary context. Existing methods typically bifurcate into sequence-based approaches, which capture evolutionary semantics but miss fine-grained spatial details, or structure-based models, which often suffer from noise in whole-protein representations or lack global context. To bridge this gap, we propose PocketGNN , a cross-modal deep learning framework that synergizes the precision of local 3D geometry with the breadth of global 1D sequence semantics. PocketGNN introduces a high-fidelity graph representation of the active pocket, enriched with a novel 24-dimensional geometric edge encoding (RBF distances, bond angles, dihedral angles) to capture the stereochemical determinants of catalysis. Crucially, this local structural view is fused with global evolutionary information extracted from pre-trained protein language models (ESM-2), creating a unified representation that spans spatial scales. Evaluated on a rigorous dataset derived from IntEnzyDB, PocketGNN achieves a Pearson correlation coefficient ( r ) of 0.98 and a coefficient of determination ( R 2 ) of 0.918 for log 10 ( k cat ) under standard random splitting. Furthermore, under a strict 40% sequence identity split designed to test zero-shot generalization to unseen families, the model maintains a robust correlation ( r = 0.67, R 2 = 0.44), significantly outperforming recent state-of-the-art methods including CatPred ( r = 0.52) and CataPro ( r = 0.50). Interpretability analysis confirms that the model successfully attends to key catalytic residues, validating its ability to learn chemically meaningful structure-function relationships rather than mere sequence memorization.
Full text 2,023 characters · extracted from oa-doi-fallback · click to expand
Abstract The enzyme turnover number (kcat) is a pivotal kinetic parameter for understanding bio-catalytic efficiency, yet its accurate prediction remains a grand challenge due to the complex interplay between local physicochemical constraints and global evolutionary context. Existing methods typically bifurcate into sequence-based approaches, which capture evolutionary semantics but miss fine-grained spatial details, or structure-based models, which often suffer from noise in whole-protein representations or lack global context. To bridge this gap, we propose PocketGNN, a cross-modal deep learning framework that synergizes the precision of local 3D geometry with the breadth of global 1D sequence semantics. PocketGNN introduces a high-fidelity graph representation of the active pocket, enriched with a novel 24-dimensional geometric edge encoding (RBF distances, bond angles, dihedral angles) to capture the stereochemical determinants of catalysis. Crucially, this local structural view is fused with global evolutionary information extracted from pre-trained protein language models (ESM-2), creating a unified representation that spans spatial scales. Evaluated on a rigorous dataset derived from IntEnzyDB, PocketGNN achieves a Pearson correlation coefficient (r) of 0.98 and a coefficient of determination (R2) of 0.918 for log10(kcat) under standard random splitting. Furthermore, under a strict 40% sequence identity split designed to test zero-shot generalization to unseen families, the model maintains a robust correlation (r = 0.67, R2 = 0.44), significantly outperforming recent state-of-the-art methods including CatPred (r = 0.52) and CataPro (r = 0.50). Interpretability analysis confirms that the model successfully attends to key catalytic residues, validating its ability to learn chemically meaningful structure-function relationships rather than mere sequence memorization. Competing Interest Statement The authors have declared no competing interest. Footnotes l-zh21{at}mails.tsinghua.edu.cn

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0