Methods
The Michaelis-Menten constant (K m) and enzyme turnover number (k cat) were
considered as the primary measures reflecting enzyme-substrate interaction kinetics for
this study. To construct the database, the K
m and k cat values of wild-type and mutant
enzyme-substrate interactions were curated from the BRENDA database (Chang et al.,
2021). The PDB structures of enzymes were mapped based on BRENDA-derived
UniProtKB annotations, wherever available. The substrate and cofactor binding sites were
mapped for each enzyme based on the bound ligand or its functional homologue. Mutant
enzymes were modelled from their wild-type structures and the protonation states of all
enzymes were corrected based on the experimental pH and temperature from BRENDA.
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
The pre-processed enzyme and substrate structures were docked to obtain the final
database of the enzyme-substrate complex structures. The complete workflow followed for
development of the database is shown below (Fig. 1).
Figure 1 : Complete workflow followed for the development of SKiD. Each step of data
processing and the external databases and tools utilized for the workflow are also
provided.
Enzyme-substrate interaction kinetics data curation
Data on experimentally measured K
m and k cat values of enzyme-substrate pairs
were collected from the BRENDA (v2023) database (Chang et al., 2021). In-house scripts
were used to process the raw data from the database and arrive at a uniform format for
database development. Redundancy in the dataset was resolved through extensive
comparisons of the annotations extracted for each datapoint including, enzyme
commission (EC) number, UniProtKB ID of the enzyme, substrate SMILES, wild-type or
mutant form of the enzyme for which the kinetics data was measured, experimental
conditions such as pH and temperature, and the cited reference for the K
m and/or k cat
values. When multiple datapoints for the same enzyme-substrate complex under same
experimental conditions were found to have differing K m and k cat values, their geometric
mean was reported in the database in accordance with previous studies (Heckmann et al.,
2018; Kroll et al., 2023; Yu et al., 2023). In case of redundant datapoints with a wide range
of values (minimum value was lower than half of the maximum value), the geometric mean
was computed post manual verification of the values from supporting literature. An outlier
analysis was also performed on the K
m and k cat values to prune datapoints with values
outside thrice the standard deviation of the log-transformed parameter distributions. These
datapoints are provided in the Supplementary Information 2.
Enzyme and substrate annotations
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
Next, enzyme and substrate annotations were performed. This is an important step
to identify their 3-dimensional structure. BRENDA can contain multiple annotations specific
to the enzyme and substrate. Datapoints were also found to have two different annotation
formats namely, comments and named columns (Supplementary Information 1 - Fig. S1).
For enzymes, the annotations were extracted from comments and other named entries in
the database through custom Python scripts. Experimental conditions and mutations in the
enzyme were only available in the comments section from which, regular expressions
were used to extract and standardize this data. The PDB IDs of the enzymes were
extracted from their UniProtKB IDs to map the enzyme structure to each datapoint. The
IUPAC names of substrates in BRENDA were extracted and annotated with their isomeric
SMILES using the OPSIN (Lowe et al., 2011) and PubChemPy libraries. Non-small
molecular substrates such as proteins, peptides, nucleic acids, polymers and peptide-
sugar conjugates were omitted from the dataset. Since several substrates had non-
standard nomenclature in BRENDA, extensive manual annotation from several standard
databases including PubChem (Kim et al., 2023), ChEMBL (Zdrazil et al., 2024), ZINC
(Tingle et al., 2023), and ChEBI (Hastings et al., 2013) was also undertaken to increase
substrate coverage of the database. Additionally, for the substrates that remained
unresolved from the previous step, SMILES were generated by drawing structures using
the GChemPaint tool (Bréfort, 2001) based on the product catalogue of laboratory
suppliers like Thermo Fisher Scientific. Finally, the three-dimensional structure of
substrates was obtained from their SMILES using the RDKit ( https://www.rdkit.org
) and
OpenBabel (O’Boyle et al., 2011) packages. Explicit hydrogens were added, and the
structures were subjected to a short energy minimization using the MMFF94 force field
(Halgren, 1996).
Mapping the structure of known enzyme-ligand complexes to the kinetics dataset
The crystallographic structures of the enzyme-substrate pairs might not be available
for majority of the kinetic parameters collected in the previous step. Various strategies
were adopted to obtain these structures. Initially, the available structural information was
collected by mapping the PDB structures based on UniProtKB annotations for each
enzyme in the dataset. These structures were classified into four categories:
substrate+cofactor structures, substrate-only structures, cofactor-only structures and apo
structures since they need to be treated differently for the mapping. Crystal additives and
ions present in the structures were not considered as substrates or cofactors during the
analysis. The substrates were differentiated from cofactors in each PDB structure using
the mapping available from the EMBL CoFactor database (Fischer et al., 2010). After the
automated discrimination of substrates from cofactors, manual verification was performed
and cases where cofactors can act as substrates were identified and corrected. Several
entries in CoFactor database were also found to be incomplete, leading to wrong
categorization of the structures. Such cases were corrected manually using the literature
associated with the structure, wherever available.
Once the PDB structures were categorized, each datapoint in the database was
associated to multiple structures from different categories. In such cases, datapoints with
at least one substrate+cofactor or substrate-only structure mapped from PDB were
considered as successful entries. Few hundreds of entries were however associated to
only apo structures or cofactor-only structures. In such cases, the structures were mapped
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
to the closest functional homologues with bound substrate using BLASTp (Altschul et al.,
1990) against the PDB database with an e-value cut-off of 0.01. If such close homologues
could not be identified, distant functional homologues with bound substrate were mapped
through PDB advanced search and manual structural alignment using backbone RMSD
cut-off of 3 Å. Only those cases were considered where at least 90% of the binding site
residues remain conserved between the parent structure and the homologue. In case of
enzymes with multiple known PDB structures, preference was assigned based on
resolution and completeness of the structure. The structure mapping obtained for all the
enzymes present in BRENDA database along with their categories are available in the
Supplementary Information 2.
Extraction of cognate substrate and cofactor binding sites from mapped structures
In accordance with the sc-PDB (Desaphy et al., 2015) convention, the substrate and
cofactor binding sites were defined as residues with at least one atom within 6.5 Å from
the substrate and cofactor atoms, respectively. For enzymes mapped to at least one
substrate+cofactor or substrate-only structures, the binding sites were directly identified
from the crystal structure using BioPython (Cock et al., 2009). For enzymes that had
cofactor-only or apo structures with close or distant functional homologues, the bound
substrate from the homologous structure was utilized to extract the binding sites. This was
achieved by alignment of the homologous structure to the parent structure using the
PyMOL API from Python, and saving the coordinates of the substrate post-alignment with
the coordinates of the parent structure. In cofactor-only structures, the cofactor from the
parent structure and the substrate from the functional homologue were used to map the
binding sites. In apo structures, the binding sites were mapped depending on the ligands
bound to the functional homologue chosen from PDB.
Mapping the mutations reported in BRENDA to the PDB structures
The dataset collected from BRENDA includes both wild-type and mutant enzymes.
However, the position of the mutated residue reported in BRENDA is predominantly based
on the UniProtKB sequence position, which does not necessarily coincide with the residue
numbering followed in PDB structures. Hence, the EMBL SIFTS database (Dana et al.,
2019) was used to derive the residue position maps between UniProtKB sequences and
PDB structures. Based on the unresolved chain-level SIFTS numbering maps, 15 error
types were identified during mutant position mapping in enzyme structures (Table 1) and
resolved using custom scripts. The distribution of these error types across BRENDA is
provided below (Fig. 2). Few of these error types and their corresponding examples from
BRENDA are discussed in the Results section. The error type mapping for all the mutant
enzymes reported in BRENDA database is also provided in the Supplementary Information
2. After mapping the mutated residue positions from BRENDA to PDB structures,
mutations were further classified into binding site and non-binding site mutations based on
the binding site definition discussed earlier.
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
Figure 2 : Percentage of occurrence of th e error types identified in BRENDA database
during mutation mapping from BRENDA to PDB structures.
Table 1: Errors encountered during mapping of mutated residue positions from BRENDA
dataset to enzyme structures from PDB. For each error type, an error cod e and the
solution type followed are also provided.
Error code Error type Solution type
1 Wrong UniProtKB ID linked in BRENDA Automated
2 Wrong mutation data curation in BRENDA Manual
3 UniProtKB ID mismatch between BRENDA and PDB Automated
4 Isoforms and INS/DEL variants in BRENDA comments Manual
5 No UniProtKB ID linked to PDB (or) no SIFTS mapping available Automated
6 Typo in mutation – regex issues in mutation extraction Automated
7 Sequence mismatch between UniProtKB and BRENDA/Reference
article Manual
8 Reverse alignment issue – UniProtKB sequence offset from PDB
sequence Automated
9 Residue numbering in non-standard fashion Manual
10 Mutated PDB structure – WT residue not matching to the UniProtKB
in reverse alignment Automated
11 Partial PDB structure – mutated residue not part of resolved
structure Automated
12 Residue numbered in sequential order – not based on author-
provided residue numbering Automated
13 Modified residue (non-standard residue or L-amino acid) used for
mutation Automated
14 Continuous residue numbering between PDB chains in multimeric
enzymes Manual
15 Unexplained Automated
se
A
e
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
Modelling mutant structures with FASPR program
While several mutant enzyme structures mapped from PDB had the exact mutation
reported in the BRENDA database, all the mutant structures were modelled from their wild-
type structures to maintain the dataset consistency. Two computational tools for mutant
structure generation from the experimental wild-type structure – FASPR (Huang et al.,
2020) and SCWRL4 (Krivov et al., 2010), were tested on a set of 10 experimentally
resolved mutant enzyme structures. The wild type PDB structures were used as the
starting conformation and mutant conformations were generated using either position
specific side chain repacking or all positions side chain repacking. The performance of
FASPR and SCWRL4 was compared using all atoms RMSD value of position specific in
silico mutated structures with their experimentally resolved structures as reference. ProFit
server was used to compute all atoms RMSD values (McLachlan, 1982), the fit was
centred around mutant site; owing to the higher flexibility of terminal residues, 10 terminal
residues from both sides of the polypeptide chain were uniformly excluded from the
calculations. Based on the comparison and another existing benchmark study (McPartlon
and Xu, 2023), FASPR was used for mutant structure generation. It was also used to
replace missing atoms and residues using the OpenMM package (Eastman et al., 2024) in
the background for repairing the enzyme structures. FASPR can perform multiple
mutations at once in the same structure, which was effectively leveraged to model the
mutant enzymes reported in BRENDA database. For enzymes without wild-type structures
in PDB, the available mutant structure was reverted to the wild-type structure by
systematically mutating using FASPR with the UniProtKB-derived sequence as the wild-
type reference.
Modification of the protonation state of enzyme structures based on experimental
data from BRENDA
The experimental data included for each enzyme-substrate interaction in BRENDA
includes the organism, pH and temperature of interest. While the UniProtKB ID is unique
to an enzyme-organism pair, the pH effect was also incorporated into the enzyme structure
through the protonation state. The PROPKA program (Olsson et al., 2011) available as
part of the PDB2PQR suite (Dolinsky et al., 2007) was used to modify the protonation state
of polar residues in the wild-type and mutant enzyme structures. Manual literature curation
was performed to retrieve pH information if the information was missing and for the
remaining cases pH 7 was used to determine the default protonation state. Since chemical
reactions involved in the biosynthesis of organic molecules is highly dependent on proton
addition and removal, this pre-processing step is invaluable for atomistic computational
studies such as docking and molecular dynamics simulations on the enzymes.
Modelling the enzyme-substrate interactions through docking
With the pre-processed enzyme and substrate structures, the enzyme-substrate
complexes were modelled through molecular docking. If the enzyme was mapped to at
least one substrate+cofactor complex in PDB, the cofactor was retained during the
docking calculations to effectively constrain the docking grid covering the substrate binding
site so that the possible substrate-cofactor interactions can be captured. To perform the
docking calculations, the GNINA (McNutt et al., 2021) and SMINA (Koes et al., 2013)
programs were initially benchmarked using the Platinum database (Pires et al., 2015),
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
which contains experimentally determined wild-type and mutant protein-ligand complexes
along with their binding affinity values. Two different experiments were performed wherein
either a single docking pose or multiple docking poses were sampled by the programs.
The similarity between the docking pose(s) and the ligand pose from the crystal structure
were compared using two measures: RMSD and ligand centroid distance. While RMSD
captures the overall agreement between the poses, the ligand centroid distance can
capture the position of the docking grid accurately, even if the orientation of the ligand is
different due to the numerous rotatable bonds present in flexible ligands. More details on
the benchmarking are provided in the Results section. Based on the benchmarking,
GNINA program with multiple pose sampling and CNN re-scoring was chosen. The
docking grids were set using the autobox_ligand option, which uses the already bound
substrate as the reference ligand for grid setting. The top-ranked docking pose out of 10
generated poses was utilized for generation of the final enzyme-substrate complex
structures.
Data Records
The SKiD database is freely accessible as a collection of spreadsheets and a
structure archive via Zenodo: https://doi.org/10.5281/zenodo.15355031
. The database
includes two spreadsheets containing the k cat and Km datasets separately, since BRENDA
contains hundreds of instances with either only the k cat value or only the K m value. The
unique enzymes and substrates present in the database are also provided with
appropriate annotations as additional spreadsheets. Each datapoint in the database is
described using 16 columns (Table 2), preserving the annotations available in the
BRENDA JSON file. Depending on the application of interest, users can filter subsets of
data from the spreadsheets along with their corresponding structure models from the
archive. Each entry in SKiD is provided a separate entry number, which is repurposed as a
folder label in the structure archive for ease of use. Each folder in the structure archive
includes the FASPR-optimized enzyme structure protonated as per the experimental pH
value, the substrate docking pose in SDF format, and a meta-data file including details of
the organism, substrate SMILES, UniProtKB ID, kinetic parameters etc. The file name in
each folder includes the PDB ID, pH value and molecule ID to facilitate comparison of
same substrate binding to multiple enzymes and enzyme-substrate complexes tested
under different pH conditions. A separate archive of numbered references is also provided
for each EC number to preserve the bibliography system used in BRENDA.
Table 2: Columns describing each datapoint in the SKiD database with a brief description
of the contents within each column.
Column name Description
Entry_ID A unique identifier provided to each data point in the SKiD database
EC_number Enzyme Commission (EC) number of the enzyme
Substrate Generic or IUPAC name of the substrate. The substrate name provided in
BRENDA is preserved to map back to the reaction easily.
UniProt_ID UniProtKB identifier of the enzyme
Protein_file Name of the enzyme structure file in the structure archive of the SKiD database
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
Organism_name Name of the parent organism (enzyme source)
kcat_value (or)
Km_value Value of the kinetic parameter of interest (kcat or Km)
Mutant A binary identifier to discriminate mutant enzymes from wild-type enzymes
Mutation The mutation(s) present in the enzyme as reported in the BRENDA database.
Multiple mutations are separated by a forward slash (/) symbol.
pH pH of the reaction as reported in BRENDA database or supporting literature
Temperature Temperature of the reaction (in °C) as reported in BRENDA database or
supporting literature
References
Chang, A.; Jeske, L.; Ulbrich, S.; Hofmann, J.; Koblitz, J.; Schomburg, I.; Neumann-Schaal, M.; Jahn,
D.; Schomburg, D. BRENDA, the ELIXIR core data resource in 2021: new developments and updates.
Nucleic Acids Res. 2021, 49(D1), D498-D508. doi: 10.1093/nar/gkaa1025
Heckmann, D.; Lloyd, C. J.; Mih, N.; Ha, Y.; Zielinski, D. C.; Haiman, Z. B.; Desouki, A. A.; Lercher, M.
J.; Palsson, B. O. Machine learning applied to enzyme turnover numbers reveals protein structural correlates
and improves metabolic models. Nat. Commun. 2018, 9(1), 5252. doi: 10.1038/s41467-018-07652-6
Kroll, A.; Rousset, Y .; Hu, X. P.; Liebrand, N. A.; Lercher, M. J. Turnover number predictions for
kinetically uncharacterized enzymes using machine and deep learning. Nat. Commun. 2023, 14(1), 4139.
doi: 10.1038/s41467-023-39840-4
Yu, H.; Deng, H.; He, J.; Keasling, J. D.; Luo, X. UniKP: a unified framework for the prediction of
enzyme kinetic parameters. Nat. Commun. 2023, 14(1), 8211. doi: 10.1038/s41467-023-44113-1
Lowe, D. M.; Corbett, P . T.; Murray-Rust, P.; Glen, R. C. Chemical name to structure: OPSIN, an open
source solution. J. Chem. Inf. Model. 2011, 51(3), 739-53. doi: 10.1021/ci100384d
Kim, S.; Chen, J.; Cheng, T.; Gindulyte, A.; He, J.; He, S.; Li, Q.; Shoemaker, B. A.; Thiessen, P. A.;
Yu, B.; Zaslavsky, L.; Zhang, J.; Bolton, E. E. PubChem 2023 update. Nucleic Acids Res . 2023, 51(D1),
D1373-D1380. doi: 10.1093/nar/gkac956
Zdrazil, B.; Felix, E.; Hunter, F .; Manners, E. J.; Blackshaw, J.; Corbett, S.; de Veij, M.; Ioannidis, H.;
Lopez, D. M.; Mosquera, J. F.; Magarinos, M. P .; Bosc, N.; Arcila, R.; Kizilören, T.; Gaulton, A.; Bento, A. P .;
Adasme, M. F.; Monecke, P.; Landrum, G. A.; Leach, A. R. The ChEMBL Database in 2023: a drug discovery
platform spanning multiple bioactivity data types and time periods. Nucleic Acids Res. 2024, 52(D1), D1180-
D1192. doi: 10.1093/nar/gkad1004
Tingle, B. I.; Tang, K. G.; Castanon, M.; Gutierrez, J. J.; Khurelbaatar, M.; Dandarchuluun, C.; Moroz,
Y . S.; Irwin, J. J. ZINC-22─ A Free Multi-Billion-Scale Database of Tangible Compounds for Ligand Discovery.
J. Chem. Inf. Model. 2023, 63(4), 1166-1176. doi: 10.1021/acs.jcim.2c01253
Hastings, J.; de Matos, P .; Dekker, A.; Ennis, M.; Harsha, B.; Kale, N.; Muthukrishnan, V.; Owen, G.;
Turner, S.; Williams, M.; Steinbeck, C. The ChEBI reference database and ontology for biologically relevant
chemistry: enhancements for 2013. Nucleic Acids Res . 2013, 41(Database issue), D456-63. doi:
10.1093/nar/gks1146
RDKit: Open-source cheminformatics. https://www.rdkit.org
O'Boyle, N. M.; Banck, M.; James, C. A.; Morley, C.; Vandermeersch, T.; Hutchison, G. R. Open Babel:
An open chemical toolbox. J. Cheminform. 2011, 3, 33. doi: 10.1186/1758-2946-3-33
Fischer, J. D.; Holliday, G. L.; Thornton, J. M. The CoFactor database: organic cofactors in enzyme
catalysis. Bioinformatics. 2010, 26(19), 2496-7. doi: 10.1093/bioinformatics/btq442
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
Altschul, S. F.; Gish, W.; Miller, W.; Myers, E. W.; Lipman, D. J. Basic local alignment search tool. J.
Mol. Biol. 1990, 215(3), 403-10. doi: 10.1016/S0022-2836(05)80360-2
Desaphy, J.; Bret, G.; Rognan, D.; Kellenberger, E. sc-PDB: a 3D-database of ligandable binding
sites--10 years on. Nucleic Acids Res. 2015, 43(Database issue), D399-404. doi: 10.1093/nar/gku928
Cock, P . J.; Antao, T.; Chang, J. T.; Chapman, B. A.; Cox, C. J.; Dalke, A.; Friedberg, I.; Hamelryck, T.;
Kauff, F.; Wilczynski, B.; de Hoon, M. J. Biopython: freely available Python tools for computational molecular
biology and bioinformatics. Bioinformatics. 2009, 25(11), 1422-3. doi: 10.1093/bioinformatics/btp163
Dana, J. M.; Gutmanas, A.; Tyagi, N.; Qi, G.; O'Donovan, C.; Martin, M.; Velankar, S. SIFTS: updated
Structure Integration with Function, Taxonomy and Sequences resource allows 40-fold increase in coverage
of structure-based annotations for proteins. Nucleic Acids Res . 2019, 47(D1), D482-D489. doi:
10.1093/nar/gky1114
Huang, X.; Pearce, R.; Zhang, Y. FASPR: an open-source tool for fast and accurate protein side-chain
packing. Bioinformatics. 2020, 36(12), 3758-3765. doi: 10.1093/bioinformatics/btaa234
Krivov, G. G.; Shapovalov, M. V.; Dunbrack, R. L. Jr. Improved prediction of protein side-chain
conformations with SCWRL4. Proteins. 2009, 77(4), 778-95. doi: 10.1002/prot.22488
Eastman, P .; Galvelis, R.; Peláez, R. P .; Abreu, C. R. A.; Farr, S. E.; Gallicchio, E.; Gorenko, A.; Henry,
M. M.; Hu, F.; Huang, J.; Krämer, A.; Michel, J.; Mitchell, J. A.; Pande, V. S.; Rodrigues, J. P.; Rodriguez-
Guerra, J.; Simmonett, A. C.; Singh, S.; Swails, J.; Turner, P.; Wang, Y.; Zhang, I.; Chodera, J. D.; De
Fabritiis, G.; Markland, T. E. OpenMM 8: Molecular Dynamics Simulation with Machine Learning Potentials.
J. Phys. Chem. B. 2024, 128(1), 109-116. doi: 10.1021/acs.jpcb.3c06662
Olsson, M. H.; Søndergaard, C. R.; Rostkowski, M.; Jensen, J. H. PROPKA3: Consistent Treatment of
Internal and Surface Residues in Empirical pKa Predictions. J. Chem. Theory Comput . 2011, 7(2), 525-37.
doi: 10.1021/ct100578z
Dolinsky, T. J.; Czodrowski, P.; Li, H.; Nielsen, J. E.; Jensen, J. H.; Klebe, G.; Baker, N. A. PDB2PQR:
expanding and upgrading automated preparation of biomolecular structures for molecular simulations.
Nucleic Acids Res. 2007, 35(Web Server issue), W522-5. doi: 10.1093/nar/gkm276
McNutt, A. T.; Francoeur, P .; Aggarwal, R.; Masuda, T.; Meli, R.; Ragoza, M.; Sunseri, J.; Koes, D. R.
GNINA 1.0: molecular docking with deep learning. J. Cheminform. 2021, 13(1), 43. doi: 10.1186/s13321-021-
00522-2
Koes, D. R.; Baumgartner, M. P.; Camacho, C. J. Lessons learned in empirical scoring with smina from
the CSAR 2011 benchmarking exercise. J. Chem. Inf. Model. 2013, 53(8), 1893-904. doi: 10.1021/ci300604z
Pires, D. E.; Blundell, T. L.; Ascher, D. B. Platinum: a database of experimentally measured effects of
mutations on structurally defined protein-ligand complexes. Nucleic Acids Res. 2015, 43(Database issue),
D387-91. doi: 10.1093/nar/gku966
McNutt, A. T.; Koes, D. R. Open-ComBind: harnessing unlabeled data for improved binding pose
prediction. J. Comput. Aided Mol. Des. 2023, 38(1), 3. doi: 10.1007/s10822-023-00544-y
Abramson, J.; Adler, J.; Dunger, J.; Evans, R.; Green, T.; Pritzel, A. et al. Accurate structure prediction
of biomolecular interactions with AlphaFold 3. Nature. 2024, 630(8016), 493-500. doi: 10.1038/s41586-024-
07487-w
Yasumitsu, Y.; Ohue, M. Generation of appropriate protein structures for virtual screening using
AlphaFold3 predicted protein–ligand complexes. BioRxiv. 2025. doi: 10.1101/2025.02.17.638750
McLachlan, A. D. Rapid comparison of protein structures. Acta Cryst. 1982, A38, 871-873. doi:
10.1107/S0567739482001806
McPartlon, M.; Xu, J. An end-to-end deep learning method for protein side-chain packing and inverse
folding. Proc. Natl. Acad. Sci. U. S. A. 2023, 120(23), e2216438120. doi: 10.1073/pnas.2216438120
Radley, E.; Davidson, J.; Foster, J.; Obexer, R.; Bell, E. L.; Green, A. P . Engineering Enzymes for
Environmental Sustainability. Angew. Chem. Int. Ed. Engl . 2023, 62(52), e202309305. doi:
10.1002/anie.202309305
Paul, C.; Hanefeld, U.; Hollmann, F.; Qu, G.; Yuan, B.; Sun, Z. Enzyme engineering for biocatalysis. J.
Mol. Catal. 2024, 555, 113874. doi: 10.1016/j.mcat.2024.113874
Sheldon, R. A.; Brady, D.; Bode, M. L. The Hitchhiker's guide to biocatalysis: recent advances in the
use of enzymes in organic synthesis. Chem. Sci. 2020, 11(10), 2587-2605. doi: 10.1039/c9sc05746c
Siedentop, R.; Rosenthal, K. Industrially Relevant Enzyme Cascades for Drug Synthesis and Their
Ecological Assessment. Int. J. Mol. Sci. 2022, 23(7), 3605. doi: 10.3390/ijms23073605
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
Kroll, A.; Ranjan, S.; Engqvist, M. K. M.; Lercher, M. J. A general model to predict small molecule
substrates of enzymes based on machine and deep learning. Nat. Commun . 2023, 14(1), 2787. doi:
10.1038/s41467-023-38347-2
Wang, X.; Quinn, D.; Moody, T. S.; Huang, M. ALDELE: All-Purpose Deep Learning Toolkits for
Predicting the Biocatalytic Activities of Enzymes. J. Chem. Inf. Model . 2024, 64(8), 3123-3139. doi:
10.1021/acs.jcim.4c00058
Thakur, D.; Pandit, S. B. Unusual commonality in active site structural features of substrate
promiscuous and specialist enzymes. J. Struct. Biol. 2022, 214(1), 107835. doi: 10.1016/j.jsb.2022.107835
Li, F.; Yuan, L.; Lu, H.; Li, G.; Chen, Y .; Engqvist, M. K. M. et al. Deep learning-based kcat prediction
enables improved enzyme-constrained model reconstruction. Nat. Catal . 2022, 5, 662–672. doi:
10.1038/s41929-022-00798-z
Yu, H.; Deng, H.; He, J.; Keasling, J. D.; Luo, X. UniKP: a unified framework for the prediction of
enzyme kinetic parameters. Nat. Commun. 2023, 14(1), 8211. doi: 10.1038/s41467-023-44113-1
Kroll, A.; Rousset, Y .; Hu, X. P.; Liebrand, N. A.; Lercher, M. J. Turnover number predictions for
kinetically uncharacterized enzymes using machine and deep learning. Nat. Commun. 2023, 14(1), 4139.
doi: 10.1038/s41467-023-39840-4
Gollub, M. G.; Backes, T.; Kaltenbach, H. M.; Stelling, J. ENKIE: a package for predicting enzyme
kinetic parameter values and their uncertainties. Bioinformatics. 2024, 40(11), btae652. doi:
10.1093/bioinformatics/btae652
Wang, J.; Yang, Z.; Chen, C.; Yao, G.; Wan, X.; Bao, S. et al. MPEK: a multitask deep learning
framework based on pretrained language models for enzymatic reaction kinetic parameters prediction. Brief.
Bioinform. 2024, 25(5), bbae387. doi: 10.1093/bib/bbae387
Wang, T.; Xiang, G.; He, S.; Su, L.; Wang, Y .; Yan, X.; Lu, H. DeepEnzyme: a robust deep learning
model for improved enzyme turnover number prediction by utilizing features of protein 3D-structures. Brief.
Bioinform. 2024, 25(5), bbae409. doi: 10.1093/bib/bbae409
Cai, Y.; Zhang, W.; Dou, Z.; Wang, C.; Yu, W.; Wang, L. PreTKcat: A pre-trained representation
learning and machine learning framework for predicting enzyme turnover number. Comput. Biol. Chem .
2025, 115, 108327. doi: 10.1016/j.compbiolchem.2024.108327
Chen, Y.; Gustafsson, J.; Rangel, A. T.; Anton, M.; Domenzain, I.; Kittikunapong, C.; Li, F. et al .
Reconstruction, simulation and analysis of enzyme-constrained metabolic models using GECKO Toolbox
3.0. Nat. Protoc. 2024, 19(3), 629-667. doi: 10.1038/s41596-023-00931-7
Pochapsky, T. C.; Pochapsky, S. S. What Your Crystal Structure Will Not Tell You about Enzyme
Function. Acc. Chem. Res. 2019, 52(5), 1409-1418. doi: 10.1021/acs.accounts.9b00066
Kingsley, L. J.; Lill, M. A. Substrate tunnels in enzymes: structure-function relationships and
computational methodology. Proteins. 2015, 83(4), 599-611. doi: 10.1002/prot.24772
Zhou, J.; Zhuang, Y .; Xia, J. Integration of enzyme constraints in a genome-scale metabolic model of
Aspergillus niger improves phenotype predictions. Microb. Cell Fact. 2021, 20(1), 125. doi: 10.1186/s12934-
021-01614-2
Savile, C. K.; Janey, J. M.; Mundorff, E. C.; Moore, J. C.; Tam, S.; Jarvis, W. R. et al . Biocatalytic
as
ymmetric synthesis of chiral amines from ketones applied to sitagliptin manufacture. Science. 2010,
329(5989), 305-9. doi: 10.1126/science.1188934
Wittig, U.; Rey, M.; Weidemann, A.; Kania, R.; Müller, W. SABIO-RK: an updated resource for
manually curated biochemical reaction kinetics. Nucleic Acids Res . 2018, 46(D1), D656-D660. doi:
10.1093/nar/gkx1065
Placzek, S.; Schomburg, I.; Chang, A.; Jeske, L.; Ulbrich, M.; Tillack, J.; Schomburg, D. BRENDA in
2017: new perspectives and new tools in BRENDA. Nucleic Acids Res . 2017, 45(D1), D380-D388. doi:
10.1093/nar/gkw952
Jeske, L.; Placzek, S.; Schomburg, I.; Chang, A.; Schomburg, D. BRENDA in 2019: a European
ELIXIR core data resource. Nucleic Acids Res. 2019, 47(D1), D542-D549. doi: 10.1093/nar/gky1048
Tipton, K. F.; Armstrong, R. N.; Bakker, B. M.; Bairoch, A.; Cornish-Bowden, A.; Halling, P. J. et al.
Standards for Reporting Enzyme Data: The STRENDA Consortium: What it aims to do and why it should be
helpful. Perspect. Sci. 2014, 1(1-6), 131-137. doi: 10.1016/j.pisc.2014.02.012
Swainston, N.; Baici, A.; Bakker, B. M.; Cornish-Bowden, A.; Fitzpatrick, P. F .; Halling, P . et al .
STRENDA DB: enabling the validation and sharing of enzyme kinetics data. FEBS J. 2018, 285(12), 2193-
2204. doi: 10.1111/febs.14427
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
Wang, C. Y .; Chang, P . M.; Ary, M. L.; Allen, B. D.; Chica, R. A.; Mayo, S. L.; Olafson, B. D. ProtaBank:
A repository for protein design and engineering data. Protein Sci . 2018, 27(6), 1113-1124. doi:
10.1002/pro.3406
Li, F.; Chen, Y.; Anton, M.; Nielsen, J. GotEnzymes: an extensive database of enzyme parameter
predictions. Nucleic Acids Res. 2023, 51(D1), D583-D586. doi: 10.1093/nar/gkac831
Li, F.; Yuan, L.; Lu, H.; Li, G.; Chen, Y .; Engqvist, M. K. M. et al. Deep learning-based k cat prediction
enables improved enzyme-constrained model reconstruction. Nat. Catal. 2022, 5, 662–672. doi:
10.1038/s41929-022-00798-z
Kroll, A.; Lercher, M. J. DLKcat cannot predict meaningful k cat values for mutants and unfamiliar
enzymes. Biol. Methods Protoc. 2024, 9(1), bpae061. doi: 10.1093/biomethods/bpae061
Kuznetsova, E.; Proudfoot, M.; Gonzalez, C. F.; Brown, G.; Omelchenko, M. V.; Borozan, I. et al .
Genome-wide analysis of substrate specificities of the Escherichia coli haloacid dehalogenase-like
phosphatase family. J. Biol. Chem. 2006, 281(47), 36149-61. doi: 10.1074/jbc.M605449200
Burroughs, A. M.; Allen, K. N.; Dunaway-Mariano, D.; Aravind, L. Evolutionary genomics of the HAD
superfamily: understanding the structural adaptations and catalytic diversity in a superfamily of
phosphoesterases and allied enzymes. J. Mol. Biol. 2006, 361(5), 1003-34. doi: 10.1016/j.jmb.2006.06.049
Qian, H.; Wang, Y .; Zhou, X.; Gu, T.; Wang, H.; Lyu, H. et al. ESM-Ezy: a deep learning strategy for the
mining of novel multicopper oxidases with superior properties. Nat. Commun . 2025, 16(1), 3274. doi:
10.1038/s41467-025-58521-y
Yan, B.; Ran, X.; Gollu, A.; Cheng, Z.; Zhou, X.; Chen, Y.; Yang, Z. J. IntEnzyDB: an Integrated
Structure-Kinetics Enzymology Database. J. Chem. Inf. Model. 2022, 62(22), 5841-5848. doi:
10.1021/acs.jcim.2c0113
Michaelis, L.; Menten, M. L.; Johnson, K. A.; Goody, R. S. The original Michaelis constant: translation
of the 1913 Michaelis-Menten paper. Biochemistry. 2011, 50(39), 8264-9. doi: 10.1021/bi201284u
Koshland, D. E. The application and usefulness of the ratio k(cat)/K(M). Bioorg. Chem. 2002, 30(3),
211-3. doi: 10.1006/bioo.2002.1246
Sheldon, R. A.; Woodley, J. M. Role of Biocatalysis in Sustainable Chemistry. Chem. Rev . 2018,
118(2), 801-838. doi: 10.1021/acs.chemrev.7b00203
Č esnik, M.; Sudar, M.; Hernández, K.; Charnock, S.; Vasi ć -Rač ki, Đ .; Clapés, P.; Blaževi ć , Z. F .
Cascade enzymatic synthesis of l-homoserine – mathematical modelling as a tool for process optimisation
and design. React. Chem. Eng. 2020, 5, 747-759 doi: 10.1039/C9RE00453J
Chang, A.; Jeske, L.; Ulbrich, S.; Hofmann, J.; Koblitz, J.; Schomburg, I. et al. BRENDA, the ELIXIR
core data resource in 2021: new developments and updates. Nucleic Acids Res. 2021, 49(D1), D498-D508.
doi: 10.1093/nar/gkaa1025
Caspi, R.; Billington, R.; Fulcher, C. A.; Keseler, I. M.; Kothari, A.; Krummenacker, M. et al . The
MetaCyc database of metabolic pathways and enzymes. Nucleic Acids Res. 2018, 46(D1), D633-D639. doi:
10.1093/nar/gkx935
Wittig, U.; Kania, R.; Golebiewski, M.; Rey, M.; Shi, L.; Jong, L. et al . SABIO-RK--database for
biochemical reaction kinetics. Nucleic Acids Res. 2012, 40(Database issue), D790-6. doi:
10.1093/nar/gkr1046
Perona, J. J.; Craik, C. S. Evolutionary divergence of substrate specificity within the chymotrypsin-like
serine protease fold. J. Biol. Chem. 1997, 272(48), 29987-90. doi: 10.1074/jbc.272.48.29987
Borkakoti, N.; Ribeiro, A. J. M.; Thornton, J. M. A structural perspective on enzymes and their catalytic
mechanisms. Curr. Opin. Struct. Biol. 2025, 92, 103040. doi: 10.1016/j.sbi.2025.103040
Sigrist, C. J.; De Castro, E.; Langendijk-Genevaux, P . S.; Le Saux, V.; Bairoch, A.; Hulo, N. ProRule: a
new database containing functional and structural information on PROSITE profiles. Bioinformatics. 2005,
21(21), 4060-6. doi: 10.1093/bioinformatics/bti614
Bréfort, J. (2001). GChemPaint - 2D chemical structure editor. Available from
https://www.nongnu.org/gchempaint/
Halgren, T. A. Merck molecular force field. I. Basis, form, scope, parameterization, and performance of
MMFF94. J. Comp. Chem. 1996, 17(5-6), 490-519. doi: 10.1002/(SICI)1096-987X(199604)17:5/63.0.CO;2-P
Author Contributions
.CC-BY-NC-ND 4.0 International licenseavailable under a
was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprint (whichthis version posted May 24, 2025. ; https://doi.org/10.1101/2025.05.18.654770doi: bioRxiv preprint
Sowmya Ramaswamy Krishnan : Data Curation, Validation, Formal analysis, Software,
Writing – Original Draft, Writing – Review & Editing, Visualization
Nishtha Pandey: Data Curation, Validation, Formal analysis, Software, Writing – Original
Draft, Writing – Review & Editing
Rajgopal Srinivasan: Methodology, Validation, Writing – Review & Editing, Supervision
Arijit Roy: Conceptualization, Methodology, Validation, Investigation, Resources, Writing –
Review & Editing, Supervision
Competing Interests
All authors are employed by Tata Consultancy Services Limited.