Aldo-keto reductase (AKR) superfamily website and database: An update.

OA: closed
AI-generated summary by claude@2026-07, 2026-07-28

This paper presents an updated AKR superfamily website and database, incorporating genetic, functional, structural, and bioinformatics data for over 200 confirmed members.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

The aldo-keto reductase (AKR) superfamily is a large family of proteins found across the kingdoms of life. Shared features of the family include 1) structural similarities such as an (α/β)8-barrel structure, disordered loop structure, cofactor binding site, and a catalytic tetrad, and 2) the ability to catalyze the nicotinamide adenine dinucleotide (phosphate) reduced (NAD(P)H)-dependent reduction of a carbonyl group. A criteria of family membership is that the protein must have a measured function, and thus, genomic sequences suggesting the transcription of potential AKR proteins are considered pseudo-members until evidence of a functionally expressed protein is available. Currently, over 200 confirmed AKR superfamily members are reported to exist. A systematic nomenclature for the AKR superfamily exists to facilitate family and subfamily designations of the member to be communicated easily. Specifically, protein names include the root "AKR", followed by the family represented by an Arabic number, the subfamily-if one exists-represented by a letter, and finally, the individual member represented by an Arabic number. The AKR superfamily database has been dedicated to tracking and reporting the current knowledge of the AKRs since 1997, and the website was last updated in 2003. Here, we present an updated version of the website and database that were released in 2023. The database contains genetic, functional, and structural data drawn from various sources, while the website provides alignment information and family tree structure derived from bioinformatics analyses.
Full text 19,377 characters · extracted from pmc-nxml · 4 sections · click to expand

Results

The website is organized through five major tabs: ‘About AKR’, ‘AKR Members’, ‘Phylogenies’, ‘Multiple Sequence Alignment’, and ‘Submit AKR Sequences’. Figure 1 provides an overview of the website organization and a screenshot of its homepage. The about AKR tab has a section on nomenclature, protein structure and function, families. The AKR superfamily naming conventions follow a specific nomenclature adopted by the 8 th international Workshop on Enzymology and Molecular Biology of Carbonyl Metabolism in 1997 [ 63 ]. All members of the superfamily start with the root “AKR,” followed by an Arabic number denoting the family. In families where a subfamily is present, the subfamily is represented by a letter. Lastly, the final Arabic number denotes the enzyme identity, which is numbered chronologically in order of discovery. An example of this nomenclature can be seen in Figure 2 . For genes encoding the AKR proteins, the gene name is identical to the protein name, but denoted in italics. Each AKR protein has its own unique name to avoid misclassifying proteins from other species as orthologs when this is not known with certainty. The use of lower case AKR names italicized or unitalicized to denote genes and proteins, respectively, is discouraged since this assumes that the homology that exists predicts conservation of function. Families are defined by its members having 40% amino acid sequence identity, meaning two members of the family should be at least 40% identical. Currently, there are 17 AKR families according to phylogenetic tree analysis. Within families exist subfamilies that are defined by their members having 60% to 97% amino acid sequence identity [ 63 ]. Members of the same subfamily with over 97% amino acid sequence identity are considered alleles of the same gene unless they have distinct activities, are encoded by different mRNA transcripts, and come from structurally different genes. An example of this exception is seen with AKR1C1 and AKR1C2, two AKR 1C members with 96% sequence identity which differ by only seven amino acids but are still considered two different members since they are coded by different genes and have distinct functions [ 64 – 67 ]. Most AKRs exist as monomeric proteins, however, some members of the superfamily have been observed to form multimers. In this case, the naming should contain the composition of proteins, and stoichiometry. A tetramer containing one AKR7A1 monomer and three AKR7A4 monomers should be denoted as AKR7A1: AKR7A4 (1:3) [ 1 ]. We note that with the updated AKR database, the nomenclature of some members appears out of sequence, and that is largely due to the fact that as new members are discovered, some of the relationships among existing members have changed. Previously named AKR proteins’ names have been kept to preserve consistency with the published record. AKRs function as phase I enzymes, catalyzing the carbonyl reduction on a variety of endogenous and xenobiotic substrates. Thus, aldehydes are reduced to primary alcohols and ketones are reduced to secondary alcohols. The alcohol functional group is then available for conjugation reactions so that the reactive carbonyl containing compound can be eliminated. All AKRs catalyze a sequential ordered bi bi reaction in which the cofactor binds first and leaves last [ 68 , 69 ]. AKRs catalyze the nicotinamide adenine dinucleotide (phosphate) reduced (NAD(P)H)-dependent reduction of carbonyl groups and the reverse oxidation reaction, thus classifying AKRs as oxidoreductases [ 3 ]. However, in vivo , these enzymes act as reductases due to their high affinity for NADPH and favorable K eg . [ 70 , 71 ]. Due to the similar reaction catalyzed, members of the AKR superfamily possess structural similarities [ 72 ]. These features include an (α/β) 8 -barrel fold ( Figure 3.A ), also known as a triose-phosphate isomerase TIM barrel with two additional helices and loop structures at the back of the barrel. A novel NADP(H)-binding motif is located in the elliptical pocket at the C-terminal end of the β-sheet [ 68 , 72 – 74 ]. Interestingly, AKRs have stereoselectivity for 4-pro-R hydride transfer from NADPH to the acceptor carbonyl group [ 68 ]. The animo acids that bind NADPH are highly conserved [ 5 , 68 , 75 ] (T24, D50, S166, N167, Q190, Y216, L219, S221, R270, S271, F272, R276, E279, and N280 in AKR1C9 numbering) ( Figure 4 ). Interaction with S166, N167 and D50 ensure that the carboxamide side-chain is tethered so that the nicotinamide head group is in the anti -configuration and the nicotinamide ring pi-stacks with Y216. To define substrate specificity, three large loops exist behind the barrel motif [ 4 ] depicted in Figure 3.B . The binding of the NADPH coenzyme causes a conformational change that reorients the loops [ 76 ]. Tight binding of the cofactor is due to anchoring the 2’ phosphate of AMP by R276 or equivalent residue [ 77 ]. In some AKRs, the tight binding is enhanced by a clamping loop which acts as a safety belt across the pyrophosphate bridge of the cofactor. The carbonyl-containing substrate binds perpendicularly to the cofactor. The preference for NADP(H) over NAD(H), can be reversed if R276 is replaced with an acidic group to repel the 2’-phopsphate of AMP [ 78 ]. Another structural motif that exists in the AKRs is the conserved catalytic tetrad that includes the residues Y55, L84, H117, and D50 [ 5 ] (AKR1C9 numbering convention) which catalyze a push-pull mechanism for hydride transfer [ 79 ]. In some AKRs H117 is replaced by a glutamic acid to increase the acidity of the active site to promote carbonyl group enolization as seen in AKR1D1 [ 80 ]. In general, AKRs have a molecular mass between 34 to 37 kDa and are monomeric and soluble [ 1 , 5 , 68 ]. The AKR superfamily has 17 separate families, denoted by their first Arabic number ( Table 1 ). The AKR members tab lists existing and potential members that are grouped by PDB structures. There are 207 AKR member entries in our current database. Each entry contains the nomenclature name, National Center of Biotechnology Information (NCBI) accession number, species expressed, enzyme name, link to protein database (PDB) entry, and alternative splicing transcripts. The NCBI accession number is an identifier for the protein in the GenBank database and links to the protein’s NCBI entry, containing alternative names, FASTA file, source organisms, amino acid sequence, and references [ 81 ]. The species expressed column reports the species where the protein was found to be present. The enzyme name column reports the type of protein (e.g., oxidoreductase, dehydrogenase) and common or alternative names for the protein. The PBD column contains a link to the structure of the AKR member in the PDB database when available. Information contained in the PBD includes the 3D structure, the depositing authors, the expression system, and experimental data and validation of the structure [ 82 ]. When multiple structures are available in the PDB, the chronologically first published structure is used. The other structures are available through the Grouped by PDB Structure section. The alt splicing column links to the Ensembl database for the AKR member. Ensembl contains genetic information, including the summary of the gene name, location, and its transcripts [ 83 ]. Physiological functions of human AKR members are described in Table 2 . To have full status as an AKR superfamily member, the member must be a functional protein associated with the gene. If the member is derived from a partial cDNA sequence or genomics project then it is not included in the database, but some are listed as potential members to be included pending functional analyses. There are currently 34 potential members listed in the database. Each entry contains species expressed , name , and NCBI accession number. Unlike the existing members, potential members are grouped together by the species in which they were detected. The actual number of potential AKR members is much larger than the ones submitted to the AKR database as evidenced by the large number of AKR genetic sequences observed across all species, many of which do not have a known function. Multiple PDB entries can exist for a single AKR member, often because structures reflect diverse conformational states that occur under differing conditions or upon binding to different ligands. The purpose of the Grouped by PDB Structure section is to provide all available PBD structures and provide relevant comments regarding the structure. Our table contains the AKR family member name, the taxonomy, resolution, complex, and PDB link. In the AKR column, the AKR nomenclature name is reported alongside any common names. If multiple protein names exist, all are listed. The taxonomy column lists the species in which the AKR protein or complex was purified from to obtain the structure. The column Res. (Å) lists the resolution of the structure reported by the PDB entry. In simple terms, the resolution of a protein structure is the distance of the smallest observable feature, thus a smaller value indicates higher resolution. Resolutions are reported in angstroms (Å), equal to 10 −10 meters [ 84 ]. The complex column provides a description of the contents of the structure in more detail. Finally, the PBD column is a direct link to the PBD entry. Most structures contain a bound cofactor, and apoenzyme structures are scarcely available, likely due to the intrinsic disorder of the loops when NAD(P)H is not bound [ 73 , 74 , 85 ]. The phylogeny tab provides information on the evolutionary relationships among AKRs. The AKR superfamily is thought to be a product of divergent evolution due to its members having common 3D structures and a highly conserved NADPH binding pocket to accomodate diverse substrates. There is also evidence of convergent evolution because AKRs are distinct from other oxidoreductase superfamilies such as long chain alcohol dehydrogenase and short chain dehydrogenase/reductase [ 5 ]. That is, genetic alignment studies have not found significant similarities between these oxidoreductase superfamilies [ 3 , 86 ]. Furthermore, the presence of AKRs across diverse life domains suggests that AKRs are an ancient superfamily [ 1 ]. To examine evolutionary relationships within the AKR superfamily, phylogenic trees were created using the multialign program, an update to the previous versions constructed with the GCG program [ 1 ]. AKR phylogeny dendrograms provide an overview of the entire AKR superfamily ( Figure 5.A ), each AKR family with at least three members ( Figure 5.B contains an example), and for each of the following taxonomic groups: Animalia, Bacteria, Fungi, Plantae, Insecta, Mammalia, Langomorpha, Rodentia , and Homo sapiens . Multiple sequence alignment (MSA) is a bioinformatics technique used to align biological sequences, such as DNA, RNA, or amino acid sequences, in order to compare similarities and differences. Visualization of such data is paramount to MSA. We provide users the ability to visualize aligned AKR protein sequences from various families and taxonomic groups in our database. The alignments are generated using MSAViewer, offering an interactive JavaScript-based representation of multiple sequence alignment [ 55 ]. To use the tool, first the group of AKRs a user would like to compare are selected. The groups include all, by family, and by taxonomic group, consistent with the phylogenies available. All families are present regardless of the number of members. For families with multiple subfamilies, each entry is listed alphabetically and numerically (e.g., AKR1C2 is listed before AKR1C3, and both are listed before AKR1D1). The default visualization can be adjusted according to options for Importing, Sorting, Filter, Selection, Visual elements, Color scheme, Extras, Exporting, and More as listed in Table 3 . As an example, the alignment of the catalytic tetrad residues for AKR1C1 members 1–35 using the MSAviewer available on the website are depicted in Figure 6 . Sharing the discovery of new members to the AKR superfamily is highly encouraged. To facilitate this process, a section of the website is dedicated to this activity. Submitted AKRs should have a functionally expressed protein and an amino acid sequence determined by cDNA or other direct methods. The protein must be purified or overexpressed from either its natural source or recombinantly. Mutant AKRs are not represented in this database. The submission should include the following information: Protein sequence from cDNA or direct methods Trivial name Species of origin Expression system Enzyme activity substrate Accession number (GenBank, Swiss-Prot, PIR) Status of publication Citations Contact information of submitter Protein sequence from cDNA or direct methods Trivial name Species of origin Expression system Enzyme activity substrate Accession number (GenBank, Swiss-Prot, PIR) Status of publication Citations Contact information of submitter Once submitted, the proposed member will be matched against the current AKRs and placed within the evolutionary tree. The location within the superfamily cluster will determine the nomenclature designation. If needed, new families and subfamilies will be created. Upon completion of the cluster analysis, the assigned designation and location in the AKR superfamily are communicated to the submitter. The new AKR will be made available on the website once the submission has been published.

Material

MAFFT [ 52 ] was used to perform multiple sequence alignment. Protein sequences of the AKR superfamily were aligned via the L-INS-i algorithm, an iterative refinement method that employs a local pairwise alignment with the affine gap cost [ 53 ]. The aligned sequences were visualized using the msaR R package [ 54 ] that provides an interface to MSAViewer [ 55 ] for web visualization. Percent identity, which measures the number of matches in relation to the length of the alignment, was calculated using the seqinr R package [ 56 ]. Gapped positions were excluded from the identity calculations. Identity measures were used to delineate AKR subfamilies. IQ-TREE [ 57 ] was used to infer maximum-likelihood phylogenies, incorporating ModelFinder [ 58 ] to improve the accuracy of phylogenetic estimates by identifying the best-fitting model of sequence evolution. The resulting phylogenies were visualized using the ggtree R package [ 59 ]. The revamped AKR website uses the Shiny web framework [ 60 ] to enable real-time user interaction with data, including filtering, selection, and manipulation of tables and visuals. The website’s tables are rendered using the DT R package [ 61 ] that leverages the JavaScript DataTables library to deliver a responsive user experience. A new sequence submission form was created with the help of the shinyjs R package [ 62 ] that provides common JavaScript operations within Shiny web applications.

Discussion

The AKR superfamily is an extensive family with much interest regarding its over 200 members. The superfamily benefits from having a centralized portal of information, a function that the website has performed for over 20 years. Since this time, the scientific research landscape has undergone many changes, and the new website has been updated accordingly. The updated cluster analysis to designate the AKR families and subfamilies, resulted in some families needing to be restructured by the analysis designation. However, our analysis is also regularly updated and upon the discovery of other members, the families might undergo other changes, representing the most updated knowledge. A new feature of the current AKR superfamily website includes the retention of past iterations, allowing for outdated information to be clearly documented. The AKR superfamily database also links to other databases such as GenBank, Ensembl, and PDB, facilitating access to carefully curated and up-to-date information regarding the members.

Introduction

The aldo-keto reductase (AKR) superfamily consists of proteins found across all forms of life, archebacteria, prokaryotes and eukaroyotes. These proteins form a group based on their enzyme function to catalyze the reduction of carbonyl groups and their similar three-dimensional structure [ 1 ]. The superfamily is distinct from related functional proteins that belong to the short-chain dehydrogenase/reductase family and the medium chain alcohol dehydrogenase family [ 2 ]. The AKR superfamily contains over 200 confirmed members and over 30 potential members as of writing. AKRs are phase I enzymes that catalyze the reduction of carbonyl groups of various substrates. This function enables the resulting alcohol to undergo conjugation reactions for elimination. AKRs conduct oxidoreduction by using the cofactor nicotinamide adenine dinucleotide (phosphate) reduced (NAD(P) (H)) [ 3 ]. The protein structure of AKRs is characterized by an (α/β) 8 -barrel structure, three additional large loops, a cofactor binding site, and a catalytic tetrad [ 1 , 4 , 5 ]. Despite the similarities, some members of the AKR superfamily have additional functions such as the reduction of nitro-groups in nitro containing xenobiotics (AKR1C1-AKR1C4) [ 6 – 10 ], the reduction of steroid double bonds (AKR1D) [ 11 – 13 ], the oxidation of proximate carcinogen trans -dihydrodiol polycyclic aromatic hydrocarbons [ 14 , 15 ], and the activation of β-subunits of potassium gated ion channels (AKR6 family) [ 16 , 17 ]. Given the diverse roles AKRs have in many important and distinct biological processes, they have been the subject of much research for decades. In the five-year period from 2019 to 2023, over 900 AKR-related research articles accessible via PubMed from its underlying database Medline were published. Among these contributions were publications in high impact journals [ 18 , 19 ], as well as discoveries regarding the role of AKRs in oncology [ 20 – 24 ], chemotherapeutic drug resistance [ 25 – 27 ], endocrinology [ 28 – 31 ], toxicology [ 32 – 37 ], prognostic and diagnostic biomarker identification [ 38 – 49 ], the inactivation of glyphosate [ 50 ], and the presence of substrate cooperativity and allosteric sites that may extend across multiple family members [ 51 ]. Due to the size of, and interest in the superfamily, the need for a centralized database regarding the AKRs was evident and began in 1997, and an AKR superfamily webpage was created in 2003 to provide access to the database [ 1 ]. Since then, the database has been updated continuously via investigator-initiated submission. As many new family members have become available and tools for their analysis have improved, a newly configured website was released in 2023 ( https://akrsuperfamily.org// ). The database acts as a central location to access information regarding the AKR superfamily. Information gathered elsewhere such as protein database (PDB) structures and genetic sequences (NCBI) are directly linked from the website. New information that is generated by compiling all the members to provide multiple sequence alignments and family dendrograms is also provided. The AKR Superfamily webpage is publicly available and maintained by the Center of Excellence in Environmental Toxicology at the University of Pennsylvania.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-07-31T06:09:14.520117+00:00
unpaywall
last seen: 2026-07-31T06:42:51.797318+00:00