Abstract
S9 family proteins are serine proteases, divided into four subfamilies, involved in several functions associated with cell signalling, defence response and development. However, the annotation, characterization and statistical information are lacking. This is compounded by the huge number of bacterial genomes available. Hence, we have performed computational searches for S9 family peptidases as a step towards organising, curating the sequences and classifying into subfamilies. We have analysed annotated S9 family sequences from ∼32000 bacterial genomes/proteomes. All the curated information are presented as S9BacDB database ( http://caps.ncbs.res.in/S9BactDB ), provided in a user friendly way. The database provides various features such as information on the assemblies used (assembly, BioProject and BioSample details of each strain), the annotated POPs and their statistical distribution, unique domain architectures and the associated sequences and curated phylogenetic analysis. In addition, it provides the unique motifs and the associated information for Prolyl OligoPeptidases (POP) subtypes. An ML model is also integrated that can classify recognised sequences into a S9 sub family category. In conclusion, S9BactDb is a comprehensive platform that provides meticulously curated data on statistical/bioinformatics analysis of all S9 family proteins originating from fully sequenced bacterial genomes in the RefSeq database.
Full text
1,718 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
S9 family proteins are serine proteases, divided into four subfamilies, involved in several functions associated with cell signalling, defence response and development. However, the annotation, characterization and statistical information are lacking. This is compounded by the huge number of bacterial genomes available. Hence, we have performed computational searches for S9 family peptidases as a step towards organising, curating the sequences and classifying into subfamilies. We have analysed annotated S9 family sequences from ∼32000 bacterial genomes/proteomes. All the curated information are presented as S9BacDB database (http://caps.ncbs.res.in/S9BactDB), provided in a user friendly way. The database provides various features such as information on the assemblies used (assembly, BioProject and BioSample details of each strain), the annotated POPs and their statistical distribution, unique domain architectures and the associated sequences and curated phylogenetic analysis. In addition, it provides the unique motifs and the associated information for Prolyl OligoPeptidases (POP) subtypes. An ML model is also integrated that can classify recognised sequences into a S9 sub family category. In conclusion, S9BactDb is a comprehensive platform that provides meticulously curated data on statistical/bioinformatics analysis of all S9 family proteins originating from fully sequenced bacterial genomes in the RefSeq database.
Competing Interest Statement
The authors have declared no competing interest.
NList of abbreviations
- POP
- Prolyl Oligopeptidase
- OPB
- Oligopeptidase B
- DPP IV
- Dipeptidyl peptidase-4
- HMM
- Hidden Markov Model
- ML
- Machine Learning
- DA
- Domain Architecture
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.