{"paper_id":"406b98a8-ec1e-4b09-9403-c1f609729d79","body_text":"Subtractive Genome Analysis for Identification and Characterization of Novel Drug Targets in the Streptococcus pneumoniae by Using In-silico Approach | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Subtractive Genome Analysis for Identification and Characterization of Novel Drug Targets in the Streptococcus pneumoniae by Using In-silico Approach Sumit Sheoran, Swati Arora, Samson Raj R, Prachi Singh, Shashi kala, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2267778/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract As per WHO, the pneumococcus causes one million fatalities each year because of their underdeveloped immune systems, children are the most vulnerable to pneumococcal infections. The rise of S. pneumoniae resistance to antibiotics is causing widespread alarm all across the globe. Since the last couple of years, a recently developed technique is being used to overcome resistant pathogens. One of these is the computational subtractive genomics technique, in which the bacterial pathogen full proteins are effectively decreased to a limited amount of probable therapeutic targets. The procedures employed in this strategy are to locate human non-homologs targets, proteins that are vital to the illness producing agent, and participation of the identified proteins in pathogen metabolic pathways that are required for bacterial survival. In this work, we applied computational subtractive genomics on the proteins of the S. pneumonia strain D39 and came up with three cytoplasmic proteins that can serve as a promising candidate for novel drug to control the pathogenicity caused by S. pneumoniae. General Microbiology Bioinformatics S. pneumonia D39 Draggability SGA In-silico Characterization Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction The human pathogen Streptococcus pneumoniae (pneumococcus) is a Gram-positive bacterium. Streptococcus pneumoniae is one of the leading causes of illness and mortality. Pneumococcus is primarily a human pathogen, and in the United States, asymptomatic carriage in the nasopharynx occurs at least once by the age of two years [ 1 ]. When the pneumococcus enters typically sterile bodily areas, however, immunological dysregulation and illness might result. Pneumococcus is a prevalent pathogen that causes bacterial meningitis, pneumonia, otitis media (OM), sinusitis, and conjunctivitis [ 2 ]. The pneumococcus is liable for 1 million fatalities annually, as per the WHO (World Health Organization). Almost 14 million children under the age of five were identified with pneumococcal disease globally in 2000, with Africa having the greatest prevalence. In the globe, Nigeria, the most densely populated country in Africa, has one of the highest mortality rates owing to this disease; among 2000, 86,000 fatalities were predicted in kids under the age of five, the second highest of any country worldwide [ 3 ]. The pneumococcus is the most prevalent cause of bacterial OM. In the United States, OM is the leading cause of paediatric clinical visits and antibiotic prescriptions. Because of their underdeveloped immune systems, children are the human population most vulnerable to pneumococcal infections [ 4 ]. S. pneumoniae was previously thought to be totally sensitive to penicillin and other beta-lactam antibiotics. However, since the 1980s, there has been a remarkable growth in antibiotic resistance among S. pneumoniae in many regions of the world. Antibiotic resistance in S. pneumoniae is a serious problem across the world [ 5 ]. There is widespread worry about increased levels of antibiotic resistance, as well as concerns that the efficacy of antimicrobial therapy may be jeopardised, resulting in treatment failure and diminished value of older medicines [ 6 ]. Strains of S.pneumoniae are classified as serotypes based upon their antigen surface. Streptococci having strain D39 is a serotype 2 strain. D39 is a historically crucial strain as it was first used by Avery and co-worker experiment on DNA as a genetic material and it is extremely virulent and lethal in murine infection model. Moreover, this strain is almost used in current studies of pneumococcal pathogenesis This study was conducted to achieve and identify new putative drug targets against S. pneumoniae strain D39. We had studied S.pneumoniae D39 strain with detail in this research and applied the recent computational based approach subtractive genomics to recognize the possible putative drug targets against to this strain only. The current study uses hierarchical in silico method, i.e. subtractive proteomics approach and several bioinformatics tools in the discovery of new druggable targets in S. pneumoniae D39 strain. 2. Methodology 2.1. Protein sequence retrieval of pathogen: The full proteins of S. pneumoniae strain D39 were obtained from NCBI [ 7 ], Uniport and VarDb. 2.2. Finding non paralogous sequences: CD-HIT with a sequence identity criterion of 0.8 (i.e., 80%) was utilised to identify paralogous or duplicate protein successions [ 8 ]. Duplicate sequences were eliminated from S. pneumoniae strain D39 entire proteins, leaving only non-paralogous sequences. 2.3. Elimination of Homologous Protein Sequences from Pathogens Proteome: For the elimination of homologous proteins from non-paralogous datasets of pathogen, Blastp [ 9 ] was used against Reference Sequence Homo sapiens by selecting the 10 − 3 E-value that was basically the cut-off expectation. Blastp was showing two different results. The first one was for those protein sequences which were similar to the host sequences showed significant similarity without asterisk and the second one was for those sequences which were non-homologous sequences with asterisk and showing “No significance” as in result. Non-homologous Sequences were left for further analysis and homologous sequences were eliminated. 2.4. Determination of Essential Genes of Pathogen from BLASTp: Essential genes are the vital genes which are necessary for the normal working of cell including protein formation, replication, division for cell and metabolism etc. that are important for the structural sand functional support of microorganisms and these genes are crucial for the existence of the pathogen [ 10 ]. The information about Essential genes were collected from publicly available databases i.e. BLASTp(DEG). DEG is developed for crucial genes and proteins, which are extracted from different reports, scientific papers and experimental procedures [ 11 – 12 ]. DEG databases were utilized for the determination of non-host crucial genes that were present in it. We found 372 essential genes for the pathogen. 2.5. Analysis Metabolic Pathways in S. pneumonia : This step helped us to find the unique pathways which were not present in the host. So, that the pathways of host did not get disturbed or disrupted due to drug. So in this way, pathways of the pathogen were manually compared with human pathways by selecting Homo sapiens name as an Organism and searched saved pathway numbers (without prefix) [ 13 ] in KEGG. 2.6. Subcellular localization prediction : Unique proteins have been analysed through KEGG pathway analysis that were then used for the prediction of localization by most commonly used online localization tool called PSORTb. This step helped us in determining whether a protein was vaccine or drug target. For finding the subcellular localization of unidentified proteins latest version of PSORTb i.e. 3.0 was utilized. It is freely available on http://www.psort.org/psortb/ . We selected appropriate strain for our pathogen and used the unique pathway proteins sequences in PSORTb to identify their localization [ 14 ]. 2.6. Finding of Functional Family: Functional family prediction is compulsory for the identification of functional classes of hypothetical essential proteins that provides help in identification whether non-host proteins were vaccine or drug targeted. Interproscan was used for the functional family prediction of the separated imaginary protein sequences. The functional family must be known for a protein to be druggable so that we can easily predict the function of that target protein could design drug for it [ 15 ]. 2.7. Draggability Targets analysis: For the identification of druggable targets of host’s pathogen proteins, Drug Bank database was used [ 16 ] The proteins that were only belonging to unique pathway or unique protein were subjected to Drug Bank to find their targets. In this report, the 28 proteins that were screened out by pathway analysis were subjected to Drug Bank one by one to check druggability of those proteins. The screening of all essential, non-homologous that are involved in unique metabolic pathways was evaluated by Blastpv3 comparing against database of Drug Bank which consists of number of Protein targets which serve as a novel Drug target against the invasive disease caused by S.pneumoniae [ 16 ]. 2.8. Protein-Multiple ligand docking: The rational structure-based drug designing is a method that is used to cut down the time and cost involve in the designing and finding a drug against drug targets [ 17 ]. Structure based drug designing involve the 3D structure of drug target and the structure of ligand molecule (Drug) for drug discovery process. The docking was performed by using GOLD to find out the attraction between the drug targets and ligand and Biovia Discovery studio software was used for the visualization of docking pose. The structure of all the drug targets were downloaded from PDB database ( https://www.rcsb.org/ ) in PDB format The structure of all the ligands (Drugs) that were approved by FDA was used as a ligand for docking purpose so the structure of all the ligands were downloaded from PubChem ( https://pubchem.ncbi.nlm.nih.gov/ ) in 3D SDF format. Now the preparation of all the drug targets was done by using chimera, all the water molecules, co-ligands, heteroatoms were eliminated and hydrogen bonds were added from the crystal structure of proteins [ 18 ] so that they will not create any hindrance while docking. For docking purpose GOLD software was used to find out the binding affinities. For the visualization of docking poses Biovia discovery studio was used. 3. Result & Discussion The major and chief interest of the present study was to find out new drug targets in S.pneumoniae strain D39. The serious clinically syndromes of pneumococcal diseases are meningitis, bacteremia and pneumonia. In the developing world, five million children under the age of five die each year from acute lower respiratory infection of which S.pneumoniae is probably the main causative agent. Furthermore, the overall increase in penicillin resistance in pneumococci and the finite use of the pneumococcal vaccine indicate that morbidity and mortality from pneumococcal disease may expand. The step at which new drugs are being identified is much slow as compared to the demand for new anti-microbial drugs in market because it needs a high investment of money, technology, expertise and also the newly developing drug resistance agents are inhibiting and delaying the way towards new drug discovery. New and recent enhancements in Bioinformatics and its tools have been made very easy, helping and facilitating ways for the researchers to simplify the process of drug discovery against certain human pathogens. Among many approaches used in Bioinformatics, one of the most powerful ways mostly used by current researchers is subtractive proteomics approach for determining and identifying pathogen-specific drug targets. This approach is totally an in silico based approach that involves proteome analysis to sort out the essential proteins of a pathogen as unique without creating any disturbance such as destruction of host proteome or changing functions of host proteome, newer, precise and recent drugs against the host pathogen are designed by proper cross-checking of pathogen proteome with host proteome to avoid any toxicity in the host. A protein is said to be a good target if that protein is crucial and essential for the existence and survival of pathogens and in the absence of this protein, the life of pathogen is compromised and accommodated. We used subtractive proteomics approach to recognize and discover new beneficial targets against S.pneumoniae D39. This approach is reported in a number of published literatures for the identification, characterization and awareness of unique possible targets in human pathogens [ 19 ]. The number of output protein sequences from first to last step of the subtractive proteomics approach is illustrated in precise detail (Fig. 1 ). 3.1. Identification of Non-paralogous Sequences: The complete proteome of S. pneumoniae D39 strain was retrieved from UniProtKB in FASTA canonical format and downloaded an Excel file of the proteome. The file of proteome was containing total 1947 protein sequences of genome. The main goal of this study is to attain essential, non-homologous unique proteins of pathogen. The reason behind obtaining these proteins is that we have to gain only those crucial and important proteins of pathogen which are vital and typically necessary for the life of the pathogen and that proteins should be absent in the host proteome. This would be helpful in the development of drug targets for a particular strain. That is why essential proteins of a pathogen must be identified and investigated and is important to disturb and distract the bacterial survival in the host. After the proteome retrieval, we proceeded toward the next step. It was important to remove paralogous sequences from the genome of the pathogen in order to attack a specific target developing of drugs and vaccines. For this purpose, CD-HIT tool was used after retrieving the complete proteome of pathogen from UniProtKB. After the removal of paralogous sequences, the result of CD-HIT returned 1887 proteins out of 1947 proteins. The pathogen was containing 60 paralogous sequences in its proteome file that were removed easily by CD-HIT tool. The proteome file obtained after utilizing CD-HIT tool was totally free from paralogous sequences. And this non-paralogous dataset was further used in the next steps. 3.2. Determination and Identification of Non-Homologous Protein Sequences in the Pathogen: There should be a chance that the proteins of host and pathogen are homologous and can be present in both of them; therefore, it is very important to identify them and to remove these hosts homologous protein sequences from the pathogen proteome so that host cell toxicity is restricted. The other advantages this step providing us is securing the proteins of pathogen we are targeting and attacking. Proteins gained after the removal of paralogous sequences from CD-HIT result were then used in Blastp by selecting H. sapiens in the box of organism and RefSeq database with an expectation value (E-value) = 10 − 3 and resulted in the determination and identification of sequences that were specifically present in the pathogen only not in the host. In this way, we found 1423 Non-homologous sequences which were belonging to only pathogen after compiling the results of Blastp. The interested protein sequences in drug targeting and designing were Non-homologous sequences (without an asterisk) as they were only belonging to the pathogen while the Homologous protein (along an asterisk) sequences were present in both host and pathogen. 3.3. Identification of Essential Proteins: DEG Databases was used to identify essential proteins of the pathogen that play important and crucial role in the survival and existence of this pathogen [ 12 ]. To find and know about essential proteins, each non-homologous protein of pathogen that was obtained from the previous step was searched by using its gene name in DEG and j browser one by one. 372 essential proteins were identified by DEG out of 1423 non-host protein sequences of S. pneumonia D39 strain that is 19.10% of the total whole proteome. One of the searched results for a gene “ murE ” in DEG. 3.4. Metabolic Pathways Analysis by using KEGG: The resulted 372 essential proteins obtained from DEG databases were then subjected to the KEGG Pathway Database to identify the main and important biological pathways of bacteria in which these essential proteins work and contribute KEGG Pathway Database results are shown in Fig. 2 . In detail, 372 essential protein sequences of pathogen were subjected and checked in the KEGG Pathway Database by selecting the organism as S.pneumoniae D39 . Out of total 372 essential proteins, only 28 proteins were established to be unique in S.pneumoniae D39 from that 5 essential proteins were found as a part of peptidoglycan biosynthesis, 11 essential proteins were found to be involved in two component system,3 proteins were involved in Vancomycin resistance, 1 was involved in Microbial metabolism in diverse environments systems, 1 proteins was involved in Antigen nucleotide biosynthesis, 4 proteins were involved in Quorum sensing, 6 proteins were involved in Biosynthesis of Secondary metabolites, 3 proteins were involved in Cationic antimicrobial peptide (CAMP) resistance. Essential functions performed by these proteins in the pathogenic pathways are shown in Fig. 10. At the end of this step, we obtained two different types of data from the results of KEGG pathway analysis that included: unique pathways and common pathways. The result of this comparison, on the basis of pathway number and pathway name, produced three types of datasets, one was of unique proteins when no hits found in KEGG, second was when both the host and pathogen were containing the same proteins called common proteins/pathways and the last one was when the protein was only present in the pathogen not in the human, were known as unique pathways. Unique proteins have their vital role in unique pathways. The unique pathway data was encouraging and helpful in finding the suitable and possible drug targets because it consisted of pathways that were specific to the pathogen only. Table.1. Unique Pathways of Bacteria. Sr.no. Pathways 1 Peptidoglycan Biosynthesis 2 Lysine Biosynthesis 3 Lipopolysaccharide Biosynthesis 4 Biosynthesis of Secondary Metabolites 5 Microbial Metabolism in Diverse Environments 6 PhosphoTransferase System (PTS) 7 Beta-Lactam Resistance 8 Quorum Sensing 9 Bacterial Secretion System 10 Two-Component System 11 Methane Metabolism 12 Cationic Antimicrobial Peptide (CAMP) Resistance 13 Pyrimidine metabolism 14 Xylene degradation 15 Benzoate degradation 3.5. Analysis and Exploration of Localization Results: The protein should be in proper place within the living cell for its appropriate and regular work. The identification and analysis for the localization of target proteins was important because of the drug requires attaching with the target in order to show its action and the target present in the cell helps finally to make proper drug compound. Proteins should be at the suitable and proper section of the cell for carrying out their specific function at specific location. Subcellular localization is one of the typical characteristics for drug targeting. The unique and non-homologous proteins obtained from PSORTb v3.0 were used for the determination of their subcellular localization 28 essential, unique and non-homologous proteins gained from KEGG Pathway were subjected into PSORTb to determine their subcellular localization. The results obtained from PSORTb showed that 71% proteins were in cytoplasmic region of the cell, 25% proteins were found to be in integral plasma membrane proteins and 4% present in extracellular space. Figure 3 shows the graphical representation of outcomes obtained from PSORTb. 3.6. Analysis of Functional Family of Essential, Non-homologous and Unique Proteins of Pathogen: Interproscan was used for the prediction of the functional family classification of the non-homologous, unique and essential protein sequences of S.pneumoniae D39 strain (Zdobnov & Apweiler, 2001). The Twenty-eight unique protein sequences were pasted in Interproscan to get the family prediction of each protein one by one for these 28 query sequences, many functional classes were predicted. The results obtained from Interproscan with the possible classification of functional families of some essential and unique short-listed proteins are illustrated in table (2). Table 2 Annotation of Function of Proteins of Pathogen S.r. number Entry ID Protein family names 1 A0A0H2ZMF9 Transpeptidase-like 2 Q04JL8 Ligases 3 A0A0H2ZPR0 phosphonate metabolism 4 A0A0H2ZNR2 Firmicutes 5 A0A0H2ZLN7 Transcriptional regulatory protein WalR-like 6 A0A0H2ZLY7 D-alanyl carrier protein 7 A0A0H2ZMU1 Diacylglycerol kinase like 8 A0A0H2ZMP6 SUF system FeS cluster assembly 9 Q04MC4 Glycosyl transferase, family 4 10 Q04LK0 transferase family 11 A0A0H2ZP52 HPr-like superfamily 12 A0A0H2ZQA1 RNA polymerase sigma factor 13 A0A0H2ZPN8 Acyl carrier protein family 14 A0A0H2ZR18 Nitroreductase 15 A0A0H2ZMW4 All DNA binding proteins 16 A0A0H2ZNP9 nucleoside triphosphate hydrolase like 17 A0A0H2ZMY1 Oxidoreductase 18 A0A0H2ZML1 Class I glutamine aminotransferases-like 19 A0A0H2ZNH9 ATPase superfamily 20 Q04N63 replication initiator 3.7. Analysis of Drug ability Potential of Unique Proteins of Pathogen: The identification and analysis of druggable proteins of S. pneumonia D39 strain was carried out in this last step by using Drug Bank Database. For this reason, the 28 shortlisted unique, essential and non-homologous protein names were searched in Drug Bank Database one by one to identify druggability of these proteins. For possible drug target it showed the identification of 5 proteins that can be useful and they integrated to show similarity with the previous existing targets of drugs in the Drug Bank database [ 21 ]. All these 5 druggable proteins were found to be in different and specific locations of the cell. 3 out of 5 proteins were present in cytoplasmic region of the cell whereas, 2 proteins were present in the plasma membrane The total obtained 5 druggable target proteins were involved in different and highly specific metabolic pathways that were totally essential for the under studied pathogen S. pneumoniae D39 strain. Strain specific drug targeting the non-homologous essential proteins that involved in unique metabolic pathways ensure the eradication of invasive and non-invasive diseases cause by S.pneumoniae with fewer side effects to the host. As we know that cytoplasmic proteins are considered more suitable one for Drug target as it is difficult to purify the proteins that are present in the membrane of the cell [ 20 ]. Moreover, surface proteins have direct contact so these protein targets have ability to obtain immune response when defined to the host Table 3 Information about 3 Druggable Proteins of S. pneumoniae D39 involved in Cytoplasmic region of the cell Gene Name Subcellular localization DrugBank ID ciaR SPD_0701 Cytoplasmic DB01972 PtsH SPD_1040 Cytoplasmic DB01899 Ptsl SPD_1039 Cytoplasmic DB08357 All the 3 proteins can be regarded as a potent Drug target because all have one homologous in Drug Bank database having at least 30% sequence similarity and serve as a promising candidate to fight against the infectious diseases caused by Streptococcus pneumoniae D39 (Table 3 ). 3.8. Analysis of protein-Multiple ligand docking: Structural based virtual screening also called as target based virtual screening that is used to find out the interaction between the receptor and ligands (Lyne, 2002). As results all the ligands were ranked according to their binding affinity with the drug target. Scoring function is very important in docking as it is responsible for the prediction of binding affinity. Thus, scoring function results are the main reason for the failure or success of docking. The more positive a binding affinity better the ligand has been bound to the protein so for the first drug target that is ciaR SPD_0701 shown the maximum binding affinity with the Piperacillin anhydrous having the best binding affinity and lower Root mean square deviation (RMSD) (Fig. 4 ) For the second drug target that is ptsH SPD_1040 shows the maximum binding affinity with penicillin G (act as a ligand) and yellow color shows the binding of penicillin G with the drug target. (Fig. 5 ) 4. Conclusion The genome sequencing technology at a large-scale has been improved. The world of drug-designing has also played a revolutionary role in the domains of microbiology and pathology. As compared to the earlier decades, the genomic and proteomic data is also increasing day by day, hence, creating a massive increase in the biological databases. The researchers find the successful treatments against the infections caused by those pathogens whose sequence information has been reported. The novel pathogenic strains with high resistance against antibacterial agents are a great threat towards the management of infections and diseases. Due to a significant role of Bioinformatics, the researchers become enabled to identify the novel drug targets against the highly resistant bacterial strains which have been emerged currently. The approach of subtractive proteomics has gained high popularity in the field of drug-designing against human pathogens. This method enables the researchers for rapid identification of potential targets in pathogens and targeting them as treatment for that particular infection. This study is aimed to use the latest and updated databases of proteins, advanced computational tools and subtractive proteomics approach to characterize the novel drug targets against the pathogenic proteins for treatment purpose. To identify the putative drug targets in any pathogen requires a reliable, short-term, fast, valuable and inexpensive approach that not only predicts drug targets but also predicts multiple drug resistant abilities of pathogens by using computational based tools to find subtraction in their whole genome. Similarly, subtractive proteomics approach has been used for the last decades to design highly specific drug targets in opposition to certain bacterial diseases or human pathogens for decreasing partially or completely stops the pathogenicity in the host. By using this in silico based modern hierarchical approach, the use of conventional approaches of drug development has reduced to a maximum level. This study was also managed by using subtractive proteomics approach and 3 non-homologous, unique and essential targets out of 1947 proteins were shortlisted in S. pneumoniae D39 strain only. While this approach has many advantages but it also has a little disadvantage due to some predictive tools used in this approach. The work of all of the tools used in this modern approach is based on prediction and algorithms. The errors may occur in these tools as the algorithms for these tools have been designed and coded by human and their specificity is limited for a specific task. Basically, that was the big limitation of this modern computational based approach. References Barry M. Gray, George M. Converse, III, Hugh C. Dillon, Jr., Epidemiologic Studies of Streptococcus pneumoniae in Infants: Acquisition, Carriage, and Infection during the First 24 Months of Life, The Journal of Infectious Diseases , Volume 142, Issue 6, December 1980, Pages 923–933, https://doi.org/10.1093/infdis/142.6.923 Daniel M. Musher, Infections Caused by Streptococcus pneumoniae : Clinical Spectrum, Pathogenesis, Immunity, and Treatment, Clinical Infectious Diseases , Volume 14, Issue 4, April 1992, Pages 801–809, https://doi.org/10.1093/clinids/14.4.801 Nancy Y. Yu, James R. Wagner, Matthew R. Laird, Gabor Melli, Sébastien Rey, Raymond Lo, Phuong Dao, S. Cenk Sahinalp, Martin Ester, Leonard J. Foster, Fiona S. L. Brinkman, PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes, Bioinformatics , Volume 26, Issue 13, 1 July 2010, Pages 1608–1615, https://doi.org/10.1093/bioinformatics/btq249 Nancy Y. Yu, James R. Wagner, Matthew R. Laird, Gabor Melli, Sébastien Rey, Raymond Lo, Phuong Dao, S. Cenk Sahinalp, Martin Ester, Leonard J. Foster, Fiona S. L. Brinkman, PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes, Bioinformatics , Volume 26, Issue 13, 1 July 2010, Pages 1608–1615, https://doi.org/10.1093/bioinformatics/btq249 Barh, D., Tiwari, S., Jain, N., Ali, A., Santos, A. R., Misra, A. N., … Kumar, A.(2010). In silico subtractive genomics for target identification in human bacterial pathogens.Drug Development Research, 72(2), 162–177. doi:10.1002/ddr.20413 Katherine L O'Brien, Lara J Wolfson, James P Watt, Emily Henkle, Maria Deloria-Knoll, Natalie McCall, Ellen Lee, Kim Mulholland, Orin S Levine, Thomas Cherian,Burden of disease caused by Streptococcus pneumoniae in children younger than 5 years: global estimates,The Lancet, Volume 374, Issue 9693, 2009,Pages 893–902,ISSN 0140–6736, https://doi.org/10.1016/S0140-6736(09)61204-6 . Ruch, P., Teodoro, D., Consortium, U., & others. (2021). UniProt . Sáez-Llorens, X., & McCracken Jr, G. H. (2003). Bacterial meningitis in children. The Lancet 361 (9375), 2139–2148. Weizhong Li, Adam Godzik, Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences, Bioinformatics , Volume 22, Issue 13, 1 July 2006, Pages 1658–1659, https://doi.org/10.1093/bioinformatics/btl158 Lavigne, R., Seto, D., Mahadevan, P., Ackermann, H.-W., & Kropinski, A. M.(2008). Unifying classical and molecular taxonomic classification: analysis of the Podoviridae using BLASTP-based tools. Research in Microbiology , 159 (5), 406–414. Zhang, R., & Lin, Y. (2009). DEG 5.0, a database of essential genes in both prokaryotes and eukaryotes. Nucleic Acids Research , 37 (suppl\\_1), D455–D458. Luo, H., Lin, Y., Liu, T., Lai, F.-L., Zhang, C.-T., Gao, F., & Zhang, R. (2021). DEG 15, an update of the Database of Essential Genes that includes built-in analysis tools. Nucleic Acids Research , 49 (D1), D677–D686. Slager, J., Aprianto, R., & Veening, J.-W. (2018). Deep genome annotation of the opportunistic human pathogen Streptococcus pneumoniae D39. Nucleic Acids Research , 46 (19), 9971–9989. Kanehisa, M., & Goto, S. (2000). KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Research , 28 (1), 27–30. Yu, N. Y., Wagner, J. R., Laird, M. R., Melli, G., Rey, S., Lo, R., Dao, P., Sahinalp, S. C., Ester, M., Foster, L. J., & others. (2010). PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes. Bioinformatics , 26 (13), 1608–1615. Quevillon, E., Silventoinen, V., Pillai, S., Harte, N., Mulder, N., Apweiler, R., & Lopez, R. (2005). InterProScan: protein domains identifier. Nucleic Acids Research , 33 (suppl\\_2), W116–W120. Sood, N., Ribero, R., Ryan, M., & Van Nuys, K. (2020). The association between drug rebates and list prices. Los Angeles: University of Southern California Leonard D. Schaeffer Center for Health Policy and Economics , 500 . Maia, E. H. B., Assis, L. C., de Oliveira, T. A., da Silva, A. M., & Taranto, A. G. (2020). Structure-based virtual screening: from classical to artificial intelligence. Frontiers in Chemistry , 8 , 343. Ibrahim, M. T., Uzairu, A., Shallangwa, G. A., & Uba, S. (2020). In-silico activity prediction and docking studies of some 2, 9-disubstituted 8-phenylthio/phenylsulfinyl-9h-purine derivatives as Anti-proliferative agents. Heliyon , 6 (1), e03158. Wadhwani, A., & Khanna, V. (2016). In silico identification of novel potential vaccine candidates in Streptococcus pneumoniae. Global J Technol Optim , 7 (109), 2. Shahid, F., Ashfaq, U. A., Saeed, S., Munir, S., Almatroudi, A., & Khurshid, M. (2020). In Silico Subtractive Proteomics Approach for Identification of Potential Drug Targets in Staphylococcus saprophyticus. International Journal of Environmental Research and Public Health , 17 (10), 3644. Uddin, R., Arif, A., Zahra, N.-A., & Sufian, M. (2021). Comparative proteome-wide study for in-silico identification and characterization of indispensable hypothetical proteins of food bornepathogen Campylobacter jejuni (CJJ) by subtractive genomics approach. Pakistan Journal of Pharmaceutical Sciences , 34 (4). Additional Declarations Competing interests: The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-2267778\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":true,\"archivedVersions\":[],\"articleType\":\"Research Article\",\"associatedPublications\":[],\"authors\":[{\"id\":151639226,\"identity\":\"4f777f6b-f14e-4bdf-bfcf-dd26fe59614e\",\"order_by\":0,\"name\":\"Sumit Sheoran\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"pdm university, Bahadurgarh\",\"correspondingAuthor\":false,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Sumit\",\"middleName\":\"\",\"lastName\":\"Sheoran\",\"suffix\":\"\"},{\"id\":151639227,\"identity\":\"50808aae-bd9a-4c11-989e-8c13bafa0da9\",\"order_by\":1,\"name\":\"Swati Arora\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABAElEQVRIiWNgGAWjYPACCQYGZuYDBz5UMIAYDURqYW9LfDjjDEgLI1FagIDnjLExbxuIRUCL/IzkZx9+7rCQl5+RYCbBO682mr8dqOVHxTacWgxupBnP7D0jYbjhRkKahOS247kzDjM2MPacuY1bi3SCMQNvm0SCgUTCMQnDbcdyG4BamBnbcGuRn53+mfEvUIv8jMQ2icQ5x3LnE9LCcDvHmBlkC8OZw8wGBxtqcjcQ0mJw/00xsyzIL8fbGB82HDuQuxGo5SA+v8j3HN/M+HZHnbx8M/+Hw39q6nLnnT988MGPCjwOAwGkiDgMJg/gV4+qpY6g4lEwCkbBKBh5AABOvl8JxLkbzQAAAABJRU5ErkJggg==\",\"orcid\":\"\",\"institution\":\"\",\"correspondingAuthor\":true,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Swati\",\"middleName\":\"\",\"lastName\":\"Arora\",\"suffix\":\"\"},{\"id\":151639228,\"identity\":\"37e55000-b40a-434c-bf3e-399ae2a6e5be\",\"order_by\":2,\"name\":\"Samson Raj R\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"\",\"correspondingAuthor\":false,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Samson\",\"middleName\":\"Raj\",\"lastName\":\"R\",\"suffix\":\"\"},{\"id\":151639229,\"identity\":\"d24dcd3d-131d-48df-bb85-44e9c62446c0\",\"order_by\":3,\"name\":\"Prachi Singh\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"\",\"correspondingAuthor\":false,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Prachi\",\"middleName\":\"\",\"lastName\":\"Singh\",\"suffix\":\"\"},{\"id\":151639230,\"identity\":\"4c2158f5-9ad2-4ef9-a831-a8c28c24ecf0\",\"order_by\":4,\"name\":\"Shashi kala\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"\",\"correspondingAuthor\":false,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Shashi\",\"middleName\":\"\",\"lastName\":\"kala\",\"suffix\":\"\"},{\"id\":151639231,\"identity\":\"79b1055d-7393-447a-b127-6b0093cbb496\",\"order_by\":5,\"name\":\"Aayushi Velingkar\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"\",\"correspondingAuthor\":false,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Aayushi\",\"middleName\":\"\",\"lastName\":\"Velingkar\",\"suffix\":\"\"},{\"id\":151639232,\"identity\":\"2026cb0c-53df-4c17-8fdd-cf954059b643\",\"order_by\":6,\"name\":\"Ahmed Khaireh Mohamed\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABI0lEQVRIiWNgGAWjYFACHhDBzGAAJCUbG0BcHsYHIHE+UrQwgzg8bMRqASuWAFG4tOi29x78dKPCWt5c7PDBmzN32MmYt/ceq/yaYyfDxsD88NENTC1mZ84lS+ecSTfcOTst2XLjmWQemTPn0m7LbksGOozN2DgHi5YbOQbSuW2HGTfczjGTfNjGzCMhkWN2W3IbM1ALD5s0di3Gv3P/HbbfcDv/G1BLPY+E/BuzYslt9fi0mEnnNhxOBNrCJrmx7TDQFh4zxo/bDuPWcuaMmXXOsfTkDbfTjC1nth3nkeDJMZZm3Hach40Zh1+O9xjfzqmxtt1wO/nhzd62ansJ9jOGH39uq7bnZ29++BiLFuyAGRpZJADGH6SoHgWjYBSMguEOAKOxZD8e6H1/AAAAAElFTkSuQmCC\",\"orcid\":\"\",\"institution\":\"\",\"correspondingAuthor\":true,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Ahmed\",\"middleName\":\"Khaireh\",\"lastName\":\"Mohamed\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2022-11-13 07:14:46\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-2267778/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-2267778/v1\",\"draftVersion\":[],\"editorialEvents\":[],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":29148053,\"identity\":\"8cae0f70-b9da-4252-ad00-855a99201e03\",\"added_by\":\"auto\",\"created_at\":\"2022-11-16 17:26:47\",\"extension\":\"png\",\"order_by\":1,\"title\":\"Figure 1\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":8160,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eIndicating No. of proteins left behind after the end of each step of the Subtractive proteomics Approach.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Onlinedrawingimage1.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-2267778/v1/7fba03e86e098beacbedfc11.png\"},{\"id\":29148052,\"identity\":\"a3090212-4df6-462c-8771-70db51be3670\",\"added_by\":\"auto\",\"created_at\":\"2022-11-16 17:26:47\",\"extension\":\"png\",\"order_by\":2,\"title\":\"Figure 2\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":6239,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eIndication and Describing of Unique and Common Proteins of \\u003cem\\u003eS. pneumoniae\\u003c/em\\u003e D39 strain from KEGG pathway results.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Onlinedrawingimage2.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-2267778/v1/29e0af54ec817a2d572bf119.png\"},{\"id\":29148060,\"identity\":\"b964079f-ab15-4a4e-b7ba-57a1a2b30935\",\"added_by\":\"auto\",\"created_at\":\"2022-11-16 17:26:47\",\"extension\":\"png\",\"order_by\":3,\"title\":\"Figure 3\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":8286,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eIllustration of Subcellular Localization for Unique proteins of S.pneumoniae D39 Strain.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Onlinedrawingimage3.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-2267778/v1/7cb2819f07b54fff584ca1cc.png\"},{\"id\":29148068,\"identity\":\"4a850c52-d1a4-41b6-a957-bb6a3fca2963\",\"added_by\":\"auto\",\"created_at\":\"2022-11-16 17:26:47\",\"extension\":\"jpeg\",\"order_by\":4,\"title\":\"Figure 4\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":248872,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003e3D \\u0026amp; 2D structure of Molecular docking of the target \\u003cem\\u003eciaR \\u003c/em\\u003eSPD_0701 observed to be interactive with (Piperacillin anhydrous) as a ligand by visualizing it on Biovia discovery studio\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"floatimage1.jpeg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-2267778/v1/3fe670d00718c2e321cdbc3a.jpeg\"},{\"id\":29148661,\"identity\":\"1cf54733-7cec-4a01-88cc-ad639d7aba43\",\"added_by\":\"auto\",\"created_at\":\"2022-11-16 17:34:47\",\"extension\":\"jpeg\",\"order_by\":5,\"title\":\"Figure 5\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":240976,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003e3D \\u0026amp; 2D structure of Molecular docking of the target \\u003cem\\u003epts H \\u003c/em\\u003eSPD_1040 and the ligand (Penicillin G) by visualizing it on Biovia discovery studio and finally the last drug target that is \\u003cem\\u003eptsI \\u003c/em\\u003eSPD_1039 shown the maximum binding affinity with the oxacillin. (Fig. 6.)\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"floatimage2.jpeg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-2267778/v1/4d7cca4d5a6666cc9447f243.jpeg\"},{\"id\":29148069,\"identity\":\"ae846836-f963-4b91-9c29-72e5ce4c95c1\",\"added_by\":\"auto\",\"created_at\":\"2022-11-16 17:26:48\",\"extension\":\"jpeg\",\"order_by\":6,\"title\":\"Figure 6\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":164255,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003e3D \\u0026amp; 2D structure of Molecular docking of the target \\u003cem\\u003eptsI \\u003c/em\\u003eSPD_1039 and the ligand (Oxacillin) by visualizing it on Biovia discovery studio\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"floatimage3.jpeg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-2267778/v1/8164f55c61a0655c96063b1f.jpeg\"},{\"id\":29148662,\"identity\":\"672a405d-ae8c-4b80-ba66-9209e9bf1afc\",\"added_by\":\"auto\",\"created_at\":\"2022-11-16 17:34:53\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":637697,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"manuscript.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-2267778/v1/8b4615ec-8655-406f-9f7d-ede9fcdefbb8.pdf\"}],\"financialInterests\":\"\\u003cp\\u003eCompeting interests: The authors declare no competing interests.\\u003c/p\\u003e\",\"formattedTitle\":\"\\u003cp\\u003eSubtractive Genome Analysis for Identification and Characterization of Novel Drug Targets in the \\u003cem\\u003eStreptococcus pneumoniae\\u003c/em\\u003e by Using \\u003cem\\u003eIn-silico\\u003c/em\\u003e Approach \\u003c/p\\u003e\",\"fulltext\":[{\"header\":\"1. Introduction\",\"content\":\"\\u003cp\\u003eThe human pathogen Streptococcus pneumoniae (pneumococcus) is a Gram-positive bacterium. Streptococcus pneumoniae is one of the leading causes of illness and mortality. Pneumococcus is primarily a human pathogen, and in the United States, asymptomatic carriage in the nasopharynx occurs at least once by the age of two years [\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e].\\u003c/p\\u003e \\u003cp\\u003eWhen the pneumococcus enters typically sterile bodily areas, however, immunological dysregulation and illness might result. Pneumococcus is a prevalent pathogen that causes bacterial meningitis, pneumonia, otitis media (OM), sinusitis, and conjunctivitis [\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e]. The pneumococcus is liable for 1\\u0026nbsp;million fatalities annually, as per the WHO (World Health Organization). Almost 14\\u0026nbsp;million children under the age of five were identified with pneumococcal disease globally in 2000, with Africa having the greatest prevalence. In the globe, Nigeria, the most densely populated country in Africa, has one of the highest mortality rates owing to this disease; among 2000, 86,000 fatalities were predicted in kids under the age of five, the second highest of any country worldwide [\\u003cspan citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e]. The pneumococcus is the most prevalent cause of bacterial OM. In the United States, OM is the leading cause of paediatric clinical visits and antibiotic prescriptions. Because of their underdeveloped immune systems, children are the human population most vulnerable to pneumococcal infections [\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e]. S. pneumoniae was previously thought to be totally sensitive to penicillin and other beta-lactam antibiotics. However, since the 1980s, there has been a remarkable growth in antibiotic resistance among S. pneumoniae in many regions of the world. Antibiotic resistance in S. pneumoniae is a serious problem across the world [\\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e5\\u003c/span\\u003e]. There is widespread worry about increased levels of antibiotic resistance, as well as concerns that the efficacy of antimicrobial therapy may be jeopardised, resulting in treatment failure and diminished value of older medicines [\\u003cspan citationid=\\\"CR6\\\" class=\\\"CitationRef\\\"\\u003e6\\u003c/span\\u003e].\\u003c/p\\u003e \\u003cp\\u003eStrains of \\u003cem\\u003eS.pneumoniae\\u003c/em\\u003e are classified as serotypes based upon their antigen surface. Streptococci having strain D39 is a serotype 2 strain. D39 is a historically crucial strain as it was first used by Avery and co-worker experiment on DNA as a genetic material and it is extremely virulent and lethal in murine infection model. Moreover, this strain is almost used in current studies of pneumococcal pathogenesis\\u003c/p\\u003e \\u003cp\\u003eThis study was conducted to achieve and identify new putative drug targets against \\u003cem\\u003eS. pneumoniae\\u003c/em\\u003e strain D39. We had studied \\u003cem\\u003eS.pneumoniae\\u003c/em\\u003e D39 strain with detail in this research and applied the recent computational based approach subtractive genomics to recognize the possible putative drug targets against to this strain only. The current study uses hierarchical \\u003cem\\u003ein silico\\u003c/em\\u003e method, i.e. subtractive proteomics approach and several bioinformatics tools in the discovery of new druggable targets in \\u003cem\\u003eS. pneumoniae\\u003c/em\\u003e D39 strain.\\u003c/p\\u003e\"},{\"header\":\"2. Methodology\",\"content\":\"\\u003cdiv id=\\\"Sec3\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.1. Protein sequence retrieval of pathogen:\\u003c/h2\\u003e \\u003cp\\u003eThe full proteins of S. pneumoniae strain D39 were obtained from NCBI [\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e], Uniport and VarDb.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec4\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.2. Finding non paralogous sequences:\\u003c/h2\\u003e \\u003cp\\u003eCD-HIT with a sequence identity criterion of 0.8 (i.e., 80%) was utilised to identify paralogous or duplicate protein successions [\\u003cspan citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e8\\u003c/span\\u003e]. Duplicate sequences were eliminated from S. pneumoniae strain D39 entire proteins, leaving only non-paralogous sequences.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec5\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.3. Elimination of Homologous Protein Sequences from Pathogens Proteome:\\u003c/h2\\u003e \\u003cp\\u003eFor the elimination of homologous proteins from non-paralogous datasets of pathogen, Blastp [\\u003cspan citationid=\\\"CR9\\\" class=\\\"CitationRef\\\"\\u003e9\\u003c/span\\u003e] was used against Reference Sequence Homo sapiens by selecting the 10\\u003csup\\u003e\\u0026minus;\\u0026thinsp;3\\u003c/sup\\u003e E-value that was basically the cut-off expectation. Blastp was showing two different results. The first one was for those protein sequences which were similar to the host sequences showed significant similarity without asterisk and the second one was for those sequences which were non-homologous sequences with asterisk and showing \\u0026ldquo;No significance\\u0026rdquo; as in result. Non-homologous Sequences were left for further analysis and homologous sequences were eliminated.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec6\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.4. Determination of Essential Genes of Pathogen from BLASTp:\\u003c/h2\\u003e \\u003cp\\u003eEssential genes are the vital genes which are necessary for the normal working of cell including protein formation, replication, division for cell and metabolism etc. that are important for the structural sand functional support of microorganisms and these genes are crucial for the existence of the pathogen [\\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e]. The information about Essential genes were collected from publicly available databases i.e. BLASTp(DEG). DEG is developed for crucial genes and proteins, which are extracted from different reports, scientific papers and experimental procedures [\\u003cspan citationid=\\\"CR11\\\" class=\\\"CitationRef\\\"\\u003e11\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e12\\u003c/span\\u003e]. DEG databases were utilized for the determination of non-host crucial genes that were present in it. We found 372 essential genes for the pathogen.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec7\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.5. Analysis Metabolic Pathways in \\u003cem\\u003eS. pneumonia\\u003c/em\\u003e:\\u003c/h2\\u003e \\u003cp\\u003eThis step helped us to find the unique pathways which were not present in the host. So, that the pathways of host did not get disturbed or disrupted due to drug. So in this way, pathways of the pathogen were manually compared with human pathways by selecting Homo sapiens name as an Organism and searched saved pathway numbers (without prefix) [\\u003cspan citationid=\\\"CR13\\\" class=\\\"CitationRef\\\"\\u003e13\\u003c/span\\u003e] in KEGG.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec8\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e\\u003cb\\u003e2.6. Subcellular localization prediction\\u003c/b\\u003e:\\u003c/h2\\u003e \\u003cp\\u003eUnique proteins have been analysed through KEGG pathway analysis that were then used for the prediction of localization by most commonly used online localization tool called PSORTb. This step helped us in determining whether a protein was vaccine or drug target. For finding the subcellular localization of unidentified proteins latest version of PSORTb i.e. 3.0 was utilized. It is freely available on \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttp://www.psort.org/psortb/\\u003c/span\\u003e\\u003cspan address=\\\"http://www.psort.org/psortb/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e. We selected appropriate strain for our pathogen and used the unique pathway proteins sequences in PSORTb to identify their localization [\\u003cspan citationid=\\\"CR14\\\" class=\\\"CitationRef\\\"\\u003e14\\u003c/span\\u003e].\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec9\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.6. Finding of Functional Family:\\u003c/h2\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"BlockQuote\\\"\\u003e \\u003cp\\u003eFunctional family prediction is compulsory for the identification of functional classes of hypothetical essential proteins that provides help in identification whether non-host proteins were vaccine or drug targeted. Interproscan was used for the functional family prediction of the separated imaginary protein sequences. The functional family must be known for a protein to be druggable so that we can easily predict the function of that target protein could design drug for it [\\u003cspan citationid=\\\"CR15\\\" class=\\\"CitationRef\\\"\\u003e15\\u003c/span\\u003e].\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec10\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.7. Draggability Targets analysis:\\u003c/h2\\u003e \\u003cp\\u003eFor the identification of druggable targets of host\\u0026rsquo;s pathogen proteins, Drug Bank database was used [\\u003cspan citationid=\\\"CR16\\\" class=\\\"CitationRef\\\"\\u003e16\\u003c/span\\u003e] The proteins that were only belonging to unique pathway or unique protein were subjected to Drug Bank to find their targets. In this report, the 28 proteins that were screened out by pathway analysis were subjected to Drug Bank one by one to check druggability of those proteins. The screening of all essential, non-homologous that are involved in unique metabolic pathways was evaluated by Blastpv3 comparing against database of Drug Bank which consists of number of Protein targets which serve as a novel Drug target against the invasive disease caused by \\u003cem\\u003eS.pneumoniae\\u003c/em\\u003e [\\u003cspan citationid=\\\"CR16\\\" class=\\\"CitationRef\\\"\\u003e16\\u003c/span\\u003e].\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec11\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.8. Protein-Multiple ligand docking:\\u003c/h2\\u003e \\u003cp\\u003eThe rational structure-based drug designing is a method that is used to cut down the time and cost involve in the designing and finding a drug against drug targets [\\u003cspan citationid=\\\"CR17\\\" class=\\\"CitationRef\\\"\\u003e17\\u003c/span\\u003e]. Structure based drug designing involve the 3D structure of drug target and the structure of ligand molecule (Drug) for drug discovery process.\\u003c/p\\u003e \\u003cp\\u003eThe docking was performed by using GOLD to find out the attraction between the drug targets and ligand and Biovia Discovery studio software was used for the visualization of docking pose. The structure of all the drug targets were downloaded from PDB database (\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://www.rcsb.org/\\u003c/span\\u003e\\u003cspan address=\\\"https://www.rcsb.org/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003cspan type=\\\"Underline\\\" class=\\\"Underline\\\" name=\\\"Emphasis\\\"\\u003e)\\u003c/span\\u003e in PDB format\\u003c/p\\u003e \\u003cp\\u003eThe structure of all the ligands (Drugs) that were approved by FDA was used as a ligand for docking purpose so the structure of all the ligands were downloaded from PubChem (\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://pubchem.ncbi.nlm.nih.gov/\\u003c/span\\u003e\\u003cspan address=\\\"https://pubchem.ncbi.nlm.nih.gov/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e) in 3D SDF format. Now the preparation of all the drug targets was done by using chimera, all the water molecules, co-ligands, heteroatoms were eliminated and hydrogen bonds were added from the crystal structure of proteins [\\u003cspan citationid=\\\"CR18\\\" class=\\\"CitationRef\\\"\\u003e18\\u003c/span\\u003e] so that they will not create any hindrance while docking. For docking purpose GOLD software was used to find out the binding affinities. For the visualization of docking poses Biovia discovery studio was used.\\u003c/p\\u003e \\u003c/div\\u003e\"},{\"header\":\"3. Result \\u0026 Discussion\",\"content\":\"\\u003cp\\u003eThe major and chief interest of the present study was to find out new drug targets in \\u003cem\\u003eS.pneumoniae\\u003c/em\\u003e strain D39. The serious clinically syndromes of pneumococcal diseases are meningitis, bacteremia and pneumonia. In the developing world, five million children under the age of five die each year from acute lower respiratory infection of which \\u003cem\\u003eS.pneumoniae\\u003c/em\\u003e is probably the main causative agent. Furthermore, the overall increase in penicillin resistance in pneumococci and the finite use of the pneumococcal vaccine indicate that morbidity and mortality from pneumococcal disease may expand. The step at which new drugs are being identified is much slow as compared to the demand for new anti-microbial drugs in market because it needs a high investment of money, technology, expertise and also the newly developing drug resistance agents are inhibiting and delaying the way towards new drug discovery. New and recent enhancements in Bioinformatics and its tools have been made very easy, helping and facilitating ways for the researchers to simplify the process of drug discovery against certain human pathogens. Among many approaches used in Bioinformatics, one of the most powerful ways mostly used by current researchers is subtractive proteomics approach for determining and identifying pathogen-specific drug targets. This approach is totally an \\u003cem\\u003ein silico\\u003c/em\\u003e based approach that involves proteome analysis to sort out the essential proteins of a pathogen as unique without creating any disturbance such as destruction of host proteome or changing functions of host proteome, newer, precise and recent drugs against the host pathogen are designed by proper cross-checking of pathogen proteome with host proteome to avoid any toxicity in the host. A protein is said to be a good target if that protein is crucial and essential for the existence and survival of pathogens and in the absence of this protein, the life of pathogen is compromised and accommodated. We used subtractive proteomics approach to recognize and discover new beneficial targets against \\u003cem\\u003eS.pneumoniae\\u003c/em\\u003e D39. This approach is reported in a number of published literatures for the identification, characterization and awareness of unique possible targets in human pathogens [\\u003cspan citationid=\\\"CR19\\\" class=\\\"CitationRef\\\"\\u003e19\\u003c/span\\u003e].\\u003c/p\\u003e \\u003cp\\u003eThe number of output protein sequences from first to last step of the subtractive proteomics approach is illustrated in precise detail (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e).\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003cdiv id=\\\"Sec13\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.1. Identification of Non-paralogous Sequences:\\u003c/h2\\u003e \\u003cp\\u003eThe complete proteome of \\u003cem\\u003eS. pneumoniae D39\\u003c/em\\u003e strain was retrieved from UniProtKB in FASTA canonical format and downloaded an Excel file of the proteome. The file of proteome was containing total 1947 protein sequences of genome. The main goal of this study is to attain essential, non-homologous unique proteins of pathogen. The reason behind obtaining these proteins is that we have to gain only those crucial and important proteins of pathogen which are vital and typically necessary for the life of the pathogen and that proteins should be absent in the host proteome. This would be helpful in the development of drug targets for a particular strain. That is why essential proteins of a pathogen must be identified and investigated and is important to disturb and distract the bacterial survival in the host. After the proteome retrieval, we proceeded toward the next step. It was important to remove paralogous sequences from the genome of the pathogen in order to attack a specific target developing of drugs and vaccines. For this purpose, CD-HIT tool was used after retrieving the complete proteome of pathogen from UniProtKB. After the removal of paralogous sequences, the result of CD-HIT returned 1887 proteins out of 1947 proteins. The pathogen was containing 60 paralogous sequences in its proteome file that were removed easily by CD-HIT tool. The proteome file obtained after utilizing CD-HIT tool was totally free from paralogous sequences. And this non-paralogous dataset was further used in the next steps.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec14\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.2. Determination and Identification of Non-Homologous Protein Sequences in the Pathogen:\\u003c/h2\\u003e \\u003cp\\u003eThere should be a chance that the proteins of host and pathogen are homologous and can be present in both of them; therefore, it is very important to identify them and to remove these hosts homologous protein sequences from the pathogen proteome so that host cell toxicity is restricted. The other advantages this step providing us is securing the proteins of pathogen we are targeting and attacking. Proteins gained after the removal of paralogous sequences from CD-HIT result were then used in Blastp by selecting H. sapiens in the box of organism and RefSeq database with an expectation value (E-value)\\u0026thinsp;=\\u0026thinsp;10\\u003csup\\u003e\\u0026minus;\\u0026thinsp;3\\u003c/sup\\u003e and resulted in the determination and identification of sequences that were specifically present in the pathogen only not in the host. In this way, we found 1423 Non-homologous sequences which were belonging to only pathogen after compiling the results of Blastp. The interested protein sequences in drug targeting and designing were Non-homologous sequences (without an asterisk) as they were only belonging to the pathogen while the Homologous protein (along an asterisk) sequences were present in both host and pathogen.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec15\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.3. Identification of Essential Proteins:\\u003c/h2\\u003e \\u003cp\\u003eDEG Databases was used to identify essential proteins of the pathogen that play important and crucial role in the survival and existence of this pathogen [\\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e12\\u003c/span\\u003e]. To find and know about essential proteins, each non-homologous protein of pathogen that was obtained from the previous step was searched by using its gene name in DEG and j browser one by one. 372 essential proteins were identified by DEG out of 1423 non-host protein sequences of \\u003cem\\u003eS. pneumonia D39\\u003c/em\\u003e strain that is 19.10% of the total whole proteome. One of the searched results for a gene \\u0026ldquo;\\u003cem\\u003emurE\\u003c/em\\u003e\\u0026rdquo; in DEG.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec16\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.4. Metabolic Pathways Analysis by using KEGG:\\u003c/h2\\u003e \\u003cp\\u003eThe resulted 372 essential proteins obtained from DEG databases were then subjected to the KEGG Pathway Database to identify the main and important biological pathways of bacteria in which these essential proteins work and contribute KEGG Pathway Database results are shown in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e .\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003cp\\u003eIn detail, 372 essential protein sequences of pathogen were subjected and checked in the KEGG Pathway Database by selecting the organism as \\u003cem\\u003eS.pneumoniae D39\\u003c/em\\u003e. Out of total 372 essential proteins, only 28 proteins were established to be unique in \\u003cem\\u003eS.pneumoniae D39\\u003c/em\\u003e from that 5 essential proteins were found as a part of peptidoglycan biosynthesis, 11 essential proteins were found to be involved in two component system,3 proteins were involved in Vancomycin resistance, 1 was involved in Microbial metabolism in diverse environments systems, 1 proteins was involved in Antigen nucleotide biosynthesis, 4 proteins were involved in Quorum sensing, 6 proteins were involved in Biosynthesis of Secondary metabolites, 3 proteins were involved in Cationic antimicrobial peptide (CAMP) resistance. Essential functions performed by these proteins in the pathogenic pathways are shown in Fig.\\u0026nbsp;10.\\u003c/p\\u003e \\u003cp\\u003eAt the end of this step, we obtained two different types of data from the results of KEGG pathway analysis that included: unique pathways and common pathways. The result of this comparison, on the basis of pathway number and pathway name, produced three types of datasets, one was of unique proteins when no hits found in KEGG, second was when both the host and pathogen were containing the same proteins called common proteins/pathways and the last one was when the protein was only present in the pathogen not in the human, were known as unique pathways. Unique proteins have their vital role in unique pathways.\\u003c/p\\u003e \\u003cp\\u003eThe unique pathway data was encouraging and helpful in finding the suitable and possible drug targets because it consisted of pathways that were specific to the pathogen only.\\u003c/p\\u003e \\u003cp\\u003e \\u003cb\\u003eTable.1. Unique Pathways of Bacteria.\\u003c/b\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"No\\\" id=\\\"Taba\\\" border=\\\"1\\\"\\u003e \\u003ccolgroup cols=\\\"2\\\"\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e \\u003cthead\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eSr.no.\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ePathways\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003c/thead\\u003e \\u003ctbody\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e1\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ePeptidoglycan Biosynthesis\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e2\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eLysine Biosynthesis\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e3\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eLipopolysaccharide Biosynthesis\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e4\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eBiosynthesis of Secondary Metabolites\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e5\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eMicrobial Metabolism in Diverse Environments\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e6\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ePhosphoTransferase System (PTS)\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e7\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eBeta-Lactam Resistance\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e8\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eQuorum Sensing\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e9\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eBacterial Secretion System\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003e10\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eTwo-Component System\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e11\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eMethane Metabolism\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e12\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eCationic Antimicrobial Peptide (CAMP) Resistance\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e13\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ePyrimidine metabolism\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e14\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eXylene degradation\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e15\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eBenzoate degradation\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003c/tbody\\u003e \\u003c/colgroup\\u003e \\u003c/table\\u003e\\u003c/div\\u003e \\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec17\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.5. Analysis and Exploration of Localization Results:\\u003c/h2\\u003e \\u003cp\\u003eThe protein should be in proper place within the living cell for its appropriate and regular work. The identification and analysis for the localization of target proteins was important because of the drug requires attaching with the target in order to show its action and the target present in the cell helps finally to make proper drug compound. Proteins should be at the suitable and proper section of the cell for carrying out their specific function at specific location. Subcellular localization is one of the typical characteristics for drug targeting. The unique and non-homologous proteins obtained from PSORTb v3.0 were used for the determination of their subcellular localization 28 essential, unique and non-homologous proteins gained from KEGG Pathway were subjected into PSORTb to determine their subcellular localization.\\u003c/p\\u003e \\u003cp\\u003eThe results obtained from PSORTb showed that 71% proteins were in cytoplasmic region of the cell, 25% proteins were found to be in integral plasma membrane proteins and 4% present in extracellular space. Figure\\u0026nbsp;\\u003cspan refid=\\\"Fig3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e shows the graphical representation of outcomes obtained from PSORTb.\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec18\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.6. Analysis of Functional Family of Essential, Non-homologous and Unique Proteins of Pathogen:\\u003c/h2\\u003e \\u003cp\\u003eInterproscan was used for the prediction of the functional family classification of the non-homologous, unique and essential protein sequences of \\u003cem\\u003eS.pneumoniae D39\\u003c/em\\u003e strain (Zdobnov \\u0026amp; Apweiler, 2001). The Twenty-eight unique protein sequences were pasted in Interproscan to get the family prediction of each protein one by one for these 28 query sequences, many functional classes were predicted. The results obtained from Interproscan with the possible classification of functional families of some essential and unique short-listed proteins are illustrated in table (2).\\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab1\\\" border=\\\"1\\\"\\u003e \\u003ccaption language=\\\"En\\\"\\u003e \\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 2\\u003c/div\\u003e \\u003cdiv class=\\\"CaptionContent\\\"\\u003e \\u003cp\\u003eAnnotation of Function of Proteins of Pathogen\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/caption\\u003e \\u003ccolgroup cols=\\\"3\\\"\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e \\u003cthead\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eS.r. number\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eEntry ID\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eProtein family names\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003c/thead\\u003e \\u003ctbody\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e1\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZMF9\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTranspeptidase-like\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e2\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eQ04JL8\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eLigases\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e3\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZPR0\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003ephosphonate metabolism\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e4\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZNR2\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eFirmicutes\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e5\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZLN7\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTranscriptional regulatory protein WalR-like\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e6\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZLY7\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eD-alanyl carrier protein\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e7\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZMU1\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eDiacylglycerol kinase like\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e8\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZMP6\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eSUF system FeS cluster assembly\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e9\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eQ04MC4\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eGlycosyl transferase, family 4\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e10\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eQ04LK0\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003etransferase family\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e11\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZP52\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eHPr-like superfamily\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e12\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZQA1\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eRNA polymerase sigma factor\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e13\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZPN8\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eAcyl carrier protein family\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e14\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZR18\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNitroreductase\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e15\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZMW4\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eAll DNA binding proteins\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e16\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZNP9\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003enucleoside triphosphate hydrolase like\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e17\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZMY1\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eOxidoreductase\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e18\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZML1\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eClass I glutamine aminotransferases-like\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e19\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eA0A0H2ZNH9\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eATPase superfamily\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e20\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eQ04N63\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003ereplication initiator\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003c/tbody\\u003e \\u003c/colgroup\\u003e \\u003c/table\\u003e\\u003c/div\\u003e \\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec19\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.7. Analysis of Drug ability Potential of Unique Proteins of Pathogen:\\u003c/h2\\u003e \\u003cp\\u003eThe identification and analysis of druggable proteins of \\u003cem\\u003eS. pneumonia\\u003c/em\\u003e D39 strain was carried out in this last step by using Drug Bank Database. For this reason, the 28 shortlisted unique, essential and non-homologous protein names were searched in Drug Bank Database one by one to identify druggability of these proteins. For possible drug target it showed the identification of 5 proteins that can be useful and they integrated to show similarity with the previous existing targets of drugs in the Drug Bank database [\\u003cspan citationid=\\\"CR21\\\" class=\\\"CitationRef\\\"\\u003e21\\u003c/span\\u003e]. All these 5 druggable proteins were found to be in different and specific locations of the cell. 3 out of 5 proteins were present in cytoplasmic region of the cell whereas, 2 proteins were present in the plasma membrane The total obtained 5 druggable target proteins were involved in different and highly specific metabolic pathways that were totally essential for the under studied pathogen \\u003cem\\u003eS. pneumoniae D39\\u003c/em\\u003e strain.\\u003c/p\\u003e \\u003cp\\u003eStrain specific drug targeting the non-homologous essential proteins that involved in unique metabolic pathways ensure the eradication of invasive and non-invasive diseases cause by \\u003cem\\u003eS.pneumoniae\\u003c/em\\u003e with fewer side effects to the host. As we know that cytoplasmic proteins are considered more suitable one for Drug target as it is difficult to purify the proteins that are present in the membrane of the cell [\\u003cspan citationid=\\\"CR20\\\" class=\\\"CitationRef\\\"\\u003e20\\u003c/span\\u003e]. Moreover, surface proteins have direct contact so these protein targets have ability to obtain immune response when defined to the host\\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab2\\\" border=\\\"1\\\"\\u003e \\u003ccaption language=\\\"En\\\"\\u003e \\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 3\\u003c/div\\u003e \\u003cdiv class=\\\"CaptionContent\\\"\\u003e \\u003cp\\u003eInformation about 3 Druggable Proteins of \\u003cem\\u003eS. pneumoniae\\u003c/em\\u003e D39 involved in Cytoplasmic region of the cell\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/caption\\u003e \\u003ccolgroup cols=\\\"3\\\"\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e \\u003cthead\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eGene Name\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eSubcellular localization\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eDrugBank ID\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003c/thead\\u003e \\u003ctbody\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003eciaR\\u003c/em\\u003e SPD_0701\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eCytoplasmic\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eDB01972\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003ePtsH\\u003c/em\\u003e SPD_1040\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eCytoplasmic\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eDB01899\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003ePtsl\\u003c/em\\u003e SPD_1039\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eCytoplasmic\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e\\u003cem\\u003eDB08357\\u003c/em\\u003e\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003c/tbody\\u003e \\u003c/colgroup\\u003e \\u003c/table\\u003e\\u003c/div\\u003e \\u003c/p\\u003e \\u003cp\\u003eAll the 3 proteins can be regarded as a potent Drug target because all have one homologous in Drug Bank database having at least 30% sequence similarity and serve as a promising candidate to fight against the infectious diseases caused by Streptococcus pneumoniae D39 (Table\\u0026nbsp;\\u003cspan refid=\\\"Tab2\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e).\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec20\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.8. Analysis of protein-Multiple ligand docking:\\u003c/h2\\u003e \\u003cp\\u003eStructural based virtual screening also called as target based virtual screening that is used to find out the interaction between the receptor and ligands (Lyne, 2002). As results all the ligands were ranked according to their binding affinity with the drug target. Scoring function is very important in docking as it is responsible for the prediction of binding affinity. Thus, scoring function results are the main reason for the failure or success of docking. The more positive a binding affinity better the ligand has been bound to the protein so for the first drug target that is \\u003cem\\u003eciaR\\u003c/em\\u003e SPD_0701 shown the maximum binding affinity with the Piperacillin anhydrous having the best binding affinity and lower Root mean square deviation (RMSD) (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig4\\\" class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003e)\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003cp\\u003eFor the second drug target that is \\u003cem\\u003eptsH\\u003c/em\\u003e SPD_1040 shows the maximum binding affinity with penicillin G (act as a ligand) and yellow color shows the binding of penicillin G with the drug target. (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig5\\\" class=\\\"InternalRef\\\"\\u003e5\\u003c/span\\u003e)\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003c/div\\u003e\"},{\"header\":\"4. Conclusion\",\"content\":\"\\u003cp\\u003eThe genome sequencing technology at a large-scale has been improved. The world of drug-designing has also played a revolutionary role in the domains of microbiology and pathology. As compared to the earlier decades, the genomic and proteomic data is also increasing day by day, hence, creating a massive increase in the biological databases. The researchers find the successful treatments against the infections caused by those pathogens whose sequence information has been reported. The novel pathogenic strains with high resistance against antibacterial agents are a great threat towards the management of infections and diseases.\\u003c/p\\u003e \\u003cp\\u003eDue to a significant role of Bioinformatics, the researchers become enabled to identify the novel drug targets against the highly resistant bacterial strains which have been emerged currently. The approach of subtractive proteomics has gained high popularity in the field of drug-designing against human pathogens. This method enables the researchers for rapid identification of potential targets in pathogens and targeting them as treatment for that particular infection.\\u003c/p\\u003e \\u003cp\\u003eThis study is aimed to use the latest and updated databases of proteins, advanced computational tools and subtractive proteomics approach to characterize the novel drug targets against the pathogenic proteins for treatment purpose. To identify the putative drug targets in any pathogen requires a reliable, short-term, fast, valuable and inexpensive approach that not only predicts drug targets but also predicts multiple drug resistant abilities of pathogens by using computational based tools to find subtraction in their whole genome. Similarly, subtractive proteomics approach has been used for the last decades to design highly specific drug targets in opposition to certain bacterial diseases or human pathogens for decreasing partially or completely stops the pathogenicity in the host. By using this \\u003cem\\u003ein silico\\u003c/em\\u003e based modern hierarchical approach, the use of conventional approaches of drug development has reduced to a maximum level. This study was also managed by using subtractive proteomics approach and 3 non-homologous, unique and essential targets out of 1947 proteins were shortlisted in \\u003cem\\u003eS. pneumoniae\\u003c/em\\u003e D39 strain only. While this approach has many advantages but it also has a little disadvantage due to some predictive tools used in this approach. The work of all of the tools used in this modern approach is based on prediction and algorithms. The errors may occur in these tools as the algorithms for these tools have been designed and coded by human and their specificity is limited for a specific task. Basically, that was the big limitation of this modern computational based approach.\\u003c/p\\u003e\"},{\"header\":\"References\",\"content\":\"\\u003col\\u003e\\u003cli\\u003e\\u003cspan\\u003eBarry M. Gray, George M. Converse, III, Hugh C. Dillon, Jr., Epidemiologic Studies of \\u003cem\\u003eStreptococcus pneumoniae\\u003c/em\\u003e in Infants: Acquisition, Carriage, and Infection during the First 24 Months of Life, \\u003cem\\u003eThe Journal of Infectious Diseases\\u003c/em\\u003e, Volume\\u0026nbsp;142, Issue 6, December 1980, Pages 923\\u0026ndash;933, \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1093/infdis/142.6.923\\u003c/span\\u003e\\u003cspan address=\\\"10.1093/infdis/142.6.923\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eDaniel M. Musher, Infections Caused by \\u003cem\\u003eStreptococcus pneumoniae\\u003c/em\\u003e: Clinical Spectrum, Pathogenesis, Immunity, and Treatment, \\u003cem\\u003eClinical Infectious Diseases\\u003c/em\\u003e, Volume\\u0026nbsp;14, Issue 4, April 1992, Pages 801\\u0026ndash;809, \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1093/clinids/14.4.801\\u003c/span\\u003e\\u003cspan address=\\\"10.1093/clinids/14.4.801\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eNancy Y. Yu, James R. Wagner, Matthew R. Laird, Gabor Melli, S\\u0026eacute;bastien Rey, Raymond Lo, Phuong Dao, S. Cenk Sahinalp, Martin Ester, Leonard J. Foster, Fiona S. L. Brinkman, PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes, \\u003cem\\u003eBioinformatics\\u003c/em\\u003e, Volume\\u0026nbsp;26, Issue 13, 1 July 2010, Pages 1608\\u0026ndash;1615, \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1093/bioinformatics/btq249\\u003c/span\\u003e\\u003cspan address=\\\"10.1093/bioinformatics/btq249\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eNancy Y. Yu, James R. Wagner, Matthew R. Laird, Gabor Melli, S\\u0026eacute;bastien Rey, Raymond Lo, Phuong Dao, S. Cenk Sahinalp, Martin Ester, Leonard J. Foster, Fiona S. L. Brinkman, PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes, \\u003cem\\u003eBioinformatics\\u003c/em\\u003e, Volume\\u0026nbsp;26, Issue 13, 1 July 2010, Pages 1608\\u0026ndash;1615, \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1093/bioinformatics/btq249\\u003c/span\\u003e\\u003cspan address=\\\"10.1093/bioinformatics/btq249\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eBarh, D., Tiwari, S., Jain, N., Ali, A., Santos, A. R., Misra, A. N., \\u0026hellip; Kumar, A.(2010). \\u003cem\\u003eIn silico subtractive genomics for target identification in human bacterial pathogens.Drug Development Research, 72(2), 162\\u0026ndash;177.\\u003c/em\\u003e doi:10.1002/ddr.20413\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eKatherine L O'Brien, Lara J Wolfson, James P Watt, Emily Henkle, Maria Deloria-Knoll, Natalie McCall, Ellen Lee, Kim Mulholland, Orin S Levine, Thomas Cherian,Burden of disease caused by Streptococcus pneumoniae in children younger than 5 years: global estimates,The Lancet, Volume\\u0026nbsp;374, Issue 9693, 2009,Pages 893\\u0026ndash;902,ISSN 0140\\u0026ndash;6736, \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/S0140-6736(09)61204-6\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/S0140-6736(09)61204-6\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eRuch, P., Teodoro, D., Consortium, U., \\u0026amp; others. (2021). \\u003cem\\u003eUniProt\\u003c/em\\u003e. S\\u0026aacute;ez-Llorens, X., \\u0026amp; McCracken Jr, G. H. (2003). Bacterial meningitis in children. \\u003cem\\u003eThe Lancet 361\\u003c/em\\u003e(9375), 2139\\u0026ndash;2148.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eWeizhong Li, Adam Godzik, Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences, \\u003cem\\u003eBioinformatics\\u003c/em\\u003e, Volume\\u0026nbsp;22, Issue 13, 1 July 2006, Pages 1658\\u0026ndash;1659, \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1093/bioinformatics/btl158\\u003c/span\\u003e\\u003cspan address=\\\"10.1093/bioinformatics/btl158\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eLavigne, R., Seto, D., Mahadevan, P., Ackermann, H.-W., \\u0026amp; Kropinski, A. M.(2008). Unifying classical and molecular taxonomic classification: analysis of the Podoviridae using BLASTP-based tools. \\u003cem\\u003eResearch in Microbiology\\u003c/em\\u003e, \\u003cem\\u003e159\\u003c/em\\u003e(5), 406\\u0026ndash;414.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eZhang, R., \\u0026amp; Lin, Y. (2009). DEG 5.0, a database of essential genes in both prokaryotes and eukaryotes. \\u003cem\\u003eNucleic Acids Research\\u003c/em\\u003e, \\u003cem\\u003e37\\u003c/em\\u003e(suppl\\\\_1), D455\\u0026ndash;D458.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eLuo, H., Lin, Y., Liu, T., Lai, F.-L., Zhang, C.-T., Gao, F., \\u0026amp; Zhang, R. (2021). DEG 15, an update of the Database of Essential Genes that includes built-in analysis tools. \\u003cem\\u003eNucleic Acids Research\\u003c/em\\u003e, \\u003cem\\u003e49\\u003c/em\\u003e(D1), D677\\u0026ndash;D686.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSlager, J., Aprianto, R., \\u0026amp; Veening, J.-W. (2018). Deep genome annotation of the opportunistic human pathogen Streptococcus pneumoniae D39. \\u003cem\\u003eNucleic Acids Research\\u003c/em\\u003e, \\u003cem\\u003e46\\u003c/em\\u003e(19), 9971\\u0026ndash;9989.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eKanehisa, M., \\u0026amp; Goto, S. (2000). KEGG: kyoto encyclopedia of genes and genomes. \\u003cem\\u003eNucleic Acids Research\\u003c/em\\u003e, \\u003cem\\u003e28\\u003c/em\\u003e(1), 27\\u0026ndash;30.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eYu, N. Y., Wagner, J. R., Laird, M. R., Melli, G., Rey, S., Lo, R., Dao, P., Sahinalp, S. C., Ester, M., Foster, L. J., \\u0026amp; others. (2010). PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes. \\u003cem\\u003eBioinformatics\\u003c/em\\u003e, \\u003cem\\u003e26\\u003c/em\\u003e(13), 1608\\u0026ndash;1615.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eQuevillon, E., Silventoinen, V., Pillai, S., Harte, N., Mulder, N., Apweiler, R., \\u0026amp; Lopez, R. (2005). InterProScan: protein domains identifier. \\u003cem\\u003eNucleic Acids Research\\u003c/em\\u003e, \\u003cem\\u003e33\\u003c/em\\u003e(suppl\\\\_2), W116\\u0026ndash;W120.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSood, N., Ribero, R., Ryan, M., \\u0026amp; Van Nuys, K. (2020). The association between drug rebates and list prices. \\u003cem\\u003eLos Angeles: University of Southern California Leonard D. Schaeffer Center for Health Policy and Economics\\u003c/em\\u003e, \\u003cem\\u003e500\\u003c/em\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eMaia, E. H. B., Assis, L. C., de Oliveira, T. A., da Silva, A. M., \\u0026amp; Taranto, A. G. (2020). Structure-based virtual screening: from classical to artificial intelligence. \\u003cem\\u003eFrontiers in Chemistry\\u003c/em\\u003e, \\u003cem\\u003e8\\u003c/em\\u003e, 343.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eIbrahim, M. T., Uzairu, A., Shallangwa, G. A., \\u0026amp; Uba, S. (2020). In-silico activity prediction and docking studies of some 2, 9-disubstituted 8-phenylthio/phenylsulfinyl-9h-purine derivatives as Anti-proliferative agents. \\u003cem\\u003eHeliyon\\u003c/em\\u003e, \\u003cem\\u003e6\\u003c/em\\u003e(1), e03158.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eWadhwani, A., \\u0026amp; Khanna, V. (2016). In silico identification of novel potential vaccine candidates in Streptococcus pneumoniae. \\u003cem\\u003eGlobal J Technol Optim\\u003c/em\\u003e, \\u003cem\\u003e7\\u003c/em\\u003e(109), 2.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eShahid, F., Ashfaq, U. A., Saeed, S., Munir, S., Almatroudi, A., \\u0026amp; Khurshid, M. (2020). In Silico Subtractive Proteomics Approach for Identification of Potential Drug Targets in Staphylococcus saprophyticus. \\u003cem\\u003eInternational Journal of Environmental Research and Public Health\\u003c/em\\u003e, \\u003cem\\u003e17\\u003c/em\\u003e(10), 3644.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eUddin, R., Arif, A., Zahra, N.-A., \\u0026amp; Sufian, M. (2021). Comparative proteome-wide study for in-silico identification and characterization of indispensable hypothetical proteins of food bornepathogen Campylobacter jejuni (CJJ) by subtractive genomics approach. \\u003cem\\u003ePakistan Journal of Pharmaceutical Sciences\\u003c/em\\u003e, \\u003cem\\u003e34\\u003c/em\\u003e(4).\\u003c/span\\u003e\\u003c/li\\u003e\\u003c/ol\\u003e\"}],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":true,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":true,\"hideJournal\":true,\"highlight\":\"\",\"institution\":\"PDM university, Bahadurgarh\",\"isAcceptedByJournal\":false,\"isAuthorSuppliedPdf\":false,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":false,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true},\"keywords\":\"S. pneumonia D39, Draggability, SGA, In-silico, Characterization\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-2267778/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-2267778/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"\\u003cp\\u003eAs per WHO, the \\u003cem\\u003epneumococcus\\u003c/em\\u003e causes one million fatalities each year because of their underdeveloped immune systems, children are the most vulnerable to pneumococcal infections. The rise of \\u003cem\\u003eS. pneumoniae\\u003c/em\\u003e resistance to antibiotics is causing widespread alarm all across the globe. Since the last couple of years, a recently developed technique is being used to overcome resistant pathogens. One of these is the computational subtractive genomics technique, in which the bacterial pathogen full proteins are effectively decreased to a limited amount of probable therapeutic targets. The procedures employed in this strategy are to locate human non-homologs targets, proteins that are vital to the illness producing agent, and participation of the identified proteins in pathogen metabolic pathways that are required for bacterial survival. In this work, we applied computational subtractive genomics on the proteins of the S. pneumonia strain D39 and came up with three cytoplasmic proteins that can serve as a promising candidate for novel drug to control the pathogenicity caused by \\u003cem\\u003eS. pneumoniae.\\u003c/em\\u003e\\u003c/p\\u003e\",\"manuscriptTitle\":\"Subtractive Genome Analysis for Identification and Characterization of Novel Drug Targets in the Streptococcus pneumoniae by Using In-silico Approach\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2022-11-16 17:26:42\",\"doi\":\"10.21203/rs.3.rs-2267778/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true}}],\"origin\":\"\",\"ownerIdentity\":\"6c6d0fab-806e-483a-ac07-a03812ab0bbf\",\"owner\":[],\"postedDate\":\"November 16th, 2022\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"posted\",\"subjectAreas\":[{\"id\":16989859,\"name\":\"General Microbiology\"},{\"id\":16989860,\"name\":\"Bioinformatics\"}],\"tags\":[],\"updatedAt\":\"2022-11-16T17:26:43+00:00\",\"versionOfRecord\":[],\"versionCreatedAt\":\"2022-11-16 17:26:42\",\"video\":\"\",\"vorDoi\":\"\",\"vorDoiUrl\":\"\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-2267778\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-2267778\",\"identity\":\"rs-2267778\",\"version\":[\"v1\"]},\"buildId\":\"7rjqhiLT3MXkJMwkYKINL\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}