Generative Deep Neural Networks for Estimating Hypervariability in Hepatitis B and C Virus Genomes | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Generative Deep Neural Networks for Estimating Hypervariability in Hepatitis B and C Virus Genomes Sharmeen Saqib, Zilwa Mumtaz, Hania Ahmed, Ashiq Ali, Obaidullah Qazi, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5560102/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Hepatitis B virus (HBV) and Hepatitis C virus (HCV) have always remained a greater global concern. Approximately 1.3 million deaths occur each year due to HBV and HCV. Due to the diverse genotypes and drug resistance, diagnostic challenges are being faced to treat these viruses. Therefore, the success ratio of the antiviral therapies has been decreasing with time in the last few decades. By deep learning predictive model, the pattern of evolution in hypervariable regions of HBV and HCV genes can be foreseen. In HCV, the hypervariable region is the Envelope glycoprotein (E2) gene, while in HBV, it includes the S1 and S2 genes. Generative models in deep learning have been used for evolutionary studies, but the application of these models is limited in viral research for predicting the evolving genotypes of viruses. The Long Short-Term Memory (LSTM) model represented a satisfactory outcome in predicting the sequences of the hypervariable genes of the evolving genotypes of the HCV and HBV genes that might be of a great help in diagnosis and vaccine design. We collected data from databases like NCBI and BVBRC. Our proposed LSTM generative model was trained on 1500 sequences of hypervariable genes of the present 7 genotypes of Hepatitis C and 10 genotypes of HBV. Apart from the traditional generative models like Recurrent Neural Network (RNN), our model not only generates the sequence but also learns and develops the relationship between various parts of the virus’s genetic code. In this study, three generative models were compared, Simple RNN, 1-Dimensional Convolutional Neural Network (ConV1d) and Long Short-Term Memory (LSTM). Among these three, LSTM demonstrated the least error rate with the highest efficiency and accuracy. While simple RNN and ConV1d illustrated relatively higher error rate and lower accuracy. LSTM gained success in reading long dependencies, hence, the proposed LSTM models are efficient at handling the sequential data along with preventing the conventional issue of losing the important information from the data, which happens frequently in generative models like Simple RNN and ConV1d. LSTM Generative model RNN Hepatitis B and Hepatitis C Evolutionary studies Sequence generation ConV1d Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Introduction Hepatitis B virus (HBV) and Hepatitis C virus (HCV) have always been considered the major health concern around the world. Among all the cancers, Liver cancer ranks third with approximately 1.3 million deaths each year by Hepatitis B virus (HBV) and Hepatitis C virus (HCV) (Rumgay et al, 2022 ). Presently, 70–80% of hepatocellular carcinoma (HCC) lead to liver cancer. Both HBV and HCV have played a significant role in the etiological traits of HCC. In all cases of HCV related to HCC, cirrhosis is present. Whereas, in HBV cirrhosis is absent in majority of cases (Sung et al, 2021 ). HCV is a single stranded RNA enveloped virus with the genome size of 9.6 kb, flanked by 5` and 3` untranslated regions. Over three thousand amino acids are transcribed into a single polyprotein, which is then broken down into structural and nonstructural proteins (Liang et al., 2021 ). HCV is the primary source of liver cirrhosis that leads to hepatocellular carcinoma (HCC) with a global decline in survival and an increase in mortality (Rich, 2024 ). Hepatic cancer can be prevented by treating the patient with newly developed antivirals for HCV. Higher death rate is reported after liver transplantation in HCV-positive recipients, and allograft rejection is caused by HCV infection. Knowing about the virus and host interaction can help in the management and prevention of HCV infection. HCV has been found to infect only humans and chimpanzees within the restricted host range (Stroffolini & Stroffolini, 2024 ). Hepatitis B is a DNA enveloped virus that belongs to the Hepadnaviridae family of viruses. The size of the genome is 3kb which is wrapped in an envelope known as nucleocapsid. The genome of 90–95% viral particles is present in relaxed circular DNA (rcDNA) that is occupied by the unified overlap sequence. There are 334 amino acids present in the genome of HBV (Zoulim Et al, 2024 ). Lamivudine (LAM) was introduced in 1995, for the treatment of HBV and it developed the resistance among 75% of the patients after 5 years of the treatment and gave rise to a resistant strain of HBV. For treating these resistant strains of HBV, Adefovir (ADV) was introduced in 2002. In 2008, ADV was replaced by a more effective drug known as Tenofovir (TDF). The latest drug that is approved for treating HBV is tenofovir Alafenamide (TAF) in 2015 and it can also be used for patients with other issues like bone, heart, kidneys and lungs (Broquetas et al, 2023). Deep learning (DL) is a strong tool that is used in the field of genomic analysis, specifically when it comes to shifting through complex genomic data to isolate features and patterns. Likewise, the machine learning has also become popular when it comes to dealing with infectious diseases like HCV, HBV, Dengue, Coronavirus etc. using epidemiological and clinical data for developing a predictive model. Viral genome classification can be done in a variety of ways using machine learning and alignment techniques (Ningthoujam et al., 2024 ). Detection of the viral sequences are carried out by tools like USEARCH, REGA and SCUEAL, that rely on the score of alignment for classification of genome. Although these methods have some constraints, particularly their execution depends upon the initial alignment that is selected (Hanke et al., 2024 ). However, there are numerous techniques that were proposed before the machine learning approach was used for the classification of viral genome sequences. But when it comes to identifying the viral genome contig and obtaining the useful hidden data these techniques have the limitations and face difficulty in executing the task. Moreover, they were trained on nucleotide sequence data that further restricts their usefulness for genomic research. In the natural language processing areas, Recurrent Neural Network (RNN’s) a deep learning model has been proven to be effective. Deep learning applications in computational biology mostly focus on genome sequencing and analysis (Padminivalli et al., 2023 ). However, because of the model's intricacy in deciphering its concealed state, RNNs are regarded as black boxes. Currently, for many clinical classifications task machine learning have been used. Deep learning is the subgroup of machine learning that nearly mimics the structure of the neurons. Clinical decision-making is enhanced by radiomics, a non-invasive technique that extracts a multitude of medical imaging features (both hand-crafted and deep-learning features). It is important for the medical practitioners to use the quantified scores along with the other clinical symptoms to improve the model while they evaluate lesions. Convolutional Neural Network (CNN) has been proved to be effective for interpreting the real-world picture, these are some major tools that have been used in majority of these techniques, and they only concentrate on deep learning or classic radiomics (Duan et al., 2022 ). For many different RNA/ DNA- based prediction tasks, convolutional neural networks (CNNs) and long short-term memory architectures (LSTMs) are proved to be effective. Previous studies mainly emphasized the regulation of human gene expression, and this topic is now being actively studied. Deep learning models have been directly trained on RNA/ DNA sequences in the field of genomic pathogens to envisage the infective potential of new pathogens, as well as the host ranges of three multi-host virus species. While ViraMiner and DeepVirFinder may identify viral sequences in metagenomic materials, they are limited to previously identified species and are unable to predict the host (Bartoszewicz et al., 2021 ). Viral genome sequences have been detected using Long Short-Term Memory (LSTM) and Convolutional Neural Networks (CNN) by training the pattern and frequency of branching. On the other hand, one kind of CNN model that excels in analyzing patterns in images while considering the inherent heterogeneity in the data is the Feedforward Neural Network (FNN). This network is trained using a backpropagation algorithm and is intended to translate fixed-length inputs to a fixed-size output. CNN, which consists of several layers, efficiently learns the complex relationship between input and output and stores and updates information in filter weights. The main objective of this study is to predict the sequences of the hypervariable genes of HBV (S1 and S2) and HCV envelop protein (E2) of evolving genotypes to have deep knowledge and better understanding about the evolution of HBV and HCV. Instead of expressing character meanings, the model learns patterns in tokens by using simple Recurrent Neural Networks (Simple RNNs) trained on the hypervariable gene sequences of already present genotypes. When it comes to processing sequential data and addressing the vanishing gradient issue that standard RNNs have, LSTM is the preferred option for genomic prediction tasks (Mumtaz et al, 2024 ). We did the comparison between the three deep learning generative models (Simple RNN, 1-dimensional Convolutional Neural Network and Long-Short Term Memory). The focus of this study was that the sequence generation of the hypervariable regions of HBV and HCV these genes are most mutated genes and play a significant role in changing the genotype of the virus. When we get the sequence of hypervariable gene of evolving genotype it would be much easier to develop the strain specific drug/vaccine even before the emergence of the evolving genotype. Material and Methods Data Collection and pre-processing: A data set consisting of a total of 1500 sequences of hypervariable genes of Hepatitis B Virus (HBV) S1 and S2 gene. And 1500 sequences of Hepatitis C Virus (HCV) hypervariable E2 gene of each genotype were retrieved from BV-BRC (Kent, 2002) and NCBI databases (NCBI, 1988). Preparation of the data was done before the analysis to make sure that the data is in the appropriate format using MEGA 11 software (Tamura, 2021). Hypervariable genes of HBV are S1 and S2. S1 is 835 base pairs while S2 is 333 base pairs. Whereas E2 of HCV is 1106 base pairs. The gene sequence data from each genotype of HBV and HCV was compiled in a separate file and labelled accordingly. Deep learning Model Architecture The generative deep learning model used the Bio python Seq IO library for loading the data from FASTA file containing HBV and HCV hypervariable gene sequences. The function load_fasta_file extracts and read sequences from the files including the sequences and subsequent virus types. It then shows all the sequences in the file along with the number of sequences (Zuvanov et al., 2021). After loading the data, the data is then prepared for training the deep learning model by tokenization, of the data at the character level and then producing the n gram input sequences. This tokenizer was used from the library known as TensorFlow Keras (Bhandari et al., 2022). By tokenizer each sequence present in the data is converted into integers, where each character is represented as an exclusive integer (Song et al., 2020 ). For standardizing the size of input data, padding was performed to match the length of all the sequences. All the data in the output is split into the predictor except the last character in the sequence and it labels that last character which then aids in training the data for model (Alrasheedi et al., 2023 ). Then the LSTM deep learning model was constructed using TensorFlow’s Keras library to predict the gene sequences of the evolving genotype. The architecture of this model consists of two dense embedded layers of Long Short-Term Memory (LSTM) designed to apprehend sequential dependencies in the data set (Sunny et al., 2020 ). The final layer is the thick dense layer, it helps in generating the probability distribution for the feasible characters in vocabulary, making it a suitable fit for multi-class categorization which generates probability distributions for all possible characters in the vocabulary, making it appropriate for multi-class classification. For compiling the model sparse categorical cross entropy is used as a loss function, and for optimizing the model weight Adam optimizer is used (Shahade et al., 2023 ). This model is designed to learn and then generate hypervariable gene sequences of the evolving genotype. The training of the model was done on nucleotide sequences using the fit function of TensorFlow’s Keras library. The training was done on 15 epochs, with the choice to fix the value based on data. This method of training guarantees effective training, sustaining the finest performance of the model along with eluding unnecessary overfitting in the data being trained (Martínez-Llop et al., 2023 ). For the generation of hypervariable gene sequences, the deep learning model is pre-trained. The seed text use the generate_sequence function to predict the successive character and builds the gene sequence found on the possibility that developed from output of the model. The performance of the trained deep learning model was evaluated by a sequence predicted by based on the tested dataset and producing the confusion matrix for assessing the accuracy of the generative model (Ferruz et al., 2023 ). The prediction of next character in the sequence used the probabilistic method by using the feature of randomness for enhancing diversity in the sequences that are generated. The generated sequences are produced by diverse nucleotides, and they are analyzed for genetic variations (Wang et al., 2023). Embedding layer transformation $$\:E\left(xt\right)=We\cdot\:1xt$$ LSTM update $$\:{h}_{t,}{c}_{t}=LSTM\left(\varkappa\:t,{h}_{t}-1s,{c}_{t}-1\right)$$ SoftMax Prediction $$\:P\left(\left.{y}_{t}\right|{h}_{t}\right)=\frac{{e}^{{W}_{0}}\cdot\:ht+{b}_{0}}{{\sum\:}_{k=1}{V}_{e}{W}_{0}\cdot\:ht+{b}_{0}}$$ Comparison of the other generative models: In this study, we compared the LSTM model with other generative models like Simple Recurrent Neural Network (Simple RNN), 1-Dimensional Convolutional Neural Network (ConV1d). For training all these models Adam Optimizer and sparse categorical cross entropy were used for generating sequences. Simple Recurrent Neural Network (Simple RNN) A neural network that is designed for processing sequential data by using the cycles in its design is known as Simple RNN. RNNs preserve an obscure state that depicts data about the prior inputs, allowing them to remember the previous dataset. For prediction of time series, language models, and speech recognition the memory mechanism of RNN is used. RNN can be laborious due to their long dependency problems and exploding gradients, that can impede their functioning when working for a longer period. Despite these challenges, innovative architectures like GRU and LSTM Simple RNN is still used as the base to tackle these constraints (Krauss, 2024 ). 1-Dimensional Convolution Neural Network (ConV1d) This deep learning model is used for generative and classification purposes based on the numerical data that is provided in the form of 1 dimensional sequence (Kareem et al, 2023 ). The input data is given to the model in the form of sequences and then it applies the chain of layers to produce a solid demonstration of the input given to the model and predict the sequence. The layers that are used in 1 dimensional generative model are input layers, convolutional layers, activation layers, pooling layers and fully connected layers (Choi et al, 2024 ). Results We worked on two other models other than our proposed model; these include Simple RNN and ConV1d for generating the hypervariable gene sequences of HBV and HCV. But LSTM, our proposed model proves to be the most efficient and accurate model for the generation of gene sequences. At first, we trained each model on fifteen epochs, as shown in Table 1 . ConV1D at fifteen epochs gave 25% accuracy and the length of the sequence generated for E2 is 654, S1 is 498 and S2 is 132 base pairs. But at twenty epochs its accuracy increased to 27% and the length of the sequence generated is 660 for E2, 510 for S1, and 190 for S2. The Simple RNN model gave 36% accuracy at fifteen epoch and 40% accuracy at twenty. The drawback of these two models was that these do not produce accurate sequence. In Simple RNN model there were some gaps that are seen in the generated sequence and ConV1d repeatedly generated the four characters in the same sequence with the gaps in between the sequence. The LSTM model, when trained at fifteen epochs, generated 1106 base pair sequence of hypervariable gene E2 from HCV genome and 835 base pairs of S1 and 333 base pairs of S2 hypervariable genes from HBV genome these are the complete sizes of the genes. Whereas the Simple RNN and ConV1d does not produce the complete sequences. The sequence length had a considerable impact on the predictive capabilities of the model and its biological importance. For further, increasing the accuracy of the model we trained the model by increasing the epochs from 15 to 20 and then checked the accuracy that appeared to increase the model’s accuracy along with the epoch. The LSTM model that was trained at 15 epochs gave 60% accuracy and at 20 epochs it showed 85% accuracy. Hence this proved that increasing the epoch increases the accuracy and the predictive abilities of the model. Table 1 Accuracy in percentage of the model on 15 and 20 Epochs, and length of the sequences they generated in base pairs. Model Epoch Accuracy (%) Sequence Generated (Base pairs) ConV1d 15 25 E2 = 654 S1 = 498 S2 = 132 20 27 E2 = 660 S1 = 510 S2 = 190 Simple RNN (Approximately) Gaps at the end of the sequence 15 36 E2 = 998 S1 = 734 S2 = 267 20 40 E2 = 1000 S1 = 789 S2 = 290 LSTM (Accurately) 15 60 E2 = 1106 S1 = 835 S2 = 333 20 85 E2 = 1106 S1 = 835 S2 = 333 We performed pairwise2 similarity on the sequences generated by the deep leaning generative model. The hypervariable gene sequences that are generated by our LSTM model have a greater accuracy with genotype A that is 64.52% of S1 HBV gene (Fig. 2 ) and with genotype A, B, C, D, E which is 61.56% of S2 of HBV. (Fig. 3 ) Whereas the model shows the highest accuracy rate with genotype 7 in case of E2 gene of HCV which is 65.01%. (Fig. 1 ) Results of the pairwise2 similarity demonstrated that the generated sequence was not closely related to the known genotypes of HBV and HCV. The sequences that were generated by the LSTM model were of perfect length of the specific gene size. Which proves that our model predicted the gene sequences completely based on the randomization without any biasness. The confusion matrix is the powerful tool for assessing the classification task, analyzing the categories of the output sequences and helps in performing the error analysis and ensuring the quality of the sequences. It is used for comparing the generated sequences with the data given to the model, providing intuitions of the false positive, false negatives, true positives and true negatives that evaluates the performance of the model. Figures 4 , 5 and 6 illustrates the confusion matrix for E2, S1 and S2 genes respectively. We check the loss and accuracy of the model. The direct relationship is observed between the accuracy and the epochs whereas, indirect relation is observed between the loss and epochs. Figure 7 represents the accuracy and loss of S2 gene of HBV. Figure 8 shows the loss accuracy for S1 gene of HBV and Fig. 9 demonstrates the loss and accuracy of the E2 gene of HCV. Discussion Deep learning model and Artificial neural networks are used and endured prompt evolution in the last decade and its predictive abilities in the field of life sciences transforming across different fields. The importance of LSTM model is present in their efficiency for predicting sequential data from the lager data and long-range dependencies. They are well-designed with mechanisms to reduce the issue of vanishing gradients, allowing them to efficiently retain and process information over extended sequences. LSTM, CNN and GRU generative models have been used for predicting and evolving complete genome of dengue virus and coronavirus (Mumtaz et al, 2024 ). Furthermore, CNN and LSTM have also been used for the predicting the cases of the Coronavirus and Dengue virus. There is a study in which Recurrent Neural Network (RNN), possibly integrates LSTM models that proves to be advantageous, it helps in predicting the HCV cases that have developed HCC due to liver cirrhosis, especially when they are compared with the logistic regression model (Ioannou et al., 2020). For generating the innovative drugs and capturing the sequential patterns in molecular structures, LSTM-ProGen model uses the LSTM model architecture. It helps with designing the molecules that interrelate efficiently with HIV-1 protease that is the target protein, with its ability to handle long range dependencies. To accelerate the development of drug discovery, LSTM ProGen produce chemically compelling target specific molecules, by augmenting the properties of drugs. LSTM helps in developing the customized medicines (Albrijawi et al., 2024). For analyzing sequential biological data and identifying antiviral peptides (AVPs) LSTM model is used. There are some specific sequences in AVPs that are captured by LSTM meritoriously, it can also improve the accuracy of the prediction model. The model permits the knowledge for improving the features of presentation and understanding their relationship with amino acids. Accommodatively learning from the trained data, LSTMs recognize perilous characteristics that can be spared by the already existing predictors. This feature commonly helps in decreasing overfitting and improves generality within the distinct datasets. Using LSTM model is one of the accurate and promising approaches used for developing the antivirals and identification of AVP (Ali et al., 2023 ). For developing vaccines and predicting cross immunoreactivity (CR) amongst diverse epitopes of HCV LSTM was employed. LSTM improves the accuracy of the prediction by recognizing crucial factors of CR and they efficiently capture the sequential dependencies in sequences of the amino acids. The ability to learn from the previous data and adjust according to the viral adaptability increases understandings of immunologic specificity. Eventually incorporating LSTM models can substantially advance the strategies in development of the vaccines for HCV and HIV (Tayebi, 2020 ) Limitations: This LSTM generative model experiences some limitations. One of the limitations is that it gives the best results with long sequences, it is not suitable for the short sequences. LSTMs are designed in such a way that they lessen the vanishing gradient issue, and still face the challenge in pertaining to the data for longer sequences leading to the prediction of the impractical variants of the virus. A considerable amount of data is required for training the LSTM model and explaining the picture of the genetic variation, this can be a limitation for the HBV and HCV genotypes. When dealing with the complex adaptabilities of the hypervariable regions of the genome, overfitting can be a biggest challenge. Furthermore, the high computational power and resources are used for the large LSTM datasets, this can also become a hurdle in versality for wide range of genetic studies. Conclusion In this study, LSTM generative model was used for generating the hypervariable gene sequences of HBV (S1 and S2) and HCV (E2), primarily emphasizing the mutation rates in the sequences of the different genotypes, which can be the source of genetic variability in the progression of the evolving genotype. These variations can be challenging for the antiviral treatments and efficiency of vaccines, because the evolving genotype might elude the response of the immune system. The ability of our LSTM model to predict the sequence of the hypervariable regions of the evolving genotype can be a great advancement in developing the targeted therapies. Predicting the precise sequences of the evolving genotype before it even appears can be the primitive approach that helps in developing the strain specific treatments that will have a positive impact on viral evolution and helps the healthcare professional to know about the virus beforehand. Overall, the LSTM generative model demonstrates significant potential for predicting evolving genotypes of HBV and HCV, ultimately improving drug and vaccine development and strengthening global public health responses to viral infections. Declarations Financial and non-financial competing interests’ declaration. The authors declare that they have no competing interest in this paper. Ethical Approval and consent to participate: It is not applicable. Funding: To write this research paper no funding has been received. Authors Contributions: Author’s Role: Sharmeen Saqib: Writing Reviewing and Editing, Methodology, Software, Data Curation, Zilwa Mumtaz : Conceptualization, Validation, Hania Ahmed: Helped in data collection, Dr Ashiq Ali: Reviewing manuscript and data analysis, Dr Obaidullah Qazi: Reviewing manuscript and data analysis, Dr Muhammad Zubair Yousaf: Formal Analysis and Supervision Acknowledgement: I would like to express my gratitude for my supervisor Dr Muhammad Zubair Yousaf for his unwavering support and guidance throughout my research. I also thank all my fellows for motivating and supporting me throughout my journey. References Sung, H., Ferlay, J., Siegel, R. L., Laversanne, M., Soerjomataram, I., Jemal, A., & Bray, F. (2021). Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: a cancer journal for clinicians , 71 (3), 209-249. https://acsjournals.onlinelibrary.wiley.com/doi/full/10.3322/caac.21660 Rumgay, H., Ferlay, J., de Martel, C., Georges, D., Ibrahim, A. S., Zheng, R., ... & Soerjomataram, I. (2022). Global, regional and national burden of primary liver cancer by subtype. European Journal of Cancer , 161 , 108-118. https://www.sciencedirect.com/science/article/abs/pii/S0959804921012430 Mumtaz, Z., Rashid, Z., Saif, R., & Yousaf, M. Z. (2024). Deep Learning guided prediction modeling of dengue virus evolving serotype. Heliyon , 10 (11), e32061. https://doi.org/10.1016/j.heliyon.2024.e32061 Liang, Y., Zhang, G., Li, Q., Han, L., Hu, X., Guo, Y., Tao, W., Zhao, X., Guo, M., Gan, T., Tong, Y., Xu, Y., Zhou, Z., Ding, Q., Wei, W., & Zhong, J. (2021). TRIM26 is a critical host factor for HCV replication and contributes to host tropism. Science Advances , 7 (2). https://doi.org/10.1126/sciadv.abd9732 Rich, N. E. (2024). Changing epidemiology of hepatocellular carcinoma within the United States and worldwide. Surgical Oncology Clinics of North America , 33 (1), 1–12. https://doi.org/10.1016/j.soc.2023.06.004 Stroffolini, T., & Stroffolini, G. (2024). Prevalence and modes of transmission of Hepatitis C virus infection: A Historical Worldwide review. Viruses , 16 (7), 1115. https://doi.org/10.3390/v16071115 Ningthoujam, S. S., Nath, R., Sarker, S. D., Nahar, L., Nath, D., & Talukdar, A. D. (2024). Prediction of medicinal properties using mathematical models and computation, and selection of plant materials. In Elsevier eBooks (pp. 91–123). https://doi.org/10.1016/b978-0-443-16102-5.00011-0 Hanke, K., Rykalina, V., Koppe, U., Gunsenheimer-Bartmeyer, B., Heuer, D., & Meixenberger, K. (2024). Developing a next level Integrated Genomic Surveillance: Advances in the Molecular Epidemiology of HIV in Germany. International Journal of Medical Microbiology , 314 , 151606. https://doi.org/10.1016/j.ijmm.2024.151606 Padminivalli, S. J. R. K., V., Rao, M. V. P. C. S., & Narne, N. S. R. (2023). Sentiment based emotion classification in unstructured textual data using dual stage deep model. Multimedia Tools and Applications , 83 (8), 22875–22907. https://doi.org/10.1007/s11042-023-16314-9 Duan, Y., Qin, J., Qiu, W., Li, S., Li, C., Liu, A., Chen, X., & Zhang, C. (2022). Performance of a generative adversarial network using ultrasound images to stage liver fibrosis and predict cirrhosis based on a deep-learning radiomics nomogram. Clinical Radiology , 77 (10), e723–e731. https://doi.org/10.1016/j.crad.2022.06.003 Bartoszewicz, J. M., Seidel, A., & Renard, B. Y. (2021). Interpretable detection of novel human viruses from genome sequencing data. NAR Genomics and Bioinformatics , 3 (1). https://doi.org/10.1093/nargab/lqab004 Zoulim, F., Chen, P. J., Dandri, M., Kennedy, P., & Seeger, C. (2024). Hepatitis B Virus DNA integration: Implications for diagnostics, therapy, and outcome. Journal of Hepatology . https://www.sciencedirect.com/science/article/pii/S0168827824023432 Broquetas, T., & Carrión, J. A. (2023). Past, present, and future of long-term treatment for hepatitis B virus. World Journal of Gastroenterology , 29 (25), 3964. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10354584/ Kareem, A. K., AL-Ani, M. M., & Nafea, A. A. (2023). Detection of autism spectrum disorder using a 1-dimensional convolutional neural network. Baghdad Science Journal , 20 (3 (Suppl.)), 1182-1182. https://www.iasj.net/iasj/download/97dacb06a351ba74 Choi, J. G., Kim, D. C., Chung, M., Lim, S., & Park, H. W. (2024). Multimodal 1D CNN for delamination prediction in CFRP drilling process with industrial robots. Computers & Industrial Engineering , 190 , 110074. https://www.sciencedirect.com/science/article/abs/pii/S0360835224001955 Krauss, P. (2024). Recurrent Neural Networks. In Artificial Intelligence and Brain Research: Neural Networks, Deep Learning and the Future of Cognition (pp. 131-137). Berlin, Heidelberg: Springer Berlin Heidelberg. https://link.springer.com/chapter/10.1007/978-3-662-68980-6_14 Choi, J. G., Kim, D. C., Chung, M., Lim, S., & Park, H. W. (2024). Multimodal 1D CNN for delamination prediction in CFRP drilling process with industrial robots. Computers & Industrial Engineering , 190 , 110074. https://www.sciencedirect.com/science/article/abs/pii/S0360835224001955 Zuvanov, L., Basso Garcia, A. L., Correr, F. H., Bizarria Jr, R., Filho, A. P. D. C., Da Costa, A. H., ... & Corrêa dos Santos, R. A. (2021). The experience of teaching introductory programming skills to bioscientists in Brazil. PLoS computational biology , 17 (11), e1009534. https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1009534 , H. N., Rimal, B., Pokhrel, N. R., Rimal, R., & Dahal, K. R. (2022). LSTM-SDM: An integrated framework of LSTM implementation for sequential data modeling. Software Impacts , 14 , 100396. https://www.sciencedirect.com/science/article/pii/S2665963822000902 Song, X., Salcianu, A., Song, Y., Dopson, D., & Zhou, D. (2020). Fast wordpiece tokenization. arXiv preprint arXiv:2012.15524 . https://arxiv.org/abs/2012.15524 Alrasheedi, F., Zhong, X., & Huang, P. C. (2023). Padding module: Learning the padding in deep neural networks. IEEE Access , 11 , 7348-7357. https://ieeexplore.ieee.org/abstract/document/10021573/ Sunny, M. A. I., Maswood, M. M. S., & Alharbi, A. G. (2020, October). Deep learning-based stock price prediction using LSTM and bi-directional LSTM model. In 2020 2nd novel intelligent and leading emerging sciences conference (NILES) (pp. 87-92). IEEE. https://ieeexplore.ieee.org/abstract/document/9257950 Shahade, A. K., Walse, K. H., Thakare, V. M., & Atique, M. (2023). Multi-lingual opinion mining for social media discourses: An approach using deep learning-based hybrid fine-tuned smith algorithm with adam optimizer. International Journal of Information Management Data Insights , 3 (2), 100182. https://www.sciencedirect.com/science/article/pii/S2667096823000290 Martínez-Llop, P. G., Bobi, J. D. D. S., & Ortega, M. O. (2023). Time consideration in machine learning models for train comfort prediction using LSTM networks. Engineering Applications of Artificial Intelligence , 123 , 106303. https://www.sciencedirect.com/science/article/pii/S0952197623004876 Ferruz, N., Heinzinger, M., Akdel, M., Goncearenco, A., Naef, L., & Dallago, C. (2023). From sequence to function through structure: Deep learning for protein design. Computational and Structural Biotechnology Journal , 21 , 238-250. https://www.sciencedirect.com/science/article/pii/S2001037022005086 Wang, R., Jiang, Y., Jin, J., Yin, C., Yu, H., Wang, F., ... & Wei, L. (2023). DeepBIO: an automated and interpretable deep-learning platform for high-throughput biological sequence prediction, functional annotation and visualization analysis. Nucleic acids research , 51 (7), 3017-3029. https://academic.oup.com/nar/article/51/7/3017/7041952#google_vignette Ioannou, G. N., Tang, W., Beste, L. A., Tincopa, M. A., Su, G. L., Van, T., ... & Waljee, A. K. (2020). Assessment of a deep learning model to predict hepatocellular carcinoma in patients with hepatitis C cirrhosis. JAMA network open , 3 (9), e2015626-e2015626. https://jamanetwork.com/journals/jamanetworkopen/article-abstract/2770062 Albrijawi, M. T., & Alhajj, R. (2024). LSTM-driven drug design using SELFIES for target-focused de novo generation of HIV-1 protease inhibitor candidates for AIDS treatment. PloS one , 19 (6), e0303597. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0303597 Ali, F., Kumar, H., Alghamdi, W., Kateb, F. A., & Alarfaj, F. K. (2023). Recent advances in machine learning-based models for prediction of antiviral peptides. Archives of Computational Methods in Engineering , 30 (7), 4033-4044. https://link.springer.com/article/10.1007/s11831-023-09933-w Tayebi, Z. (2020). Machine learning and deep learning to predict cross-immunoreactivity of viral epitopes. https://scholarworks.gsu.edu/cs_theses/96/ National Center for Biotechnology Information (NCBI). (n.d.). National Center for Biotechnology Information . U.S. National Library of Medicine. https://www.ncbi.nlm.nih.gov/ Kent, W. J., Sugnet, C. W., Furey, T. S., Roskin, K. M., Pringle, T. H., Zahler, A. M., & Haussler, D. (2002). The human genome browser at UCSC. Genome Research , 12(6), 996-1006. https://doi.org/10.1101/gr.229102 Tamura, K., Stecher, G., & Kumar, S. (2021). MEGA11: Molecular Evolutionary Genetics Analysis version 11. Molecular Biology and Evolution , 38(7), 3022-3027. https://doi.org/10.1093/molbev/msab120 Additional Declarations No competing interests reported. Supplementary Files floatimage1.png Graphical Abstract Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5560102","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":391993527,"identity":"1794d4a1-e1ac-4f91-aab2-23fdafb795ad","order_by":0,"name":"Sharmeen Saqib","email":"","orcid":"","institution":"Forman Christian College","correspondingAuthor":false,"prefix":"","firstName":"Sharmeen","middleName":"","lastName":"Saqib","suffix":""},{"id":391993528,"identity":"86fcb38a-dabd-4f2f-b3ac-ec7d8f240053","order_by":1,"name":"Zilwa Mumtaz","email":"","orcid":"","institution":"Forman Christian College","correspondingAuthor":false,"prefix":"","firstName":"Zilwa","middleName":"","lastName":"Mumtaz","suffix":""},{"id":391993529,"identity":"066c3b79-226f-4053-bd9b-2dc3d2be1bea","order_by":2,"name":"Hania Ahmed","email":"","orcid":"","institution":"Forman Christian College","correspondingAuthor":false,"prefix":"","firstName":"Hania","middleName":"","lastName":"Ahmed","suffix":""},{"id":391993530,"identity":"fa09128d-741a-437a-a8b6-0debe162894a","order_by":3,"name":"Ashiq Ali","email":"","orcid":"","institution":"Chinese Academy of Sciences","correspondingAuthor":false,"prefix":"","firstName":"Ashiq","middleName":"","lastName":"Ali","suffix":""},{"id":391993531,"identity":"836c268f-aa19-4d6c-abd9-03af872fbedd","order_by":4,"name":"Obaidullah Qazi","email":"","orcid":"","institution":"Abdul Wali Khan University","correspondingAuthor":false,"prefix":"","firstName":"Obaidullah","middleName":"","lastName":"Qazi","suffix":""},{"id":391993532,"identity":"47da878c-253b-463b-8031-4507faeaf111","order_by":5,"name":"Muhammad Zubair Yousaf","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABNklEQVRIiWNgGAWjYHACNgYGAyizokIOSDI3HoBwE4D4AAEtZ84YA0nGBiK0wMDZNiK06PafMXvwo+BeYr9E+sMPB+cZyPO3NzYc+PDLhoGfPceAueAMhhazA2fMDXsMihNnzsgxlji4zcBwxpmDDQdn9qUxSPa8MWCecQNTy8EeMwkeg4TEDTdyGKQ/bvuTYCCR2HCYt+cwg8ENoC08HzC1HOYxk/wD1pL++MfBOQYQLX97/jPY49JyjMdMGmJLgpnEwQaoFoYfBxgMJEBasDjsDFu5sYxBgvHMnjdmFgeOQf3S25DMI3HmWcHhGVi8f/7wtodv/iTI9rOnP75xoAYUYs0HH/z4YyfH35688XHBMQwtMODYgMJlbGPgAdGHcWpgYLBH4/+BUMx4tIyCUTAKRsGIAQDwx3vxDSd44AAAAABJRU5ErkJggg==","orcid":"","institution":"Forman Christian College","correspondingAuthor":true,"prefix":"","firstName":"Muhammad","middleName":"Zubair","lastName":"Yousaf","suffix":""}],"badges":[],"createdAt":"2024-12-01 21:08:03","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5560102/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5560102/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":72037660,"identity":"02a5a6b5-2ccb-463a-83d9-7418698b7d6d","added_by":"auto","created_at":"2024-12-21 00:28:27","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":216836,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePairwise 2 similarities between the sequences given and the generated sequence in percentage of HCV E2.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/5b0d4994172b7f64117e2d0a.png"},{"id":72037951,"identity":"0762535d-4b48-49ee-a7e8-f77793410348","added_by":"auto","created_at":"2024-12-21 00:36:28","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":217628,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePairwise 2 similarities between HBV S1 sequences given and generated sequences.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/bcc3f47342c5e24019e7d940.png"},{"id":72039061,"identity":"a7052d0a-350e-4c7d-92db-0b148d756468","added_by":"auto","created_at":"2024-12-21 00:52:27","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":244321,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePairwise 2 similarities between HBV S2 sequences given and generated sequences\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/7c79164ae80895dec7f00c35.png"},{"id":72037948,"identity":"56359a3b-05e7-4a77-94f0-56d2174f0ce7","added_by":"auto","created_at":"2024-12-21 00:36:28","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":200188,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eConfusion matrix for HCV E2\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/295b54ca494821d6e6c7ff0d.png"},{"id":72037663,"identity":"421bc48b-89c4-4b95-9440-9cfdecd81b0f","added_by":"auto","created_at":"2024-12-21 00:28:28","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":232721,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eConfusion matrix for HBV S1\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/c6c21cf7144c038019cb62c2.png"},{"id":72037673,"identity":"71a45fee-0828-4b94-a7af-e104ed33c2fc","added_by":"auto","created_at":"2024-12-21 00:28:28","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":254522,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eConfusion matrix for HBV S2\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/6d64d2e467c4095990aae29a.png"},{"id":72037675,"identity":"fa89f35c-0c37-4ba1-b53d-aca75913688a","added_by":"auto","created_at":"2024-12-21 00:28:28","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":153404,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eModel Accuracy and loss for S2 HCV\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/6b8a5cf16ef01bde21644f66.png"},{"id":72037667,"identity":"d6c0c3a6-463e-476a-977c-1a47140ff2ef","added_by":"auto","created_at":"2024-12-21 00:28:28","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":106274,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eModel Accuracy and Model Loss for S1 HBV.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/a4e65d1c7a63fe292525a250.png"},{"id":72037669,"identity":"87cf9eb4-aaf5-4fdc-abd2-883385fbfeed","added_by":"auto","created_at":"2024-12-21 00:28:28","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":117235,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eModel Accuracy and Model Loss for S2 HBV.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/4f54e4fb921177ab9fea5ef7.png"},{"id":72074685,"identity":"98f29abc-81eb-4700-b9f4-a4a37708cc1a","added_by":"auto","created_at":"2024-12-21 15:16:47","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2501284,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/bd752c5f-45a2-4505-b4c5-c6866cd44ec1.pdf"},{"id":72037946,"identity":"f4322cc3-369d-44a6-8b5c-f5fa0ffdf531","added_by":"auto","created_at":"2024-12-21 00:36:27","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":103414,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGraphical Abstract\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-5560102/v1/846dd2ede45e2f674697442c.png"}],"financialInterests":"No competing interests reported.","formattedTitle":"Generative Deep Neural Networks for Estimating Hypervariability in Hepatitis B and C Virus Genomes","fulltext":[{"header":"Introduction","content":"\u003cp\u003eHepatitis B virus (HBV) and Hepatitis C virus (HCV) have always been considered the major health concern around the world. Among all the cancers, Liver cancer ranks third with approximately 1.3\u0026nbsp;million deaths each year by Hepatitis B virus (HBV) and Hepatitis C virus (HCV) (Rumgay et al, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Presently, 70\u0026ndash;80% of hepatocellular carcinoma (HCC) lead to liver cancer. Both HBV and HCV have played a significant role in the etiological traits of HCC. In all cases of HCV related to HCC, cirrhosis is present. Whereas, in HBV cirrhosis is absent in majority of cases (Sung et al, \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). HCV is a single stranded RNA enveloped virus with the genome size of 9.6 kb, flanked by 5` and 3` untranslated regions. Over three thousand amino acids are transcribed into a single polyprotein, which is then broken down into structural and nonstructural proteins (Liang et al., \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). HCV is the primary source of liver cirrhosis that leads to hepatocellular carcinoma (HCC) with a global decline in survival and an increase in mortality (Rich, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Hepatic cancer can be prevented by treating the patient with newly developed antivirals for HCV. Higher death rate is reported after liver transplantation in HCV-positive recipients, and allograft rejection is caused by HCV infection. Knowing about the virus and host interaction can help in the management and prevention of HCV infection. HCV has been found to infect only humans and chimpanzees within the restricted host range (Stroffolini \u0026amp; Stroffolini, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Hepatitis B is a DNA enveloped virus that belongs to the \u003cem\u003eHepadnaviridae\u003c/em\u003e family of viruses. The size of the genome is 3kb which is wrapped in an envelope known as nucleocapsid. The genome of 90\u0026ndash;95% viral particles is present in relaxed circular DNA (rcDNA) that is occupied by the unified overlap sequence. There are 334 amino acids present in the genome of HBV (Zoulim Et al, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Lamivudine (LAM) was introduced in 1995, for the treatment of HBV and it developed the resistance among 75% of the patients after 5 years of the treatment and gave rise to a resistant strain of HBV. For treating these resistant strains of HBV, Adefovir (ADV) was introduced in 2002. In 2008, ADV was replaced by a more effective drug known as Tenofovir (TDF). The latest drug that is approved for treating HBV is tenofovir Alafenamide (TAF) in 2015 and it can also be used for patients with other issues like bone, heart, kidneys and lungs (Broquetas et al, 2023).\u003c/p\u003e \u003cp\u003eDeep learning (DL) is a strong tool that is used in the field of genomic analysis, specifically when it comes to shifting through complex genomic data to isolate features and patterns. Likewise, the machine learning has also become popular when it comes to dealing with infectious diseases like HCV, HBV, Dengue, Coronavirus etc. using epidemiological and clinical data for developing a predictive model. Viral genome classification can be done in a variety of ways using machine learning and alignment techniques (Ningthoujam et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Detection of the viral sequences are carried out by tools like USEARCH, REGA and SCUEAL, that rely on the score of alignment for classification of genome. Although these methods have some constraints, particularly their execution depends upon the initial alignment that is selected (Hanke et al., \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). However, there are numerous techniques that were proposed before the machine learning approach was used for the classification of viral genome sequences. But when it comes to identifying the viral genome contig and obtaining the useful hidden data these techniques have the limitations and face difficulty in executing the task. Moreover, they were trained on nucleotide sequence data that further restricts their usefulness for genomic research. In the natural language processing areas, Recurrent Neural Network (RNN\u0026rsquo;s) a deep learning model has been proven to be effective. Deep learning applications in computational biology mostly focus on genome sequencing and analysis (Padminivalli et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eHowever, because of the model's intricacy in deciphering its concealed state, RNNs are regarded as black boxes. Currently, for many clinical classifications task machine learning have been used. Deep learning is the subgroup of machine learning that nearly mimics the structure of the neurons. Clinical decision-making is enhanced by radiomics, a non-invasive technique that extracts a multitude of medical imaging features (both hand-crafted and deep-learning features). It is important for the medical practitioners to use the quantified scores along with the other clinical symptoms to improve the model while they evaluate lesions. Convolutional Neural Network (CNN) has been proved to be effective for interpreting the real-world picture, these are some major tools that have been used in majority of these techniques, and they only concentrate on deep learning or classic radiomics (Duan et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFor many different RNA/ DNA- based prediction tasks, convolutional neural networks (CNNs) and long short-term memory architectures (LSTMs) are proved to be effective. Previous studies mainly emphasized the regulation of human gene expression, and this topic is now being actively studied. Deep learning models have been directly trained on RNA/ DNA sequences in the field of genomic pathogens to envisage the infective potential of new pathogens, as well as the host ranges of three multi-host virus species. While ViraMiner and DeepVirFinder may identify viral sequences in metagenomic materials, they are limited to previously identified species and are unable to predict the host (Bartoszewicz et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eViral genome sequences have been detected using Long Short-Term Memory (LSTM) and Convolutional Neural Networks (CNN) by training the pattern and frequency of branching. On the other hand, one kind of CNN model that excels in analyzing patterns in images while considering the inherent heterogeneity in the data is the Feedforward Neural Network (FNN). This network is trained using a backpropagation algorithm and is intended to translate fixed-length inputs to a fixed-size output. CNN, which consists of several layers, efficiently learns the complex relationship between input and output and stores and updates information in filter weights.\u003c/p\u003e \u003cp\u003eThe main objective of this study is to predict the sequences of the hypervariable genes of HBV (S1 and S2) and HCV envelop protein (E2) of evolving genotypes to have deep knowledge and better understanding about the evolution of HBV and HCV. Instead of expressing character meanings, the model learns patterns in tokens by using simple Recurrent Neural Networks (Simple RNNs) trained on the hypervariable gene sequences of already present genotypes. When it comes to processing sequential data and addressing the vanishing gradient issue that standard RNNs have, LSTM is the preferred option for genomic prediction tasks (Mumtaz et al, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). We did the comparison between the three deep learning generative models (Simple RNN, 1-dimensional Convolutional Neural Network and Long-Short Term Memory). The focus of this study was that the sequence generation of the hypervariable regions of HBV and HCV these genes are most mutated genes and play a significant role in changing the genotype of the virus. When we get the sequence of hypervariable gene of evolving genotype it would be much easier to develop the strain specific drug/vaccine even before the emergence of the evolving genotype.\u003c/p\u003e"},{"header":"Material and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eData Collection and pre-processing:\u003c/h2\u003e \u003cp\u003eA data set consisting of a total of 1500 sequences of hypervariable genes of Hepatitis B Virus (HBV) S1 and S2 gene. And 1500 sequences of Hepatitis C Virus (HCV) hypervariable E2 gene of each genotype were retrieved from BV-BRC (Kent, 2002) and NCBI databases (NCBI, 1988). Preparation of the data was done before the analysis to make sure that the data is in the appropriate format using MEGA 11 software (Tamura, 2021). Hypervariable genes of HBV are S1 and S2. S1 is 835 base pairs while S2 is 333 base pairs. Whereas E2 of HCV is 1106 base pairs. The gene sequence data from each genotype of HBV and HCV was compiled in a separate file and labelled accordingly.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eDeep learning Model Architecture\u003c/h3\u003e\n\u003cp\u003eThe generative deep learning model used the Bio python Seq IO library for loading the data from FASTA file containing HBV and HCV hypervariable gene sequences. The function load_fasta_file extracts and read sequences from the files including the sequences and subsequent virus types. It then shows all the sequences in the file along with the number of sequences (Zuvanov et al., 2021). After loading the data, the data is then prepared for training the deep learning model by tokenization, of the data at the character level and then producing the n gram input sequences. This tokenizer was used from the library known as TensorFlow Keras (Bhandari et al., 2022). By tokenizer each sequence present in the data is converted into integers, where each character is represented as an exclusive integer (Song et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). For standardizing the size of input data, padding was performed to match the length of all the sequences. All the data in the output is split into the predictor except the last character in the sequence and it labels that last character which then aids in training the data for model (Alrasheedi et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Then the LSTM deep learning model was constructed using TensorFlow\u0026rsquo;s Keras library to predict the gene sequences of the evolving genotype. The architecture of this model consists of two dense embedded layers of Long Short-Term Memory (LSTM) designed to apprehend sequential dependencies in the data set (Sunny et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). The final layer is the thick dense layer, it helps in generating the probability distribution for the feasible characters in vocabulary, making it a suitable fit for multi-class categorization which generates probability distributions for all possible characters in the vocabulary, making it appropriate for multi-class classification. For compiling the model sparse categorical cross entropy is used as a loss function, and for optimizing the model weight Adam optimizer is used (Shahade et al., \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). This model is designed to learn and then generate hypervariable gene sequences of the evolving genotype. The training of the model was done on nucleotide sequences using the fit function of TensorFlow\u0026rsquo;s Keras library. The training was done on 15 epochs, with the choice to fix the value based on data. This method of training guarantees effective training, sustaining the finest performance of the model along with eluding unnecessary overfitting in the data being trained (Mart\u0026iacute;nez-Llop et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). For the generation of hypervariable gene sequences, the deep learning model is pre-trained. The seed text use the generate_sequence function to predict the successive character and builds the gene sequence found on the possibility that developed from output of the model. The performance of the trained deep learning model was evaluated by a sequence predicted by based on the tested dataset and producing the confusion matrix for assessing the accuracy of the generative model (Ferruz et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). The prediction of next character in the sequence used the probabilistic method by using the feature of randomness for enhancing diversity in the sequences that are generated. The generated sequences are produced by diverse nucleotides, and they are analyzed for genetic variations (Wang et al., 2023).\u003c/p\u003e\n\u003ch3\u003eEmbedding layer transformation\u003c/h3\u003e\n\u003cp\u003e \u003cdiv id=\"Equa\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:E\\left(xt\\right)=We\\cdot\\:1xt$$\u003c/div\u003e \u003c/div\u003e \u003c/p\u003e\n\u003ch3\u003eLSTM update\u003c/h3\u003e\n\u003cp\u003e \u003cdiv id=\"Equb\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:{h}_{t,}{c}_{t}=LSTM\\left(\\varkappa\\:t,{h}_{t}-1s,{c}_{t}-1\\right)$$\u003c/div\u003e \u003c/div\u003e \u003c/p\u003e\n\u003ch3\u003eSoftMax Prediction\u003c/h3\u003e\n\u003cp\u003e \u003cdiv id=\"Equc\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:P\\left(\\left.{y}_{t}\\right|{h}_{t}\\right)=\\frac{{e}^{{W}_{0}}\\cdot\\:ht+{b}_{0}}{{\\sum\\:}_{k=1}{V}_{e}{W}_{0}\\cdot\\:ht+{b}_{0}}$$\u003c/div\u003e \u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eComparison of the other generative models:\u003c/h2\u003e \u003cp\u003eIn this study, we compared the LSTM model with other generative models like Simple Recurrent Neural Network (Simple RNN), 1-Dimensional Convolutional Neural Network (ConV1d). For training all these models Adam Optimizer and sparse categorical cross entropy were used for generating sequences.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eSimple Recurrent Neural Network (Simple RNN)\u003c/h3\u003e\n\u003cp\u003eA neural network that is designed for processing sequential data by using the cycles in its design is known as Simple RNN. RNNs preserve an obscure state that depicts data about the prior inputs, allowing them to remember the previous dataset. For prediction of time series, language models, and speech recognition the memory mechanism of RNN is used. RNN can be laborious due to their long dependency problems and exploding gradients, that can impede their functioning when working for a longer period. Despite these challenges, innovative architectures like GRU and LSTM Simple RNN is still used as the base to tackle these constraints (Krauss, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e\n\u003ch3\u003e1-Dimensional Convolution Neural Network (ConV1d)\u003c/h3\u003e\n\u003cp\u003eThis deep learning model is used for generative and classification purposes based on the numerical data that is provided in the form of 1 dimensional sequence (Kareem et al, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). The input data is given to the model in the form of sequences and then it applies the chain of layers to produce a solid demonstration of the input given to the model and predict the sequence. The layers that are used in 1 dimensional generative model are input layers, convolutional layers, activation layers, pooling layers and fully connected layers (Choi et al, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eWe worked on two other models other than our proposed model; these include Simple RNN and ConV1d for generating the hypervariable gene sequences of HBV and HCV. But LSTM, our proposed model proves to be the most efficient and accurate model for the generation of gene sequences. At first, we trained each model on fifteen epochs, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. ConV1D at fifteen epochs gave 25% accuracy and the length of the sequence generated for E2 is 654, S1 is 498 and S2 is 132 base pairs. But at twenty epochs its accuracy increased to 27% and the length of the sequence generated is 660 for E2, 510 for S1, and 190 for S2. The Simple RNN model gave 36% accuracy at fifteen epoch and 40% accuracy at twenty. The drawback of these two models was that these do not produce accurate sequence. In Simple RNN model there were some gaps that are seen in the generated sequence and ConV1d repeatedly generated the four characters in the same sequence with the gaps in between the sequence. The LSTM model, when trained at fifteen epochs, generated 1106 base pair sequence of hypervariable gene E2 from HCV genome and 835 base pairs of S1 and 333 base pairs of S2 hypervariable genes from HBV genome these are the complete sizes of the genes. Whereas the Simple RNN and ConV1d does not produce the complete sequences. The sequence length had a considerable impact on the predictive capabilities of the model and its biological importance. For further, increasing the accuracy of the model we trained the model by increasing the epochs from 15 to 20 and then checked the accuracy that appeared to increase the model\u0026rsquo;s accuracy along with the epoch. The LSTM model that was trained at 15 epochs gave 60% accuracy and at 20 epochs it showed 85% accuracy. Hence this proved that increasing the epoch increases the accuracy and the predictive abilities of the model.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracy in percentage of the model on 15 and 20 Epochs, and length of the sequences they generated in base pairs.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eEpoch\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAccuracy (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSequence Generated (Base pairs)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eConV1d\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE2\u0026thinsp;=\u0026thinsp;654\u003c/p\u003e \u003cp\u003eS1\u0026thinsp;=\u0026thinsp;498\u003c/p\u003e \u003cp\u003eS2\u0026thinsp;=\u0026thinsp;132\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE2\u0026thinsp;=\u0026thinsp;660\u003c/p\u003e \u003cp\u003eS1\u0026thinsp;=\u0026thinsp;510\u003c/p\u003e \u003cp\u003eS2\u0026thinsp;=\u0026thinsp;190\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSimple RNN\u003c/p\u003e \u003cp\u003e(Approximately)\u003c/p\u003e \u003cp\u003eGaps at the end of the sequence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE2\u0026thinsp;=\u0026thinsp;998\u003c/p\u003e \u003cp\u003eS1\u0026thinsp;=\u0026thinsp;734\u003c/p\u003e \u003cp\u003eS2\u0026thinsp;=\u0026thinsp;267\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE2\u0026thinsp;=\u0026thinsp;1000\u003c/p\u003e \u003cp\u003eS1\u0026thinsp;=\u0026thinsp;789\u003c/p\u003e \u003cp\u003eS2\u0026thinsp;=\u0026thinsp;290\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eLSTM\u003c/p\u003e \u003cp\u003e(Accurately)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE2\u0026thinsp;=\u0026thinsp;1106\u003c/p\u003e \u003cp\u003eS1\u0026thinsp;=\u0026thinsp;835\u003c/p\u003e \u003cp\u003eS2\u0026thinsp;=\u0026thinsp;333\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE2\u0026thinsp;=\u0026thinsp;1106\u003c/p\u003e \u003cp\u003eS1\u0026thinsp;=\u0026thinsp;835\u003c/p\u003e \u003cp\u003eS2\u0026thinsp;=\u0026thinsp;333\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eWe performed pairwise2 similarity on the sequences generated by the deep leaning generative model. The hypervariable gene sequences that are generated by our LSTM model have a greater accuracy with genotype A that is 64.52% of S1 HBV gene (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) and with genotype A, B, C, D, E which is 61.56% of S2 of HBV. (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e) Whereas the model shows the highest accuracy rate with genotype 7 in case of E2 gene of HCV which is 65.01%. (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e) Results of the pairwise2 similarity demonstrated that the generated sequence was not closely related to the known genotypes of HBV and HCV. The sequences that were generated by the LSTM model were of perfect length of the specific gene size. Which proves that our model predicted the gene sequences completely based on the randomization without any biasness.\u003c/p\u003e \u003cp\u003eThe confusion matrix is the powerful tool for assessing the classification task, analyzing the categories of the output sequences and helps in performing the error analysis and ensuring the quality of the sequences. It is used for comparing the generated sequences with the data given to the model, providing intuitions of the false positive, false negatives, true positives and true negatives that evaluates the performance of the model. Figures\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e and \u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e illustrates the confusion matrix for E2, S1 and S2 genes respectively. We check the loss and accuracy of the model. The direct relationship is observed between the accuracy and the epochs whereas, indirect relation is observed between the loss and epochs. Figure\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e represents the accuracy and loss of S2 gene of HBV. Figure\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e shows the loss accuracy for S1 gene of HBV and Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e demonstrates the loss and accuracy of the E2 gene of HCV.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eDeep learning model and Artificial neural networks are used and endured prompt evolution in the last decade and its predictive abilities in the field of life sciences transforming across different fields. The importance of LSTM model is present in their efficiency for predicting sequential data from the lager data and long-range dependencies. They are well-designed with mechanisms to reduce the issue of vanishing gradients, allowing them to efficiently retain and process information over extended sequences. LSTM, CNN and GRU generative models have been used for predicting and evolving complete genome of dengue virus and coronavirus (Mumtaz et al, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Furthermore, CNN and LSTM have also been used for the predicting the cases of the Coronavirus and Dengue virus. There is a study in which Recurrent Neural Network (RNN), possibly integrates LSTM models that proves to be advantageous, it helps in predicting the HCV cases that have developed HCC due to liver cirrhosis, especially when they are compared with the logistic regression model (Ioannou et al., 2020). For generating the innovative drugs and capturing the sequential patterns in molecular structures, LSTM-ProGen model uses the LSTM model architecture. It helps with designing the molecules that interrelate efficiently with HIV-1 protease that is the target protein, with its ability to handle long range dependencies. To accelerate the development of drug discovery, LSTM ProGen produce chemically compelling target specific molecules, by augmenting the properties of drugs. LSTM helps in developing the customized medicines (Albrijawi et al., 2024). For analyzing sequential biological data and identifying antiviral peptides (AVPs) LSTM model is used. There are some specific sequences in AVPs that are captured by LSTM meritoriously, it can also improve the accuracy of the prediction model. The model permits the knowledge for improving the features of presentation and understanding their relationship with amino acids. Accommodatively learning from the trained data, LSTMs recognize perilous characteristics that can be spared by the already existing predictors. This feature commonly helps in decreasing overfitting and improves generality within the distinct datasets. Using LSTM model is one of the accurate and promising approaches used for developing the antivirals and identification of AVP (Ali et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). For developing vaccines and predicting cross immunoreactivity (CR) amongst diverse epitopes of HCV LSTM was employed. LSTM improves the accuracy of the prediction by recognizing crucial factors of CR and they efficiently capture the sequential dependencies in sequences of the amino acids. The ability to learn from the previous data and adjust according to the viral adaptability increases understandings of immunologic specificity. Eventually incorporating LSTM models can substantially advance the strategies in development of the vaccines for HCV and HIV (Tayebi, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2020\u003c/span\u003e)\u003c/p\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eLimitations:\u003c/h2\u003e \u003cp\u003eThis LSTM generative model experiences some limitations. One of the limitations is that it gives the best results with long sequences, it is not suitable for the short sequences. LSTMs are designed in such a way that they lessen the vanishing gradient issue, and still face the challenge in pertaining to the data for longer sequences leading to the prediction of the impractical variants of the virus. A considerable amount of data is required for training the LSTM model and explaining the picture of the genetic variation, this can be a limitation for the HBV and HCV genotypes. When dealing with the complex adaptabilities of the hypervariable regions of the genome, overfitting can be a biggest challenge. Furthermore, the high computational power and resources are used for the large LSTM datasets, this can also become a hurdle in versality for wide range of genetic studies.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn this study, LSTM generative model was used for generating the hypervariable gene sequences of HBV (S1 and S2) and HCV (E2), primarily emphasizing the mutation rates in the sequences of the different genotypes, which can be the source of genetic variability in the progression of the evolving genotype. These variations can be challenging for the antiviral treatments and efficiency of vaccines, because the evolving genotype might elude the response of the immune system. The ability of our LSTM model to predict the sequence of the hypervariable regions of the evolving genotype can be a great advancement in developing the targeted therapies. Predicting the precise sequences of the evolving genotype before it even appears can be the primitive approach that helps in developing the strain specific treatments that will have a positive impact on viral evolution and helps the healthcare professional to know about the virus beforehand. Overall, the LSTM generative model demonstrates significant potential for predicting evolving genotypes of HBV and HCV, ultimately improving drug and vaccine development and strengthening global public health responses to viral infections.\u003c/p\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFinancial and non-financial competing interests\u0026rsquo; declaration.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interest in this paper.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical Approval and consent to participate:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIt is not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo write this research paper no funding has been received.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors Contributions:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor\u0026rsquo;s Role:\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSharmeen Saqib:\u003c/strong\u003e Writing Reviewing and Editing, Methodology, Software, Data Curation, \u003cstrong\u003eZilwa Mumtaz\u003c/strong\u003e: Conceptualization, Validation, \u003cstrong\u003eHania Ahmed:\u0026nbsp;\u003c/strong\u003eHelped in data collection, \u003cstrong\u003eDr Ashiq Ali:\u0026nbsp;\u003c/strong\u003eReviewing manuscript and data analysis, \u003cstrong\u003eDr Obaidullah Qazi:\u0026nbsp;\u003c/strong\u003eReviewing manuscript and data analysis, \u003cstrong\u003eDr Muhammad Zubair Yousaf:\u003c/strong\u003e Formal Analysis and Supervision\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgement:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eI would like to express my gratitude for my supervisor Dr Muhammad Zubair Yousaf for his unwavering support and guidance throughout my research. I also thank all my fellows for motivating and supporting me throughout my journey.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eSung, H., Ferlay, J., Siegel, R. L., Laversanne, M., Soerjomataram, I., Jemal, A., \u0026amp; Bray, F. (2021). Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. \u003cem\u003eCA: a cancer journal for clinicians\u003c/em\u003e, \u003cem\u003e71\u003c/em\u003e(3), 209-249.\u003cbr\u003ehttps://acsjournals.onlinelibrary.wiley.com/doi/full/10.3322/caac.21660 \u003c/li\u003e\n\u003cli\u003eRumgay, H., Ferlay, J., de Martel, C., Georges, D., Ibrahim, A. S., Zheng, R., ... \u0026amp; Soerjomataram, I. (2022). Global, regional and national burden of primary liver cancer by subtype. \u003cem\u003eEuropean Journal of Cancer\u003c/em\u003e, \u003cem\u003e161\u003c/em\u003e, 108-118.\u003cbr\u003ehttps://www.sciencedirect.com/science/article/abs/pii/S0959804921012430 \u003c/li\u003e\n\u003cli\u003eMumtaz, Z., Rashid, Z., Saif, R., \u0026amp; Yousaf, M. Z. (2024). Deep Learning guided prediction modeling of dengue virus evolving serotype. \u003cem\u003eHeliyon\u003c/em\u003e, \u003cem\u003e10\u003c/em\u003e(11), e32061. https://doi.org/10.1016/j.heliyon.2024.e32061 \u003c/li\u003e\n\u003cli\u003eLiang, Y., Zhang, G., Li, Q., Han, L., Hu, X., Guo, Y., Tao, W., Zhao, X., Guo, M., Gan, T., Tong, Y., Xu, Y., Zhou, Z., Ding, Q., Wei, W., \u0026amp; Zhong, J. (2021). TRIM26 is a critical host factor for HCV replication and contributes to host tropism. \u003cem\u003eScience Advances\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e(2). https://doi.org/10.1126/sciadv.abd9732 \u003c/li\u003e\n\u003cli\u003eRich, N. E. (2024). Changing epidemiology of hepatocellular carcinoma within the United States and worldwide. \u003cem\u003eSurgical Oncology Clinics of North America\u003c/em\u003e, \u003cem\u003e33\u003c/em\u003e(1), 1\u0026ndash;12. https://doi.org/10.1016/j.soc.2023.06.004\u003c/li\u003e\n\u003cli\u003eStroffolini, T., \u0026amp; Stroffolini, G. (2024). Prevalence and modes of transmission of Hepatitis C virus infection: A Historical Worldwide review. \u003cem\u003eViruses\u003c/em\u003e, \u003cem\u003e16\u003c/em\u003e(7), 1115. https://doi.org/10.3390/v16071115\u003c/li\u003e\n\u003cli\u003eNingthoujam, S. S., Nath, R., Sarker, S. D., Nahar, L., Nath, D., \u0026amp; Talukdar, A. D. (2024). Prediction of medicinal properties using mathematical models and computation, and selection of plant materials. In \u003cem\u003eElsevier eBooks\u003c/em\u003e (pp. 91\u0026ndash;123). https://doi.org/10.1016/b978-0-443-16102-5.00011-0\u003c/li\u003e\n\u003cli\u003eHanke, K., Rykalina, V., Koppe, U., Gunsenheimer-Bartmeyer, B., Heuer, D., \u0026amp; Meixenberger, K. (2024). Developing a next level Integrated Genomic Surveillance: Advances in the Molecular Epidemiology of HIV in Germany. \u003cem\u003eInternational Journal of Medical Microbiology\u003c/em\u003e, \u003cem\u003e314\u003c/em\u003e, 151606. https://doi.org/10.1016/j.ijmm.2024.151606\u003c/li\u003e\n\u003cli\u003ePadminivalli, S. J. R. K., V., Rao, M. V. P. C. S., \u0026amp; Narne, N. S. R. (2023). Sentiment based emotion classification in unstructured textual data using dual stage deep model. \u003cem\u003eMultimedia Tools and Applications\u003c/em\u003e, \u003cem\u003e83\u003c/em\u003e(8), 22875\u0026ndash;22907. https://doi.org/10.1007/s11042-023-16314-9\u003c/li\u003e\n\u003cli\u003eDuan, Y., Qin, J., Qiu, W., Li, S., Li, C., Liu, A., Chen, X., \u0026amp; Zhang, C. (2022). Performance of a generative adversarial network using ultrasound images to stage liver fibrosis and predict cirrhosis based on a deep-learning radiomics nomogram. \u003cem\u003eClinical Radiology\u003c/em\u003e, \u003cem\u003e77\u003c/em\u003e(10), e723\u0026ndash;e731. https://doi.org/10.1016/j.crad.2022.06.003 \u003c/li\u003e\n\u003cli\u003eBartoszewicz, J. M., Seidel, A., \u0026amp; Renard, B. Y. (2021). Interpretable detection of novel human viruses from genome sequencing data. \u003cem\u003eNAR Genomics and Bioinformatics\u003c/em\u003e, \u003cem\u003e3\u003c/em\u003e(1). https://doi.org/10.1093/nargab/lqab004 \u003c/li\u003e\n\u003cli\u003eZoulim, F., Chen, P. J., Dandri, M., Kennedy, P., \u0026amp; Seeger, C. (2024). Hepatitis B Virus DNA integration: Implications for diagnostics, therapy, and outcome. \u003cem\u003eJournal of Hepatology\u003c/em\u003e. https://www.sciencedirect.com/science/article/pii/S0168827824023432\u003c/li\u003e\n\u003cli\u003eBroquetas, T., \u0026amp; Carri\u0026oacute;n, J. A. (2023). Past, present, and future of long-term treatment for hepatitis B virus. \u003cem\u003eWorld Journal of Gastroenterology\u003c/em\u003e, \u003cem\u003e29\u003c/em\u003e(25), 3964. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10354584/\u003c/li\u003e\n\u003cli\u003eKareem, A. K., AL-Ani, M. M., \u0026amp; Nafea, A. A. (2023). Detection of autism spectrum disorder using a 1-dimensional convolutional neural network. \u003cem\u003eBaghdad Science Journal\u003c/em\u003e, \u003cem\u003e20\u003c/em\u003e(3 (Suppl.)), 1182-1182. https://www.iasj.net/iasj/download/97dacb06a351ba74\u003c/li\u003e\n\u003cli\u003eChoi, J. G., Kim, D. C., Chung, M., Lim, S., \u0026amp; Park, H. W. (2024). Multimodal 1D CNN for delamination prediction in CFRP drilling process with industrial robots. \u003cem\u003eComputers \u0026amp; Industrial Engineering\u003c/em\u003e, \u003cem\u003e190\u003c/em\u003e, 110074. https://www.sciencedirect.com/science/article/abs/pii/S0360835224001955\u003c/li\u003e\n\u003cli\u003eKrauss, P. (2024). Recurrent Neural Networks. In \u003cem\u003eArtificial Intelligence and Brain Research: Neural Networks, Deep Learning and the Future of Cognition\u003c/em\u003e (pp. 131-137). Berlin, Heidelberg: Springer Berlin Heidelberg. https://link.springer.com/chapter/10.1007/978-3-662-68980-6_14\u003c/li\u003e\n\u003cli\u003eChoi, J. G., Kim, D. C., Chung, M., Lim, S., \u0026amp; Park, H. W. (2024). Multimodal 1D CNN for delamination prediction in CFRP drilling process with industrial robots. \u003cem\u003eComputers \u0026amp; Industrial Engineering\u003c/em\u003e, \u003cem\u003e190\u003c/em\u003e, 110074. https://www.sciencedirect.com/science/article/abs/pii/S0360835224001955\u003c/li\u003e\n\u003cli\u003eZuvanov, L., Basso Garcia, A. L., Correr, F. H., Bizarria Jr, R., Filho, A. P. D. C., Da Costa, A. H., ... \u0026amp; Corr\u0026ecirc;a dos Santos, R. A. (2021). The experience of teaching introductory programming skills to bioscientists in Brazil. \u003cem\u003ePLoS computational biology\u003c/em\u003e, \u003cem\u003e17\u003c/em\u003e(11), e1009534. https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1009534\u003c/li\u003e\n\u003cli\u003e, H. N., Rimal, B., Pokhrel, N. R., Rimal, R., \u0026amp; Dahal, K. R. (2022). LSTM-SDM: An integrated framework of LSTM implementation for sequential data modeling. \u003cem\u003eSoftware Impacts\u003c/em\u003e, \u003cem\u003e14\u003c/em\u003e, 100396. https://www.sciencedirect.com/science/article/pii/S2665963822000902\u003c/li\u003e\n\u003cli\u003eSong, X., Salcianu, A., Song, Y., Dopson, D., \u0026amp; Zhou, D. (2020). Fast wordpiece tokenization. \u003cem\u003earXiv preprint arXiv:2012.15524\u003c/em\u003e. https://arxiv.org/abs/2012.15524\u003c/li\u003e\n\u003cli\u003eAlrasheedi, F., Zhong, X., \u0026amp; Huang, P. C. (2023). Padding module: Learning the padding in deep neural networks. \u003cem\u003eIEEE Access\u003c/em\u003e, \u003cem\u003e11\u003c/em\u003e, 7348-7357. https://ieeexplore.ieee.org/abstract/document/10021573/\u003c/li\u003e\n\u003cli\u003eSunny, M. A. I., Maswood, M. M. S., \u0026amp; Alharbi, A. G. (2020, October). Deep learning-based stock price prediction using LSTM and bi-directional LSTM model. In \u003cem\u003e2020 2nd novel intelligent and leading emerging sciences conference (NILES)\u003c/em\u003e (pp. 87-92). IEEE. https://ieeexplore.ieee.org/abstract/document/9257950\u003c/li\u003e\n\u003cli\u003eShahade, A. K., Walse, K. H., Thakare, V. M., \u0026amp; Atique, M. (2023). Multi-lingual opinion mining for social media discourses: An approach using deep learning-based hybrid fine-tuned smith algorithm with adam optimizer. \u003cem\u003eInternational Journal of Information Management Data Insights\u003c/em\u003e, \u003cem\u003e3\u003c/em\u003e(2), 100182. https://www.sciencedirect.com/science/article/pii/S2667096823000290\u003c/li\u003e\n\u003cli\u003eMart\u0026iacute;nez-Llop, P. G., Bobi, J. D. D. S., \u0026amp; Ortega, M. O. (2023). Time consideration in machine learning models for train comfort prediction using LSTM networks. \u003cem\u003eEngineering Applications of Artificial Intelligence\u003c/em\u003e, \u003cem\u003e123\u003c/em\u003e, 106303. https://www.sciencedirect.com/science/article/pii/S0952197623004876\u003c/li\u003e\n\u003cli\u003eFerruz, N., Heinzinger, M., Akdel, M., Goncearenco, A., Naef, L., \u0026amp; Dallago, C. (2023). From sequence to function through structure: Deep learning for protein design. \u003cem\u003eComputational and Structural Biotechnology Journal\u003c/em\u003e, \u003cem\u003e21\u003c/em\u003e, 238-250. https://www.sciencedirect.com/science/article/pii/S2001037022005086\u003c/li\u003e\n\u003cli\u003eWang, R., Jiang, Y., Jin, J., Yin, C., Yu, H., Wang, F., ... \u0026amp; Wei, L. (2023). DeepBIO: an automated and interpretable deep-learning platform for high-throughput biological sequence prediction, functional annotation and visualization analysis. \u003cem\u003eNucleic acids research\u003c/em\u003e, \u003cem\u003e51\u003c/em\u003e(7), 3017-3029. https://academic.oup.com/nar/article/51/7/3017/7041952#google_vignette\u003c/li\u003e\n\u003cli\u003eIoannou, G. N., Tang, W., Beste, L. A., Tincopa, M. A., Su, G. L., Van, T., ... \u0026amp; Waljee, A. K. (2020). Assessment of a deep learning model to predict hepatocellular carcinoma in patients with hepatitis C cirrhosis. \u003cem\u003eJAMA network open\u003c/em\u003e, \u003cem\u003e3\u003c/em\u003e(9), e2015626-e2015626. https://jamanetwork.com/journals/jamanetworkopen/article-abstract/2770062\u003c/li\u003e\n\u003cli\u003eAlbrijawi, M. T., \u0026amp; Alhajj, R. (2024). LSTM-driven drug design using SELFIES for target-focused de novo generation of HIV-1 protease inhibitor candidates for AIDS treatment. \u003cem\u003ePloS one\u003c/em\u003e, \u003cem\u003e19\u003c/em\u003e(6), e0303597. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0303597\u003c/li\u003e\n\u003cli\u003eAli, F., Kumar, H., Alghamdi, W., Kateb, F. A., \u0026amp; Alarfaj, F. K. (2023). Recent advances in machine learning-based models for prediction of antiviral peptides. \u003cem\u003eArchives of Computational Methods in Engineering\u003c/em\u003e, \u003cem\u003e30\u003c/em\u003e(7), 4033-4044. https://link.springer.com/article/10.1007/s11831-023-09933-w\u003c/li\u003e\n\u003cli\u003eTayebi, Z. (2020). Machine learning and deep learning to predict cross-immunoreactivity of viral epitopes. https://scholarworks.gsu.edu/cs_theses/96/ \u003c/li\u003e\n\u003cli\u003eNational Center for Biotechnology Information (NCBI). (n.d.). \u003cem\u003eNational Center for Biotechnology Information\u003c/em\u003e. U.S. National Library of Medicine. https://www.ncbi.nlm.nih.gov/\u003c/li\u003e\n\u003cli\u003eKent, W. J., Sugnet, C. W., Furey, T. S., Roskin, K. M., Pringle, T. H., Zahler, A. M., \u0026amp; Haussler, D. (2002). The human genome browser at UCSC. \u003cem\u003eGenome Research\u003c/em\u003e, 12(6), 996-1006. https://doi.org/10.1101/gr.229102\u003c/li\u003e\n\u003cli\u003eTamura, K., Stecher, G., \u0026amp; Kumar, S. (2021). MEGA11: Molecular Evolutionary Genetics Analysis version 11. \u003cem\u003eMolecular Biology and Evolution\u003c/em\u003e, 38(7), 3022-3027. https://doi.org/10.1093/molbev/msab120\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"LSTM, Generative model, RNN, Hepatitis B and Hepatitis C, Evolutionary studies, Sequence generation, ConV1d","lastPublishedDoi":"10.21203/rs.3.rs-5560102/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5560102/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eHepatitis B virus (HBV) and Hepatitis C virus (HCV) have always remained a greater global concern. Approximately 1.3 million deaths occur each year due to HBV and HCV. Due to the diverse genotypes and drug resistance, diagnostic challenges are being faced to treat these viruses. Therefore, the success ratio of the antiviral therapies has been decreasing with time in the last few decades. By deep learning predictive model, the pattern of evolution in hypervariable regions of HBV and HCV genes can be foreseen. In HCV, the hypervariable region is the Envelope glycoprotein (E2) gene, while in HBV, it includes the S1 and S2 genes. Generative models in deep learning have been used for evolutionary studies, but the application of these models is limited in viral research for predicting the evolving genotypes of viruses. The Long Short-Term Memory (LSTM) model represented a satisfactory outcome in predicting the sequences of the hypervariable genes of the evolving genotypes of the HCV and HBV genes that might be of a great help in diagnosis and vaccine design. We collected data from databases like NCBI and BVBRC. Our proposed LSTM generative model was trained on 1500 sequences of hypervariable genes of the present 7 genotypes of Hepatitis C and 10 genotypes of HBV. Apart from the traditional generative models like Recurrent Neural Network (RNN), our model not only generates the sequence but also learns and develops the relationship between various parts of the virus’s genetic code. In this study, three generative models were compared, Simple RNN, 1-Dimensional Convolutional Neural Network (ConV1d) and Long Short-Term Memory (LSTM). Among these three, LSTM demonstrated the least error rate with the highest efficiency and accuracy. While simple RNN and ConV1d illustrated relatively higher error rate and lower accuracy. LSTM gained success in reading long dependencies, hence, the proposed LSTM models are efficient at handling the sequential data along with preventing the conventional issue of losing the important information from the data, which happens frequently in generative models like Simple RNN and ConV1d.\u003c/p\u003e","manuscriptTitle":"Generative Deep Neural Networks for Estimating Hypervariability in Hepatitis B and C Virus Genomes","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-12-21 00:28:23","doi":"10.21203/rs.3.rs-5560102/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"1ee2ae0b-6a41-4ce4-8029-77858247273a","owner":[],"postedDate":"December 21st, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-01-07T08:53:07+00:00","versionOfRecord":[],"versionCreatedAt":"2024-12-21 00:28:23","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5560102","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5560102","identity":"rs-5560102","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.