VaxBot-HPV: A GPT-based Chatbot for Answering HPV Vaccine-related Questions

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
⚙ AI-generated summary by qwen3.7-flash, 2026-09-07 ⓘ

VaxBot-HPV, a GPT-based chatbot trained on HPV vaccine literature, demonstrated superior answer relevancy and faithfulness compared to baseline models in addressing vaccine-related queries.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

Abstract Background: HPV vaccine is an effective measure to prevent and control the diseases caused by Human Papillomavirus (HPV). This study addresses the development of VaxBot-HPV, a chatbot aimed at improving health literacy and promoting vaccination uptake by providing information and answering questions about the HPV vaccine; Methods: We constructed the knowledge base (KB) for VaxBot-HPV, which consists of 451 documents from biomedical literature and web sources on the HPV vaccine. We extracted 202 question-answer pairs from the KB and 39 questions generated by GPT-4 for training and testing purposes. To comprehensively understand the capabilities and potential of GPT-based chatbots, three models were involved in this study : GPT-3.5, VaxBot-HPV, and GPT-4. The evaluation criteria included answer relevancy and faithfulness; Results: VaxBot-HPV demonstrated superior performance in answer relevancy and faithfulness compared to baselines (Answer relevancy: 0.85; Faithfulness: 0.97) for the test questions in KB, (Answer relevancy: 0.85; Faithfulness: 0.96) for GPT generated questions; Conclusions: This study underscores the importance of leveraging advanced language models and fine-tuning techniques in the development of chatbots for healthcare applications, with implications for improving medical education and public health communication.
Full text 125,712 characters · extracted from preprint-html · click to expand
VaxBot-HPV: A GPT-based Chatbot for Answering HPV Vaccine-related Questions | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article VaxBot-HPV: A GPT-based Chatbot for Answering HPV Vaccine-related Questions Cui Tao, Yiming Li, Jianfu Li, Manqi Li, Evan Yu, Muhammad Amith, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4876692/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 3 You are reading this latest preprint version Abstract Background : HPV vaccine is an effective measure to prevent and control the diseases caused by Human Papillomavirus (HPV). This study addresses the development of VaxBot-HPV, a chatbot aimed at improving health literacy and promoting vaccination uptake by providing information and answering questions about the HPV vaccine; Methods : We constructed the knowledge base (KB) for VaxBot-HPV, which consists of 451 documents from biomedical literature and web sources on the HPV vaccine. We extracted 202 question-answer pairs from the KB and 39 questions generated by GPT-4 for training and testing purposes. To comprehensively understand the capabilities and potential of GPT-based chatbots, three models were involved in this study : GPT-3.5, VaxBot-HPV, and GPT-4. The evaluation criteria included answer relevancy and faithfulness; Results : VaxBot-HPV demonstrated superior performance in answer relevancy and faithfulness compared to baselines (Answer relevancy: 0.85; Faithfulness: 0.97) for the test questions in KB, (Answer relevancy: 0.85; Faithfulness: 0.96) for GPT generated questions; Conclusions : This study underscores the importance of leveraging advanced language models and fine-tuning techniques in the development of chatbots for healthcare applications, with implications for improving medical education and public health communication. Biological sciences/Computational biology and bioinformatics Biological sciences/Computational biology and bioinformatics/Machine learning Vaccine HPV vaccine Cervical Cancer GPT Large Language model QA system Chatbot Medical education Figures Figure 1 Figure 2 1. Introduction Human Papillomavirus (HPV) is a group of viruses that infect the skin and mucous membranes, with over 100 types identified [ 1 ]. HPV is primarily transmitted through sexual contact and can infect the genital area, leading to genital warts and various cancers, including cervical, anal, penile, vaginal, vulvar, and oropharyngeal cancers [ 2 ], [ 3 ], [ 4 ]. Among these, cervical cancer stands out as the most common HPV-related cancer and a leading cause of cancer-related deaths in women worldwide, contributing to an estimated 266,000 cervical cancer deaths annually due to HPV infection [ 5 ], [ 6 ], [ 7 ], [ 8 ], [ 9 ]. This burden is especially pronounced in low- and middle-income countries where access to screening and treatment is limited [ 5 ]. Similar to other infectious diseases, the development of HPV vaccines has also been a significant advancement in preventive healthcare [ 10 ], [ 11 ], [ 12 ]. HPV vaccines primarily target HPV types 16 and 18, which are responsible for approximately 70% of cervical cancers and a significant proportion of other HPV-related cancers [ 13 ]. By preventing HPV infection, these vaccines can effectively reduce the incidence of HPV-related diseases, including cervical cancer [ 13 ]. Clinical trials have demonstrated the high efficacy of HPV vaccines in preventing HPV infection and related diseases [ 14 ]. Furthermore, population-based studies have shown a substantial decline in HPV infections and HPV-related outcomes in countries with high HPV vaccination coverage, highlighting the real-world effectiveness of these vaccines [ 15 ]. Overall, HPV vaccines are a crucial tool in the prevention of HPV-related diseases, particularly cervical cancer. Widespread vaccination has the potential to significantly reduce the burden of HPV-related cancers and improve the overall health outcomes of populations globally [ 16 ]. Despite the proven benefits of HPV vaccination, there are various concerns and forms of hesitancy surrounding its use [ 17 ]. Some individuals and communities are hesitant due to insufficient and inadequate information about HPV vaccination or misinformation about the vaccine's safety and efficacy, often fueled by misinformation spread through social media and other channels [ 18 ], [ 19 ], [ 20 ]. Concerns about the long-term effects of the vaccine and its perceived necessity for individuals who may not consider themselves to be at high risk for HPV-related diseases also contribute to hesitancy [ 20 ]. Additionally, cultural or religious beliefs, distrust of pharmaceutical companies, and concerns about the vaccination's affordability and accessibility in low-resource settings can all play a role in vaccine hesitancy [ 21 ]. Addressing these concerns through accurate information, targeted education campaigns, and improved access to vaccination services is crucial in increasing HPV vaccination rates and reducing the burden of HPV-related diseases. Traditionally, question answering (QA) systems have been developed using rule-based approaches, information retrieval techniques, deep learning-based approaches, or hybrid methods [ 22 ], [ 23 ]. Rule-based QA systems rely on predefined rules and patterns to extract relevant information from a knowledge base or document collection in response to a question [ 24 ]. Tsampos and Marakakis, for example, developed a rule-based medical question answering system in Python using spaCy for natural language processing and Neo4j for graph database management [ 25 ]. They used Cypher queries to retrieve information from the graph database to answer user questions, and the system can handle complex questions by searching for relations between remote nodes and using synonyms to match nodes or paths [ 25 ]. Cairns et al. developed MiPACQ, a rule-based question answering system, by first retrieving candidate answer paragraphs using a paragraph-level baseline system based on the Lucene search engine [ 26 ]. The paragraphs were then re-ranked using a fixed formula that incorporated semantic annotations from the MiPACQ annotation pipeline [ 26 ]. This method utilized a scoring function that combined original paragraph scores with bag-of-words and UMLS entity components, ensuring that relevant paragraphs were prioritized for better question answering performance [ 26 ]. Information retrieval-based QA systems use keyword matching and ranking algorithms to retrieve documents or passages likely to contain the answer [ 27 ]. For example, Guo et al. developed a retrieval-based medical question answering system that efficiently retrieves answers using Elasticsearch and enhances them with semantic matching and knowledge graphs [ 28 ]. The system's novel siamese-based answer selection architecture outperformed baseline models and systems in both Chinese and English datasets, demonstrating consistent improvements in quantification and qualification evaluations [ 28 ]. Deep learning-based QA systems have emerged as a more flexible and adaptable approach, leveraging techniques such as powerful neural network architectures to automatically learn to understand and respond to questions [ 29 ]. Yin et al. developed Evebot, a conversational system for detecting negative emotions and preventing depression through positive suggestions [ 30 ]. It uses deep-learning models including a Bi-LSTM for emotion detection and an anti-language sequence-to-sequence neural network for counseling [ 30 ]. While these traditional QA systems have been effective for certain types of questions and domains, they have several limitations. One major limitation is the reliance of rule-based approaches on predefined rules or keywords, which makes them less flexible and adaptable to new or complex questions [ 31 ]. These systems also struggle with understanding natural language queries and context, often leading to inaccurate or incomplete answers. Additionally, traditional QA systems are limited by the quality and coverage of their underlying knowledge base or document collection, which can affect the accuracy and relevance of their answers [ 32 ]. For deep learning-based QA systems, one major limitation is their dependency on large amounts of labeled training data [ 33 ], [ 34 ], [ 35 ], [ 36 ]. These systems require vast datasets to learn patterns in language and develop accurate models, which can be challenging and resource-intensive to obtain, especially for specialized domains or languages [ 33 ]. Additionally, deep learning-based QA systems may struggle with out-of-domain or adversarial examples, where the input falls outside the scope of the training data, leading to errors or inaccurate responses [ 29 ], [ 37 ], [ 38 ]. Another limitation of traditional QA systems is their inability to provide explanations or reasoning behind their answers [ 39 ]. These systems typically return a single answer without any supporting context or evidence, making it challenging for users to understand how the answer was derived [ 40 ], [ 41 ]. This lack of transparency can reduce user trust and confidence in the system, especially in critical applications such as healthcare or legal domains [ 42 ]. Overall, while traditional QA systems have been valuable in certain contexts, their limitations have led to the development of more advanced approaches In recent years, the advent of powerful language models, such as the Generative Pre-trained Transformer (GPT), has revolutionized the field of natural language processing (NLP) and opened up new possibilities for conversational agents [ 35 ], [ 43 ], [ 44 ], [ 45 ]. GPT, developed by OpenAI, is a state-of-the-art deep learning model capable of generating human-like text based on the input it receives [ 35 ], [ 43 ], [ 44 ], [ 45 ]. The latest iteration, GPT-4, is distinguished by its ability to learn from vast amounts of text data, supported by its billions of parameters, enabling it to capture complex patterns in language and generate highly coherent and informative text [ 46 ], [ 47 , p. 4], [ 48 ]. However, a significant challenge with GPT models, including ChatGPT, is their tendency to produce hallucinations or responses that, while plausible, are factually incorrect [ 49 ]. This issue has raised concerns about the reliability of these models, especially in critical applications such as healthcare [ 50 ]. To address this problem, researchers and developers are investigating the use of well-curated knowledge bases (KBs) to refine the models. By integrating authenticated and reliable information from KBs, the goal is to enhance the model's capability to generate pertinent and accurate responses, thereby decreasing the risk of hallucinations. This has led to the development of chatbots and question answering systems powered by GPT that can provide information and assistance across various domains [ 48 ]. In the context of healthcare, the potential of GPT-powered question answering systems and chatbots is particularly promising [ 51 ]. Seenivasan et al. developed an end-to-end trainable Language-Vision GPT (LV-GPT) model to leverage GPT-based LLMs for Visual Question Answering (VQA) in robotic surgery [ 52 ]. The LV-GPT model extends GPT2 to process vision input (images) by incorporating a vision tokenizer and vision token embedding [ 52 ]. The model outperforms other state-of-the-art VQA models on public surgical-VQA datasets and a newly annotated dataset, demonstrating its effectiveness in capturing context from both language and vision modalities [ 52 ]. Shi et al. developed a GPT-based Question Answering System for Fundus Fluorescein Angiography (FFA) with an image-text alignment module and a GPT-based interactive QA module [ 53 ]. The system showed satisfactory performance in automatic evaluation and high accuracy and completeness in manual assessments, facilitating dynamic communication between ophthalmologists and patients for enhanced diagnostic processes [ 53 ]. Although GPT-powered question answering systems and chatbots in healthcare hold significant promise, we found that these systems exhibit hallucination issues because they use pre-trained GPT models directly without fine-tuning [ 53 ].In the case of HPV vaccination, where inadequate information and misconceptions are prevalent, leveraging fine-tuning techniques with advanced GPT models can significantly enhance the accuracy and reliability of information provided. A GPT-powered chatbot, when properly fine-tuned, could play a crucial role in educating the public and increasing awareness about the importance of vaccination. In this paper, we present the development and evaluation of a GPT-powered chatbot (VaxBot-HPV) designed to provide information and answer questions about the HPV vaccine. We also describe the design and implementation of the chatbot, its capabilities and limitations, as well as its potential impact on public health. Overall, this paper highlights the potential of GPT-powered question answering systems and chatbots in healthcare, particularly in the context of HPV vaccination, and demonstrates how these systems can be leveraged to improve health literacy and promote vaccination uptake. 2. Materials and Methods The study is structured around three primary stages. Initially, we constructed a KB and collected question-answer pairs relevant to the HPV vaccine within the KB to develop the benchmark. Subsequently, we inferred answers for the questions in the test benchmark using both pretrained GPT models and GPT models fine-tuned on the benchmark. Finally, we assessed the results in terms of faithfulness and answer relevancy. Figure 1 shows the overview of the study framework. 2.1. KB and Gold Standard Construction To build VaxBot-HPV, a chatbot designed to offer reliable information about the HPV vaccine, we first developed a KB deriving from peer-reviewed biomedical literature and web sources, resulting in a total of 451 documents on the HPV vaccine. To construct the question-answer pairs, we extracted 202 pairs of frequently asked questions and their answers related to the HPV vaccine from the collected webpages in the KB. The gold standard of question-answer sets was meticulously reviewed by domain experts to ensure their relevance and accuracy. 2.2. Models We utilized two state-of-the-art LLMs, GPT-3.5 and GPT-4, developed by OpenAI, as the key components of this study. GPT-3.5: GPT-3.5 is the iteration in OpenAI's series of large-scale language models, following the groundbreaking GPT-3. With an even larger model size (175 billion parameters) and enhanced capabilities, GPT-3.5 builds on the success of its predecessors in natural language processing (NLP) [ 54 ]. This advanced model exhibits impressive proficiency in understanding and generating human-like text, showcasing its potential for a wide range of applications including chatbots, content creation, and language translation [ 55 ]. GPT-4: GPT-4, the advancement in OpenAI's renowned Generative Pre-trained Transformer series, marks a significant milestone in the field of natural language processing (NLP). With its remarkable increase to 170 trillion parameters, GPT-4 surpasses its predecessor, GPT-3, enabling it to tackle even more complex language tasks with improved accuracy and understanding [ 54 ]. This model represents a significant leap forward in NLP capabilities, holding the potential to revolutionize various fields, from conversational AI to content generation and beyond [ 55 ]. 2.3. Experiment Setup In this study, VaxBot-HPV, fine-tuned on GPT-3.5, was developed using question-answer pairs from both a knowledge base and GPT-generated questions. The question-answer pairs derived from the knowledge base were divided into 162 samples for training purposes and 40 for testing purposes. To enhance question diversity and ensure the generalizability of our findings, we employed GPT-4 models to generate 80 question-answer pairs. After careful review, we included 39 questions based on their relevance to the HPV vaccine and manually updated their answers for inclusion into our study. Among the GPT-generated questions, 28 question-pairs were randomly selected for training and the rest for testing. The parameters of the VaxBot-HPV are outlined in Table 1 . We used the following prompt to instruct the GPT models in answering the query: “ You are an expert Q&A system that is trusted around the world. Always answer the query using the provided context information, and not prior knowledge. Some rules to follow : 1. Never directly reference the given context in your answer. 2. Avoid statements like 'Based on the context, ...' or 'The context information ...' or anything along those lines.” The prompts for GPT-4 to generate questions were as follows: Using the provided context from referencing articles on HPV vaccine, formulate a question that captures an important fact from the context. Restrict the question to the context information provided. Please only output the question. VaxBot-HPV's development involved the performance comparison of various models, including GPT-3.5, as well as GPT-4, for each experimental set. Table 1 Parameters of VaxBot-HPV. Parameter Value n_epochs 2 batch_size 1 learning_rate_multiplier 1 temperature 0.3 context_window 2,048 Token limit 4,096 The experiments were carried out using a high-performance server containing 8 Nvidia A100 GPUs, each with a memory capacity of 80GB. This server configuration facilitated the effective training and evaluation of the models, ensuring the production of dependable and precise results. 2.4. Evaluation The evaluation involved answer relevancy and faithfulness. Both are critical aspects in assessing the quality of generated responses. Answer relevancy gauges the extent to which the answers align with the questions, while faithfulness ensures factual accuracy, a fundamental requirement for reliable information retrieval. These metrics collectively provide a comprehensive evaluation of the model's performance in understanding and responding to user queries. Additionally, evaluations were conducted to thoroughly assess the system's effectiveness. The assessment of all outcomes was carried out using the Ragas metrics, which are GPT-supported measures widely adopted in NLP tasks to evaluate the quality of generated text [ 56 ]. Specifically, the RAGAS metrics calculate answer relevancy and faithfulness through a detailed process. For answer relevancy, the ground truth answer and the generated answer are vectorized using the specified embedding model, and their cosine similarity is computed to determine alignment [ 56 ]. For answer faithfulness, the process involves quantifying factual correctness by identifying true positives (facts present in both the ground truth and the generated answer), false positives (facts present in the generated answer but not in the ground truth), and false negatives (facts present in the ground truth but not in the generated answer) [ 56 ]. The F1 score is then used to quantify correctness based on these values [ 56 ]. A weighted average of factual correctness and semantic similarity provides the final score [ 56 ]. 3. Results Table 2 illustrates the automatic performance evaluation of different GPT models in answer relevancy and faithfulness on the questions extracted from the KB. The results indicate that the VaxBot-HPV outperformed both the GPT-3.5 andGPT-4 models in terms of answer relevancy, achieving a score of 0.85 compared to 0.80 and 0.83, respectively. Similarly, the VaxBot-HPV exhibited higher faithfulness, scoring 0.97, compared to 0.92 for the GPT-3.5 model and 0.91 for the GPT-4 model. These results suggest that fine-tuning the GPT-3.5 model leads to improved performance in both answer relevancy and faithfulness compared to using the models in their pretrained states. Table 2 Performance evaluation of different GPT models in answer relevancy and faithfulness on the questions extracted from the knowledge base. Model Answer Relevancy Faithfulness GPT-3.5 0.80 0.92 VaxBot-HPV 0.85 0.97 GPT-4 0.83 0.91 Table 3 presents the performance evaluation of different GPT models in terms of answer relevancy and faithfulness on questions generated by GPT-4. The GPT-3.5 model achieved an answer relevancy score of 0.80 and a faithfulness score of 0.90. In comparison, VaxBot-HPV showed improved performance with an answer relevancy score of 0.85 and a faithfulness score of 0.96. These results highlight the benefits of fine-tuning the GPT model, demonstrating its broader generalizability, applicability and robustness. Table 3 Performance evaluation of different GPT models in answer relevancy and faithfulness on the questions generated by GPT-4. Model Answer Relevancy Faithfulness GPT-3.5 0.80 0.90 VaxBot-HPV 0.85 0.96 Figure 2 shows two samples of questions and its generated questions by the four systems. We selected two questions. One question (“What are the risks of cervical cancer besides pregnancy at an early age?”) is generated by GPT, and another question (“What are the risks of the HPV vaccine?”) is from the test benchmark. VaxBot-HPV demonstrates an advantage in providing comprehensive and accurate responses to health-related inquiries compared to other systems. For instance, when asked about the risks of cervical cancer besides early pregnancy, VaxBot-HPV effectively listed multiple risk factors, including having multiple sexual partners, weakened immune systems, and specific health conditions. In contrast, the GPT-3.5 failed to identify any additional risk factors, while the GPT-4 provided information not directly to the question, such as “genital warts occurred most in adolescents and young adults”, which could be misleading. Additionally, ChatGPT, although comprehensive, was not succinct and failed to answer the question directly. Furthermore, when it comes to the question ”What are the risks of the HPV vaccine?”, VaxBot-HPV effectively summarized over 12 years of safety monitoring, highlighted common and rare side effects, and provided actionable advice on preventing fainting-related injuries, all while maintaining a clear and concise format. In contrast, the GPT-3.5 and GPT-4, though accurate, lacked depth, information sources and reassurance, merely listing side effects without addressing common myths or providing detailed context. ChatGPT-4, despite its comprehensiveness, often failed to deliver succinct answers, resulting in verbose responses that lacked focus. These examples illustrate that VaxBot-HPV not only enhances the specificity and clarity of responses but also ensures that users receive accurate, reliable and actionable health information efficiently. 4. Discussion The development and evaluation of VaxBot-HPV, a chatbot designed to provide information and answer questions about the HPV vaccine, demonstrates the potential of advanced language models, particularly GPT-3.5 and GPT-4, in healthcare applications. Compared to traditional QA systems, VaxBot-HPV leverages the capabilities of GPT models, especially after fine-tuning, to generate relevant and accurate responses to user queries. VaxBot-HPV has a substantial advantage over existing pre-trained GPT models. This extensive pre-training allows VaxBot-HPV to have a deeper understanding of language and context, enabling it to provide more relevant and accurate answers to user queries. Unlike rule-based systems, which rely on predefined rules and patterns, and retrieval-based systems, which use keyword matching and ranking algorithms, VaxBot-HPV's pretrained model allows it to generate responses based on a broader understanding of the topic. This capability enhances the chatbot's ability to address a wide variety of questions and provide more informative and helpful responses to users. Moreover, VaxBot-HPV allows the answers to be dynamically generated, potentially offering more tailored responses to users compared to standard, one-size-fits-all answers. The fine-tuning process further enhances VaxBot-HPV's performance, particularly in the context of HPV vaccination, by adapting it to the specific domain. This adaptation improves answer relevancy and faithfulness, addressing common issues of ChatGPT such as hallucinations, where the model generates plausible but inaccurate responses. By fine-tuning on a dataset specific to HPV vaccination, VaxBot-HPV can learn the nuances of the topic, including relevant terminology, common misconceptions, and specific concerns that users may have. This specificity allows the chatbot to provide more accurate and tailored responses, increasing its overall effectiveness in addressing user queries related to the HPV vaccine. Furthermore, the fine-tuning process helps mitigate bias and misinformation that may be present in generic language models, ensuring that VaxBot-HPV provides reliable and trustworthy information to users seeking information about HPV vaccination. Additionally, the specific fine-tuning, which includes context in addition to question-answer pairs, enables VaxBot-HPV to extend beyond just answering questions. These models can also provide explanations or reasoning behind their answers, increasing transparency and user trust. This feature is particularly important in healthcare applications, where understanding the rationale behind medical advice is crucial for informed decision-making. In terms of evaluations, incorporating multiple sources, including questions generated by GPT models, strengthens the credibility and reliability of our findings regarding VaxBot-HPV's performance. By leveraging questions from diverse sources, we were able to assess the chatbot's ability to handle a wide range of queries beyond those explicitly included in the knowledge base. This comprehensive evaluation approach not only ensures the robustness of our results but also demonstrates VaxBot-HPV's versatility in addressing various user inquiries. Overall, the use of multiple evaluation metrics underscores the effectiveness and adaptability of VaxBot-HPV in providing reliable information and support to users. While VaxBot-HPV demonstrates promising performance, there are several limitations to consider. First, the chatbot's effectiveness is contingent on the quality and comprehensiveness of the underlying knowledge base. Incomplete or inaccurate information in the KB could lead to erroneous or insufficient responses from the chatbot. Additionally, the chatbot's reliance on text-based interactions may limit its accessibility to individuals with visual or cognitive impairments who may benefit from alternative communication methods. Moreover, the inclusion of manual evaluations is needed to provide a holistic assessment of the chatbot's performance, enhancing the depth and accuracy of our conclusions. Furthermore, the evaluation of VaxBot-HPV was primarily based on its performance in answering questions, overlooking other aspects of user interaction such as ease of use, user satisfaction or engagement. Finally, the generalizability of our findings may be limited to the specific domain of HPV vaccination and may not extend to other healthcare contexts. Future research could focus on several areas to enhance the capabilities and impact of VaxBot-HPV. First, expanding the knowledge base to include a broader range of topics related to HPV vaccination and addressing emerging concerns or misconceptions could improve the chatbot's effectiveness and relevance. Second, integrating multimedia capabilities, such as image or video recognition, could enhance the chatbot's ability to provide information and support in a more interactive and engaging manner. Third, incorporating feedback mechanisms to gather user input and improve the chatbot's responses over time could enhance its usability and user satisfaction. Furthermore, exploring the integration of VaxBot-HPV with existing healthcare systems or platforms could facilitate its adoption and integration into clinical workflows, potentially improving access to information and promoting HPV vaccination uptake. Fourth, extending the chatbot to cover other types of vaccines and medical domains could broaden its applicability and utility, making it a more versatile tool for addressing public health concerns. Lastly, we need to add a user interface to VaxBot-HPV to make it more accessible and user-friendly, enhancing the overall user experience and encouraging more people to use the chatbot for reliable information on HPV vaccination and related topics. 5. Conclusions In conclusion, the development of VaxBot-HPV demonstrates the potential of GPT-powered chatbots in healthcare, particularly in promoting vaccination uptake and addressing common concerns and misconceptions. The study also underscores the importance of leveraging advanced language models and fine-tuning techniques in healthcare chatbot development. The efficacy of VaxBot-HPV highlights the transformative impact of such technologies on medical education, healthcare communication and information dissemination. Declarations Acknowledgements: C.T. discloses support for the research of this work from National Institute of Allergy And Infectious Diseases of the National Institutes of Health [grant number R01AI130460 and U24AI171008], and CPRIT [grant number RP220244]. Author Contributions: Methodology, Y.L., J.L., C.T.; software, Y.L., and J.L.; validation, L.T., and L.S.; formal analysis, M.L., Y.L., and J.L.; investigation, M.A., E.Y.; resources, C.T.; data curation, Y.L.; writing—original draft preparation, Y.L.; writing—review and editing, C.T.; visualization, Y.L.; supervision, C.T.; project administration, C.T.; funding acquisition, C.T., L.C. All authors have read and agreed to the published version of the manuscript. Competing Interests: The authors declare no conflicts of interest. Data Availability: Dataset available on request from the authors. Computer Code: Computer code available on request from the authors. Institutional Review Board Statement: Not applicable. Informed Consent Statement: Not applicable. References M. das G. P. Leto, G. F. dos Santos Júnior, A. M. Porro, and J. Tomimori, “Human papillomavirus infection: etiopathogenesis, molecular biology and clinical manifestations,” An. Bras. Dermatol. , vol. 86, pp. 306–317, Apr. 2011, doi: 10.1590/S0365-05962011000200014 . P. Brianti, E. De Flammineis, and S. R. Mercuri, “Review of HPV-related diseases and cancers,” New Microbiol , vol. 40, no. 2, pp. 80–85, Apr. 2017. F. S. Alhamlan, M. B. Alfageeh, M. A. Al Mushait, I. A. Al-Badawi, and M. N. Al-Ahdal, “Human Papillomavirus-Associated Cancers,” in Microbial Pathogenesis: Infection and Immunity , U. Kishore, Ed., Cham: Springer International Publishing, 2021, pp. 1–14. doi: 10.1007/978-3-030-67452-6_1 . C. Chelimo, T. A. Wouldes, L. D. Cameron, and J. M. Elwood, “Risk factors for and prevention of human papillomaviruses (HPV), genital warts and cervical cancer,” Journal of Infection , vol. 66, no. 3, pp. 207–217, Mar. 2013, doi: 10.1016/j.jinf.2012.10.024 . R. Hull et al. , “Cervical cancer in low and middle–income countries (Review),” Oncology Letters , vol. 20, no. 3, pp. 2058–2074, Sep. 2020, doi: 10.3892/ol.2020.11754 . K. S. Okunade, “Human papillomavirus and cervical cancer,” Journal of Obstetrics and Gynaecology , Jul. 2020, Accessed: Mar. 27, 2024. [Online]. Available: https://www.tandfonline.com/doi/abs/ 10.1080/01443615.2019.1634030 M. Arbyn et al. , “Estimates of incidence and mortality of cervical cancer in 2018: a worldwide analysis,” The Lancet Global Health , vol. 8, no. 2, pp. e191–e203, Feb. 2020, doi: 10.1016/S2214-109X(19)30482-6 . M. Arbyn et al. , “Worldwide burden of cervical cancer in 2008,” Annals of Oncology , vol. 22, no. 12, pp. 2675–2686, Dec. 2011, doi: 10.1093/annonc/mdr015 . E. Tesfaye et al. , “Prevalence of human papillomavirus infection and associated factors among women attending cervical cancer screening in setting of Addis Ababa, Ethiopia,” Scientific Reports , vol. 14, no. 1, p. 4053, Feb. 2024, doi: 10.1038/s41598-024-54754-x . S. S. Ali, A. Y. Nirupama, S. Chaudhuri, and G. V. S. Murthy, “Therapeutic HPV Vaccination: A Strategy for Cervical Cancer Elimination in India,” Indian J Gynecol Oncolog , vol. 22, no. 2, p. 38, Mar. 2024, doi: 10.1007/s40944-024-00800-5 . Y. Li et al. , “Unpacking adverse events and associations post COVID-19 vaccination: a deep dive into vaccine adverse event reporting system data,” Expert Review of Vaccines , vol. 23, no. 1, pp. 53–59, Dec. 2024, doi: 10.1080/14760584.2023.2292203 . Li Y, Li J, Dang Y, Chen Y, and Tao C, “Temporal and Spatial Analysis of COVID-19 Vaccines Using Reports from Vaccine Adverse Event Reporting System,” JMIR Preprints , doi: 10.2196/preprints.51007 . L. Iqbal, M. Jehan, and S. Azam, “Advancements in mRNA Vaccines: A Promising Approach for Combating Human Papillomavirus-Related Cancers,” Cancer Control , vol. 31, p. 10732748241238628, Jan. 2024, doi: 10.1177/10732748241238629 . C. A. Gonçalves, G. Pereira-da-Silva, R. C. C. P. Silveira, P. C. M. Mayer, A. Zilly, and L. C. Lopes-Júnior, “Safety, Efficacy, and Immunogenicity of Therapeutic Vaccines for Patients with High-Grade Cervical Intraepithelial Neoplasia (CIN 2/3) Associated with Human Papillomavirus: A Systematic Review,” Cancers , vol. 16, no. 3, Art. no. 3, Jan. 2024, doi: 10.3390/cancers16030672 . E. M. Webster et al. , “Building knowledge using a novel web-based intervention to promote HPV vaccination in a diverse, low-income population,” Gynecologic Oncology , vol. 181, pp. 102–109, Feb. 2024, doi: 10.1016/j.ygyno.2023.12.005 . D. R. Lowy and J. T. Schiller, “Reducing HPV-Associated Cancer Globally,” Cancer Prevention Research , vol. 5, no. 1, pp. 18–23, Jan. 2012, doi: 10.1158/1940-6207.CAPR-11-0542 . P. G. Szilagyi et al. , “Prevalence and characteristics of HPV vaccine hesitancy among parents of adolescents across the US,” Vaccine , vol. 38, no. 38, pp. 6027–6037, Aug. 2020, doi: 10.1016/j.vaccine.2020.06.074 . K. H. Nguyen et al. , “Parental vaccine hesitancy and its association with adolescent HPV vaccination,” Vaccine , vol. 39, no. 17, pp. 2416–2423, Apr. 2021, doi: 10.1016/j.vaccine.2021.03.048 . W. Jennings et al. , “Lack of Trust, Conspiracy Beliefs, and Social Media Use Predict COVID-19 Vaccine Hesitancy,” Vaccines , vol. 9, no. 6, Art. no. 6, Jun. 2021, doi: 10.3390/vaccines9060593 . F. Gauna, P. Verger, L. Fressard, M. Jardin, J. K. Ward, and P. Peretti-Watel, “Vaccine hesitancy about the HPV vaccine among French young women and their parents: a telephone survey,” BMC Public Health , vol. 23, no. 1, p. 628, Apr. 2023, doi: 10.1186/s12889-023-15334-2 . G. Adeyanju, “Behavioral Insights into Vaccine Hesitancy Determinants in Sub-Saharan Africa,” Sep. 2022, Accessed: Mar. 28, 2024. [Online]. Available: https://www.db-thueringen.de/receive/dbt_mods_00053424 Y. Chen and F. Zulkernine, “BIRD-QA: A BERT-based Information Retrieval Approach to Domain Specific Question Answering,” in 2021 IEEE International Conference on Big Data (Big Data) , Dec. 2021, pp. 3503–3510. doi: 10.1109/BigData52589.2021.9671523 . G. Vanitha, S. Sanampudi, and M. I.LAKSHMI, “APPROCHES FOR QUESTION ANSWERING SYSTEMS,” International Journal of Engineering Science and Technology , vol. 3, Feb. 2011. I. Thalib, Widyawan, and I. Soesanti, “A Review on Question Analysis, Document Retrieval and Answer Extraction Method in Question Answering System,” in 2020 International Conference on Smart Technology and Applications (ICoSTA) , Feb. 2020, pp. 1–5. doi: 10.1109/ICoSTA48221.2020.1570614175 . I. Tsampos and E. Marakakis, “A Medical Question Answering System with NLP and graph database”. B. L. Cairns et al. , “The MiPACQ Clinical Question Answering System,” AMIA Annu Symp Proc , vol. 2011, pp. 171–180, 2011. X. Feng, Q. Liu, C. Lao, and D. Sun, “Design and Implementation of Automatic Question Answering System in Information Retrieval,” in Proceedings of the 7th International Conference on Informatics, Environment, Energy and Applications , in IEEA ’18. New York, NY, USA: Association for Computing Machinery, Mar. 2018, pp. 207–211. doi: 10.1145/3208854.3208862 . Q. Guo, S. Cao, and Z. Yi, “A medical question answering system using large language models and knowledge graphs,” International Journal of Intelligent Systems , vol. 37, no. 11, pp. 8548–8564, 2022, doi: 10.1002/int.22955 . N. Saeed, humaira ashraf, and N. Jhanjhi, “DEEP LEARNING BASED QUESTION ANSWERING SYSTEM (SURVEY),” Preprints , Dec. 2023, doi: 10.20944/preprints202312.1739.v1 . J. Yin, Z. Chen, K. Zhou, and C. Yu, “A Deep Learning Based Chatbot for Campus Psychological Therapy.” 2019. F. Khennouche, Y. Elmir, Y. Himeur, N. Djebari, and A. Amira, “Revolutionizing generative pre-traineds: Insights and challenges in deploying ChatGPT and generative chatbots for FAQs,” Expert Systems with Applications , vol. 246, p. 123224, Jul. 2024, doi: 10.1016/j.eswa.2024.123224 . P. Yin, N. Duan, B. Kao, J. Bao, and M. Zhou, “Answering Questions with Complex Semantic Constraints on Open Knowledge Bases,” in Proceedings of the 24th ACM International on Conference on Information and Knowledge Management , in CIKM ’15. New York, NY, USA: Association for Computing Machinery, Oct. 2015, pp. 1301–1310. doi: 10.1145/2806416.2806542 . A. Abdallah, B. Piryani, and A. Jatowt, “Exploring the state of the art in legal QA systems,” J Big Data , vol. 10, no. 1, p. 127, Aug. 2023, doi: 10.1186/s40537-023-00802-8 . Y. Li et al. , “Development of a Natural Language Processing Tool to Extract Acupuncture Point Location Terms,” in 2023 IEEE 11th International Conference on Healthcare Informatics (ICHI) , Jun. 2023, pp. 344–351. doi: 10.1109/ICHI57859.2023.00053 . Y. Li et al. , “Artificial intelligence-powered pharmacovigilance: A review of machine and deep learning in clinical text-based adverse drug event detection for benchmark datasets,” Journal of Biomedical Informatics , vol. 152, p. 104621, Apr. 2024, doi: 10.1016/j.jbi.2024.104621 . J. He et al. , “Prompt Tuning in Biomedical Relation Extraction,” J Healthc Inform Res , Feb. 2024, doi: 10.1007/s41666-024-00162-9 . E. Stroh and P. Mathur, “Question Answering Using Deep Learning”. J. Li et al. , “Mapping Vaccine Names in Clinical Trials to Vaccine Ontology using Cascaded Fine-Tuned Domain-Specific Language Models,” Res Sq , p. rs.3.rs-3362256, Sep. 2023, doi: 10.21203/rs.3.rs-3362256/v1 . P. Lu et al. , “Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering”. J. Lin et al. , “What Makes a Good Answer? The Role of Context in Question Answering”. S. Min, V. Zhong, R. Socher, and C. Xiong, “Efficient and Robust Question Answering from Minimal Context over Documents.” 2018. N. Goyal et al. , “What Else Do I Need to Know? The Effect of Background Information on Users’ Reliance on QA Systems.” 2023. Y. Li, J. Li, J. He, and C. Tao, “AE-GPT: Using Large Language Models to extract adverse events from surveillance reports-A use case with influenza vaccine adverse events,” PLOS ONE , vol. 19, no. 3, p. e0300919, Mar. 2024, doi: 10.1371/journal.pone.0300919 . Y. Hu et al. , “Zero-shot Clinical Entity Recognition using ChatGPT,” arXiv.org , 2023, doi: 10.48550/arXiv.2303.16416 . Y. Li et al. , “Relation Extraction Using Large Language Models: A Case Study on Acupuncture Point Locations,” arXiv.org , 2024, doi: 10.48550/arXiv.2404.05415 . E. Chang, Examining GPT-4’s Capabilities and Enhancement with SocraSynth . 2023. K. S. Kalyan, “A survey of GPT-3 family large language models including ChatGPT and GPT-4,” Natural Language Processing Journal , vol. 6, p. 100048, Mar. 2024, doi: 10.1016/j.nlp.2023.100048 . T. M. Al-Hasan, A. N. Sayed, F. Bensaali, Y. Himeur, I. Varlamis, and G. Dimitrakopoulos, “From Traditional Recommender Systems to GPT-Based Chatbots: A Survey of Recent Developments and Future Directions,” Big Data and Cognitive Computing , vol. 8, no. 4, Art. no. 4, Apr. 2024, doi: 10.3390/bdcc8040036 . J. Li, X. Cheng, X. Zhao, J.-Y. Nie, and J.-R. Wen, “HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , H. Bouamor, J. Pino, and K. Bali, Eds., Singapore: Association for Computational Linguistics, Dec. 2023, pp. 6449–6464. doi: 10.18653/v1/2023.emnlp-main.397 . T. R. McIntosh, T. Liu, T. Susnjak, P. Watters, A. Ng, and M. N. Halgamuge, “A Culturally Sensitive Test to Evaluate Nuanced GPT Hallucination,” IEEE Transactions on Artificial Intelligence , pp. 1–13, 2023, doi: 10.1109/TAI.2023.3332837 . S. García-Méndez and F. de Arriba-Pérez, “Large Language Models and Healthcare Alliance: Potential and Challenges of Two Representative Use Cases,” Ann Biomed Eng , Feb. 2024, doi: 10.1007/s10439-024-03454-8 . L. Seenivasan, M. Islam, G. Kannan, and H. Ren, “SurgicalGPT: End-to-End Language-Vision GPT for Visual Question Answering in Surgery,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2023 , H. Greenspan, A. Madabhushi, P. Mousavi, S. Salcudean, J. Duncan, T. Syeda-Mahmood, and R. Taylor, Eds., Cham: Springer Nature Switzerland, 2023, pp. 281–290. doi: 10.1007/978-3-031-43996-4_27 . D. Shi et al. , FFA-GPT: an Interactive Visual Question Answering System for Fundus Fluorescein Angiography . 2023. doi: 10.21203/rs.3.rs-3307492/v1 . A. Koubaa, “GPT-4 vs. GPT-3.5: A Concise Showdown,” Preprints , Mar. 2023, doi: 10.20944/preprints202303.0422.v1 . K. Nayanam and V. Sharma, “TOWARDS ARCHITECTING RESEARCH PERSPECTIVE FUTURE SCOPE WITH CHAT GPT,” Jul. 2024. “Metrics | Ragas.” Accessed: Mar. 08, 2024. [Online]. Available: https://docs.ragas.io/en/latest/concepts/metrics/index.html Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editor assigned by journal 14 Aug, 2024 Submission checks completed at journal 14 Aug, 2024 First submitted to journal 07 Aug, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4876692","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":340267656,"identity":"96fc2d59-4669-4b97-ac6e-7d99dfc2cdc3","order_by":0,"name":"Cui Tao","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAwElEQVRIiWNgGAWjYBAC+wNA4geDhByUz0xYCxsQM/YwSBiTpoWZhYEhsYF4LeyHDzAztlmkr20/Y/iAocIaphePFp60BKAWidxtZ3KMDRjOpBOhhSHHAKLlQI6ZBGPbYSK08L8Ba0k3O/8GqOUfMVokgLawtkkkmN0A2dJAlJZnCQd7zkkYbrvxrNgg4Vi6MREOSz744EdZnbzZ+eSNDz7UWMsS1AICByAUhwFDAjHKkQD7AxI1jIJRMApGwUgBABOeOUGF03KAAAAAAElFTkSuQmCC","orcid":"","institution":"Mayo Clinic","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Cui","middleName":"","lastName":"Tao","suffix":""},{"id":340267657,"identity":"26c287d1-1b60-4ca0-adaa-809287aa67ee","order_by":1,"name":"Yiming Li","email":"","orcid":"","institution":"The University of Texas Health Science Center at Houston","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yiming","middleName":"","lastName":"Li","suffix":""},{"id":340267658,"identity":"07767304-bc34-4c37-8f8d-3b85ffa21d3e","order_by":2,"name":"Jianfu Li","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jianfu","middleName":"","lastName":"Li","suffix":""},{"id":340267659,"identity":"e71f43d2-ba90-4ffd-8601-becf84051462","order_by":3,"name":"Manqi Li","email":"","orcid":"","institution":"The University of Texas Health Science Center at Houston","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Manqi","middleName":"","lastName":"Li","suffix":""},{"id":340267660,"identity":"4ce9a589-b137-4ff6-a757-efd1033dd582","order_by":4,"name":"Evan Yu","email":"","orcid":"","institution":"The University of Texas Health Science Center at Houston","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Evan","middleName":"","lastName":"Yu","suffix":""},{"id":340267661,"identity":"bc3fc3e2-9b89-41a5-9354-47a0eaf00bc7","order_by":5,"name":"Muhammad Amith","email":"","orcid":"","institution":"The University of Texas Medical Branch at Galveston","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Muhammad","middleName":"","lastName":"Amith","suffix":""},{"id":340267662,"identity":"ae88c5e0-f293-40a1-9867-260cb9102b4e","order_by":6,"name":"Lu Tang","email":"","orcid":"","institution":"Texas A\u0026M University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Lu","middleName":"","lastName":"Tang","suffix":""},{"id":340267663,"identity":"b54ebc52-2b08-448d-83f7-c808c0857bac","order_by":7,"name":"Lara Savas","email":"","orcid":"","institution":"The University of Texas Health Science Center at Houston","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Lara","middleName":"","lastName":"Savas","suffix":""},{"id":340267664,"identity":"55ebac6e-46f1-4e12-b95c-a8429709d231","order_by":8,"name":"Licong Cui","email":"","orcid":"","institution":"The University of Texas Health Science Center at Houston","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Licong","middleName":"","lastName":"Cui","suffix":""}],"badges":[],"createdAt":"2024-08-07 19:08:33","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4876692/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4876692/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":64325704,"identity":"2eb6399a-01ef-47d7-a827-4561536cce2e","added_by":"auto","created_at":"2024-09-11 16:30:52","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":114147,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of the framework.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-4876692/v1/1d56cf9e734442e7e8f5af6c.png"},{"id":64326063,"identity":"f0a18285-94cc-4a03-935f-9d2cd652b3ae","added_by":"auto","created_at":"2024-09-11 16:38:52","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1710735,"visible":true,"origin":"","legend":"\u003cp\u003eSamples of questions and answers (a) GPT generated question (b) question in the test benchmark over systems\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-4876692/v1/bc3d471f427957194b01a370.png"},{"id":64326089,"identity":"d5f36361-1fe5-4541-9478-ea6e0a8cc447","added_by":"auto","created_at":"2024-09-11 16:38:59","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2193654,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4876692/v1/c1333511-3b87-46a9-b9eb-8c426c1caf63.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"VaxBot-HPV: A GPT-based Chatbot for Answering HPV Vaccine-related Questions","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eHuman Papillomavirus (HPV) is a group of viruses that infect the skin and mucous membranes, with over 100 types identified [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. HPV is primarily transmitted through sexual contact and can infect the genital area, leading to genital warts and various cancers, including cervical, anal, penile, vaginal, vulvar, and oropharyngeal cancers [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e], [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Among these, cervical cancer stands out as the most common HPV-related cancer and a leading cause of cancer-related deaths in women worldwide, contributing to an estimated 266,000 cervical cancer deaths annually due to HPV infection [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. This burden is especially pronounced in low- and middle-income countries where access to screening and treatment is limited [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSimilar to other infectious diseases, the development of HPV vaccines has also been a significant advancement in preventive healthcare [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. HPV vaccines primarily target HPV types 16 and 18, which are responsible for approximately 70% of cervical cancers and a significant proportion of other HPV-related cancers [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. By preventing HPV infection, these vaccines can effectively reduce the incidence of HPV-related diseases, including cervical cancer [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Clinical trials have demonstrated the high efficacy of HPV vaccines in preventing HPV infection and related diseases [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Furthermore, population-based studies have shown a substantial decline in HPV infections and HPV-related outcomes in countries with high HPV vaccination coverage, highlighting the real-world effectiveness of these vaccines [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Overall, HPV vaccines are a crucial tool in the prevention of HPV-related diseases, particularly cervical cancer. Widespread vaccination has the potential to significantly reduce the burden of HPV-related cancers and improve the overall health outcomes of populations globally [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eDespite the proven benefits of HPV vaccination, there are various concerns and forms of hesitancy surrounding its use [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Some individuals and communities are hesitant due to insufficient and inadequate information about HPV vaccination or misinformation about the vaccine's safety and efficacy, often fueled by misinformation spread through social media and other channels [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e], [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Concerns about the long-term effects of the vaccine and its perceived necessity for individuals who may not consider themselves to be at high risk for HPV-related diseases also contribute to hesitancy [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Additionally, cultural or religious beliefs, distrust of pharmaceutical companies, and concerns about the vaccination's affordability and accessibility in low-resource settings can all play a role in vaccine hesitancy [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Addressing these concerns through accurate information, targeted education campaigns, and improved access to vaccination services is crucial in increasing HPV vaccination rates and reducing the burden of HPV-related diseases.\u003c/p\u003e \u003cp\u003eTraditionally, question answering (QA) systems have been developed using rule-based approaches, information retrieval techniques, deep learning-based approaches, or hybrid methods [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Rule-based QA systems rely on predefined rules and patterns to extract relevant information from a knowledge base or document collection in response to a question [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Tsampos and Marakakis, for example, developed a rule-based medical question answering system in Python using spaCy for natural language processing and Neo4j for graph database management [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. They used Cypher queries to retrieve information from the graph database to answer user questions, and the system can handle complex questions by searching for relations between remote nodes and using synonyms to match nodes or paths [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Cairns et al. developed MiPACQ, a rule-based question answering system, by first retrieving candidate answer paragraphs using a paragraph-level baseline system based on the Lucene search engine [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. The paragraphs were then re-ranked using a fixed formula that incorporated semantic annotations from the MiPACQ annotation pipeline [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. This method utilized a scoring function that combined original paragraph scores with bag-of-words and UMLS entity components, ensuring that relevant paragraphs were prioritized for better question answering performance [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. Information retrieval-based QA systems use keyword matching and ranking algorithms to retrieve documents or passages likely to contain the answer [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. For example, Guo et al. developed a retrieval-based medical question answering system that efficiently retrieves answers using Elasticsearch and enhances them with semantic matching and knowledge graphs [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. The system's novel siamese-based answer selection architecture outperformed baseline models and systems in both Chinese and English datasets, demonstrating consistent improvements in quantification and qualification evaluations [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. Deep learning-based QA systems have emerged as a more flexible and adaptable approach, leveraging techniques such as powerful neural network architectures to automatically learn to understand and respond to questions [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. Yin et al. developed Evebot, a conversational system for detecting negative emotions and preventing depression through positive suggestions [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. It uses deep-learning models including a Bi-LSTM for emotion detection and an anti-language sequence-to-sequence neural network for counseling [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWhile these traditional QA systems have been effective for certain types of questions and domains, they have several limitations. One major limitation is the reliance of rule-based approaches on predefined rules or keywords, which makes them less flexible and adaptable to new or complex questions [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. These systems also struggle with understanding natural language queries and context, often leading to inaccurate or incomplete answers. Additionally, traditional QA systems are limited by the quality and coverage of their underlying knowledge base or document collection, which can affect the accuracy and relevance of their answers [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. For deep learning-based QA systems, one major limitation is their dependency on large amounts of labeled training data [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e], [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. These systems require vast datasets to learn patterns in language and develop accurate models, which can be challenging and resource-intensive to obtain, especially for specialized domains or languages [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Additionally, deep learning-based QA systems may struggle with out-of-domain or adversarial examples, where the input falls outside the scope of the training data, leading to errors or inaccurate responses [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e], [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAnother limitation of traditional QA systems is their inability to provide explanations or reasoning behind their answers [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. These systems typically return a single answer without any supporting context or evidence, making it challenging for users to understand how the answer was derived [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e], [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. This lack of transparency can reduce user trust and confidence in the system, especially in critical applications such as healthcare or legal domains [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. Overall, while traditional QA systems have been valuable in certain contexts, their limitations have led to the development of more advanced approaches\u003c/p\u003e \u003cp\u003eIn recent years, the advent of powerful language models, such as the Generative Pre-trained Transformer (GPT), has revolutionized the field of natural language processing (NLP) and opened up new possibilities for conversational agents [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e], [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e], [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. GPT, developed by OpenAI, is a state-of-the-art deep learning model capable of generating human-like text based on the input it receives [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e], [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e], [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. The latest iteration, GPT-4, is distinguished by its ability to learn from vast amounts of text data, supported by its billions of parameters, enabling it to capture complex patterns in language and generate highly coherent and informative text [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e], [\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e, p. 4], [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. However, a significant challenge with GPT models, including ChatGPT, is their tendency to produce hallucinations or responses that, while plausible, are factually incorrect [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e]. This issue has raised concerns about the reliability of these models, especially in critical applications such as healthcare [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]. To address this problem, researchers and developers are investigating the use of well-curated knowledge bases (KBs) to refine the models. By integrating authenticated and reliable information from KBs, the goal is to enhance the model's capability to generate pertinent and accurate responses, thereby decreasing the risk of hallucinations. This has led to the development of chatbots and question answering systems powered by GPT that can provide information and assistance across various domains [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn the context of healthcare, the potential of GPT-powered question answering systems and chatbots is particularly promising [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e]. Seenivasan et al. developed an end-to-end trainable Language-Vision GPT (LV-GPT) model to leverage GPT-based LLMs for Visual Question Answering (VQA) in robotic surgery [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e]. The LV-GPT model extends GPT2 to process vision input (images) by incorporating a vision tokenizer and vision token embedding [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e]. The model outperforms other state-of-the-art VQA models on public surgical-VQA datasets and a newly annotated dataset, demonstrating its effectiveness in capturing context from both language and vision modalities [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e]. Shi et al. developed a GPT-based Question Answering System for Fundus Fluorescein Angiography (FFA) with an image-text alignment module and a GPT-based interactive QA module [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. The system showed satisfactory performance in automatic evaluation and high accuracy and completeness in manual assessments, facilitating dynamic communication between ophthalmologists and patients for enhanced diagnostic processes [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. Although GPT-powered question answering systems and chatbots in healthcare hold significant promise, we found that these systems exhibit hallucination issues because they use pre-trained GPT models directly without fine-tuning [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e].In the case of HPV vaccination, where inadequate information and misconceptions are prevalent, leveraging fine-tuning techniques with advanced GPT models can significantly enhance the accuracy and reliability of information provided. A GPT-powered chatbot, when properly fine-tuned, could play a crucial role in educating the public and increasing awareness about the importance of vaccination.\u003c/p\u003e \u003cp\u003eIn this paper, we present the development and evaluation of a GPT-powered chatbot (VaxBot-HPV) designed to provide information and answer questions about the HPV vaccine. We also describe the design and implementation of the chatbot, its capabilities and limitations, as well as its potential impact on public health.\u003c/p\u003e \u003cp\u003eOverall, this paper highlights the potential of GPT-powered question answering systems and chatbots in healthcare, particularly in the context of HPV vaccination, and demonstrates how these systems can be leveraged to improve health literacy and promote vaccination uptake.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cp\u003eThe study is structured around three primary stages. Initially, we constructed a KB and collected question-answer pairs relevant to the HPV vaccine within the KB to develop the benchmark. Subsequently, we inferred answers for the questions in the test benchmark using both pretrained GPT models and GPT models fine-tuned on the benchmark. Finally, we assessed the results in terms of faithfulness and answer relevancy. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the overview of the study framework.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. KB and Gold Standard Construction\u003c/h2\u003e \u003cp\u003eTo build VaxBot-HPV, a chatbot designed to offer reliable information about the HPV vaccine, we first developed a KB deriving from peer-reviewed biomedical literature and web sources, resulting in a total of 451 documents on the HPV vaccine.\u003c/p\u003e \u003cp\u003eTo construct the question-answer pairs, we extracted 202 pairs of frequently asked questions and their answers related to the HPV vaccine from the collected webpages in the KB. The gold standard of question-answer sets was meticulously reviewed by domain experts to ensure their relevance and accuracy.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Models\u003c/h2\u003e \u003cp\u003eWe utilized two state-of-the-art LLMs, GPT-3.5 and GPT-4, developed by OpenAI, as the key components of this study.\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eGPT-3.5: GPT-3.5 is the iteration in OpenAI's series of large-scale language models, following the groundbreaking GPT-3. With an even larger model size (175\u0026nbsp;billion parameters) and enhanced capabilities, GPT-3.5 builds on the success of its predecessors in natural language processing (NLP) [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]. This advanced model exhibits impressive proficiency in understanding and generating human-like text, showcasing its potential for a wide range of applications including chatbots, content creation, and language translation [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e].\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eGPT-4: GPT-4, the advancement in OpenAI's renowned Generative Pre-trained Transformer series, marks a significant milestone in the field of natural language processing (NLP). With its remarkable increase to 170 trillion parameters, GPT-4 surpasses its predecessor, GPT-3, enabling it to tackle even more complex language tasks with improved accuracy and understanding [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]. This model represents a significant leap forward in NLP capabilities, holding the potential to revolutionize various fields, from conversational AI to content generation and beyond [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e].\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. Experiment Setup\u003c/h2\u003e \u003cp\u003eIn this study, VaxBot-HPV, fine-tuned on GPT-3.5, was developed using question-answer pairs from both a knowledge base and GPT-generated questions. The question-answer pairs derived from the knowledge base were divided into 162 samples for training purposes and 40 for testing purposes. To enhance question diversity and ensure the generalizability of our findings, we employed GPT-4 models to generate 80 question-answer pairs. After careful review, we included 39 questions based on their relevance to the HPV vaccine and manually updated their answers for inclusion into our study. Among the GPT-generated questions, 28 question-pairs were randomly selected for training and the rest for testing.\u003c/p\u003e \u003cp\u003eThe parameters of the VaxBot-HPV are outlined in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. We used the following prompt to instruct the GPT models in answering the query:\u003c/p\u003e \u003cp\u003e \u003cem\u003e\u0026ldquo; You are an expert Q\u0026amp;A system that is trusted around the world.\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eAlways answer the query using the provided context information, and not prior knowledge.\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eSome rules to follow\u003c/em\u003e:\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003e1. Never directly reference the given context in your answer.\u003c/h3\u003e\n\u003cp\u003e \u003cem\u003e2. Avoid statements like 'Based on the context, ...' or 'The context information ...' or anything along those lines.\u0026rdquo;\u003c/em\u003e \u003c/p\u003e \u003cp\u003eThe prompts for GPT-4 to generate questions were as follows:\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eUsing the provided context from referencing articles on HPV vaccine, formulate a question that captures an important fact from the context. Restrict the question to the context information provided. Please only output the question.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eVaxBot-HPV's development involved the performance comparison of various models, including GPT-3.5, as well as GPT-4, for each experimental set.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eParameters of VaxBot-HPV.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eParameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eValue\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003en_epochs\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ebatch_size\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elearning_rate_multiplier\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003etemperature\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003econtext_window\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2,048\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eToken limit\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4,096\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe experiments were carried out using a high-performance server containing 8 Nvidia A100 GPUs, each with a memory capacity of 80GB. This server configuration facilitated the effective training and evaluation of the models, ensuring the production of dependable and precise results.\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.4. Evaluation\u003c/h2\u003e \u003cp\u003eThe evaluation involved answer relevancy and faithfulness. Both are critical aspects in assessing the quality of generated responses. Answer relevancy gauges the extent to which the answers align with the questions, while faithfulness ensures factual accuracy, a fundamental requirement for reliable information retrieval. These metrics collectively provide a comprehensive evaluation of the model's performance in understanding and responding to user queries. Additionally, evaluations were conducted to thoroughly assess the system's effectiveness.\u003c/p\u003e \u003cp\u003eThe assessment of all outcomes was carried out using the Ragas metrics, which are GPT-supported measures widely adopted in NLP tasks to evaluate the quality of generated text [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. Specifically, the RAGAS metrics calculate answer relevancy and faithfulness through a detailed process. For answer relevancy, the ground truth answer and the generated answer are vectorized using the specified embedding model, and their cosine similarity is computed to determine alignment [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. For answer faithfulness, the process involves quantifying factual correctness by identifying true positives (facts present in both the ground truth and the generated answer), false positives (facts present in the generated answer but not in the ground truth), and false negatives (facts present in the ground truth but not in the generated answer) [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. The F1 score is then used to quantify correctness based on these values [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. A weighted average of factual correctness and semantic similarity provides the final score [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e illustrates the automatic performance evaluation of different GPT models in answer relevancy and faithfulness on the questions extracted from the KB. The results indicate that the VaxBot-HPV outperformed both the GPT-3.5 andGPT-4 models in terms of answer relevancy, achieving a score of 0.85 compared to 0.80 and 0.83, respectively. Similarly, the VaxBot-HPV exhibited higher faithfulness, scoring 0.97, compared to 0.92 for the GPT-3.5 model and 0.91 for the GPT-4 model. These results suggest that fine-tuning the GPT-3.5 model leads to improved performance in both answer relevancy and faithfulness compared to using the models in their pretrained states.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance evaluation of different GPT models in answer relevancy and faithfulness on the questions extracted from the knowledge base.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAnswer Relevancy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFaithfulness\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-3.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVaxBot-HPV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e presents the performance evaluation of different GPT models in terms of answer relevancy and faithfulness on questions generated by GPT-4. The GPT-3.5 model achieved an answer relevancy score of 0.80 and a faithfulness score of 0.90. In comparison, VaxBot-HPV showed improved performance with an answer relevancy score of 0.85 and a faithfulness score of 0.96. These results highlight the benefits of fine-tuning the GPT model, demonstrating its broader generalizability, applicability and robustness.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance evaluation of different GPT models in answer relevancy and faithfulness on the questions generated by GPT-4.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAnswer Relevancy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFaithfulness\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-3.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVaxBot-HPV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows two samples of questions and its generated questions by the four systems. We selected two questions. One question (\u0026ldquo;What are the risks of cervical cancer besides pregnancy at an early age?\u0026rdquo;) is generated by GPT, and another question (\u0026ldquo;What are the risks of the HPV vaccine?\u0026rdquo;) is from the test benchmark. VaxBot-HPV demonstrates an advantage in providing comprehensive and accurate responses to health-related inquiries compared to other systems. For instance, when asked about the risks of cervical cancer besides early pregnancy, VaxBot-HPV effectively listed multiple risk factors, including having multiple sexual partners, weakened immune systems, and specific health conditions. In contrast, the GPT-3.5 failed to identify any additional risk factors, while the GPT-4 provided information not directly to the question, such as \u0026ldquo;genital warts occurred most in adolescents and young adults\u0026rdquo;, which could be misleading. Additionally, ChatGPT, although comprehensive, was not succinct and failed to answer the question directly. Furthermore, when it comes to the question \u0026rdquo;What are the risks of the HPV vaccine?\u0026rdquo;, VaxBot-HPV effectively summarized over 12 years of safety monitoring, highlighted common and rare side effects, and provided actionable advice on preventing fainting-related injuries, all while maintaining a clear and concise format. In contrast, the GPT-3.5 and GPT-4, though accurate, lacked depth, information sources and reassurance, merely listing side effects without addressing common myths or providing detailed context. ChatGPT-4, despite its comprehensiveness, often failed to deliver succinct answers, resulting in verbose responses that lacked focus. These examples illustrate that VaxBot-HPV not only enhances the specificity and clarity of responses but also ensures that users receive accurate, reliable and actionable health information efficiently.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThe development and evaluation of VaxBot-HPV, a chatbot designed to provide information and answer questions about the HPV vaccine, demonstrates the potential of advanced language models, particularly GPT-3.5 and GPT-4, in healthcare applications. Compared to traditional QA systems, VaxBot-HPV leverages the capabilities of GPT models, especially after fine-tuning, to generate relevant and accurate responses to user queries.\u003c/p\u003e \u003cp\u003eVaxBot-HPV has a substantial advantage over existing pre-trained GPT models. This extensive pre-training allows VaxBot-HPV to have a deeper understanding of language and context, enabling it to provide more relevant and accurate answers to user queries. Unlike rule-based systems, which rely on predefined rules and patterns, and retrieval-based systems, which use keyword matching and ranking algorithms, VaxBot-HPV's pretrained model allows it to generate responses based on a broader understanding of the topic. This capability enhances the chatbot's ability to address a wide variety of questions and provide more informative and helpful responses to users. Moreover, VaxBot-HPV allows the answers to be dynamically generated, potentially offering more tailored responses to users compared to standard, one-size-fits-all answers. The fine-tuning process further enhances VaxBot-HPV's performance, particularly in the context of HPV vaccination, by adapting it to the specific domain. This adaptation improves answer relevancy and faithfulness, addressing common issues of ChatGPT such as hallucinations, where the model generates plausible but inaccurate responses. By fine-tuning on a dataset specific to HPV vaccination, VaxBot-HPV can learn the nuances of the topic, including relevant terminology, common misconceptions, and specific concerns that users may have. This specificity allows the chatbot to provide more accurate and tailored responses, increasing its overall effectiveness in addressing user queries related to the HPV vaccine. Furthermore, the fine-tuning process helps mitigate bias and misinformation that may be present in generic language models, ensuring that VaxBot-HPV provides reliable and trustworthy information to users seeking information about HPV vaccination. Additionally, the specific fine-tuning, which includes context in addition to question-answer pairs, enables VaxBot-HPV to extend beyond just answering questions. These models can also provide explanations or reasoning behind their answers, increasing transparency and user trust. This feature is particularly important in healthcare applications, where understanding the rationale behind medical advice is crucial for informed decision-making.\u003c/p\u003e \u003cp\u003eIn terms of evaluations, incorporating multiple sources, including questions generated by GPT models, strengthens the credibility and reliability of our findings regarding VaxBot-HPV's performance. By leveraging questions from diverse sources, we were able to assess the chatbot's ability to handle a wide range of queries beyond those explicitly included in the knowledge base. This comprehensive evaluation approach not only ensures the robustness of our results but also demonstrates VaxBot-HPV's versatility in addressing various user inquiries. Overall, the use of multiple evaluation metrics underscores the effectiveness and adaptability of VaxBot-HPV in providing reliable information and support to users.\u003c/p\u003e \u003cp\u003eWhile VaxBot-HPV demonstrates promising performance, there are several limitations to consider. First, the chatbot's effectiveness is contingent on the quality and comprehensiveness of the underlying knowledge base. Incomplete or inaccurate information in the KB could lead to erroneous or insufficient responses from the chatbot. Additionally, the chatbot's reliance on text-based interactions may limit its accessibility to individuals with visual or cognitive impairments who may benefit from alternative communication methods. Moreover, the inclusion of manual evaluations is needed to provide a holistic assessment of the chatbot's performance, enhancing the depth and accuracy of our conclusions. Furthermore, the evaluation of VaxBot-HPV was primarily based on its performance in answering questions, overlooking other aspects of user interaction such as ease of use, user satisfaction or engagement. Finally, the generalizability of our findings may be limited to the specific domain of HPV vaccination and may not extend to other healthcare contexts.\u003c/p\u003e \u003cp\u003eFuture research could focus on several areas to enhance the capabilities and impact of VaxBot-HPV. First, expanding the knowledge base to include a broader range of topics related to HPV vaccination and addressing emerging concerns or misconceptions could improve the chatbot's effectiveness and relevance. Second, integrating multimedia capabilities, such as image or video recognition, could enhance the chatbot's ability to provide information and support in a more interactive and engaging manner. Third, incorporating feedback mechanisms to gather user input and improve the chatbot's responses over time could enhance its usability and user satisfaction. Furthermore, exploring the integration of VaxBot-HPV with existing healthcare systems or platforms could facilitate its adoption and integration into clinical workflows, potentially improving access to information and promoting HPV vaccination uptake. Fourth, extending the chatbot to cover other types of vaccines and medical domains could broaden its applicability and utility, making it a more versatile tool for addressing public health concerns. Lastly, we need to add a user interface to VaxBot-HPV to make it more accessible and user-friendly, enhancing the overall user experience and encouraging more people to use the chatbot for reliable information on HPV vaccination and related topics.\u003c/p\u003e"},{"header":"5. Conclusions","content":"\u003cp\u003eIn conclusion, the development of VaxBot-HPV demonstrates the potential of GPT-powered chatbots in healthcare, particularly in promoting vaccination uptake and addressing common concerns and misconceptions. The study also underscores the importance of leveraging advanced language models and fine-tuning techniques in healthcare chatbot development. The efficacy of VaxBot-HPV highlights the transformative impact of such technologies on medical education, healthcare communication and information dissemination.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements:\u0026nbsp;\u003c/strong\u003eC.T. discloses support for the research of this work from National Institute of Allergy And Infectious Diseases of the National Institutes of Health [grant number R01AI130460 and U24AI171008], and CPRIT [grant number RP220244].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u003c/strong\u003e\u0026nbsp; \u0026nbsp;Methodology, Y.L., J.L., C.T.; software, Y.L., and J.L.; validation, L.T., and L.S.; formal analysis, M.L., Y.L., and J.L.; investigation, M.A., E.Y.; resources, C.T.; data curation, Y.L.; writing\u0026mdash;original draft preparation, Y.L.; writing\u0026mdash;review and editing, C.T.; visualization, Y.L.; supervision, C.T.; project administration, C.T.; funding acquisition, C.T., L.C. All authors have read and agreed to the published version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests:\u003c/strong\u003e The authors declare no conflicts of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability:\u003c/strong\u003e Dataset available on request from the authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eComputer Code:\u0026nbsp;\u003c/strong\u003eComputer code available on request from the authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInstitutional Review Board Statement:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInformed Consent Statement:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eM. das G. P. Leto, G. F. dos Santos J\u0026uacute;nior, A. M. Porro, and J. Tomimori, \u0026ldquo;Human papillomavirus infection: etiopathogenesis, molecular biology and clinical manifestations,\u0026rdquo; \u003cem\u003eAn. Bras. Dermatol.\u003c/em\u003e, vol. 86, pp. 306\u0026ndash;317, Apr. 2011, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1590/S0365-05962011000200014\u003c/span\u003e\u003cspan address=\"10.1590/S0365-05962011000200014\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. Brianti, E. De Flammineis, and S. R. Mercuri, \u0026ldquo;Review of HPV-related diseases and cancers,\u0026rdquo; \u003cem\u003eNew Microbiol\u003c/em\u003e, vol. 40, no. 2, pp. 80\u0026ndash;85, Apr. 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eF. S. Alhamlan, M. B. Alfageeh, M. A. Al Mushait, I. A. Al-Badawi, and M. N. Al-Ahdal, \u0026ldquo;Human Papillomavirus-Associated Cancers,\u0026rdquo; in \u003cem\u003eMicrobial Pathogenesis: Infection and Immunity\u003c/em\u003e, U. Kishore, Ed., Cham: Springer International Publishing, 2021, pp. 1\u0026ndash;14. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/978-3-030-67452-6_1\u003c/span\u003e\u003cspan address=\"10.1007/978-3-030-67452-6_1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eC. Chelimo, T. A. Wouldes, L. D. Cameron, and J. M. Elwood, \u0026ldquo;Risk factors for and prevention of human papillomaviruses (HPV), genital warts and cervical cancer,\u0026rdquo; \u003cem\u003eJournal of Infection\u003c/em\u003e, vol. 66, no. 3, pp. 207\u0026ndash;217, Mar. 2013, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jinf.2012.10.024\u003c/span\u003e\u003cspan address=\"10.1016/j.jinf.2012.10.024\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eR. Hull \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Cervical cancer in low and middle\u0026ndash;income countries (Review),\u0026rdquo; \u003cem\u003eOncology Letters\u003c/em\u003e, vol. 20, no. 3, pp. 2058\u0026ndash;2074, Sep. 2020, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3892/ol.2020.11754\u003c/span\u003e\u003cspan address=\"10.3892/ol.2020.11754\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK. S. Okunade, \u0026ldquo;Human papillomavirus and cervical cancer,\u0026rdquo; \u003cem\u003eJournal of Obstetrics and Gynaecology\u003c/em\u003e, Jul. 2020, Accessed: Mar. 27, 2024. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.tandfonline.com/doi/abs/\u003c/span\u003e\u003cspan address=\"https://www.tandfonline.com/doi/abs/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1080/01443615.2019.1634030\u003c/span\u003e\u003cspan address=\"10.1080/01443615.2019.1634030\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM. Arbyn \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Estimates of incidence and mortality of cervical cancer in 2018: a worldwide analysis,\u0026rdquo; \u003cem\u003eThe Lancet Global Health\u003c/em\u003e, vol. 8, no. 2, pp. e191\u0026ndash;e203, Feb. 2020, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/S2214-109X(19)30482-6\u003c/span\u003e\u003cspan address=\"10.1016/S2214-109X(19)30482-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM. Arbyn \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Worldwide burden of cervical cancer in 2008,\u0026rdquo; \u003cem\u003eAnnals of Oncology\u003c/em\u003e, vol. 22, no. 12, pp. 2675\u0026ndash;2686, Dec. 2011, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/annonc/mdr015\u003c/span\u003e\u003cspan address=\"10.1093/annonc/mdr015\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eE. Tesfaye \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Prevalence of human papillomavirus infection and associated factors among women attending cervical cancer screening in setting of Addis Ababa, Ethiopia,\u0026rdquo; \u003cem\u003eScientific Reports\u003c/em\u003e, vol. 14, no. 1, p. 4053, Feb. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-024-54754-x\u003c/span\u003e\u003cspan address=\"10.1038/s41598-024-54754-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. S. Ali, A. Y. Nirupama, S. Chaudhuri, and G. V. S. Murthy, \u0026ldquo;Therapeutic HPV Vaccination: A Strategy for Cervical Cancer Elimination in India,\u0026rdquo; \u003cem\u003eIndian J Gynecol Oncolog\u003c/em\u003e, vol. 22, no. 2, p. 38, Mar. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s40944-024-00800-5\u003c/span\u003e\u003cspan address=\"10.1007/s40944-024-00800-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eY. Li \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Unpacking adverse events and associations post COVID-19 vaccination: a deep dive into vaccine adverse event reporting system data,\u0026rdquo; \u003cem\u003eExpert Review of Vaccines\u003c/em\u003e, vol. 23, no. 1, pp. 53\u0026ndash;59, Dec. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1080/14760584.2023.2292203\u003c/span\u003e\u003cspan address=\"10.1080/14760584.2023.2292203\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Y, Li J, Dang Y, Chen Y, and Tao C, \u0026ldquo;Temporal and Spatial Analysis of COVID-19 Vaccines Using Reports from Vaccine Adverse Event Reporting System,\u0026rdquo; \u003cem\u003eJMIR Preprints\u003c/em\u003e, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/preprints.51007\u003c/span\u003e\u003cspan address=\"10.2196/preprints.51007\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eL. Iqbal, M. Jehan, and S. Azam, \u0026ldquo;Advancements in mRNA Vaccines: A Promising Approach for Combating Human Papillomavirus-Related Cancers,\u0026rdquo; \u003cem\u003eCancer Control\u003c/em\u003e, vol. 31, p. 10732748241238628, Jan. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1177/10732748241238629\u003c/span\u003e\u003cspan address=\"10.1177/10732748241238629\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eC. A. Gon\u0026ccedil;alves, G. Pereira-da-Silva, R. C. C. P. Silveira, P. C. M. Mayer, A. Zilly, and L. C. Lopes-J\u0026uacute;nior, \u0026ldquo;Safety, Efficacy, and Immunogenicity of Therapeutic Vaccines for Patients with High-Grade Cervical Intraepithelial Neoplasia (CIN 2/3) Associated with Human Papillomavirus: A Systematic Review,\u0026rdquo; \u003cem\u003eCancers\u003c/em\u003e, vol. 16, no. 3, Art. no. 3, Jan. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/cancers16030672\u003c/span\u003e\u003cspan address=\"10.3390/cancers16030672\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eE. M. Webster \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Building knowledge using a novel web-based intervention to promote HPV vaccination in a diverse, low-income population,\u0026rdquo; \u003cem\u003eGynecologic Oncology\u003c/em\u003e, vol. 181, pp. 102\u0026ndash;109, Feb. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ygyno.2023.12.005\u003c/span\u003e\u003cspan address=\"10.1016/j.ygyno.2023.12.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eD. R. Lowy and J. T. Schiller, \u0026ldquo;Reducing HPV-Associated Cancer Globally,\u0026rdquo; \u003cem\u003eCancer Prevention Research\u003c/em\u003e, vol. 5, no. 1, pp. 18\u0026ndash;23, Jan. 2012, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1158/1940-6207.CAPR-11-0542\u003c/span\u003e\u003cspan address=\"10.1158/1940-6207.CAPR-11-0542\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. G. Szilagyi \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Prevalence and characteristics of HPV vaccine hesitancy among parents of adolescents across the US,\u0026rdquo; \u003cem\u003eVaccine\u003c/em\u003e, vol. 38, no. 38, pp. 6027\u0026ndash;6037, Aug. 2020, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.vaccine.2020.06.074\u003c/span\u003e\u003cspan address=\"10.1016/j.vaccine.2020.06.074\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK. H. Nguyen \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Parental vaccine hesitancy and its association with adolescent HPV vaccination,\u0026rdquo; \u003cem\u003eVaccine\u003c/em\u003e, vol. 39, no. 17, pp. 2416\u0026ndash;2423, Apr. 2021, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.vaccine.2021.03.048\u003c/span\u003e\u003cspan address=\"10.1016/j.vaccine.2021.03.048\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eW. Jennings \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Lack of Trust, Conspiracy Beliefs, and Social Media Use Predict COVID-19 Vaccine Hesitancy,\u0026rdquo; \u003cem\u003eVaccines\u003c/em\u003e, vol. 9, no. 6, Art. no. 6, Jun. 2021, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/vaccines9060593\u003c/span\u003e\u003cspan address=\"10.3390/vaccines9060593\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eF. Gauna, P. Verger, L. Fressard, M. Jardin, J. K. Ward, and P. Peretti-Watel, \u0026ldquo;Vaccine hesitancy about the HPV vaccine among French young women and their parents: a telephone survey,\u0026rdquo; \u003cem\u003eBMC Public Health\u003c/em\u003e, vol. 23, no. 1, p. 628, Apr. 2023, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12889-023-15334-2\u003c/span\u003e\u003cspan address=\"10.1186/s12889-023-15334-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eG. Adeyanju, \u0026ldquo;Behavioral Insights into Vaccine Hesitancy Determinants in Sub-Saharan Africa,\u0026rdquo; Sep. 2022, Accessed: Mar. 28, 2024. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.db-thueringen.de/receive/dbt_mods_00053424\u003c/span\u003e\u003cspan address=\"https://www.db-thueringen.de/receive/dbt_mods_00053424\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eY. Chen and F. Zulkernine, \u0026ldquo;BIRD-QA: A BERT-based Information Retrieval Approach to Domain Specific Question Answering,\u0026rdquo; in \u003cem\u003e2021 IEEE International Conference on Big Data (Big Data)\u003c/em\u003e, Dec. 2021, pp. 3503\u0026ndash;3510. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/BigData52589.2021.9671523\u003c/span\u003e\u003cspan address=\"10.1109/BigData52589.2021.9671523\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eG. Vanitha, S. Sanampudi, and M. I.LAKSHMI, \u0026ldquo;APPROCHES FOR QUESTION ANSWERING SYSTEMS,\u0026rdquo; \u003cem\u003eInternational Journal of Engineering Science and Technology\u003c/em\u003e, vol. 3, Feb. 2011.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eI. Thalib, Widyawan, and I. Soesanti, \u0026ldquo;A Review on Question Analysis, Document Retrieval and Answer Extraction Method in Question Answering System,\u0026rdquo; in \u003cem\u003e2020 International Conference on Smart Technology and Applications (ICoSTA)\u003c/em\u003e, Feb. 2020, pp. 1\u0026ndash;5. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ICoSTA48221.2020.1570614175\u003c/span\u003e\u003cspan address=\"10.1109/ICoSTA48221.2020.1570614175\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eI. Tsampos and E. Marakakis, \u0026ldquo;A Medical Question Answering System with NLP and graph database\u0026rdquo;.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eB. L. Cairns \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;The MiPACQ Clinical Question Answering System,\u0026rdquo; \u003cem\u003eAMIA Annu Symp Proc\u003c/em\u003e, vol. 2011, pp. 171\u0026ndash;180, 2011.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eX. Feng, Q. Liu, C. Lao, and D. Sun, \u0026ldquo;Design and Implementation of Automatic Question Answering System in Information Retrieval,\u0026rdquo; in \u003cem\u003eProceedings of the 7th International Conference on Informatics, Environment, Energy and Applications\u003c/em\u003e, in IEEA \u0026rsquo;18. New York, NY, USA: Association for Computing Machinery, Mar. 2018, pp. 207\u0026ndash;211. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/3208854.3208862\u003c/span\u003e\u003cspan address=\"10.1145/3208854.3208862\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQ. Guo, S. Cao, and Z. Yi, \u0026ldquo;A medical question answering system using large language models and knowledge graphs,\u0026rdquo; \u003cem\u003eInternational Journal of Intelligent Systems\u003c/em\u003e, vol. 37, no. 11, pp. 8548\u0026ndash;8564, 2022, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1002/int.22955\u003c/span\u003e\u003cspan address=\"10.1002/int.22955\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eN. Saeed, humaira ashraf, and N. Jhanjhi, \u0026ldquo;DEEP LEARNING BASED QUESTION ANSWERING SYSTEM (SURVEY),\u0026rdquo; \u003cem\u003ePreprints\u003c/em\u003e, Dec. 2023, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.20944/preprints202312.1739.v1\u003c/span\u003e\u003cspan address=\"10.20944/preprints202312.1739.v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ. Yin, Z. Chen, K. Zhou, and C. Yu, \u0026ldquo;A Deep Learning Based Chatbot for Campus Psychological Therapy.\u0026rdquo; 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eF. Khennouche, Y. Elmir, Y. Himeur, N. Djebari, and A. Amira, \u0026ldquo;Revolutionizing generative pre-traineds: Insights and challenges in deploying ChatGPT and generative chatbots for FAQs,\u0026rdquo; \u003cem\u003eExpert Systems with Applications\u003c/em\u003e, vol. 246, p. 123224, Jul. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.eswa.2024.123224\u003c/span\u003e\u003cspan address=\"10.1016/j.eswa.2024.123224\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. Yin, N. Duan, B. Kao, J. Bao, and M. Zhou, \u0026ldquo;Answering Questions with Complex Semantic Constraints on Open Knowledge Bases,\u0026rdquo; in \u003cem\u003eProceedings of the 24th ACM International on Conference on Information and Knowledge Management\u003c/em\u003e, in CIKM \u0026rsquo;15. New York, NY, USA: Association for Computing Machinery, Oct. 2015, pp. 1301\u0026ndash;1310. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/2806416.2806542\u003c/span\u003e\u003cspan address=\"10.1145/2806416.2806542\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eA. Abdallah, B. Piryani, and A. Jatowt, \u0026ldquo;Exploring the state of the art in legal QA systems,\u0026rdquo; \u003cem\u003eJ Big Data\u003c/em\u003e, vol. 10, no. 1, p. 127, Aug. 2023, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s40537-023-00802-8\u003c/span\u003e\u003cspan address=\"10.1186/s40537-023-00802-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eY. Li \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Development of a Natural Language Processing Tool to Extract Acupuncture Point Location Terms,\u0026rdquo; in \u003cem\u003e2023 IEEE 11th International Conference on Healthcare Informatics (ICHI)\u003c/em\u003e, Jun. 2023, pp. 344\u0026ndash;351. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ICHI57859.2023.00053\u003c/span\u003e\u003cspan address=\"10.1109/ICHI57859.2023.00053\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eY. Li \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Artificial intelligence-powered pharmacovigilance: A review of machine and deep learning in clinical text-based adverse drug event detection for benchmark datasets,\u0026rdquo; \u003cem\u003eJournal of Biomedical Informatics\u003c/em\u003e, vol. 152, p. 104621, Apr. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jbi.2024.104621\u003c/span\u003e\u003cspan address=\"10.1016/j.jbi.2024.104621\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ. He \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Prompt Tuning in Biomedical Relation Extraction,\u0026rdquo; \u003cem\u003eJ Healthc Inform Res\u003c/em\u003e, Feb. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s41666-024-00162-9\u003c/span\u003e\u003cspan address=\"10.1007/s41666-024-00162-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eE. Stroh and P. Mathur, \u0026ldquo;Question Answering Using Deep Learning\u0026rdquo;.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ. Li \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Mapping Vaccine Names in Clinical Trials to Vaccine Ontology using Cascaded Fine-Tuned Domain-Specific Language Models,\u0026rdquo; \u003cem\u003eRes Sq\u003c/em\u003e, p. rs.3.rs-3362256, Sep. 2023, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.21203/rs.3.rs-3362256/v1\u003c/span\u003e\u003cspan address=\"10.21203/rs.3.rs-3362256/v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. Lu \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering\u0026rdquo;.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ. Lin \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;What Makes a Good Answer? The Role of Context in Question Answering\u0026rdquo;.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. Min, V. Zhong, R. Socher, and C. Xiong, \u0026ldquo;Efficient and Robust Question Answering from Minimal Context over Documents.\u0026rdquo; 2018.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eN. Goyal \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;What Else Do I Need to Know? The Effect of Background Information on Users\u0026rsquo; Reliance on QA Systems.\u0026rdquo; 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eY. Li, J. Li, J. He, and C. Tao, \u0026ldquo;AE-GPT: Using Large Language Models to extract adverse events from surveillance reports-A use case with influenza vaccine adverse events,\u0026rdquo; \u003cem\u003ePLOS ONE\u003c/em\u003e, vol. 19, no. 3, p. e0300919, Mar. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0300919\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0300919\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eY. Hu \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Zero-shot Clinical Entity Recognition using ChatGPT,\u0026rdquo; \u003cem\u003earXiv.org\u003c/em\u003e, 2023, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arXiv.2303.16416\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2303.16416\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eY. Li \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Relation Extraction Using Large Language Models: A Case Study on Acupuncture Point Locations,\u0026rdquo; \u003cem\u003earXiv.org\u003c/em\u003e, 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arXiv.2404.05415\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2404.05415\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eE. Chang, \u003cem\u003eExamining GPT-4\u0026rsquo;s Capabilities and Enhancement with SocraSynth\u003c/em\u003e. 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK. S. Kalyan, \u0026ldquo;A survey of GPT-3 family large language models including ChatGPT and GPT-4,\u0026rdquo; \u003cem\u003eNatural Language Processing Journal\u003c/em\u003e, vol. 6, p. 100048, Mar. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.nlp.2023.100048\u003c/span\u003e\u003cspan address=\"10.1016/j.nlp.2023.100048\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eT. M. Al-Hasan, A. N. Sayed, F. Bensaali, Y. Himeur, I. Varlamis, and G. Dimitrakopoulos, \u0026ldquo;From Traditional Recommender Systems to GPT-Based Chatbots: A Survey of Recent Developments and Future Directions,\u0026rdquo; \u003cem\u003eBig Data and Cognitive Computing\u003c/em\u003e, vol. 8, no. 4, Art. no. 4, Apr. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/bdcc8040036\u003c/span\u003e\u003cspan address=\"10.3390/bdcc8040036\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ. Li, X. Cheng, X. Zhao, J.-Y. Nie, and J.-R. Wen, \u0026ldquo;HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models,\u0026rdquo; in \u003cem\u003eProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing\u003c/em\u003e, H. Bouamor, J. Pino, and K. Bali, Eds., Singapore: Association for Computational Linguistics, Dec. 2023, pp. 6449\u0026ndash;6464. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.18653/v1/2023.emnlp-main.397\u003c/span\u003e\u003cspan address=\"10.18653/v1/2023.emnlp-main.397\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eT. R. McIntosh, T. Liu, T. Susnjak, P. Watters, A. Ng, and M. N. Halgamuge, \u0026ldquo;A Culturally Sensitive Test to Evaluate Nuanced GPT Hallucination,\u0026rdquo; \u003cem\u003eIEEE Transactions on Artificial Intelligence\u003c/em\u003e, pp. 1\u0026ndash;13, 2023, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/TAI.2023.3332837\u003c/span\u003e\u003cspan address=\"10.1109/TAI.2023.3332837\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. Garc\u0026iacute;a-M\u0026eacute;ndez and F. de Arriba-P\u0026eacute;rez, \u0026ldquo;Large Language Models and Healthcare Alliance: Potential and Challenges of Two Representative Use Cases,\u0026rdquo; \u003cem\u003eAnn Biomed Eng\u003c/em\u003e, Feb. 2024, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s10439-024-03454-8\u003c/span\u003e\u003cspan address=\"10.1007/s10439-024-03454-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eL. Seenivasan, M. Islam, G. Kannan, and H. Ren, \u0026ldquo;SurgicalGPT: End-to-End Language-Vision GPT for Visual Question Answering in Surgery,\u0026rdquo; in \u003cem\u003eMedical Image Computing and Computer Assisted Intervention \u0026ndash; MICCAI 2023\u003c/em\u003e, H. Greenspan, A. Madabhushi, P. Mousavi, S. Salcudean, J. Duncan, T. Syeda-Mahmood, and R. Taylor, Eds., Cham: Springer Nature Switzerland, 2023, pp. 281\u0026ndash;290. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/978-3-031-43996-4_27\u003c/span\u003e\u003cspan address=\"10.1007/978-3-031-43996-4_27\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eD. Shi \u003cem\u003eet al.\u003c/em\u003e, \u003cem\u003eFFA-GPT: an Interactive Visual Question Answering System for Fundus Fluorescein Angiography\u003c/em\u003e. 2023. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.21203/rs.3.rs-3307492/v1\u003c/span\u003e\u003cspan address=\"10.21203/rs.3.rs-3307492/v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eA. Koubaa, \u0026ldquo;GPT-4 vs. GPT-3.5: A Concise Showdown,\u0026rdquo; \u003cem\u003ePreprints\u003c/em\u003e, Mar. 2023, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.20944/preprints202303.0422.v1\u003c/span\u003e\u003cspan address=\"10.20944/preprints202303.0422.v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK. Nayanam and V. Sharma, \u0026ldquo;TOWARDS ARCHITECTING RESEARCH PERSPECTIVE FUTURE SCOPE WITH CHAT GPT,\u0026rdquo; Jul. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\u0026ldquo;Metrics | Ragas.\u0026rdquo; Accessed: Mar. 08, 2024. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://docs.ragas.io/en/latest/concepts/metrics/index.html\u003c/span\u003e\u003cspan address=\"https://docs.ragas.io/en/latest/concepts/metrics/index.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"npj-biomedical-innovations","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [npj Biomedical Innovations](https://www.nature.com/npjbiomedinnov/)","snPcode":"44385","submissionUrl":"https://submission.springernature.com/new-submission/44385/3","title":"npj Biomedical Innovations","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Vaccine, HPV vaccine, Cervical Cancer, GPT, Large Language model, QA system, Chatbot, Medical education","lastPublishedDoi":"10.21203/rs.3.rs-4876692/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4876692/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e: HPV vaccine is an effective measure to prevent and control the diseases caused by Human Papillomavirus (HPV). This study addresses the development of VaxBot-HPV, a chatbot aimed at improving health literacy and promoting vaccination uptake by providing information and answering questions about the HPV vaccine;\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e: We constructed the knowledge base (KB) for VaxBot-HPV, which consists of 451 documents from biomedical literature and web sources on the HPV vaccine. We extracted 202 question-answer pairs from the KB and 39 questions generated by GPT-4 for training and testing purposes. To comprehensively understand the capabilities and potential of GPT-based chatbots, three models were involved in this study : GPT-3.5, VaxBot-HPV, and GPT-4. The evaluation criteria included answer relevancy and faithfulness;\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e: VaxBot-HPV demonstrated superior performance in answer relevancy and faithfulness compared to baselines (Answer relevancy: 0.85; Faithfulness: 0.97) for the test questions in KB, (Answer relevancy: 0.85; Faithfulness: 0.96) for GPT generated questions;\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e: This study underscores the importance of leveraging advanced language models and fine-tuning techniques in the development of chatbots for healthcare applications, with implications for improving medical education and public health communication.\u003c/p\u003e","manuscriptTitle":"VaxBot-HPV: A GPT-based Chatbot for Answering HPV Vaccine-related Questions","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-09-11 16:30:47","doi":"10.21203/rs.3.rs-4876692/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorAssigned","content":"","date":"2024-08-14T13:18:02+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-08-14T09:48:49+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Biomedical Innovations","date":"2024-08-07T19:07:15+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"npj-biomedical-innovations","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [npj Biomedical Innovations](https://www.nature.com/npjbiomedinnov/)","snPcode":"44385","submissionUrl":"https://submission.springernature.com/new-submission/44385/3","title":"npj Biomedical Innovations","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4dd56c78-c656-4c41-a3c0-c5700b0cb2a2","owner":[],"postedDate":"September 11th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":36049300,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":36049301,"name":"Biological sciences/Computational biology and bioinformatics/Machine learning"}],"tags":[],"updatedAt":"2024-09-11T16:30:47+00:00","versionOfRecord":[],"versionCreatedAt":"2024-09-11 16:30:47","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4876692","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4876692","identity":"rs-4876692","version":["v1"]},"buildId":"zQwnuV7TCBrMSSSToR1PI","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: preprint-html ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-08-12T06:43:03.944938+00:00
License: CC-BY-4.0