DoctorBOT: An AI-powered Chatbot for Evidence-Based Obstetric Hemorrhage Management in Low-Resource Settings | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Short Report DoctorBOT: An AI-powered Chatbot for Evidence-Based Obstetric Hemorrhage Management in Low-Resource Settings Sandra Jaramillo-Rincón, Ambre Mychalski, Juan Yepes-Nuñez, Ruben Manrique This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6580160/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 12 You are reading this latest preprint version Abstract Background: Postpartum hemorrhage (PPH) remains a leading cause of maternal mortality, especially in low-resource settings. Despite the existence of high-quality clinical practice guidelines (CPGs), their implementation in rural areas is limited due to lack of training and access. Artificial intelligence (AI)-based tools may bridge this gap by supporting real-time clinical decision-making. Objective: This study presents the development and evaluation of DoctorBOT, an AI-powered chatbot using Retrieval-Augmented Generation (RAG) to deliver evidence-based recommendations for PPH management. Methods: DoctorBOT was trained using vectorized high-quality obstetric CPGs. Clinical questions were collected from 26 rural healthcare workers and evaluated by large language models (LLMs) ChatGPT-3.5 and GPT-4. Performance was assessed through expert evaluation (Likert scale across five domains) and automated semantic comparison using BERTScore. Results: GPT-3.5 outperformed GPT-4 in human evaluations for accuracy and clinical utility. Conversely, GPT-4 achieved a higher BERTScore (0.83 vs. 0.79). LLaMA2 achieved the highest score (0.92) after additional fine-tuning. In a user survey, 85% of rural healthcare providers rated DoctorBOT as “very satisfactory.” Conclusion: DoctorBOT demonstrates strong potential to support evidence-based clinical decision-making in obstetric emergencies. Its integration could reduce maternal mortality and improve health equity in underserved areas. Postpartum Hemorrhage Maternal Mortality Artificial Intelligence Clinical Decision Support Chatbot Low-Resource Settings Retrieval-Augmented Generation Large Language Models Evidence-Based Medicine Figures Figure 1 Figure 2 Introduction Maternal mortality is a significant global health issue, especially in rural areas with limited access to care. About 94% of maternal deaths occur in low- and lower-middle-income countries, with rural regions most affected. ( 1 , 2 ) Obstetric hemorrhage, a leading cause of these deaths, requires prompt clinical decision-making to improve outcomes. Innovative solutions are needed to meet Sustainable Development Goal (SDG) 3.1, which aims to reduce the global maternal mortality ratio to under 70 per-100,000 live births by 2030. (( 3 ) AI-driven healthcare systems, including Clinical Decision Support Systems (CDSS) and chatbots, are currently being used to provide information to expectant mothers and support prenatal care and labor management. (( 4 ) Methods DoctorBOT is an AI chatbot developed by a multidisciplinary team from the Engineering and Medicine faculties at Universidad de Los Andes, Colombia. The team focused on accuracy and adaptability to help manage obstetric emergencies. Utilizing a Retrieval-Augmented Generation (RAG) framework (( 5 ), DoctorBOT dynamically accesses authoritative clinical guideline databases to provide evidence-based recommendations tailored to specific scenarios. To improve reliability, the prompts and retrieval engine were optimized to minimize inaccuracies and customize responses to the complexities of obstetric emergencies. (Fig. 1 ) Twenty-six healthcare professionals, working in rural settings without specialized training, provided common clinical questions encountered during obstetric emergencies. These questions were then evaluated using two large language models (LLMs): ChatGPT 3.5 and ChatGPT 4. A dual evaluation approach assessed the quality of their responses. Human Evaluation Human Evaluation A panel of seven experts (four obstetricians and three anesthesiologists) with extensive experience managing obstetric emergencies evaluated the responses from each LLM. Responses were assessed across five key domains using a 10-point Likert scale. Experts also provided qualitative feedback on each response. (Supplement material 1: Evaluation of the Language Model Prototype (LLM-1.3) for Postpartum Hemorrhage Care) Automated Evaluation BERTScore,( 6 ) a metric that evaluates the semantic similarity between two texts, was used to assess the alignment of the chatbot's responses with established clinical guidelines. This provided a quantitative assessment across key domains: accuracy, relevance, comprehensibility, coherence, and clinical utility. Results Human Evaluation Human evaluations show GPT-3.5 turbo outperformed GPT-4, particularly in expert assessments of accuracy and clinical utility, with scores of 8.44 and 9.03, respectively, compared to GPT-4's scores of 6.5 for accuracy and 7.0 for clinical utility (Fig. 2 .A). In a preliminary user survey involving 22 healthcare professionals from rural maternal care settings, 85% rated DoctorBOT's performance as "very satisfactory." Automated Evaluation GPT-4 achieved a higher average BERTScore (0.83) than GPT-3.5 (0.79). LLaMA2 also showed notable improvements with additional training; after 12 epochs, LLaMA2 achieved a BERTScore of 0.92 (Fig. 2 .B). Discussion DoctorBOT is an AI chatbot that offers real-time clinical guidance for obstetric emergencies in resource-limited settings. Unlike traditional chatbots, it uses Retrieval-Augmented Generation (RAG) to deliver accurate and context-specific responses. Our evaluation identified a trade-off between accuracy and speed among the tested LLMs. GPT-4 showed superior accuracy and a 90% reduction in hallucinations, while GPT-3.5 Turbo was more effective for real-world obstetric emergencies due to its faster response time. This highlights the need to balance accuracy and speed in urgent medical decisions. Healthcare professionals often require quick, context-dependent answers, making GPT-3.5 Turbo more practical in those situations. However, for complex, evidence-based clinical decisions, GPT-4's accuracy is crucial. This suggests that a hybrid approach may be ideal, using the speed of GPT-3.5 turbo for quick queries and the accuracy of GPT-4 for complex decisions. The 85% satisfaction score from healthcare professionals highlights DoctorBOT's potential as a valuable decision-making tool. The next research phase will focus on validating DoctorBOT's effectiveness in clinical settings through: Clinical Validation: A pilot study to evaluate its impact on decision-making, provider confidence, and patient outcomes. Usability Testing: Collecting feedback on the user interface, ease of use, and workflow integration. Ethical Considerations: Addressing ethical concerns around AI use in healthcare, including data privacy, bias, and transparency. Conclusion DoctorBOT prototype model represents a promising advancement toward achieving SDG 3.1. With continued development and successful implementation, this AI-powered tool has the potential to significantly enhance obstetric emergency care, saving lives and contributing to improved maternal health worldwide, mainly by reducing maternal mortality rates in resource-limited settings. By advancing these areas, DoctorBOT aims to significantly enhance obstetric emergency care, save lives, and improve maternal health globally. Declarations Conflict of Interest Disclosures: The authors report no conflicts of interest. Funding/Support: This research received no external funding. Role of the Funder/Sponsor: Not applicable. Prior Presentations: An abstract of this work was accepted for a short oral presentation titled “Enhancing Obstetric Emergency Care with AI-Integrated Clinical Guidelines” at the Global Evidence Summit 2024 . Related Manuscripts: This manuscript has not been published, posted, or submitted elsewhere. Ethics approval and consent to participate The Institutional Research Ethics Committee of the Faculty of Medicine at Universidad de los Andes reviewed and approved this study. The first approval (Act No. 1909, May 31, 2024) covered the development of the AI prototype. It determined that no informed consent was required as no personal data or interactions with human subjects were involved. The second approval (Act No. 2025012801 SJ, January 28, 2025) covered the validation process with healthcare professionals. For this phase, informed consent was obtained from all participants before their inclusion in the study. Both approvals were granted by national ethical regulations (Resolution 008430 of 1993).” Acknowledgments We sincerely thank the team of obstetricians and anesthesiologists — Dr. Mauricio Vasco, Dr. Camilo Fonseca, Dr. Amparo Ramírez, and Dr. Luis Carlos Franco — for their generous support in validating the prototype of DoctorBOT. Their clinical expertise was instrumental in assessing the relevance, accuracy, and utility of the chatbot's responses. We also extend our deep gratitude to the rural physicians, nurses, and traditional birth attendants who contributed to this project by sharing real-life clinical questions and scenarios from their daily practice. Their insights, shaped by firsthand experience managing obstetric emergencies in low-resource settings, were essential to building a model grounded in the realities of frontline care. This work reflects not only technological innovation, but also the lived experiences and professional dedication of those who strive to save lives under challenging conditions. Data Sharing Statement The data, code, and supporting materials for this study are openly available on GitHub at: https://github.com/AmbreMychalski/Tesis-1. The repository includes: Structured prompt examples The chatbot’s retrieval pipeline (RAG framework) Evaluation scripts for BERTScore analysis Supporting documentation for replication A data dictionary defining key fields in the prototype evaluation data These materials are available immediately and without restriction for non-commercial use. Researchers may use the code and data for further analysis, adaptation, or model benchmarking. No individual patient data were collected or used in this study. No additional restrictions apply. AI Disclosure We declare that we used artificial intelligence-based tools to support the translation and language editing of the manuscript, ensuring clarity and consistency throughout the text. Ethical Statement The Research Ethics Committee at Universidad de los Andes approved the protocol. In accordance with current Colombian regulations — Resolution 008430 of 1993 and Resolution 2378 of 2008 — this study is classified as minimal risk . References Goldenberg RL, McClure EM, Saleem S. Improving pregnancy outcomes in low- and middle-income countries. Reproductive Health. Volume 15. BioMed Central Ltd.; 2018. WHO UUWBG and UD. Trends in maternal mortality 2000 to 2020. Inform [Internet]. 2023 [cited 2025 Mar 10]; Available from: https://www.who.int/publications/i/item/9789240068759 United Nations. https://sdgs.un.org/goals/goal3#targets_and_indicators . THE 17 GOALS. Bajwa J, Munir U, Nori A, Williams B. Artificial intelligence in healthcare: transforming the practice of medicine. Future Healthc J. 2021;8(2):e188–94. Siriwardhana S, Weerasekera R, Wen E, Kaluarachchi T, Rajib R†, Nanayakkara S. Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering. Available from: https://doi.org/10.1162/tacl Zhang T, Kishore V, Wu F, Weinberger KQ, Artzi Y, BERTScore. Evaluating Text Generation with BERT. 2019; Available from: http://arxiv.org/abs/1904.09675 Additional Declarations No competing interests reported. Supplementary Files EvaluationPostpartumHemorrhageShortEnglish.pdf Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 11 Mar, 2026 Reviews received at journal 19 Feb, 2026 Reviewers agreed at journal 14 Feb, 2026 Reviews received at journal 12 Sep, 2025 Reviewers agreed at journal 16 Aug, 2025 Reviewers agreed at journal 06 Aug, 2025 Reviewers agreed at journal 07 Jun, 2025 Reviewers invited by journal 01 Jun, 2025 Editor assigned by journal 01 Jun, 2025 Editor invited by journal 23 May, 2025 Submission checks completed at journal 22 May, 2025 First submitted to journal 22 May, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6580160","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Short Report","associatedPublications":[],"authors":[{"id":465118014,"identity":"a4e70210-696f-4911-bc3e-fef0c1f3b036","order_by":0,"name":"Sandra Jaramillo-Rincón","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9klEQVRIiWNgGAWjYHCCBBDBw8fOwPiANC1szAzMBqTZBdTCJkGUSvkGhoefK2q2ybAxMx+r/FFjl8/PwHzs4xcGOzlcWgwOMCRLnjl2G+gwtrTbPMeSLWc2sCXPlmFINsapBegXyQY2kBYes9uMDcwGBgd4jJklGA4kNuB2WPLPhn8QLYU/G+oN7AlpYTjAkCbZ2AbRwsDbcNjAgIHHmPEDHi0GhxnSLBv7wH5JluY5dtxA4jBbMjODAW6/yLf3JN9s+Hbbnp+9+eDHHzXVBvztzYcZf1TgDjEGZp4EdBGQIN5YZT+AKcb4A5+OUTAKRsEoGGkAAOMmRuF724CFAAAAAElFTkSuQmCC","orcid":"","institution":"Universidad de los Andes","correspondingAuthor":true,"prefix":"","firstName":"Sandra","middleName":"","lastName":"Jaramillo-Rincón","suffix":""},{"id":465118015,"identity":"8a5205d2-5630-4ba8-b743-07948fb68af9","order_by":1,"name":"Ambre Mychalski","email":"","orcid":"","institution":"Universidad de los Andes","correspondingAuthor":false,"prefix":"","firstName":"Ambre","middleName":"","lastName":"Mychalski","suffix":""},{"id":465118016,"identity":"30cb8779-bcb9-43fc-8ac9-cbf980bfb9c4","order_by":2,"name":"Juan Yepes-Nuñez","email":"","orcid":"","institution":"Universidad de los Andes","correspondingAuthor":false,"prefix":"","firstName":"Juan","middleName":"","lastName":"Yepes-Nuñez","suffix":""},{"id":465118018,"identity":"f687dcfd-d0a9-45e3-a287-93e9da972c6d","order_by":3,"name":"Ruben Manrique","email":"","orcid":"","institution":"Universidad de los Andes","correspondingAuthor":false,"prefix":"","firstName":"Ruben","middleName":"","lastName":"Manrique","suffix":""}],"badges":[],"createdAt":"2025-05-02 17:38:15","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6580160/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6580160/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":83906558,"identity":"1d96b55f-108a-46f4-85d1-4083981ae6cc","added_by":"auto","created_at":"2025-06-04 10:25:24","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":29552,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eChatbot Architecture\u003c/strong\u003e\u003cbr\u003e\nOverview of the Retrieval-Augmented Generation (RAG) process for addressing physician queries using obstetric guidelines.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e1.\u003c/em\u003e \u003cem\u003e\u003cstrong\u003eVectorial Database Construction:\u003c/strong\u003e\u003c/em\u003e\u003cem\u003e The database used for the RAG process contains vectorized forms of obstetric guidelines sourced from various global health organizations.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e2.\u003c/em\u003e \u003cem\u003e\u003cstrong\u003eHyDE Methodology:\u003c/strong\u003e\u003c/em\u003e\u003cem\u003e The physician’s question is processed through the HyDE (Hypothetical Document Embeddings) framework, generating an initial naive answer using conversation history and the query itself. This step uses GPT-3.5 and GPT-4 models.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e3.\u003c/em\u003e \u003cem\u003e\u003cstrong\u003eRAG Framework:\u003c/strong\u003e\u003c/em\u003e\u003cem\u003e The naive answer is vectorized, and cosine similarity is used to retrieve relevant fragments from the guideline database.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e4.\u003c/em\u003e \u003cem\u003e\u003cstrong\u003eAnswer Generation:\u003c/strong\u003e\u003c/em\u003e\u003cem\u003e The original question and the retrieved content are provided to the model to generate the final evidence-based answer.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-6580160/v1/24582fa3455c777266ac36a3.png"},{"id":83906555,"identity":"af289aec-0e5f-4b00-964e-e245558d08d6","added_by":"auto","created_at":"2025-06-04 10:25:23","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":47884,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePerformance Evaluation of LLMs for DoctorBOT\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA. \u003cstrong\u003eHuman Evaluation of GPT Models:\u003c/strong\u003e Performance comparison between GPT-3.5 and GPT-4 across five evaluation domains as assessed by clinical experts.\u003cbr\u003e\nB. \u003cstrong\u003eAutomated Evaluation Using BERTScore:\u003c/strong\u003e Semantic similarity scores for three fine-tuned LLaMA2 models, demonstrating consistent performance across training epochs.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-6580160/v1/3cf54234a9ef79ec8a04558e.png"},{"id":83906786,"identity":"f9d2a1b5-8f81-4ecd-a8f4-bed08e05eb17","added_by":"auto","created_at":"2025-06-04 10:33:28","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":765894,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6580160/v1/92ed4ef1-6daa-4688-87bf-0bdabd892aa3.pdf"},{"id":83906559,"identity":"3ab195bc-f7bd-4a4c-a119-9e0495769f03","added_by":"auto","created_at":"2025-06-04 10:25:24","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":230409,"visible":true,"origin":"","legend":"","description":"","filename":"EvaluationPostpartumHemorrhageShortEnglish.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6580160/v1/88b65499be278ccf16c499e9.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"DoctorBOT: An AI-powered Chatbot for Evidence-Based Obstetric Hemorrhage Management in Low-Resource Settings","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMaternal mortality is a significant global health issue, especially in rural areas with limited access to care. About 94% of maternal deaths occur in low- and lower-middle-income countries, with rural regions most affected. (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) Obstetric hemorrhage, a leading cause of these deaths, requires prompt clinical decision-making to improve outcomes. Innovative solutions are needed to meet Sustainable Development Goal (SDG) 3.1, which aims to reduce the global maternal mortality ratio to under 70 per-100,000 live births by 2030. ((\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eAI-driven healthcare systems, including Clinical Decision Support Systems (CDSS) and chatbots, are currently being used to provide information to expectant mothers and support prenatal care and labor management. ((\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e)\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eDoctorBOT is an AI chatbot developed by a multidisciplinary team from the Engineering and Medicine faculties at Universidad de Los Andes, Colombia. The team focused on accuracy and adaptability to help manage obstetric emergencies.\u003c/p\u003e\n\u003cp\u003eUtilizing a Retrieval-Augmented Generation (RAG) framework ((\u003cspan class=\"CitationRef\"\u003e5\u003c/span\u003e), DoctorBOT dynamically accesses authoritative clinical guideline databases to provide evidence-based recommendations tailored to specific scenarios. To improve reliability, the prompts and retrieval engine were optimized to minimize inaccuracies and customize responses to the complexities of obstetric emergencies. (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e)\u003c/p\u003e\n\u003cp\u003eTwenty-six healthcare professionals, working in rural settings without specialized training, provided common clinical questions encountered during obstetric emergencies. These questions were then evaluated using two large language models (LLMs): ChatGPT 3.5 and ChatGPT 4. A dual evaluation approach assessed the quality of their responses. Human Evaluation\u003c/p\u003e\n\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003eHuman Evaluation\u003c/h2\u003e\n\u003cp\u003eA panel of seven experts (four obstetricians and three anesthesiologists) with extensive experience managing obstetric emergencies evaluated the responses from each LLM. Responses were assessed across five key domains using a 10-point Likert scale. Experts also provided qualitative feedback on each response. (Supplement material 1: Evaluation of the Language Model Prototype (LLM-1.3) for Postpartum Hemorrhage Care)\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003eAutomated Evaluation\u003c/h3\u003e\n\u003cp\u003eBERTScore,(\u003cspan class=\"CitationRef\"\u003e6\u003c/span\u003e) a metric that evaluates the semantic similarity between two texts, was used to assess the alignment of the chatbot's responses with established clinical guidelines. This provided a quantitative assessment across key domains: accuracy, relevance, comprehensibility, coherence, and clinical utility.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n\u003ch2\u003eHuman Evaluation\u003c/h2\u003e\n\u003cp\u003eHuman evaluations show GPT-3.5 turbo outperformed GPT-4, particularly in expert assessments of accuracy and clinical utility, with scores of 8.44 and 9.03, respectively, compared to GPT-4's scores of 6.5 for accuracy and 7.0 for clinical utility (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e.A).\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn a preliminary user survey involving 22 healthcare professionals from rural maternal care settings, 85% rated DoctorBOT's performance as \"very satisfactory.\"\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003eAutomated Evaluation\u003c/h3\u003e\n\u003cp\u003eGPT-4 achieved a higher average BERTScore (0.83) than GPT-3.5 (0.79). LLaMA2 also showed notable improvements with additional training; after 12 epochs, LLaMA2 achieved a BERTScore of 0.92 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e.B).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eDoctorBOT is an AI chatbot that offers real-time clinical guidance for obstetric emergencies in resource-limited settings. Unlike traditional chatbots, it uses Retrieval-Augmented Generation (RAG) to deliver accurate and context-specific responses.\u003c/p\u003e \u003cp\u003eOur evaluation identified a trade-off between accuracy and speed among the tested LLMs. GPT-4 showed superior accuracy and a 90% reduction in hallucinations, while GPT-3.5 Turbo was more effective for real-world obstetric emergencies due to its faster response time. This highlights the need to balance accuracy and speed in urgent medical decisions.\u003c/p\u003e \u003cp\u003eHealthcare professionals often require quick, context-dependent answers, making GPT-3.5 Turbo more practical in those situations. However, for complex, evidence-based clinical decisions, GPT-4's accuracy is crucial.\u003c/p\u003e \u003cp\u003eThis suggests that a hybrid approach may be ideal, using the speed of GPT-3.5 turbo for quick queries and the accuracy of GPT-4 for complex decisions. The 85% satisfaction score from healthcare professionals highlights DoctorBOT's potential as a valuable decision-making tool.\u003c/p\u003e \u003cp\u003eThe next research phase will focus on validating DoctorBOT's effectiveness in clinical settings through:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eClinical Validation: A pilot study to evaluate its impact on decision-making, provider confidence, and patient outcomes.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eUsability Testing: Collecting feedback on the user interface, ease of use, and workflow integration.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eEthical Considerations: Addressing ethical concerns around AI use in healthcare, including data privacy, bias, and transparency.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eDoctorBOT prototype model represents a promising advancement toward achieving SDG 3.1. With continued development and successful implementation, this AI-powered tool has the potential to significantly enhance obstetric emergency care, saving lives and contributing to improved maternal health worldwide, mainly by reducing maternal mortality rates in resource-limited settings. By advancing these areas, DoctorBOT aims to significantly enhance obstetric emergency care, save lives, and improve maternal health globally.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eConflict of Interest Disclosures:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors report no conflicts of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding/Support:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no external funding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRole of the Funder/Sponsor:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePrior Presentations:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAn abstract of this work was accepted for a short oral presentation titled \u003cem\u003e\u0026ldquo;Enhancing Obstetric Emergency Care with AI-Integrated Clinical Guidelines\u0026rdquo;\u003c/em\u003e at the \u003cstrong\u003eGlobal Evidence Summit 2024\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRelated Manuscripts:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis manuscript has not been published, posted, or submitted elsewhere.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Institutional Research Ethics Committee of the Faculty of Medicine at Universidad de los Andes reviewed and approved this study. The first approval (Act No. 1909, May 31, 2024) covered the development of the AI prototype. It determined that no informed consent was required as no personal data or interactions with human subjects were involved. The second approval (Act No. 2025012801 SJ, January 28, 2025) covered the validation process with healthcare professionals. For this phase, informed consent was obtained from all participants before their inclusion in the study. Both approvals were granted by national ethical regulations (Resolution 008430 of 1993).\u0026rdquo;\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eWe sincerely thank the team of obstetricians and anesthesiologists \u0026mdash; \u003cstrong\u003eDr. Mauricio Vasco, Dr. Camilo Fonseca, Dr. Amparo Ram\u0026iacute;rez, and Dr. Luis Carlos Franco\u003c/strong\u003e \u0026mdash; for their generous support in validating the prototype of DoctorBOT. Their clinical expertise was instrumental in assessing the relevance, accuracy, and utility of the chatbot\u0026apos;s responses.\u003c/p\u003e\n\u003cp\u003eWe also extend our deep gratitude to the rural physicians, nurses, and traditional birth attendants who contributed to this project by sharing real-life clinical questions and scenarios from their daily practice. Their insights, shaped by firsthand experience managing obstetric emergencies in low-resource settings, were essential to building a model grounded in the realities of frontline care.\u003c/p\u003e\n\u003cp\u003eThis work reflects not only technological innovation, but also the lived experiences and professional dedication of those who strive to save lives under challenging conditions.\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eData Sharing Statement\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eThe data, code, and supporting materials for this study are openly available on GitHub at: https://github.com/AmbreMychalski/Tesis-1.\u003c/p\u003e\n\u003cp\u003eThe repository includes:\u003c/p\u003e\n\u003cul type=\"disc\"\u003e\n \u003cli\u003eStructured prompt examples\u003c/li\u003e\n \u003cli\u003eThe chatbot\u0026rsquo;s retrieval pipeline (RAG framework)\u003c/li\u003e\n \u003cli\u003eEvaluation scripts for BERTScore analysis\u003c/li\u003e\n \u003cli\u003eSupporting documentation for replication\u003c/li\u003e\n \u003cli\u003eA data dictionary defining key fields in the prototype evaluation data\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThese materials are available immediately and without restriction for non-commercial use. Researchers may use the code and data for further analysis, adaptation, or model benchmarking. No individual patient data were collected or used in this study. No additional restrictions apply.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAI Disclosure\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe declare that we used artificial intelligence-based tools to support the translation and language editing of the manuscript, ensuring clarity and consistency throughout the text.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Research Ethics Committee at Universidad de los Andes approved the protocol. In accordance with current Colombian regulations \u0026mdash; Resolution 008430 of 1993 and Resolution 2378 of 2008 \u0026mdash; this study is classified as \u003cstrong\u003eminimal risk\u003c/strong\u003e\u003cstrong\u003e.\u003c/strong\u003e\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eGoldenberg RL, McClure EM, Saleem S. Improving pregnancy outcomes in low- and middle-income countries. Reproductive Health. Volume 15. BioMed Central Ltd.; 2018.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWHO UUWBG and UD. Trends in maternal mortality 2000 to 2020. Inform [Internet]. 2023 [cited 2025 Mar 10]; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.who.int/publications/i/item/9789240068759\u003c/span\u003e\u003cspan address=\"https://www.who.int/publications/i/item/9789240068759\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUnited Nations. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://sdgs.un.org/goals/goal3#targets_and_indicators\u003c/span\u003e\u003cspan address=\"https://sdgs.un.org/goals/goal3#targets_and_indicators\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. THE 17 GOALS.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBajwa J, Munir U, Nori A, Williams B. Artificial intelligence in healthcare: transforming the practice of medicine. Future Healthc J. 2021;8(2):e188\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSiriwardhana S, Weerasekera R, Wen E, Kaluarachchi T, Rajib R\u0026dagger;, Nanayakkara S. Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1162/tacl\u003c/span\u003e\u003cspan address=\"10.1162/tacl\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang T, Kishore V, Wu F, Weinberger KQ, Artzi Y, BERTScore. Evaluating Text Generation with BERT. 2019; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/1904.09675\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/1904.09675\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-research-notes","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"resn","sideBox":"Learn more about [BMC Research Notes](http://bmcresnotes.biomedcentral.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/resn/default.aspx","title":"BMC Research Notes","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Postpartum Hemorrhage, Maternal Mortality, Artificial Intelligence, Clinical Decision Support, Chatbot, Low-Resource Settings, Retrieval-Augmented Generation, Large Language Models, Evidence-Based Medicine","lastPublishedDoi":"10.21203/rs.3.rs-6580160/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6580160/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e Postpartum hemorrhage (PPH) remains a leading cause of maternal mortality, especially in low-resource settings. Despite the existence of high-quality clinical practice guidelines (CPGs), their implementation in rural areas is limited due to lack of training and access. Artificial intelligence (AI)-based tools may bridge this gap by supporting real-time clinical decision-making.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eObjective:\u003c/strong\u003e This study presents the development and evaluation of DoctorBOT, an AI-powered chatbot using Retrieval-Augmented Generation (RAG) to deliver evidence-based recommendations for PPH management.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods:\u003c/strong\u003e DoctorBOT was trained using vectorized high-quality obstetric CPGs. Clinical questions were collected from 26 rural healthcare workers and evaluated by large language models (LLMs) ChatGPT-3.5 and GPT-4. Performance was assessed through expert evaluation (Likert scale across five domains) and automated semantic comparison using BERTScore.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults:\u003c/strong\u003e GPT-3.5 outperformed GPT-4 in human evaluations for accuracy and clinical utility. Conversely, GPT-4 achieved a higher BERTScore (0.83 vs. 0.79). LLaMA2 achieved the highest score (0.92) after additional fine-tuning. In a user survey, 85% of rural healthcare providers rated DoctorBOT as “very satisfactory.”\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion:\u003c/strong\u003e DoctorBOT demonstrates strong potential to support evidence-based clinical decision-making in obstetric emergencies. Its integration could reduce maternal mortality and improve health equity in underserved areas.\u003c/p\u003e","manuscriptTitle":"DoctorBOT: An AI-powered Chatbot for Evidence-Based Obstetric Hemorrhage Management in Low-Resource Settings","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-04 10:25:19","doi":"10.21203/rs.3.rs-6580160/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-03-11T09:45:51+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-19T14:14:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"155921124295119376519037207318311088123","date":"2026-02-14T07:57:37+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-09-12T22:43:52+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"232861133208522289541885059279174006563","date":"2025-08-16T14:23:07+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"327873889771106857542528078972733825128","date":"2025-08-06T21:05:47+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"11047931572471035910997138352253167629","date":"2025-06-07T13:28:53+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-06-02T03:39:37+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-06-02T03:38:51+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-05-23T12:34:16+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-05-22T14:38:08+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Research Notes","date":"2025-05-22T14:37:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-research-notes","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"resn","sideBox":"Learn more about [BMC Research Notes](http://bmcresnotes.biomedcentral.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/resn/default.aspx","title":"BMC Research Notes","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e70dcc0f-d369-4b50-9fc5-e5bcddf4ba2a","owner":[],"postedDate":"June 4th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[],"tags":[],"updatedAt":"2026-03-11T09:56:09+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-04 10:25:19","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6580160","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6580160","identity":"rs-6580160","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.