Effect of Prompt Structure on Large Language Model-Recommended Triage Levels in Urologic Patient Messages | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Effect of Prompt Structure on Large Language Model-Recommended Triage Levels in Urologic Patient Messages Ahmet Olgun, Salim Zengin This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9453460/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract Purpose. To evaluate whether prompt structure changes large language model-recommended triage in urologic patient messages and whether safety-focused prompting reduces unsafe undertriage without increasing overtriage. Methods. In this scenario-based, cross-sectional, comparative study, 50 synthetic urologic patient messages were tested in three prompt arms: a neutral base prompt (P1), a red-flag-focused prompt (P2), and a safety- and triage-calibration-focused prompt (P3). This yielded 150 first-response outputs in Turkish from GPT-5.4 Thinking. For each case, two urologists predefined the gold-standard triage level and required core warnings. Primary outcomes were complete concordance, unsafe undertriage, and overtriage; omission of critical warnings was secondary. Results. Complete concordance was 45/50 (90.0%) in P1, 42/50 (84.0%) in P2, and 41/50 (82.0%) in P3. No unsafe undertriage or omission of critical warnings occurred in any arm. Overtriage was 2/50 (4.0%) in P1, 8/50 (16.0%) in P2, and 9/50 (18.0%) in P3. Compared with P1, overtriage was significantly higher in P2 and P3 (p = 0.031 and p = 0.016, respectively), with no difference between P2 and P3. Conclusion. In this dataset, safety-focused prompting did not improve unsafe undertriage but increased overtriage. Evaluations of patient-facing artificial intelligence should therefore assess not only red-flag recognition but also whether the recommended level of care is appropriate. Large language models Prompt structure Triage Urology Patient messages Overtriage Introduction Large language models are becoming increasingly visible in clinical communication as tools for answering patient questions and explaining health information. Studies across different medical fields have explored their potential in patient education and counseling, while also documenting important limitations in accuracy, reliability, and clinical appropriateness [ 1 ]. A similar trend has emerged in urology. Studies on stone disease, benign prostatic hyperplasia, and common urologic complaints suggest that these systems may have value for patient communication, but that response quality and safety still require careful evaluation [ 2 – 5 ]. However, most studies in the urologic literature have focused on information quality. Response accuracy, clarity, completeness, and user preference have often been assessed, whereas the urgency level to which the model directs the patient has been examined less often [ 2 – 5 ]. Yet the clinical value of a patient-facing response depends not only on the information it provides, but also on whether it directs the patient to an appropriate level and timing of care. Data on triage guidance remain limited. Early urology-specific studies suggest that the model can produce technically acceptable recommendations in acute urologic scenarios, but that the recommended level of care may vary across scenarios [ 6 ]. Studies outside urology likewise show that triage performance is not fixed and that the content of the prompt can influence the recommended urgency level [ 7 , 8 ]. Accordingly, the effect of prompt structure on model-recommended triage in urologic patient messages warrants separate evaluation. In this study, the same synthetic case set was presented to the model using three approaches: a neutral base prompt, a red-flag-focused prompt, and a safety- and triage-calibration-focused prompt. The aim was to compare the effects of these prompt structures on concordance with the gold-standard triage level, unsafe undertriage, overtriage, and omission of critical warnings. Materials and Methods Study Design This was a scenario-based, cross-sectional, comparative study designed to evaluate how the large language model–recommended triage level changed across different prompt structures when applied to synthetic urologic patient messages. The sample size was not based on a formal power calculation. Instead, the case set was developed purposively to allow comparison across a range of clinical scenarios and triage levels. Case Set and Gold-Standard Definition A total of 50 synthetic patient messages were used. The scenarios were predefined to represent different urgency levels in routine urologic practice and were grouped under eight clinical categories: acute scrotum/torsion, renal colic/urolithiasis, urinary tract infection, hematuria/clots, retention/catheter-related presentations, priapism, genital-perineal infection/Fournier gangrene, and stent- or nephrostomy-related complaints. The case set was designed to represent clinically distinctive situations encountered in daily urologic practice across low-, intermediate-, and high-acuity triage levels. For each case, the gold-standard triage level and the required core warning messages were jointly defined by two urologists before data collection, after the study scenarios had been finalized. In this process, the clinical priority of each case and the key safety messages that should be conveyed to the patient were specified in advance. Model outputs were then evaluated against this predefined framework. The gold-standard triage system had three levels: level 1, immediate emergency evaluation; level 2, urgent evaluation or same-day/short-term in-person evaluation; and level 3, routine or planned outpatient evaluation. Code 9 was reserved for responses in which triage timing truly could not be inferred from the model output. However, because all final responses could be assigned to level 1, 2, or 3, code 9 was not used in the final dataset. Prompt Arms Each case was presented to the model in three prompt arms while preserving the same clinical content. The first arm was defined as the neutral base prompt (P1) and was intended to generate a short, general, practical, non-diagnostic response to the patient; this arm did not provide the model with any additional red-flag or safety-oriented framing. The second arm was designed as the red-flag-focused prompt (P2) and asked the model to assess possible emergency warning signs before responding. The third arm was designed as the safety- and triage-calibration-focused prompt (P3); in this arm, the model was asked to state the urgency level clearly, explicitly describe situations that might require emergency evaluation, and avoid recommending unnecessary emergency care when no emergency features were present. The full text of the P1, P2, and P3 prompt templates, together with the coding framework used for triage classification and omission of critical warnings, is provided in Online Resource 1. Generation of Model Responses Responses were obtained using temporary chats within the same paid ChatGPT account. During data collection, memory, prior conversation effects, project context, and custom instructions were disabled. A separate new temporary chat was opened for each case and each prompt arm; web search, file upload, and additional tools were not used. All outputs were recorded in Turkish and based on the first response only. Throughout the data collection process, the model was manually set to GPT-5.4 Thinking. Instant mode or automatic routing was not used. Only the first response was recorded for each case; no follow-up questions were asked, and no requests were made for rewriting, shortening, or correction. All responses were obtained on April 13, 2026. Coding Approach Each model response was coded according to predefined rules using the three-level triage system. Coding was based not on exact wording alone, but on the practical level of care to which the response directed the patient. Statements such as “immediately,” “emergency,” “call emergency services,” or “do not wait” were coded as level 1. Statements such as “today,” “the same day,” “without delay,” “as soon as possible,” or “in-person evaluation within a short time” were coded as level 2. Statements such as “routine,” “planned,” “outpatient clinic,” or “appointment-based evaluation” were coded as level 3. When a response contained more than one timing expression, coding was based on the dominant recommendation that would most strongly direct the patient's behavior. Accordingly, coding reflected the functional triage effect of the response rather than surface-level lexical similarity. Triage coding and assessment of omission of critical warnings were performed independently by two urologists using the predefined study rules. Initial inter-rater agreement was also examined. Although joint review had been planned for responses with disagreement, complete agreement was achieved for all responses at the initial assessment. Outcomes and Definitions The primary outcomes were complete concordance, unsafe undertriage, and overtriage. The secondary outcome was omission of critical warnings. Complete concordance was defined as an exact match between the model-assigned three-level triage code and the predefined gold-standard triage code for that case. Unsafe undertriage was defined as a model recommendation of level 3 when the gold standard was level 1 or 2. Overtriage was defined as the model recommending a more urgent triage level than the gold standard. Omission of critical warnings was evaluated at the level of meaning rather than word-for-word matching. Thus, even when a model response appeared appropriate in terms of triage level, this outcome was coded as positive if the response did not meaningfully include the predefined key safety messages for that case. Statistical Analysis For each prompt arm, complete concordance, unsafe undertriage, overtriage, and omission of critical warnings were summarized as counts and percentages. Classification distributions according to gold-standard triage level were also reported. Initial inter-rater agreement for triage coding was assessed using Cohen's kappa coefficient. To make the magnitude of differences between prompt arms easier to interpret, absolute rate differences were also reported for the relevant outcomes. For pairwise comparisons between matched prompt arms, exact McNemar tests were used for outcomes with event counts suitable for comparison. Because unsafe undertriage and omission of critical warnings were not observed in any arm, no comparative statistical tests were applied for these outcomes. Statistical significance was defined as p < 0.05. Final Analysis Set The final analysis included a total of 150 model responses derived from 50 cases evaluated across the three prompt arms. All comparisons were performed using the paired data structure in which all three prompt arms were available for each case. Results A total of 50 synthetic urologic patient cases were evaluated across the three prompt arms, yielding 150 model responses for analysis. The prompt arms consisted of the neutral base prompt (P1), the red-flag-focused prompt (P2), and the safety- and triage-calibration-focused prompt (P3). For each arm, the model-assigned three-level triage code was compared with the gold-standard triage code. Unsafe undertriage, overtriage, and omission of critical warnings were also recorded as separate outcomes. Initial inter-rater agreement for triage coding was 100% (150/150; Cohen's kappa = 1.00). Initial agreement for omission of critical warnings was also 100% (150/150). Complete concordance was 45/50 (90.0%) in P1, 42/50 (84.0%) in P2, and 41/50 (82.0%) in P3. No unsafe undertriage or omission of critical warnings was observed in any prompt arm. Overtriage was 2/50 (4.0%) in P1, 8/50 (16.0%) in P2, and 9/50 (18.0%) in P3. Classification distributions showed that the differences between prompt arms were driven mainly by cases with a gold-standard triage level of 3; in P2 and P3, these lower-acuity cases were more often shifted to level 2. By contrast, no case with a gold-standard triage level of 1 or 2 was classified as level 3 in any arm. Overtriage was higher in P2 and P3 than in P1. This corresponded to 6 additional cases and an absolute increase of 12 percentage points for P2, and 7 additional cases and an absolute increase of 14 percentage points for P3 (p = 0.031 and p = 0.016, respectively). No significant difference was observed between P2 and P3 (p = 1.00). Table 1 Overall performance by prompt arm Prompt arm Complete concordance n/N (%) Unsafe undertriage n/N (%) Overtriage n/N (%) Omission of critical warnings n/N (%) P1 (neutral base prompt) 45/50 (90.0) 0/50 (0.0) 2/50 (4.0) 0/50 (0.0) P2 (red-flag-focused prompt) 42/50 (84.0) 0/50 (0.0) 8/50 (16.0) 0/50 (0.0) P3 (safety- and triage-calibration-focused prompt) 41/50 (82.0) 0/50 (0.0) 9/50 (18.0) 0/50 (0.0) Table 2 Classification distribution according to gold-standard triage level Gold-standard triage P1 P2 P3 1 (n = 12) 9 cases as level 1, 3 cases as level 2, 0 cases as level 3 12 cases as level 1, 0 cases as level 2, 0 cases as level 3 12 cases as level 1, 0 cases as level 2, 0 cases as level 3 2 (n = 12) 0 cases as level 1, 12 cases as level 2, 0 cases as level 3 1 case as level 1, 11 cases as level 2, 0 cases as level 3 1 case as level 1, 11 cases as level 2, 0 cases as level 3 3 (n = 26) 0 cases as level 1, 2 cases as level 2, 24 cases as level 3 0 cases as level 1, 7 cases as level 2, 19 cases as level 3 0 cases as level 1, 8 cases as level 2, 18 cases as level 3 Discussion The main finding of this study is that changes in prompt structure can materially alter model triage behavior. The neutral base prompt showed the highest concordance with the gold standard, whereas the red-flag-focused and safety-focused prompts were not associated with any unsafe undertriage but were associated with higher overtriage. In this dataset, emphasizing safety did not provide an additional safety benefit and may have shifted the triage threshold toward recommending a more urgent level of care. The differences between prompt arms were most evident in cases with a gold-standard triage level of 3. While the neutral base prompt kept most of these responses at the routine level, the red-flag-focused and safety-focused prompts more often shifted these lower-acuity cases to level 2. In contrast, no case with a gold-standard triage level of 1 or 2 was shifted down to level 3 in any prompt arm. The prompt effect in this study therefore became visible primarily as more cautious upward triage rather than as missed urgency. Taken together with the existing literature, these findings suggest that triage output is not a fixed property of the model. In the same case set, changing only the prompt structure changed the recommended level of care. This study therefore supports the view that evaluations of patient-facing artificial intelligence responses should consider not only factual correctness or the presence of red-flag content, but also the appropriateness of the recommended care level. A strength of this study is its design, which was well suited to isolate the effect of prompt structure from case content. The same 50 cases were evaluated in a paired manner across three prompt structures, allowing the observed differences to be linked directly to prompt framing. The gold-standard triage levels and required core warnings were jointly defined by two urologists before data collection, which clarified the evaluation framework. Methodological consistency was further supported by obtaining all responses from the same account, using the same model, working in temporary chats, and recording only the first output for each case. Taken together, these features provide a solid framework for interpreting the effect of prompt structure on triage behavior in urologic patient messages. This study also has limitations. First, the evaluation was performed with a single model in a specific version, and the findings therefore cannot be directly generalized to other models, versions, or platform conditions. Second, the case set consisted of synthetic patient messages. Although this approach provided a controlled setting for comparison, it does not fully capture the variation in real patient wording, unexpected linguistic ambiguity, or the dynamics of interactive follow-up questioning. Third, the study was based only on the first responses generated in Turkish. The findings may therefore not apply in the same way to other languages or to multi-turn interactions. Finally, this study evaluated the model's textual guidance behavior rather than real clinical outcomes. The actual effects of the observed differences on patient behavior, health care use, and clinical outcomes remain outside the scope of this work. No unsafe undertriage or omission of critical warnings was observed in any prompt arm. This suggests that no clear between-arm difference was detected for these outcomes in this dataset; however, the case set, sample size, or limited discriminatory capacity of the outcome definitions may also have contributed to this result. Conclusion In conclusion, prompt structure meaningfully influences model triage behavior in urologic patient messages. In this study, safety-focused prompts did not show an added advantage for unsafe undertriage, but they were associated with higher overtriage rates. These findings indicate that evaluations of patient-facing artificial intelligence responses should consider not only the presence of red flags, but also the appropriateness of the recommended level of care. In lower-risk urologic messages, safety-framed prompts may be associated with recommendations for a higher level of care. Declarations Supplementary Information Online Resource 1 contains the full prompt templates, English back-translations, coding framework, and original Turkish synthetic case messages used in the study. Competing Interests The authors declare that they have no competing interests. Research Involving Human Participants and/or Animals This study used only synthetic patient messages and did not involve human participants, human data, human tissue, or animals. Formal ethics committee approval was therefore not required. Informed Consent Not applicable. Funding This study received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Author Contribution AO: Protocol/project development, data collection or management, data analysis, manuscript writing/editing. SZ: Protocol/project development, data collection or management, data analysis, manuscript writing/editing. Both authors approved the final manuscript and accept accountability for all aspects of the work. Data Availability The synthetic case set, prompt framework, and coded output data used in this study are available from the corresponding author on reasonable request. References Sequi MB, Pastore AL, Al Salhi Y et al (2026) Current role of artificial intelligence in the management of benign prostatic hyperplasia: a systematic review. Minerva Urol Nephrol 78(1):40–52. https://doi.org/10.23736/S2724-6051.25.06658-3 Javid M, Bhandari M, Parameshwari P et al (2024) Evaluation of ChatGPT for patient counseling in kidney stone clinic: a prospective study. J Endourol 38(4):377–383. https://doi.org/10.1089/end.2023.0571 Warren CJ, Payne NG, Edmonds VS et al (2025) Quality of chatbot information related to benign prostatic hyperplasia. Prostate 85(2):175–180. https://doi.org/10.1002/pros.24814 Robinson EJ, Qiu C, Sands S et al (2025) Physician vs. AI-generated messages in urology: evaluation of accuracy, completeness, and preference by patients and physicians. World J Urol 43:48. https://doi.org/10.1007/s00345-024-05399-y Szczesniewski JJ, Tellez Fouz C, Ramos Alba A et al (2023) ChatGPT and most frequent urological diseases: analysing the quality of information and potential risks for patients. World J Urol 41(11):3149–3153. https://doi.org/10.1007/s00345-023-04563-0 Hirtsiefer C, Nestler T, Eckrich J et al (2024) Capabilities of ChatGPT-3.5 as a urological triage system. Eur Urol Open Sci 70:148–153. https://doi.org/10.1016/j.euros.2024.10.015 Masanneck L, Schmidt L, Seifert A et al (2024) Triage performance across large language models, ChatGPT, and untrained doctors in emergency medicine: comparative study. J Med Internet Res 26:e53297. https://doi.org/10.2196/53297 Wang C, Wang F, Li S et al (2025) Patient triage and guidance in emergency departments using large language models: multimetric study. J Med Internet Res 27:e71613. https://doi.org/10.2196/71613 Ortaç M, Ergül RB, Yazılı HB et al (2025) ChatGPT's competence in responding to urological emergencies. Turk J Trauma Emerg Surg 31(3):291–295. https://doi.org/10.14744/tjtes.2024.03377 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 16 May, 2026 Reviewers invited by journal 22 Apr, 2026 Editor assigned by journal 22 Apr, 2026 Submission checks completed at journal 22 Apr, 2026 First submitted to journal 17 Apr, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9453460","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":627448648,"identity":"87db6cd0-5142-460b-9e17-fafb932ff24d","order_by":0,"name":"Ahmet Olgun","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABAElEQVRIiWNgGAWjYDACCRBRYcPAB6I/ADEbOzFaDpxJY2AD0owzQFqYidFysO0wWAszD0iEkBb+2d2Jjz+cOSzHxr/G7LHNr23yfMwMjB8+5uCx5M7ZzQYHKtKN2STemBvn9t02bGNmYJacuQ2PNTdyt0kcOGOd2CZxxkw6t+c2I1ALGzMvHi3yN3K3/zjYxgzRYtlz256gFgOgLUDvOye28feYSTP8uJ1IUIvhjdzNEmfOpAH9wlYm2dtwO7mNmbEZr1/kbuRu/FBRYSPHz394m8SPP7dt57c3H/zwEZ/34UAiARiXbSAWYwMx6oGA/wCQ+EOk4lEwCkbBKBhRAAA35FMbfhT7qAAAAABJRU5ErkJggg==","orcid":"","institution":"","correspondingAuthor":true,"prefix":"","firstName":"Ahmet","middleName":"","lastName":"Olgun","suffix":""},{"id":627448649,"identity":"7c1d6e89-44c6-4519-b73c-8bc02dfb31df","order_by":1,"name":"Salim Zengin","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Salim","middleName":"","lastName":"Zengin","suffix":""}],"badges":[],"createdAt":"2026-04-18 02:08:29","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9453460/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9453460/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108491502,"identity":"f94a9325-e574-4f76-ab31-2d47e26fd600","added_by":"auto","created_at":"2026-05-05 09:54:12","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":163735,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9453460/v1/f708cf70-c14e-49cc-8ae8-013fdfa777c5.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Effect of Prompt Structure on Large Language Model-Recommended Triage Levels in Urologic Patient Messages","fulltext":[{"header":"Introduction","content":"\u003cp\u003eLarge language models are becoming increasingly visible in clinical communication as tools for answering patient questions and explaining health information. Studies across different medical fields have explored their potential in patient education and counseling, while also documenting important limitations in accuracy, reliability, and clinical appropriateness [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. A similar trend has emerged in urology. Studies on stone disease, benign prostatic hyperplasia, and common urologic complaints suggest that these systems may have value for patient communication, but that response quality and safety still require careful evaluation [\u003cspan additionalcitationids=\"CR3 CR4\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eHowever, most studies in the urologic literature have focused on information quality. Response accuracy, clarity, completeness, and user preference have often been assessed, whereas the urgency level to which the model directs the patient has been examined less often [\u003cspan additionalcitationids=\"CR3 CR4\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Yet the clinical value of a patient-facing response depends not only on the information it provides, but also on whether it directs the patient to an appropriate level and timing of care.\u003c/p\u003e \u003cp\u003eData on triage guidance remain limited. Early urology-specific studies suggest that the model can produce technically acceptable recommendations in acute urologic scenarios, but that the recommended level of care may vary across scenarios [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Studies outside urology likewise show that triage performance is not fixed and that the content of the prompt can influence the recommended urgency level [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAccordingly, the effect of prompt structure on model-recommended triage in urologic patient messages warrants separate evaluation. In this study, the same synthetic case set was presented to the model using three approaches: a neutral base prompt, a red-flag-focused prompt, and a safety- and triage-calibration-focused prompt. The aim was to compare the effects of these prompt structures on concordance with the gold-standard triage level, unsafe undertriage, overtriage, and omission of critical warnings.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Design\u003c/h2\u003e \u003cp\u003eThis was a scenario-based, cross-sectional, comparative study designed to evaluate how the large language model\u0026ndash;recommended triage level changed across different prompt structures when applied to synthetic urologic patient messages. The sample size was not based on a formal power calculation. Instead, the case set was developed purposively to allow comparison across a range of clinical scenarios and triage levels.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eCase Set and Gold-Standard Definition\u003c/h3\u003e\n\u003cp\u003eA total of 50 synthetic patient messages were used. The scenarios were predefined to represent different urgency levels in routine urologic practice and were grouped under eight clinical categories: acute scrotum/torsion, renal colic/urolithiasis, urinary tract infection, hematuria/clots, retention/catheter-related presentations, priapism, genital-perineal infection/Fournier gangrene, and stent- or nephrostomy-related complaints. The case set was designed to represent clinically distinctive situations encountered in daily urologic practice across low-, intermediate-, and high-acuity triage levels.\u003c/p\u003e \u003cp\u003eFor each case, the gold-standard triage level and the required core warning messages were jointly defined by two urologists before data collection, after the study scenarios had been finalized. In this process, the clinical priority of each case and the key safety messages that should be conveyed to the patient were specified in advance. Model outputs were then evaluated against this predefined framework.\u003c/p\u003e \u003cp\u003eThe gold-standard triage system had three levels: level 1, immediate emergency evaluation; level 2, urgent evaluation or same-day/short-term in-person evaluation; and level 3, routine or planned outpatient evaluation. Code 9 was reserved for responses in which triage timing truly could not be inferred from the model output. However, because all final responses could be assigned to level 1, 2, or 3, code 9 was not used in the final dataset.\u003c/p\u003e\n\u003ch3\u003ePrompt Arms\u003c/h3\u003e\n\u003cp\u003eEach case was presented to the model in three prompt arms while preserving the same clinical content. The first arm was defined as the neutral base prompt (P1) and was intended to generate a short, general, practical, non-diagnostic response to the patient; this arm did not provide the model with any additional red-flag or safety-oriented framing. The second arm was designed as the red-flag-focused prompt (P2) and asked the model to assess possible emergency warning signs before responding. The third arm was designed as the safety- and triage-calibration-focused prompt (P3); in this arm, the model was asked to state the urgency level clearly, explicitly describe situations that might require emergency evaluation, and avoid recommending unnecessary emergency care when no emergency features were present.\u003c/p\u003e \u003cp\u003eThe full text of the P1, P2, and P3 prompt templates, together with the coding framework used for triage classification and omission of critical warnings, is provided in Online Resource 1.\u003c/p\u003e\n\u003ch3\u003eGeneration of Model Responses\u003c/h3\u003e\n\u003cp\u003e Responses were obtained using temporary chats within the same paid ChatGPT account. During data collection, memory, prior conversation effects, project context, and custom instructions were disabled. A separate new temporary chat was opened for each case and each prompt arm; web search, file upload, and additional tools were not used. All outputs were recorded in Turkish and based on the first response only.\u003c/p\u003e \u003cp\u003eThroughout the data collection process, the model was manually set to GPT-5.4 Thinking. Instant mode or automatic routing was not used. Only the first response was recorded for each case; no follow-up questions were asked, and no requests were made for rewriting, shortening, or correction. All responses were obtained on April 13, 2026.\u003c/p\u003e\n\u003ch3\u003eCoding Approach\u003c/h3\u003e\n\u003cp\u003eEach model response was coded according to predefined rules using the three-level triage system. Coding was based not on exact wording alone, but on the practical level of care to which the response directed the patient. Statements such as \u0026ldquo;immediately,\u0026rdquo; \u0026ldquo;emergency,\u0026rdquo; \u0026ldquo;call emergency services,\u0026rdquo; or \u0026ldquo;do not wait\u0026rdquo; were coded as level 1. Statements such as \u0026ldquo;today,\u0026rdquo; \u0026ldquo;the same day,\u0026rdquo; \u0026ldquo;without delay,\u0026rdquo; \u0026ldquo;as soon as possible,\u0026rdquo; or \u0026ldquo;in-person evaluation within a short time\u0026rdquo; were coded as level 2. Statements such as \u0026ldquo;routine,\u0026rdquo; \u0026ldquo;planned,\u0026rdquo; \u0026ldquo;outpatient clinic,\u0026rdquo; or \u0026ldquo;appointment-based evaluation\u0026rdquo; were coded as level 3.\u003c/p\u003e \u003cp\u003eWhen a response contained more than one timing expression, coding was based on the dominant recommendation that would most strongly direct the patient's behavior. Accordingly, coding reflected the functional triage effect of the response rather than surface-level lexical similarity. Triage coding and assessment of omission of critical warnings were performed independently by two urologists using the predefined study rules. Initial inter-rater agreement was also examined. Although joint review had been planned for responses with disagreement, complete agreement was achieved for all responses at the initial assessment.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eOutcomes and Definitions\u003c/h2\u003e \u003cp\u003eThe primary outcomes were complete concordance, unsafe undertriage, and overtriage. The secondary outcome was omission of critical warnings.\u003c/p\u003e \u003cp\u003eComplete concordance was defined as an exact match between the model-assigned three-level triage code and the predefined gold-standard triage code for that case.\u003c/p\u003e \u003cp\u003eUnsafe undertriage was defined as a model recommendation of level 3 when the gold standard was level 1 or 2. Overtriage was defined as the model recommending a more urgent triage level than the gold standard. Omission of critical warnings was evaluated at the level of meaning rather than word-for-word matching. Thus, even when a model response appeared appropriate in terms of triage level, this outcome was coded as positive if the response did not meaningfully include the predefined key safety messages for that case.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eFor each prompt arm, complete concordance, unsafe undertriage, overtriage, and omission of critical warnings were summarized as counts and percentages. Classification distributions according to gold-standard triage level were also reported. Initial inter-rater agreement for triage coding was assessed using Cohen's kappa coefficient. To make the magnitude of differences between prompt arms easier to interpret, absolute rate differences were also reported for the relevant outcomes.\u003c/p\u003e \u003cp\u003eFor pairwise comparisons between matched prompt arms, exact McNemar tests were used for outcomes with event counts suitable for comparison. Because unsafe undertriage and omission of critical warnings were not observed in any arm, no comparative statistical tests were applied for these outcomes. Statistical significance was defined as p\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eFinal Analysis Set\u003c/h3\u003e\n\u003cp\u003eThe final analysis included a total of 150 model responses derived from 50 cases evaluated across the three prompt arms. All comparisons were performed using the paired data structure in which all three prompt arms were available for each case.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eA total of 50 synthetic urologic patient cases were evaluated across the three prompt arms, yielding 150 model responses for analysis. The prompt arms consisted of the neutral base prompt (P1), the red-flag-focused prompt (P2), and the safety- and triage-calibration-focused prompt (P3). For each arm, the model-assigned three-level triage code was compared with the gold-standard triage code. Unsafe undertriage, overtriage, and omission of critical warnings were also recorded as separate outcomes. Initial inter-rater agreement for triage coding was 100% (150/150; Cohen's kappa\u0026thinsp;=\u0026thinsp;1.00). Initial agreement for omission of critical warnings was also 100% (150/150).\u003c/p\u003e \u003cp\u003eComplete concordance was 45/50 (90.0%) in P1, 42/50 (84.0%) in P2, and 41/50 (82.0%) in P3. No unsafe undertriage or omission of critical warnings was observed in any prompt arm. Overtriage was 2/50 (4.0%) in P1, 8/50 (16.0%) in P2, and 9/50 (18.0%) in P3. Classification distributions showed that the differences between prompt arms were driven mainly by cases with a gold-standard triage level of 3; in P2 and P3, these lower-acuity cases were more often shifted to level 2. By contrast, no case with a gold-standard triage level of 1 or 2 was classified as level 3 in any arm. Overtriage was higher in P2 and P3 than in P1. This corresponded to 6 additional cases and an absolute increase of 12 percentage points for P2, and 7 additional cases and an absolute increase of 14 percentage points for P3 (p\u0026thinsp;=\u0026thinsp;0.031 and p\u0026thinsp;=\u0026thinsp;0.016, respectively). No significant difference was observed between P2 and P3 (p\u0026thinsp;=\u0026thinsp;1.00).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eOverall performance by prompt arm\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePrompt arm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eComplete concordance\u003c/p\u003e \u003cp\u003en/N (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUnsafe undertriage\u003c/p\u003e \u003cp\u003en/N (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOvertriage\u003c/p\u003e \u003cp\u003en/N (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eOmission of critical warnings\u003c/p\u003e \u003cp\u003en/N (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eP1 (neutral base prompt)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45/50 (90.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0/50 (0.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2/50 (4.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0/50 (0.0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eP2 (red-flag-focused prompt)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e42/50 (84.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0/50 (0.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e8/50 (16.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0/50 (0.0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eP3 (safety- and triage-calibration-focused prompt)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e41/50 (82.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0/50 (0.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9/50 (18.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0/50 (0.0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eClassification distribution according to gold-standard triage level\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGold-standard triage\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eP1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eP2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eP3\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1 (n\u0026thinsp;=\u0026thinsp;12)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e9 cases as level 1,\u003c/p\u003e \u003cp\u003e3 cases as level 2,\u003c/p\u003e \u003cp\u003e0 cases as level 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e12 cases as level 1,\u003c/p\u003e \u003cp\u003e0 cases as level 2,\u003c/p\u003e \u003cp\u003e0 cases as level 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e12 cases as level 1,\u003c/p\u003e \u003cp\u003e0 cases as level 2,\u003c/p\u003e \u003cp\u003e0 cases as level 3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2 (n\u0026thinsp;=\u0026thinsp;12)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0 cases as level 1,\u003c/p\u003e \u003cp\u003e12 cases as level 2,\u003c/p\u003e \u003cp\u003e0 cases as level 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1 case as level 1,\u003c/p\u003e \u003cp\u003e11 cases as level 2,\u003c/p\u003e \u003cp\u003e0 cases as level 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1 case as level 1,\u003c/p\u003e \u003cp\u003e11 cases as level 2,\u003c/p\u003e \u003cp\u003e0 cases as level 3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3 (n\u0026thinsp;=\u0026thinsp;26)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0 cases as level 1,\u003c/p\u003e \u003cp\u003e2 cases as level 2,\u003c/p\u003e \u003cp\u003e24 cases as level 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0 cases as level 1,\u003c/p\u003e \u003cp\u003e7 cases as level 2,\u003c/p\u003e \u003cp\u003e19 cases as level 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0 cases as level 1,\u003c/p\u003e \u003cp\u003e8 cases as level 2,\u003c/p\u003e \u003cp\u003e18 cases as level 3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe main finding of this study is that changes in prompt structure can materially alter model triage behavior. The neutral base prompt showed the highest concordance with the gold standard, whereas the red-flag-focused and safety-focused prompts were not associated with any unsafe undertriage but were associated with higher overtriage. In this dataset, emphasizing safety did not provide an additional safety benefit and may have shifted the triage threshold toward recommending a more urgent level of care.\u003c/p\u003e \u003cp\u003eThe differences between prompt arms were most evident in cases with a gold-standard triage level of 3. While the neutral base prompt kept most of these responses at the routine level, the red-flag-focused and safety-focused prompts more often shifted these lower-acuity cases to level 2. In contrast, no case with a gold-standard triage level of 1 or 2 was shifted down to level 3 in any prompt arm. The prompt effect in this study therefore became visible primarily as more cautious upward triage rather than as missed urgency.\u003c/p\u003e \u003cp\u003eTaken together with the existing literature, these findings suggest that triage output is not a fixed property of the model. In the same case set, changing only the prompt structure changed the recommended level of care. This study therefore supports the view that evaluations of patient-facing artificial intelligence responses should consider not only factual correctness or the presence of red-flag content, but also the appropriateness of the recommended care level.\u003c/p\u003e \u003cp\u003eA strength of this study is its design, which was well suited to isolate the effect of prompt structure from case content. The same 50 cases were evaluated in a paired manner across three prompt structures, allowing the observed differences to be linked directly to prompt framing. The gold-standard triage levels and required core warnings were jointly defined by two urologists before data collection, which clarified the evaluation framework. Methodological consistency was further supported by obtaining all responses from the same account, using the same model, working in temporary chats, and recording only the first output for each case. Taken together, these features provide a solid framework for interpreting the effect of prompt structure on triage behavior in urologic patient messages.\u003c/p\u003e \u003cp\u003eThis study also has limitations. First, the evaluation was performed with a single model in a specific version, and the findings therefore cannot be directly generalized to other models, versions, or platform conditions. Second, the case set consisted of synthetic patient messages. Although this approach provided a controlled setting for comparison, it does not fully capture the variation in real patient wording, unexpected linguistic ambiguity, or the dynamics of interactive follow-up questioning. Third, the study was based only on the first responses generated in Turkish. The findings may therefore not apply in the same way to other languages or to multi-turn interactions. Finally, this study evaluated the model's textual guidance behavior rather than real clinical outcomes. The actual effects of the observed differences on patient behavior, health care use, and clinical outcomes remain outside the scope of this work. No unsafe undertriage or omission of critical warnings was observed in any prompt arm. This suggests that no clear between-arm difference was detected for these outcomes in this dataset; however, the case set, sample size, or limited discriminatory capacity of the outcome definitions may also have contributed to this result.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn conclusion, prompt structure meaningfully influences model triage behavior in urologic patient messages. In this study, safety-focused prompts did not show an added advantage for unsafe undertriage, but they were associated with higher overtriage rates. These findings indicate that evaluations of patient-facing artificial intelligence responses should consider not only the presence of red flags, but also the appropriateness of the recommended level of care. In lower-risk urologic messages, safety-framed prompts may be associated with recommendations for a higher level of care.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eSupplementary Information\u003c/h2\u003e \u003cp\u003eOnline Resource 1 contains the full prompt templates, English back-translations, coding framework, and original Turkish synthetic case messages used in the study.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eCompeting Interests\u003c/h2\u003e \u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eResearch Involving Human Participants and/or Animals\u003c/h2\u003e \u003cp\u003e This study used only synthetic patient messages and did not involve human participants, human data, human tissue, or animals. Formal ethics committee approval was therefore not required.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eInformed Consent\u003c/h2\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis study received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAO: Protocol/project development, data collection or management, data analysis, manuscript writing/editing. SZ: Protocol/project development, data collection or management, data analysis, manuscript writing/editing. Both authors approved the final manuscript and accept accountability for all aspects of the work.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe synthetic case set, prompt framework, and coded output data used in this study are available from the corresponding author on reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSequi MB, Pastore AL, Al Salhi Y et al (2026) Current role of artificial intelligence in the management of benign prostatic hyperplasia: a systematic review. Minerva Urol Nephrol 78(1):40\u0026ndash;52. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.23736/S2724-6051.25.06658-3\u003c/span\u003e\u003cspan address=\"10.23736/S2724-6051.25.06658-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJavid M, Bhandari M, Parameshwari P et al (2024) Evaluation of ChatGPT for patient counseling in kidney stone clinic: a prospective study. J Endourol 38(4):377\u0026ndash;383. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1089/end.2023.0571\u003c/span\u003e\u003cspan address=\"10.1089/end.2023.0571\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWarren CJ, Payne NG, Edmonds VS et al (2025) Quality of chatbot information related to benign prostatic hyperplasia. Prostate 85(2):175\u0026ndash;180. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/pros.24814\u003c/span\u003e\u003cspan address=\"10.1002/pros.24814\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRobinson EJ, Qiu C, Sands S et al (2025) Physician vs. AI-generated messages in urology: evaluation of accuracy, completeness, and preference by patients and physicians. World J Urol 43:48. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s00345-024-05399-y\u003c/span\u003e\u003cspan address=\"10.1007/s00345-024-05399-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSzczesniewski JJ, Tellez Fouz C, Ramos Alba A et al (2023) ChatGPT and most frequent urological diseases: analysing the quality of information and potential risks for patients. World J Urol 41(11):3149\u0026ndash;3153. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s00345-023-04563-0\u003c/span\u003e\u003cspan address=\"10.1007/s00345-023-04563-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHirtsiefer C, Nestler T, Eckrich J et al (2024) Capabilities of ChatGPT-3.5 as a urological triage system. Eur Urol Open Sci 70:148\u0026ndash;153. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.euros.2024.10.015\u003c/span\u003e\u003cspan address=\"10.1016/j.euros.2024.10.015\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMasanneck L, Schmidt L, Seifert A et al (2024) Triage performance across large language models, ChatGPT, and untrained doctors in emergency medicine: comparative study. J Med Internet Res 26:e53297. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2196/53297\u003c/span\u003e\u003cspan address=\"10.2196/53297\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang C, Wang F, Li S et al (2025) Patient triage and guidance in emergency departments using large language models: multimetric study. J Med Internet Res 27:e71613. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2196/71613\u003c/span\u003e\u003cspan address=\"10.2196/71613\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOrta\u0026ccedil; M, Erg\u0026uuml;l RB, Yazılı HB et al (2025) ChatGPT's competence in responding to urological emergencies. Turk J Trauma Emerg Surg 31(3):291\u0026ndash;295. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.14744/tjtes.2024.03377\u003c/span\u003e\u003cspan address=\"10.14744/tjtes.2024.03377\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"world-journal-of-urology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"wjur","sideBox":"Learn more about [World Journal of Urology](https://link.springer.com/journal/345)","snPcode":"345","submissionUrl":"https://submission.nature.com/new-submission/345/3","title":"World Journal of Urology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Large language models, Prompt structure, Triage, Urology, Patient messages, Overtriage","lastPublishedDoi":"10.21203/rs.3.rs-9453460/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9453460/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003ePurpose. To evaluate whether prompt structure changes large language model-recommended triage in urologic patient messages and whether safety-focused prompting reduces unsafe undertriage without increasing overtriage.\u003c/p\u003e \u003cp\u003eMethods. In this scenario-based, cross-sectional, comparative study, 50 synthetic urologic patient messages were tested in three prompt arms: a neutral base prompt (P1), a red-flag-focused prompt (P2), and a safety- and triage-calibration-focused prompt (P3). This yielded 150 first-response outputs in Turkish from GPT-5.4 Thinking. For each case, two urologists predefined the gold-standard triage level and required core warnings. Primary outcomes were complete concordance, unsafe undertriage, and overtriage; omission of critical warnings was secondary.\u003c/p\u003e \u003cp\u003eResults. Complete concordance was 45/50 (90.0%) in P1, 42/50 (84.0%) in P2, and 41/50 (82.0%) in P3. No unsafe undertriage or omission of critical warnings occurred in any arm. Overtriage was 2/50 (4.0%) in P1, 8/50 (16.0%) in P2, and 9/50 (18.0%) in P3. Compared with P1, overtriage was significantly higher in P2 and P3 (p\u0026thinsp;=\u0026thinsp;0.031 and p\u0026thinsp;=\u0026thinsp;0.016, respectively), with no difference between P2 and P3.\u003c/p\u003e \u003cp\u003eConclusion. In this dataset, safety-focused prompting did not improve unsafe undertriage but increased overtriage. Evaluations of patient-facing artificial intelligence should therefore assess not only red-flag recognition but also whether the recommended level of care is appropriate.\u003c/p\u003e","manuscriptTitle":"Effect of Prompt Structure on Large Language Model-Recommended Triage Levels in Urologic Patient Messages","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-30 11:06:57","doi":"10.21203/rs.3.rs-9453460/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"226711902506117477295498057611234514038","date":"2026-05-16T10:04:30+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-22T04:54:13+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-04-22T04:22:29+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-04-22T04:21:31+00:00","index":"","fulltext":""},{"type":"submitted","content":"World Journal of Urology","date":"2026-04-18T02:01:06+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"world-journal-of-urology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"wjur","sideBox":"Learn more about [World Journal of Urology](https://link.springer.com/journal/345)","snPcode":"345","submissionUrl":"https://submission.nature.com/new-submission/345/3","title":"World Journal of Urology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"723f73a9-c56c-4e2a-b972-8a044bc7e59a","owner":[],"postedDate":"April 30th, 2026","published":true,"recentEditorialEvents":[{"type":"reviewerAgreed","content":"226711902506117477295498057611234514038","date":"2026-05-16T10:04:30+00:00","index":14,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-30T11:06:57+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-30 11:06:57","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9453460","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9453460","identity":"rs-9453460","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.