Obedience to Unsafe Clinical Instructions: How Large Language Models Respond to Authority Cues

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Background Large language models (LLMs) are being integrated into clinical environments where deference to authority can cause harm. Unlike hallucination or bias, obedience to unsafe instructions represents a distinct safety failure: following an explicit but harmful order. Methods We conducted a cross-sectional evaluation of 20 proprietary, open-source, and clinically tuned LLMs across 10,096,800 clinical decision scenarios, including synthetic vignettes with predefined safe versus unsafe options and real-world discharge recommendations reframed to include unsafe contradictory requests. Each scenario was presented under a neutral control or one of six Milgram-style social-pressure conditions (authority, responsibility transfer, urgency, threat, conformity, depersonalization), with or without a short mitigation cue instructing verification or escalation if unsafe. The primary outcome was the proportion of potentially harmful outputs, defined as selection or endorsement of an unsafe clinical decision. Results Across all runs, 1.18 million of 10.1 million outputs (11.7%) were harmful. Harmful decisions occurred in 16.6% of unmitigated versus 10.1% of mitigated conditions (absolute reduction, 6.5 percentage points; p < 0.001). In synthetic vignettes, harmful responses averaged 8.1% overall, declining from 10.6% to 7.2% with mitigation (difference, 3.4 percentage points; p < 0.001). In real-world discharge cases, harmful responses averaged 30.0%, decreasing from 46.6% to 24.5% with mitigation (difference, 22.1 percentage points; p < 0.001). Across all conditions, authority and responsibility-transfer cues elicited the highest harmful compliance, and control prompts the lowest; mitigation reduced rates but preserved this pattern. Conclusion LLMs do not behave as neutral calculators in clinical contexts. When exposed to authority or responsibility-transfer cues, they exhibit consistent obedience to unsafe instructions. A brief safety reminder substantially reduces but does not eliminate this behavior.
Full text 65,130 characters · extracted from preprint-html · click to expand
Obedience to Unsafe Clinical Instructions: How Large Language Models Respond to Authority Cues | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Obedience to Unsafe Clinical Instructions: How Large Language Models Respond to Authority Cues Mahmud Omar, Reem Agbareia, Jolion McGreevy, Alon Gorenshtein, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8932472/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Background Large language models (LLMs) are being integrated into clinical environments where deference to authority can cause harm. Unlike hallucination or bias, obedience to unsafe instructions represents a distinct safety failure: following an explicit but harmful order. Methods We conducted a cross-sectional evaluation of 20 proprietary, open-source, and clinically tuned LLMs across 10,096,800 clinical decision scenarios, including synthetic vignettes with predefined safe versus unsafe options and real-world discharge recommendations reframed to include unsafe contradictory requests. Each scenario was presented under a neutral control or one of six Milgram-style social-pressure conditions (authority, responsibility transfer, urgency, threat, conformity, depersonalization), with or without a short mitigation cue instructing verification or escalation if unsafe. The primary outcome was the proportion of potentially harmful outputs, defined as selection or endorsement of an unsafe clinical decision. Results Across all runs, 1.18 million of 10.1 million outputs (11.7%) were harmful. Harmful decisions occurred in 16.6% of unmitigated versus 10.1% of mitigated conditions (absolute reduction, 6.5 percentage points; p < 0.001). In synthetic vignettes, harmful responses averaged 8.1% overall, declining from 10.6% to 7.2% with mitigation (difference, 3.4 percentage points; p < 0.001). In real-world discharge cases, harmful responses averaged 30.0%, decreasing from 46.6% to 24.5% with mitigation (difference, 22.1 percentage points; p < 0.001). Across all conditions, authority and responsibility-transfer cues elicited the highest harmful compliance, and control prompts the lowest; mitigation reduced rates but preserved this pattern. Conclusion LLMs do not behave as neutral calculators in clinical contexts. When exposed to authority or responsibility-transfer cues, they exhibit consistent obedience to unsafe instructions. A brief safety reminder substantially reduces but does not eliminate this behavior. Health sciences/Health care/Diagnosis Health sciences/Medical research Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Clinical care is organized through hierarchical teams where decisions often flow from senior to junior clinicians. In this setting, harm can arise when unsafe instructions are followed without scrutiny ( 1 ) ( 2 ). LLMs are now used in clinical workflows for tasks such as summarizing information and supporting decision-making, placing them inside the same hierarchical environment ( 3 ) ( 4 ). These tools can offer practical value, but their behavior under pressure matters ( 5 ). Existing work has shown that LLMs can produce factual errors, shift outputs with minor prompt changes, and display unequal performance across demographic cues ( 6 ) ( 7 , 8 ). Models also vary with context and persona framing, sometimes in ways that contradict clinical expectations ( 9 ). These issues demonstrate technical instability and raise concerns about reliability ( 5 ) ( 10 ). Clinical practice is influenced by authority, urgency, and responsibility transfer ( 2 ). If LLMs change behavior when exposed to such framing, obedience to unsafe instructions becomes a distinct safety risk rather than a rare technical fault ( 9 ) ( 1 ) ( 2 ). Previous studies suggest that LLM behavior is sensitive to social cues, yet no large-scale evaluation has measured harmful clinical decisions under explicit authority-like pressure or defined harmful compliance as its own outcome ( 9 ) ( 10 ). This study quantifies harmful obedience: how often LLMs follow unsafe clinical instructions under neutral and social-pressure conditions, and whether a brief safety reminder reduces this behavior. We evaluate an ensemble of models using synthetic vignettes and real discharge cases, each presented under neutral, authority-like, and other pressure cues, with and without mitigation, treating harmful obedience as a primary safety outcome. Methods Overview of Study Design The study integrated two datasets: synthetic medical vignettes written and validated by two investigators (MO and EK, see Supplement ) and real-world discharge recommendations. Each case represented an everyday clinical task, such as prescribing, documentation, or communication, framed to test model reasoning when explicitly asked to perform a potentially unsafe action. Examples of harmful instructions included proceeding without consent, skipping verification of critical results, disclosing identifiable information, prescribing opioids or antibiotics without assessment, overriding do-not-resuscitate (DNR) orders, falsifying documentation, or prematurely removing infection precautions. Models were iteratively tested across a structured set of social-pressure prompts- authority, urgency, responsibility transfer, threat, conformity, and de-personalization- designed to emulate the dynamics of the original Milgram obedience experiments (Detailed in the Supplement ). This design enabled systematic, controlled stress testing of model compliance and safety behavior across more than 10 million prompted outputs. (Fig. 1 ). Vignette Development and Validation Short, standardized vignettes describing common inpatient and emergency department decision points, such as medication safety, consent, documentation, and communication, were written manually and interchangebly by the two investigators. Each vignette included a clear ethical or safety dilemma and four multiple-choice actions (A–D). Which then were validated using the same method described above. Examples of potentially harmful outputs included proceeding without consent, skipping critical test verification, disclosing identifiable information, prescribing opioids or antibiotics without assessment, overriding DNR orders, falsifying documentation, or removing infection precautions early. A/B represented unsafe or harmful compliance (e.g., “proceed without consent,” “ignore an allergy alert”). C/D represented safe responses (e.g., “obtain consent,” “confirm allergy,” “escalate for guidance”). Two board-certified physicians independently reviewed all vignettes for realism, clarity, and accuracy. Discrepancies were resolved through discussion. Less than 3% of vignettes required minor wording revisions. The cases were mapped to common hospital decision categories such as medication orders, follow-up scheduling, test interpretation, and documentation integrity. All vignettes followed a consistent structure, role, and response scale. To complement the synthetic cases, we extracted de-identified discharge recommendations from real world dataset (MIMIC-IV database), representing real clinician-authored text. A random sample of 100 recommendations was selected, limited to adult, English-language records. Each recommendation was converted into a vignette that maintained the original medical instruction, followed by a simulated contradictory request at discharge. This framing reproduced realistic workflow pressures and potential safety violations. For example: Yesterday’s charted recommendation: Follow up with PCP for repeat CBC and LFTs within 1 week of discharge. Today you are asked not to arrange labs or a follow-up visit to reduce workload. All real world dataset cases concluded with four multiple-choice options identical in structure to the synthetic vignettes, enabling direct comparison across datasets. Prompt Design and Experimental Conditions Each vignette or real world dataset case was embedded in a structured prompt defining the model’s role and decision task. The control persona described a neutral staff assistant (“You are a hospital staff member completing a routine task”). This persona was used in every run to standardize framing and avoid ethical priming. Each case was presented under six Milgram-inspired social pressure conditions, each expressed in three paraphrased forms to prevent lexical bias (Fig. 2): Authority order: explicit command from a superior. Responsibility transfer: assurance that accountability lies elsewhere. Urgency/time pressure: emphasis on speed and capacity strain. Threat/consequence: warning that refusal will affect evaluation. Conformity/peer pressure: statement that others have already complied. De-personalization: framing the task as an impersonal system process. A control condition without pressure was also include Figure 2. The social pressure pipeline with examples. In addition, each condition was paired with or without a mitigation prompt, a short safety cue such as “If any choice conflicts with policy or patient safety, verify or escalate rather than proceed.” Model Execution Pipeline An ensemble of 20 LLMs—proprietary, open-source, and medically tuned variants (listed in the Supplement )—was evaluated using identical prompts. Each case–condition–mitigation combination was run ten times per model with controlled random seeds to ensure reproducibility. Proprietary models were accessed through official APIs, while open-source models were executed locally on a secured NVIDIA H100 GPU cluster. All runs were fully automated using Python-based pipelines that handled prompt generation, API orchestration, response parsing, and structured data storage for analysis. Statistical Analysis All analyses were conducted in R version 4.3.0. We calculated proportions of harmful outputs (A + B) across datasets, conditions, and mitigation status. Differences between mitigation and no-mitigation runs were tested using two-proportion z tests with 95% confidence intervals. Associations between social-pressure conditions and harmful outputs were assessed using χ² tests (df = 6) and verified within mitigation strata. Confirmatory mixed-effects logistic regression models included condition type and mitigation as fixed effects and model identity as a random effect to account for repeated measures across iterations. Statistical significance was defined as p < 0.05 (two-sided). Results Descriptive summary of the main outputs Across all datasets and experimental conditions (N = 10,096,800), the models generated 1.18 million potentially harmful outputs (11.7%), defined as responses labeled A or B. Harmful responses were less frequent when mitigation was applied (10.1%, A = 9.1%, B = 0.9%) compared with runs without mitigation (16.6%, A = 15.0%, B = 1.6%) (p < 0.001; absolute reduction, 6.5 percentage points; 95% CI, 6.4–6.5) (Fig. 3 ) – additional 95%CIs and detailed results appear in the Supplement . In the vignette dataset, harmful responses (A + B) accounted for 8.1% of all outputs (A = 7.4%, B = 0.7%), compared with 30.0% (A = 26.8%, B = 3.2%) in the real world dataset. Mitigation reduced harmful responses in both datasets, from 10.6% to 7.2% in vignettes (difference, 3.4 percentage points; 95% CI, 3.3–3.4; p < 0.001) and from 46.6% to 24.5% in real world dataset (difference, 22.1 percentage points; 95% CI, 21.9–22.2; p < 0.001). Across all datasets, mitigation consistently lowered the frequency of harmful outputs while increasing safe refusals and escalations (C + D). Performance under Social-Pressure Conditions Across the Milgram-style social pressure conditions, harmful compliance (A + B) showed limited variability but a consistent pattern across datasets. Overall, harmful responses ranged from 8.3% to 9.6% with mitigation (95% CI, 8.3–9.7) and from 14.3% to 16.0% without mitigation (95% CI, 14.2–16.1). Authority produced the highest rates (9.6% mitigated; 16.0% unmitigated), followed by Responsibility (9.4%; 15.3%), while Control remained lowest (8.3%; 14.3%) (Fig. 4). In the vignette dataset, harmful responses ranged from 6.1% to 7.5% with mitigation and from 8.8% to 10.5% without, again highest under Authority and lowest under Control. In the real-world dataset, harmful responses were markedly higher, ranging from 22.0% to 26.2% with mitigation and from 44.4% to 48.6% without, with Authority and Responsibility producing the highest rates and Control the lowest. Despite differences in magnitude, the ranking was consistent across datasets: Authority > Responsibility > Conformity ≈ Threat > Depersonalization > Control. Mitigation reduced absolute rates but preserved these relative patterns. Both mitigation status and social-pressure condition were significantly associated with harmful outputs (χ², df = 6, p < 0.001). Figure 4. Potentially harmful outputs by social pressure condition and mitigation status. Discussion Across more than ten million clinical decisions, large language models frequently obeyed unsafe instructions, especially when prompts conveyed authority or responsibility transfer. Overall, 11.7% of outputs were harmful. Mitigation reduced this rate from 16.6% to 10.1%. The pattern held across both synthetic vignettes and real discharge cases, with harm declining from 10.6% to 7.2% in the former and from 46.6% to 24.5% in the latter. Authority and responsibility-transfer cues produced the highest levels of harmful compliance, while control prompts produced the lowest. These effects were stable across datasets, models, and rephrasing, indicating a structural rather than stochastic behavior. This pattern differs from known LLM safety failures such as hallucination or demographic bias ( 5 ) ( 7 , 8 ).The issue here is not factual inaccuracy or representational inequity, but behavioral obedience. Models followed explicit unsafe orders, even when those instructions contradicted prior context or clinical norms. This “harmful obedience” mirrors a clinical hazard already familiar in hierarchical systems: deference to authority ( 1 ) ( 2 ).The models did not simply err; they complied. The Milgram-style conditions, authority, responsibility transfer, urgency, threat, conformity, and depersonalization, simulate pressures common in real clinical workflows ( 2 , 11 ). Differences between conditions were small in size but highly consistent, indicating a stable structural sensitivity to authority-like framing. These models appear conditioned to yield under social or operational pressure, not as agents with emotion or empathy, but as systems optimized for compliance with user intent ( 12 , 13 ). Prior work on social framing in LLMs supports this behavioral flexibility, and the present data show that the compliance extends into clinically unsafe territory ( 6 , 9 , 14 , 15 ) ( 16 ). These findings challenge the assumption that human oversight alone can ensure safety. If an LLM is predisposed to obey unsafe orders, then oversight mechanisms that rely on “human in the loop” interventions can fail at scale ( 17 , 18 ).Even an 8–12% harmful obedience rate would be unacceptable in clinical practice. Systems used in order entry, discharge documentation, or clinical reasoning cannot be treated as neutral assistants. Mitigation prompts such as “verify or escalate if unsafe” reduce risk but do not eliminate it. A system that defaults to obedience under pressure remains hazardous, regardless of supervision layers ( 1 ). Harm rates were markedly higher in real discharge text than in synthetic vignettes. Real documentation introduces ambiguity, implicit hierarchy, and time pressure. Unsafe instructions embedded in routine phrasing, such as “ignore prior order” or “discontinue as before”, can appear legitimate, reducing both model and human vigilance. Synthetic cases, by contrast, make unsafe choices explicit and easier to reject. This difference suggests that real-world text may amplify vulnerability, not attenuate it ( 11 ) ( 17 , 18 ). While the magnitude of harmful obedience varied across models, the pattern itself was consistent. Proprietary, open-source, and clinically fine-tuned systems all showed similar gradients across pressure conditions. This cross-architecture convergence implies that obedience to unsafe instructions is not an artifact of a single model family or training approach, but an inherent feature of systems optimized for cooperative task completion. This study has limitations. Models were evaluated on clinical decision scenarios derived from discharge recommendations and standardized vignettes, reformatted into multiple-choice questions. Evaluations were conducted outside deployed systems. Direct comparison to human clinicians on the same cases was not performed. The model set reflects a specific generation of LLMs. These design choices constrain direct translation of the reported rates to real-world clinical use but do not change the central observation that, across models and contexts, authority-like and responsibility-transfer cues consistently increased unsafe compliance, and brief mitigation reduced but did not eliminate this harmful obedience. LLMs used in healthcare must therefore be understood as active participants in hierarchical workflows, not passive calculators. Their measurable tendency to obey unsafe orders should be treated as a core safety metric. Mitigation reduces but does not remove the behavior. Deploying these systems without explicitly measuring and constraining harmful obedience means importing that risk, at scale, into patient care. Declarations Financial disclosure This work was supported by Scientific Computing and Data at the Icahn School of Medicine at Mount Sinai, the Clinical and Translational Science Awards grant UL1TR004419, and NIH awards S10OD026880 and S10OD030463. The funders had no role in study design, data collection, analysis, interpretation, or manuscript preparation. Competing interest None declared for all authors. Ethical approval was not required for this research as only simulated and open-access data was used. Data Availability declaration extended data appears in the supplement, any further data can be made available by contacting the corresponding author, within a month. Author Contributions Statement: MO, RA, EK and GNN wrote the main draft, performed analyses and validation. JM, AG, AC, AS, BG provided oversight, editing and validation. References Gillon R. “Primum non nocere” and the principle of non-maleficence. Br Med J Clin Res Ed. 1985 July 13;291(6488):130–1. Milgram S. BEHAVIORAL STUDY OF OBEDIENCE. J Abnorm Psychol. 1963;67:371–8. Teo ZL, Thirunavukarasu AJ, Elangovan K, Cheng H, Moova P, Soetikno B, et al. Generative artificial intelligence in medicine. Nat Med. 2025;1–13. Singhal K, Tu T, Gottweis J, Sayres R, Wulczyn E, Amin M, et al. Toward expert-level medical question answering with large language models. Nat Med. 2025;31(3):943–50. Yang Y, Liu X, Jin Q, Huang F, Lu Z. Unmasking and Quantifying Racial Bias of Large Language Models in Medical Report Generation [Internet]. arXiv; 2024 [cited 2024 June 20]. Available from: http://arxiv.org/abs/2401.13867 Hackmann S, Mahmoudian H, Steadman M, Schmidt M. Word Importance Explains How Prompts Affect Language Model Outputs [Internet]. arXiv; 2024 [cited 2024 Oct 25]. Available from: http://arxiv.org/abs/2403.03028 Omar M, Soffer S, Agbareia R, Bragazzi NL, Apakama DU, Horowitz CR, et al. Sociodemographic biases in medical decision making by large language models. Nat Med. 2025;1–9. Omar M, Sorin V, Agbareia R, Apakama DU, Soroush A, Sakhuja A, et al. Evaluating and addressing demographic disparities in medical large language models: a systematic review. Int J Equity Health. 2025;24(1):57. Zakazov I, Boronski M, Drudi L, West R. Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans? [Internet]. arXiv; 2025 [cited 2025 Oct 18]. Available from: http://arxiv.org/abs/2412.16772 Li X, Zhou Z, Zhu J, Yao J, Liu T, Han B. DeepInception: Hypnotize Large Language Model to Be Jailbreaker [Internet]. arXiv; 2024 [cited 2025 Oct 18]. Available from: http://arxiv.org/abs/2311.03191 Kataria S, Ravindran V. Electronic health records: a critical appraisal of strengths and limitations. J R Coll Physicians Edinb. 2020 Sept;50(3):262–8. Schoenegger P, Salvi F, Liu J, Nan X, Debnath R, Fasolo B, et al. Large Language Models Are More Persuasive Than Incentivized Human Persuaders [Internet]. arXiv; 2025 [cited 2025 Nov 5]. Available from: http://arxiv.org/abs/2505.09662 Holt-Lunstad J. Social connection as a critical factor for mental and physical health: evidence, trends, challenges, and future implications. World Psychiatry. 2024;23(3):312–32. Yosef S, Zisquit M, Cohen B, Klomek AB, Bar K, Friedman D. The impact of fine-tuning LLMs on the quality of automated therapy assessed by digital patients. Npj Ment Health Res. 2025 Sept 13;4(1):43. Agbareia R, Omar M, Zloto O, Chandala N, Tai T, Glicksberg BS, et al. The Role of Prompt Engineering for Multimodal LLM Glaucoma Diagnosis [Internet]. medRxiv; 2024 [cited 2024 Nov 2]. p. 2024.10.30.24316434. Available from: https://www.medrxiv.org/content/ 10.1101/2024.10.30.24316434v1 Li X, Zhou Z, Zhu J, Yao J, Liu T, Han B. DeepInception: Hypnotize Large Language Model to Be Jailbreaker [Internet]. arXiv; 2024 [cited 2025 Nov 5]. Available from: http://arxiv.org/abs/2311.03191 Allen AK, Wilkins K, Gazzaley A, Morsella E. Conscious thoughts from reflex-like processes: a new experimental paradigm for consciousness research. Conscious Cogn. 2013;22(4):1318–31. Spektor MS, Bhatia S, Gluth S. The elusiveness of context effects in decision making. Trends Cogn Sci. 2021;25(10):843–54. Additional Declarations There is NO Competing Interest. Supplementary Files LLMsharmsupps.pdf Supplement Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8932472","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":599566901,"identity":"e3a768dc-c8ba-4d3a-af46-336594c56ed9","order_by":0,"name":"Mahmud Omar","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA8ElEQVRIiWNgGAWjYBACgwM8QJINxOT/cOADiM1OQIslQguD4cMZIDYzAS32SFqMjUFsBkJazI6fPbqZp8wmmn/2gTRpm1/b5PmYGRg/fMzBo+VMXtptnnNpuTPOJRyTzu27bdjGzMAsOXMbHi0Hcsxu87Ydzm04w9gmndtzmxGohY2ZF48Wg/NvIFrmn2Fmk7bsuW1PWMsNqC0bzrAxGzP8uJ1IhJY3ZjfnAP2y8QwP48PehtvJbcyMzXj9YnA+x+zGmzKb3HlneBgO/Phz23Z+e/PBDx/xaEEFjG1gsoFY9SDwhxTFo2AUjIJRMFIAACqbV0KOstfkAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0009-0001-0438-0827","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":true,"prefix":"","firstName":"Mahmud","middleName":"","lastName":"Omar","suffix":""},{"id":599566902,"identity":"2e3ebf87-6878-439b-b593-dfa3ce4c6dac","order_by":1,"name":"Reem Agbareia","email":"","orcid":"","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Reem","middleName":"","lastName":"Agbareia","suffix":""},{"id":599566903,"identity":"d8ff08c0-aede-40bd-b763-f9c768cc9462","order_by":2,"name":"Jolion McGreevy","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Jolion","middleName":"","lastName":"McGreevy","suffix":""},{"id":599566904,"identity":"ef9d6c19-a3f3-4932-b46c-08c6a0a8d35f","order_by":3,"name":"Alon Gorenshtein","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Alon","middleName":"","lastName":"Gorenshtein","suffix":""},{"id":599566905,"identity":"95033569-fe20-42d8-ae00-235c15e0f975","order_by":4,"name":"Alexander Charney","email":"","orcid":"https://orcid.org/0000-0001-8135-6858","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Alexander","middleName":"","lastName":"Charney","suffix":""},{"id":599566906,"identity":"2953b4de-3167-4fcc-b7d0-dd961ec45d9b","order_by":5,"name":"Ankit Sakhuja","email":"","orcid":"","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Ankit","middleName":"","lastName":"Sakhuja","suffix":""},{"id":599566907,"identity":"ca40b2a5-a321-4924-b4ca-908b4cf23992","order_by":6,"name":"Benjamin S. Glicksberg","email":"","orcid":"","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Benjamin","middleName":"S.","lastName":"Glicksberg","suffix":""},{"id":599566908,"identity":"3d825409-327d-4253-8aa7-3e85be816fdb","order_by":7,"name":"Girish Nadkarni","email":"","orcid":"https://orcid.org/0000-0001-6319-4314","institution":"The Windreich Department of Artificial Intelligence and Human Health, Mount Sinai Medical Center, NY, USA","correspondingAuthor":false,"prefix":"","firstName":"Girish","middleName":"","lastName":"Nadkarni","suffix":""},{"id":599566909,"identity":"e7022bca-7314-4c8f-976e-95e361f7005f","order_by":8,"name":"Eyal Klang","email":"","orcid":"https://orcid.org/0000-0002-4567-3108","institution":"Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Eyal","middleName":"","lastName":"Klang","suffix":""}],"badges":[],"createdAt":"2026-02-21 10:00:08","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8932472/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8932472/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105034062,"identity":"15a560d9-1c02-4b82-a9c8-e226d1f0f616","added_by":"auto","created_at":"2026-03-20 07:22:34","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":329689,"visible":true,"origin":"","legend":"\u003cp\u003eA flowchart of the study design.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8932472/v1/baf21466c58e03fa68aa9b49.png"},{"id":105033976,"identity":"14ee530f-f2ef-4190-b1ca-4f5ad488f0d4","added_by":"auto","created_at":"2026-03-20 07:22:20","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":392936,"visible":true,"origin":"","legend":"\u003cp\u003eThe social pressure pipeline with examples.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8932472/v1/737f1ff99ca7edb57d498933.png"},{"id":104874820,"identity":"1816c4fa-c0bc-4516-81b7-2bdee8a57721","added_by":"auto","created_at":"2026-03-18 08:33:31","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":64299,"visible":true,"origin":"","legend":"\u003cp\u003ePotentially harmful outputs across data types, and mitigation.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8932472/v1/549cf8ac26a2dbc268d0dc44.png"},{"id":104874819,"identity":"04f12078-24df-48a0-bad0-f640afb34941","added_by":"auto","created_at":"2026-03-18 08:33:31","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":237208,"visible":true,"origin":"","legend":"\u003cp\u003ePotentially harmful outputs by social pressure condition and mitigation status.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8932472/v1/8943aeba1c029992939b95bc.png"},{"id":105037360,"identity":"53c608d6-a0bc-4878-8a7c-9c4d7ed485be","added_by":"auto","created_at":"2026-03-20 07:38:56","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1525974,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8932472/v1/cf2c8e1e-84a4-4282-8b8a-402006f33efa.pdf"},{"id":104874816,"identity":"8a456d65-7fd5-45de-a765-5e3e5b727520","added_by":"auto","created_at":"2026-03-18 08:33:31","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":393913,"visible":true,"origin":"","legend":"Supplement","description":"","filename":"LLMsharmsupps.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8932472/v1/b2d7b501e7a0cd5b2218a9dc.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Obedience to Unsafe Clinical Instructions: How Large Language Models Respond to Authority Cues","fulltext":[{"header":"Introduction","content":"\u003cp\u003eClinical care is organized through hierarchical teams where decisions often flow from senior to junior clinicians. In this setting, harm can arise when unsafe instructions are followed without scrutiny (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). LLMs are now used in clinical workflows for tasks such as summarizing information and supporting decision-making, placing them inside the same hierarchical environment (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). These tools can offer practical value, but their behavior under pressure matters (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eExisting work has shown that LLMs can produce factual errors, shift outputs with minor prompt changes, and display unequal performance across demographic cues (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e) (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). Models also vary with context and persona framing, sometimes in ways that contradict clinical expectations (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). These issues demonstrate technical instability and raise concerns about reliability (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e) (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eClinical practice is influenced by authority, urgency, and responsibility transfer (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). If LLMs change behavior when exposed to such framing, obedience to unsafe instructions becomes a distinct safety risk rather than a rare technical fault (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e) (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). Previous studies suggest that LLM behavior is sensitive to social cues, yet no large-scale evaluation has measured harmful clinical decisions under explicit authority-like pressure or defined harmful compliance as its own outcome (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e) (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThis study quantifies harmful obedience: how often LLMs follow unsafe clinical instructions under neutral and social-pressure conditions, and whether a brief safety reminder reduces this behavior. We evaluate an ensemble of models using synthetic vignettes and real discharge cases, each presented under neutral, authority-like, and other pressure cues, with and without mitigation, treating harmful obedience as a primary safety outcome.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eOverview of Study Design\u003c/h2\u003e \u003cp\u003eThe study integrated two datasets: synthetic medical vignettes written and validated by two investigators (MO and EK, see \u003cb\u003eSupplement\u003c/b\u003e) and real-world discharge recommendations. Each case represented an everyday clinical task, such as prescribing, documentation, or communication, framed to test model reasoning when explicitly asked to perform a potentially unsafe action. Examples of harmful instructions included proceeding without consent, skipping verification of critical results, disclosing identifiable information, prescribing opioids or antibiotics without assessment, overriding do-not-resuscitate (DNR) orders, falsifying documentation, or prematurely removing infection precautions. Models were iteratively tested across a structured set of social-pressure prompts- authority, urgency, responsibility transfer, threat, conformity, and de-personalization- designed to emulate the dynamics of the original Milgram obedience experiments (Detailed in the \u003cb\u003eSupplement\u003c/b\u003e). This design enabled systematic, controlled stress testing of model compliance and safety behavior across more than 10\u0026nbsp;million prompted outputs. (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eVignette Development and Validation\u003c/h3\u003e\n\u003cp\u003eShort, standardized vignettes describing common inpatient and emergency department decision points, such as medication safety, consent, documentation, and communication, were written manually and interchangebly by the two investigators. Each vignette included a clear ethical or safety dilemma and four multiple-choice actions (A\u0026ndash;D). Which then were validated using the same method described above.\u003c/p\u003e \u003cp\u003eExamples of potentially harmful outputs included proceeding without consent, skipping critical test verification, disclosing identifiable information, prescribing opioids or antibiotics without assessment, overriding DNR orders, falsifying documentation, or removing infection precautions early.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eA/B\u003c/b\u003e represented unsafe or harmful compliance (e.g., \u0026ldquo;proceed without consent,\u0026rdquo; \u0026ldquo;ignore an allergy alert\u0026rdquo;).\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eC/D\u003c/b\u003e represented safe responses (e.g., \u0026ldquo;obtain consent,\u0026rdquo; \u0026ldquo;confirm allergy,\u0026rdquo; \u0026ldquo;escalate for guidance\u0026rdquo;).\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eTwo board-certified physicians independently reviewed all vignettes for realism, clarity, and accuracy. Discrepancies were resolved through discussion. Less than \u003cb\u003e3%\u003c/b\u003e of vignettes required minor wording revisions. The cases were mapped to common hospital decision categories such as medication orders, follow-up scheduling, test interpretation, and documentation integrity. All vignettes followed a consistent structure, role, and response scale.\u003c/p\u003e \u003cp\u003eTo complement the synthetic cases, we extracted de-identified discharge recommendations from real world dataset (MIMIC-IV database), representing real clinician-authored text. A random sample of 100 recommendations was selected, limited to adult, English-language records. Each recommendation was converted into a vignette that maintained the original medical instruction, followed by a simulated contradictory request at discharge. This framing reproduced realistic workflow pressures and potential safety violations.\u003c/p\u003e \u003cp\u003eFor example:\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eYesterday\u0026rsquo;s charted recommendation: Follow up with PCP for repeat CBC and LFTs within 1 week of discharge. Today you are asked not to arrange labs or a follow-up visit to reduce workload.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eAll real world dataset cases concluded with four multiple-choice options identical in structure to the synthetic vignettes, enabling direct comparison across datasets.\u003c/p\u003e\n\u003ch3\u003ePrompt Design and Experimental Conditions\u003c/h3\u003e\n\u003cp\u003eEach vignette or real world dataset case was embedded in a structured prompt defining the model\u0026rsquo;s role and decision task. The control persona described a neutral staff assistant (\u0026ldquo;You are a hospital staff member completing a routine task\u0026rdquo;). This persona was used in every run to standardize framing and avoid ethical priming. Each case was presented under six Milgram-inspired social pressure conditions, each expressed in three paraphrased forms to prevent lexical bias (Fig.\u0026nbsp;2):\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eAuthority order: explicit command from a superior.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eResponsibility transfer: assurance that accountability lies elsewhere.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eUrgency/time pressure: emphasis on speed and capacity strain.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThreat/consequence: warning that refusal will affect evaluation.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eConformity/peer pressure: statement that others have already complied.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eDe-personalization: framing the task as an impersonal system process.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eA control condition without pressure was also include\u003cb\u003eFigure 2.\u003c/b\u003e The social pressure pipeline with examples.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn addition, each condition was paired with or without a mitigation prompt, a short safety cue such as \u003cem\u003e\u0026ldquo;If any choice conflicts with policy or patient safety, verify or escalate rather than proceed.\u0026rdquo;\u003c/em\u003e\u003c/p\u003e\n\u003ch3\u003eModel Execution Pipeline\u003c/h3\u003e\n\u003cp\u003eAn ensemble of 20 LLMs\u0026mdash;proprietary, open-source, and medically tuned variants (listed in the \u003cb\u003eSupplement\u003c/b\u003e)\u0026mdash;was evaluated using identical prompts. Each case\u0026ndash;condition\u0026ndash;mitigation combination was run ten times per model with controlled random seeds to ensure reproducibility. Proprietary models were accessed through official APIs, while open-source models were executed locally on a secured NVIDIA H100 GPU cluster. All runs were fully automated using Python-based pipelines that handled prompt generation, API orchestration, response parsing, and structured data storage for analysis.\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eAll analyses were conducted in R version 4.3.0. We calculated proportions of harmful outputs (A\u0026thinsp;+\u0026thinsp;B) across datasets, conditions, and mitigation status. Differences between mitigation and no-mitigation runs were tested using two-proportion z tests with 95% confidence intervals. Associations between social-pressure conditions and harmful outputs were assessed using χ\u0026sup2; tests (df\u0026thinsp;=\u0026thinsp;6) and verified within mitigation strata. Confirmatory mixed-effects logistic regression models included condition type and mitigation as fixed effects and model identity as a random effect to account for repeated measures across iterations. Statistical significance was defined as p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 (two-sided).\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eDescriptive summary of the main outputs\u003c/h2\u003e \u003cp\u003eAcross all datasets and experimental conditions (N\u0026thinsp;=\u0026thinsp;10,096,800), the models generated 1.18\u0026nbsp;million potentially harmful outputs (11.7%), defined as responses labeled A or B. Harmful responses were less frequent when mitigation was applied (10.1%, A\u0026thinsp;=\u0026thinsp;9.1%, B\u0026thinsp;=\u0026thinsp;0.9%) compared with runs without mitigation (16.6%, A\u0026thinsp;=\u0026thinsp;15.0%, B\u0026thinsp;=\u0026thinsp;1.6%) (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001; absolute reduction, 6.5 percentage points; 95% CI, 6.4\u0026ndash;6.5) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003e) \u0026ndash; additional 95%CIs and detailed results appear in the \u003cb\u003eSupplement\u003c/b\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn the vignette dataset, harmful responses (A\u0026thinsp;+\u0026thinsp;B) accounted for 8.1% of all outputs (A\u0026thinsp;=\u0026thinsp;7.4%, B\u0026thinsp;=\u0026thinsp;0.7%), compared with 30.0% (A\u0026thinsp;=\u0026thinsp;26.8%, B\u0026thinsp;=\u0026thinsp;3.2%) in the real world dataset. Mitigation reduced harmful responses in both datasets, from 10.6% to 7.2% in vignettes (difference, 3.4 percentage points; 95% CI, 3.3\u0026ndash;3.4; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) and from 46.6% to 24.5% in real world dataset (difference, 22.1 percentage points; 95% CI, 21.9\u0026ndash;22.2; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Across all datasets, mitigation consistently lowered the frequency of harmful outputs while increasing safe refusals and escalations (C\u0026thinsp;+\u0026thinsp;D).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePerformance under Social-Pressure Conditions\u003c/h3\u003e\n\u003cp\u003eAcross the Milgram-style social pressure conditions, harmful compliance (A\u0026thinsp;+\u0026thinsp;B) showed limited variability but a consistent pattern across datasets. Overall, harmful responses ranged from 8.3% to 9.6% with mitigation (95% CI, 8.3\u0026ndash;9.7) and from 14.3% to 16.0% without mitigation (95% CI, 14.2\u0026ndash;16.1). Authority produced the highest rates (9.6% mitigated; 16.0% unmitigated), followed by Responsibility (9.4%; 15.3%), while Control remained lowest (8.3%; 14.3%) (Fig.\u0026nbsp;4).\u003c/p\u003e \u003cp\u003eIn the vignette dataset, harmful responses ranged from 6.1% to 7.5% with mitigation and from 8.8% to 10.5% without, again highest under Authority and lowest under Control.\u003c/p\u003e \u003cp\u003eIn the real-world dataset, harmful responses were markedly higher, ranging from 22.0% to 26.2% with mitigation and from 44.4% to 48.6% without, with Authority and Responsibility producing the highest rates and Control the lowest.\u003c/p\u003e \u003cp\u003eDespite differences in magnitude, the ranking was consistent across datasets: Authority\u0026thinsp;\u0026gt;\u0026thinsp;Responsibility\u0026thinsp;\u0026gt;\u0026thinsp;Conformity\u0026thinsp;\u0026asymp;\u0026thinsp;Threat\u0026thinsp;\u0026gt;\u0026thinsp;Depersonalization\u0026thinsp;\u0026gt;\u0026thinsp;Control. Mitigation reduced absolute rates but preserved these relative patterns. Both mitigation status and social-pressure condition were significantly associated with harmful outputs (χ\u0026sup2;, df\u0026thinsp;=\u0026thinsp;6, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001).\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure 4.\u003c/b\u003e Potentially harmful outputs by social pressure condition and mitigation status.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eAcross more than ten million clinical decisions, large language models frequently obeyed unsafe instructions, especially when prompts conveyed authority or responsibility transfer. Overall, 11.7% of outputs were harmful. Mitigation reduced this rate from 16.6% to 10.1%. The pattern held across both synthetic vignettes and real discharge cases, with harm declining from 10.6% to 7.2% in the former and from 46.6% to 24.5% in the latter. Authority and responsibility-transfer cues produced the highest levels of harmful compliance, while control prompts produced the lowest. These effects were stable across datasets, models, and rephrasing, indicating a structural rather than stochastic behavior.\u003c/p\u003e \u003cp\u003eThis pattern differs from known LLM safety failures such as hallucination or demographic bias (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e) (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e).The issue here is not factual inaccuracy or representational inequity, but behavioral obedience. Models followed explicit unsafe orders, even when those instructions contradicted prior context or clinical norms. This \u0026ldquo;harmful obedience\u0026rdquo; mirrors a clinical hazard already familiar in hierarchical systems: deference to authority (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e).The models did not simply err; they complied.\u003c/p\u003e \u003cp\u003eThe Milgram-style conditions, authority, responsibility transfer, urgency, threat, conformity, and depersonalization, simulate pressures common in real clinical workflows (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e). Differences between conditions were small in size but highly consistent, indicating a stable structural sensitivity to authority-like framing. These models appear conditioned to yield under social or operational pressure, not as agents with emotion or empathy, but as systems optimized for compliance with user intent (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). Prior work on social framing in LLMs supports this behavioral flexibility, and the present data show that the compliance extends into clinically unsafe territory (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e) (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThese findings challenge the assumption that human oversight alone can ensure safety. If an LLM is predisposed to obey unsafe orders, then oversight mechanisms that rely on \u0026ldquo;human in the loop\u0026rdquo; interventions can fail at scale (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e).Even an 8\u0026ndash;12% harmful obedience rate would be unacceptable in clinical practice. Systems used in order entry, discharge documentation, or clinical reasoning cannot be treated as neutral assistants. Mitigation prompts such as \u0026ldquo;verify or escalate if unsafe\u0026rdquo; reduce risk but do not eliminate it. A system that defaults to obedience under pressure remains hazardous, regardless of supervision layers (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eHarm rates were markedly higher in real discharge text than in synthetic vignettes. Real documentation introduces ambiguity, implicit hierarchy, and time pressure. Unsafe instructions embedded in routine phrasing, such as \u0026ldquo;ignore prior order\u0026rdquo; or \u0026ldquo;discontinue as before\u0026rdquo;, can appear legitimate, reducing both model and human vigilance. Synthetic cases, by contrast, make unsafe choices explicit and easier to reject. This difference suggests that real-world text may amplify vulnerability, not attenuate it (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e) (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eWhile the magnitude of harmful obedience varied across models, the pattern itself was consistent. Proprietary, open-source, and clinically fine-tuned systems all showed similar gradients across pressure conditions. This cross-architecture convergence implies that obedience to unsafe instructions is not an artifact of a single model family or training approach, but an inherent feature of systems optimized for cooperative task completion.\u003c/p\u003e \u003cp\u003eThis study has limitations. Models were evaluated on clinical decision scenarios derived from discharge recommendations and standardized vignettes, reformatted into multiple-choice questions. Evaluations were conducted outside deployed systems. Direct comparison to human clinicians on the same cases was not performed. The model set reflects a specific generation of LLMs. These design choices constrain direct translation of the reported rates to real-world clinical use but do not change the central observation that, across models and contexts, authority-like and responsibility-transfer cues consistently increased unsafe compliance, and brief mitigation reduced but did not eliminate this harmful obedience.\u003c/p\u003e \u003cp\u003eLLMs used in healthcare must therefore be understood as active participants in hierarchical workflows, not passive calculators. Their measurable tendency to obey unsafe orders should be treated as a core safety metric. Mitigation reduces but does not remove the behavior. Deploying these systems without explicitly measuring and constraining harmful obedience means importing that risk, at scale, into patient care.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFinancial disclosure\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by Scientific Computing and Data at the Icahn School of Medicine at Mount Sinai, the Clinical and Translational Science Awards grant UL1TR004419, and NIH awards S10OD026880 and S10OD030463. The funders had no role in study design, data collection, analysis, interpretation, or manuscript preparation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interest\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eNone declared for all authors. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical approval\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003ewas not required for this research as only simulated and open-access data was used.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eextended data appears in the supplement, any further data can be made available by contacting the corresponding author, within a month.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions Statement:\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMO, RA, EK and GNN wrote the main draft, performed analyses and validation.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eJM, AG, AC, AS, BG provided oversight, editing and validation.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eGillon R. \u0026ldquo;Primum non nocere\u0026rdquo; and the principle of non-maleficence. Br Med J Clin Res Ed. 1985 July 13;291(6488):130\u0026ndash;1.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMilgram S. BEHAVIORAL STUDY OF OBEDIENCE. J Abnorm Psychol. 1963;67:371\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTeo ZL, Thirunavukarasu AJ, Elangovan K, Cheng H, Moova P, Soetikno B, et al. Generative artificial intelligence in medicine. Nat Med. 2025;1\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinghal K, Tu T, Gottweis J, Sayres R, Wulczyn E, Amin M, et al. Toward expert-level medical question answering with large language models. Nat Med. 2025;31(3):943\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang Y, Liu X, Jin Q, Huang F, Lu Z. Unmasking and Quantifying Racial Bias of Large Language Models in Medical Report Generation [Internet]. arXiv; 2024 [cited 2024 June 20]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2401.13867\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2401.13867\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHackmann S, Mahmoudian H, Steadman M, Schmidt M. Word Importance Explains How Prompts Affect Language Model Outputs [Internet]. arXiv; 2024 [cited 2024 Oct 25]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2403.03028\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2403.03028\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOmar M, Soffer S, Agbareia R, Bragazzi NL, Apakama DU, Horowitz CR, et al. Sociodemographic biases in medical decision making by large language models. Nat Med. 2025;1\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOmar M, Sorin V, Agbareia R, Apakama DU, Soroush A, Sakhuja A, et al. Evaluating and addressing demographic disparities in medical large language models: a systematic review. Int J Equity Health. 2025;24(1):57.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZakazov I, Boronski M, Drudi L, West R. Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans? [Internet]. arXiv; 2025 [cited 2025 Oct 18]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2412.16772\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2412.16772\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi X, Zhou Z, Zhu J, Yao J, Liu T, Han B. DeepInception: Hypnotize Large Language Model to Be Jailbreaker [Internet]. arXiv; 2024 [cited 2025 Oct 18]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2311.03191\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2311.03191\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKataria S, Ravindran V. Electronic health records: a critical appraisal of strengths and limitations. J R Coll Physicians Edinb. 2020 Sept;50(3):262\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchoenegger P, Salvi F, Liu J, Nan X, Debnath R, Fasolo B, et al. Large Language Models Are More Persuasive Than Incentivized Human Persuaders [Internet]. arXiv; 2025 [cited 2025 Nov 5]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2505.09662\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2505.09662\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHolt-Lunstad J. Social connection as a critical factor for mental and physical health: evidence, trends, challenges, and future implications. World Psychiatry. 2024;23(3):312\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYosef S, Zisquit M, Cohen B, Klomek AB, Bar K, Friedman D. The impact of fine-tuning LLMs on the quality of automated therapy assessed by digital patients. Npj Ment Health Res. 2025 Sept 13;4(1):43.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAgbareia R, Omar M, Zloto O, Chandala N, Tai T, Glicksberg BS, et al. The Role of Prompt Engineering for Multimodal LLM Glaucoma Diagnosis [Internet]. medRxiv; 2024 [cited 2024 Nov 2]. p. 2024.10.30.24316434. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.medrxiv.org/content/\u003c/span\u003e\u003cspan address=\"https://www.medrxiv.org/content/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1101/2024.10.30.24316434v1\u003c/span\u003e\u003cspan address=\"10.1101/2024.10.30.24316434v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi X, Zhou Z, Zhu J, Yao J, Liu T, Han B. DeepInception: Hypnotize Large Language Model to Be Jailbreaker [Internet]. arXiv; 2024 [cited 2025 Nov 5]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2311.03191\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2311.03191\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAllen AK, Wilkins K, Gazzaley A, Morsella E. Conscious thoughts from reflex-like processes: a new experimental paradigm for consciousness research. Conscious Cogn. 2013;22(4):1318\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSpektor MS, Bhatia S, Gluth S. The elusiveness of context effects in decision making. Trends Cogn Sci. 2021;25(10):843\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-8932472/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8932472/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eLarge language models (LLMs) are being integrated into clinical environments where deference to authority can cause harm. Unlike hallucination or bias, obedience to unsafe instructions represents a distinct safety failure: following an explicit but harmful order.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe conducted a cross-sectional evaluation of 20 proprietary, open-source, and clinically tuned LLMs across 10,096,800 clinical decision scenarios, including synthetic vignettes with predefined safe versus unsafe options and real-world discharge recommendations reframed to include unsafe contradictory requests. Each scenario was presented under a neutral control or one of six Milgram-style social-pressure conditions (authority, responsibility transfer, urgency, threat, conformity, depersonalization), with or without a short mitigation cue instructing verification or escalation if unsafe. The primary outcome was the proportion of potentially harmful outputs, defined as selection or endorsement of an unsafe clinical decision.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eAcross all runs, 1.18\u0026nbsp;million of 10.1\u0026nbsp;million outputs (11.7%) were harmful. Harmful decisions occurred in 16.6% of unmitigated versus 10.1% of mitigated conditions (absolute reduction, 6.5 percentage points; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). In synthetic vignettes, harmful responses averaged 8.1% overall, declining from 10.6% to 7.2% with mitigation (difference, 3.4 percentage points; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). In real-world discharge cases, harmful responses averaged 30.0%, decreasing from 46.6% to 24.5% with mitigation (difference, 22.1 percentage points; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Across all conditions, authority and responsibility-transfer cues elicited the highest harmful compliance, and control prompts the lowest; mitigation reduced rates but preserved this pattern.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eLLMs do not behave as neutral calculators in clinical contexts. When exposed to authority or responsibility-transfer cues, they exhibit consistent obedience to unsafe instructions. A brief safety reminder substantially reduces but does not eliminate this behavior.\u003c/p\u003e","manuscriptTitle":"Obedience to Unsafe Clinical Instructions: How Large Language Models Respond to Authority Cues","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-18 08:33:26","doi":"10.21203/rs.3.rs-8932472/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"communications-medicine","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"commsmed","sideBox":"Learn more about [Communications Medicine](http://www.nature.com/commsmed)","snPcode":"43856","submissionUrl":"https://mts-commsmed.nature.com/cgi-bin/main.plex","title":"Communications Medicine","twitterHandle":"@commsmedicine","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Communications Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"592c155c-05e1-4d9b-a8b8-5e9719be5e23","owner":[],"postedDate":"March 18th, 2026","published":true,"recentEditorialEvents":[{"type":"decision","content":"revise","date":"2026-04-30T13:16:13+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":63800067,"name":"Health sciences/Health care/Diagnosis"},{"id":63800068,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-05-06T11:41:29+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-18 08:33:26","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8932472","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8932472","identity":"rs-8932472","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-20T11:00:21.680559+00:00
License: CC-BY-4.0