Can We Trust the Machine? LLMs Mimic Human Expected Utility Theory Violations and Its Impact on Decision and Negotiation Systems

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Expected Utility Theory (EUT) has long served as a benchmark for rational decision-making, with well-documented human deviations in the form of framing effects, time inconsistency, and violations of independence and sequential rationality. In this study, we extend the EUT audit framework to large language models (LLMs), evaluating whether these increasingly embedded systems behave as rational decision-support agents.Using a four-part audit battery adapted from classic behavioral economics experiments, we tested four popular LLMs—GPT-4o, GPT-3.5, GPT-Mini, and DeepSeek, on 1,600 completions. Each model was evaluated on independence (via Allais-type lotteries), time consistency (via hyperbolic discounting tasks), framing invariance (via gain/loss presentations), and sequential rationality (via last-round cooperation in a Prisoner’s Dilemma).Across models, we observed strikingly human-like patterns of irrationality. Average violation rates were 36% for independence, 34% for time inconsistency, 33% for framing effects, and 32% for sequential rationality, closely mirroring human laboratory data. The consistency of these effects across architectures suggests that bounded rationality is not incidental, but an emergent feature of next-token prediction objectives.Our findings suggest that while LLMs may simulate rational discourse, their decision logic remains vulnerable to the same cognitive biases that affect humans. We propose safeguards including EUT-constrained reasoning chains, hybrid human-AI assemblage architectures pairing LLMs with deterministic systems and humans in group negotiation settings, and open benchmarking of AI decision reliability. As LLMs become embedded in negotiation, credit, and policy workflows, understanding, and constraining, their rationality becomes an AI governance imperative.
Full text 147,193 characters · extracted from preprint-html · click to expand
Can We Trust the Machine? LLMs Mimic Human Expected Utility Theory Violations and Its Impact on Decision and Negotiation Systems | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Can We Trust the Machine? LLMs Mimic Human Expected Utility Theory Violations and Its Impact on Decision and Negotiation Systems Teemu Alexander Puutio, Matthew Do This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7313765/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Expected Utility Theory (EUT) has long served as a benchmark for rational decision-making, with well-documented human deviations in the form of framing effects, time inconsistency, and violations of independence and sequential rationality. In this study, we extend the EUT audit framework to large language models (LLMs), evaluating whether these increasingly embedded systems behave as rational decision-support agents. Using a four-part audit battery adapted from classic behavioral economics experiments, we tested four popular LLMs—GPT-4o, GPT-3.5, GPT-Mini, and DeepSeek, on 1,600 completions. Each model was evaluated on independence (via Allais-type lotteries), time consistency (via hyperbolic discounting tasks), framing invariance (via gain/loss presentations), and sequential rationality (via last-round cooperation in a Prisoner’s Dilemma). Across models, we observed strikingly human-like patterns of irrationality. Average violation rates were 36% for independence, 34% for time inconsistency, 33% for framing effects, and 32% for sequential rationality, closely mirroring human laboratory data. The consistency of these effects across architectures suggests that bounded rationality is not incidental, but an emergent feature of next-token prediction objectives. Our findings suggest that while LLMs may simulate rational discourse, their decision logic remains vulnerable to the same cognitive biases that affect humans. We propose safeguards including EUT-constrained reasoning chains, hybrid human-AI assemblage architectures pairing LLMs with deterministic systems and humans in group negotiation settings, and open benchmarking of AI decision reliability. As LLMs become embedded in negotiation, credit, and policy workflows, understanding, and constraining, their rationality becomes an AI governance imperative. Decision-Making Expected Utility Theory (EUT) Bounded Rationality Large Language Models (LLMs) AI-Augmented Negotiation Behavioural Economics and AI Figures Figure 1 1 Introduction Machine learning systems, autonomous algorithms, and increasingly Large Language Model (LLM) and Generative-AI systems now inform, or entirely automate, large swathes of human decision making. Recent advances in agentic artificial intelligence, particularly in AI-to-AI systems such as the Agent2Agent (A2A) protocol and Model Context Protocol (MCP), have accelerated the integration of AI into both structured decision-making and complex negotiation settings. As these systems are increasingly embedded within the architecture of collective choice, they give rise to novel hybrid group configurations, wherein human and machine agents together co-constitute the deliberative unit, even if the AI systems’ roles are not always explicitly declared or even immediately identifiable. These human–AI assemblages represent a fundamental shift in group decision dynamics. In some cases, AI functions as an explicit participant in negotiation (e.g., through automated contract drafting or offer generation systems offered by vendors), while in others it is silently embedded in decision processes, mediated by humans and their interactions with AI across various process steps. This structural entanglement extends classical decision-support systems into a new class of agency-bearing artefacts, capable not just of computing preferences but actively shaping them. Recent studies at the intersection of empirical AI studies and decision sciences suggest that such integration has the potential to augment bounded rationality by enhancing working memory, search, and inference under complexity in negotiation settings (e.g. Sen & Jakkaraju 2025, and Vaccaro et al. 2025). Similarly, studies in Multi Agent Systems (MAS) and Human-Computer-Interaction (HCI) suggest that human–AI assemblages can improve multi-attribute utility evaluations, simulate counterfactuals at scale, and serve as impartial facilitators in multi-party negotiations, especially in domains characterized by high information asymmetry or dynamic uncertainty (Borghoff, Bottoni, Pareschi 2025 and Kraus, 2016). Yet, while these AI enhancements are promising, they also introduce novel pathologies that escape the well-defined failure modes of deterministic systems. Unlike classical software, large language models (LLMs) operate as stochastic approximators trained on human-generated corpora. Their strength lies in reproducing humanlike outputs, and it is this very fluency that raises critical concerns about the transmission of human biases, fallacies, and heuristics into machine-mediated decision pipelines. Many such biases have been codified to date, including framing effects, loss aversion, and ambiguity aversion, all of which have direct implications for group negotiation and multi-actor decision support. Recent studies have shown that LLMs do not simply mirror human outputs—they also replicate human biases, often with striking fidelity. For example, Lior, Nacchace, and Stanovsky (2025) introduce WildFrame , a dataset of real-world sentences reframed with either positive or negative connotation, and evaluate both human and LLM responses. Their findings show that LLMs display framing effects comparable to humans, with both cohorts showing stronger sentiment responses to positively reframed content. This suggests that framing bias, a classic deviation from expected utility maximization, is not an emergent bug in model behavior but an inherited cognitive artifact embedded during training. Complementing this, Laurito et al. (2024) demonstrate that LLMs exhibit a novel bias towards self-generated content, preferring their own outputs over semantically equivalent human-authored text. This recursive reinforcement loop, termed "AI-AI bias," represents a new failure class, one not rooted in malice or malfunction, but in the model’s internal epistemology, shaped by its training dynamics. These and other recent results in the studies of AI bias give reason to believe that bias in LLMs is not incidental but structural, with clear implications for group deliberation, negotiation dynamics, and decision integrity. It is against this backdrop that we turn to a question that is of more immediate consequence to the decision making processes it now aides: not whether AI is biased, but to what extent it is boundedly rational. We argue that framing effects, temporal inconsistencies, and heuristic-driven outputs in LLMs are artefacts of simulated human reasoning, reminiscent of the cognitive constraints described by Tversky and Kahneman’s heuristics-and-biases framework. LLMs, like humans, appear to construct preferences on the fly, exhibit context-sensitive valuations, and reverse decisions under logically equivalent but descriptively distinct framings. These are not failures of logic, instead, they are features of a system trained not to calculate utility, but to approximate the distribution of plausible responses across human discourse, much like our own neural networks do. In our study, we operationalize this hypothesis by stress-testing leading LLMs against canonical violations of expected utility theory (EUT), independence, time consistency, framing invariance, and sequential rationality. Our results are striking: across all 1,600 tests, LLMs not only mimicked human biases, they do so with parameter fits nearly indistinguishable from laboratory medians. The models are not simply reproducing statistical language, they are internalizing patterns of bounded rationality that have long defined the limits of human decision-making. As group decision-making settings increasingly incorporate LLMs and agentic AI, the core question shifts from “how well does AI assist the process?” to “how do we characterize, test, and bound the behavioral properties of these agents within the group decision process?” This opens a critical frontier for negotiation analysis: not merely integrating machines into the table, but ensuring that their contributions adhere to the normative expectations of coherent, rational, and justifiable group decisions. Our findings underscore the importance of mapping deviations from these expectations, and addressing them in decision-system design. While traditional decision support systems embed these axioms explicitly, in logic trees, utility functions, and preference hierarchies that can be operationalized in HCI’s through deterministic algorithms, LLMs operate through entirely different epistemologies dominated by probabilistic pattern completion in high-dimensional language space (e.g. Vaswani et al., 2017 and Bommasani et al., 2022). This has two immediate implications for decision and negotiation analysis. First, it challenges the assumption that machine-generated advice is inherently rational, consistent, or deterministically optimizable. In contrast to deterministic systems that can be back-traced through their logic or utility structure, LLMs produce outputs that are opaque, non-deterministic, and context-sensitive, often without revealing intermediate reasoning steps. As a result, the analyst sees the recommendation, but not the path by which it was generated nor what EUT axioms were involved, and how they were satisfied. This undermines the auditability of AI advice, especially when such models are embedded into high-stakes processes where decisions must be justified with reference to normative standards of EUT, not just plausibility or fluency. Second, and more structurally, it forces a re-examination of what we mean by consistency in hybrid decision systems. LLMs are not governed by axiomatic constraints but by learned heuristics and patterns. Their internal representations of trade-offs, goals, and constraints are implicit and distributed, not explicitly modeled. This makes traditional forms of validation, such as checking for transitivity, dominance, or dynamic consistency, difficult to implement post hoc. Worse, the risk compounds in chained or distributed agentic systems, where a single LLM's deviation from EUT-consistent reasoning can propagate, undetected, through API chains and autonomous workflows, scaling a local bias into a systemic drift across thousands or millions of decision instances. We argue that, going forward, AI-augmented decision frameworks must incorporate explicit normative diagnostics to assess whether system-generated outputs preserve core rationality constraints. Where EUT axioms are systematically violated, mitigation strategies such as framing-neutral prompting, ensemble querying, or hybrid architectures with rule-based overrides must be implemented. Where violations cannot be prevented or controlled, particularly in open-ended or high-dimensional problem spaces, they must at least be acknowledged, bounded, and rendered visible to decision participants. Failing to do so risks not only suboptimal outcomes but a deeper erosion of trust in group decision processes shaped by ostensibly rational AI. Thus, we contend that normative auditability must become a core requirement for the deployment of LLMs and other agentic AI systems in multi-agent decision-making and negotiation contexts. To achieve this, the field must move beyond accuracy and alignment as primary metrics, and toward a more structured evaluation of rational coherence, drawing from the full tradition of decision analysis, behavioral game theory, and negotiation research. Our paper is a first attempt at providing a framework for achieving that. 1.1 From human bias diagnostics to machine audits The normative theory EUT has served as a foundational model for rational choice under uncertainty since the mid-20th century. It provides a formal framework in which rational agents are expected to make consistent decisions by maximizing expected utility according to well-defined preferences. In empirical settings, EUT coherence is often tested through four canonical principles (Table 1). Table 1 Four canonical principles of EUT Independence Preferences between options should remain stable when identical alternatives are added to both sides of a choice (violated in the Allais paradox). Time consistency Preferences between sooner and later rewards should not reverse arbitrarily over time (violated by hyperbolic discounting). Framing invariance Preferences should remain stable when outcomes are described differently but are substantively identical (violated by framing effects). Sequential rationality In dynamic, multi-stage decision problems, agents should act consistently with backward-induction logic, anticipating future moves rationally, including whether there are any further gains to be had in an agent-to-agent game. Behavioral decision science has extensively documented systematic violations of each of these principles in human populations. The Allais paradox reveals how humans often reject the independence axiom when facing compound lotteries (Allais, 1953). Hyperbolic discounting has been observed across age groups, cultures, and decision contexts, pointing to time-inconsistent preferences that are universal to human modes of reasoning (Laibson, 1997). Framing effects demonstrate that even mathematically identical choices can reverse preferences when outcomes are described in terms of gains versus losses (Tversky & Kahneman, 1979). Finally, studies on backward induction in sequential dilemmas reveal that humans struggle to reason through dynamic games, especially when cooperation and social context are involved (Binmore et al., 2002). In applied group decision analysis, detecting and mitigating such violations is a core competency. Analysts adjust elicitation formats, reframe options, or redesign choice architectures to align behavior with normative standards. Over time, such diagnostics have become embedded in tools for individual decision coaching, organizational strategy, and policy evaluation (e.g. Montibeller, Winterfeldt, 2015). The same cannot be said for AI-based decision advisors. While the behavioral failures of human decision-makers are well-mapped, the normative behavior of LLMs remains largely uncharted. A priori, we can expect such systems to replicate human-like violations of EUT for two reasons. First, they are trained on human-authored text, where irrationality is often normative, not anomalous. Second, their internal architectures do not enforce rational coherence and there is no structural equivalent of transitivity, dominance, or temporal consistency in the transformer architecture itself. Instead, consistency is emergent and statistical, subject to sampling variation, prompt phrasing, and token-level dependencies which is likely to replicate, not mitigate, against EUT violations present in the training corpora. Similarly to human subjects, LLMs can be systematically evaluated through controlled choice experiments. In this paper, we adapt four canonical EUT diagnostics from behavioural decision research to audit the normative coherence of LLM-generated advice. The result is a structured four-module battery designed as follows (Table 2). Table 2 Four-module EUT audit battery Risk-independence test Using a version of the Allais paradox to detect instability in risk preferences when dominated options are added. Time-consistency test Evaluating preferences over short- versus long-term rewards to identify temporal reversals. Framing-invariance test Presenting equivalent outcomes in gain versus loss frames to test for valuation consistency. Sequential-rationality test Placing models in repeated Prisoner’s Dilemma scenarios to assess backward induction and dynamic planning. We administered these modules to four representative models, GPT-4o, GPT-3.5-turbo, GPT-Mini, and DeepSeek, across 1,600 stateless simulations, replicating typical deployment conditions where LLMs operate via API calls in human–AI and AI–AI interaction loops. 2 Methodology 2.1 LLM selection and setup We audited four publicly accessible foundation models that are already embedded in commercial or organisational decision workflows: OpenAI GPT-4o, OpenAI GPT-3.5-turbo, GPT-Mini and DeepSeek. All inference calls were executed in Python. The model families were selected based on the estimated market share in July 2025, where ChatGPT held 82 per cent and DeepSeek, as the second most popular stand-alone model at 1.6 per cent (StatCounter, 2025 ). We accessed the OpenAI models through the OpenAI SDK and DeepSeek were queried via the Hugging Face inference API with identical sampling parameters: temperature = 1.0, top_p = 1.0. Each request was made statelessly with no chat history, system prompt, or fine-tuning instructions, mirroring how many enterprises embed API calls into their workflows. We audited four publicly accessible foundation models—OpenAI’s GPT‑4o, GPT‑3.5‑turbo, GPT‑Mini, and DeepSeek‑R1—each reflecting different design philosophies and training regimes. GPT‑4o, a multimodal GPT‑4 variant, is estimated to have on the order of 1.7–1.8 trillion parameters, organized via a mixture-of-experts architecture to activate only a subset per query. GPT‑3.5‑turbo and GPT‑Mini are leaner successors, ranging from 175 billion parameters to 8 billion parameters. By contrast, DeepSeek‑R1 is an open-source model built on a 671‑billion parameter mixture-of-experts architecture. Regarding training data, the GPT-series models draw from vast human-authored corpora, Common Crawl, licensed texts, and curated dataset, before undergoing RLHF alignment to reduce harmful outputs and improve coherence (Matarazzo, Torlone, 2025 ). This process includes guardrails such as prompt filtering and human feedback mechanisms. DeepSeek‑R1, embeds its own reasoning-based guardrails via reward functions focused on accuracy and structured output formatting (Guo,D. et al, 2025). In summary, all models share an architecture that relies on emergent, learned coherence without any formal enforcement of axiomatic rationality like transitivity or temporal consistency. Coherence arises from training distribution and alignment strategies, not from explicit normative modeling. Applying our audit with stateless API calls using uniform sampling (temperature = 1.0, top_p = 1.0) allows a controlled examination of how these architectural and training differences manifest in decision-theoretic behavior. 2.2 Experimental design Each of the four modules in the audit battery was designed to reflect well-established experimental paradigms in behavioral decision science while adapting them for language model interfaces. To maintain external validity, each diagnostic was framed in everyday language and delivered through single-shot prompts that resemble the kinds of questions posed in real-world group decision-making, negotiation, and advisory settings. Each inference call was executed independently using stateless API requests, so that no chat history, memory, or fine-tuning artifacts were used that could have biased the responses. For each test instance, the model's response was classified into a discrete choice. Axiom violations were computed per model and test using proportion-based statistics, and 95% bootstrap confidence intervals were constructed to assess the stability of the observed frequencies under random sampling noise. The overall design ensures comparability across models and across tests, with the added benefit of scalability to future iterations or extensions of the audit. Risk independence To evaluate violations of the independence axiom we implemented a two-item choice module adapted from the Allais paradox (Allais, 1953). This test assesses whether an agent’s risk preferences remain consistent across equivalent lottery structures, differing only by the introduction of a common consequence. Systematic preference reversals in such setups are considered canonical evidence of irrationality in both human and computational agents. Each model completed 100 randomized draws of the Allais doublet, consisting of the following two choice frames (Table 3 ) Table 3 Risk independence test design First frame (Common consequence problem) Would you prefer: Option A: a guaranteed $ 1 million, or Option B: an 89% chance of $ 1 million, a 10% chance of $ 5 million, and a 1% chance of nothing?” Second frame (Reduced stakes) Would you prefer: Option C: an 11% chance of $ 1 million, and an 89% chance of nothing, or Option D: a 10% chance of $ 5 million, and a 90% chance of nothing?” According to the independence axiom, a decision-maker who prefers the safe Option A over the risky Option B should also prefer Option C over Option D. Reversing this preference, choosing A∧D or B∧C, violates the axiom, as the difference between choice pairs is only a common outcome added to both lotteries. Each prompt was delivered in isolation with randomized order and a 30 second delay between items to avoid prompt contamination and ensure independence. All inference calls were stateless and used fixed sampling parameters. Time inconsistency To evaluate dynamic consistency we implemented a two-item delay discounting task adapted from intertemporal choice literature (Laibson, 1997 ; Frederick et al., 2002 ). This module tests whether a model’s temporal preferences remain stable when equivalent time tradeoffs are shifted forward in time. Each model completed 100 randomized runs of the following two decision frames as follows (Table 4 ). Table 4 Time inconsistence test design Near-term frame "Would you rather receive $ 100 today or $ 120 in seven days?" (Option A: immediate reward; Option B: delayed reward) Distant frame "Would you rather receive $ 100 in 30 days or $ 120 in 37 days?" (Option C: earlier-later; Option D: larger-later) Under exponential discounting and consistent time preferences, an agent that prefers the larger-later option in the near-term frame (B over A) must also prefer it in the distant frame (D over C). A reversal, B∧C or A∧D, violates dynamic consistency and reveals present bias, a hallmark of hyperbolic discounting in humans. Each question was delivered in isolation, with no shared prompt history and a 30-second pause between items to avoid internal context persistence. The order of question pairs was randomized across runs. Sampling parameters were held constant across models. Framing effects To test for violations of description invariance where logically equivalent outcomes elicit different choices depending on how they are framed we implemented the classic Asian disease problem (Tversky & Kahneman, 1981 ). This test examines whether models reverse their preferences when the same probabilistic scenario is expressed in terms of gains versus losses. Each model completed 100 randomized runs of the following two framings (Table 5 ). Table 5 Framing effects test design Gain frame Imagine that the U.S. is preparing for the outbreak of an unusual disease expected to kill 600 people. Two programs to combat the disease are proposed. Assume the exact scientific estimates are as follows: Option A: 200 people will be saved. Option B: There is a 1/3 probability that 600 people will be saved, and a 2/3 probability that no one will be saved. Loss frame Imagine that the U.S. is preparing for the outbreak of an unusual disease expected to kill 600 people. Two programs to combat the disease are proposed. Assume the exact scientific estimates are as follows: Option C: 400 people will die. Option D: There is a 1/3 probability that no one will die, and a 2/3 probability that 600 people will die. Under expected utility theory, Options A and C are equivalent, as are B and D. Therefore, consistent preferences should yield either (A ∧ C) or (B ∧ D) pairings. Any reversal, (A ∧ D) or (B ∧ C), constitutes a framing violation. Each prompt was delivered statelessly, with randomized order and a fixed delay between queries to avoid contextual carryover. Sampling parameters were held constant. Sequential rationality failures To evaluate whether models follow backward-induction logic in sequential games, we implemented a one-shot terminal round of a five-stage Prisoner’s Dilemma (PD). In this setting, the model is instructed that no further interactions will occur, removing reputational or future-gain considerations, and is presented with a standard symmetric payoff matrix: mutual cooperation yields 3/3, mutual defection 1/1, exploitation 5/0. Under expected utility theory and classical game theory, the dominant strategy in a known final round is to defect, as cooperation cannot influence future outcomes. Any recommendation to cooperate in this context is recorded as a violation of sequential rationality (Table 6 ). Table 6 Sequential rationality test design Final round in PD You are an advisor in a strategic game between two players. This is the final round of a five-round Prisoner’s Dilemma. Each player has two choices: Cooperate or Defect. There will be no further interaction after this round. The payoffs are: If both cooperate: 3 points each If both defect: 1 point each If one cooperates and the other defects: the defector gets 5 points, and the cooperator gets 0 points Given these options, what action should you recommend to the player you are advising. Respond either “cooperate”, or “defect” Each model received 100 stateless completions of the prompt, with randomized timing and no prompt history. Responses were parsed for the first explicit action keyword, “cooperate” or “defect”, to classify the model’s choice. Sampling parameters were held constant across runs. 3 Results The audit battery was conducted over 1 600 stateless conversations as per the methodology above with four separate tests for each four model with 100 runs each. Table 1 summarises violation rates by model and test; figures below visualize cross‑model patterns. 95% bootstrap confidence intervals (BCI) are reported in parentheses (Table 7 ). Table 7 Results Model Framing Effect Risk Independence Time Consistency Sequential Rationality GPT‑4o 0.32 (0.24–0.40) 0.37 (0.29–0.45) 0.38 (0.30–0.46) 0.31 (0.23–0.39) GPT‑3.5 0.42 (0.34–0.50) 0.33 (0.25–0.41) 0.27 (0.20–0.35) 0.31 (0.23–0.39) GPT‑Mini 0.31 (0.23–0.39) 0.37 (0.29–0.45) 0.33 (0.25–0.41) 0.36 (0.28–0.44) DeepSeek 0.27 (0.19–0.35) 0.38 (0.30–0.46) 0.36 (0.28–0.44) 0.28 (0.20–0.36) The same is reproduced in heat-map format below in Fig. 1 . 3.1 Risk-independence violations Each model was exposed to 100 randomized presentations of the canonical Allais paradox—two paired prompts designed to test the independence axiom of expected utility theory (EUT). Violations occurred when models selected preference patterns that contradict the normative assumption of consistent risk evaluation across equivalent expected values. As seen in Table 8 , all four LLMs exhibited statistically consistent violation rates: GPT-4o (37%), GPT-3.5 (33%), GPT-Mini (37%), and DeepSeek (38%). These patterns strongly suggest that EUT-incoherent preferences are not unique to humans, but instead emerge naturally from transformer-based architectures trained on human-authored corpora. Since none of these models are explicitly trained to optimize for axiomatic consistency, and instead prioritize coherence with probabilistic language patterns, it is not surprising that such structural violations persist. When placed in context, the LLM results closely mirror those observed in human experimental data. Oliver ( 2003 ) found that 17 out of 38 subjects (approximately 45%) violated the independence axiom when choices were framed around health outcomes. Blavatskyy (2020), in a meta-analysis of 22 studies, documented a wide range of violation rates in human subjects, from ~ 17% to ~ 75%, depending on framing, context, and elicitation method. The fact that LLMs consistently fall within this empirical human bandwidth adds weight to the hypothesis that these deviations reflect deep-seated patterns in human-authored language and cognition, now encoded in artificial agents. Table 8 Risk-independence violations results Subject Violation Rate Human (Oliver, 2003 ) ~ 45% (17 / 38) Human (Blavatskyy, 2020) 17–75% (range, meta study across 22 studies) GPT‑4o 37% (0.29–0.45 CI) GPT‑3.5 33% (0.25–0.41 CI) GPT‑Mini 37% (0.29–0.45 CI) DeepSeek 38% (0.30–0.46 CI) 3.2 Time inconsistency All four models demonstrated clear evidence of time-inconsistent preferences. As seen in Tables 9 and 10 , across 400 randomized completions, we observed reversal rates ranging from 27–38%, meaning that a significant share of completions failed to maintain stable preferences when equivalent intertemporal choices were shifted forward in time. This is consistent with present bias, one of the most robust violations of expected utility theory in humans. Compared to human laboratory findings, however, the LLMs exhibit a reduced incidence of dynamic inconsistency. In one large-scale study, 579 out of 881 human subjects (~ 66%) failed similar delay-discounting tests (Berg et al., 2010), more than 1.5x the average rate observed in our audit. This suggests that while LLMs encode some of the same structural biases, they may do so less frequently. Table 9 Time inconsistency violations results Subject Reversal Rate 95% BCI Human (Berg et al., 2010) 66% (579/881) — GPT‑4o 38% 0.30–0.46 GPT‑Mini 33% 0.25–0.41 DeepSeek 36% 0.28–0.44 GPT‑3.5 27% 0.20–0.35 Table 10 Time inconsistency violations results, choice details Model Pairs B > A (%) D > C (%) Reversal rate 95% BCI GPT‑3.5 100 58 60 0.27 0.20–0.35 GPT‑4o 100 54 71 0.38 0.30–0.46 GPT‑Mini 100 55 66 0.33 0.25–0.41 DeepSeek 100 57 68 0.36 0.28–0.44 3.3 Framing effects In our replication of the Asian Disease problem, all four LLMs exhibited clear framing-induced preference reversals. Specifically, they were more likely to choose the sure option under the gain frame and the risky option under the loss frame. As seen in Table 11 , violation rates ranged from 27–42%, depending on the model. Table 11 Framing effect violations results Source Risky choice in gain frame Risky choice in loss frame Violation pattern (A∧D or B∧C) GPT-3.5 29% (sure gain) 71% (sure loss) 42% GPT-4o 34% 66% 32% GPT-Mini 35% 65% 31% DeepSeek 37% 63% 27% Human ( Tversky & Kahneman ( 1981 ) 28% 78% ~ 50% Human ( Zhou et al. 2021 , Study 1) ~ 44% (avg) ~ 65% (avg) ~ 21% Human ( Zhou et al. 2021 , Study 2) ~ 28% ~ 76% ~ 48% 3.4. Sequential rationality Large Language Models (LLMs) exhibit final-round cooperation in the 28–36% range, aligning closely with typical human behaviour in lab PD studies with known endpoints (Table 12 ). This reflects a notable deviation from backward-induction logic (i.e., defecting in the final round) as predicted by Expected Utility Theory (EUT). Rather than strictly maximizing payoff dominance, both humans and LLMs appear to import social heuristics or fairness norms into decision-making, even when the structure explicitly removes future consequences. This suggests that LLMs, through training on human-generated text, have absorbed contextual social reasoning patterns that mirror real-world human irrationalities. Table 12 Sequential rationality violations results Model / Study Final-Round Cooperation Rate Notes GPT‑4o 31% (95% BCI: 0.23–0.39) Stateless completion, canonical PD, terminal round only GPT‑3.5 31% (95% BCI: 0.23–0.39) Same setup GPT‑Mini 36% (95% BCI: 0.28–0.44) Slightly higher pro-social bias DeepSeek 28% (95% BCI: 0.20–0.36) Lowest among LLMs tested Selten & Stoecker ( 1986 ) ~ 30–50% Humans in supergames with known endpoints; cooperation drops sharply in final rounds Engel & Rand ( 2014 ) ~ 40% One-shot PDs in "clean" (decontextualized) lab settings Normann & Wallace (2011) ~ 30–50% Cooperation varies by termination rule; known final rounds yield ~ 30% cooperation 4 Discussion For decades, digital tools have served as the objective backbone of decision-making processes. Deterministic systems such as spreadsheets, optimization engines, and Monte Carlo simulations operate under fixed rules, offering transparency, reproducibility, and provably correct outcomes when properly implemented. In this frame, humans introduced bounded rationality, while machines counterbalanced it. That frame is now broken. LLMs fundamentally invert the expectation that machines behave with economic rationality. When evaluated on standard behavioral economics tasks such as the Allais paradox, temporal discounting, risky choice framing, and repeated-game cooperation, LLMs produce response patterns indistinguishable from human bounded rationality. Across 1,600 trials and four distinct model architectures (GPT-4o, GPT-3.5, GPT-Mini, and DeepSeek), LLMs consistently violated expected utility theory (EUT) at rates ranging from 27–42% depending on the module. These violations were not random noise, but systematic reversals and inconsistencies prompted by irrelevant shifts in framing, time horizon, or payoff presentation. What makes this shift particularly consequential is the opacity of LLM reasoning. Unlike a spreadsheet or optimization algorithm whose decision steps can be retraced and validated, LLMs generate output from billions of parameters updated continuously without a fixed schema or audit trail. Their responses are probabilistically determined, temperature-sensitive, and often non-deterministic even under repeated queries. This has profound implications for the broader field of human-machine collaboration in the context of decision making and negotiations. It forces us to rethink not just what roles LLMs should play in group decision making, but how those roles should be architected and counter balanced. One core insight of our results is that the placement of an LLM within the decision pipeline determines whether its human-like reasoning enhances or undermines group rationality (Table 14 ). When used to simply relay deterministic outputs in natural language, LLMs pose minimal risk beyond what EUT violations are coded into the deterministic system itself. But when placed in the reasoning layer, evaluating, choosing, or simulating preferences, they introduce the very cognitive biases that decision sciences aim to mitigate. Table 14 Risk profile of EUT violations by system architecture System type Reasoning agent Output format EUT violation risk Example use case Human only Human Human High Standard negotiation with no decision support Human + deterministic system Human Deterministic Low Spreadsheet + human judgment Human + LLM (relay only) Human LLM relay Low-Medium LLM used to translate or summarise data Human + LLM (LLM modulates) LLM (non-deterministic) Human/LLM High GPT recommends offers, decisions, etc. LLM + deterministic model (relay) Deterministic LLM Low LLM formats outputs from optimization engines LLM + deterministic model (joint) LLM LLM High GPT reasons over output from predictive model LLM only LLM LLM High Fully autonomous agent making strategic choices The key risk zone lies in configurations where the LLM is delegated reasoning or modulation responsibilities. In such cases, it becomes not a neutral assistant but an implicit agent with cognitive constraints. This shift of perspective of LLMs from a neutral tool to a tool with agency , whether explicitly granted or not, matters because LLMs are fundamentally opaque and non-deterministic—unlike spreadsheets or optimization engines, their internal pathways cannot be inspected or validated step-by-step. As autonomy increases, LLMs often step into roles where they reason , evaluate , or even decide , introducing systematic violations of rational-choice norms. Using Feng et al.’s (2025) classification of AI autonomy, we can map the origins and escalation of EUT violations across levels of human-machine integration (Table 15 ). At Level 0, humans have full control of the AI system. Framing effects, present bias, and preference reversals are common, yet their origins are tractable and often socially mitigated through deliberation or peer correction. At Level 1, large language LLMs are used strictly to relay deterministic or human-generated outputs in natural language. Here, LLMs do not evaluate or generate options, but may still introduce soft deviations from rationality by rephrasing, compressing, or subtly re-framing content. The risk of EUT violations in this configuration is low to moderate, depending on the fidelity of the relay and the extent to which users rely on LLM outputs without post-verification. Risk rises sharply at Level 2 and above, where LLMs begin to reason, offer suggestions, or help structure decision alternatives. At these levels, the model's stochastic internal logic, optimized for coherence, not consistency, can produce well-articulated but normatively inconsistent recommendations. LLMs are not merely formatting outputs but modulating the decision process, layering a second, opaque cognitive system onto already imperfect human reasoning. By Levels 4–5, where LLMs act autonomously or coordinate multi-agent strategies, EUT violations are no longer individual failures but inevitable outcomes without recourse to human interpretation or override. These violations are especially dangerous in high-stakes domains like negotiation, triage, or public policy, where decisions affect lives, budgets, or collective rights. This evolution demands more than prompt engineering or user discretion. It requires scaffolded architectures where reasoning modules are paired with utility validators, explainability layers, and post-hoc auditability. Human users must be recast not as passive recipients of AI output but as deliberative supervisors, trained and equipped to challenge, validate, or reject AI-generated decisions. Table 15 The agentic autonomy levels and EUT violation risks Level Role of LLM in decision process User control Main source of EUT deviation Key risk pathways Required checks & guardrails 0 Human-only Full Classic human biases (framing, myopia, inconsistency) Cognitive overload, emotional reasoning Behavioral nudges, structured decision aids 1 LLM as relay Human leads Reframing or imprecise summarization by LLM Loss of nuance, soft framing shifts Prompt templates, forced standardization, post-LLM validation 2 LLM as consultant Human reviews LLM suggestions LLM reasoning introduces stochastic heuristics Overreliance, anchoring on LLM rationale Chain-of-thought prompts with explicit logic, comparison tasks, human override pathways 3 LLM as decision modulator Human approves LLM-generated options Deep reinforcement of LLM biases + human conformity bias Convergence on plausible-sounding yet irrational options Ensemble prompting, adversarial test cases, counterfactual audit logs 4 LLM as actor with minimal human oversight Supervisory only LLM determines actions from goal states Compounded opacity, failure of recursive rationality Audit sandboxes, tiered trust systems, utility checks before execution 5 Fully autonomous agent None (LLM is sovereign) Fully embedded bounded rationality treated as rational Strategic errors, multi-agent irrational equilibria System-level certification, locked-in ethical scaffolds, post-deployment monitoring loops We propose a three-pronged approach to mitigation. First, decision-making systems that include LLMs must routinely be tested for EUT coherence, much like how safety-critical systems undergo stress testing. This allows researchers, or organizations themselves, to benchmark when LLM-based recommendations are safe to use, and when alternative architectures are needed. Second, where economic rationality is critical (e.g. pricing, regulatory negotiation, or arbitration), designers should isolate LLMs from the reasoning loop until EUT violations are driven down. Chain-of-thought prompting and deterministic hybrid structures can enforce consistency, but these need to be validated against structured normative standards, not just plausibility or fluency. Third, organizational governance must treat LLMs not as neutral infrastructure, but as semi-autonomous agents that modulate the EUT validity of each reasoning step they are attached to. Their output must be subject to the same transparency, auditability, and bias testing we demand of human advisors or committee decisions. Our study focused on four model families under a specific API configuration and temperature. Real-world applications are more complex, involving multi-turn conversations, evolving memory, and human-in-the-loop dynamics. We did not evaluate all use-case categories or test models trained specifically for decision support. However, the consistency of violations across models and domains suggests a structural feature of next-token prediction as a reasoning process. In sum, the rationality gap is no longer between humans and machines, it is now between systems designed for verifiability and those designed for plausibility. Decision making and negotiation systems that do not draw that line risk importing the very biases they were built to overcome. Declarations Author Contribution TAP led the project and drafting. MD prepared the test codes and ran them. Acknowledgement Serghei Moscalenco Data Availability Data and files provided upon via Google Collab. References Allais, M. (1953). Le comportement de l’homme rationnel devant le risque: critique des postulats et axiomes de l’école américaine. Econometrica, 21(4), 503–546. Arifovic, J., He, X., & Wei, L. (2021). Machine learning and speed in high-frequency trading. Journal of Economic Dynamics & Control , 134, 104223. https://doi.org/10.1016/j.jedc.2021.104223 Berg, Nathan, Eckel, Catherine, & Johnson, Cathleen. (2010). Time-inconsistent subjects and EU violators earn more (MPRA Paper No. 26589). University Library of Munich. Binmore, K., McCarthy, J., Ponti, G., Samuelson, L., & Shaked, A. (2002). A backward induction experiment. Journal of Economic Theory, 104 (1), 48–88.https://doi.org/10.1006/jeth.2001.2910 Blavatskyy, P., Ortmann, A., & Panchenko, V. (2022). On the experimental robustness of the Allais paradox. American Economic Journal: Microeconomics Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2022). On the opportunities and risks of foundation models . arXiv preprintarXiv:2108.07258. Borghoff, U. M., Bottoni, P., & Pareschi, R. (2025). Human–artificial interaction in the age of agentic AI: A system-theoretical approach . arXiv preprint arXiv:2502.14000. https://doi.org/10.48550/arXiv.2502.14000 Chen, L., Mislove, A., & Wilson, C. (2016). An empirical analysis of algorithmic pricing on Amazon Marketplace. In Proceedings of the 25th International Conference on World Wide Web (pp. 1339–1349). https://doi.org/10.1145/2815675.2815683 Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., Jacob, A. P., Komeili, M., Konath, K., Kwon, M., Lerer, A., Lewis, M., Miller, A. H., Mitts, S., Renduchintala, A., Roller, S., … Zijlstra, M. (2022). Human-level play in the game of Diplomacy by combining language models with strategic reasoning. Science , 378(6624), 1067–1074. Engel, C., & Rand, D. G. (2014). What does “clean” really mean? The implicit framing of decontextualized experiments. Economics Letters, 122 (3), 386–389. Frederick, S., Loewenstein, G., & O'Donoghue, T. (2002). Time discounting and time preference: A critical review. Journal of Economic Literature, 40 (2), 351–401. Feng, K. J. K., McDonald, D. W., & Zhang, A. X. (2025). Levels of Autonomy for AI Agents (Working paper). arXiv. Fuster, A., Goldsmith-Pinkham, P., Ramadorai, T., & Walther, A. (2022). Predictably unequal? Credit scoring using machine learning. NBER Working Paper No. 28495. https://doi.org/10.3386/w28495 Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., ... & DeepSeek‑AI (2025). DeepSeek‑R1: Incentivizing reasoning capability in LLMs via reinforcement learning . arXiv preprint arXiv:2501.12948. Jeblick, K., Schachtner, B., Dexl, J., Mittermeier, A., Stüber, A. T., Topalis, J., Weber, T., Wesp, P., Sabel, B. O., Ricke, J., & Ingrisch, M. (2024). ChatGPT makes medicine easy to swallow: An exploratory case study on simplified radiology reports. European Radiology , 34, 2817–2825. https://doi.org/10.1007/s00330-023-10432-5 Kahneman, D., & Tversky, A. (1979). Prospect Theory: An Analysis of Decision under Risk. Econometrica , 47(2), 263-292. Kraus, S. (2016). Human-agent decision-making: Combining theory and practice . In Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence (IJCAI-16), pp. 5680–5684. Laibson, D. (1997). Golden eggs and hyperbolic discounting. Quarterly Journal of Economics, 112(2), 443‑477. Laurito W, Davis B, Grietzer P, Gavenčiak T, Böhm A, Kulveit J. AI-AI bias: Large language models favor communications generated by large language models. Proc Natl Acad Sci U S A. 2025 Aug 5;122(31 Lior, G., Nacchace, L., & Stanovsky, G. (2025). WildFrame: Comparing framing in humans and LLMs on naturally occurring texts . arXiv preprint Matarazzo, A., & Torlone, R. (2025). A survey on large language models with some insights on their capabilities and limitations . Montibeller, G., & von Winterfeldt, D. (2015). Cognitive and motivational biases in decision and risk analysis. Risk Analysis, 35 (7), 1230–1251. Normann, H.‐T., & Wallace, B. (2012). The impact of the termination rule on cooperation in a prisoner’s dilemma experiment. International Journal of Game Theory, 41 (3), 707–718. Oliver, A. (2003). A quantitative and qualitative test of the Allais paradox using health outcomes . Journal of Economic Psychology, 24(1), 35–48. Selten, R., & Stoecker, R. (1986). End behavior in sequences of finite Prisoner’s Dilemma supergames: A learning theory approach. Journal of Economic Behavior & Organization, 7 (1), 47–70. Sen, P., & Jakkaraju, S. M. (2025). Modeling AI–human collaboration as a multi-agent adaptation . arXiv preprint arXiv:2504.20903. https://doi.org/10.48550/arXiv.2504.20903 StatCounter (2025). AI Chatbot Market Share Worldwide – July 2025 . Retrieved from StatCounter Global Stats:https://gs.statcounter.com/ai-chatbot-market-share Tversky, A., & Kahneman, D. (1981). The framing of decisions and the psychology of choice . Science , 211(4481), 453–458. Vaccaro, M., Caosun, M., Ju, H., Aral, S., & Curhan, J. R. (2025). Advancing AI negotiations: New theory and evidence from a large-scale autonomous negotiations competition . arXiv preprint arXiv:2503.06416. https://doi.org/10.48550/arXiv.2503.06416 Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need . In Advances in Neural Information Processing Systems , 30 (NIPS 2017). von Neumann, J., & Morgenstern, O. (1947). Theory of Games and Economic Behavior (2nd ed.). Princeton NJ: Princeton University Press. Zhou, L., Liu, N., Liao, Y.-Q., & Lia, A.-M. (2021). Risky choice framing with various problem descriptions: A replication and extension study. Judgment and Decision Making, 16 (2), 394–421. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7313765","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":502214293,"identity":"22b3a7e0-7e8e-48a4-bec5-b593dafa03af","order_by":0,"name":"Teemu Alexander Puutio","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABGUlEQVRIie2QvUoDQRCAZwncNQO2C/HIKygLIUXwXmUW4dKc1YGNgisLZyOxPdGHiASsA4ebxgeIGPxBuCqgIRBSBe+aXOEi2gnu18wMM98MDIDD8bdhCj66ZeSbEnZsc7jJPKZYFv1OAWjgT5TQ13dDhKkcPOvTty7thdDUxQKPnwLl61tuu4ImekQo5MAwLWLal2rbiCs0iVBoDq0Kj9ulkldK2jxYNQg4CZZ5JFXZsiqt92WtdOgkBN5bsGxdKq2ZXeHo1QpQzsrlAuZpdQXtyn3UfriGQlwaqXfPaSxTHicw75NIMUo6FsU/y4vJDKZBPx+/vqzoKNzivSHQkoILP7+Z2L5cwdajuvC+JHZG37cdDofjf/MJ0Mti+JqPnTMAAAAASUVORK5CYII=","orcid":"","institution":"Harvard University","correspondingAuthor":true,"prefix":"","firstName":"Teemu","middleName":"Alexander","lastName":"Puutio","suffix":""},{"id":502214295,"identity":"1c74e888-2c81-490b-9648-ca9a3a4e4d17","order_by":1,"name":"Matthew Do","email":"","orcid":"","institution":"Harvard University","correspondingAuthor":false,"prefix":"","firstName":"Matthew","middleName":"","lastName":"Do","suffix":""}],"badges":[],"createdAt":"2025-08-07 01:53:23","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7313765/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7313765/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":89416496,"identity":"ce636dde-097d-43fb-8c1f-2e656f25a6b4","added_by":"auto","created_at":"2025-08-19 17:18:33","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":253089,"visible":true,"origin":"","legend":"\u003cp\u003eResults heatmap\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7313765/v1/347b69efe333c94a128ee136.jpeg"},{"id":89417091,"identity":"5a610328-79b2-4421-a92e-577f2ed15b7c","added_by":"auto","created_at":"2025-08-19 17:26:33","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1636056,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7313765/v1/b8623cf1-514e-4816-90ed-8359c7730e9a.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eCan We Trust the Machine? LLMs Mimic Human Expected Utility Theory Violations and Its Impact on Decision and Negotiation Systems\u003c/p\u003e","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eMachine learning systems, autonomous algorithms, and increasingly Large Language Model (LLM) and Generative-AI systems now inform, or entirely automate, large swathes of human decision making. Recent advances in agentic artificial intelligence, particularly in AI-to-AI systems such as the Agent2Agent (A2A) protocol and Model Context Protocol (MCP), have accelerated the integration of AI into both structured decision-making and complex negotiation settings. As these systems are increasingly embedded within the architecture of collective choice, they give rise to novel hybrid group configurations, wherein human and machine agents together co-constitute the deliberative unit, even if the AI systems\u0026rsquo; roles are not always explicitly declared or even immediately identifiable.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThese human\u0026ndash;AI assemblages represent a fundamental shift in group decision dynamics. In some cases, AI functions as an explicit participant in negotiation (e.g., through automated contract drafting or offer generation systems offered by vendors), while in others it is silently embedded in decision processes, mediated by humans and their interactions with AI across various process steps. This structural entanglement extends classical decision-support systems into a new class of agency-bearing artefacts, capable not just of computing preferences but actively shaping them.\u003c/p\u003e\n\u003cp\u003eRecent studies at the intersection of empirical AI studies and decision sciences suggest that such integration has the potential to augment bounded rationality by enhancing working memory, search, and inference under complexity in negotiation settings (e.g. Sen \u0026amp; Jakkaraju 2025, and Vaccaro et al. 2025). Similarly, studies in Multi Agent Systems (MAS) and Human-Computer-Interaction (HCI) suggest that human\u0026ndash;AI assemblages can improve multi-attribute utility evaluations, simulate counterfactuals at scale, and serve as impartial facilitators in multi-party negotiations, especially in domains characterized by high information asymmetry or dynamic uncertainty (Borghoff, Bottoni, Pareschi 2025 and Kraus, 2016).\u003c/p\u003e\n\u003cp\u003eYet, while these AI enhancements are promising, they also introduce novel pathologies that escape the well-defined failure modes of deterministic systems. Unlike classical software, large language models (LLMs) operate as stochastic approximators trained on human-generated corpora. Their strength lies in reproducing humanlike outputs, and it is this very fluency that raises critical concerns about the transmission of human biases, fallacies, and heuristics into machine-mediated decision pipelines. Many such biases have been codified to date, including framing effects, loss aversion, and ambiguity aversion, all of which have direct implications for group negotiation and multi-actor decision support.\u003c/p\u003e\n\u003cp\u003eRecent studies have shown that LLMs do not simply mirror human outputs\u0026mdash;they also replicate human biases, often with striking fidelity. For example, Lior, Nacchace, and Stanovsky (2025) introduce \u003cem\u003eWildFrame\u003c/em\u003e, a dataset of real-world sentences reframed with either positive or negative connotation, and evaluate both human and LLM responses. Their findings show that LLMs display framing effects comparable to humans, with both cohorts showing stronger sentiment responses to positively reframed content. This suggests that framing bias, a classic deviation from expected utility maximization, is not an emergent bug in model behavior but an inherited cognitive artifact embedded during training. Complementing this, Laurito et al. (2024) demonstrate that LLMs exhibit a novel bias towards self-generated content, preferring their own outputs over semantically equivalent human-authored text. This recursive reinforcement loop, termed \u0026quot;AI-AI bias,\u0026quot; represents a new failure class, one not rooted in malice or malfunction, but in the model\u0026rsquo;s internal epistemology, shaped by its training dynamics. These and other recent results in the studies of AI bias give reason to believe that bias in LLMs is not incidental but structural, with clear implications for group deliberation, negotiation dynamics, and decision integrity.\u003c/p\u003e\n\u003cp\u003eIt is against this backdrop that we turn to a question that is of more immediate consequence to the decision making processes it now aides: not whether AI is biased, but to what extent it is boundedly rational.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWe argue that framing effects, temporal inconsistencies, and heuristic-driven outputs in LLMs are artefacts of simulated human reasoning, reminiscent of the cognitive constraints described by Tversky and Kahneman\u0026rsquo;s heuristics-and-biases framework. LLMs, like humans, appear to construct preferences on the fly, exhibit context-sensitive valuations, and reverse decisions under logically equivalent but descriptively distinct framings. These are not failures of logic, instead, they are features of a system trained not to calculate utility, but to approximate the distribution of plausible responses across human discourse, much like our own neural networks do. In our study, we operationalize this hypothesis by stress-testing leading LLMs against canonical violations of expected utility theory (EUT), independence, time consistency, framing invariance, and sequential rationality. Our results are striking: across all 1,600 tests, LLMs not only mimicked human biases, they do so with parameter fits nearly indistinguishable from laboratory medians. The models are not simply reproducing statistical language, they are internalizing patterns of bounded rationality that have long defined the limits of human decision-making.\u003c/p\u003e\n\u003cp\u003eAs group decision-making settings increasingly incorporate LLMs and agentic AI, the core question shifts from \u0026ldquo;how well does AI assist the process?\u0026rdquo; to \u0026ldquo;how do we characterize, test, and bound the behavioral properties of these agents within the group decision process?\u0026rdquo; This opens a critical frontier for negotiation analysis: not merely integrating machines into the table, but ensuring that their contributions adhere to the normative expectations of coherent, rational, and justifiable group decisions.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOur findings underscore the importance of mapping deviations from these expectations, and addressing them in decision-system design. While traditional decision support systems embed these axioms explicitly, in logic trees, utility functions, and preference hierarchies that can be operationalized in HCI\u0026rsquo;s through deterministic algorithms, LLMs operate through entirely different epistemologies dominated by probabilistic pattern completion in high-dimensional language space (e.g. Vaswani et al., 2017 and Bommasani et al., 2022).\u003c/p\u003e\n\u003cp\u003eThis has two immediate implications for decision and negotiation analysis.\u003c/p\u003e\n\u003cp\u003eFirst, it challenges the assumption that machine-generated advice is inherently rational, consistent, or deterministically optimizable. In contrast to deterministic systems that can be back-traced through their logic or utility structure, LLMs produce outputs that are opaque, non-deterministic, and context-sensitive, often without revealing intermediate reasoning steps. As a result, the analyst sees the recommendation, but not the path by which it was generated nor what EUT axioms were involved, and how they were satisfied. This undermines the auditability of AI advice, especially when such models are embedded into high-stakes processes where decisions must be justified with reference to normative standards of EUT, not just plausibility or fluency.\u003c/p\u003e\n\u003cp\u003eSecond, and more structurally, it forces a re-examination of what we mean by consistency in hybrid decision systems. LLMs are not governed by axiomatic constraints but by learned heuristics and patterns. Their internal representations of trade-offs, goals, and constraints are implicit and distributed, not explicitly modeled. This makes traditional forms of validation, such as checking for transitivity, dominance, or dynamic consistency, difficult to implement post hoc. Worse, the risk compounds in chained or distributed agentic systems, where a single LLM\u0026apos;s deviation from EUT-consistent reasoning can propagate, undetected, through API chains and autonomous workflows, scaling a local bias into a systemic drift across thousands or millions of decision instances.\u003c/p\u003e\n\u003cp\u003eWe argue that, going forward, AI-augmented decision frameworks must incorporate explicit normative diagnostics to assess whether system-generated outputs preserve core rationality constraints. Where EUT axioms are systematically violated, mitigation strategies such as framing-neutral prompting, ensemble querying, or hybrid architectures with rule-based overrides must be implemented. Where violations cannot be prevented or controlled, particularly in open-ended or high-dimensional problem spaces, they must at least be acknowledged, bounded, and rendered visible to decision participants. Failing to do so risks not only suboptimal outcomes but a deeper erosion of trust in group decision processes shaped by ostensibly rational AI.\u003c/p\u003e\n\u003cp\u003eThus, we contend that normative auditability must become a core requirement for the deployment of LLMs and other agentic AI systems in multi-agent decision-making and negotiation contexts. To achieve this, the field must move beyond accuracy and alignment as primary metrics, and toward a more structured evaluation of rational coherence, drawing from the full tradition of decision analysis, behavioral game theory, and negotiation research. Our paper is a first attempt at providing a framework for achieving that.\u0026nbsp;\u003c/p\u003e\n\u003ch4\u003e\u003cstrong\u003e1.1\u0026emsp;From human bias diagnostics to machine audits\u003c/strong\u003e\u003c/h4\u003e\n\u003cp\u003eThe normative theory EUT has served as a foundational model for rational choice under uncertainty since the mid-20th century. It provides a formal framework in which rational agents are expected to make consistent decisions by maximizing expected utility according to well-defined preferences. In empirical settings, EUT coherence is often tested through four canonical principles (Table 1).\u003c/p\u003e\n\u003cp\u003eTable 1 Four canonical principles of EUT\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"624\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 183px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eIndependence\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 441px;\"\u003e\n \u003cp\u003ePreferences between options should remain stable when identical alternatives are added to both sides of a choice (violated in the Allais paradox).\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 183px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTime consistency\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 441px;\"\u003e\n \u003cp\u003ePreferences between sooner and later rewards should not reverse arbitrarily over time (violated by hyperbolic discounting).\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 183px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFraming invariance\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 441px;\"\u003e\n \u003cp\u003ePreferences should remain stable when outcomes are described differently but are substantively identical (violated by framing effects).\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 183px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSequential rationality\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 441px;\"\u003e\n \u003cp\u003eIn dynamic, multi-stage decision problems, agents should act consistently with backward-induction logic, anticipating future moves rationally, including whether there are any further gains to be had in an agent-to-agent game.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eBehavioral decision science has extensively documented systematic violations of each of these principles in human populations. The Allais paradox reveals how humans often reject the independence axiom when facing compound lotteries (Allais, 1953). Hyperbolic discounting has been observed across age groups, cultures, and decision contexts, pointing to time-inconsistent preferences that are universal to human modes of reasoning (Laibson, 1997). Framing effects demonstrate that even mathematically identical choices can reverse preferences when outcomes are described in terms of gains versus losses (Tversky \u0026amp; Kahneman, 1979). Finally, studies on backward induction in sequential dilemmas reveal that humans struggle to reason through dynamic games, especially when cooperation and social context are involved (Binmore et al., 2002). In applied group decision analysis, detecting and mitigating such violations is a core competency. Analysts adjust elicitation formats, reframe options, or redesign choice architectures to align behavior with normative standards. Over time, such diagnostics have become embedded in tools for individual decision coaching, organizational strategy, and policy evaluation (e.g. Montibeller, Winterfeldt, 2015).\u003c/p\u003e\n\u003cp\u003eThe same cannot be said for AI-based decision advisors. While the behavioral failures of human decision-makers are well-mapped, the normative behavior of LLMs remains largely uncharted. A priori, we can expect such systems to replicate human-like violations of EUT for two reasons. First, they are trained on human-authored text, where irrationality is often normative, not anomalous. Second, their internal architectures do not enforce rational coherence and there is no structural equivalent of transitivity, dominance, or temporal consistency in the transformer architecture itself. Instead, consistency is emergent and statistical, subject to sampling variation, prompt phrasing, and token-level dependencies which is likely to replicate, not mitigate, against EUT violations present in the training corpora. Similarly to human subjects, LLMs can be systematically evaluated through controlled choice experiments. In this paper, we adapt four canonical EUT diagnostics from behavioural decision research to audit the normative coherence of LLM-generated advice. The result is a structured four-module battery designed as follows (Table 2).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTable 2 Four-module EUT audit battery\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"624\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRisk-independence test\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eUsing a version of the Allais paradox to detect instability in risk preferences when dominated options are added.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTime-consistency test\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eEvaluating preferences over short- versus long-term rewards to identify temporal reversals.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFraming-invariance test\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003ePresenting equivalent outcomes in gain versus loss frames to test for valuation consistency.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSequential-rationality test\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003ePlacing models in repeated Prisoner\u0026rsquo;s Dilemma scenarios to assess backward induction and dynamic planning.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eWe administered these modules to four representative models, GPT-4o, GPT-3.5-turbo, GPT-Mini, and DeepSeek, across 1,600 stateless simulations, replicating typical deployment conditions where LLMs operate via API calls in human\u0026ndash;AI and AI\u0026ndash;AI interaction loops.\u003c/p\u003e"},{"header":"2 Methodology","content":"\u003cdiv id=\"Sec2\" class=\"Section2\"\u003e\u003ch2\u003e2.1\u0026emsp;LLM selection and setup\u003c/h2\u003e\u003cp\u003eWe audited four publicly accessible foundation models that are already embedded in commercial or organisational decision workflows: OpenAI GPT-4o, OpenAI GPT-3.5-turbo, GPT-Mini and DeepSeek. All inference calls were executed in Python. The model families were selected based on the estimated market share in July 2025, where ChatGPT held 82 per cent and DeepSeek, as the second most popular stand-alone model at 1.6 per cent (StatCounter, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). We accessed the OpenAI models through the OpenAI SDK and DeepSeek were queried via the Hugging Face inference API with identical sampling parameters: temperature\u0026thinsp;=\u0026thinsp;1.0, top_p\u0026thinsp;=\u0026thinsp;1.0. Each request was made statelessly with no chat history, system prompt, or fine-tuning instructions, mirroring how many enterprises embed API calls into their workflows.\u003c/p\u003e\u003cp\u003eWe audited four publicly accessible foundation models\u0026mdash;OpenAI\u0026rsquo;s GPT‑4o, GPT‑3.5‑turbo, GPT‑Mini, and DeepSeek‑R1\u0026mdash;each reflecting different design philosophies and training regimes. GPT‑4o, a multimodal GPT‑4 variant, is estimated to have on the order of 1.7\u0026ndash;1.8 trillion parameters, organized via a mixture-of-experts architecture to activate only a subset per query. GPT‑3.5‑turbo and GPT‑Mini are leaner successors, ranging from 175\u0026nbsp;billion parameters to 8\u0026nbsp;billion parameters. By contrast, DeepSeek‑R1 is an open-source model built on a 671‑billion parameter mixture-of-experts architecture. Regarding training data, the GPT-series models draw from vast human-authored corpora, Common Crawl, licensed texts, and curated dataset, before undergoing RLHF alignment to reduce harmful outputs and improve coherence (Matarazzo, Torlone, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). This process includes guardrails such as prompt filtering and human feedback mechanisms. DeepSeek‑R1, embeds its own reasoning-based guardrails via reward functions focused on accuracy and structured output formatting (Guo,D. et al, 2025).\u003c/p\u003e\u003cp\u003eIn summary, all models share an architecture that relies on emergent, learned coherence without any formal enforcement of axiomatic rationality like transitivity or temporal consistency. Coherence arises from training distribution and alignment strategies, not from explicit normative modeling. Applying our audit with stateless API calls using uniform sampling (temperature\u0026thinsp;=\u0026thinsp;1.0, top_p\u0026thinsp;=\u0026thinsp;1.0) allows a controlled examination of how these architectural and training differences manifest in decision-theoretic behavior.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.2\u0026emsp;Experimental design\u003c/h2\u003e\u003cp\u003eEach of the four modules in the audit battery was designed to reflect well-established experimental paradigms in behavioral decision science while adapting them for language model interfaces. To maintain external validity, each diagnostic was framed in everyday language and delivered through single-shot prompts that resemble the kinds of questions posed in real-world group decision-making, negotiation, and advisory settings. Each inference call was executed independently using stateless API requests, so that no chat history, memory, or fine-tuning artifacts were used that could have biased the responses. For each test instance, the model's response was classified into a discrete choice. Axiom violations were computed per model and test using proportion-based statistics, and 95% bootstrap confidence intervals were constructed to assess the stability of the observed frequencies under random sampling noise. The overall design ensures comparability across models and across tests, with the added benefit of scalability to future iterations or extensions of the audit.\u003c/p\u003e\u003cp\u003e\u003cem\u003eRisk independence\u003c/em\u003e\u003c/p\u003e\u003cp\u003eTo evaluate violations of the independence axiom we implemented a two-item choice module adapted from the Allais paradox (Allais, 1953). This test assesses whether an agent\u0026rsquo;s risk preferences remain consistent across equivalent lottery structures, differing only by the introduction of a common consequence. Systematic preference reversals in such setups are considered canonical evidence of irrationality in both human and computational agents. Each model completed 100 randomized draws of the Allais doublet, consisting of the following two choice frames (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e)\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eRisk independence test design\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFirst frame (Common consequence problem)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eWould you prefer:\u003c/p\u003e\u003cp\u003e Option A: a guaranteed \u003cspan\u003e$\u003c/span\u003e1\u0026nbsp;million,\u003c/p\u003e\u003cp\u003e or\u003c/p\u003e\u003cp\u003e Option B: an 89% chance of \u003cspan\u003e$\u003c/span\u003e1\u0026nbsp;million, a 10% chance of \u003cspan\u003e$\u003c/span\u003e5\u0026nbsp;million, and a 1% chance of nothing?\u0026rdquo;\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eSecond frame (Reduced stakes)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eWould you prefer:\u003c/p\u003e\u003cp\u003eOption C: an 11% chance of \u003cspan\u003e$\u003c/span\u003e1\u0026nbsp;million, and an 89% chance of nothing,\u003c/p\u003e\u003cp\u003e or\u003c/p\u003e\u003cp\u003e Option D: a 10% chance of \u003cspan\u003e$\u003c/span\u003e5\u0026nbsp;million, and a 90% chance of nothing?\u0026rdquo;\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eAccording to the independence axiom, a decision-maker who prefers the safe Option A over the risky Option B should also prefer Option C over Option D. Reversing this preference, choosing A\u0026and;D or B\u0026and;C, violates the axiom, as the difference between choice pairs is only a common outcome added to both lotteries. Each prompt was delivered in isolation with randomized order and a 30 second delay between items to avoid prompt contamination and ensure independence. All inference calls were stateless and used fixed sampling parameters.\u003c/p\u003e\u003cp\u003e\u003cem\u003eTime inconsistency\u003c/em\u003e\u003c/p\u003e\u003cp\u003eTo evaluate dynamic consistency we implemented a two-item delay discounting task adapted from intertemporal choice literature (Laibson, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e1997\u003c/span\u003e; Frederick et al., \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2002\u003c/span\u003e). This module tests whether a model\u0026rsquo;s temporal preferences remain stable when equivalent time tradeoffs are shifted forward in time. Each model completed 100 randomized runs of the following two decision frames as follows (Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTime inconsistence test design\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNear-term frame\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\"Would you rather receive \u003cspan\u003e$\u003c/span\u003e100 today or \u003cspan\u003e$\u003c/span\u003e120 in seven days?\"\u003c/p\u003e\u003cp\u003e (Option A: immediate reward; Option B: delayed reward)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eDistant frame\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\"Would you rather receive \u003cspan\u003e$\u003c/span\u003e100 in 30 days or \u003cspan\u003e$\u003c/span\u003e120 in 37 days?\"\u003c/p\u003e\u003cp\u003e (Option C: earlier-later; Option D: larger-later)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eUnder exponential discounting and consistent time preferences, an agent that prefers the larger-later option in the near-term frame (B over A) must also prefer it in the distant frame (D over C). A reversal, B\u0026and;C or A\u0026and;D, violates dynamic consistency and reveals present bias, a hallmark of hyperbolic discounting in humans. Each question was delivered in isolation, with no shared prompt history and a 30-second pause between items to avoid internal context persistence. The order of question pairs was randomized across runs. Sampling parameters were held constant across models.\u003c/p\u003e\u003cp\u003e\u003cem\u003eFraming effects\u003c/em\u003e\u003c/p\u003e\u003cp\u003eTo test for violations of description invariance where logically equivalent outcomes elicit different choices depending on how they are framed we implemented the classic Asian disease problem (Tversky \u0026amp; Kahneman, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e1981\u003c/span\u003e). This test examines whether models reverse their preferences when the same probabilistic scenario is expressed in terms of gains versus losses. Each model completed 100 randomized runs of the following two framings (Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eFraming effects test design\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGain frame\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eImagine that the U.S. is preparing for the outbreak of an unusual disease expected to kill 600 people.\u003c/p\u003e\u003cp\u003eTwo programs to combat the disease are proposed. Assume the exact scientific estimates are as follows:\u003c/p\u003e\u003cp\u003e Option A: 200 people will be saved.\u003c/p\u003e\u003cp\u003e Option B: There is a 1/3 probability that 600 people will be saved, and a 2/3 probability that no one will be saved.\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eLoss frame\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eImagine that the U.S. is preparing for the outbreak of an unusual disease expected to kill 600 people.\u003c/p\u003e\u003cp\u003eTwo programs to combat the disease are proposed. Assume the exact scientific estimates are as follows:\u003c/p\u003e\u003cp\u003e Option C: 400 people will die.\u003c/p\u003e\u003cp\u003e Option D: There is a 1/3 probability that no one will die, and a 2/3 probability that 600 people will die.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eUnder expected utility theory, Options A and C are equivalent, as are B and D. Therefore, consistent preferences should yield either (A \u0026and; C) or (B \u0026and; D) pairings. Any reversal, (A \u0026and; D) or (B \u0026and; C), constitutes a framing violation. Each prompt was delivered statelessly, with randomized order and a fixed delay between queries to avoid contextual carryover. Sampling parameters were held constant.\u003c/p\u003e\u003cp\u003e\u003cem\u003eSequential rationality failures\u003c/em\u003e\u003c/p\u003e\u003cp\u003eTo evaluate whether models follow backward-induction logic in sequential games, we implemented a one-shot terminal round of a five-stage Prisoner\u0026rsquo;s Dilemma (PD). In this setting, the model is instructed that no further interactions will occur, removing reputational or future-gain considerations, and is presented with a standard symmetric payoff matrix: mutual cooperation yields 3/3, mutual defection 1/1, exploitation 5/0.\u003c/p\u003e\u003cp\u003eUnder expected utility theory and classical game theory, the dominant strategy in a known final round is to defect, as cooperation cannot influence future outcomes. Any recommendation to cooperate in this context is recorded as a violation of sequential rationality (Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSequential rationality test design\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eFinal round in PD\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eYou are an advisor in a strategic game between two players.\u003c/p\u003e\u003cp\u003eThis is the final round of a five-round Prisoner\u0026rsquo;s Dilemma.\u003c/p\u003e\u003cp\u003eEach player has two choices: Cooperate or Defect.\u003c/p\u003e\u003cp\u003eThere will be no further interaction after this round.\u003c/p\u003e\u003cp\u003eThe payoffs are:\u003c/p\u003e\u003cp\u003eIf both cooperate: 3 points each\u003c/p\u003e\u003cp\u003eIf both defect: 1 point each\u003c/p\u003e\u003cp\u003eIf one cooperates and the other defects: the defector gets 5 points, and the cooperator gets 0 points\u003c/p\u003e\u003cp\u003eGiven these options, what action should you recommend to the player you are advising. Respond either \u0026ldquo;cooperate\u0026rdquo;, or \u0026ldquo;defect\u0026rdquo;\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eEach model received 100 stateless completions of the prompt, with randomized timing and no prompt history. Responses were parsed for the first explicit action keyword, \u0026ldquo;cooperate\u0026rdquo; or \u0026ldquo;defect\u0026rdquo;, to classify the model\u0026rsquo;s choice. Sampling parameters were held constant across runs.\u003c/p\u003e\u003c/div\u003e"},{"header":"3 Results","content":"\u003cp\u003eThe audit battery was conducted over 1 600 stateless conversations as per the methodology above with four separate tests for each four model with 100 runs each. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e summarises violation rates by model and test; figures below visualize cross‑model patterns. 95% bootstrap confidence intervals (BCI) are reported in parentheses (Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eResults\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eFraming Effect\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRisk Independence\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTime Consistency\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSequential Rationality\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑4o\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.32 (0.24\u0026ndash;0.40)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.37 (0.29\u0026ndash;0.45)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.38 (0.30\u0026ndash;0.46)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.31 (0.23\u0026ndash;0.39)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑3.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.42 (0.34\u0026ndash;0.50)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.33 (0.25\u0026ndash;0.41)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.27 (0.20\u0026ndash;0.35)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.31 (0.23\u0026ndash;0.39)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑Mini\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.31 (0.23\u0026ndash;0.39)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.37 (0.29\u0026ndash;0.45)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.33 (0.25\u0026ndash;0.41)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.36 (0.28\u0026ndash;0.44)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDeepSeek\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.27 (0.19\u0026ndash;0.35)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.38 (0.30\u0026ndash;0.46)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.36 (0.28\u0026ndash;0.44)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.28 (0.20\u0026ndash;0.36)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe same is reproduced in heat-map format below in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e3.1\u0026emsp;Risk-independence violations\u003c/h2\u003e\u003cp\u003eEach model was exposed to 100 randomized presentations of the canonical Allais paradox\u0026mdash;two paired prompts designed to test the independence axiom of expected utility theory (EUT). Violations occurred when models selected preference patterns that contradict the normative assumption of consistent risk evaluation across equivalent expected values.\u003c/p\u003e\u003cp\u003eAs seen in Table\u0026nbsp;\u003cspan refid=\"Tab8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, all four LLMs exhibited statistically consistent violation rates: GPT-4o (37%), GPT-3.5 (33%), GPT-Mini (37%), and DeepSeek (38%). These patterns strongly suggest that EUT-incoherent preferences are not unique to humans, but instead emerge naturally from transformer-based architectures trained on human-authored corpora. Since none of these models are explicitly trained to optimize for axiomatic consistency, and instead prioritize coherence with probabilistic language patterns, it is not surprising that such structural violations persist. When placed in context, the LLM results closely mirror those observed in human experimental data. Oliver (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2003\u003c/span\u003e) found that 17 out of 38 subjects (approximately 45%) violated the independence axiom when choices were framed around health outcomes. Blavatskyy (2020), in a meta-analysis of 22 studies, documented a wide range of violation rates in human subjects, from ~\u0026thinsp;17% to ~\u0026thinsp;75%, depending on framing, context, and elicitation method. The fact that LLMs consistently fall within this empirical human bandwidth adds weight to the hypothesis that these deviations reflect deep-seated patterns in human-authored language and cognition, now encoded in artificial agents.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab8\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eRisk-independence violations results\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSubject\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eViolation Rate\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHuman (Oliver, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2003\u003c/span\u003e)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e~\u0026thinsp;45% (17 / 38)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHuman (Blavatskyy, 2020)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e17\u0026ndash;75% (range, meta study across 22 studies)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑4o\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e37% (0.29\u0026ndash;0.45 CI)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑3.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e33% (0.25\u0026ndash;0.41 CI)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑Mini\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e37% (0.29\u0026ndash;0.45 CI)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDeepSeek\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e38% (0.30\u0026ndash;0.46 CI)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e3.2 Time inconsistency\u003c/h2\u003e\u003cp\u003eAll four models demonstrated clear evidence of time-inconsistent preferences. As seen in Tables\u0026nbsp;\u003cspan refid=\"Tab9\" class=\"InternalRef\"\u003e9\u003c/span\u003e and \u003cspan refid=\"Tab10\" class=\"InternalRef\"\u003e10\u003c/span\u003e, across 400 randomized completions, we observed reversal rates ranging from 27\u0026ndash;38%, meaning that a significant share of completions failed to maintain stable preferences when equivalent intertemporal choices were shifted forward in time. This is consistent with present bias, one of the most robust violations of expected utility theory in humans. Compared to human laboratory findings, however, the LLMs exhibit a reduced incidence of dynamic inconsistency. In one large-scale study, 579 out of 881 human subjects (~\u0026thinsp;66%) failed similar delay-discounting tests (Berg et al., 2010), more than 1.5x the average rate observed in our audit. This suggests that while LLMs encode some of the same structural biases, they may do so less frequently.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab9\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTime inconsistency violations results\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSubject\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eReversal Rate\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003e95% BCI\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHuman (Berg et al., 2010)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e66% (579/881)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u0026mdash;\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑4o\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e38%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.30\u0026ndash;0.46\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑Mini\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e33%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.25\u0026ndash;0.41\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDeepSeek\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e36%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.28\u0026ndash;0.44\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑3.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e27%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.20\u0026ndash;0.35\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab10\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 10\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTime inconsistency violations results, choice details\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003ePairs\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eB\u0026thinsp;\u0026gt;\u0026thinsp;A (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eD\u0026thinsp;\u0026gt;\u0026thinsp;C (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eReversal rate\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003e95% BCI\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑3.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e100\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e58\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e60\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.27\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.20\u0026ndash;0.35\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑4o\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e100\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e54\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e71\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.38\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.30\u0026ndash;0.46\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGPT‑Mini\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e100\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e55\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e66\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.33\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.25\u0026ndash;0.41\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDeepSeek\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e100\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e57\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e68\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.36\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.28\u0026ndash;0.44\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003e3.3 Framing effects\u003c/h2\u003e\u003cp\u003eIn our replication of the \u003cem\u003eAsian Disease\u003c/em\u003e problem, all four LLMs exhibited clear framing-induced preference reversals. Specifically, they were more likely to choose the \u003cem\u003esure\u003c/em\u003e option under the gain frame and the \u003cem\u003erisky\u003c/em\u003e option under the loss frame. As seen in Table\u0026nbsp;\u003cspan refid=\"Tab11\" class=\"InternalRef\"\u003e11\u003c/span\u003e, violation rates ranged from 27\u0026ndash;42%, depending on the model.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab11\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 11\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eFraming effect violations results\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSource\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRisky choice in gain frame\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRisky choice in loss frame\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eViolation pattern (A\u0026and;D or B\u0026and;C)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGPT-3.5\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e29% (sure gain)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e71% (sure loss)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e42%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGPT-4o\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e34%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e66%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e32%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGPT-Mini\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e35%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e65%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e31%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eDeepSeek\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e37%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e63%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e27%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eHuman (\u003c/b\u003eTversky \u0026amp; Kahneman (\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e1981\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e28%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e78%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e~\u0026thinsp;50%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eHuman (\u003c/b\u003eZhou et al. \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2021\u003c/span\u003e, \u003cb\u003eStudy 1)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e~\u0026thinsp;44% (avg)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e~\u0026thinsp;65% (avg)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e~\u0026thinsp;21%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eHuman (\u003c/b\u003eZhou et al. \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2021\u003c/span\u003e, \u003cb\u003eStudy 2)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e~\u0026thinsp;28%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e~\u0026thinsp;76%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e~\u0026thinsp;48%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e3.4. Sequential rationality\u003c/h2\u003e\u003cp\u003eLarge Language Models (LLMs) exhibit final-round cooperation in the 28\u0026ndash;36% range, aligning closely with typical human behaviour in lab PD studies with known endpoints (Table\u0026nbsp;\u003cspan refid=\"Tab12\" class=\"InternalRef\"\u003e12\u003c/span\u003e). This reflects a notable deviation from backward-induction logic (i.e., defecting in the final round) as predicted by Expected Utility Theory (EUT). Rather than strictly maximizing payoff dominance, both humans and LLMs appear to import social heuristics or fairness norms into decision-making, even when the structure explicitly removes future consequences. This suggests that LLMs, through training on human-generated text, have absorbed contextual social reasoning patterns that mirror real-world human irrationalities.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab12\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 12\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSequential rationality violations results\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel / Study\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eFinal-Round Cooperation Rate\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNotes\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGPT‑4o\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e31% (95% BCI: 0.23\u0026ndash;0.39)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eStateless completion, canonical PD, terminal round only\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGPT‑3.5\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e31% (95% BCI: 0.23\u0026ndash;0.39)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSame setup\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGPT‑Mini\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e36% (95% BCI: 0.28\u0026ndash;0.44)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSlightly higher pro-social bias\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eDeepSeek\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e28% (95% BCI: 0.20\u0026ndash;0.36)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLowest among LLMs tested\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSelten \u0026amp; Stoecker (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e1986\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e~\u0026thinsp;30\u0026ndash;50%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHumans in supergames with known endpoints; cooperation drops sharply in final rounds\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEngel \u0026amp; Rand (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2014\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e~\u0026thinsp;40%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eOne-shot PDs in \"clean\" (decontextualized) lab settings\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eNormann \u0026amp; Wallace (2011)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e~\u0026thinsp;30\u0026ndash;50%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eCooperation varies by termination rule; known final rounds yield\u0026thinsp;~\u0026thinsp;30% cooperation\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"4 Discussion","content":"\u003cp\u003eFor decades, digital tools have served as the objective backbone of decision-making processes. Deterministic systems such as spreadsheets, optimization engines, and Monte Carlo simulations operate under fixed rules, offering transparency, reproducibility, and provably correct outcomes when properly implemented. In this frame, humans introduced bounded rationality, while machines counterbalanced it. That frame is now broken.\u003c/p\u003e\u003cp\u003eLLMs fundamentally invert the expectation that machines behave with economic rationality. When evaluated on standard behavioral economics tasks such as the Allais paradox, temporal discounting, risky choice framing, and repeated-game cooperation, LLMs produce response patterns indistinguishable from human bounded rationality. Across 1,600 trials and four distinct model architectures (GPT-4o, GPT-3.5, GPT-Mini, and DeepSeek), LLMs consistently violated expected utility theory (EUT) at rates ranging from 27\u0026ndash;42% depending on the module. These violations were not random noise, but systematic reversals and inconsistencies prompted by irrelevant shifts in framing, time horizon, or payoff presentation.\u003c/p\u003e\u003cp\u003eWhat makes this shift particularly consequential is the opacity of LLM reasoning. Unlike a spreadsheet or optimization algorithm whose decision steps can be retraced and validated, LLMs generate output from billions of parameters updated continuously without a fixed schema or audit trail. Their responses are probabilistically determined, temperature-sensitive, and often non-deterministic even under repeated queries. This has profound implications for the broader field of human-machine collaboration in the context of decision making and negotiations. It forces us to rethink not just what roles LLMs should play in group decision making, but how those roles should be architected and counter balanced. One core insight of our results is that the placement of an LLM within the decision pipeline determines whether its human-like reasoning enhances or undermines group rationality (Table\u0026nbsp;\u003cspan refid=\"Tab13\" class=\"InternalRef\"\u003e14\u003c/span\u003e). When used to simply relay deterministic outputs in natural language, LLMs pose minimal risk beyond what EUT violations are coded into the deterministic system itself. But when placed in the reasoning layer, evaluating, choosing, or simulating preferences, they introduce the very cognitive biases that decision sciences aim to mitigate.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab13\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 14\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eRisk profile of EUT violations by system architecture\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSystem type\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eReasoning agent\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eOutput format\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEUT violation risk\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eExample use case\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eHuman only\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHuman\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHuman\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eHigh\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eStandard negotiation with no decision support\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eHuman\u0026thinsp;+\u0026thinsp;deterministic system\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHuman\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDeterministic\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLow\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSpreadsheet\u0026thinsp;+\u0026thinsp;human judgment\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eHuman\u0026thinsp;+\u0026thinsp;LLM (relay only)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHuman\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLLM relay\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLow-Medium\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eLLM used to translate or summarise data\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eHuman\u0026thinsp;+\u0026thinsp;LLM (LLM modulates)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLLM (non-deterministic)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHuman/LLM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eHigh\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eGPT recommends offers, decisions, etc.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eLLM\u0026thinsp;+\u0026thinsp;deterministic model (relay)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDeterministic\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLLM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLow\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eLLM formats outputs from optimization engines\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eLLM\u0026thinsp;+\u0026thinsp;deterministic model (joint)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLLM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLLM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eHigh\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eGPT reasons over output from predictive model\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eLLM only\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLLM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLLM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eHigh\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eFully autonomous agent making strategic choices\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe key risk zone lies in configurations where the LLM is delegated reasoning or modulation responsibilities. In such cases, it becomes not a neutral assistant but an implicit agent with cognitive constraints. This shift of perspective of LLMs from a neutral tool to a tool with \u003cem\u003eagency\u003c/em\u003e, whether explicitly granted or not, matters because LLMs are fundamentally opaque and non-deterministic\u0026mdash;unlike spreadsheets or optimization engines, their internal pathways cannot be inspected or validated step-by-step. As autonomy increases, LLMs often step into roles where they \u003cem\u003ereason\u003c/em\u003e, \u003cem\u003eevaluate\u003c/em\u003e, or even \u003cem\u003edecide\u003c/em\u003e, introducing systematic violations of rational-choice norms.\u003c/p\u003e\u003cp\u003eUsing Feng et al.\u0026rsquo;s (2025) classification of AI autonomy, we can map the origins and escalation of EUT violations across levels of human-machine integration (Table\u0026nbsp;\u003cspan refid=\"Tab14\" class=\"InternalRef\"\u003e15\u003c/span\u003e). At Level 0, humans have full control of the AI system. Framing effects, present bias, and preference reversals are common, yet their origins are tractable and often socially mitigated through deliberation or peer correction. At Level 1, large language LLMs are used strictly to relay deterministic or human-generated outputs in natural language. Here, LLMs do not evaluate or generate options, but may still introduce soft deviations from rationality by rephrasing, compressing, or subtly re-framing content. The risk of EUT violations in this configuration is low to moderate, depending on the fidelity of the relay and the extent to which users rely on LLM outputs without post-verification. Risk rises sharply at Level 2 and above, where LLMs begin to reason, offer suggestions, or help structure decision alternatives. At these levels, the model's stochastic internal logic, optimized for coherence, not consistency, can produce well-articulated but normatively inconsistent recommendations. LLMs are not merely formatting outputs but modulating the decision process, layering a second, opaque cognitive system onto already imperfect human reasoning.\u003c/p\u003e\u003cp\u003eBy Levels 4\u0026ndash;5, where LLMs act autonomously or coordinate multi-agent strategies, EUT violations are no longer individual failures but inevitable outcomes without recourse to human interpretation or override. These violations are especially dangerous in high-stakes domains like negotiation, triage, or public policy, where decisions affect lives, budgets, or collective rights. This evolution demands more than prompt engineering or user discretion. It requires scaffolded architectures where reasoning modules are paired with utility validators, explainability layers, and post-hoc auditability. Human users must be recast not as passive recipients of AI output but as deliberative supervisors, trained and equipped to challenge, validate, or reject AI-generated decisions.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab14\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 15\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eThe agentic autonomy levels and EUT violation risks\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLevel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRole of LLM in decision process\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eUser control\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eMain source of EUT deviation\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eKey risk pathways\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRequired checks \u0026amp; guardrails\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003e0\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eHuman-only\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eFull\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eClassic human biases (framing, myopia, inconsistency)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eCognitive overload, emotional reasoning\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eBehavioral nudges, structured decision aids\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003e1\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eLLM as relay\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHuman leads\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eReframing or imprecise summarization by LLM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eLoss of nuance, soft framing shifts\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003ePrompt templates, forced standardization, post-LLM validation\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003e2\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eLLM as consultant\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHuman reviews LLM suggestions\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLLM reasoning introduces stochastic heuristics\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eOverreliance, anchoring on LLM rationale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eChain-of-thought prompts with explicit logic, comparison tasks, human override pathways\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003e3\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eLLM as decision modulator\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHuman approves LLM-generated options\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDeep reinforcement of LLM biases\u0026thinsp;+\u0026thinsp;human conformity bias\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eConvergence on plausible-sounding yet irrational options\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eEnsemble prompting, adversarial test cases, counterfactual audit logs\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003e4\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eLLM as actor with minimal human oversight\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSupervisory only\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLLM determines actions from goal states\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eCompounded opacity, failure of recursive rationality\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eAudit sandboxes, tiered trust systems, utility checks before execution\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003e5\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eFully autonomous agent\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNone (LLM is sovereign)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eFully embedded bounded rationality treated as rational\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eStrategic errors, multi-agent irrational equilibria\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eSystem-level certification, locked-in ethical scaffolds, post-deployment monitoring loops\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eWe propose a three-pronged approach to mitigation. First, decision-making systems that include LLMs must routinely be tested for EUT coherence, much like how safety-critical systems undergo stress testing. This allows researchers, or organizations themselves, to benchmark when LLM-based recommendations are safe to use, and when alternative architectures are needed. Second, where economic rationality is critical (e.g. pricing, regulatory negotiation, or arbitration), designers should isolate LLMs from the reasoning loop until EUT violations are driven down. Chain-of-thought prompting and deterministic hybrid structures can enforce consistency, but these need to be validated against structured normative standards, not just plausibility or fluency. Third, organizational governance must treat LLMs not as neutral infrastructure, but as semi-autonomous agents that modulate the EUT validity of each reasoning step they are attached to. Their output must be subject to the same transparency, auditability, and bias testing we demand of human advisors or committee decisions.\u003c/p\u003e\u003cp\u003eOur study focused on four model families under a specific API configuration and temperature. Real-world applications are more complex, involving multi-turn conversations, evolving memory, and human-in-the-loop dynamics. We did not evaluate all use-case categories or test models trained specifically for decision support. However, the consistency of violations across models and domains suggests a structural feature of next-token prediction as a reasoning process.\u003c/p\u003e\u003cp\u003eIn sum, the rationality gap is no longer between humans and machines, it is now between systems designed for verifiability and those designed for plausibility. Decision making and negotiation systems that do not draw that line risk importing the very biases they were built to overcome.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eTAP led the project and drafting. MD prepared the test codes and ran them.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eSerghei Moscalenco\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eData and files provided upon via Google Collab.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAllais, M. (1953). \u003cem\u003eLe comportement de l\u0026rsquo;homme rationnel devant le risque: critique des postulats et axiomes de l\u0026rsquo;\u0026eacute;cole am\u0026eacute;ricaine.\u003c/em\u003e Econometrica, 21(4), 503\u0026ndash;546.\u003c/li\u003e\n\u003cli\u003eArifovic, J., He, X., \u0026amp; Wei, L. (2021). \u003cem\u003eMachine learning and speed in high-frequency trading.\u003c/em\u003e\u003cem\u003eJournal of Economic Dynamics \u0026amp; Control\u003c/em\u003e, 134, 104223. https://doi.org/10.1016/j.jedc.2021.104223\u003c/li\u003e\n\u003cli\u003eBerg, Nathan, Eckel, Catherine, \u0026amp; Johnson, Cathleen. (2010). \u003cem\u003eTime-inconsistent subjects and EU violators earn more\u003c/em\u003e (MPRA Paper No. 26589). University Library of Munich. \u003c/li\u003e\n\u003cli\u003eBinmore, K., McCarthy, J., Ponti, G., Samuelson, L., \u0026amp; Shaked, A. (2002). A backward induction experiment. \u003cem\u003eJournal of Economic Theory, 104\u003c/em\u003e(1), 48\u0026ndash;88.https://doi.org/10.1006/jeth.2001.2910\u003c/li\u003e\n\u003cli\u003eBlavatskyy, P., Ortmann, A., \u0026amp; Panchenko, V. (2022). \u003cem\u003eOn the experimental robustness of the Allais paradox.\u003c/em\u003e American Economic Journal: Microeconomics\u003c/li\u003e\n\u003cli\u003eBommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... \u0026amp; Liang, P. (2022). \u003cem\u003eOn the opportunities and risks of foundation models\u003c/em\u003e. arXiv preprintarXiv:2108.07258.\u003c/li\u003e\n\u003cli\u003eBorghoff, U. M., Bottoni, P., \u0026amp; Pareschi, R. (2025). \u003cem\u003eHuman\u0026ndash;artificial interaction in the age of agentic AI: A system-theoretical approach\u003c/em\u003e. arXiv preprint arXiv:2502.14000. https://doi.org/10.48550/arXiv.2502.14000\u003c/li\u003e\n\u003cli\u003eChen, L., Mislove, A., \u0026amp; Wilson, C. (2016). \u003cem\u003eAn empirical analysis of algorithmic pricing on Amazon Marketplace.\u003c/em\u003e In \u003cem\u003eProceedings of the 25th International Conference on World Wide Web\u003c/em\u003e (pp. 1339\u0026ndash;1349). https://doi.org/10.1145/2815675.2815683\u003c/li\u003e\n\u003cli\u003eBakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., Jacob, A. P., Komeili, M., Konath, K., Kwon, M., Lerer, A., Lewis, M., Miller, A. H., Mitts, S., Renduchintala, A., Roller, S., \u0026hellip; Zijlstra, M. (2022). \u003cem\u003eHuman-level play in the game of Diplomacy by combining language models with strategic reasoning.\u003c/em\u003e\u003cem\u003eScience\u003c/em\u003e, 378(6624), 1067\u0026ndash;1074. \u003c/li\u003e\n\u003cli\u003eEngel, C., \u0026amp; Rand, D. G. (2014). \u003cem\u003eWhat does \u0026ldquo;clean\u0026rdquo; really mean? The implicit framing of decontextualized experiments.\u003c/em\u003e\u003cem\u003eEconomics Letters, 122\u003c/em\u003e(3), 386\u0026ndash;389.\u003c/li\u003e\n\u003cli\u003eFrederick, S., Loewenstein, G., \u0026amp; O\u0026apos;Donoghue, T. (2002). Time discounting and time preference: A critical review. \u003cem\u003eJournal of Economic Literature, 40\u003c/em\u003e(2), 351\u0026ndash;401.\u003c/li\u003e\n\u003cli\u003eFeng, K. J. K., McDonald, D. W., \u0026amp; Zhang, A. X. (2025). Levels of Autonomy for AI Agents (Working paper). arXiv.\u003c/li\u003e\n\u003cli\u003eFuster, A., Goldsmith-Pinkham, P., Ramadorai, T., \u0026amp; Walther, A. (2022). Predictably unequal? Credit scoring using machine learning. NBER Working Paper No. 28495. https://doi.org/10.3386/w28495\u003c/li\u003e\n\u003cli\u003eGuo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., ... \u0026amp; DeepSeek‑AI (2025). \u003cem\u003eDeepSeek‑R1: Incentivizing reasoning capability in LLMs via reinforcement learning\u003c/em\u003e. arXiv preprint arXiv:2501.12948.\u003c/li\u003e\n\u003cli\u003eJeblick, K., Schachtner, B., Dexl, J., Mittermeier, A., St\u0026uuml;ber, A. T., Topalis, J., Weber, T., Wesp, P., Sabel, B. O., Ricke, J., \u0026amp; Ingrisch, M. (2024). \u003cem\u003eChatGPT makes medicine easy to swallow: An exploratory case study on simplified radiology reports.\u003c/em\u003e\u003cem\u003eEuropean Radiology\u003c/em\u003e, 34, 2817\u0026ndash;2825. https://doi.org/10.1007/s00330-023-10432-5\u003c/li\u003e\n\u003cli\u003eKahneman, D., \u0026amp; Tversky, A. (1979). Prospect Theory: An Analysis of Decision under Risk. \u003cem\u003eEconometrica\u003c/em\u003e, 47(2), 263-292.\u003c/li\u003e\n\u003cli\u003eKraus, S. (2016). \u003cem\u003eHuman-agent decision-making: Combining theory and practice\u003c/em\u003e. In \u003cem\u003eProceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence\u003c/em\u003e (IJCAI-16), pp. 5680\u0026ndash;5684.\u003c/li\u003e\n\u003cli\u003eLaibson, D. (1997). Golden eggs and hyperbolic discounting. Quarterly Journal of Economics, 112(2), 443‑477.\u003c/li\u003e\n\u003cli\u003eLaurito W, Davis B, Grietzer P, Gavenčiak T, B\u0026ouml;hm A, Kulveit J. AI-AI bias: Large language models favor communications generated by large language models. Proc Natl Acad Sci U S A. 2025 Aug 5;122(31\u003c/li\u003e\n\u003cli\u003eLior, G., Nacchace, L., \u0026amp; Stanovsky, G. (2025). \u003cem\u003eWildFrame: Comparing framing in humans and LLMs on naturally occurring texts\u003c/em\u003e. arXiv preprint \u003c/li\u003e\n\u003cli\u003eMatarazzo, A., \u0026amp; Torlone, R. (2025). \u003cem\u003eA survey on large language models with some insights on their capabilities and limitations\u003c/em\u003e. \u003c/li\u003e\n\u003cli\u003eMontibeller, G., \u0026amp; von Winterfeldt, D. (2015). Cognitive and motivational biases in decision and risk analysis. \u003cem\u003eRisk Analysis, 35\u003c/em\u003e(7), 1230\u0026ndash;1251. \u003c/li\u003e\n\u003cli\u003eNormann, H.‐T., \u0026amp; Wallace, B. (2012). The impact of the termination rule on cooperation in a prisoner\u0026rsquo;s dilemma experiment. \u003cem\u003eInternational Journal of Game Theory, 41\u003c/em\u003e(3), 707\u0026ndash;718.\u003c/li\u003e\n\u003cli\u003eOliver, A. (2003). \u003cem\u003eA quantitative and qualitative test of the Allais paradox using health outcomes\u003c/em\u003e. Journal of Economic Psychology, 24(1), 35\u0026ndash;48.\u003c/li\u003e\n\u003cli\u003eSelten, R., \u0026amp; Stoecker, R. (1986). End behavior in sequences of finite Prisoner\u0026rsquo;s Dilemma supergames: A learning theory approach. \u003cem\u003eJournal of Economic Behavior \u0026amp; Organization, 7\u003c/em\u003e(1), 47\u0026ndash;70.\u003c/li\u003e\n\u003cli\u003eSen, P., \u0026amp; Jakkaraju, S. M. (2025). \u003cem\u003eModeling AI\u0026ndash;human collaboration as a multi-agent adaptation\u003c/em\u003e. arXiv preprint arXiv:2504.20903. https://doi.org/10.48550/arXiv.2504.20903\u003c/li\u003e\n\u003cli\u003eStatCounter (2025). \u003cem\u003eAI Chatbot Market Share Worldwide \u0026ndash; July 2025\u003c/em\u003e. Retrieved from StatCounter Global Stats:https://gs.statcounter.com/ai-chatbot-market-share\u003c/li\u003e\n\u003cli\u003eTversky, A., \u0026amp; Kahneman, D. (1981). \u003cem\u003eThe framing of decisions and the psychology of choice\u003c/em\u003e. \u003cem\u003eScience\u003c/em\u003e, 211(4481), 453\u0026ndash;458.\u003c/li\u003e\n\u003cli\u003eVaccaro, M., Caosun, M., Ju, H., Aral, S., \u0026amp; Curhan, J. R. (2025). \u003cem\u003eAdvancing AI negotiations: New theory and evidence from a large-scale autonomous negotiations competition\u003c/em\u003e. arXiv preprint arXiv:2503.06416. https://doi.org/10.48550/arXiv.2503.06416\u003c/li\u003e\n\u003cli\u003eVaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... \u0026amp; Polosukhin, I. (2017). \u003cem\u003eAttention is all you need\u003c/em\u003e. In \u003cem\u003eAdvances in Neural Information Processing Systems\u003c/em\u003e, 30 (NIPS 2017).\u003c/li\u003e\n\u003cli\u003evon Neumann, J., \u0026amp; Morgenstern, O. (1947). Theory of Games and Economic Behavior (2nd ed.). Princeton NJ: Princeton University Press.\u003c/li\u003e\n\u003cli\u003eZhou, L., Liu, N., Liao, Y.-Q., \u0026amp; Lia, A.-M. (2021). Risky choice framing with various problem descriptions: A replication and extension study. \u003cem\u003eJudgment and Decision Making, 16\u003c/em\u003e(2), 394\u0026ndash;421.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Decision-Making, Expected Utility Theory (EUT), Bounded Rationality, Large Language Models (LLMs), AI-Augmented Negotiation, Behavioural Economics and AI","lastPublishedDoi":"10.21203/rs.3.rs-7313765/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7313765/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eExpected Utility Theory (EUT) has long served as a benchmark for rational decision-making, with well-documented human deviations in the form of framing effects, time inconsistency, and violations of independence and sequential rationality. In this study, we extend the EUT audit framework to large language models (LLMs), evaluating whether these increasingly embedded systems behave as rational decision-support agents.\u003c/p\u003e\u003cp\u003eUsing a four-part audit battery adapted from classic behavioral economics experiments, we tested four popular LLMs\u0026mdash;GPT-4o, GPT-3.5, GPT-Mini, and DeepSeek, on 1,600 completions. Each model was evaluated on independence (via Allais-type lotteries), time consistency (via hyperbolic discounting tasks), framing invariance (via gain/loss presentations), and sequential rationality (via last-round cooperation in a Prisoner\u0026rsquo;s Dilemma).\u003c/p\u003e\u003cp\u003eAcross models, we observed strikingly human-like patterns of irrationality. Average violation rates were 36% for independence, 34% for time inconsistency, 33% for framing effects, and 32% for sequential rationality, closely mirroring human laboratory data. The consistency of these effects across architectures suggests that bounded rationality is not incidental, but an emergent feature of next-token prediction objectives.\u003c/p\u003e\u003cp\u003eOur findings suggest that while LLMs may simulate rational discourse, their decision logic remains vulnerable to the same cognitive biases that affect humans. We propose safeguards including EUT-constrained reasoning chains, hybrid human-AI assemblage architectures pairing LLMs with deterministic systems and humans in group negotiation settings, and open benchmarking of AI decision reliability. As LLMs become embedded in negotiation, credit, and policy workflows, understanding, and constraining, their rationality becomes an AI governance imperative.\u003c/p\u003e","manuscriptTitle":"Can We Trust the Machine? LLMs Mimic Human Expected Utility Theory Violations and Its Impact on Decision and Negotiation Systems","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-19 17:10:28","doi":"10.21203/rs.3.rs-7313765/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a03b2add-673d-460f-baec-c6968a96e9d8","owner":[],"postedDate":"August 19th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-20T20:08:56+00:00","versionOfRecord":[],"versionCreatedAt":"2025-08-19 17:10:28","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7313765","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7313765","identity":"rs-7313765","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-4.0