When Do LLMs Say “Some”? An Investigation of Scalar Implicature and Politeness Mitigation in Large Language Models

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Scalar implicatures link semantic meaning to pragmatic reasoning: although some is logically compatible with all, listeners often enrich some to “some but not all.” Yet some also functions as a politeness hedge that mitigates face-threatening acts, potentially canceling the usual “not all” inference. This study investigates whether large language models (LLMs) exhibit the same context-sensitive trade-off between informativeness and politeness as humans, and whether chain-of-thought (CoT) optimization yields more human-like pragmatic flexibility. Mandarin-speaking participants completed parallel production and judgment tasks that manipulated face threat (face-threatening vs. neutral) and factual state ( All vs. Most ). In Experiment 1, participants completed sentences among seven quantifiers, including some in Experiment 2, they evaluated the acceptability of under-informative some statements versus fully informative factual statements, with reaction times recorded. Three LLMs (DeepSeek-V3.2, DeepSeek-R1, GPT-4o) were tested with identical materials. Humans showed a robust informativeness–politeness reweighting: face threat increased production and acceptance of some while sharply reducing acceptance of blunt, fully informative statements. LLMs generally captured the face-based licensing of vagueness but did not reliably penalize socially costly maximal informativeness. The CoT-optimized model showed greater context sensitivity and more human-like distributional patterns than the non-CoT models. Together, these findings are consistent with partial alignment with human mitigation but incomplete social-norm calibration, with CoT optimization offering a modest reduction in the mismatch.
Full text 238,144 characters · extracted from preprint-html · click to expand
When Do LLMs Say “Some”? An Investigation of Scalar Implicature and Politeness Mitigation in Large Language Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article When Do LLMs Say “Some”? An Investigation of Scalar Implicature and Politeness Mitigation in Large Language Models Yuying Wang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8763414/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 7 You are reading this latest preprint version Abstract Scalar implicatures link semantic meaning to pragmatic reasoning: although some is logically compatible with all, listeners often enrich some to “some but not all.” Yet some also functions as a politeness hedge that mitigates face-threatening acts, potentially canceling the usual “not all” inference. This study investigates whether large language models (LLMs) exhibit the same context-sensitive trade-off between informativeness and politeness as humans, and whether chain-of-thought (CoT) optimization yields more human-like pragmatic flexibility. Mandarin-speaking participants completed parallel production and judgment tasks that manipulated face threat (face-threatening vs. neutral) and factual state ( All vs. Most ). In Experiment 1, participants completed sentences among seven quantifiers, including some in Experiment 2, they evaluated the acceptability of under-informative some statements versus fully informative factual statements, with reaction times recorded. Three LLMs (DeepSeek-V3.2, DeepSeek-R1, GPT-4o) were tested with identical materials. Humans showed a robust informativeness–politeness reweighting: face threat increased production and acceptance of some while sharply reducing acceptance of blunt, fully informative statements. LLMs generally captured the face-based licensing of vagueness but did not reliably penalize socially costly maximal informativeness. The CoT-optimized model showed greater context sensitivity and more human-like distributional patterns than the non-CoT models. Together, these findings are consistent with partial alignment with human mitigation but incomplete social-norm calibration, with CoT optimization offering a modest reduction in the mismatch. Humanities/Philosophy Biological sciences/Psychology Social science/Psychology Social science/Science technology and society scalar implicature politeness mitigation quantifiers pragmatics large language models chain-of-thought Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Scalar implicatures represent a key interface between logical meaning and pragmatic reasoning (Grice, 1975; Horn, 1972, 2006; Levinson, 2000), arising when a speaker’s use of a weak scalar term invites the inference that a stronger alternative does not hold (Horn, 1972). The quantifier some offers the canonical illustration: although it semantically denotes a non-zero subset and is logically compatible with the universal quantifier all , comprehenders reliably interpret some as conveying the strengthened meaning “ some but not all ” (Chemla & Bott, 2014; Chemla & Singh, 2014; Katsos & Cummins, 2010). Consider the contrast in (1). (1a) Some of the experimental materials were human-created. (1b) All of the experimental materials were human-created. From a purely logical standpoint, if (1b) is true, (1a) must be true as well; thus, (1a) and (1b) are congruent with respect to their truth conditions. However, human comprehenders typically infer from (1a) that not all of the experimental materials were human-created. This enriched interpretation is emerges from expectations regarding communicative cooperativity and informativeness within the Gricean framework (Grice, 1975). Within the Gricean Cooperative Principle, communication is assumed to be rational and goal-directed. If a speaker knows that a stronger, more informative statement (e.g., all ) is true but nevertheless opts for a weaker expression (e.g., some ), the listener is warranted in inferring that the speaker is deliberately avoiding the stronger alternative. The use of some thus implicates that all does not hold, giving rise to the scalar implicature not all . Politeness and face threat mitigation Speakers typically follow Gricean maxims by being as informative as possible. Pragmatic theory, however, has long emphasized that social context can override this default. Brown and Levinson (1987) define face as a person’s public self-image and show that speakers use mitigation strategies to avoid threatening face. A direct criticism is a prototypical face-threatening acts (FTAs), and speakers often soften such acts with indirect or imprecise, vague language. Particularly, in such contexts, scalar quantifier some can function as a hedge that narrows the scope of the negative assessment, thereby reducing its face-threat relative to a blunt, fully factual statement. Consequently, the pragmatic use of the vague quantifier some often extends beyond the Gricean pressure toward maximal informativeness. For example, when offering critical feedback after the party, knowing that all of the guests found the joke offensive, a speaker might choose between the following formulations: (2a) All guests didn’t like your joke. (2b) Some guests didn’t like your joke. In this setting, the use of some in (2b) does not invite the scalar inference not all but to mitigation of FTAs. In fact, although saying some when all will interpret as underinformative but it will be a polite softening of the statement. Thus, the trade-off between informativeness and politeness is highly context-dependent. Bonnefon et al. (2009) provide experimental evidence for this effect. They found that when an utterance would threaten the listener’s face, listeners are less likely to draw the usual implicature from some to not all . Importantly, further experiments further refine the boundary conditions: when the negative evaluation directly targets the addressee’s own face (rather than involving a third party), scalar enrichment is selectively weakened, indicating that face concerns can modulate the derivation of scalar meaning. Subsequent work corroborates this broader efficiency–politeness trade-off across scalar expressions. Politeness-oriented contexts tend to reduce listeners’ endorsement of enriched interpretations (Feeney & Bonnefon, 2013), and may yield “polite misunderstandings” by increasing ambiguity about whether a speaker’s wording is intended to maximize clarity or to manage face (Bonnefon et al., 2011). Extending this account to Mandarin, Zhang and Wu (2020) show that listeners’ interpretation of some is modulated by motive attribution: when the utterance is construed as informativeness-driven, listeners are more likely to endorse scalar enrichment; when it is identified as a politeness/face-management strategy, they tend to prefer a literal, all-compatible reading, thereby suppressing the scalar inference. Large language models and pragmatic flexibility The advent of large language models (LLMs) opens new avenues for studying pragmatics, but their pragmatic capabilities are still need being understood (Cong, 2024; Ma et al., 2025; Shulginov et al., 2025; Wu et al., 2024; Yu et al., 2025; Yue et al., 2024; Zhao & Hawkins, 2025). LLMs are trained on massive text corpora and often generate fluent, contextually appropriate language, yet it remains unclear whether they exhibit human-like pragmatic flexibility, particularly in domains such as scalar implicature and politeness. Recent studies offer a nuanced picture of how LLMs handle scalar implicatures triggered by the quantifier some . On one hand, certain LLMs do generate the expected implicature in simple contexts, aligning with human intuitions. For instance, Cho and Kim (2024) found that even without any context, both BERT and GPT-2 tend to treat some pragmatically as implying not all , suggesting an inherent grasp of the implicature at a basic level. BERT in particular appeared to encode this inference by default. On the other hand, noticeable differences from human pragmatics emerge once contextual factors are introduced. Building on the classic paradigm of Bonnefon et al. (2009), Qiu, Duan, and Cai (2023) found that GPT-3.5 shows a highly uniform pragmatic interpretation preference for some across both face-threatening and face-enhancing contexts. Its interpretations are entirely unaffected by face considerations, indicating a lack of the human-like ability to flexibly suppress the “not all” scalar inference in response to different pragmatic cues. In contrast, Ku (2025) notes that a single GPT-based agent will often interpret “ Some students passed” in a strictly semantic way that permits all students passing. However, when the same model was placed in a multi-agent interactive framework, the system is able to derive scalar inferences that are much closer to human-like interpretations. This gap underscores that, despite their impressive linguistic fluency, traditional LLMs fall short of internalizing the context-sensitive nature of scalar implicature as observed in human communication. From output to reasoning: a Chain-of-Thought perspective on pragmatic flexibility in LLMs Recent advances in model training have enabled LLMs to generate not only answers but also explicit reasoning traces when instructed to do so. Chain-of-thought (CoT) prompting has been shown to elicit more structured, stepwise solutions and to improve performance on tasks requiring multi-step computation (Wei et al., 2022). The emergence of reasoning-oriented models such as DeepSeek-R1 reflects a growing research trend toward training LLMs with reinforcement learning objectives that directly incentivize stepwise reasoning behaviors rather than relying solely on scale and instruction tuning (Wang, 2025). In the DeepSeek-R1 framework, pure reinforcement learning is used to encourage advanced reasoning patterns such as self-reflection and verification, which has been shown to enhance performance on verifiable reasoning tasks including mathematics and coding relative to conventional supervised approaches (Guo et al., 2025). This shift is methodologically consequential not only for its performance implications, but also because CoT-oriented models make aspects of their decision process observable. The length and structure of reasoning traces can be treated as a proxy for computational cost, enabling comparisons not only of final outputs but also of the effort required to produce them. de Varda et al. (2025) show that reasoning cost—operationalized as the number of reasoning steps or tokens—scales with task difficulty in ways that parallel human cognitive effort. At the same time, it remains unclear whether such traces correspond to the model’s internal inference processes or instead reflect learned strategies for explanation rather than transparent readouts of underlying computation. Nevertheless, CoT-oriented training provides a new empirical window for systematically characterizing pragmatic flexibility in LLMs. Current study and research questions To address these issues, we conducted two experiments in parallel with human participants and LLMs. In the quantifier production task, agents were asked to complete sentences as a speaker. In the acceptability judgment task, humans and LLMs judged how appropriate utterances with some or stronger quantifiers were in those contexts as a listener. All stimuli, procedure, and response formats are matched across humans and LLMs. We included both chain-of-thought–trained and standard LLMs (GPT-4o, DeepSeek-V3.2, and DeepSeek-R1) to examine how reasoning optimization shapes pragmatic behavior. This rigorous design allows us to test, within a single unified framework, whether the resulting patterns of quantifier production and acceptability judgments in LLMs align with those observed in human participants. At the same time, by contrasting reasoning-optimized and non–Chain-of-Thought models, we assess whether explicit reasoning supervision yields more human-like sensitivity to contextual and social constraints on informativeness, or whether such pragmatic adaptation can be accounted for by surface-level language modeling alone. These questions will clarify the extent to which LLMs replicate human pragmatic strategies and illuminate whether their outputs reflect genuine evaluative processes or mere statistical tendencies. Methods Participants Forty-eight native utterances of Mandarin Chinese (females = 24; age M = 20.8 ± 2.10 years) were recruited for Experiment 1. Another forty-eight native utterances of Mandarin Chinese for Experiment 2 (females = 24; age M = 21.38 ± 2.11 years). All participants had normal or corrected-to-normal vision and gave informed consent. Each participant was a native Chinese speaker to ensure they understood the linguistic nuances of the materials. Participants were tested individually in a quiet loom room and were compensated for their time. This research was conducted in accordance with the Declaration of Helsinki and was approved by the Ethics Committee of the author’s institution. Language models We evaluated three Large Language Models using the same experimental materials as the human participants. The models were accessed via the PsyLingLLM toolbox, which presents linguistic stimuli to LLMs in a controlled manner. The models included three contemporary LLMs spanning distinct inference and deployment regimes. (1) GPT-4o is a proprietary, instruction-tuned OpenAI GPT-4–class model evaluated in a direct-answer setting. (2) DeepSeek-V3.2 is a state-of-the-art open-source Mixture-of-Experts LLM (671B parameters) evaluated in the same direct-answer setting. (3) DeepSeek-R1 is a reasoning-optimized derivative of DeepSeek-V3 (671B Mixture-of-Experts), explicitly optimized via reinforcement learning to generate CoT style reasoning; it was evaluated in a trace-emitting setting in which an explicit multi-step reasoning trace precedes the final answer. For analysis, we distinguish three model categories: a trace-emitting CoT model (DeepSeek-R1), a proprietary direct-answer GPT model (GPT-4o), and an open-source direct-answer non-CoT model (DeepSeek-V3.2). Each model was queried via API access through the PsyLingLLM framework using identical prompts, ensuring consistency in input presentation, output format, and generation parameters across models. Materials and design Experiment 1: Pragmatic Completion. This experiment employed a 2 × 2 within-subjects design. We manipulated Face Threat (face-threatening vs. neutral context) and Factual State (All vs. Most, the ground-truth prevalence in the context) as within-item factors. The materials consisted of sentence contexts and prompts designed to elicit a quantifier. Each trial presented a short vignette describing a scenario involving a group of people or items, along with a prompt sentence to be completed with a quantifier. The Face Threat context manipulation provided a social motivation for under-informative answers. Specifically, Face-threatening scenarios involved dyadic interactions in which the speaker addressed the potentially affected individual directly, thereby creating face-threatening pressure (e.g. Mei performed a piece at the evening party. Afterward, she asked her friend how many parts of the piece she had played wrong. Her friend heard that all parts were off, but not wanting Mei to lose face.), whereas Neutral scenarios described third-party situations in which the speaker reported information about others, with no direct interpersonal stakes (e.g. Mei and her friend went to a show where they heard an audience member perform a piano piece. Afterward, she asked her friend how many parts of the piece were played wrong. Her friend heard that all parts were played incorrectly.). The Factual State manipulation set the ground truth of the scenario: in All contexts, 100% of the group possessed a certain property or did a certain action; in Most contexts, the relevant property held for a majority of cases (more than half, but not all). After reading the context, participants saw an incomplete sentence such as “[Context]. When asked, the person replied: ‘___ (of the X did Y).’” Seven quantifier options were available for the blank, spanning a range of informational strengths. These included all (全部) and none (没有) as fully informative endpoints; most (多数) and a few (少数) as intermediate quantifiers; more informationally specific alternatives relative to most/few—namely the great majority (大多数) and very few (极少数); and the vague, pragmatically flexible quantifier some (一些). Participants were instructed to select the quantifier that best completed the sentence given the context. Thus, the primary dependent variable was the choice of quantifier. For analysis, we focused on whether the quantifier “some” was chosen (binary outcome: some vs. other quantifier). This choice of “some” is critical because “some” implicates “not all,” making it a key indicator of pragmatic under-informativeness. Each scenario was instantiated in four versions crossing Face Threat and Factual State. To ensure that each participant saw only one version of each scenario, we distributed the four versions of all scenarios across four material lists using a Latin-square design. Each list contained exactly one condition version of each scenario, and participants were randomly assigned to one list. To balance the a priori plausibility of the seven quantifier options and to obscure the critical manipulation, we embedded two types of filler trials: 12 fillers designed to make each option a factually licensed and natural response at least some of the time, and 14 descriptive fillers that did not involve interpersonal considerations, thereby disrupting the target contextual structure. The 32 critical trials and 26 filler trials were pseudo-randomized such that no more than two trials from the same experimental condition occurred consecutively, mitigating potential order effects. The same 32 critical scenarios were used for LLM evaluation. All four versions of each scenario were presented to each model. To estimate response variability, each stimulus was independently sampled 10 times per model with different random seeds, yielding a total of 1,280 observations per model. All trials were fully randomized. Each trial was run in isolation, with no conversational history or cross-trial context retained between prompts. Experiment 2: Pragmatic Judgment. This experiment used a 2 × 2 × 2 within-subjects design. The independent variables were Face Threat (face-threatening/neutral), Factual State (Most vs. All), and Statement Type (“Some” statement vs. fully informative “Factual” statement). Materials consisted of brief scenarios and a target statement to be evaluated. Each scenario set up a context similar to Experiment 1. Following the scenario, participants were presented with a target sentence that was either an under-informative “some” statement or a fully informative factual statement. For example, in a context where indeed all members of a group did something, the “Some” statement might present the sentence “ Some of the Y did X .” (e.g. Some were played wrong.); the “Factual” statement in that context would present “ All of the Y did X ” (e.g. All were played wrong.). In a Most context, the factual statement would be “ Most of the Y did X .” Participants were instructed to judge whether the given statement was an acceptable or appropriate description of the situation, given the context. They responded “Yes” (acceptable) or “No” (unacceptable) for each statement. The primary dependent variable was the binary acceptability judgment for the “some” statement – i.e., whether an under-informative “some” was accepted as an appropriate utterance. By design, fully informative statements in factual conditions are truthfully correct; we expected those to be generally accepted, serving as a control. Each participant encountered all combinations of the three factors across trials. For instance, each participant might see a scenario of each type: Face Threat × Factual State, each with both a “some” version and a “factual” version of the statement, for a total of 8 trial types. There were same 32 unique scenarios in total, each participant seeing each scenario in one of its possible forms. Trials were randomized and balanced to prevent any ordering or priming effects. As in Experiment 1, the same scenarios were used for LLM evaluation; with all four versions of each scenario presented and each item sampled 10 times, this yielded a total of 2,560 observations per model. Procedure and Measures Human Behavioral Procedure. Participants were tested using a computer-based task programmed in MATLAB 2022b (MathWorks Inc.). Experiment 1 (Production task) and Experiment 2 (Judgment task) were run as separate sessions with different participant groups. All instructions and materials were presented in Chinese. Participants were first given on-screen instructions explaining the task with examples. For Experiment 1, participants were told they would read scenarios and complete sentences with an appropriate quantifier. They practiced with 20 of example scenarios to familiarize themselves with the seven quantifier options. During the actual trials, each scenario text was presented on the screen. After reading the scenario, participants pressed the space bar to proceed to the response screen. Below the question, the seven quantifier options were displayed and mapped to the number keys 1–7; their left–right positions were randomized on each trial, so the option–key mapping varied across trials. Participants selected the quantifier that best completed the sentence. They were instructed to respond as naturally and accurately as possible, and that there were no strictly “correct” answers but some might be more appropriate given the context. Once a selection was made, the next trial began. For Experiment 2, participants were instructed that they would read a scenario and then see a statement, and they should judge whether saying that statement in that context was acceptable or not. Trials began with the context description, followed by the target statement. Participants pressed one of two keys to record their judgment (e.g., “J” = acceptable, “F” = unacceptable); for the remaining half of participants, this key–response mapping was reversed to counterbalance motor-response associations across participants. They were encouraged to rely on their pragmatic intuition about whether the response was appropriate, not just whether it was true. Participants completed a brief practice for Experiment 2 as well, to ensure they understood the concept of pragmatic acceptability. Reaction times (RTs) were recorded for each trial in Experiment 2 in two components. Contextual reading reaction time is measured from the beginning to the end of a contextual short story. Judgment RT was measured from target-statement onset to the keypress indicating the participant’s acceptability judgment. The entire session lasted approximately 30 minutes per experiment. LLM Evaluation Procedure. Each experiment’s stimuli were presented to the language models in a parallel manner. We constructed standardized prompts in Chinese to ensure the models received the same information as human participants. For Experiment 1, the prompt for each trial consisted of the scenario description followed by an incomplete sentence. Generation was performed with a sampling temperature set to 1.0. LLM responses were collected using the PsyLingLLM framework, which allows systematic logging of model outputs and associated generation metadata. For the non-CoT models (DeepSeek-V3.2 and GPT-4o), each trial yielded a single-token answer corresponding to the selected quantifier. For these models, we recorded the chosen answer, the overall response latency (from prompt submission to token return), and the log probability associated with the returned output token. For the CoT model (DeepSeek-R1), token-level log probabilities were not available. Instead, we separately extracted the final answer quantifier and the accompanying chain-of-thought reasoning trace. For each CoT trial, we recorded the total elapsed time from prompt submission to final response, as well as the length of the chain-of-thought in tokens, which served as an index of reasoning extent. For Experiment 2, the LLM evaluation procedure closely paralleled that of Experiment 1. Each trial prompt consisted of the scenario description followed by a target statement and a simple binary acceptability question (yes/no), mirroring the human judgment task. The target statement was either an under-informative some statement or a fully informative factual statement, depending on condition. The full Chinese prompts and example materials used for LLM evaluation are provided in Supplementary Materials S1 and S2. Statistical analysis We analyzed the data using regression models that respect the repeated-measures structure of the experiments. All analyses were conducted in R (version 4.3.2). Human data were analyzed with mixed-effects models implemented in lme4 (Bates et al., 2015) to account for repeated observations from the same participants and repeated use of the same items. For LLM data, because there is no participant-level sampling and each item was evaluated repeatedly via independent API calls, inference was conducted at the item level using generalized linear/linear models with item fixed effects and item-clustered robust standard errors. For choice-level outcomes, we analyzed binary responses with logit-link models. In Experiment 1, the dependent variable was whether the response option corresponding to some was selected on each trial. In Experiment 2, the dependent variable was whether the statement was judged acceptable. Fixed effects included the experimental manipulations—Face Threat, Factual State, and Statement Type—and their factorial interactions. For human data, we fit generalized linear mixed-effects models (GLMMs; binomial family with logit link) with crossed random intercepts for participant and item; where counterbalanced lists were used, Version was included as a fixed-effect control (Face Threat × Factual State × Statement Type + Material Version, with crossed random intercepts for Subject and Item). For LLM data, each model produced multiple independent responses per item×condition cell (10 API calls; temperature = 1), so we avoided treating calls as independent observations. Instead, we aggregated each item×condition (and statement type, when applicable) to a binomial count (k successes out of n = 10) and fit binomial-logit generalized linear models (GLMs) including item fixed effects to control for item-specific baselines; statistical inference used item-clustered robust standard errors (sandwich estimator), and we report odds ratios (OR), 95% confidence intervals, and Wald statistics for the focal effects. To quantify distributional similarity between humans and models in the full response-option distributions, we also computed Jensen–Shannon (JS) divergence between the empirical human distribution and each model distribution within each condition. Confidence intervals were obtained via cluster bootstrap resampling that respected the dependence structure. JS divergence was treated as descriptive evidence of distributional similarity and does not replace the regression-based inference. For continued cost-level outcomes in experiment 2, human processing costs were operationalized using reaction times at the context-reading stage (ContextRT) and the judgment stage (JudgmentRT). To reduce positive skew, RTs were log-transformed using the natural logarithm (log), after excluding non-positive observations. We then fit linear mixed-effects models of the form log(RT) ~ Face Threat × Factual State × Statement Type + Material Version, with crossed random intercepts for Subject and Item. Fixed-effect estimates are reported as regression coefficients (b) with 95% confidence intervals, and inference relied on Satterthwaite-approximated degrees of freedom as implemented in lmerTest. For LLM data, to maintain item-level inference and avoid pseudo-replication across repeated API calls, we aggregated each proxy to the Item × Face Threat × Factual State × Statement Type level by averaging across the 10 calls. We then fit linear models including the full experimental interaction and item fixed effects (CostValue ~ Face Threat × Factual State × Statement Type + Item). Estimates are reported as b with 95% confidence intervals; inference used item-clustered robust standard errors, with Wald t tests computed using degrees of freedom equal to the number of item clusters minus one. Because human RTs and model proxies are expressed in fundamentally different units, we did not compare absolute magnitudes across systems; instead, we compared the pattern of experimental effects within each system. Finally, for cost–acceptance alignment, we conducted condition-wise, item-level correlations restricted to “Some” statements. For humans, we first computed item × Face Threat × Factual State cell means for log-transformed RT, and mean acceptability, ensuring inference was driven by between-item variation rather than trial-level pseudo-replication. For LLMs, we computed the corresponding item × Face Threat × Factual State cell means for each model’s cost proxy. Correlation uncertainty was summarized with 95% confidence intervals derived via Fisher’s z transformation. All scripts for preprocessing and analysis are available in the project repository (OSF link). Results CoT reasoning yields more human-like mitigation profiles in the production task In the completion task (Experiment 1), Human participants exhibited a clear pragmatic adaptation in their quantifier use (Fig. 2 , Table S1 ). In neutral contexts (no interpersonal stakes), utterances almost always provided fully informative answers: for instance, when all individuals in the scenario had done the action, nearly every participant chose all (全部), and when only most had done it, most (多数) was overwhelmingly used. In contrast, Face-threatening contexts (a context pointing out that “not all did it” would embarrass someone) elicited far more under-informative responses. In these scenarios, utterances often avoided the stronger quantifiers even when they were true. The vague quantifier some (一些) became the single most frequent choice, accounting for roughly 44.53% and 46.88% of responses in both Most and All contexts with face threat. For example, when in reality all members of a group had done the task, only a minority of utterances actually said “ all did it”; nearly half responded with “ some (did it)”. The remaining responses in Face-threatening conditions were spread across other weaker terms – e.g. few (少数) or even none (没有) – effectively understating the outcome to avoid embarrassment to any individual. Consequently, fully informative quantifiers were rarely used in Face-threatening contexts (1.82% most responses in Most contexts, and 2.08% all responses in All contexts; compared with 81.51% and 83.60%, respectively, under the neutral condition). This distributional shift confirms that human utterances often sacrificed informativeness for politeness, frequently opting for “some” instead of a more informative term when a truthful answer would be socially awkward, replicates the classic finding that face threat blocks the usual scalar inference. The LLMs’ responses showed qualitatively similar context effects, but with notable differences in degree. All three models learned to use some more often in Face-threatening situations, but the extent of this shift varied. Notably, DeepSeek-V3.2 in Face-threatening contexts produced some almost categorically – for instance, 93.75% of its responses were some in the Most condition, and over 95.63% in All condition. This exceeded the human rate and indicates an exaggerated politeness bias in DeepSeek-V3.2. GPT-4o also displayed a similar over-reliance on some under face threat (92.50% and 73.43% of responses some ) and in fact produced less varied output than humans in those conditions. By contrast, DeepSeek-R1 more closely mimicked the human distribution: in Face-threatening contexts it chose some on roughly one-third to one-half of trials (32.18% in Most , 46.88% in All ). Beyond some frequency, it occasionally used other mitigating quantifiers ( few , etc.), yielding a broader response spread. Quantitatively, DeepSeek-R1’s full response distributions closely matched those of humans, with low JS divergence (0.038 and 0.048 in the Most and All conditions, under face-threatening contexts) and relatively high entropy (1.04 and 1.37, compared with human values of 1.30 and 1.52, see Table S3). Whereas GPT-4o and especially DeepSeek-V3.2 diverged more. For DeepSeek-V3.2, the JS divergence increased to 0.18 and 0.19 in face-threatening contexts (under the Most and All conditions, respectively), alongside low entropy values (0.36 and 0.27), indicating a near-deterministic preference for some . GPT-4o with JS divergence values of 0.11 and 0.17 and entropy levels of 0.39 and 1.02 under face-threatening conditions, indicating a strong bias toward “some” while retaining slightly variance. In Neutral contexts, all systems were near-ceiling on the contextually appropriate quantifier (≥ 81.56%) and thus closely matched human behavior (JS ≤ 0.05). We modeled the binary likelihood of choosing some (vs. any other quantifier) using logistic regression (Table 1 ). The human GLMM data revealed significant main effects of both experimental factors and their interaction. Specifically, there was a huge effect of Face Threat: utterances were far more likely to choose some in face-threatening situations than in neutral ones (OR = 12.80, 95% CI [8.17, 20.20], p < .001). There was a reliable effect of the Factual State: overall, scenarios where only Most of the group had done the action led to fewer some responses than scenarios where All had done it (OR = 0.35, 95% CI [0.18, 0.68], p = .002). In other words, when factuality is an All condition, participants were somewhat less inclined to hedge with some. We also observed a significant Face Threat × Factual State interaction (OR = 3.23, 95% CI [1.53, 6.82], p = .002), indicating that the effect of Face Threat on the likelihood of choosing some was stronger in the All context than in the Most context. The LLMs exhibited the same qualitative effects pattern in the GLM analysis, with extreme coefficients consistent with near-ceiling shifts. DeepSeek-V3.2 (OR = 2.03 × 10²⁰, 95% CI [5.86 × 10¹⁹, 7.04 × 10²⁰], p < .001) showed an extremely large estimated odds ratio for Face Threat, indicating that under face-threatening conditions it almost never produced other responses when some was licensed. A similarly amplified Face Threat effect was observed for GPT-4o (OR = 5.16 × 10¹⁹, 95% CI [1.03 × 10¹⁹, 2.57 × 10²⁰], p < .001), and a more moderate but reliable effect for DeepSeek-R1 (OR = 73.10, 95% CI [8.46, 632], z = 3.90, p < .001). All three models showed dramatically reduced odds of producing some in the Most relative to All contexts under the current coding scheme (DeepSeek-V3.2: OR = 2.08 × 10⁻¹⁸, 95% CI [2.16 × 10⁻²¹, 2.01 × 10⁻¹⁵], p < .001; GPT-4o: OR = 1.10 × 10⁻⁸, 95% CI [1.92 × 10⁻⁹, 6.33 × 10⁻⁸], p < .001; DeepSeek-R1: OR = 3.68 × 10⁻⁸, 95% CI [3.76 × 10⁻⁹, 3.60 × 10⁻⁷], p < .001), consistent with near-deterministic sensitivity to the underlying factual state. Critically, a significant Face Threat × Factual State interaction was observed for all systems, indicating that the licensing effect of face threat on some production depended on whether the true state was All or Most (DeepSeek-V3.2: OR = 1.12 × 10¹⁸, 95% CI [6.20 × 10¹⁶, 2.03 × 10¹⁹], p < .001; GPT-4o: OR = 1.08 × 10⁷, 95% CI [7.18 × 10⁵, 1.63 × 10⁸], p < .001; DeepSeek-R1: OR = 5.85 × 10⁷, 95% CI [7.29 × 10⁶, 4.69 × 10⁸], p < .001). Collectively, Experiment 1 establishes a clear human–LLM contrast: face threat elicits a systematic trade-off between informativeness and social appropriateness in production. Humans markedly reduced fully informative quantifiers and increased under-informative responses—most prominently some, but also a broader downward-weakening pattern among non-some alternatives—indicating a general mitigation strategy under interpersonal risk. GPT-4o and DeepSeek-V3.2 often realized mitigation in an almost deterministic, some-centric way, thereby increasing divergence from human response distributions. By contrast, DeepSeek-R1 most closely approximated human behavior, both in overall distributional similarity and in exhibiting a human-like weakening strategy beyond some. These results suggest that DeepSeek-R1 may deploy a dual mitigation repertoire like human participants—vagueness plus systematic weakening among non-some choices—whereas the non-CoT models rely predominantly on a vagueness-dominant strategy. Face Threat Increases Acceptance of Some in Judgment Task In the judgment task (Experiment 2), we tested whether listeners judge under-informative statements to be acceptable descriptions given the context (Table S2). Figure 3 a gray bar displays acceptance rates for target statements as a function of face-threat and factual state. Human judgments revealed a mirror-image of the production results: acceptability of a some statement (which is true but not maximally informative) was low in neutral contexts but rose dramatically in face-threatening contexts. Specifically, in the absence of any face threat, participants largely dispreferred a some statement if a more informative statement was warranted – especially in the All context. When indeed all members had done the action, saying “ Some of Y did X ” was often viewed as infelicitous or odd (mean acceptability 35.94%). Even when the truth was most did it, a some statement was accepted only about 57.81% of the time without a face motive, reflecting a moderate penalty for under-informativeness. However, when face is threatened, listeners become more tolerant of the use of the word some . In the face-threatening scenarios, the vast majority of participants found the under-informative some statement acceptable. Acceptance of some statements jumped to about 94.27% in the Most context and 93.23% in the All context when a face threat was present. Thus, humans strongly contextualized their pragmatic judgments: a sentence like “Some of them succeeded,” which would ordinarily sound infelicitous if in fact all of them succeeded, becomes entirely appropriate when the speaker has a polite reason to remain vague. Figure 3 a illustrates this crossover: the Face Threat effect on some -statement acceptability is huge, whereas the factual context effect ( All vs. Most ) is smaller. A complementary pattern emerged for fully informative factual statements (Fig. 3 b). In neutral contexts, these statements were endorsed at high rates ( All : 88.02% and Most : 93.22%), consistent with their truth and informativeness. Under face threat, however, participants often judged the same maximally informative utterances as pragmatically inappropriate. In face-threatening Most contexts, only about 18.75% accepted “ Most of Y did X ,” despite its factual correctness, plausibly because it foregrounds that not everyone succeeded. Likewise, in face-threatening All contexts, “ All of Y did X ” was accepted by only 10.94% of participants, suggesting that even correct maximal informativeness can be dispreferred when it conflicts with interpersonal considerations. These findings imply a strong interaction: face threat substantially increases the acceptability of under-informative some statements, while decreasing the acceptability of fully informative statements that may be socially insensitive. A logistic mixed-effects model confirmed these patterns. Across all trials, there were significant main effects of Face Threat (face-threatening vs. neutral: OR = 14.90, 95% CI [7.42, 30. 00], z = 7.59, p < .001), Factual State ( Most vs. All : OR = 0.36, 95% CI [0.23, 0.56], p < .001), and Statement Type ( some vs. factual : OR = 13.5, 95% CI [6.87, 26.5], p < .001). Critically, Face Threat interacted strongly with Statement Type (OR = 0.00066, 95% CI [0.00023, 0.0019], p < .001), capturing the observed crossover: face threat increased acceptance of some while decreasing acceptance of fully informative statements. To make this reversal transparent, we also modeled the two statement types separately. For some statements, face threat substantially increased acceptance (OR = 25.10, 95% CI [11.20, 56. 00], p < .001) and acceptance was lower in Most than All contexts (OR = 0.27, 95% CI [0.16, 0.46], p < .001), with a modest Face Threat × Factual State interaction (OR = 2.85, 95% CI [1.01, 8.00], p = .047). For factual statements, face threat sharply reduced acceptance (OR = 0.0043, 95% CI [0.0017, 0.011], p < .001), while Factual State had a smaller effect (OR = 0.44, 95% CI [0.20, 0.97], p = .041), and Face Threat × Factual State was not significant (OR = 0.99, 95% CI [0.35, 2.80], p = .989). Together, these models show that human participants’ overall three-way pattern are driven primarily by the selective licensing of under-informativeness under face threat, and the selective penalization of blunt informativeness in the same contexts. All three LLMs were partially converged with, but also differentiated from, this human profile (Fig. 3 a). Across systems, the Face Threat manipulation affected the acceptability of some in the same direction as in human judgments; however, the models differed markedly in calibration, both in their baseline tolerance for under-informativeness in neutral contexts and in the magnitude with which face threat licensed vagueness. DeepSeek-R1 exhibited the same directional profile but with a markedly stricter baseline and a larger face shift: it accepted some rarely in neutral contexts ( Most 22.50% and All 7.81%) while approaching ceiling under face threat ( Most 99.38% and All 99.06%), indicating a strong informativeness expectation unless vagueness is pragmatically licensed. GPT-4o likewise increased acceptance under face threat but was substantially more permissive in neutral contexts, especially in the Most condition ( Most 88.44% and All 44.06%), yielding a comparatively smaller face-related increase for Most but a pronounced increase for All ( Most 99.69% and All 100%). DeepSeek-V3.2 also showed higher acceptance under face threat, yet the modulation was weaker and remained far below the human ceiling ( Most 54.06% and All 15.63% for neutral contexts; Most 65.00% and All 43.13% for neutral contexts). Overall, the descriptive patterns suggest that face threat licenses some across systems, but models vary in how strongly they enforce informativeness in neutral contexts and how fully they capitalize on face-based licensing in face-threatening contexts. For fully informative statements, the models showed only partial pragmatic sensitivity. Like humans, they accepted factual statements at ceiling in neutral contexts (all three models: 100% in Most and 100% in All ). Under face threat, however, the models remained far more tolerant than humans of these blunt, maximally informative utterances. Whereas GPT-4o still accepted them at high rates ( Most 70.00% and All 89.38%), and DeepSeek-R1 likewise remained largely accepting ( Most 76.56% and All 96.56%). DeepSeek-V3.2 showed the largest drop among the models, yet still exceeded human acceptance ( Most 82.81% and All 52.19%). Thus, although the models captured that face threat licenses vagueness, they did not reproduce the human tendency to treat full informativeness itself as socially inappropriate in face-threatening contexts. Because LLMs produced near-deterministic responses for the fully informative statements (acceptance rates reached ceiling), the full Face Threat × Factual State × Statement Type model exhibited quasi-complete separation (full three-way results see Table 2 a), yielding unstable or effectively infinite estimates for the three-way interaction. To ensure interpretable inference and to target the theoretically critical contrast—whether face threat licenses under-informative utterances—we therefore focus our model-based analyses on trials with some statement type only (Table 2 b). In the some -only models, Face Threat robustly increased acceptance for GPT-4o (OR = 73.90, 95% CI [8.18, 668], p < .001), and for DeepSeek-R1 (OR = 2.24 × 10 19 , 95% CI [4.20 × 10 18 , 1.19 × 10 20 ], p < .001), whereas DeepSeek-V3.2 did not show a reliable Face Threat main effect for some (OR = 2.11, 95% CI [0.78, 5.71], p = .141). Factual State also reliably reduced some acceptance across systems (GPT-4o: OR = 0.013, 95% CI [0.0022, 0.073], p < .001; DeepSeek-R1: OR = 0.22, 95% CI [0.13, 0.37], p < .001; DeepSeek-V3.2: OR = 0.031, 95% CI [0.0097, 0.099], p < .001). Finally, Face Threat × Factual State for some was non-significant for DeepSeek-R1 (OR = 2.93, 95% CI [0.45, 19.10], p = .259), and pronounced for GPT-4o and DeepSeek-V3.2 (GPT-4o: OR = 1.33 × 10 9 , 95% CI [8.56 × 10 7 , 2.06 × 10 10 ], p < .001; DeepSeek-V3.2: OR = 7.36, 95% CI [2.18, 24.80], p = .001). We next asked whether processing cost tracks the pragmatic trade-off between informativeness and face management. We analyzed log-transformed response times (RTs) separately for the context-reading stage (ContextRT) and the judgment stage (JudgmentRT) as a function of Face Threat, Factual State, and Statement Type. For LLMs, we used parallel cost proxies: log-thinking tokens for DeepSeek-R1 and log-probability for non-CoT models DeepSeek-V3.2 and GPT-4o. As a result, Human ContextRT showed little systematic pragmatic modulation: face threat produced only a marginal trend toward faster reading (b = − 0.091, 95% CI [− 0.19, 0.006], p = .066), with no reliable higher-order interactions. This absence of early-stage effects suggests that the pragmatic manipulation does not primarily alter initial comprehension demands. Instead, the cost signature emerged downstream at the evaluative stage. In JudgmentRT, participants judged statements faster under face-threatening (b = − 0.47, 95% CI [− 0.59, − 0.35], p < .001), and—at baseline—factual responses were judged faster than some (b = − 0.38, 95% CI [− 0.498, − 0.262], p < .001), consistent with some incurring conflict when full informativeness is expected. Critically, this factual advantage was substantially reduced under face-threatening (b = 0.44, 95% CI [0.28, 0.61], p < .001), indicating that face threat licenses vagueness and attenuates the processing penalty for under-informativeness. Moreover, a significant three-way interaction showed that the magnitude of this reweighting depends on whether the underlying state was All versus Most (b = − 0.29, 95% CI [− 0.53, − 0.055], p = .016). Turning to LLM-based cost proxies, the CoT model DeepSeek-R1 showed a highly interaction-rich token-cost profile. DeepSeek-R1’s thinking token cost was lower overall under face threat (b = − 0.28, 95% CI [− 0.38, − 0.17], p < .001) and when the true state was All (b = − 0.094, 95% CI [− 0.15, − 0.036], p = .003), and factual responses required substantially fewer tokens at baseline (b = − 0.75, 95% CI [− 0.85, − 0.64], p < .001). Crucially, however, this baseline “factual-is-easy” advantage was strongly context-dependent: face threat markedly attenuated the token advantage for factual responses (Face Threat × Statement Type: b = 0.55, 95% CI [0.37, 0.74], p < .001), and this modulation further varied with the All/Most distinction (Face Threat × Factual State × Statement Type: b = 0.51, 95% CI [0.34, 0.67], p < .001). Notably, DeepSeek-R1 also showed reliable two-way interactions involving Factual State (Face Threat × Factual State: b = 0.096, 95% CI [0.0070, 0.19], p = .036; Factual State × Statement Type: b = − 0.14, 95% CI [− 0.23, − 0.042], p = .006), underscoring that its explicit reasoning length is jointly shaped by informational state and social context. These effects suggest that when maximal informativeness can be socially risky, the CoT model’s reasoning expands from a default informativeness-driven mode toward more conditional, context-sensitive justification. By comparison, DeepSeek-V3.2’s log-probability primarily reflected an overall preference for factual responses (b = 0.27, 95% CI [0.14, 0.40], p < .001), with no reliable Face Threat, Factual State, or interaction effects. GPT-4o, in contrast, showed a targeted pragmatic reallocation: although both face threat (b = 0.17, 95% CI [0.077, 0.26], p < .001) and factual type increased log-probability (b = 0.17, 95% CI [0.082, 0.26], p < .001), face threat sharply reduced the relative log-probability advantage of factual responses (Face Threat × Statement Type: b = − 0.34, 95% CI [− 0.47, − 0.20], t = − 4.96, p < .001). This interaction indicates that face threat changes which type of response GPT-4o favors, rather than uniformly affecting response confidence. Finally, to assess whether humans and models align at the level of item-by-item variation, we computed within-cell correlations across items between human measures and model cost proxies for some statements (Fig. 4 , Table S4). Overall, most correlations were small and non-significant, indicating limited fine-grained alignment beyond condition-level means. Nevertheless, three localized correspondences emerged. In face threat neutral contexts, human acceptance of some correlated positively with DeepSeek-R1 token cost in both the Most condition ( r = 0.44, 95% CI [0.11, 0.68], p = .011) and the All condition ( r = 0.47, 95% CI [0.14, 0.70], p = .007). In addition, in the face-threatening All condition, human JudgmentRT correlated negatively with GPT-4o log-probability ( r = − 0.51, 95% CI [− 0.73, − 0.20], p = .003), suggesting shared item-level difficulty in the most pragmatically conflicted context. Given the number of tests conducted, we treat these correlations as exploratory. Taken together, the acceptance-rate, cost, and alignment analyses converge on a conclusion: LLMs capture the direction of human pragmatic adaptation under face threat, but remain miscalibrated in the social stakes of informativeness and differ in how their internal cost signals resemble human evaluative processing. At the behavioral level, human acceptance rates exhibited a clear crossover: in neutral contexts, listeners penalized under-informative some when a stronger alternative was warranted—most strongly when the underlying state supported maximal informativeness—whereas under face threat, the same some statements became broadly acceptable. In complementary fashion, maximally informative factual statements were readily endorsed in neutral contexts but were strongly downgraded as pragmatically inappropriate under face threat, indicating that humans treat full informativeness itself as socially costly when it foregrounds a potentially embarrassing implication. Against this benchmark, all three models increased acceptance of some under face threat, but they did so with sharply different calibration profiles. DeepSeek-R1 implemented a stringent informativeness prior in neutral contexts and then shifted toward a human like near-ceiling acceptance under face threat. GPT-4o showed the opposite calibration: it was already permissive in neutral contexts, and thus exhibited a smaller face-driven increase in those cells while still moving toward ceiling under face threat. DeepSeek-V3.2 displayed the weakest contextualization, with only modest increases under face threat and acceptance rates remaining far below the human ceiling. Critically, none of the models reproduced the human-level rejection of blunt factual utterances in face-threatening settings; all three remained substantially more tolerant of maximally informative statements even when they are socially insensitive. Consequently, the systems capture “face threat licenses vagueness” more readily than the complementary human norm that “face threat penalizes blunt truth.” The cost analyses provide converging process-level constraints. In humans, processing cost was concentrated at the judgment stage rather than during context reading, consistent with a late evaluative reweighting of informativeness against interpersonal considerations. DeepSeek-R1 exhibited a strongly context-conditioned token-cost profile, consistent with engaging in explicit justification jointly shaped by informational state and interpersonal risk; however, its higher-order interaction pattern need not mirror human RT structure, given that CoT length and decision time are different computational currencies. GPT-4o showed a more targeted cost signature consistent with probability mass reallocating toward some specifically when face threat makes mitigation relevant, whereas DeepSeek-V3.2 showed comparatively weak contextual modulation. Finally, exploratory item-level correlations indicate that any human–model coupling is localized: where correspondence appears, it tends to surface in theoretically diagnostic regions of the design rather than reflecting a global item-by-item match. Overall, in the judgement task, the models approximate the human licensing of under-informativeness under face threat, but they do not fully reproduce the human tendency to treat maximal informativeness itself as socially costly, highlighting a gap between pragmatic directionality and socially calibrated evaluative norms. Discussion Scalar expressions such as some are traditionally analyzed through the lens of informativity-based competition with stronger alternatives, yielding scalar implicatures like not all (Grice, 1975; Levinson, 2000). At the same time, everyday communication routinely exploits some as a face-saving device, allowing speakers to soften potentially awkward or socially risky truths (Brown & Levinson, 1987). By leveraging this dual function, the present study examined whether humans and LLMs exhibit similar context-dependent reweighting of informativeness and politeness, and whether any surface-level alignment reflects shared underlying evaluative mechanisms. Across two experiments, our results converge on a clear asymmetry. Humans systematically modulate both production and judgment in response to face threat, treating informativeness not as an absolute requirement but as a socially negotiable norm. LLMs, in contrast, capture the direction of this pragmatic adaptation, most notably, the increased tolerance for under-informative expressions under face-threatening conditions, but fail to reproduce its full social logic. Informativeness–Politeness Tradeoff in Quantifier Use Experiment 1 showed that human speakers flexibly balance informational accuracy against interpersonal concerns. In contexts without social stakes, speakers overwhelmingly adhered to informativeness, producing quantifiers that precisely matched the factual state. When face was threatened, however, they systematically weakened their utterances, most prominently by deploying some but also by shifting toward other weaker quantifiers. This pattern reflects a classic pragmatic insight: speakers may voluntarily forgo informativeness when stating the full truth risks embarrassing the addressee (Bonnefon et al., 2009; Feeney & Bonnefon, 2013). The LLMs exhibited only partial convergence with this human strategy. While all models increased their reliance on some under face threat, their mitigation strategies were markedly narrower. Non-CoT models largely collapsed onto a single vague option, whereas the CoT model DeepSeek-R1 more closely resembled humans in distributing responses across a range of weaker quantifiers. This suggests that explicit reasoning mechanisms can support more graded adjustments, but even then without the full diversity observed in human speech. Implicature Suppression and the Social Cost of Directness Judgment data from Experiment 2 extend this picture by showing that humans treat informativeness itself as socially evaluable. When no face threat was present, under-informative some statements were dispreferred, consistent with standard expectations of cooperative informativeness. Under face threat, however, the same statements became broadly acceptable, reflecting listeners’ willingness to suspend scalar expectations to preserve interpersonal harmony. Crucially, this tolerance was paired with a complementary effect: fully informative statements, though factually correct, were judged inappropriate when they foregrounded an embarrassing implication. Human listeners thus appear sensitive not only to when vagueness is licensed, but also to when blunt truthfulness incurs social cost (Yoon et al., 2020). The LLMs showed only partial alignment with this evaluative pattern. All three models increased acceptance of some under face threat, indicating sensitivity to the idea that vagueness can be appropriate in socially delicate contexts. However, none of the models reliably penalized maximally informative statements in the same situations. In other words, the models captured the heuristic that “face threat licenses vagueness,” but failed to internalize the complementary human norm that “face threat can make full informativeness inappropriate.” This asymmetry suggests that current LLMs treat politeness primarily as an additive strategy—introducing hedging when warranted—rather than as a constraint that can actively disfavor directness. From Informativeness to Face Management: Mechanism-Level Divergence Across completion and judgment, the human data point to a unified mechanism in which scalar enrichment is not a fixed by-product of lexical alternatives, but a contextually gated inference that is reweighted by social stakes. In neutral contexts, participants behaved as predicted by informativity-driven accounts: weak utterances some were treated as pragmatically deficient when a stronger true alternative was available, consistent with Gricean competition and scalar strengthening. Under face threat, however, both speakers and listeners systematically shifted the weighting of constraints: speakers licensed underinformativeness as a face-saving hedge, and listeners suppressed the usual “not all” pressure and instead treated vagueness as socially appropriate. This pattern is consistent with broader evidence that scalar implicatures are effortful and context-sensitive rather than default and obligatory (Bott & Noveck, 2004), and with classic politeness accounts in which speakers manage face via strategic attenuation of commitment (Bonnefon et al., 2009; Brown & Levinson, 1987). The critical question is whether LLMs implement a similar contextually gated inference. Mechanistically, this predicts two coupled signatures under face threat: an increased tendency to produce and accept some, and a concomitant decrease in the perceived appropriateness of fully informative alternatives. The models reliably showed the first signature. They were far less reliable on the second—and this is arguably the harder part. Suppressing full informativeness is less overt and harder to cue than producing vagueness: it requires an implicit social norm that “saying the whole truth” can be dispreferred when it threatens face, even though it is factually correct. Consistent with this, the models often continued to endorse blunt factual statements as acceptable in face-threatening contexts, indicating limited sensitivity to the social penalty attached to directness. This pattern suggests that models may be learning stylistic mitigation without learning the deeper pragmatic norm that makes hedging rational for humans. More fundamentally, these patterns are consistent with the view that contemporary LLMs lack stable, causally grounded representations of interlocutors’ mental states and social incentives (Bender et al., 2021). Their apparent politeness adjustments plausibly reflect learned associations between linguistic forms and contextual cues in training data, rather than an underlying social-motivational model. As a result, when contextual cues are sparse or when interaction requires higher-order perspective-taking beyond the training distribution, models may revert to more literal, semantics-biased responses. This raises an important generalization challenge: whether surface-level pragmatic sensitivity observed in controlled, single-shot prompts can scale to real-world dialogue, where face concerns are often implicit, distributed across turns, and negotiated dynamically. On this view, current LLMs exhibit shallow social-pragmatic adaptation, but do not yet match the human capacity for deep, socially regulated pragmatic control. Chain-of-Thought Optimization and Pragmatic Flexibility The contrast between the CoT-based DeepSeek-R1 and the non-CoT DeepSeek-V3.2 is informative not because R1 merely produces more human-like outputs across production, judgment, and processing cost, but because it reveals how different computational regimes support pragmatic flexibility. Specifically, compared with non-CoT architectures, DeepSeek-R1 does not simply amplify shifts in output preferences; instead, it exhibits qualitatively different patterns in how contextual information constrains the generation process. Variation in reasoning length across contexts indicates that CoT generation adapts to how tightly the context constrains the decision space. When pragmatic considerations strongly favor one response type (e.g., face-threatening contexts that license vagueness), the model converges more quickly; when multiple responses remain plausible (e.g., neutral contexts prioritizing informativeness), longer reasoning trajectories are observed. This pattern is more like a property of search and constraint satisfaction in the generation process, rather than as evidence of human-like deliberation (De Varda et al., 2025). From this perspective, the advantage of CoT lies not in enhancing “reasoning” itself, but in providing intermediate representations that support multiple context-sensitive paths through the decision space. Importantly, this flexibility does not require assuming that the model explicitly represents social norms or reasons about face in a human-like way. Instead, explicit reasoning trajectories function as a structural scaffold through which informational and interpersonal cues can jointly constrain generation. In this sense, CoT optimization makes it easier for contextual factors to exert differentiated influence, yielding behavior that more closely approximates human pragmatic flexibility without implying shared underlying cognition (Wei et al., 2022). Limitations and Future Directions Our study has several limitations that suggest avenues for future work. First, our understanding of LLMs’ internal decision processes remains indirect. It will be important to develop better interpretability tools to see whether concepts like “face” or “implicature” are truly represented. Second, our experiments were conducted in Mandarin Chinese with contrived scenarios; pragmatic norms can vary across languages and cultures. It remains to be seen whether LLMs trained in other languages show similar or different patterns, and whether they can capture culture-specific politeness strategies. Third, real-world conversation is richer than our one-shot tasks. Humans use intonation, social context, and history to guide politeness; current LLMs mostly rely on a static prompt. Future work should test models in multi-turn dialogues or interactive settings, perhaps with simulated social feedback. Adding an explicit “theory of mind” component, or training on dialogues with nuanced interpersonal goals, could also narrow the gap. More broadly, our human–model comparison framework can be applied to other pragmatic phenomena. Just as LLMs serve as engineering tools, they are now testbeds for cognitive theories. By examining when models exhibit or fail at pragmatics, we can both refine our understanding of human language and identify ways to improve AI. Our results reinforce the view that human implicature is deeply influenced by social motives, whereas LLMs currently treat pragmatic effects more superficially. As language models continue to advance, we hope they will increasingly grasp when to be tactful and when to be direct – and perhaps even explain their choices. Achieving that level of social-intelligence in machines would not only enhance human–AI communication, but would also mark a significant step toward mirroring human pragmatic competence. Table 1 Experiment 1: regression results. System Effect OR 95% CI z p Human Face threat 12.80 [8.17, 20.2] 11.08 < .001 Human Factual state 0.35 [0.178, 0.684] -3.06 .002 Human Face threat × Factual state 3.23 [1.53, 6.82] 3.07 .002 DeepSeek-R1 Face threat 73.10 [8.46, 632.] 3.90 < .001 DeepSeek-R1 Factual state < 0.001 [3.760e-09, 3.600e-07] -14.71 < .001 DeepSeek-R1 Face threat × Factual state 5.85e + 07 [7.290e + 06, 4.690e + 08] 16.84 < .001 DeepSeek-V3.2 Face threat 2.03e + 20 [5.860e + 19, 7.040e + 20] 73.70 < .001 DeepSeek-V3.2 Factual state < 0.001 [2.160e-21, 2.010e-15] -11.61 < .001 DeepSeek-V3.2 Face threat × Factual state 1.12e + 18 [6.200e + 16, 2.030e + 19] 28.12 < .001 GPT-4o Face threat 5.16e + 19 [1.030e + 19, 2.570e + 20] 55.36 < .001 GPT-4o Factual state < 0.001 [1.920e-09, 6.330e-08] -20.54 < .001 GPT-4o Face threat × Factual state 1.08e + 07 [7.180e + 05, 1.630e + 08] 11.70 < .001 Note. OR = odds ratio. CI = 95% confidence interval. ORs are rounded to two decimals; very large ORs are shown in scientific notation. p values are two-tailed. Table 2 a Experiment 2: three-way regression result s. System Effect OR 95% CI z p Human Face threat 14.90 [7.42, 30.0] 7.58 < .001 Human Factual state 0.36 [0.230, 0.559] -4.52 < .001 Human Statement type 13.50 [6.87, 26.5] 7.55 < .001 Human Face threat × Factual state 2.29 [0.883, 5.92] 1.70 .089 Human Face threat × Statement type 6.59e-04 [0.000234, 0.00186] -13.88 < .001 Human Factual state × Statement type 1.31 [0.553, 3.12] 0.62 .536 Human Face threat × Factual state × Statement type 0.46 [0.119, 1.79] -1.12 .263 DeepSeek-R1 Face threat 1.32e + 03 [148., 1.180e + 04] 6.44 < .001 DeepSeek-R1 Factual state 0.25 [0.154, 0.400] -5.74 < .001 DeepSeek-R1 Statement type 9.63e + 09 [6.010e + 09, 1.540e + 10] 95.65 < .001 DeepSeek-R1 Face threat × Factual state 2.66 [0.471, 15.0] 1.11 .268 DeepSeek-R1 Face threat × Statement type 1.73e-11 [1.490e-12, 2.000e-10] -19.81 < .001 DeepSeek-R1 Factual state × Statement type 4.03 [2.50, 6.48] 5.74 < .001 DeepSeek-R1 Face threat × Factual state × Statement type 0.03 [0.00370, 0.279] -3.12 .002 DeepSeek-V3.2 Face threat 1.74 [0.833, 3.62] 1.47 .141 DeepSeek-V3.2 Factual state 0.11 [0.0500, 0.236] -5.61 < .001 DeepSeek-V3.2 Statement type 4.89e + 08 [2.670e + 08, 8.960e + 08] 64.72 < .001 DeepSeek-V3.2 Face threat × Factual state 3.11 [1.40, 6.90] 2.79 .005 DeepSeek-V3.2 Face threat × Statement type 6.29e-09 [2.710e-09, 1.460e-08] -43.89 < .001 DeepSeek-V3.2 Factual state × Statement type 9.21 [NA, NA] —— —— DeepSeek-V3.2 Face threat × Factual state × Statement type 0.05 [0.0378, 0.0798] -15.22 < .001 GPT-4o Face threat 50.90 [7.30, 356.] 3.96 < .001 GPT-4o Factual state 0.05 [0.0169, 0.148] -5.40 < .001 GPT-4o Statement type 1.61e + 08 [NA, NA] —— —— GPT-4o Face threat × Factual state 6.31e + 07 [8.040e + 06, 4.950e + 08] 17.09 < .001 GPT-4o Face threat × Statement type 1.36e-10 [2.150e-11, 8.660e-10] -24.08 < .001 GPT-4o Factual state × Statement type 20.00 [NA, NA] —— —— GPT-4o Face threat × Factual state × Statement type 3.28e-09 [NA, NA] —— —— Note. OR = odds ratio. CI = 95% confidence interval. ORs are rounded to two decimals; very large (≥ 1,000) or very small (< 0.01) ORs are shown in scientific notation. p values are two-tailed. Table 2 b Experiment 2: two-way regression results (Statement Type = some). System Effect OR 95% CI z p Human Face threat 25.10 [11.2, 56.0] 7.86 < .001 Human Factual state 0.27 [0.162, 0.457] -4.91 < .001 Human Face threat × Factual state 2.85 [1.01, 8.00] 1.98 .047 DeepSeek-R1 Face threat 2.24e + 19 [4.200e + 18, 1.190e + 20] 52.23 < .001 DeepSeek-R1 Factual state 0.22 [0.126, 0.371] -5.55 < .001 DeepSeek-R1 Face threat × Factual state 2.93 [0.452, 19.1] 1.13 .259 DeepSeek-V3.2 Face threat 2.11 [0.780, 5.71] 1.47 .141 DeepSeek-V3.2 Factual state 0.03 [0.00968, 0.0985] -5.88 < .001 DeepSeek-V3.2 Face threat × Factual state 7.36 [2.18, 24.8] 3.22 .001 GPT-4o Face threat 73.90 [8.18, 668.] 3.83 < .001 GPT-4o Factual state 0.01 [0.00217, 0.0727] -4.89 < .001 GPT-4o Face threat × Factual state 1.33e + 09 [8.560e + 07, 2.060e + 10] 15.02 < .001 Note. OR = odds ratio. CI = 95% confidence interval. ORs are rounded to two decimals; very large (≥ 1,000) or very small (< 0.01) ORs are shown in scientific notation. p values are two-tailed. Table 3 Experiment 2 cost indicators: regression results (effect labels corrected). System Cost Indicator Effect b CI t p Human ContextRT Intercept 8.191 [7.785, 8.598] 40.69 < .001 Human ContextRT Face threat -0.091 [-0.187, 0.006] -1.84 .066 Human ContextRT Factual state 0.062 [-0.035, 0.158] 1.26 .209 Human ContextRT Response type 0.079 [-0.017, 0.176] 1.61 .108 Human ContextRT Face threat × Factual state -0.011 [-0.148, 0.125] -0.16 .870 Human ContextRT Face threat × Response type -0.083 [-0.219, 0.054] -1.19 .234 Human ContextRT Factual state × Response type -0.094 [-0.23, 0.043] -1.35 .177 Human ContextRT Face threat × Factual state × Response type 0.026 [-0.167, 0.219] 0.27 .788 Human JudgmentRT Intercept 7.294 [6.967, 7.621] 44.88 < .001 Human JudgmentRT Face threat -0.47 [-0.588, -0.351] -7.79 < .001 Human JudgmentRT Factual state -0.046 [-0.164, 0.072] -0.76 .445 Human JudgmentRT Response type -0.38 [-0.498, -0.262] -6.3 < .001 Human JudgmentRT Face threat × Factual state 0.059 [-0.108, 0.226] 0.69 .488 Human JudgmentRT Face threat × Response type 0.444 [0.276, 0.611] 5.2 < .001 Human JudgmentRT Factual state × Response type 0.163 [-0.005, 0.33] 1.9 .057 Human JudgmentRT Face threat × Factual state × Response type -0.291 [-0.528, -0.055] -2.41 .016 DeepSeek-R1 logTokens Face threat 8.644 [8.353, 8.934] 59.97 < .001 DeepSeek-R1 logTokens Factual state -0.276 [-0.352, -0.2] -7.13 < .001 DeepSeek-R1 logTokens Response type -0.002 [-0.078, 0.074] -0.04 .967 DeepSeek-R1 logTokens Face threat× Factual state -0.093 [-0.169, -0.017] -2.41 .016 DeepSeek-R1 logTokens Face threat × Response type 0.042 [-0.065, 0.149] 0.77 .442 DeepSeek-R1 logTokens Factual state × Response type 0.112 [0.005, 0.22] 2.05 .040 DeepSeek-R1 logTokens Face threat × Factual state × Response type 0.014 [-0.093, 0.122] 0.26 .794 DeepSeek-V3.2 logprob Face threat -0.103 [-0.255, 0.049] -1.33 .183 DeepSeek-V3.2 logprob Factual state -0.275 [-0.376, -0.174] -5.55 < .001 DeepSeek-V3.2 logprob Response type -0.095 [-0.153, -0.036] -3.29 .003 DeepSeek-V3.2 logprob Face threat × Factual state -0.748 [-0.856, -0.64] -14.11 < .001 DeepSeek-V3.2 logprob Face threat × Response type 0.096 [0.007, 0.185] 2.2 .036 DeepSeek-V3.2 logprob Factual state × Response type 0.553 [0.366, 0.741] 6.01 < .001 DeepSeek-V3.2 logprob Face threat × Factual state × Response type -0.135 [-0.229, -0.042] -2.96 .006 GPT-4o logprob Face threat 0.507 [0.342, 0.672] 6.26 < .001 GPT-4o logprob Factual state -0.093 [-0.247, 0.061] -1.24 .226 GPT-4o logprob Response type 0.088 [-0.083, 0.259] 1.05 .303 GPT-4o logprob Face threat × Factual state 0.268 [0.138, 0.399] 4.19 < .001 GPT-4o logprob Face threat × Response type 0.004 [-0.196, 0.204] 0.04 .965 GPT-4o logprob Factual state × Response type -0.053 [-0.257, 0.152] -0.53 .603 GPT-4o logprob Face threat × Factual state × Response type -0.088 [-0.259, 0.083] -1.05 .303 Note. b = fixed-effect estimate. CI = 95% confidence interval. Face threat is coded as Yes vs No; Factual state is All vs Most. Response type contrasts factual vs some. p values are two-tailed. Declarations Additional Information The author declares no competing interests. Funding The author received no specific funding for this work. Author Contribution Yuying Wang solely conceived and designed the study, conducted the experiments, collected human behavioral data, simulated and analyzed LLM outputs, and wrote and revised the manuscript. Data Availability The datasets generated and analyzed during the current study are available at OSF [https://osf.io/m57jg/overview?view\_only=fc899334d57948e7b886da614e92688f], and can be accessed upon reasonable request. The data include human behavioral data, LLM output data, and analysis scripts used in this study. References Bates, D., Mächler, M., Bolker, B., & Walker, S. (2015). Fitting Linear Mixed-Effects Models Using lme4. Journal of Statistical Software , 67 (1). https://doi.org/10.18637/jss.v067.i01 Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21 , 610–623. https://doi.org/10.1145/3442188.3445922 Bonnefon, J.-F., Feeney, A., & De Neys, W. (2011). The Risk of Polite Misunderstandings. Current Directions in Psychological Science , 20 (5), 321–324. https://doi.org/10.1177/0963721411418472 Bonnefon, J.-F., Feeney, A., & Villejoubert, G. (2009). When some is actually all: Scalar inferences in face-threatening contexts. Cognition , 112 (2), 249–258. https://doi.org/10.1016/j.cognition.2009.05.005 Bott, L., & Noveck, I. A. (2004). Some utterances are underinformative: The onset and time course of scalar inferences. Journal of Memory and Language , 51 (3), 437–457. https://doi.org/10.1016/j.jml.2004.05.006 Brown, P., & Levinson, S. C. (1987). Politeness: Some universals in language usage (Vol. 4). Cambridge university press. Chemla, E., & Bott, L. (2014). Processing inferences at the semantics/pragmatics frontier: Disjunctions and free choice. Cognition , 130 (3), 380–396. https://doi.org/10.1016/j.cognition.2013.11.013 Chemla, E., & Singh, R. (2014). Remarks on the Experimental Turn in the Study of Scalar Implicature, Part I. Language and Linguistics Compass , 8 (9), 373–386. https://doi.org/10.1111/lnc3.12081 Cho, Y., & Kim, S. mook. (2024). Pragmatic inference of scalar implicature by LLMs (Version 1). arXiv. https://doi.org/10.48550/ARXIV.2408.06673 Cong, Y. (2024). Manner implicatures in large language models. Scientific Reports, 14(1), 29113. https://doi.org/10.1038/s41598-024-80571-3 De Varda, A. G., D’Elia, F. P., Kean, H., Lampinen, A., & Fedorenko, E. (2025). The cost of thinking is similar between large reasoning models and humans. Proceedings of the National Academy of Sciences , 122 (47), e2520077122. https://doi.org/10.1073/pnas.2520077122 Feeney, A., & Bonnefon, J.-F. (2013). Politeness and Honesty Contribute Additively to the Interpretation of Scalar Expressions. Journal of Language and Social Psychology , 32 (2), 181–190. https://doi.org/10.1177/0261927X12456840 Grice, H. P. (1975). Logic and conversation. Syntax and Semantics , 3 , 43–58. Guo, D., Yang, D., Zhang, Haowei, Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, Ruoyu, Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., … Zhang, Z. (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature , 645 (8081), 633–638. https://doi.org/10.1038/s41586-025-09422-z Horn, L. R. (1972). On the semantic properties of logical operators in English . University of California, Los Angeles. Horn, L. R. (2006). The Border Wars: A Neo-Gricean Perspective. In K. Von Heusinger & K. Turner (Eds.), Where Semantics meets Pragmatics (pp. 21–48). BRILL. https://doi.org/10.1163/9780080462608_006 Katsos, N., & Cummins, C. (2010). Pragmatics: From Theory to Experiment and Back Again. Language and Linguistics Compass , 4 (5), 282–295. https://doi.org/10.1111/j.1749-818X.2010.00203.x Ku, H. B. (2025). Scaling Implicature via Structured Interaction in Multi-Agent LLMs. In Proc. 1st Workshop on Integrating NLP and Psychology to Study Social Interactions at AAAI Int. Conf. Weblogs and Social Media (ICWSM) . In Proc. 1st Workshop on Integrating NLP and Psychology to Study Social Interactions at AAAI Int. Conf. Weblogs and Social Media (ICWSM). Levinson, S. C. (2000). Presumptive meanings: The theory of generalized conversational implicature . MIT press. Ma, B., Li, Y., Zhou, W., Gong, Z., Liu, Y. J., Jasinskaja, K., Friedrich, A., Hirschberg, J., Kreuter, F., & Plank, B. (2025). Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (arXiv:2502.12378). arXiv. https://doi.org/10.48550/arXiv.2502.12378 Qiu, Z., Duan, X., & Cai, Z. G. (2023). Pragmatic Implicature Processing in ChatGPT . PsyArXiv. https://doi.org/10.31234/osf.io/qtbh9 Shulginov, V., Şimşek, H. B., Kudriashov, S., Randautsova, R., & Shevela, S. A. (2025). Evaluating the Pragmatic Competence of Large Language Models in Detecting Mitigated and Unmitigated Types of Disagreement . 2025 . Wang, J. (2025). A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1 (arXiv:2502.10867). arXiv. https://doi.org/10.48550/arXiv.2502.10867 Wei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, brian, Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Advances in Neural Information Processing Systems (Vol. 35, pp. 24824–24837). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf Wu, S., Yang, S., Chen, Z., & Su, Q. (2024). Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference Tuning. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 22583–22599. https://doi.org/10.18653/v1/2024.emnlp-main.1258 Yoon, E. J., Tessler, M. H., Goodman, N. D., & Frank, M. C. (2020). Polite Speech Emerges From Competing Social Goals. Open Mind , 4 , 71–87. https://doi.org/10.1162/opmi_a_00035 Yu, K., Zeng, Q., Xuan, W., Li, W., Wu, J., & Voigt, R. (2025). The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models (Version 3). arXiv. https://doi.org/10.48550/ARXIV.2505.18497 Yue, S., Song, S., Cheng, X., & Hu, H. (2024). Do Large Language Models Understand Conversational Implicature—A case study with a chinese sitcom (arXiv:2404.19509). arXiv. https://doi.org/10.48550/arXiv.2404.19509 Zhang, J., & Wu, Y. (2020). Only youxie think it is a nice thing to say: Interpreting scalar items in face-threatening contexts by native Chinese speakers. Journal of Pragmatics , 168 , 19–35. https://doi.org/10.1016/j.pragma.2020.06.008 Zhao, H., & Hawkins, R. D. (2025). Comparing human and LLM politeness strategies in free production. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , 16199–16227. https://doi.org/10.18653/v1/2025.emnlp-main.820 Additional Declarations No competing interests reported. Supplementary Files SupplementalInformation.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 10 Apr, 2026 Reviewers agreed at journal 15 Mar, 2026 Reviewers invited by journal 11 Mar, 2026 Editor invited by journal 05 Feb, 2026 Editor assigned by journal 03 Feb, 2026 Submission checks completed at journal 03 Feb, 2026 First submitted to journal 02 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8763414","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":606359586,"identity":"c7d1e03f-8471-4309-a27e-ccf351bb57fd","order_by":0,"name":"Yuying Wang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA4UlEQVRIiWNgGAWjYPACGzDJ2AAiDxCnJY2Bh1Qth0nQYnAjx/Bzwa/z9vYSCWwPZ7YxyPHdSGD8XIBfi7H0zL7biT0SCeyGG9sYjCVvJDBLz8CrJXeDNG/P7QQeoC2SD9sYEjfcSGBj5sGvZfNv3p5z9jAt9cRo2SbN8+MAYw9IC9BhCQaEtEieef/NmrchObHnzMN2wxnnJAxnnnnYLI1PC9/xtOTbPH/s7Nnbk4897Cmzkec7nnzwMz4tChcSgPHRBmKCSQkGWPTgBPL9B4DkHzCbDa/KUTAKRsEoGLkAAKHATu1P+dFHAAAAAElFTkSuQmCC","orcid":"","institution":"Peking University","correspondingAuthor":true,"prefix":"","firstName":"Yuying","middleName":"","lastName":"Wang","suffix":""}],"badges":[],"createdAt":"2026-02-02 10:08:31","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8763414/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8763414/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104867380,"identity":"eae62818-95b5-4390-8b34-f11da39c57e1","added_by":"auto","created_at":"2026-03-18 07:12:32","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":68910,"visible":true,"origin":"","legend":"\u003cp\u003eExperimental paradigm and design overview. Participants (humans and LLMs) read short scenarios manipulating Context (face-threatening vs. neutral) and Factual State (all vs. most). In Experiment 1 (production), they completed a target sentence by choosing a quantifier (choice completion). In Experiment 2 (judgment), they evaluated the acceptability (Y/N) of a given response, with Statement Type manipulated between a vague statement (e.g., “some”) and a fully factual statement consistent with the context’s factual state.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8763414/v1/e4e444438238b754f6945d67.jpeg"},{"id":104867511,"identity":"a09b8d6d-dfb1-460b-b325-6812386f2426","added_by":"auto","created_at":"2026-03-18 07:12:59","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":98586,"visible":true,"origin":"","legend":"\u003cp\u003eQuantifier choice distributions in Experiment 1. Proportion of responses for each quantifier option as a function of Face Threat (face-threatening vs. neutral) and Factual State for human participants and LLMs. Panels show the two factual states: (a) Factual State = most and (b) Factual State = all. Within each panel and source, bars indicate the proportion of trials on which each quantifier was selected; error bars show 95% confidence intervals. Colors encode the chosen quantifier.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8763414/v1/3654b41662507d2403ae5d85.png"},{"id":104867508,"identity":"e41eacae-3a28-4630-b67d-dc3b202ffe42","added_by":"auto","created_at":"2026-03-18 07:12:59","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":76165,"visible":true,"origin":"","legend":"\u003cp\u003eAcceptance judgments in Experiment 2. (a) Acceptance of vague responses \u003cem\u003esome\u003c/em\u003e and (b) acceptance of factual responses as a function of face threat (neutral vs. face-threatening) and context factual state (\u003cem\u003eMost\u003c/em\u003e vs. \u003cem\u003eAll\u003c/em\u003e) for human participants and LLMs. Bars show mean acceptance proportions in each condition separately for each source; error bars indicate 95% confidence intervals. Colors denote source.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8763414/v1/5226700f60fcda2c1e0e3015.png"},{"id":104867535,"identity":"48ad0b3c-437a-4baa-b4af-a5bf71d189f8","added_by":"auto","created_at":"2026-03-18 07:13:08","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":97818,"visible":true,"origin":"","legend":"\u003cp\u003eHuman–LLM item-level correlations of cost proxy in Experiment 2. Heatmaps show Pearson correlations \u003cem\u003er\u003c/em\u003e between human processing/behavioral measures (rows: log Judgment RT, log Context RT, and acceptance rate) and model-specific cost proxies (columns: DeepSeek-R1 log thinking-token cost; DeepSeek-V3.2 log probability; GPT-4o log probability). Correlations are computed across items within each condition cell and are plotted separately by Face Threat and Factual State. Cell colors encode the direction and magnitude of r (blue = positive; red = negative), with the numeric r overlaid. Asterisks mark significance (* \u003cem\u003ep\u003c/em\u003e \u0026lt; .05, ** \u003cem\u003ep\u003c/em\u003e \u0026lt; .01).\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8763414/v1/85ca6eda8fabdf5b6b91bd44.png"},{"id":104867639,"identity":"288acb3d-55a4-430f-bcdc-68bd559f7fa2","added_by":"auto","created_at":"2026-03-18 07:13:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1397378,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8763414/v1/8fed2fbd-a3d4-4967-b9d6-3dc1d9eb4351.pdf"},{"id":104867408,"identity":"b76424f7-a3dc-47b0-a341-e71146802a84","added_by":"auto","created_at":"2026-03-18 07:12:39","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":57738,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementalInformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-8763414/v1/5f4d0e2aec72818a273dbc20.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"When Do LLMs Say “Some”? An Investigation of Scalar Implicature and Politeness Mitigation in Large Language Models","fulltext":[{"header":"Introduction","content":"\u003cp\u003eScalar implicatures represent a key interface between logical meaning and pragmatic reasoning (Grice, 1975; Horn, 1972, 2006; Levinson, 2000), arising when a speaker’s use of a weak scalar term invites the inference that a stronger alternative does not hold (Horn, 1972). The quantifier \u003cem\u003esome\u003c/em\u003e offers the canonical illustration: although it semantically denotes a non-zero subset and is logically compatible with the universal quantifier \u003cem\u003eall\u003c/em\u003e, comprehenders reliably interpret \u003cem\u003esome\u003c/em\u003e as conveying the strengthened meaning “\u003cem\u003esome\u003c/em\u003e but \u003cem\u003enot all\u003c/em\u003e” (Chemla \u0026amp; Bott, 2014; Chemla \u0026amp; Singh, 2014; Katsos \u0026amp; Cummins, 2010).\u003c/p\u003e \u003cp\u003eConsider the contrast in (1).\u003c/p\u003e \u003cp\u003e(1a)\u003cem\u003eSome\u003c/em\u003e of the experimental materials were human-created.\u003c/p\u003e \u003cp\u003e(1b)\u003cem\u003eAll\u003c/em\u003e of the experimental materials were human-created.\u003c/p\u003e \u003cp\u003eFrom a purely logical standpoint, if (1b) is true, (1a) must be true as well; thus, (1a) and (1b) are congruent with respect to their truth conditions. However, human comprehenders typically infer from (1a) that \u003cem\u003enot all\u003c/em\u003e of the experimental materials were human-created. This enriched interpretation is emerges from expectations regarding communicative cooperativity and informativeness within the Gricean framework (Grice, 1975). Within the Gricean Cooperative Principle, communication is assumed to be rational and goal-directed. If a speaker knows that a stronger, more informative statement (e.g., \u003cem\u003eall\u003c/em\u003e) is true but nevertheless opts for a weaker expression (e.g., \u003cem\u003esome\u003c/em\u003e), the listener is warranted in inferring that the speaker is deliberately avoiding the stronger alternative. The use of \u003cem\u003esome\u003c/em\u003e thus implicates that \u003cem\u003eall\u003c/em\u003e does not hold, giving rise to the scalar implicature \u003cem\u003enot all\u003c/em\u003e.\u003c/p\u003e \u003cp\u003ePoliteness and face threat mitigation\u003c/p\u003e \u003cp\u003eSpeakers typically follow Gricean maxims by being as informative as possible. Pragmatic theory, however, has long emphasized that social context can override this default. Brown and Levinson (1987) define face as a person’s public self-image and show that speakers use mitigation strategies to avoid threatening face. A direct criticism is a prototypical face-threatening acts (FTAs), and speakers often soften such acts with indirect or imprecise, vague language. Particularly, in such contexts, scalar quantifier \u003cem\u003esome\u003c/em\u003e can function as a hedge that narrows the scope of the negative assessment, thereby reducing its face-threat relative to a blunt, fully factual statement. Consequently, the pragmatic use of the vague quantifier \u003cem\u003esome\u003c/em\u003e often extends beyond the Gricean pressure toward maximal informativeness. For example, when offering critical feedback after the party, knowing that \u003cem\u003eall\u003c/em\u003e of the guests found the joke offensive, a speaker might choose between the following formulations:\u003c/p\u003e \u003cp\u003e(2a) \u003cem\u003eAll\u003c/em\u003e guests didn’t like your joke.\u003c/p\u003e \u003cp\u003e(2b) \u003cem\u003eSome\u003c/em\u003e guests didn’t like your joke.\u003c/p\u003e \u003cp\u003eIn this setting, the use of \u003cem\u003esome\u003c/em\u003e in (2b) does not invite the scalar inference \u003cem\u003enot all\u003c/em\u003e but to mitigation of FTAs. In fact, although saying \u003cem\u003esome\u003c/em\u003e when \u003cem\u003eall\u003c/em\u003e will interpret as underinformative but it will be a polite softening of the statement. Thus, the trade-off between informativeness and politeness is highly context-dependent.\u003c/p\u003e \u003cp\u003eBonnefon et al. (2009) provide experimental evidence for this effect. They found that when an utterance would threaten the listener’s face, listeners are less likely to draw the usual implicature from \u003cem\u003esome\u003c/em\u003e to \u003cem\u003enot all\u003c/em\u003e. Importantly, further experiments further refine the boundary conditions: when the negative evaluation directly targets the addressee’s own face (rather than involving a third party), scalar enrichment is selectively weakened, indicating that face concerns can modulate the derivation of scalar meaning.\u003c/p\u003e \u003cp\u003eSubsequent work corroborates this broader efficiency–politeness trade-off across scalar expressions. Politeness-oriented contexts tend to reduce listeners’ endorsement of enriched interpretations (Feeney \u0026amp; Bonnefon, 2013), and may yield “polite misunderstandings” by increasing ambiguity about whether a speaker’s wording is intended to maximize clarity or to manage face (Bonnefon et al., 2011). Extending this account to Mandarin, Zhang and Wu (2020) show that listeners’ interpretation of \u003cem\u003esome\u003c/em\u003e is modulated by motive attribution: when the utterance is construed as informativeness-driven, listeners are more likely to endorse scalar enrichment; when it is identified as a politeness/face-management strategy, they tend to prefer a literal, all-compatible reading, thereby suppressing the scalar inference.\u003c/p\u003e \u003cp\u003eLarge language models and pragmatic flexibility\u003c/p\u003e \u003cp\u003eThe advent of large language models (LLMs) opens new avenues for studying pragmatics, but their pragmatic capabilities are still need being understood (Cong, 2024; Ma et al., 2025; Shulginov et al., 2025; Wu et al., 2024; Yu et al., 2025; Yue et al., 2024; Zhao \u0026amp; Hawkins, 2025). LLMs are trained on massive text corpora and often generate fluent, contextually appropriate language, yet it remains unclear whether they exhibit human-like pragmatic flexibility, particularly in domains such as scalar implicature and politeness.\u003c/p\u003e \u003cp\u003eRecent studies offer a nuanced picture of how LLMs handle scalar implicatures triggered by the quantifier \u003cem\u003esome\u003c/em\u003e. On one hand, certain LLMs do generate the expected implicature in simple contexts, aligning with human intuitions. For instance, Cho and Kim (2024) found that even without any context, both BERT and GPT-2 tend to treat \u003cem\u003esome\u003c/em\u003e pragmatically as implying \u003cem\u003enot all\u003c/em\u003e, suggesting an inherent grasp of the implicature at a basic level. BERT in particular appeared to encode this inference by default. On the other hand, noticeable differences from human pragmatics emerge once contextual factors are introduced. Building on the classic paradigm of Bonnefon et al. (2009), Qiu, Duan, and Cai (2023) found that GPT-3.5 shows a highly uniform pragmatic interpretation preference for some across both face-threatening and face-enhancing contexts. Its interpretations are entirely unaffected by face considerations, indicating a lack of the human-like ability to flexibly suppress the “not all” scalar inference in response to different pragmatic cues. In contrast, Ku (2025) notes that a single GPT-based agent will often interpret “\u003cem\u003eSome\u003c/em\u003e students passed” in a strictly semantic way that permits all students passing. However, when the same model was placed in a multi-agent interactive framework, the system is able to derive scalar inferences that are much closer to human-like interpretations.\u003c/p\u003e \u003cp\u003eThis gap underscores that, despite their impressive linguistic fluency, traditional LLMs fall short of internalizing the context-sensitive nature of scalar implicature as observed in human communication.\u003c/p\u003e \u003cp\u003eFrom output to reasoning: a Chain-of-Thought perspective on pragmatic flexibility in LLMs\u003c/p\u003e \u003cp\u003eRecent advances in model training have enabled LLMs to generate not only answers but also explicit reasoning traces when instructed to do so. Chain-of-thought (CoT) prompting has been shown to elicit more structured, stepwise solutions and to improve performance on tasks requiring multi-step computation (Wei et al., 2022).\u003c/p\u003e \u003cp\u003eThe emergence of reasoning-oriented models such as DeepSeek-R1 reflects a growing research trend toward training LLMs with reinforcement learning objectives that directly incentivize stepwise reasoning behaviors rather than relying solely on scale and instruction tuning (Wang, 2025). In the DeepSeek-R1 framework, pure reinforcement learning is used to encourage advanced reasoning patterns such as self-reflection and verification, which has been shown to enhance performance on verifiable reasoning tasks including mathematics and coding relative to conventional supervised approaches (Guo et al., 2025).\u003c/p\u003e \u003cp\u003eThis shift is methodologically consequential not only for its performance implications, but also because CoT-oriented models make aspects of their decision process observable. The length and structure of reasoning traces can be treated as a proxy for computational cost, enabling comparisons not only of final outputs but also of the effort required to produce them. de Varda et al. (2025) show that reasoning cost—operationalized as the number of reasoning steps or tokens—scales with task difficulty in ways that parallel human cognitive effort. At the same time, it remains unclear whether such traces correspond to the model’s internal inference processes or instead reflect learned strategies for explanation rather than transparent readouts of underlying computation. Nevertheless, CoT-oriented training provides a new empirical window for systematically characterizing pragmatic flexibility in LLMs.\u003c/p\u003e \u003cp\u003eCurrent study and research questions\u003c/p\u003e \u003cp\u003eTo address these issues, we conducted two experiments in parallel with human participants and LLMs. In the quantifier production task, agents were asked to complete sentences as a speaker. In the acceptability judgment task, humans and LLMs judged how appropriate utterances with some or stronger quantifiers were in those contexts as a listener. All stimuli, procedure, and response formats are matched across humans and LLMs. We included both chain-of-thought–trained and standard LLMs (GPT-4o, DeepSeek-V3.2, and DeepSeek-R1) to examine how reasoning optimization shapes pragmatic behavior.\u003c/p\u003e \u003cp\u003eThis rigorous design allows us to test, within a single unified framework, whether the resulting patterns of quantifier production and acceptability judgments in LLMs align with those observed in human participants. At the same time, by contrasting reasoning-optimized and non–Chain-of-Thought models, we assess whether explicit reasoning supervision yields more human-like sensitivity to contextual and social constraints on informativeness, or whether such pragmatic adaptation can be accounted for by surface-level language modeling alone. These questions will clarify the extent to which LLMs replicate human pragmatic strategies and illuminate whether their outputs reflect genuine evaluative processes or mere statistical tendencies.\u003c/p\u003e "},{"header":"Methods","content":"\u003cp\u003eParticipants\u003c/p\u003e\u003cp\u003eForty-eight native utterances of Mandarin Chinese (females = 24; age M = 20.8 ± 2.10 years) were recruited for Experiment 1. Another forty-eight native utterances of Mandarin Chinese for Experiment 2 (females = 24; age M = 21.38 ± 2.11 years). All participants had normal or corrected-to-normal vision and gave informed consent. Each participant was a native Chinese speaker to ensure they understood the linguistic nuances of the materials. Participants were tested individually in a quiet loom room and were compensated for their time. This research was conducted in accordance with the Declaration of Helsinki and was approved by the Ethics Committee of the author’s institution.\u003c/p\u003e\u003cp\u003eLanguage models\u003c/p\u003e\u003cp\u003eWe evaluated three Large Language Models using the same experimental materials as the human participants. The models were accessed via the PsyLingLLM toolbox, which presents linguistic stimuli to LLMs in a controlled manner. The models included three contemporary LLMs spanning distinct inference and deployment regimes. (1) GPT-4o is a proprietary, instruction-tuned OpenAI GPT-4–class model evaluated in a direct-answer setting. (2) DeepSeek-V3.2 is a state-of-the-art open-source Mixture-of-Experts LLM (671B parameters) evaluated in the same direct-answer setting. (3) DeepSeek-R1 is a reasoning-optimized derivative of DeepSeek-V3 (671B Mixture-of-Experts), explicitly optimized via reinforcement learning to generate CoT style reasoning; it was evaluated in a trace-emitting setting in which an explicit multi-step reasoning trace precedes the final answer. For analysis, we distinguish three model categories: a trace-emitting CoT model (DeepSeek-R1), a proprietary direct-answer GPT model (GPT-4o), and an open-source direct-answer non-CoT model (DeepSeek-V3.2). Each model was queried via API access through the PsyLingLLM framework using identical prompts, ensuring consistency in input presentation, output format, and generation parameters across models.\u003c/p\u003e\u003cp\u003eMaterials and design\u003c/p\u003e\u003cp\u003e\u003cb\u003eExperiment 1: Pragmatic Completion.\u003c/b\u003e This experiment employed a 2 × 2 within-subjects design. We manipulated Face Threat (face-threatening vs. neutral context) and Factual State (All vs. Most, the ground-truth prevalence in the context) as within-item factors. The materials consisted of sentence contexts and prompts designed to elicit a quantifier. Each trial presented a short vignette describing a scenario involving a group of people or items, along with a prompt sentence to be completed with a quantifier. The Face Threat context manipulation provided a social motivation for under-informative answers. Specifically, Face-threatening scenarios involved dyadic interactions in which the speaker addressed the potentially affected individual directly, thereby creating face-threatening pressure (e.g. Mei performed a piece at the evening party. Afterward, she asked her friend how many parts of the piece she had played wrong. Her friend heard that all parts were off, but not wanting Mei to lose face.), whereas Neutral scenarios described third-party situations in which the speaker reported information about others, with no direct interpersonal stakes (e.g. Mei and her friend went to a show where they heard an audience member perform a piano piece. Afterward, she asked her friend how many parts of the piece were played wrong. Her friend heard that all parts were played incorrectly.). The Factual State manipulation set the ground truth of the scenario: in All contexts, 100% of the group possessed a certain property or did a certain action; in Most contexts, the relevant property held for a majority of cases (more than half, but not all). After reading the context, participants saw an incomplete sentence such as “[Context]. When asked, the person replied: ‘___ (of the X did Y).’” Seven quantifier options were available for the blank, spanning a range of informational strengths. These included all (全部) and none (没有) as fully informative endpoints; most (多数) and a few (少数) as intermediate quantifiers; more informationally specific alternatives relative to most/few—namely the great majority (大多数) and very few (极少数); and the vague, pragmatically flexible quantifier some (一些). Participants were instructed to select the quantifier that best completed the sentence given the context. Thus, the primary dependent variable was the choice of quantifier. For analysis, we focused on whether the quantifier “some” was chosen (binary outcome: some vs. other quantifier). This choice of “some” is critical because “some” implicates “not all,” making it a key indicator of pragmatic under-informativeness. Each scenario was instantiated in four versions crossing Face Threat and Factual State. To ensure that each participant saw only one version of each scenario, we distributed the four versions of all scenarios across four material lists using a Latin-square design. Each list contained exactly one condition version of each scenario, and participants were randomly assigned to one list. To balance the a priori plausibility of the seven quantifier options and to obscure the critical manipulation, we embedded two types of filler trials: 12 fillers designed to make each option a factually licensed and natural response at least some of the time, and 14 descriptive fillers that did not involve interpersonal considerations, thereby disrupting the target contextual structure. The 32 critical trials and 26 filler trials were pseudo-randomized such that no more than two trials from the same experimental condition occurred consecutively, mitigating potential order effects. The same 32 critical scenarios were used for LLM evaluation. All four versions of each scenario were presented to each model. To estimate response variability, each stimulus was independently sampled 10 times per model with different random seeds, yielding a total of 1,280 observations per model. All trials were fully randomized. Each trial was run in isolation, with no conversational history or cross-trial context retained between prompts.\u003c/p\u003e\u003cp\u003e\u003cb\u003eExperiment 2: Pragmatic Judgment.\u003c/b\u003e This experiment used a 2 × 2 × 2 within-subjects design. The independent variables were Face Threat (face-threatening/neutral), Factual State (Most vs. All), and Statement Type (“Some” statement vs. fully informative “Factual” statement). Materials consisted of brief scenarios and a target statement to be evaluated. Each scenario set up a context similar to Experiment 1. Following the scenario, participants were presented with a target sentence that was either an under-informative “some” statement or a fully informative factual statement. For example, in a context where indeed all members of a group did something, the “Some” statement might present the sentence “\u003cem\u003eSome\u003c/em\u003e of the \u003cem\u003eY\u003c/em\u003e did \u003cem\u003eX\u003c/em\u003e.” (e.g. \u003cem\u003eSome\u003c/em\u003e were played wrong.); the “Factual” statement in that context would present “\u003cem\u003eAll\u003c/em\u003eof the \u003cem\u003eY\u003c/em\u003e did \u003cem\u003eX\u003c/em\u003e” (e.g. \u003cem\u003eAll\u003c/em\u003e were played wrong.). In a Most context, the factual statement would be “\u003cem\u003eMost\u003c/em\u003e of the \u003cem\u003eY\u003c/em\u003e did \u003cem\u003eX\u003c/em\u003e.” Participants were instructed to judge whether the given statement was an acceptable or appropriate description of the situation, given the context. They responded “Yes” (acceptable) or “No” (unacceptable) for each statement. The primary dependent variable was the binary acceptability judgment for the “some” statement – i.e., whether an under-informative “some” was accepted as an appropriate utterance. By design, fully informative statements in factual conditions are truthfully correct; we expected those to be generally accepted, serving as a control. Each participant encountered all combinations of the three factors across trials. For instance, each participant might see a scenario of each type: Face Threat × Factual State, each with both a “some” version and a “factual” version of the statement, for a total of 8 trial types. There were same 32 unique scenarios in total, each participant seeing each scenario in one of its possible forms. Trials were randomized and balanced to prevent any ordering or priming effects. As in Experiment 1, the same scenarios were used for LLM evaluation; with all four versions of each scenario presented and each item sampled 10 times, this yielded a total of 2,560 observations per model.\u003c/p\u003e\u003cp\u003eProcedure and Measures\u003c/p\u003e\u003cp\u003e \u003cb\u003eHuman Behavioral Procedure.\u003c/b\u003e Participants were tested using a computer-based task programmed in MATLAB 2022b (MathWorks Inc.). Experiment 1 (Production task) and Experiment 2 (Judgment task) were run as separate sessions with different participant groups. All instructions and materials were presented in Chinese. Participants were first given on-screen instructions explaining the task with examples. For Experiment 1, participants were told they would read scenarios and complete sentences with an appropriate quantifier. They practiced with 20 of example scenarios to familiarize themselves with the seven quantifier options. During the actual trials, each scenario text was presented on the screen. After reading the scenario, participants pressed the space bar to proceed to the response screen. Below the question, the seven quantifier options were displayed and mapped to the number keys 1–7; their left–right positions were randomized on each trial, so the option–key mapping varied across trials. Participants selected the quantifier that best completed the sentence. They were instructed to respond as naturally and accurately as possible, and that there were no strictly “correct” answers but some might be more appropriate given the context. Once a selection was made, the next trial began. For Experiment 2, participants were instructed that they would read a scenario and then see a statement, and they should judge whether saying that statement in that context was acceptable or not. Trials began with the context description, followed by the target statement. Participants pressed one of two keys to record their judgment (e.g., “J” = acceptable, “F” = unacceptable); for the remaining half of participants, this key–response mapping was reversed to counterbalance motor-response associations across participants. They were encouraged to rely on their pragmatic intuition about whether the response was appropriate, not just whether it was true. Participants completed a brief practice for Experiment 2 as well, to ensure they understood the concept of pragmatic acceptability. Reaction times (RTs) were recorded for each trial in Experiment 2 in two components. Contextual reading reaction time is measured from the beginning to the end of a contextual short story. Judgment RT was measured from target-statement onset to the keypress indicating the participant’s acceptability judgment. The entire session lasted approximately 30 minutes per experiment.\u003c/p\u003e\u003cp\u003e \u003cb\u003eLLM Evaluation Procedure.\u003c/b\u003e Each experiment’s stimuli were presented to the language models in a parallel manner. We constructed standardized prompts in Chinese to ensure the models received the same information as human participants. For Experiment 1, the prompt for each trial consisted of the scenario description followed by an incomplete sentence. Generation was performed with a sampling temperature set to 1.0. LLM responses were collected using the PsyLingLLM framework, which allows systematic logging of model outputs and associated generation metadata. For the non-CoT models (DeepSeek-V3.2 and GPT-4o), each trial yielded a single-token answer corresponding to the selected quantifier. For these models, we recorded the chosen answer, the overall response latency (from prompt submission to token return), and the log probability associated with the returned output token. For the CoT model (DeepSeek-R1), token-level log probabilities were not available. Instead, we separately extracted the final answer quantifier and the accompanying chain-of-thought reasoning trace. For each CoT trial, we recorded the total elapsed time from prompt submission to final response, as well as the length of the chain-of-thought in tokens, which served as an index of reasoning extent. For Experiment 2, the LLM evaluation procedure closely paralleled that of Experiment 1. Each trial prompt consisted of the scenario description followed by a target statement and a simple binary acceptability question (yes/no), mirroring the human judgment task. The target statement was either an under-informative some statement or a fully informative factual statement, depending on condition. The full Chinese prompts and example materials used for LLM evaluation are provided in Supplementary Materials S1 and S2.\u003c/p\u003e\u003ch2\u003eStatistical analysis\u003c/h2\u003e\u003cp\u003eWe analyzed the data using regression models that respect the repeated-measures structure of the experiments. All analyses were conducted in R (version 4.3.2). Human data were analyzed with mixed-effects models implemented in lme4 (Bates et al., 2015) to account for repeated observations from the same participants and repeated use of the same items. For LLM data, because there is no participant-level sampling and each item was evaluated repeatedly via independent API calls, inference was conducted at the item level using generalized linear/linear models with item fixed effects and item-clustered robust standard errors.\u003c/p\u003e\u003cp\u003eFor choice-level outcomes, we analyzed binary responses with logit-link models. In Experiment 1, the dependent variable was whether the response option corresponding to some was selected on each trial. In Experiment 2, the dependent variable was whether the statement was judged acceptable. Fixed effects included the experimental manipulations—Face Threat, Factual State, and Statement Type—and their factorial interactions. For human data, we fit generalized linear mixed-effects models (GLMMs; binomial family with logit link) with crossed random intercepts for participant and item; where counterbalanced lists were used, Version was included as a fixed-effect control (Face Threat × Factual State × Statement Type + Material Version, with crossed random intercepts for Subject and Item). For LLM data, each model produced multiple independent responses per item×condition cell (10 API calls; temperature = 1), so we avoided treating calls as independent observations. Instead, we aggregated each item×condition (and statement type, when applicable) to a binomial count (k successes out of n = 10) and fit binomial-logit generalized linear models (GLMs) including item fixed effects to control for item-specific baselines; statistical inference used item-clustered robust standard errors (sandwich estimator), and we report odds ratios (OR), 95% confidence intervals, and Wald statistics for the focal effects. To quantify distributional similarity between humans and models in the full response-option distributions, we also computed Jensen–Shannon (JS) divergence between the empirical human distribution and each model distribution within each condition. Confidence intervals were obtained via cluster bootstrap resampling that respected the dependence structure. JS divergence was treated as descriptive evidence of distributional similarity and does not replace the regression-based inference.\u003c/p\u003e\u003cp\u003eFor continued cost-level outcomes in experiment 2, human processing costs were operationalized using reaction times at the context-reading stage (ContextRT) and the judgment stage (JudgmentRT). To reduce positive skew, RTs were log-transformed using the natural logarithm (log), after excluding non-positive observations. We then fit linear mixed-effects models of the form log(RT) ~ Face Threat × Factual State × Statement Type + Material Version, with crossed random intercepts for Subject and Item. Fixed-effect estimates are reported as regression coefficients (b) with 95% confidence intervals, and inference relied on Satterthwaite-approximated degrees of freedom as implemented in lmerTest. For LLM data, to maintain item-level inference and avoid pseudo-replication across repeated API calls, we aggregated each proxy to the Item × Face Threat × Factual State × Statement Type level by averaging across the 10 calls. We then fit linear models including the full experimental interaction and item fixed effects (CostValue ~ Face Threat × Factual State × Statement Type + Item). Estimates are reported as b with 95% confidence intervals; inference used item-clustered robust standard errors, with Wald t tests computed using degrees of freedom equal to the number of item clusters minus one. Because human RTs and model proxies are expressed in fundamentally different units, we did not compare absolute magnitudes across systems; instead, we compared the pattern of experimental effects within each system. Finally, for cost–acceptance alignment, we conducted condition-wise, item-level correlations restricted to “Some” statements. For humans, we first computed item × Face Threat × Factual State cell means for log-transformed RT, and mean acceptability, ensuring inference was driven by between-item variation rather than trial-level pseudo-replication. For LLMs, we computed the corresponding item × Face Threat × Factual State cell means for each model’s cost proxy. Correlation uncertainty was summarized with 95% confidence intervals derived via Fisher’s z transformation.\u003c/p\u003e\u003cp\u003eAll scripts for preprocessing and analysis are available in the project repository (OSF link).\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eCoT reasoning yields more human-like mitigation profiles in the production task\u003c/p\u003e \u003cp\u003eIn the completion task (Experiment 1), Human participants exhibited a clear pragmatic adaptation in their quantifier use (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). In neutral contexts (no interpersonal stakes), utterances almost always provided fully informative answers: for instance, when \u003cem\u003eall\u003c/em\u003e individuals in the scenario had done the action, nearly every participant chose \u003cem\u003eall\u003c/em\u003e (全部), and when only most had done it, \u003cem\u003emost\u003c/em\u003e (多数) was overwhelmingly used. In contrast, Face-threatening contexts (a context pointing out that \u0026ldquo;not all did it\u0026rdquo; would embarrass someone) elicited far more under-informative responses. In these scenarios, utterances often avoided the stronger quantifiers even when they were true. The vague quantifier \u003cem\u003esome\u003c/em\u003e (一些) became the single most frequent choice, accounting for roughly 44.53% and 46.88% of responses in both \u003cem\u003eMost\u003c/em\u003e and \u003cem\u003eAll\u003c/em\u003e contexts with face threat. For example, when in reality all members of a group had done the task, only a minority of utterances actually said \u0026ldquo;\u003cem\u003eall\u003c/em\u003e did it\u0026rdquo;; nearly half responded with \u0026ldquo;\u003cem\u003esome\u003c/em\u003e (did it)\u0026rdquo;. The remaining responses in Face-threatening conditions were spread across other weaker terms \u0026ndash; e.g. \u003cem\u003efew\u003c/em\u003e (少数) or even \u003cem\u003enone\u003c/em\u003e (没有) \u0026ndash; effectively understating the outcome to avoid embarrassment to any individual. Consequently, fully informative quantifiers were rarely used in Face-threatening contexts (1.82% \u003cem\u003emost\u003c/em\u003e responses in \u003cem\u003eMost\u003c/em\u003e contexts, and 2.08% \u003cem\u003eall\u003c/em\u003e responses in \u003cem\u003eAll\u003c/em\u003e contexts; compared with 81.51% and 83.60%, respectively, under the neutral condition). This distributional shift confirms that human utterances often sacrificed informativeness for politeness, frequently opting for \u0026ldquo;some\u0026rdquo; instead of a more informative term when a truthful answer would be socially awkward, replicates the classic finding that face threat blocks the usual scalar inference.\u003c/p\u003e \u003cp\u003eThe LLMs\u0026rsquo; responses showed qualitatively similar context effects, but with notable differences in degree. All three models learned to use \u003cem\u003esome\u003c/em\u003e more often in Face-threatening situations, but the extent of this shift varied. Notably, DeepSeek-V3.2 in Face-threatening contexts produced \u003cem\u003esome\u003c/em\u003e almost categorically \u0026ndash; for instance, 93.75% of its responses were \u003cem\u003esome\u003c/em\u003e in the \u003cem\u003eMost\u003c/em\u003e condition, and over 95.63% in \u003cem\u003eAll\u003c/em\u003e condition. This exceeded the human rate and indicates an exaggerated politeness bias in DeepSeek-V3.2. GPT-4o also displayed a similar over-reliance on \u003cem\u003esome\u003c/em\u003e under face threat (92.50% and 73.43% of responses \u003cem\u003esome\u003c/em\u003e) and in fact produced less varied output than humans in those conditions. By contrast, DeepSeek-R1 more closely mimicked the human distribution: in Face-threatening contexts it chose \u003cem\u003esome\u003c/em\u003e on roughly one-third to one-half of trials (32.18% in \u003cem\u003eMost\u003c/em\u003e, 46.88% in \u003cem\u003eAll\u003c/em\u003e). Beyond \u003cem\u003esome\u003c/em\u003e frequency, it occasionally used other mitigating quantifiers (\u003cem\u003efew\u003c/em\u003e, etc.), yielding a broader response spread. Quantitatively, DeepSeek-R1\u0026rsquo;s full response distributions closely matched those of humans, with low JS divergence (0.038 and 0.048 in the \u003cem\u003eMost\u003c/em\u003e and \u003cem\u003eAll\u003c/em\u003e conditions, under face-threatening contexts) and relatively high entropy (1.04 and 1.37, compared with human values of 1.30 and 1.52, see Table S3). Whereas GPT-4o and especially DeepSeek-V3.2 diverged more. For DeepSeek-V3.2, the JS divergence increased to 0.18 and 0.19 in face-threatening contexts (under the \u003cem\u003eMost\u003c/em\u003e and \u003cem\u003eAll\u003c/em\u003e conditions, respectively), alongside low entropy values (0.36 and 0.27), indicating a near-deterministic preference for \u003cem\u003esome\u003c/em\u003e. GPT-4o with JS divergence values of 0.11 and 0.17 and entropy levels of 0.39 and 1.02 under face-threatening conditions, indicating a strong bias toward \u0026ldquo;some\u0026rdquo; while retaining slightly variance. In Neutral contexts, all systems were near-ceiling on the contextually appropriate quantifier (\u0026ge;\u0026thinsp;81.56%) and thus closely matched human behavior (JS\u0026thinsp;\u0026le;\u0026thinsp;0.05).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe modeled the binary likelihood of choosing \u003cem\u003esome\u003c/em\u003e (vs. any other quantifier) using logistic regression (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The human GLMM data revealed significant main effects of both experimental factors and their interaction. Specifically, there was a huge effect of Face Threat: utterances were far more likely to choose \u003cem\u003esome\u003c/em\u003e in face-threatening situations than in neutral ones (OR\u0026thinsp;=\u0026thinsp;12.80, 95% CI [8.17, 20.20], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001). There was a reliable effect of the Factual State: overall, scenarios where only \u003cem\u003eMost\u003c/em\u003e of the group had done the action led to fewer \u003cem\u003esome\u003c/em\u003e responses than scenarios where \u003cem\u003eAll\u003c/em\u003e had done it (OR\u0026thinsp;=\u0026thinsp;0.35, 95% CI [0.18, 0.68], \u003cem\u003ep\u003c/em\u003e = .002). In other words, when factuality is an \u003cem\u003eAll\u003c/em\u003e condition, participants were somewhat less inclined to hedge with some. We also observed a significant Face Threat \u0026times; Factual State interaction (OR\u0026thinsp;=\u0026thinsp;3.23, 95% CI [1.53, 6.82], \u003cem\u003ep\u003c/em\u003e = .002), indicating that the effect of Face Threat on the likelihood of choosing \u003cem\u003esome\u003c/em\u003e was stronger in the \u003cem\u003eAll\u003c/em\u003e context than in the \u003cem\u003eMost\u003c/em\u003e context.\u003c/p\u003e \u003cp\u003eThe LLMs exhibited the same qualitative effects pattern in the GLM analysis, with extreme coefficients consistent with near-ceiling shifts. DeepSeek-V3.2 (OR\u0026thinsp;=\u0026thinsp;2.03 \u0026times; 10\u0026sup2;⁰, 95% CI [5.86 \u0026times; 10\u0026sup1;⁹, 7.04 \u0026times; 10\u0026sup2;⁰], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001) showed an extremely large estimated odds ratio for Face Threat, indicating that under face-threatening conditions it almost never produced other responses when \u003cem\u003esome\u003c/em\u003e was licensed. A similarly amplified Face Threat effect was observed for GPT-4o (OR\u0026thinsp;=\u0026thinsp;5.16 \u0026times; 10\u0026sup1;⁹, 95% CI [1.03 \u0026times; 10\u0026sup1;⁹, 2.57 \u0026times; 10\u0026sup2;⁰], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), and a more moderate but reliable effect for DeepSeek-R1 (OR\u0026thinsp;=\u0026thinsp;73.10, 95% CI [8.46, 632], z\u0026thinsp;=\u0026thinsp;3.90, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001). All three models showed dramatically reduced odds of producing \u003cem\u003esome\u003c/em\u003e in the \u003cem\u003eMost\u003c/em\u003e relative to \u003cem\u003eAll\u003c/em\u003e contexts under the current coding scheme (DeepSeek-V3.2: OR\u0026thinsp;=\u0026thinsp;2.08 \u0026times; 10⁻\u0026sup1;⁸, 95% CI [2.16 \u0026times; 10⁻\u0026sup2;\u0026sup1;, 2.01 \u0026times; 10⁻\u0026sup1;⁵], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001; GPT-4o: OR\u0026thinsp;=\u0026thinsp;1.10 \u0026times; 10⁻⁸, 95% CI [1.92 \u0026times; 10⁻⁹, 6.33 \u0026times; 10⁻⁸], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001; DeepSeek-R1: OR\u0026thinsp;=\u0026thinsp;3.68 \u0026times; 10⁻⁸, 95% CI [3.76 \u0026times; 10⁻⁹, 3.60 \u0026times; 10⁻⁷], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), consistent with near-deterministic sensitivity to the underlying factual state. Critically, a significant Face Threat \u0026times; Factual State interaction was observed for all systems, indicating that the licensing effect of face threat on \u003cem\u003esome\u003c/em\u003e production depended on whether the true state was \u003cem\u003eAll\u003c/em\u003e or \u003cem\u003eMost\u003c/em\u003e (DeepSeek-V3.2: OR\u0026thinsp;=\u0026thinsp;1.12 \u0026times; 10\u0026sup1;⁸, 95% CI [6.20 \u0026times; 10\u0026sup1;⁶, 2.03 \u0026times; 10\u0026sup1;⁹], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001; GPT-4o: OR\u0026thinsp;=\u0026thinsp;1.08 \u0026times; 10⁷, 95% CI [7.18 \u0026times; 10⁵, 1.63 \u0026times; 10⁸], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001; DeepSeek-R1: OR\u0026thinsp;=\u0026thinsp;5.85 \u0026times; 10⁷, 95% CI [7.29 \u0026times; 10⁶, 4.69 \u0026times; 10⁸], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001).\u003c/p\u003e \u003cp\u003eCollectively, Experiment 1 establishes a clear human\u0026ndash;LLM contrast: face threat elicits a systematic trade-off between informativeness and social appropriateness in production. Humans markedly reduced fully informative quantifiers and increased under-informative responses\u0026mdash;most prominently some, but also a broader downward-weakening pattern among non-some alternatives\u0026mdash;indicating a general mitigation strategy under interpersonal risk. GPT-4o and DeepSeek-V3.2 often realized mitigation in an almost deterministic, some-centric way, thereby increasing divergence from human response distributions. By contrast, DeepSeek-R1 most closely approximated human behavior, both in overall distributional similarity and in exhibiting a human-like weakening strategy beyond some. These results suggest that DeepSeek-R1 may deploy a dual mitigation repertoire like human participants\u0026mdash;vagueness plus systematic weakening among non-some choices\u0026mdash;whereas the non-CoT models rely predominantly on a vagueness-dominant strategy.\u003c/p\u003e \u003cp\u003eFace Threat Increases Acceptance of \u003cem\u003eSome in\u003c/em\u003e Judgment Task\u003c/p\u003e \u003cp\u003eIn the judgment task (Experiment 2), we tested whether listeners judge under-informative statements to be acceptable descriptions given the context (Table S2). Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea gray bar displays acceptance rates for target statements as a function of face-threat and factual state. Human judgments revealed a mirror-image of the production results: acceptability of a \u003cem\u003esome\u003c/em\u003e statement (which is true but not maximally informative) was low in neutral contexts but rose dramatically in face-threatening contexts. Specifically, in the absence of any face threat, participants largely dispreferred a \u003cem\u003esome\u003c/em\u003e statement if a more informative statement was warranted \u0026ndash; especially in the \u003cem\u003eAll\u003c/em\u003e context. When indeed \u003cem\u003eall\u003c/em\u003e members had done the action, saying \u0026ldquo;\u003cem\u003eSome\u003c/em\u003e of \u003cem\u003eY\u003c/em\u003e did \u003cem\u003eX\u003c/em\u003e\u0026rdquo; was often viewed as infelicitous or odd (mean acceptability 35.94%). Even when the truth was \u003cem\u003emost\u003c/em\u003e did it, a \u003cem\u003esome\u003c/em\u003e statement was accepted only about 57.81% of the time without a face motive, reflecting a moderate penalty for under-informativeness. However, when face is threatened, listeners become more tolerant of the use of the word \u003cem\u003esome\u003c/em\u003e. In the face-threatening scenarios, the vast majority of participants found the under-informative \u003cem\u003esome\u003c/em\u003e statement acceptable. Acceptance of \u003cem\u003esome\u003c/em\u003e statements jumped to about 94.27% in the \u003cem\u003eMost\u003c/em\u003e context and 93.23% in the \u003cem\u003eAll\u003c/em\u003e context when a face threat was present. Thus, humans strongly contextualized their pragmatic judgments: a sentence like \u0026ldquo;Some of them succeeded,\u0026rdquo; which would ordinarily sound infelicitous if in fact all of them succeeded, becomes entirely appropriate when the speaker has a polite reason to remain vague. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea illustrates this crossover: the Face Threat effect on \u003cem\u003esome\u003c/em\u003e-statement acceptability is huge, whereas the factual context effect (\u003cem\u003eAll\u003c/em\u003e vs. \u003cem\u003eMost\u003c/em\u003e) is smaller.\u003c/p\u003e \u003cp\u003eA complementary pattern emerged for fully informative factual statements (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb). In neutral contexts, these statements were endorsed at high rates (\u003cem\u003eAll\u003c/em\u003e: 88.02% and \u003cem\u003eMost\u003c/em\u003e: 93.22%), consistent with their truth and informativeness. Under face threat, however, participants often judged the same maximally informative utterances as pragmatically inappropriate. In face-threatening \u003cem\u003eMost\u003c/em\u003e contexts, only about 18.75% accepted \u0026ldquo;\u003cem\u003eMost\u003c/em\u003e of \u003cem\u003eY\u003c/em\u003e did \u003cem\u003eX\u003c/em\u003e,\u0026rdquo; despite its factual correctness, plausibly because it foregrounds that not everyone succeeded. Likewise, in face-threatening \u003cem\u003eAll\u003c/em\u003e contexts, \u0026ldquo;\u003cem\u003eAll\u003c/em\u003e of \u003cem\u003eY\u003c/em\u003e did \u003cem\u003eX\u003c/em\u003e\u0026rdquo; was accepted by only 10.94% of participants, suggesting that even correct maximal informativeness can be dispreferred when it conflicts with interpersonal considerations.\u003c/p\u003e \u003cp\u003eThese findings imply a strong interaction: face threat substantially increases the acceptability of under-informative some statements, while decreasing the acceptability of fully informative statements that may be socially insensitive. A logistic mixed-effects model confirmed these patterns. Across all trials, there were significant main effects of Face Threat (face-threatening vs. neutral: OR\u0026thinsp;=\u0026thinsp;14.90, 95% CI [7.42, 30. 00], z\u0026thinsp;=\u0026thinsp;7.59, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), Factual State (\u003cem\u003eMost\u003c/em\u003e vs. \u003cem\u003eAll\u003c/em\u003e: OR\u0026thinsp;=\u0026thinsp;0.36, 95% CI [0.23, 0.56], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), and Statement Type (\u003cem\u003esome\u003c/em\u003e vs. \u003cem\u003efactual\u003c/em\u003e: OR\u0026thinsp;=\u0026thinsp;13.5, 95% CI [6.87, 26.5], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001). Critically, Face Threat interacted strongly with Statement Type (OR\u0026thinsp;=\u0026thinsp;0.00066, 95% CI [0.00023, 0.0019], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), capturing the observed crossover: face threat increased acceptance of some while decreasing acceptance of fully informative statements. To make this reversal transparent, we also modeled the two statement types separately. For \u003cem\u003esome\u003c/em\u003e statements, face threat substantially increased acceptance (OR\u0026thinsp;=\u0026thinsp;25.10, 95% CI [11.20, 56. 00], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001) and acceptance was lower in \u003cem\u003eMost\u003c/em\u003e than \u003cem\u003eAll\u003c/em\u003e contexts (OR\u0026thinsp;=\u0026thinsp;0.27, 95% CI [0.16, 0.46], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), with a modest Face Threat \u0026times; Factual State interaction (OR\u0026thinsp;=\u0026thinsp;2.85, 95% CI [1.01, 8.00], \u003cem\u003ep\u003c/em\u003e = .047). For factual statements, face threat sharply reduced acceptance (OR\u0026thinsp;=\u0026thinsp;0.0043, 95% CI [0.0017, 0.011], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), while Factual State had a smaller effect (OR\u0026thinsp;=\u0026thinsp;0.44, 95% CI [0.20, 0.97], \u003cem\u003ep\u003c/em\u003e = .041), and Face Threat \u0026times; Factual State was not significant (OR\u0026thinsp;=\u0026thinsp;0.99, 95% CI [0.35, 2.80], \u003cem\u003ep\u003c/em\u003e = .989). Together, these models show that human participants\u0026rsquo; overall three-way pattern are driven primarily by the selective licensing of under-informativeness under face threat, and the selective penalization of blunt informativeness in the same contexts.\u003c/p\u003e \u003cp\u003eAll three LLMs were partially converged with, but also differentiated from, this human profile (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea). Across systems, the Face Threat manipulation affected the acceptability of some in the same direction as in human judgments; however, the models differed markedly in calibration, both in their baseline tolerance for under-informativeness in neutral contexts and in the magnitude with which face threat licensed vagueness. DeepSeek-R1 exhibited the same directional profile but with a markedly stricter baseline and a larger face shift: it accepted some rarely in neutral contexts (\u003cem\u003eMost\u003c/em\u003e 22.50% and \u003cem\u003eAll\u003c/em\u003e 7.81%) while approaching ceiling under face threat (\u003cem\u003eMost\u003c/em\u003e 99.38% and \u003cem\u003eAll\u003c/em\u003e 99.06%), indicating a strong informativeness expectation unless vagueness is pragmatically licensed. GPT-4o likewise increased acceptance under face threat but was substantially more permissive in neutral contexts, especially in the Most condition (\u003cem\u003eMost\u003c/em\u003e 88.44% and \u003cem\u003eAll\u003c/em\u003e 44.06%), yielding a comparatively smaller face-related increase for Most but a pronounced increase for All (\u003cem\u003eMost\u003c/em\u003e 99.69% and \u003cem\u003eAll\u003c/em\u003e 100%). DeepSeek-V3.2 also showed higher acceptance under face threat, yet the modulation was weaker and remained far below the human ceiling (\u003cem\u003eMost\u003c/em\u003e 54.06% and \u003cem\u003eAll\u003c/em\u003e 15.63% for neutral contexts; \u003cem\u003eMost\u003c/em\u003e 65.00% and \u003cem\u003eAll\u003c/em\u003e 43.13% for neutral contexts). Overall, the descriptive patterns suggest that face threat licenses \u003cem\u003esome\u003c/em\u003e across systems, but models vary in how strongly they enforce informativeness in neutral contexts and how fully they capitalize on face-based licensing in face-threatening contexts.\u003c/p\u003e \u003cp\u003eFor fully informative statements, the models showed only partial pragmatic sensitivity. Like humans, they accepted factual statements at ceiling in neutral contexts (all three models: 100% in \u003cem\u003eMost\u003c/em\u003e and 100% in \u003cem\u003eAll\u003c/em\u003e). Under face threat, however, the models remained far more tolerant than humans of these blunt, maximally informative utterances. Whereas GPT-4o still accepted them at high rates (\u003cem\u003eMost\u003c/em\u003e 70.00% and \u003cem\u003eAll\u003c/em\u003e 89.38%), and DeepSeek-R1 likewise remained largely accepting (\u003cem\u003eMost\u003c/em\u003e 76.56% and \u003cem\u003eAll\u003c/em\u003e 96.56%). DeepSeek-V3.2 showed the largest drop among the models, yet still exceeded human acceptance (\u003cem\u003eMost\u003c/em\u003e 82.81% and \u003cem\u003eAll\u003c/em\u003e 52.19%). Thus, although the models captured that face threat licenses vagueness, they did not reproduce the human tendency to treat full informativeness itself as socially inappropriate in face-threatening contexts.\u003c/p\u003e \u003cp\u003eBecause LLMs produced near-deterministic responses for the fully informative statements (acceptance rates reached ceiling), the full Face Threat \u0026times; Factual State \u0026times; Statement Type model exhibited quasi-complete separation (full three-way results see Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e2\u003c/span\u003ea), yielding unstable or effectively infinite estimates for the three-way interaction. To ensure interpretable inference and to target the theoretically critical contrast\u0026mdash;whether face threat licenses under-informative utterances\u0026mdash;we therefore focus our model-based analyses on trials with \u003cem\u003esome\u003c/em\u003e statement type only (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e2\u003c/span\u003eb). In the \u003cem\u003esome\u003c/em\u003e-only models, Face Threat robustly increased acceptance for GPT-4o (OR\u0026thinsp;=\u0026thinsp;73.90, 95% CI [8.18, 668], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), and for DeepSeek-R1 (OR\u0026thinsp;=\u0026thinsp;2.24 \u0026times; 10\u003csup\u003e19\u003c/sup\u003e, 95% CI [4.20 \u0026times; 10\u003csup\u003e18\u003c/sup\u003e, 1.19 \u0026times; 10\u003csup\u003e20\u003c/sup\u003e], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), whereas DeepSeek-V3.2 did not show a reliable Face Threat main effect for some (OR\u0026thinsp;=\u0026thinsp;2.11, 95% CI [0.78, 5.71], \u003cem\u003ep\u003c/em\u003e = .141). Factual State also reliably reduced some acceptance across systems (GPT-4o: OR\u0026thinsp;=\u0026thinsp;0.013, 95% CI [0.0022, 0.073], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001; DeepSeek-R1: OR\u0026thinsp;=\u0026thinsp;0.22, 95% CI [0.13, 0.37], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001; DeepSeek-V3.2: OR\u0026thinsp;=\u0026thinsp;0.031, 95% CI [0.0097, 0.099], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001). Finally, Face Threat \u0026times; Factual State for \u003cem\u003esome\u003c/em\u003e was non-significant for DeepSeek-R1 (OR\u0026thinsp;=\u0026thinsp;2.93, 95% CI [0.45, 19.10], \u003cem\u003ep\u003c/em\u003e = .259), and pronounced for GPT-4o and DeepSeek-V3.2 (GPT-4o: OR\u0026thinsp;=\u0026thinsp;1.33 \u0026times; 10\u003csup\u003e9\u003c/sup\u003e, 95% CI [8.56 \u0026times; 10\u003csup\u003e7\u003c/sup\u003e, 2.06 \u0026times; 10\u003csup\u003e10\u003c/sup\u003e], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001; DeepSeek-V3.2: OR\u0026thinsp;=\u0026thinsp;7.36, 95% CI [2.18, 24.80], \u003cem\u003ep\u003c/em\u003e = .001).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe next asked whether processing cost tracks the pragmatic trade-off between informativeness and face management. We analyzed log-transformed response times (RTs) separately for the context-reading stage (ContextRT) and the judgment stage (JudgmentRT) as a function of Face Threat, Factual State, and Statement Type. For LLMs, we used parallel cost proxies: log-thinking tokens for DeepSeek-R1 and log-probability for non-CoT models DeepSeek-V3.2 and GPT-4o. As a result, Human ContextRT showed little systematic pragmatic modulation: face threat produced only a marginal trend toward faster reading (b\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.091, 95% CI [\u0026minus;\u0026thinsp;0.19, 0.006], \u003cem\u003ep\u003c/em\u003e = .066), with no reliable higher-order interactions. This absence of early-stage effects suggests that the pragmatic manipulation does not primarily alter initial comprehension demands. Instead, the cost signature emerged downstream at the evaluative stage. In JudgmentRT, participants judged statements faster under face-threatening (b\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.47, 95% CI [\u0026minus;\u0026thinsp;0.59, \u0026minus;\u0026thinsp;0.35], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), and\u0026mdash;at baseline\u0026mdash;factual responses were judged faster than \u003cem\u003esome\u003c/em\u003e (b\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.38, 95% CI [\u0026minus;\u0026thinsp;0.498, \u0026minus;\u0026thinsp;0.262], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), consistent with some incurring conflict when full informativeness is expected. Critically, this factual advantage was substantially reduced under face-threatening (b\u0026thinsp;=\u0026thinsp;0.44, 95% CI [0.28, 0.61], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), indicating that face threat licenses vagueness and attenuates the processing penalty for under-informativeness. Moreover, a significant three-way interaction showed that the magnitude of this reweighting depends on whether the underlying state was \u003cem\u003eAll\u003c/em\u003e versus \u003cem\u003eMost\u003c/em\u003e (b\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.29, 95% CI [\u0026minus;\u0026thinsp;0.53, \u0026minus;\u0026thinsp;0.055], \u003cem\u003ep\u003c/em\u003e = .016).\u003c/p\u003e \u003cp\u003eTurning to LLM-based cost proxies, the CoT model DeepSeek-R1 showed a highly interaction-rich token-cost profile. DeepSeek-R1\u0026rsquo;s thinking token cost was lower overall under face threat (b\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.28, 95% CI [\u0026minus;\u0026thinsp;0.38, \u0026minus;\u0026thinsp;0.17], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001) and when the true state was \u003cem\u003eAll\u003c/em\u003e (b\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.094, 95% CI [\u0026minus;\u0026thinsp;0.15, \u0026minus;\u0026thinsp;0.036], \u003cem\u003ep\u003c/em\u003e = .003), and factual responses required substantially fewer tokens at baseline (b\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.75, 95% CI [\u0026minus;\u0026thinsp;0.85, \u0026minus;\u0026thinsp;0.64], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001). Crucially, however, this baseline \u0026ldquo;factual-is-easy\u0026rdquo; advantage was strongly context-dependent: face threat markedly attenuated the token advantage for factual responses (Face Threat \u0026times; Statement Type: b\u0026thinsp;=\u0026thinsp;0.55, 95% CI [0.37, 0.74], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), and this modulation further varied with the \u003cem\u003eAll/Most\u003c/em\u003e distinction (Face Threat \u0026times; Factual State \u0026times; Statement Type: b\u0026thinsp;=\u0026thinsp;0.51, 95% CI [0.34, 0.67], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001). Notably, DeepSeek-R1 also showed reliable two-way interactions involving Factual State (Face Threat \u0026times; Factual State: b\u0026thinsp;=\u0026thinsp;0.096, 95% CI [0.0070, 0.19], \u003cem\u003ep\u003c/em\u003e = .036; Factual State \u0026times; Statement Type: b\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.14, 95% CI [\u0026minus;\u0026thinsp;0.23, \u0026minus;\u0026thinsp;0.042], \u003cem\u003ep\u003c/em\u003e = .006), underscoring that its explicit reasoning length is jointly shaped by informational state and social context. These effects suggest that when maximal informativeness can be socially risky, the CoT model\u0026rsquo;s reasoning expands from a default informativeness-driven mode toward more conditional, context-sensitive justification. By comparison, DeepSeek-V3.2\u0026rsquo;s log-probability primarily reflected an overall preference for factual responses (b\u0026thinsp;=\u0026thinsp;0.27, 95% CI [0.14, 0.40], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), with no reliable Face Threat, Factual State, or interaction effects. GPT-4o, in contrast, showed a targeted pragmatic reallocation: although both face threat (b\u0026thinsp;=\u0026thinsp;0.17, 95% CI [0.077, 0.26], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001) and factual type increased log-probability (b\u0026thinsp;=\u0026thinsp;0.17, 95% CI [0.082, 0.26], \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), face threat sharply reduced the relative log-probability advantage of factual responses (Face Threat \u0026times; Statement Type: b\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.34, 95% CI [\u0026minus;\u0026thinsp;0.47, \u0026minus;\u0026thinsp;0.20], t\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;4.96, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001). This interaction indicates that face threat changes which type of response GPT-4o favors, rather than uniformly affecting response confidence.\u003c/p\u003e \u003cp\u003eFinally, to assess whether humans and models align at the level of item-by-item variation, we computed within-cell correlations across items between human measures and model cost proxies for \u003cem\u003esome\u003c/em\u003e statements (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, Table S4). Overall, most correlations were small and non-significant, indicating limited fine-grained alignment beyond condition-level means. Nevertheless, three localized correspondences emerged. In face threat neutral contexts, human acceptance of some correlated positively with DeepSeek-R1 token cost in both the \u003cem\u003eMost\u003c/em\u003e condition (\u003cem\u003er\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.44, 95% CI [0.11, 0.68], \u003cem\u003ep\u003c/em\u003e = .011) and the \u003cem\u003eAll\u003c/em\u003e condition (\u003cem\u003er\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.47, 95% CI [0.14, 0.70], \u003cem\u003ep\u003c/em\u003e = .007). In addition, in the face-threatening \u003cem\u003eAll\u003c/em\u003e condition, human JudgmentRT correlated negatively with GPT-4o log-probability (\u003cem\u003er\u003c/em\u003e\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.51, 95% CI [\u0026minus;\u0026thinsp;0.73, \u0026minus;\u0026thinsp;0.20], \u003cem\u003ep\u003c/em\u003e = .003), suggesting shared item-level difficulty in the most pragmatically conflicted context. Given the number of tests conducted, we treat these correlations as exploratory.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTaken together, the acceptance-rate, cost, and alignment analyses converge on a conclusion: LLMs capture the direction of human pragmatic adaptation under face threat, but remain miscalibrated in the social stakes of informativeness and differ in how their internal cost signals resemble human evaluative processing. At the behavioral level, human acceptance rates exhibited a clear crossover: in neutral contexts, listeners penalized under-informative some when a stronger alternative was warranted\u0026mdash;most strongly when the underlying state supported maximal informativeness\u0026mdash;whereas under face threat, the same some statements became broadly acceptable. In complementary fashion, maximally informative factual statements were readily endorsed in neutral contexts but were strongly downgraded as pragmatically inappropriate under face threat, indicating that humans treat full informativeness itself as socially costly when it foregrounds a potentially embarrassing implication. Against this benchmark, all three models increased acceptance of some under face threat, but they did so with sharply different calibration profiles. DeepSeek-R1 implemented a stringent informativeness prior in neutral contexts and then shifted toward a human like near-ceiling acceptance under face threat. GPT-4o showed the opposite calibration: it was already permissive in neutral contexts, and thus exhibited a smaller face-driven increase in those cells while still moving toward ceiling under face threat. DeepSeek-V3.2 displayed the weakest contextualization, with only modest increases under face threat and acceptance rates remaining far below the human ceiling. Critically, none of the models reproduced the human-level rejection of blunt factual utterances in face-threatening settings; all three remained substantially more tolerant of maximally informative statements even when they are socially insensitive. Consequently, the systems capture \u0026ldquo;face threat licenses vagueness\u0026rdquo; more readily than the complementary human norm that \u0026ldquo;face threat penalizes blunt truth.\u0026rdquo;\u003c/p\u003e \u003cp\u003eThe cost analyses provide converging process-level constraints. In humans, processing cost was concentrated at the judgment stage rather than during context reading, consistent with a late evaluative reweighting of informativeness against interpersonal considerations. DeepSeek-R1 exhibited a strongly context-conditioned token-cost profile, consistent with engaging in explicit justification jointly shaped by informational state and interpersonal risk; however, its higher-order interaction pattern need not mirror human RT structure, given that CoT length and decision time are different computational currencies. GPT-4o showed a more targeted cost signature consistent with probability mass reallocating toward some specifically when face threat makes mitigation relevant, whereas DeepSeek-V3.2 showed comparatively weak contextual modulation. Finally, exploratory item-level correlations indicate that any human\u0026ndash;model coupling is localized: where correspondence appears, it tends to surface in theoretically diagnostic regions of the design rather than reflecting a global item-by-item match. Overall, in the judgement task, the models approximate the human licensing of under-informativeness under face threat, but they do not fully reproduce the human tendency to treat maximal informativeness itself as socially costly, highlighting a gap between pragmatic directionality and socially calibrated evaluative norms.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eScalar expressions such as \u003cem\u003esome\u003c/em\u003e are traditionally analyzed through the lens of informativity-based competition with stronger alternatives, yielding scalar implicatures like \u003cem\u003enot all\u003c/em\u003e (Grice, 1975; Levinson, 2000). At the same time, everyday communication routinely exploits \u003cem\u003esome\u003c/em\u003e as a face-saving device, allowing speakers to soften potentially awkward or socially risky truths (Brown \u0026amp; Levinson, 1987). By leveraging this dual function, the present study examined whether humans and LLMs exhibit similar context-dependent reweighting of informativeness and politeness, and whether any surface-level alignment reflects shared underlying evaluative mechanisms.\u003c/p\u003e \u003cp\u003eAcross two experiments, our results converge on a clear asymmetry. Humans systematically modulate both production and judgment in response to face threat, treating informativeness not as an absolute requirement but as a socially negotiable norm. LLMs, in contrast, capture the direction of this pragmatic adaptation, most notably, the increased tolerance for under-informative expressions under face-threatening conditions, but fail to reproduce its full social logic.\u003c/p\u003e \u003cp\u003eInformativeness–Politeness Tradeoff in Quantifier Use\u003c/p\u003e \u003cp\u003eExperiment 1 showed that human speakers flexibly balance informational accuracy against interpersonal concerns. In contexts without social stakes, speakers overwhelmingly adhered to informativeness, producing quantifiers that precisely matched the factual state. When face was threatened, however, they systematically weakened their utterances, most prominently by deploying some but also by shifting toward other weaker quantifiers. This pattern reflects a classic pragmatic insight: speakers may voluntarily forgo informativeness when stating the full truth risks embarrassing the addressee (Bonnefon et al., 2009; Feeney \u0026amp; Bonnefon, 2013).\u003c/p\u003e \u003cp\u003eThe LLMs exhibited only partial convergence with this human strategy. While all models increased their reliance on some under face threat, their mitigation strategies were markedly narrower. Non-CoT models largely collapsed onto a single vague option, whereas the CoT model DeepSeek-R1 more closely resembled humans in distributing responses across a range of weaker quantifiers. This suggests that explicit reasoning mechanisms can support more graded adjustments, but even then without the full diversity observed in human speech.\u003c/p\u003e \u003cp\u003eImplicature Suppression and the Social Cost of Directness\u003c/p\u003e \u003cp\u003eJudgment data from Experiment 2 extend this picture by showing that humans treat informativeness itself as socially evaluable. When no face threat was present, under-informative some statements were dispreferred, consistent with standard expectations of cooperative informativeness. Under face threat, however, the same statements became broadly acceptable, reflecting listeners’ willingness to suspend scalar expectations to preserve interpersonal harmony. Crucially, this tolerance was paired with a complementary effect: fully informative statements, though factually correct, were judged inappropriate when they foregrounded an embarrassing implication. Human listeners thus appear sensitive not only to when vagueness is licensed, but also to when blunt truthfulness incurs social cost (Yoon et al., 2020).\u003c/p\u003e \u003cp\u003eThe LLMs showed only partial alignment with this evaluative pattern. All three models increased acceptance of some under face threat, indicating sensitivity to the idea that vagueness can be appropriate in socially delicate contexts. However, none of the models reliably penalized maximally informative statements in the same situations. In other words, the models captured the heuristic that “face threat licenses vagueness,” but failed to internalize the complementary human norm that “face threat can make full informativeness inappropriate.” This asymmetry suggests that current LLMs treat politeness primarily as an additive strategy—introducing hedging when warranted—rather than as a constraint that can actively disfavor directness.\u003c/p\u003e \u003cp\u003eFrom Informativeness to Face Management: Mechanism-Level Divergence\u003c/p\u003e \u003cp\u003eAcross completion and judgment, the human data point to a unified mechanism in which scalar enrichment is not a fixed by-product of lexical alternatives, but a contextually gated inference that is reweighted by social stakes. In neutral contexts, participants behaved as predicted by informativity-driven accounts: weak utterances \u003cem\u003esome\u003c/em\u003e were treated as pragmatically deficient when a stronger true alternative was available, consistent with Gricean competition and scalar strengthening. Under face threat, however, both speakers and listeners systematically shifted the weighting of constraints: speakers licensed underinformativeness as a face-saving hedge, and listeners suppressed the usual “not all” pressure and instead treated vagueness as socially appropriate. This pattern is consistent with broader evidence that scalar implicatures are effortful and context-sensitive rather than default and obligatory (Bott \u0026amp; Noveck, 2004), and with classic politeness accounts in which speakers manage face via strategic attenuation of commitment (Bonnefon et al., 2009; Brown \u0026amp; Levinson, 1987).\u003c/p\u003e \u003cp\u003eThe critical question is whether LLMs implement a similar contextually gated inference. Mechanistically, this predicts two coupled signatures under face threat: an increased tendency to produce and accept some, and a concomitant decrease in the perceived appropriateness of fully informative alternatives. The models reliably showed the first signature. They were far less reliable on the second—and this is arguably the harder part. Suppressing full informativeness is less overt and harder to cue than producing vagueness: it requires an implicit social norm that “saying the whole truth” can be dispreferred when it threatens face, even though it is factually correct. Consistent with this, the models often continued to endorse blunt factual statements as acceptable in face-threatening contexts, indicating limited sensitivity to the social penalty attached to directness. This pattern suggests that models may be learning stylistic mitigation without learning the deeper pragmatic norm that makes hedging rational for humans. More fundamentally, these patterns are consistent with the view that contemporary LLMs lack stable, causally grounded representations of interlocutors’ mental states and social incentives (Bender et al., 2021). Their apparent politeness adjustments plausibly reflect learned associations between linguistic forms and contextual cues in training data, rather than an underlying social-motivational model. As a result, when contextual cues are sparse or when interaction requires higher-order perspective-taking beyond the training distribution, models may revert to more literal, semantics-biased responses. This raises an important generalization challenge: whether surface-level pragmatic sensitivity observed in controlled, single-shot prompts can scale to real-world dialogue, where face concerns are often implicit, distributed across turns, and negotiated dynamically. On this view, current LLMs exhibit shallow social-pragmatic adaptation, but do not yet match the human capacity for deep, socially regulated pragmatic control.\u003c/p\u003e \u003cp\u003eChain-of-Thought Optimization and Pragmatic Flexibility\u003c/p\u003e \u003cp\u003eThe contrast between the CoT-based DeepSeek-R1 and the non-CoT DeepSeek-V3.2 is informative not because R1 merely produces more human-like outputs across production, judgment, and processing cost, but because it reveals how different computational regimes support pragmatic flexibility. Specifically, compared with non-CoT architectures, DeepSeek-R1 does not simply amplify shifts in output preferences; instead, it exhibits qualitatively different patterns in how contextual information constrains the generation process. Variation in reasoning length across contexts indicates that CoT generation adapts to how tightly the context constrains the decision space. When pragmatic considerations strongly favor one response type (e.g., face-threatening contexts that license vagueness), the model converges more quickly; when multiple responses remain plausible (e.g., neutral contexts prioritizing informativeness), longer reasoning trajectories are observed. This pattern is more like a property of search and constraint satisfaction in the generation process, rather than as evidence of human-like deliberation (De Varda et al., 2025).\u003c/p\u003e \u003cp\u003eFrom this perspective, the advantage of CoT lies not in enhancing “reasoning” itself, but in providing intermediate representations that support multiple context-sensitive paths through the decision space. Importantly, this flexibility does not require assuming that the model explicitly represents social norms or reasons about face in a human-like way. Instead, explicit reasoning trajectories function as a structural scaffold through which informational and interpersonal cues can jointly constrain generation. In this sense, CoT optimization makes it easier for contextual factors to exert differentiated influence, yielding behavior that more closely approximates human pragmatic flexibility without implying shared underlying cognition (Wei et al., 2022).\u003c/p\u003e \u003cp\u003eLimitations and Future Directions\u003c/p\u003e \u003cp\u003eOur study has several limitations that suggest avenues for future work. First, our understanding of LLMs’ internal decision processes remains indirect. It will be important to develop better interpretability tools to see whether concepts like “face” or “implicature” are truly represented. Second, our experiments were conducted in Mandarin Chinese with contrived scenarios; pragmatic norms can vary across languages and cultures. It remains to be seen whether LLMs trained in other languages show similar or different patterns, and whether they can capture culture-specific politeness strategies. Third, real-world conversation is richer than our one-shot tasks. Humans use intonation, social context, and history to guide politeness; current LLMs mostly rely on a static prompt. Future work should test models in multi-turn dialogues or interactive settings, perhaps with simulated social feedback. Adding an explicit “theory of mind” component, or training on dialogues with nuanced interpersonal goals, could also narrow the gap.\u003c/p\u003e \u003cp\u003eMore broadly, our human–model comparison framework can be applied to other pragmatic phenomena. Just as LLMs serve as engineering tools, they are now testbeds for cognitive theories. By examining when models exhibit or fail at pragmatics, we can both refine our understanding of human language and identify ways to improve AI. Our results reinforce the view that human implicature is deeply influenced by social motives, whereas LLMs currently treat pragmatic effects more superficially. As language models continue to advance, we hope they will increasingly grasp when to be tactful and when to be direct – and perhaps even explain their choices. Achieving that level of social-intelligence in machines would not only enhance human–AI communication, but would also mark a significant step toward mirroring human pragmatic competence.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003ctable id=\"Tab1\" border=\"1\"\u003e \u003ccaption\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cem\u003eExperiment 1: regression results.\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003c/colgroup\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\"\u003e \u003cp\u003eSystem\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eEffect\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eOR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003e95% CI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003ez\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003ep\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e12.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[8.17, 20.2]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e11.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.35\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.178, 0.684]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-3.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.002\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e3.23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[1.53, 6.82]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e3.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.002\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e73.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[8.46, 632.]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e3.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; 0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[3.760e-09, 3.600e-07]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-14.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e5.85e + 07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[7.290e + 06, 4.690e + 08]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e16.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e2.03e + 20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[5.860e + 19, 7.040e + 20]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e73.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; 0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[2.160e-21, 2.010e-15]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-11.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.12e + 18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[6.200e + 16, 2.030e + 19]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e28.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e5.16e + 19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[1.030e + 19, 2.570e + 20]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e55.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; 0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[1.920e-09, 6.330e-08]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-20.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.08e + 07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[7.180e + 05, 1.630e + 08]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e11.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"6\"\u003e\u003cem\u003eNote.\u003c/em\u003e OR = odds ratio. CI = 95% confidence interval. ORs are rounded to two decimals; very large ORs are shown in scientific notation. p values are two-tailed.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003cp\u003e\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003ctable id=\"Tab2\" border=\"1\"\u003e \u003ccaption\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cb\u003ea\u003c/b\u003e \u003cem\u003eExperiment 2: three-way regression result\u003c/em\u003es.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003c/colgroup\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\"\u003e \u003cp\u003eSystem\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eEffect\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eOR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003e95% CI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003ez\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003ep\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e14.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[7.42, 30.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e7.58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.230, 0.559]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-4.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eStatement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e13.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[6.87, 26.5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e7.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e2.29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.883, 5.92]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.089\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e6.59e-04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.000234, 0.00186]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-13.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.553, 3.12]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.62\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.536\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.119, 1.79]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-1.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.263\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.32e + 03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[148., 1.180e + 04]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e6.44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.154, 0.400]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-5.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eStatement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e9.63e + 09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[6.010e + 09, 1.540e + 10]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e95.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e2.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.471, 15.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.268\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.73e-11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[1.490e-12, 2.000e-10]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-19.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e4.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[2.50, 6.48]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e5.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.00370, 0.279]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-3.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.002\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.833, 3.62]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.141\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.0500, 0.236]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-5.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eStatement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e4.89e + 08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[2.670e + 08, 8.960e + 08]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e64.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e3.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[1.40, 6.90]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e2.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.005\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e6.29e-09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[2.710e-09, 1.460e-08]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-43.89\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e9.21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[NA, NA]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e——\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e——\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.0378, 0.0798]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-15.22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e50.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[7.30, 356.]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e3.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.0169, 0.148]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-5.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eStatement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.61e + 08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[NA, NA]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e——\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e——\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e6.31e + 07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[8.040e + 06, 4.950e + 08]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e17.09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.36e-10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[2.150e-11, 8.660e-10]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e-24.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e20.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[NA, NA]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e——\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e——\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state × Statement type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e3.28e-09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[NA, NA]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e——\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e——\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"6\"\u003e\u003cem\u003eNote.\u003c/em\u003e OR = odds ratio. CI = 95% confidence interval. ORs are rounded to two decimals; very large (≥ 1,000) or very small (\u0026lt; 0.01) ORs are shown in scientific notation. p values are two-tailed.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003cp\u003e\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003ctable id=\"Tab3\" border=\"1\"\u003e \u003ccaption\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cb\u003eb\u003c/b\u003e \u003cem\u003eExperiment 2: two-way regression results (Statement Type = some).\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003c/colgroup\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\"\u003e \u003cp\u003eSystem\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eEffect\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eOR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003e95% CI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003ez\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003ep\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e25.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[11.2, 56.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e7.86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.162, 0.457]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-4.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e2.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[1.01, 8.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e1.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.047\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e2.24e + 19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[4.200e + 18, 1.190e + 20]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e52.23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.126, 0.371]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-5.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e2.93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.452, 19.1]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e1.13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.259\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e2.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.780, 5.71]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e1.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.141\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.00968, 0.0985]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-5.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e7.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[2.18, 24.8]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e3.22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e73.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[8.18, 668.]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e3.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[0.00217, 0.0727]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-4.89\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e1.33e + 09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e[8.560e + 07, 2.060e + 10]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e15.02\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"6\"\u003e\u003cem\u003eNote.\u003c/em\u003e OR = odds ratio. CI = 95% confidence interval. ORs are rounded to two decimals; very large (≥ 1,000) or very small (\u0026lt; 0.01) ORs are shown in scientific notation. p values are two-tailed.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003cp\u003e\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003ctable id=\"Tab4\" border=\"1\"\u003e \u003ccaption\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cem\u003eExperiment 2 cost indicators: regression results (effect labels corrected).\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003c/colgroup\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\"\u003e \u003cp\u003eSystem\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eCost Indicator\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eEffect\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eb\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eCI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003et\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003ep\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eContextRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eIntercept\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e8.191\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[7.785, 8.598]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e40.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eContextRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.091\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.187, 0.006]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-1.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.066\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eContextRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.062\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.035, 0.158]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e1.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.209\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eContextRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eResponse type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.079\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.017, 0.176]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e1.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.108\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eContextRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.011\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.148, 0.125]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.870\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eContextRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.083\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.219, 0.054]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-1.19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.234\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eContextRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.094\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.23, 0.043]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-1.35\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.177\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eContextRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.026\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.167, 0.219]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.788\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eJudgmentRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eIntercept\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e7.294\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[6.967, 7.621]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e44.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eJudgmentRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.588, -0.351]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-7.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eJudgmentRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.046\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.164, 0.072]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.445\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eJudgmentRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eResponse type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.38\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.498, -0.262]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-6.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eJudgmentRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.059\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.108, 0.226]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.488\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eJudgmentRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.444\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[0.276, 0.611]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e5.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eJudgmentRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.163\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.005, 0.33]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e1.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.057\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eHuman\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eJudgmentRT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.291\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.528, -0.055]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-2.41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.016\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogTokens\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e8.644\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[8.353, 8.934]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e59.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogTokens\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.276\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.352, -0.2]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-7.13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogTokens\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eResponse type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.002\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.078, 0.074]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.967\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogTokens\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat× Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.093\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.169, -0.017]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-2.41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.016\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogTokens\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.042\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.065, 0.149]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.77\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.442\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogTokens\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.112\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[0.005, 0.22]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e2.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.040\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-R1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogTokens\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.014\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.093, 0.122]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.794\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.103\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.255, 0.049]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-1.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.183\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.275\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.376, -0.174]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-5.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eResponse type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.095\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.153, -0.036]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-3.29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.003\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.748\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.856, -0.64]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-14.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.096\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[0.007, 0.185]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e2.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.036\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.553\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[0.366, 0.741]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e6.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eDeepSeek-V3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.135\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.229, -0.042]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-2.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.006\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.507\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[0.342, 0.672]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e6.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.093\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.247, 0.061]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-1.24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.226\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eResponse type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.088\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.083, 0.259]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e1.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.303\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.268\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[0.138, 0.399]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e4.19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e\u0026lt; .001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.004\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.196, 0.204]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e0.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.965\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFactual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.053\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.257, 0.152]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.603\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003elogprob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003eFace threat × Factual state × Response type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-0.088\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e[-0.259, 0.083]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\"\u003e \u003cp\u003e-1.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e.303\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"7\"\u003e\u003cem\u003eNote.\u003c/em\u003e b = fixed-effect estimate. CI = 95% confidence interval. Face threat is coded as Yes vs No; Factual state is All vs Most. Response type contrasts factual vs some. p values are two-tailed.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003cp\u003e\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eAdditional Information\u003c/h2\u003e \u003cp\u003eThe author declares no competing interests.\u003c/p\u003e \u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThe author received no specific funding for this work.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eYuying Wang solely conceived and designed the study, conducted the experiments, collected human behavioral data, simulated and analyzed LLM outputs, and wrote and revised the manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets generated and analyzed during the current study are available at OSF [https://osf.io/m57jg/overview?view\\_only=fc899334d57948e7b886da614e92688f], and can be accessed upon reasonable request. The data include human behavioral data, LLM output data, and analysis scripts used in this study.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBates, D., M\u0026auml;chler, M., Bolker, B., \u0026amp; Walker, S. (2015). Fitting Linear Mixed-Effects Models Using lme4. \u003cem\u003eJournal of Statistical Software\u003c/em\u003e, \u003cem\u003e67\u003c/em\u003e(1). https://doi.org/10.18637/jss.v067.i01\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBender, E. M., Gebru, T., McMillan-Major, A., \u0026amp; Shmitchell, S. (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? \u0026#129436;. \u003cem\u003eProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT \u0026rsquo;21\u003c/em\u003e, 610\u0026ndash;623. https://doi.org/10.1145/3442188.3445922\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBonnefon, J.-F., Feeney, A., \u0026amp; De Neys, W. (2011). The Risk of Polite Misunderstandings. \u003cem\u003eCurrent Directions in Psychological Science\u003c/em\u003e, \u003cem\u003e20\u003c/em\u003e(5), 321\u0026ndash;324. https://doi.org/10.1177/0963721411418472\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBonnefon, J.-F., Feeney, A., \u0026amp; Villejoubert, G. (2009). When some is actually all: Scalar inferences in face-threatening contexts. \u003cem\u003eCognition\u003c/em\u003e, \u003cem\u003e112\u003c/em\u003e(2), 249\u0026ndash;258. https://doi.org/10.1016/j.cognition.2009.05.005\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBott, L., \u0026amp; Noveck, I. A. (2004). Some utterances are underinformative: The onset and time course of scalar inferences. \u003cem\u003eJournal of Memory and Language\u003c/em\u003e, \u003cem\u003e51\u003c/em\u003e(3), 437\u0026ndash;457. https://doi.org/10.1016/j.jml.2004.05.006\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrown, P., \u0026amp; Levinson, S. C. (1987). \u003cem\u003ePoliteness: Some universals in language usage\u003c/em\u003e (Vol. 4). Cambridge university press.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChemla, E., \u0026amp; Bott, L. (2014). Processing inferences at the semantics/pragmatics frontier: Disjunctions and free choice. \u003cem\u003eCognition\u003c/em\u003e, \u003cem\u003e130\u003c/em\u003e(3), 380\u0026ndash;396. https://doi.org/10.1016/j.cognition.2013.11.013\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChemla, E., \u0026amp; Singh, R. (2014). Remarks on the Experimental Turn in the Study of Scalar Implicature, Part I. \u003cem\u003eLanguage and Linguistics Compass\u003c/em\u003e, \u003cem\u003e8\u003c/em\u003e(9), 373\u0026ndash;386. https://doi.org/10.1111/lnc3.12081\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCho, Y., \u0026amp; Kim, S. mook. (2024). \u003cem\u003ePragmatic inference of scalar implicature by LLMs\u003c/em\u003e (Version 1). arXiv. https://doi.org/10.48550/ARXIV.2408.06673\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCong, Y. (2024). Manner implicatures in large language models. Scientific Reports, 14(1), 29113. https://doi.org/10.1038/s41598-024-80571-3\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDe Varda, A. G., D\u0026rsquo;Elia, F. P., Kean, H., Lampinen, A., \u0026amp; Fedorenko, E. (2025). The cost of thinking is similar between large reasoning models and humans. \u003cem\u003eProceedings of the National Academy of Sciences\u003c/em\u003e, \u003cem\u003e122\u003c/em\u003e(47), e2520077122. https://doi.org/10.1073/pnas.2520077122\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFeeney, A., \u0026amp; Bonnefon, J.-F. (2013). Politeness and Honesty Contribute Additively to the Interpretation of Scalar Expressions. \u003cem\u003eJournal of Language and Social Psychology\u003c/em\u003e, \u003cem\u003e32\u003c/em\u003e(2), 181\u0026ndash;190. https://doi.org/10.1177/0261927X12456840\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrice, H. P. (1975). Logic and conversation. \u003cem\u003eSyntax and Semantics\u003c/em\u003e, \u003cem\u003e3\u003c/em\u003e, 43\u0026ndash;58.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo, D., Yang, D., Zhang, Haowei, Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, Ruoyu, Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., \u0026hellip; Zhang, Z. (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. \u003cem\u003eNature\u003c/em\u003e, \u003cem\u003e645\u003c/em\u003e(8081), 633\u0026ndash;638. https://doi.org/10.1038/s41586-025-09422-z\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHorn, L. R. (1972). \u003cem\u003eOn the semantic properties of logical operators in English\u003c/em\u003e. University of California, Los Angeles.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHorn, L. R. (2006). The Border Wars: A Neo-Gricean Perspective. In K. Von Heusinger \u0026amp; K. Turner (Eds.), \u003cem\u003eWhere Semantics meets Pragmatics\u003c/em\u003e (pp. 21\u0026ndash;48). BRILL. https://doi.org/10.1163/9780080462608_006\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKatsos, N., \u0026amp; Cummins, C. (2010). Pragmatics: From Theory to Experiment and Back Again. \u003cem\u003eLanguage and Linguistics Compass\u003c/em\u003e, \u003cem\u003e4\u003c/em\u003e(5), 282\u0026ndash;295. https://doi.org/10.1111/j.1749-818X.2010.00203.x\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKu, H. B. (2025). Scaling Implicature via Structured Interaction in Multi-Agent LLMs. \u003cem\u003eIn Proc. 1st Workshop on Integrating NLP and Psychology to Study Social Interactions at AAAI Int. Conf. Weblogs and Social Media (ICWSM)\u003c/em\u003e. In Proc. 1st Workshop on Integrating NLP and Psychology to Study Social Interactions at AAAI Int. Conf. Weblogs and Social Media (ICWSM).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLevinson, S. C. (2000). \u003cem\u003ePresumptive meanings: The theory of generalized conversational implicature\u003c/em\u003e. MIT press.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMa, B., Li, Y., Zhou, W., Gong, Z., Liu, Y. J., Jasinskaja, K., Friedrich, A., Hirschberg, J., Kreuter, F., \u0026amp; Plank, B. (2025). \u003cem\u003ePragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges\u003c/em\u003e (arXiv:2502.12378). arXiv. https://doi.org/10.48550/arXiv.2502.12378\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQiu, Z., Duan, X., \u0026amp; Cai, Z. G. (2023). \u003cem\u003ePragmatic Implicature Processing in ChatGPT\u003c/em\u003e. PsyArXiv. https://doi.org/10.31234/osf.io/qtbh9\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShulginov, V., Şimşek, H. B., Kudriashov, S., Randautsova, R., \u0026amp; Shevela, S. A. (2025). \u003cem\u003eEvaluating the Pragmatic Competence of Large Language Models in Detecting Mitigated and Unmitigated Types of Disagreement\u003c/em\u003e. \u003cem\u003e2025\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, J. (2025). \u003cem\u003eA Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1\u003c/em\u003e (arXiv:2502.10867). arXiv. https://doi.org/10.48550/arXiv.2502.10867\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, brian, Xia, F., Chi, E., Le, Q. V., \u0026amp; Zhou, D. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, \u0026amp; A. Oh (Eds.), \u003cem\u003eAdvances in Neural Information Processing Systems\u003c/em\u003e (Vol. 35, pp. 24824\u0026ndash;24837). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu, S., Yang, S., Chen, Z., \u0026amp; Su, Q. (2024). Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference Tuning. \u003cem\u003eProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing\u003c/em\u003e, 22583\u0026ndash;22599. https://doi.org/10.18653/v1/2024.emnlp-main.1258\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYoon, E. J., Tessler, M. H., Goodman, N. D., \u0026amp; Frank, M. C. (2020). Polite Speech Emerges From Competing Social Goals. \u003cem\u003eOpen Mind\u003c/em\u003e, \u003cem\u003e4\u003c/em\u003e, 71\u0026ndash;87. https://doi.org/10.1162/opmi_a_00035\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu, K., Zeng, Q., Xuan, W., Li, W., Wu, J., \u0026amp; Voigt, R. (2025). \u003cem\u003eThe Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models\u003c/em\u003e (Version 3). arXiv. https://doi.org/10.48550/ARXIV.2505.18497\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYue, S., Song, S., Cheng, X., \u0026amp; Hu, H. (2024). \u003cem\u003eDo Large Language Models Understand Conversational Implicature\u0026mdash;A case study with a chinese sitcom\u003c/em\u003e (arXiv:2404.19509). arXiv. https://doi.org/10.48550/arXiv.2404.19509\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, J., \u0026amp; Wu, Y. (2020). Only youxie think it is a nice thing to say: Interpreting scalar items in face-threatening contexts by native Chinese speakers. \u003cem\u003eJournal of Pragmatics\u003c/em\u003e, \u003cem\u003e168\u003c/em\u003e, 19\u0026ndash;35. https://doi.org/10.1016/j.pragma.2020.06.008\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, H., \u0026amp; Hawkins, R. D. (2025). Comparing human and LLM politeness strategies in free production. \u003cem\u003eProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing\u003c/em\u003e, 16199\u0026ndash;16227. https://doi.org/10.18653/v1/2025.emnlp-main.820\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"scalar implicature, politeness mitigation, quantifiers, pragmatics, large language models, chain-of-thought","lastPublishedDoi":"10.21203/rs.3.rs-8763414/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8763414/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eScalar implicatures link semantic meaning to pragmatic reasoning: although \u003cem\u003esome\u003c/em\u003e is logically compatible with all, listeners often enrich \u003cem\u003esome\u003c/em\u003e to \u0026ldquo;some but not all.\u0026rdquo; Yet \u003cem\u003esome\u003c/em\u003e also functions as a politeness hedge that mitigates face-threatening acts, potentially canceling the usual \u0026ldquo;not all\u0026rdquo; inference. This study investigates whether large language models (LLMs) exhibit the same context-sensitive trade-off between informativeness and politeness as humans, and whether chain-of-thought (CoT) optimization yields more human-like pragmatic flexibility. Mandarin-speaking participants completed parallel production and judgment tasks that manipulated face threat (face-threatening vs. neutral) and factual state (\u003cem\u003eAll\u003c/em\u003e vs. \u003cem\u003eMost\u003c/em\u003e). In Experiment 1, participants completed sentences among seven quantifiers, including \u003cem\u003esome\u003c/em\u003e in Experiment 2, they evaluated the acceptability of under-informative \u003cem\u003esome\u003c/em\u003e statements versus fully informative factual statements, with reaction times recorded. Three LLMs (DeepSeek-V3.2, DeepSeek-R1, GPT-4o) were tested with identical materials. Humans showed a robust informativeness\u0026ndash;politeness reweighting: face threat increased production and acceptance of \u003cem\u003esome\u003c/em\u003e while sharply reducing acceptance of blunt, fully informative statements. LLMs generally captured the face-based licensing of vagueness but did not reliably penalize socially costly maximal informativeness. The CoT-optimized model showed greater context sensitivity and more human-like distributional patterns than the non-CoT models. Together, these findings are consistent with partial alignment with human mitigation but incomplete social-norm calibration, with CoT optimization offering a modest reduction in the mismatch.\u003c/p\u003e","manuscriptTitle":"When Do LLMs Say “Some”? An Investigation of Scalar Implicature and Politeness Mitigation in Large Language Models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-18 07:10:16","doi":"10.21203/rs.3.rs-8763414/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-04-10T21:46:03+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"241725510776853987204700667127181830469","date":"2026-03-15T13:26:17+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-11T16:03:26+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-05T10:19:59+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-03T07:11:01+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-03T07:09:44+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-02-02T09:46:17+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7467d9bb-8d2b-4bcb-ada1-a576c206a25d","owner":[],"postedDate":"March 18th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":64526512,"name":"Humanities/Philosophy"},{"id":64526513,"name":"Biological sciences/Psychology"},{"id":64526514,"name":"Social science/Psychology"},{"id":64526515,"name":"Social science/Science technology and society"}],"tags":[],"updatedAt":"2026-03-18T07:10:21+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-18 07:10:16","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8763414","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8763414","identity":"rs-8763414","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00