Large language models outperform humans at estimating society's everyday norms—but hybrids are even better

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract As AI assistants and social robots enter human environments, their ability to navigate context-dependent social norms is essential to avoid harm and ensure successful collaboration. We evaluate six large language models (LLMs) on their ability to estimate American social norms across 555 everyday scenarios (measured in prior work) and compare these to estimates from 320 humans. LLMs achieve remarkably high accuracy, clearly outperforming the average human. However, the errors LLMs make are systematic; they are similar across runs of the same LLM and even across different LLMs. As a consequence of this homogeneity, aggregating estimates of LLMs produces little improvement. Individual humans make much worse estimates, often defaulting to extreme right-or-wrong judgments even when asked to estimate population averages, but their errors are idiosyncratic and, consequently, aggregating their estimates yields dramatic improvement through wisdom-of-crowds effects. As humans make different errors than LLMs, hybrid ensembles combining both substantially outperform either alone.
Full text 98,438 characters · extracted from preprint-html · click to expand
Large language models outperform humans at estimating society's everyday norms—but hybrids are even better | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Large language models outperform humans at estimating society's everyday norms—but hybrids are even better Kimmo Eriksson, Simon Karlsson, Irina Vartanova, Pontus Strimling This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9161239/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract As AI assistants and social robots enter human environments, their ability to navigate context-dependent social norms is essential to avoid harm and ensure successful collaboration. We evaluate six large language models (LLMs) on their ability to estimate American social norms across 555 everyday scenarios (measured in prior work) and compare these to estimates from 320 humans. LLMs achieve remarkably high accuracy, clearly outperforming the average human. However, the errors LLMs make are systematic; they are similar across runs of the same LLM and even across different LLMs. As a consequence of this homogeneity, aggregating estimates of LLMs produces little improvement. Individual humans make much worse estimates, often defaulting to extreme right-or-wrong judgments even when asked to estimate population averages, but their errors are idiosyncratic and, consequently, aggregating their estimates yields dramatic improvement through wisdom-of-crowds effects. As humans make different errors than LLMs, hybrid ensembles combining both substantially outperform either alone. Scientific community and society/Social sciences/Psychology/Human behaviour Physical sciences/Mathematics and computing/Information technology Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Social robots and AI assistants are rapidly entering human environments, providing companionship in care facilities, assisting customers in retail settings, and collaborating with workers in offices. To function effectively in these roles, such systems must navigate the vast landscape of context-dependent social norms: knowing that laughter is welcome at a party but inappropriate at a funeral, that physical proximity acceptable among friends may violate boundaries with strangers, and that behaviors appropriate in one culture may give offense in another. When AI systems misunderstand these unwritten rules, the consequences range from awkward interactions to genuine harm (1, 2). Understanding how well large language models grasp social norms and where they systematically fail has thus become a question of both theoretical and practical importance. Following quantitative approaches to norm research (3-5), we operationalize a social norm as the average appropriateness judgment in a population - how appropriate members of a society collectively consider a behavior in a given context. This operationalization captures information that is helpful for an agent to avoid conflict. An AI system that accurately estimates the population average can calibrate its behavior and advice to match typical societal expectations, minimizing the risk of causing offense in situations where the expectations are not otherwise known. By contrast, a system whose estimates deviate systematically from the population average will more often misjudge what others find acceptable. Crucially, this is an estimation task: the question is not whether an LLM has its own normative views, but whether it can accurately infer what a population of people collectively considers appropriate. For measures of what a population of people collectively considers appropriate, we draw on a large-scale dataset of appropriateness ratings for 555 everyday scenarios in the United States (6). The scenarios were constructed by systematically pairing 37 common behaviors with 15 common situations, and each scenario was rated by a large sample of U.S. participants on a scale from 0 (extremely inappropriate) to 9 (extremely appropriate). Figure 1 illustrates the combinatorial structure of the 555 scenarios, with behaviors as rows and contexts as columns, and the average appropriateness rating of each scenario. These ratings provide stable estimates of collective U.S. norms that serve as our estimation target: the measured norms that LLMs and human estimators will be tested against. Importantly, the original paper on this data did not present ratings per scenario, so the estimation target is unlikely to have been in the LLM training data. Fig. 1. The estimation target: everyday norms from a prior U.S. study (6). Heatmap of mean appropriateness ratings for 555 scenarios constructed by pairing 37 behaviors (rows) with 15 contexts (columns). Behaviors and contexts are ordered by their overall mean ratings. Brighter colors indicate more appropriate scenarios. To our knowledge, no studies address how LLMs estimate social norms, but there is a body of research on how people estimate, and sometimes misestimate, them (7). Our first question is whether LLMs are better or worse at estimating social norms than humans are. There are arguments on both sides. Theories of embodied cognition propose that social understanding fundamentally requires lived experience: having navigated awkward silences, felt the sting of social rejection, or learned through trial and error which behaviours are acceptable in which contexts (8-10). LLMs lack all such experience. However, the vast scale of training data may allow them to extract normative patterns from how behaviors are discussed across countless contexts, and distributional approaches suggest that social knowledge can emerge from co-occurrence patterns in language alone (11). Humans, for their part, encounter only a fraction of possible situations through direct experience. Moreover, their judgments may be colored by the well-documented false consensus effect, that is, the tendency to assume others share one's own views (12, 13), or they may underestimate such agreement as in pluralistic ignorance (7). In sum, two competing hypotheses can be formulated. Hypothesis 1 (Accuracy): At the task of estimating the average appropriateness ratings for everyday scenarios in the U.S., the accuracy of LLMs is (a) worse than humans, or (b) better than humans. A second question is whether ensemble estimation improves accuracy. If different LLMs share similar training distributions and optimization objectives, their errors should be correlated across scenarios, making them poor candidates for ensemble improvement. Recent work suggests this is indeed the case. Jiang et al. (14) documented a pervasive "Artificial Hivemind" effect: across more than 70 language models, outputs converged on strikingly similar responses both within models and across different architectures, even for open-ended creative tasks. This homogeneity extends beyond creative generation to domains requiring social and ethical reasoning (15, 16). It is plausible, but not known, that this homogeneity could extend also to inferences of social norms, in which case the diversity required for wisdom-of-the-crowd effects may be absent in LLM ensembles. Humans, by contrast, may show a different profile. While individuals in our study share a common cultural framework, they encounter different situations through their personal histories. If these idiosyncratic experiences generate sufficiently independent errors across scenarios, aggregating estimates should substantially outperform even the best individual. The wisdom-of-crowds literature demonstrates that aggregating independent judgments often outperforms even the best individual, because diverse errors cancel out (17-19). If humans and LLMs differ on error independence—with LLM errors being correlated and human errors being independent—the best accuracy might be achieved by human-LLM hybrids that capitalize on complementary strengths. Indeed, Abels and Lenaerts (20), in a study of news headline authenticity judgments, found that simply averaging responses from multiple LLMs failed to improve performance due to limited diversity, but that hybrid crowds combining LLMs with humans capitalized on complementary strengths. In sum, the above considerations yield a set of related predictions: Hypothesis 2 (Aggregation and Hybrids): (a) Human errors are independent across individuals. (b) LLM errors are correlated across models. (c) Aggregating human estimates yields substantial improvement in accuracy. (d) Aggregating LLM estimates yields little improvement in accuracy. (e) Human-LLM hybrids achieve the best accuracy. We test these hypotheses by giving the task of estimating the measured norm for each of the 555 scenarios to six large language models available in January 2026: three proprietary frontier models (GPT-5.2, Claude Sonnet 4.5, Gemini 3 Flash) and three smaller open-weight models (Llama 3.1 8B, Llama 3.1 70B, DeepSeek V3). We also gave the same task to 320 U.S. participants, except to avoid fatigue each participant only estimated a random subset of 50 scenarios. Our results support Hypothesis 1b and Hypothesis 2. LLMs achieve remarkably high accuracy, substantially outperforming most individual humans. At the same time, LLM errors are correlated across models, preventing improvement through aggregation, while human errors are independent, enabling powerful wisdom-of-crowds effects. Human-LLM hybrids substantially outperform either approach alone. We end the paper by discussing the broad implications this has for designing AI systems that must navigate complex social environments. Results Hypothesis 1: LLM vs Human Accuracy Our first question concerns whether LLMs are better or worse than humans at estimating population-average appropriateness ratings. To ensure a fair comparison with individual human estimators, we evaluate each LLM based on the mean error it makes in a single run (averaged over 15 runs). We find that LLMs, especially the largest ones, tend to be better (supporting Hypothesis 1b over Hypothesis 1a). Figure 2 shows mean absolute error (MAE) for each individual and each LLM. Two of the proprietary LLMs outperform even the best human participant (MAEs = 0.78 and 0.85 vs. 0.92). The accuracy of the open-weight models is not as good (MAEs = 1.22–1.62) but still better than the average human participant's (MAE = 1.77). All six LLMs significantly outperformed the average human (all p < .001, two-sample Welch t-tests; see Supplementary Table 1 for detailed statistics). Fig. 2. LLMs are more accurate than humans at estimating measured norms. The mean absolute error (MAE) of single-run (k = 1) estimates of 555 social norms made by six LLMs and 320 human participants. The dashed line indicates the average among human participants. The standard errors of MAE estimates are very small for the LLMs (mean SE = 0.01). Every human participant rated a random subset of 50 scenarios; their MAE is adjusted for subset difficulty (see Methods), and their standard errors are not as small (mean SE = 0.20). Why do humans show lower accuracy? To understand this, we examined how distributions of estimated norms compare to the distribution of measured norms (Figure 3). The estimates of LLMs (orange line) have a distribution that broadly tracks the true norms (purple line). By contrast, individual human estimators (green line) strongly overestimate extreme values and rarely make estimates in the middle range. We term this phenomenon categorical bias : even when explicitly asked to estimate continuous population averages, individuals default to binary moral judgments classifying behaviors as simply wrong or right. This is likely a false consensus effect, that is, participants tend to have a categorical view of the appropriateness of a scenario and are unaware of how common it is that others have a different view. We argue that humans' categorical bias explains why they are outperformed by LLMs. In support of this interpretation, individuals who made more endpoint estimates (0-1 or 8-9) showed significantly higher error, r(318) = 0.29, 95% CI [0.18, 0.38], p < 0.001. (For example, the outlier respondent at the right end of Fig. 2 made endpoint estimates 88% of the time.) Thus, better norm estimation accuracy is associated with resisting the categorical bias by making more graded, nuanced estimates. Fig. 3. Comparing the Distributions of Estimated and Measured Norms. The proportions falling in each unit-width bin (0–1, 1–2, ..., 8–9) for true norms (purple) are almost perfectly tracked by single-run LLM estimates (k = 1; yellow), whereas human participants grossly overestimate the frequency of norms close to the scale endpoints (green). Hypothesis 2: Aggregation and Hybrids Hypothesis 2 concerns the potential for improving norm estimates by aggregating several estimates. As discussed in the introduction, this potential hinges on whether errors are correlated or not. Supporting H2a, pairwise correlations between individual humans' signed errors tend to be very low (mean r = 0.06, 95% CI [0.06, 0.07]). By contrast, between-LLM error correlations are typically medium strong (mean r = 0.43, 95% CI [0.35, 0.51]), supporting H2b. These positive correlations indicate that different LLMs consistently overestimate or underestimate the same scenarios. For example, there were 37 scenarios where all six LLMs overestimated (Supplementary Table 2) and 59 scenarios where they all underestimated (Supplementary Table 3). Within a single model, errors are even more consistent: within-model correlations of signed errors range from r(553) = 0.73 to 0.98, all p < 0.001 (Supplementary Table 4). These differences in error correlation have direct consequences for aggregation. Figure 4 shows how MAE varies when we vary the number k of estimates aggregated from k = 1 to k = 15. Supporting H2c, human MAE drops steeply from 1.76 for an average participant down to 0.57 for an average collective of 15 participants. This is the classic wisdom-of-crowds effect: because different individuals make different errors, those errors cancel out through averaging. Supporting H2d, results for LLMs remain flat in Figure 4. Averaging the estimates from prompting the same model k times produces almost no improvement. Moreover, this holds even when aggregating across different LLMs (from this point forward, we report LLM results based on estimates aggregated across k = 15 runs): aggregating estimates from all six models achieves MAE = 0.73, not significantly different from the best individual model (Gemini 3 Flash, MAE = 0.71; Welch t(1108) = 0.43, p = .666), and selective aggregation of only the best-performing models also did not yield significant improvements (Supplementary Table 5). Thus, different LLMs share systematic blind spots that make ensemble approaches ineffective. This finding establishes that the so-called Artificial Hivemind, previously documented for open-ended generation tasks (14) and ethical reasoning (15, 16), extends to estimation of a population's social norms. Fig. 4. Aggregation of norm estimates improves norm estimation accuracy of humans but not of LLMs. The MAE of estimates of 555 social norms at different levels of aggregation k . Each LLM value is the average MAE achieved when estimates are aggregated across k runs. For humans it is the average MAE achieved aggregating estimates of k humans. Supporting H2e, the best accuracy is achieved by human-LLM hybrids. We paired each of the 320 human estimators with each LLM and calculated the accuracy of their average estimates (Figure 5). The results show a clear benefit of hybrids. For example, even though every single participant was outperformed by Claude Sonnet 4.5, as many as 23% (95% CI [18%, 28%]) of participants create hybrids with Claude Sonnet 4.5 that outperform it. The benefits of hybrids can also be illustrated using a collective of humans. Whereas no individual participant is as accurate as the best LLMs, recall that an aggregate of 15 participants is more accurate than the best LLM (Fig. 4). Averaging the estimates of the 15-person aggregate with the estimates of Gemini 3 Flash yielded MAE = 0.48, representing a 32.6% (95% CI [28.8%, 36.5%]) error reduction compared to Gemini 3 Flash alone and a 15.3% (95% CI [8.6%, 21.1%]) improvement over the human aggregate alone. In sum, these results fully support H2a–H2e. Figure 5. The accuracy of Human-LLM Hybrids. Distribution of MAE for hybrids created by pairing each of the 320 individual humans with the aggregated estimates (k = 15) of each of the six LLMs. Each panel shows results for one LLM. The proportion of individuals creating hybrids superior to the LLM alone varies by model capability: 90% for LLaMA 3.1 8B, 63% for LLaMA 3.1 70B, 72% for DeepSeek V3, 39% for GPT-5.2, 23% for Claude Sonnet 4.5, and 8% for Gemini 3 Flash. Dashed vertical lines indicate each LLM's solo performance for reference. Discussion Our investigation into how large language models estimate social norms yields insights at multiple levels. We tested hypotheses about norm estimation accuracy, error patterns, and aggregation benefits. LLMs achieve remarkable accuracy at estimating population-average appropriateness ratings, with the best LLMS outperforming all human subjects and all LLMs outperforming most humans (supporting Hypothesis 1b). Yet LLM errors are correlated across models while human errors are independent, making humans better candidates for aggregation and hybrids the most accurate (supporting Hypothesis 2). Humans provided an interesting baseline. To our knowledge, this is the first systematic study of how well people understand how everyday behavior is judged by other people in the same society. We found that people typically believe that the average judgments are much more extreme than they really are. This categorical bias likely reflects a form of false consensus effect: participants themselves tend to have categorical views of scenario appropriateness and project these categorical judgments onto the population, unaware of how common it is that others hold different views. However, when these categorical estimates are aggregated across individuals, they closely track the measured norms. This too is consistent with false consensus; if people estimate that the average rating of a scenario is the same as the rating they would personally give, the average of these estimates will be close to the average rating in the same population. Interestingly, this result simultaneously shows that everyday norms are not subject to pluralistic ignorance, the phenomenon of collective misperception of norms that have been documented in several other domains (7). Against the inaccuracy of human estimates of population-average norms for everyday behavior, the accuracy of LLM estimates is striking. LLMs trained purely on text achieve estimation errors roughly half the size of the average human, despite having no embodied experience: they have never navigated awkward silences, felt social rejection, or learned through trial and error which behaviors give offense. This finding speaks to a central question in cognitive science: does social understanding require lived experience, as theories of grounded cognition suggest (8-10)? From the perspective of grounded cognition theory, LLM knowledge is amodal (consisting of abstract statistical patterns over symbols) whereas human social knowledge is modal, rooted in perceptual and sensorimotor experience. Our results suggest that for the specific task of estimating population-average norms, knowledge that lacks direct modal grounding is still very effective. The norms governing everyday behavior are apparently encoded, implicitly but recoverable, in the patterns of how people write about social situations. This is arguably a first step toward AI being capable of acting according to social norms, although navigating real social situations will also require understanding of many other things including real-time social dynamics and individual deviations from average norms. The latter is important because, as the categorical bias shows, there is considerable individual variation in ideas about what the norms are. While LLMs estimate population-average norms with high accuracy, they still make some errors. These errors showed remarkable homogeneity. Not only does the same LLM produce nearly identical estimates across repeated queries, but different LLMs make similar errors. Aggregating estimates within or across models therefore provides little improvement in accuracy. This finding extends the Artificial Hivemind effect, recently documented for open-ended generation tasks (14) and ethical reasoning (15, 16), to social norm estimation. The practical implication is important: organizations seeking robust social AI cannot simply ensemble multiple models, because the models share too many blind spots. This brings us to hybrids. As different humans make different estimation errors, their errors also tend not to correlate with LLM errors. This allows human-LLM hybrids to capitalize on LLM accuracy while benefiting from human diversity. The mechanism is the same one that enables wisdom-of-crowds effects within human collectives: independent errors cancel out through averaging. An open question is what characterizes the norms where LLMs systematically err. The shared blind spots we identified do not appear to follow simple patterns. The 37 scenarios that all six LLMs overestimated span diverse behaviors, from physical actions like jumping and running to social behaviors like singing and arguing, as well as diverse contexts from buses to family dinners. Similarly, the 59 scenarios that all LLMs underestimated range from reading in class to laughing in church to spitting in public spaces (Supplementary Tables 2 and 3). The errors encompass both overestimations and underestimations, suggesting that LLMs do not simply have a general tendency toward permissiveness or conservatism. Understanding what features of scenarios predict systematic LLM error—whether they relate to cultural specificity, ambiguity, changing norms, or other factors—is an important direction for future research. Such understanding could inform both the development of more accurate models and the design of hybrid systems that strategically defer to human judgment for specific types of scenarios. Note that this study is limited to norms in the United States. Whether similar patterns hold across cultures remains to be tested. For example, due to the reach of American culture, it could be that LLMs are especially good at estimating American norms. It is also possible that the categorical bias we found among Americans is less pronounced in other cultures. Another limitation is that our findings may not reflect the abilities of future generations of LLMs. Our study evaluated specific LLM versions available in early 2026. Model capabilities continue to advance rapidly, and absolute performance levels will likely improve. However, we expect the complementary nature of statistical and experiential learning to remain robust. Finally, we address concerns about online survey data possibly being generated by LLMs. This concern has grown as researchers have documented that AI agents can successfully pose as human survey respondents (21, 22). However, multiple lines of evidence suggest our human data are genuine. First, the estimation target data were collected in January 2023, just after the first public release of ChatGPT and before LLMs could be used for automatic completion of online surveys. Second, our human participants who provided estimates were recruited via Prolific. A recent comprehensive assessment of seven online platforms using multiple AI detection tests found that Prolific had a very low rate of LLM contamination (23). Third, and most compellingly, if participants had used LLMs to make their estimates, they should not have been systematically worse at estimating norms than LLMs—yet they were; compared to LLMs, participants displayed much worse accuracy and a qualitatively different distribution of estimates. In conclusion, large language models have learned a great deal about population-average social norms from text alone, more than most individual humans can accurately report when asked to estimate these averages. Yet their knowledge is systematically incomplete, with shared blind spots that resist correction through aggregation. Humans, despite their categorical biases, collectively preserve the diversity needed for wisdom-of-crowds effects. Optimal estimate of population-average norms thus emerges neither from statistical learning alone nor from any individual's experience, but from integrating these complementary forms of social knowledge. Materials and Methods Experimental Design This study aimed to assess how accurately large language models estimate social norms, using human norm perception as a reference, and to test whether combining LLM and human estimates yields superior performance. We designed an estimation task in which both LLMs and human participants estimated the mean appropriateness ratings for 555 everyday scenarios (37 behaviors × 15 situations) previously measured in a U.S. sample. This design allowed us to assess estimation accuracy against the measured norms, compare performance across LLMs and humans, and test two prespecified hypotheses: (1) LLMs would be either more or less accurate than humans; (2) LLM errors would be correlated (preventing aggregation benefits), human errors would be independent (enabling wisdom-of-crowds), and hybrids would perform best. Our primary outcome measure was Mean Absolute Error (MAE) between estimates and measured norms. Establishing the Estimation Target To serve as the estimation target, we used a large-scale dataset of appropriateness ratings from Eriksson et al. (6). In that study, 555 unique scenarios were created by systematically pairing 37 common behaviors (e.g., argue, cry, laugh, read, pray audibly) with 15 common situations (e.g., in a bar, in church, at a job interview, in one's own room). In a data collection conducted in January 2023, a total of 555 U.S. participants, recruited via Prolific, rated the appropriateness of each scenario on a 10-point scale from 0 ("The behavior is extremely inappropriate in this situation") to 9 ("The behavior is extremely appropriate in this situation"). To avoid fatigue, each participant rated a random subset of 50 scenarios constructed from a random selection of 10 behaviors and 5 situations. Participants whose ratings would have been more accurate if interpreted as having used the scale in reverse were excluded, leaving 550 participants and yielding approximately 50 ratings per scenario. These ratings provided stable estimates of the collective U.S. norm for each scenario. While the framework of behaviors and contexts was published, the scenario-level appropriateness ratings were not, making it unlikely they were part of LLM training data. Determination of Sample Sizes No formal power analysis was performed. For LLMs, 15 runs per scenario allowed estimation of within-model variability. For the human estimation task, we targeted ~300 participants to ensure at least 15 independent estimates per scenario, allowing comparison between aggregating estimates from k runs of an LLM versus k estimates from different humans. LLM Estimation Task In January 2026, we evaluated six large language models selected to span a range of architectures, parameter counts, and access types, allowing us to test whether patterns such as error homogeneity hold across diverse systems. Proprietary frontier models included GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Flash. Open-weight models included Llama 3.1 8B, Llama 3.1 70B, and DeepSeek V3. All models were given an identical prompt for each of the 555 scenarios. To assess response stability, we queried each model fifteen times per scenario. The full prompt text was: "From various sources in our everyday lives we have all developed a subjective impression for the appropriateness of any given behavior in a particular situation. Imagine a number of people from the United States rated the appropriateness of "[scenario]" on a scale from 0 to 9, where 0 = extremely inappropriate and 9 = extremely appropriate. Interpret the phrase naturally (e.g., "writing on a bus" = writing while riding the bus). Your task is to estimate the average rating these respondents would give. Respond with only a number (0–9) with up to two decimals, no explanation." Human Estimation Task Human participants were given the task of estimating the mean ratings from the normative dataset. Following the original study design, each participant was assigned a random subset of 50 scenarios constructed from a random selection of 10 behaviors and 5 situations. For each scenario they were asked to "estimate the average rating that U.S. respondents would give" on the 0–9 scale. Participants We recruited 337 U.S. participants via Prolific in two waves. We initially recruited 106 participants in October 2025. Preliminary analysis revealed that while this provided sufficient data for individual-level comparisons, we required additional participants to robustly test aggregation benefits up to k=15 estimates per scenario (matching the 15 LLM runs per scenario). We therefore recruited an additional 231 participants in February 2026. Six participants did not complete the survey and were excluded. Of the 331 completers, eleven were excluded because reversing their responses reduced their MAE, suggesting scale misuse, leaving 320 participants for analysis. The analyzed sample comprised 174 women (54.4%), 145 men (45.3%), and 1 participant who selected "other" (0.3%). Mean age was 43.5 years (SD=13.2, median=42, range=19-84, n=319). Education levels were: less than high school (0.9%), high school diploma (14.1%), some college (20.0%), bachelor's degree (44.7%), master's degree (17.2%), and doctoral degree (3.1%). Estimation accuracy was almost identical between the two recruitment waves (October: mean adjusted MAE=1.75, SD=0.40; February: mean adjusted MAE=1.77, SD=0.48). Accuracy also did not vary by demographic characteristics: there was no significant difference between women and men, and no significant correlations with age or education. The insensitivity of estimation accuracy to these demographic variables supports generalizability of human norm estimation performance beyond this particular sample. Statistical Analysis Analysis was performed using R version 4.3.3 (24) with the tidyverse package (25). Performance metrics. Our primary metric for estimation accuracy was Mean Absolute Error (MAE), calculated as the average absolute difference between an estimator's estimate and the actual mean rating for each scenario. Note that human participants only estimated norms for 50 of the 555 scenarios. Because scenarios were randomly assigned, each individual's performance is an unbiased estimate of their accuracy over the full set. However, the random subsets introduce noise due to variation in the difficulty of the assigned subsets. To reduce this noise, we adjusted the MAE score of each human estimator by multiplying their raw MAE by an adjustment factor obtained by dividing the mean MAE across all 555 scenarios (calculated from all available estimates) by the mean MAE for that participant's specific 50 scenarios. This adjustment controls for subset difficulty, allowing fair comparison across all 320 human estimators and between human estimators and LLMs. Any residual noise in individual MAE estimates introduced by between-subset variation in difficulty works against the between-individual comparisons reported here, making our findings conservative. Effect sizes. Our primary outcome measure was Mean Absolute Error (MAE) between estimates and measured norms. MAE values are directly interpretable on the original 0-9 appropriateness scale and serve as our effect size metric. For context, an MAE of 1.0 represents an average estimation error of approximately one scale point. We report unstandardized MAE differences as our primary effect sizes because they retain interpretability in the original measurement units. Categorical bias. To quantify each human estimator's tendency toward extreme moral judgments, we calculated the proportion of their estimates falling at the scale endpoints (ratings of 0–1 or 8–9). We then computed the Pearson correlation between this endpoint proportion and each individual's adjusted MAE across the 320 human estimators. Aggregated estimates. To test our hypotheses about collective intelligence, we constructed several forms of aggregated estimates. For each LLM, we aggregated estimates at the scenario level by randomly sampling k runs ( k = 1–15) from the 15 available model outputs per scenario. For each k , we formed an ensemble estimate by averaging the k sampled runs and computed the MAE of this aggregate. This sampling procedure was repeated 100 times to obtain stable estimates. For humans, we used the same scenario-level procedure: for each scenario, we randomly sampled k human estimates ( k = 1–15) without replacement, averaged them to form an ensemble, and computed the MAE across iterations. To define the hybrids between LLMs and the aggregate of 15 humans, we randomly sampled 15 human estimates per scenario (without replacement) and averaged them, before averaging those aggregate estimates with the LLM estimates for the corresponding scenarios. Error correlation analysis. To understand the basis for aggregation effects, we analyzed signed estimation errors (estimated value minus measured norm). We computed pairwise Pearson correlations of signed errors between runs of the same LLM (within-model correlations), between different LLMs (between-model correlations), and between human estimators (between-human correlations, restricted to scenarios rated by both individuals in each pair). Assumptions and corrections. Pearson correlations were used to quantify error similarity across scenarios. With n=555 scenarios, the sampling distributions of correlation coefficients are well-approximated by their asymptotic distributions regardless of marginal distributions of errors. We report mean correlations across multiple model pairs (15 between-LLM pairs, 105 within-model pairs per LLM) to characterize overall patterns of error dependence rather than to test individual hypotheses. Accordingly, no corrections for multiple comparisons were applied to these descriptive summaries. Declarations Acknowledgements The manuscript has benefitted from constructive comments from Fredrik Jansson, Måns Magnusson, Kristoffer Pettersson, William Hagman, and Sebastian Krakowski. A large language model (Claude Opus 4.5) was used to generate editorial suggestions. This research was supported by a grant from the Knut and Alice Wallenberg foundation [grant number 2022.0191]. Data availability All data and analysis code generated in this study have been deposited at https://osf.io/qd7kj/overview?view_only=7aad5a1949c7416f9c6d0cb6926bd42e. Consent and ethics Participants gave informed consent and data were collected fully anonymously online. No ethics review was required for this anonymous survey according to the regulations in Sweden, the country from which the study was conducted. Competing interests The authors declare no competing interests. Author Contributions K.E. wrote the paper, S.K. performed the statistical analysis and visualizations with support from I.V., and P.S. conceived the study and provided critical revisions. All authors reviewed the manuscript. References C. Bicchieri, The Grammar of Society: The Nature and Dynamics of Social Norms (Cambridge Univ. Press, 2005). D. T. Miller, D. A. Prentice, Changing norms to change behavior. Annu. Rev. Psychol. 67 , 339–361 (2016). M. J. Gelfand, J. L. Raver, L. Nishii, L. M. Leslie, J. Lun, B. C. Lim, L. Duan, A. Almaliach, S. Ang, J. Arnadottir, Z. Aycan, K. Boehnke, P. Boski, R. Cabecinhas, D. Chan, J. Chhokar, A. D'Amato, M. Ferrer, I. C. Fischlmayr, R. Fischer, M. Fülöp, J. Georgas, E. S. Kashima, Y. Kashima, K. Kim, A. Lempereur, P. Marquez, R. Othman, B. Overlaet, P. Panagiotopoulou, K. Peltzer, L. R. Perez-Florizno, L. Ponomarenko, A. Realo, V. Schei, M. Schmitt, P. B. Smith, N. Soomro, E. Szabo, N. Taveesin, M. Toyama, E. Van de Vliert, N. Vohra, C. Ward, S. Yamaguchi, Differences between tight and loose cultures: A 33-nation study. Science 332 , 1100–1104 (2011). K. Eriksson, P. Strimling, M. Gelfand, J. Wu, J. Abernathy, C. S. Akotia, A. Arriaga, F. R. Aquino, A. Kirchner-Häusler, W.-Q. E. Chua, Z. Dorrough, F. G. Efendic, O. Elmas, Z. Findor, G. Gunsoy, R. Haugestad, M. H. Bahair, A. Mbyirukira, M. Kohút, D. Lazarević, A. G. Luukkonen, R. Nayak, T. R. Nielsen, E. Opara, I. K. Penner, A. G. Pietri, A. M. C. Reis, H. Şahin, T. M. W. Sjåstad, P. A. M. Van Lange, Perceptions of the appropriate response to norm violation in 57 societies. Nat. Commun. 12 , 1481 (2021). K. Eriksson, P. Strimling, I. Vartanova, M. Gelfand, G. Wu, Everyday norms have become more permissive over time and vary across cultures. Commun. Psychol. 3 , 145 (2025). K. Eriksson, P. Strimling, I. Vartanova, Appropriateness ratings of everyday behaviors in the United States now and 50 years ago. Front. Psychol. 14 , 1237494 (2023). R. H. Sargent, L. S. Newman, Pluralistic ignorance research in psychology: A scoping review of topic and method variation and directions for future research. Rev. Gen. Psychol. 25 , 163–184 (2021). L. W. Barsalou, Grounded cognition. Annu. Rev. Psychol. 59 , 617–645 (2008). A. M. Glenberg, Embodiment as a unifying perspective for psychology. Wiley Interdiscip. Rev. Cogn. Sci. 1 , 586–596 (2010). T. W. Schubert, G. R. Semin, Embodiment as a unifying perspective for psychology. Eur. J. Soc. Psychol. 39 , 1135–1141 (2009). T. K. Landauer, S. T. Dumais, A solution to Plato's problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge. Psychol. Rev. 104 , 211–240 (1997). L. Ross, D. Greene, P. House, The "false consensus effect": An egocentric bias in social perception and attribution processes. J. Exp. Soc. Psychol. 13 , 279–301 (1977). J. Krueger, R. W. Clement, The truly false consensus effect: An ineradicable and egocentric bias in social perception. J. Pers. Soc. Psychol. 67 , 596–610 (1994). L. Jiang, Y. Chai, M. Li, M. Liu, R. Fok, N. Dziri, Y. Tsvetkov, M. Sap, Y. Choi, Artificial Hivemind: The open-ended homogeneity of language models (and beyond). arXiv:2510.22954 [cs.CL] (2025). W. R. Neuman, C. Coleman, M. Shah, Analyzing the ethical logic of six large language models. arXiv:2501.08951 [cs.CL] (2025). C. Coleman, W. R. Neuman, A. Dasdan, S. Ali, M. Shah, The convergent ethics of AI? Analyzing moral foundation priorities in large language models with a multi-framework approach. arXiv:2504.19255 [cs.CL] (2025). F. Galton, Vox populi. Nature 75 , 450–451 (1907). R. P. Larrick, J. B. Soll, Intuitions about combining opinions: Misappreciation of the averaging principle. Manage. Sci. 52 , 111–127 (2006). J. Surowiecki, The Wisdom of Crowds (Anchor Books, 2005). A. Abels, T. Lenaerts, Wisdom from diversity: Bias mitigation through hybrid human-LLM crowds. arXiv:2505.12349 [cs.CL] (2025). F. Panizza, Y. Kyrychenko, J. Roozenbeek, How to stop the survey-taking AI chatbots that threaten to upend social science. Nature 650 , 293–295 (2026). S. J. Westwood, The potential existential threat of large language models to online survey research. Proc. Natl. Acad. Sci. U.S.A. 122 , e2417835122 (2025). G. Zhang, R. Walatka, S. Y. Chen, O. Urminsky, K. Fernandez, A. Low, J. E. Bogard, C. R. Fox, Estimating the threat of AI-agent responding across online survey platforms. Preprint athttps://osf.io/preprints/psyarxiv/xcg26 (2026). R Core Team. R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria (2023).https://www.R-project.org/ Wickham, H. et al. Welcome to the tidyverse. J. Open Source Softw. 4 , 1686 (2019). Additional Declarations There is NO Competing Interest. Supplementary Files SupplementtoNComms.docx Supplementary Information nrreportingsummaryfilled.pdf Reporting Summary Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9161239","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":612814698,"identity":"750efbfe-2912-4740-b5b8-3c6f1ed585d4","order_by":0,"name":"Kimmo Eriksson","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAsElEQVRIiWNgGAWjYBACCWYwZQPEjI0HSNGSBtLSQKQWCHUYTBKnRbKdO3XDh4rzdmvbDwNtqbGJJqhFmpl3280ZZ24nbzuTCNRyLC23gZAWOaCW27xtt5PNDgC1MDYcJlbLv3PJZucfEqlFGqyl4YCd2Q1ibZFsBvnlWHKC2Q2gLQnE+EXi/NltNz7U2NmbnU9/+OBDjQ1hLTCQCFaZQKxyELAnRfEoGAWjYBSMMAAAXlJHTSGjg58AAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-7164-0924","institution":"Mälardalen University","correspondingAuthor":true,"prefix":"","firstName":"Kimmo","middleName":"","lastName":"Eriksson","suffix":""},{"id":612814699,"identity":"8d3fd501-8ec3-4e18-8eff-6a414b24254b","order_by":1,"name":"Simon Karlsson","email":"","orcid":"","institution":"Institutet för framtidsstudier","correspondingAuthor":false,"prefix":"","firstName":"Simon","middleName":"","lastName":"Karlsson","suffix":""},{"id":612814700,"identity":"6c8a2961-6df8-46ec-aa65-f2a4847c88d9","order_by":2,"name":"Irina Vartanova","email":"","orcid":"","institution":"Institute for Futures Studies","correspondingAuthor":false,"prefix":"","firstName":"Irina","middleName":"","lastName":"Vartanova","suffix":""},{"id":612814701,"identity":"58f4cf7d-2831-4f53-a84e-d8ea698c60a8","order_by":3,"name":"Pontus Strimling","email":"","orcid":"","institution":"Institute for Futures Studies","correspondingAuthor":false,"prefix":"","firstName":"Pontus","middleName":"","lastName":"Strimling","suffix":""}],"badges":[],"createdAt":"2026-03-18 15:46:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9161239/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9161239/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105762679,"identity":"cff8e606-1624-4569-aecc-a6d7b0b3002d","added_by":"auto","created_at":"2026-03-30 19:00:57","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":200568,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe estimation target: everyday norms from a prior U.S. study (6).\u003c/strong\u003e Heatmap of mean appropriateness ratings for 555 scenarios constructed by pairing 37 behaviors (rows) with 15 contexts (columns). Behaviors and contexts are ordered by their overall mean ratings. Brighter colors indicate more appropriate scenarios.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-9161239/v1/e5af362714542add51b33945.png"},{"id":105904145,"identity":"35fcb172-9e59-40cf-b25b-ae140b4bb725","added_by":"auto","created_at":"2026-04-01 10:04:54","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":89394,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLLMs are more accurate than humans at estimating measured norms. \u003c/strong\u003eThe mean absolute error (MAE) of single-run (k = 1) estimates of 555 social norms made by six LLMs and 320 human participants. The dashed line indicates the average among human participants. The standard errors of MAE estimates are very small for the LLMs (mean SE = 0.01). Every human participant rated a random subset of 50 scenarios; their MAE is adjusted for subset difficulty (see Methods), and their standard errors are not as small (mean SE = 0.20).\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-9161239/v1/dad8d20c5d805c31f14e5cc7.png"},{"id":106092986,"identity":"33126ac1-dff6-49ce-8813-1a2e6cfd67c7","added_by":"auto","created_at":"2026-04-03 11:32:00","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":87491,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparing the Distributions of Estimated and Measured Norms. \u003c/strong\u003eThe proportions falling in each unit-width bin (0–1, 1–2, ..., 8–9) for true norms (purple) are almost perfectly tracked by single-run LLM estimates (k = 1; yellow), whereas human participants grossly overestimate the frequency of norms close to the scale endpoints (green).\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-9161239/v1/08fb669286f619816f840d58.png"},{"id":105762683,"identity":"741e8923-dddf-4180-92ac-0d3ee67717ef","added_by":"auto","created_at":"2026-03-30 19:00:58","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":123802,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAggregation of norm estimates improves norm estimation accuracy of humans but not of LLMs. \u003c/strong\u003eThe MAE of estimates of 555 social norms at different levels of aggregation \u003cem\u003ek\u003c/em\u003e. Each LLM value is the average MAE achieved when estimates are aggregated across \u003cem\u003ek\u003c/em\u003e runs. For humans it is the average MAE achieved aggregating estimates of \u003cem\u003ek \u003c/em\u003ehumans.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-9161239/v1/75ddfaf8f4c5e4fd8fc6820d.png"},{"id":107479826,"identity":"5ae7dda1-b458-4e70-b49f-9bf78b9c5f8a","added_by":"auto","created_at":"2026-04-22 01:53:42","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":118243,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe accuracy of Human-LLM Hybrids.\u003c/strong\u003e Distribution of MAE for hybrids created by pairing each of the 320 individual humans with the aggregated estimates (k = 15) of each of the six LLMs. Each panel shows results for one LLM. The proportion of individuals creating hybrids superior to the LLM alone varies by model capability: 90% for LLaMA 3.1 8B, 63% for LLaMA 3.1 70B, 72% for DeepSeek V3, 39% for GPT-5.2, 23% for Claude Sonnet 4.5, and 8% for Gemini 3 Flash. Dashed vertical lines indicate each LLM's solo performance for reference.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-9161239/v1/098238916fb9cef1b6b598c5.png"},{"id":107869526,"identity":"7c3c2d7c-57a8-4d89-9706-d7cfc5cee60a","added_by":"auto","created_at":"2026-04-27 07:37:15","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":729560,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9161239/v1/b17f7bfa-fb54-47eb-a9a7-a8719f21d0bd.pdf"},{"id":106092987,"identity":"82fb083d-d98b-4569-8ed9-cb7192b62519","added_by":"auto","created_at":"2026-04-03 11:32:00","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":41310,"visible":true,"origin":"","legend":"Supplementary Information","description":"","filename":"SupplementtoNComms.docx","url":"https://assets-eu.researchsquare.com/files/rs-9161239/v1/bac862f23312619672ee267f.docx"},{"id":105762682,"identity":"75bfe26d-46a3-47da-9710-70b583f3a9ac","added_by":"auto","created_at":"2026-03-30 19:00:58","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":1665976,"visible":true,"origin":"","legend":"Reporting Summary","description":"","filename":"nrreportingsummaryfilled.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9161239/v1/ce9359aa443ed46d9346c301.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Large language models outperform humans at estimating society's everyday norms—but hybrids are even better","fulltext":[{"header":"Introduction","content":"\u003cp\u003eSocial robots and AI assistants are rapidly entering human environments, providing companionship in care facilities, assisting customers in retail settings, and collaborating with workers in offices. To function effectively in these roles, such systems must navigate the vast landscape of context-dependent social norms: knowing that laughter is welcome at a party but inappropriate at a funeral, that physical proximity acceptable among friends may violate boundaries with strangers, and that behaviors appropriate in one culture may give offense in another. When AI systems misunderstand these unwritten rules, the consequences range from awkward interactions to genuine harm (1, 2). Understanding how well large language models grasp social norms and where they systematically fail has thus become a question of both theoretical and practical importance.\u003c/p\u003e\n\u003cp\u003eFollowing quantitative approaches to norm research (3-5), we operationalize a social norm as the average appropriateness judgment in a population - how appropriate members of a society collectively consider a behavior in a given context. This operationalization captures information that is helpful for an agent to avoid conflict. An AI system that accurately estimates the population average can calibrate its behavior and advice to match typical societal expectations, minimizing the risk of causing offense in situations where the expectations are not otherwise known. By contrast, a system whose estimates deviate systematically from the population average will more often misjudge what others find acceptable. Crucially, this is an estimation task: the question is not whether an LLM has its own normative views, but whether it can accurately infer what a population of people collectively considers appropriate.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFor measures of what a population of people collectively considers appropriate, we draw on a large-scale dataset of appropriateness ratings for 555 everyday scenarios in the United States (6). The scenarios were constructed by systematically pairing 37 common behaviors with 15 common situations, and each scenario was rated by a large sample of U.S. participants on a scale from 0 (extremely inappropriate) to 9 (extremely appropriate). Figure 1 illustrates the combinatorial structure of the 555 scenarios, with behaviors as rows and contexts as columns, and the average appropriateness rating of each scenario. These ratings provide stable estimates of collective U.S. norms that serve as our estimation target: the measured norms that LLMs and human estimators will be tested against. Importantly, the original paper on this data did not present ratings per scenario, so the estimation target is unlikely to have been in the LLM training data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFig. 1. The estimation target: everyday norms from a prior U.S. study (6).\u003c/strong\u003e Heatmap of mean appropriateness ratings for 555 scenarios constructed by pairing 37 behaviors (rows) with 15 contexts (columns). Behaviors and contexts are ordered by their overall mean ratings. Brighter colors indicate more appropriate scenarios.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo our knowledge, no studies address how LLMs estimate social norms, but there is a body of research on how people estimate, and sometimes misestimate, them (7). Our first question is whether LLMs are better or worse at estimating social norms than humans are. There are arguments on both sides. Theories of embodied cognition propose that social understanding fundamentally requires lived experience: having navigated awkward silences, felt the sting of social rejection, or learned through trial and error which behaviours are acceptable in which contexts (8-10). LLMs lack all such experience. However, the vast scale of training data may allow them to extract normative patterns from how behaviors are discussed across countless contexts, and distributional approaches suggest that social knowledge can emerge from co-occurrence patterns in language alone (11). Humans, for their part, encounter only a fraction of possible situations through direct experience. Moreover, their judgments may be colored by the well-documented false consensus effect, that is, the tendency to assume others share one\u0026apos;s own views (12, 13), or they may underestimate such agreement as in pluralistic ignorance (7). In sum, two competing hypotheses can be formulated.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHypothesis 1 (Accuracy):\u003c/strong\u003e At the task of estimating the average appropriateness ratings for everyday scenarios in the U.S., the accuracy of LLMs is (a) worse than humans, or (b) better than humans.\u003c/p\u003e\n\u003cp\u003eA second question is whether ensemble estimation improves accuracy. If different LLMs share similar training distributions and optimization objectives, their errors should be correlated across scenarios, making them poor candidates for ensemble improvement. Recent work suggests this is indeed the case. Jiang et al. (14) documented a pervasive \u0026quot;Artificial Hivemind\u0026quot; effect: across more than 70 language models, outputs converged on strikingly similar responses both within models and across different architectures, even for open-ended creative tasks. This homogeneity extends beyond creative generation to domains requiring social and ethical reasoning (15, 16). It is plausible, but not known, that this homogeneity could extend also to inferences of social norms, in which case the diversity required for wisdom-of-the-crowd effects may be absent in LLM ensembles.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHumans, by contrast, may show a different profile. While individuals in our study share a common cultural framework, they encounter different situations through their personal histories. If these idiosyncratic experiences generate sufficiently independent errors across scenarios, aggregating estimates should substantially outperform even the best individual. The wisdom-of-crowds literature demonstrates that aggregating independent judgments often outperforms even the best individual, because diverse errors cancel out (17-19). If humans and LLMs differ on error independence\u0026mdash;with LLM errors being correlated and human errors being independent\u0026mdash;the best accuracy might be achieved by human-LLM hybrids that capitalize on complementary strengths. Indeed, Abels and Lenaerts (20), in a study of news headline authenticity judgments, found that simply averaging responses from multiple LLMs failed to improve performance due to limited diversity, but that hybrid crowds combining LLMs with humans capitalized on complementary strengths.\u003c/p\u003e\n\u003cp\u003eIn sum, the above considerations yield a set of related predictions:\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHypothesis 2 (Aggregation and Hybrids):\u003c/strong\u003e (a) Human errors are independent across individuals. (b) LLM errors are correlated across models. (c) Aggregating human estimates yields substantial improvement in accuracy. \u0026nbsp;(d) Aggregating LLM estimates yields little improvement in accuracy. (e) Human-LLM hybrids achieve the best accuracy.\u003c/p\u003e\n\u003cp\u003eWe test these hypotheses by giving the task of estimating the measured norm for each of the 555 scenarios to six large language models available in January 2026: three proprietary frontier models (GPT-5.2, Claude Sonnet 4.5, Gemini 3 Flash) and three smaller open-weight models (Llama 3.1 8B, Llama 3.1 70B, DeepSeek V3). We also gave the same task to 320 U.S. participants, except to avoid fatigue each participant only estimated a random subset of 50 scenarios.\u003c/p\u003e\n\u003cp\u003eOur results support Hypothesis 1b and Hypothesis 2. LLMs achieve remarkably high accuracy, substantially outperforming most individual humans. At the same time, LLM errors are correlated across models, preventing improvement through aggregation, while human errors are independent, enabling powerful wisdom-of-crowds effects. Human-LLM hybrids substantially outperform either approach alone. We end the paper by discussing the broad implications this has for designing AI systems that must navigate complex social environments.\u003c/p\u003e"},{"header":"Results","content":"\u003ch2\u003eHypothesis 1: LLM vs Human Accuracy\u003c/h2\u003e\n\u003cp\u003eOur first question concerns whether LLMs are better or worse than humans at estimating population-average appropriateness ratings. To ensure a fair comparison with individual human estimators, we evaluate each LLM based on the mean error it makes in a single run (averaged over 15 runs). We find that LLMs, especially the largest ones, tend to be better (supporting Hypothesis 1b over Hypothesis 1a). Figure 2 shows mean absolute error (MAE) for each individual and each LLM. Two of the proprietary LLMs outperform even the best human participant (MAEs = 0.78 and 0.85 vs. 0.92). The accuracy of the open-weight models is not as good (MAEs = 1.22\u0026ndash;1.62) but still better than the average human participant\u0026apos;s (MAE = 1.77). All six LLMs significantly outperformed the average human (all p \u0026lt; .001, two-sample Welch t-tests; see Supplementary Table 1 for detailed statistics).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFig. 2. LLMs are more accurate than humans at estimating measured norms.\u0026nbsp;\u003c/strong\u003eThe mean absolute error (MAE) of single-run (k = 1) estimates of 555 social norms made by six LLMs and 320 human participants. The dashed line indicates the average among human participants. The standard errors of MAE estimates are very small for the LLMs (mean SE = 0.01). Every human participant rated a random subset of 50 scenarios; their MAE is adjusted for subset difficulty (see Methods), and their standard errors are not as small (mean SE = 0.20).\u003c/p\u003e\n\u003cp\u003eWhy do humans show lower accuracy? To understand this, we examined how distributions of estimated norms compare to the distribution of measured norms (Figure 3). The estimates of LLMs (orange line) have a distribution that broadly tracks the true norms (purple line). By contrast, individual human estimators (green line) strongly overestimate extreme values and rarely make estimates in the middle range. We term this phenomenon \u003cem\u003ecategorical bias\u003c/em\u003e: even when explicitly asked to estimate continuous population averages, individuals default to binary moral judgments classifying behaviors as simply wrong or right. This is likely a false consensus effect, that is, participants tend to have a categorical view of the appropriateness of a scenario and are unaware of how common it is that others have a different view. We argue that humans\u0026apos; categorical bias explains why they are outperformed by LLMs. In support of this interpretation, individuals who made more endpoint estimates (0-1 or 8-9) showed significantly higher error, r(318) = 0.29, 95% CI [0.18, 0.38], p \u0026lt; 0.001. (For example, the outlier respondent at the right end of Fig. 2 made endpoint estimates 88% of the time.) Thus, better norm estimation accuracy is associated with resisting the categorical bias by making more graded, nuanced estimates.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFig. 3. Comparing the Distributions of Estimated and Measured Norms.\u0026nbsp;\u003c/strong\u003eThe proportions falling in each unit-width bin (0\u0026ndash;1, 1\u0026ndash;2, ..., 8\u0026ndash;9) for true norms (purple) are almost perfectly tracked by single-run LLM estimates (k = 1; yellow), whereas human participants grossly overestimate the frequency of norms close to the scale endpoints (green).\u003c/p\u003e\n\u003ch2\u003eHypothesis 2: Aggregation and Hybrids\u003c/h2\u003e\n\u003cp\u003eHypothesis 2 concerns the potential for improving norm estimates by aggregating several estimates. As discussed in the introduction, this potential hinges on whether errors are correlated or not. Supporting H2a, pairwise correlations between individual humans\u0026apos; signed errors tend to be very low (mean r = 0.06, 95% CI [0.06, 0.07]).\u003c/p\u003e\n\u003cp\u003eBy contrast, between-LLM error correlations are typically medium strong (mean r = 0.43, 95% CI [0.35, 0.51]), supporting H2b. These positive correlations indicate that different LLMs consistently overestimate or underestimate the same scenarios. For example, there were 37 scenarios where all six LLMs overestimated (Supplementary Table 2) and 59 scenarios where they all underestimated (Supplementary Table 3). Within a single model, errors are even more consistent: within-model correlations of signed errors range from r(553) = 0.73 to 0.98, all p \u0026lt; 0.001 (Supplementary Table 4).\u003c/p\u003e\n\u003cp\u003eThese differences in error correlation have direct consequences for aggregation. Figure 4 shows how MAE varies when we vary the number k of estimates aggregated from k = 1 to k = 15. Supporting H2c, human MAE drops steeply from 1.76 for an average participant down to 0.57 for an average collective of 15 participants. This is the classic wisdom-of-crowds effect: because different individuals make different errors, those errors cancel out through averaging.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSupporting H2d, results for LLMs remain flat in Figure 4. Averaging the estimates from prompting the same model k times produces almost no improvement. Moreover, this holds even when aggregating across different LLMs (from this point forward, we report LLM results based on estimates aggregated across k = 15 runs): aggregating estimates from all six models achieves MAE = 0.73, not significantly different from the best individual model (Gemini 3 Flash, MAE = 0.71; Welch t(1108) = 0.43, p = .666), and selective aggregation of only the best-performing models also did not yield significant improvements (Supplementary Table 5). Thus, different LLMs share systematic blind spots that make ensemble approaches ineffective. This finding establishes that the so-called Artificial Hivemind, previously documented for open-ended generation tasks (14) and ethical reasoning (15, 16), extends to estimation of a population\u0026apos;s social norms.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFig. 4. Aggregation of norm estimates improves norm estimation accuracy of humans but not of LLMs.\u0026nbsp;\u003c/strong\u003eThe MAE of estimates of 555 social norms at different levels of aggregation \u003cem\u003ek\u003c/em\u003e. Each LLM value is the average MAE achieved when estimates are aggregated across \u003cem\u003ek\u003c/em\u003e runs. For humans it is the average MAE achieved aggregating estimates of \u003cem\u003ek\u0026nbsp;\u003c/em\u003ehumans.\u003c/p\u003e\n\u003cp\u003eSupporting H2e, the best accuracy is achieved by human-LLM hybrids. We paired each of the 320 human estimators with each LLM and calculated the accuracy of their average estimates (Figure 5). The results show a clear benefit of hybrids. For example, even though every single participant was outperformed by Claude Sonnet 4.5, as many as 23% (95% CI [18%, 28%]) of participants create hybrids with Claude Sonnet 4.5 that outperform it. The benefits of hybrids can also be illustrated using a collective of humans. Whereas no individual participant is as accurate as the best LLMs, recall that an aggregate of 15 participants is more accurate than the best LLM (Fig. 4). Averaging the estimates of the 15-person aggregate with the estimates of Gemini 3 Flash yielded MAE = 0.48, representing a 32.6% (95% CI [28.8%, 36.5%]) error reduction compared to Gemini 3 Flash alone and a 15.3% (95% CI [8.6%, 21.1%]) improvement over the human aggregate alone.\u003c/p\u003e\n\u003cp\u003eIn sum, these results fully support H2a\u0026ndash;H2e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFigure 5. The accuracy of Human-LLM Hybrids.\u003c/strong\u003e Distribution of MAE for hybrids created by pairing each of the 320 individual humans with the aggregated estimates (k = 15) of each of the six LLMs. Each panel shows results for one LLM. The proportion of individuals creating hybrids superior to the LLM alone varies by model capability: 90% for LLaMA 3.1 8B, 63% for LLaMA 3.1 70B, 72% for DeepSeek V3, 39% for GPT-5.2, 23% for Claude Sonnet 4.5, and 8% for Gemini 3 Flash. Dashed vertical lines indicate each LLM\u0026apos;s solo performance for reference.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eOur investigation into how large language models estimate social norms yields insights at multiple levels. We tested hypotheses about norm estimation accuracy, error patterns, and aggregation benefits. LLMs achieve remarkable accuracy at estimating population-average appropriateness ratings, with the best LLMS outperforming all human subjects and all LLMs outperforming most humans (supporting Hypothesis 1b). Yet LLM errors are correlated across models while human errors are independent, making humans better candidates for aggregation and hybrids the most accurate (supporting Hypothesis 2).\u003c/p\u003e\n\u003cp\u003eHumans provided an interesting baseline. To our knowledge, this is the first systematic study of how well people understand how everyday behavior is judged by other people in the same society. We found that people typically believe that the average judgments are much more extreme than they really are. This categorical bias likely reflects a form of false consensus effect: participants themselves tend to have categorical views of scenario appropriateness and project these categorical judgments onto the population, unaware of how common it is that others hold different views. However, when these categorical estimates are aggregated across individuals, they closely track the measured norms. This too is consistent with false consensus; if people estimate that the average rating of a scenario is the same as the rating they would personally give, the average of these estimates will be close to the average rating in the same population. Interestingly, this result simultaneously shows that everyday norms are not subject to pluralistic ignorance, the phenomenon of collective misperception of norms that have been documented in several other domains (7).\u003c/p\u003e\n\u003cp\u003eAgainst the inaccuracy of human estimates of population-average norms for everyday behavior, the accuracy of LLM estimates is striking. LLMs trained purely on text achieve estimation errors roughly half the size of the average human, despite having no embodied experience: they have never navigated awkward silences, felt social rejection, or learned through trial and error which behaviors give offense. This finding speaks to a central question in cognitive science: does social understanding require lived experience, as theories of grounded cognition suggest (8-10)? From the perspective of grounded cognition theory, LLM knowledge is amodal (consisting of abstract statistical patterns over symbols) whereas human social knowledge is modal, rooted in perceptual and sensorimotor experience. Our results suggest that for the specific task of estimating population-average norms, knowledge that lacks direct modal grounding is still very effective. The norms governing everyday behavior are apparently encoded, implicitly but recoverable, in the patterns of how people write about social situations. This is arguably a first step toward AI being capable of acting according to social norms, although navigating real social situations will also require understanding of many other things including \u0026nbsp;real-time social dynamics and individual deviations from average norms. The latter is important because, as the categorical bias shows, there is considerable individual variation in ideas about what the norms are.\u003c/p\u003e\n\u003cp\u003eWhile LLMs estimate population-average norms with high accuracy, they still make some errors. These errors showed remarkable homogeneity. Not only does the same LLM produce nearly identical estimates across repeated queries, but different LLMs make similar errors. Aggregating estimates within or across models therefore provides little improvement in accuracy. This finding extends the Artificial Hivemind effect, recently documented for open-ended generation tasks (14) and ethical reasoning (15, 16), to social norm estimation. The practical implication is important: organizations seeking robust social AI cannot simply ensemble multiple models, because the models share too many blind spots. This brings us to hybrids. As different humans make different estimation errors, their errors also tend not to correlate with LLM errors. This allows human-LLM hybrids to capitalize on LLM accuracy while benefiting from human diversity. The mechanism is the same one that enables wisdom-of-crowds effects within human collectives: independent errors cancel out through averaging.\u003c/p\u003e\n\u003cp\u003eAn open question is what characterizes the norms where LLMs systematically err. The shared blind spots we identified do not appear to follow simple patterns. The 37 scenarios that all six LLMs overestimated span diverse behaviors, from physical actions like jumping and running to social behaviors like singing and arguing, as well as diverse contexts from buses to family dinners. Similarly, the 59 scenarios that all LLMs underestimated range from reading in class to laughing in church to spitting in public spaces (Supplementary Tables 2 and 3). The errors encompass both overestimations and underestimations, suggesting that LLMs do not simply have a general tendency toward permissiveness or conservatism. Understanding what features of scenarios predict systematic LLM error—whether they relate to cultural specificity, ambiguity, changing norms, or other factors—is an important direction for future research. Such understanding could inform both the development of more accurate models and the design of hybrid systems that strategically defer to human judgment for specific types of scenarios.\u003c/p\u003e\n\u003cp\u003eNote that this study is limited to norms in the United States. Whether similar patterns hold across cultures remains to be tested. For example, due to the reach of American culture, it could be that LLMs are especially good at estimating American norms. It is also possible that the categorical bias we found among Americans is less pronounced in other cultures. Another limitation is that our findings may not reflect the abilities of future generations of LLMs. Our study evaluated specific LLM versions available in early 2026. Model capabilities continue to advance rapidly, and absolute performance levels will likely improve. However, we expect the complementary nature of statistical and experiential learning to remain robust.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFinally, we address concerns about online survey data possibly being generated by LLMs. This concern has grown as researchers have documented that AI agents can successfully pose as human survey respondents (21, 22). However, multiple lines of evidence suggest our human data are genuine. First, the estimation target data were collected in January 2023, just after the first public release of ChatGPT and before LLMs could be used for automatic completion of online surveys. Second, our human participants who provided estimates were recruited via Prolific. A recent comprehensive assessment of seven online platforms using multiple AI detection tests found that Prolific had a very low rate of LLM contamination (23). Third, and most compellingly, if participants had used LLMs to make their estimates, they should not have been systematically worse at estimating norms than LLMs—yet they were; compared to LLMs, participants displayed much worse accuracy and a qualitatively different distribution of estimates. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn conclusion, large language models have learned a great deal about population-average social norms from text alone, more than most individual humans can accurately report when asked to estimate these averages. Yet their knowledge is systematically incomplete, with shared blind spots that resist correction through aggregation. Humans, despite their categorical biases, collectively preserve the diversity needed for wisdom-of-crowds effects. Optimal estimate of population-average norms thus emerges neither from statistical learning alone nor from any individual's experience, but from integrating these complementary forms of social knowledge.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003ch2\u003eExperimental Design\u003c/h2\u003e\n\u003cp\u003eThis study aimed to assess how accurately large language models estimate social norms, using human norm perception as a reference, and to test whether combining LLM and human estimates yields superior performance. We designed an estimation task in which both LLMs and human participants estimated the mean appropriateness ratings for 555 everyday scenarios (37 behaviors \u0026times; 15 situations) previously measured in a U.S. sample. This design allowed us to assess estimation accuracy against the measured norms, compare performance across LLMs and humans, and test two prespecified hypotheses: (1) LLMs would be either more or less accurate than humans; (2) LLM errors would be correlated (preventing aggregation benefits), human errors would be independent (enabling wisdom-of-crowds), and hybrids would perform best. Our primary outcome measure was Mean Absolute Error (MAE) between estimates and measured norms.\u003c/p\u003e\n\u003ch2\u003eEstablishing the Estimation Target\u003c/h2\u003e\n\u003cp\u003eTo serve as the estimation target, we used a large-scale dataset of appropriateness ratings from Eriksson et al. (6). In that study, 555 unique scenarios were created by systematically pairing 37 common behaviors (e.g., argue, cry, laugh, read, pray audibly) with 15 common situations (e.g., in a bar, in church, at a job interview, in one\u0026apos;s own room). In a data collection conducted in January 2023, a total of 555 U.S. participants, recruited via Prolific, rated the appropriateness of each scenario on a 10-point scale from 0 (\u0026quot;The behavior is extremely inappropriate in this situation\u0026quot;) to 9 (\u0026quot;The behavior is extremely appropriate in this situation\u0026quot;). To avoid fatigue, each participant rated a random subset of 50 scenarios constructed from a random selection of 10 behaviors and 5 situations. Participants whose ratings would have been more accurate if interpreted as having used the scale in reverse were excluded, leaving 550 participants and yielding approximately 50 ratings per scenario. These ratings provided stable estimates of the collective U.S. norm for each scenario. While the framework of behaviors and contexts was published, the scenario-level appropriateness ratings were not, making it unlikely they were part of LLM training data.\u003c/p\u003e\n\u003ch2\u003eDetermination of Sample Sizes\u003c/h2\u003e\n\u003cp\u003eNo formal power analysis was performed. For LLMs, 15 runs per scenario allowed estimation of within-model variability. For the human estimation task, we targeted ~300 participants to ensure at least 15 independent estimates per scenario, allowing comparison between aggregating estimates from \u003cem\u003ek\u003c/em\u003e runs of an LLM versus \u003cem\u003ek\u0026nbsp;\u003c/em\u003eestimates from different humans.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eLLM Estimation Task\u003c/h2\u003e\n\u003cp\u003eIn January 2026, we evaluated six large language models selected to span a range of architectures, parameter counts, and access types, allowing us to test whether patterns such as error homogeneity hold across diverse systems. Proprietary frontier models included GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Flash. Open-weight models included Llama 3.1 8B, Llama 3.1 70B, and DeepSeek V3. All models were given an identical prompt for each of the 555 scenarios. To assess response stability, we queried each model fifteen times per scenario. The full prompt text was:\u003c/p\u003e\n\u003cp\u003e\u0026quot;From various sources in our everyday lives we have all developed a subjective impression for the appropriateness of any given behavior in a particular situation. Imagine a number of people from the United States rated the appropriateness of \u0026quot;[scenario]\u0026quot; on a scale from 0 to 9, where 0 = extremely inappropriate and 9 = extremely appropriate. Interpret the phrase naturally (e.g., \u0026quot;writing on a bus\u0026quot; = writing while riding the bus). Your task is to estimate the average rating these respondents would give. Respond with only a number (0\u0026ndash;9) with up to two decimals, no explanation.\u0026quot;\u003c/p\u003e\n\u003ch2\u003eHuman Estimation Task\u003c/h2\u003e\n\u003cp\u003eHuman participants were given the task of estimating the mean ratings from the normative dataset. Following the original study design, each participant was assigned a random subset of 50 scenarios constructed from a random selection of 10 behaviors and 5 situations. For each scenario they were asked to \u0026quot;estimate the average rating that U.S. respondents would give\u0026quot; on the 0\u0026ndash;9 scale.\u003c/p\u003e\n\u003ch2\u003eParticipants\u003c/h2\u003e\n\u003cp\u003eWe recruited 337 U.S. participants via Prolific in two waves. We initially recruited 106 participants in October 2025. Preliminary analysis revealed that while this provided sufficient data for individual-level comparisons, we required additional participants to robustly test aggregation benefits up to k=15 estimates per scenario (matching the 15 LLM runs per scenario). We therefore recruited an additional 231 participants in February 2026. Six participants did not complete the survey and were excluded. Of the 331 completers, eleven were excluded because reversing their responses reduced their MAE, suggesting scale misuse, leaving 320 participants for analysis. The analyzed sample comprised 174 women (54.4%), 145 men (45.3%), and 1 participant who selected \u0026quot;other\u0026quot; (0.3%). Mean age was 43.5 years (SD=13.2, median=42, range=19-84, n=319). Education levels were: less than high school (0.9%), high school diploma (14.1%), some college (20.0%), bachelor\u0026apos;s degree (44.7%), master\u0026apos;s degree (17.2%), and doctoral degree (3.1%).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eEstimation accuracy was almost identical between the two recruitment waves (October: mean adjusted MAE=1.75, SD=0.40; February: mean adjusted MAE=1.77, SD=0.48). Accuracy also did not vary by demographic characteristics: there was no significant difference between women and men, and no significant correlations with age or education. The insensitivity of estimation accuracy to these demographic variables supports generalizability of human norm estimation performance beyond this particular sample.\u003c/p\u003e\n\u003ch2\u003eStatistical Analysis\u003c/h2\u003e\n\u003cp\u003eAnalysis was performed using R version 4.3.3 (24) with the tidyverse package (25). \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePerformance metrics.\u003c/strong\u003e Our primary metric for estimation accuracy was Mean Absolute Error (MAE), calculated as the average absolute difference between an estimator\u0026apos;s estimate and the actual mean rating for each scenario. Note that human participants only estimated norms for 50 of the 555 scenarios. Because scenarios were randomly assigned, each individual\u0026apos;s performance is an unbiased estimate of their accuracy over the full set. However, the random subsets introduce noise due to variation in the difficulty of the assigned subsets. To reduce this noise, we adjusted the MAE score of each human estimator by multiplying their raw MAE by an adjustment factor obtained by dividing the mean MAE across all 555 scenarios (calculated from all available estimates) by the mean MAE for that participant\u0026apos;s specific 50 scenarios. This adjustment controls for subset difficulty, allowing fair comparison across all 320 human estimators and between human estimators and LLMs. Any residual noise in individual MAE estimates introduced by between-subset variation in difficulty works against the between-individual comparisons reported here, making our findings conservative.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEffect sizes.\u003c/strong\u003e Our primary outcome measure was Mean Absolute Error (MAE) between estimates and measured norms. MAE values are directly interpretable on the original 0-9 appropriateness scale and serve as our effect size metric. For context, an MAE of 1.0 represents an average estimation error of approximately one scale point. We report unstandardized MAE differences as our primary effect sizes because they retain interpretability in the original measurement units.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCategorical bias.\u003c/strong\u003e To quantify each human estimator\u0026apos;s tendency toward extreme moral judgments, we calculated the proportion of their estimates falling at the scale endpoints (ratings of 0\u0026ndash;1 or 8\u0026ndash;9). We then computed the Pearson correlation between this endpoint proportion and each individual\u0026apos;s adjusted MAE across the 320 human estimators.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAggregated estimates.\u003c/strong\u003e To test our hypotheses about collective intelligence, we constructed several forms of aggregated estimates. For each LLM, we aggregated estimates at the scenario level by randomly sampling \u003cem\u003ek\u003c/em\u003e runs (\u003cem\u003ek\u003c/em\u003e = 1\u0026ndash;15) from the 15 available model outputs per scenario. For each \u003cem\u003ek\u003c/em\u003e, we formed an ensemble estimate by averaging the \u003cem\u003ek\u003c/em\u003e sampled runs and computed the MAE of this aggregate. This sampling procedure was repeated 100 times to obtain stable estimates. For humans, we used the same scenario-level procedure: for each scenario, we randomly sampled \u003cem\u003ek\u0026nbsp;\u003c/em\u003ehuman estimates (\u003cem\u003ek\u003c/em\u003e = 1\u0026ndash;15) without replacement, averaged them to form an ensemble, and computed the MAE across iterations.\u003c/p\u003e\n\u003cp\u003eTo define the hybrids between LLMs and the aggregate of 15 humans, we randomly sampled 15 human estimates per scenario (without replacement) and averaged them, before averaging those aggregate estimates with the LLM estimates for the corresponding scenarios.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eError correlation analysis.\u003c/strong\u003e To understand the basis for aggregation effects, we analyzed signed estimation errors (estimated value minus measured norm). We computed pairwise Pearson correlations of signed errors between runs of the same LLM (within-model correlations), between different LLMs (between-model correlations), and between human estimators (between-human correlations, restricted to scenarios rated by both individuals in each pair).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAssumptions and corrections.\u003c/strong\u003e Pearson correlations were used to quantify error similarity across scenarios. With n=555 scenarios, the sampling distributions of correlation coefficients are well-approximated by their asymptotic distributions regardless of marginal distributions of errors. We report mean correlations across multiple model pairs (15 between-LLM pairs, 105 within-model pairs per LLM) to characterize overall patterns of error dependence rather than to test individual hypotheses. Accordingly, no corrections for multiple comparisons were applied to these descriptive summaries.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe manuscript has benefitted from constructive comments from Fredrik Jansson, M\u0026aring;ns Magnusson, Kristoffer Pettersson, William Hagman, and Sebastian Krakowski. A large language model (Claude Opus 4.5) was used to generate editorial suggestions. This research was supported by a grant from the Knut and Alice Wallenberg foundation [grant number 2022.0191].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll data and analysis code generated in this study have been deposited at https://osf.io/qd7kj/overview?view_only=7aad5a1949c7416f9c6d0cb6926bd42e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent and ethics\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eParticipants gave informed consent and data were collected fully anonymously online. No ethics review was required for this anonymous survey according to the regulations in Sweden, the country from which the study was conducted.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eK.E. wrote the paper, S.K. performed the statistical analysis and visualizations with support from I.V., and P.S. conceived the study and provided critical revisions. All authors reviewed the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col start=\"1\" type=\"1\"\u003e\n \u003cli\u003eC. Bicchieri, \u003cem\u003eThe Grammar of Society: The Nature and Dynamics of Social Norms\u003c/em\u003e (Cambridge Univ. Press, 2005).\u003c/li\u003e\n \u003cli\u003eD. T. Miller, D. A. Prentice, Changing norms to change behavior.\u0026nbsp;\u003cem\u003eAnnu. Rev. Psychol.\u003c/em\u003e \u003cstrong\u003e67\u003c/strong\u003e, 339–361 (2016).\u003c/li\u003e\n \u003cli\u003eM. J. Gelfand, J. L. Raver, L. Nishii, L. M. Leslie, J. Lun, B. C. Lim, L. Duan, A. Almaliach, S. Ang, J. Arnadottir, Z. Aycan, K. Boehnke, P. Boski, R. Cabecinhas, D. Chan, J. Chhokar, A. D'Amato, M. Ferrer, I. C. Fischlmayr, R. Fischer, M. Fülöp, J. Georgas, E. S. Kashima, Y. Kashima, K. Kim, A. Lempereur, P. Marquez, R. Othman, B. Overlaet, P. Panagiotopoulou, K. Peltzer, L. R. Perez-Florizno, L. Ponomarenko, A. Realo, V. Schei, M. Schmitt, P. B. Smith, N. Soomro, E. Szabo, N. Taveesin, M. Toyama, E. Van de Vliert, N. Vohra, C. Ward, S. Yamaguchi, Differences between tight and loose cultures: A 33-nation study. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e332\u003c/strong\u003e, 1100–1104 (2011).\u003c/li\u003e\n \u003cli\u003eK. Eriksson, P. Strimling, M. Gelfand, J. Wu, J. Abernathy, C. S. Akotia, A. Arriaga, F. R. Aquino, A. Kirchner-Häusler, W.-Q. E. Chua, Z. Dorrough, F. G. Efendic, O. Elmas, Z. Findor, G. Gunsoy, R. Haugestad, M. H. Bahair, A. Mbyirukira, M. Kohút, D. Lazarević, A. G. Luukkonen, R. Nayak, T. R. Nielsen, E. Opara, I. K. Penner, A. G. Pietri, A. M. C. Reis, H. Şahin, T. M. W. Sjåstad, P. A. M. Van Lange, Perceptions of the appropriate response to norm violation in 57 societies. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 1481 (2021).\u003c/li\u003e\n \u003cli\u003eK. Eriksson, P. Strimling, I. Vartanova, M. Gelfand, G. Wu, Everyday norms have become more permissive over time and vary across cultures.\u0026nbsp;\u003cem\u003eCommun. Psychol.\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 145 (2025).\u003c/li\u003e\n \u003cli\u003eK. Eriksson, P. Strimling, I. Vartanova, Appropriateness ratings of everyday behaviors in the United States now and 50 years ago.\u0026nbsp;\u003cem\u003eFront. Psychol.\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 1237494 (2023).\u003c/li\u003e\n \u003cli\u003eR. H. Sargent, L. S. Newman, Pluralistic ignorance research in psychology: A scoping review of topic and method variation and directions for future research.\u0026nbsp;\u003cem\u003eRev. Gen. Psychol.\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 163–184 (2021).\u003c/li\u003e\n \u003cli\u003eL. W. Barsalou, Grounded cognition.\u0026nbsp;\u003cem\u003eAnnu. Rev. Psychol.\u003c/em\u003e \u003cstrong\u003e59\u003c/strong\u003e, 617–645 (2008).\u003c/li\u003e\n \u003cli\u003eA. M. Glenberg, Embodiment as a unifying perspective for psychology.\u0026nbsp;\u003cem\u003eWiley Interdiscip. Rev. Cogn. Sci.\u003c/em\u003e \u003cstrong\u003e1\u003c/strong\u003e, 586–596 (2010).\u003c/li\u003e\n \u003cli\u003eT. W. Schubert, G. R. Semin, Embodiment as a unifying perspective for psychology.\u0026nbsp;\u003cem\u003eEur. J. Soc. Psychol.\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, 1135–1141 (2009).\u003c/li\u003e\n \u003cli\u003eT. K. Landauer, S. T. Dumais, A solution to Plato's problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge.\u0026nbsp;\u003cem\u003ePsychol. Rev.\u003c/em\u003e \u003cstrong\u003e104\u003c/strong\u003e, 211–240 (1997).\u003c/li\u003e\n \u003cli\u003eL. Ross, D. Greene, P. House, The \"false consensus effect\": An egocentric bias in social perception and attribution processes.\u0026nbsp;\u003cem\u003eJ. Exp. Soc. Psychol.\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, 279–301 (1977).\u003c/li\u003e\n \u003cli\u003eJ. Krueger, R. W. Clement, The truly false consensus effect: An ineradicable and egocentric bias in social perception.\u0026nbsp;\u003cem\u003eJ. Pers. Soc. Psychol.\u003c/em\u003e \u003cstrong\u003e67\u003c/strong\u003e, 596–610 (1994).\u003c/li\u003e\n \u003cli\u003eL. Jiang, Y. Chai, M. Li, M. Liu, R. Fok, N. Dziri, Y. Tsvetkov, M. Sap, Y. Choi, Artificial Hivemind: The open-ended homogeneity of language models (and beyond). arXiv:2510.22954 [cs.CL] (2025).\u003c/li\u003e\n \u003cli\u003eW. R. Neuman, C. Coleman, M. Shah, Analyzing the ethical logic of six large language models. arXiv:2501.08951 [cs.CL] (2025).\u003c/li\u003e\n \u003cli\u003eC. Coleman, W. R. Neuman, A. Dasdan, S. Ali, M. Shah, The convergent ethics of AI? Analyzing moral foundation priorities in large language models with a multi-framework approach. arXiv:2504.19255 [cs.CL] (2025).\u003c/li\u003e\n \u003cli\u003eF. Galton, Vox populi. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e75\u003c/strong\u003e, 450–451 (1907).\u003c/li\u003e\n \u003cli\u003eR. P. Larrick, J. B. Soll, Intuitions about combining opinions: Misappreciation of the averaging principle.\u0026nbsp;\u003cem\u003eManage. Sci.\u003c/em\u003e \u003cstrong\u003e52\u003c/strong\u003e, 111–127 (2006).\u003c/li\u003e\n \u003cli\u003eJ. Surowiecki, \u003cem\u003eThe Wisdom of Crowds\u003c/em\u003e (Anchor Books, 2005).\u003c/li\u003e\n \u003cli\u003eA. Abels, T. Lenaerts, Wisdom from diversity: Bias mitigation through hybrid human-LLM crowds. arXiv:2505.12349 [cs.CL] (2025).\u003c/li\u003e\n \u003cli\u003eF. Panizza, Y. Kyrychenko, J. Roozenbeek, How to stop the survey-taking AI chatbots that threaten to upend social science.\u0026nbsp;\u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e650\u003c/strong\u003e, 293–295 (2026).\u003c/li\u003e\n \u003cli\u003eS. J. Westwood, The potential existential threat of large language models to online survey research.\u0026nbsp;\u003cem\u003eProc. Natl. Acad. Sci. U.S.A.\u003c/em\u003e \u003cstrong\u003e122\u003c/strong\u003e, e2417835122 (2025).\u003c/li\u003e\n \u003cli\u003eG. Zhang, R. Walatka, S. Y. Chen, O. Urminsky, K. Fernandez, A. Low, J. E. Bogard, C. R. Fox, Estimating the threat of AI-agent responding across online survey platforms. Preprint athttps://osf.io/preprints/psyarxiv/xcg26 (2026).\u003c/li\u003e\n \u003cli\u003eR Core Team. R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria (2023).https://www.R-project.org/\u003c/li\u003e\n \u003cli\u003eWickham, H. et al. Welcome to the tidyverse.\u0026nbsp;\u003cem\u003eJ. Open Source Softw.\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 1686 (2019).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-9161239/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9161239/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"As AI assistants and social robots enter human environments, their ability to navigate context-dependent social norms is essential to avoid harm and ensure successful collaboration. We evaluate six large language models (LLMs) on their ability to estimate American social norms across 555 everyday scenarios (measured in prior work) and compare these to estimates from 320 humans. LLMs achieve remarkably high accuracy, clearly outperforming the average human. However, the errors LLMs make are systematic; they are similar across runs of the same LLM and even across different LLMs. As a consequence of this homogeneity, aggregating estimates of LLMs produces little improvement. Individual humans make much worse estimates, often defaulting to extreme right-or-wrong judgments even when asked to estimate population averages, but their errors are idiosyncratic and, consequently, aggregating their estimates yields dramatic improvement through wisdom-of-crowds effects. As humans make different errors than LLMs, hybrid ensembles combining both substantially outperform either alone.","manuscriptTitle":"Large language models outperform humans at estimating society's everyday norms—but hybrids are even better","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-30 19:00:53","doi":"10.21203/rs.3.rs-9161239/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"877f197e-7278-4073-84d4-bfbb52fc00ce","owner":[],"postedDate":"March 30th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":65209312,"name":"Scientific community and society/Social sciences/Psychology/Human behaviour"},{"id":65209313,"name":"Physical sciences/Mathematics and computing/Information technology"}],"tags":[],"updatedAt":"2026-04-24T15:26:05+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-30 19:00:53","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9161239","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9161239","identity":"rs-9161239","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-23T02:00:01.238055+00:00
License: CC-BY-4.0