Socio-cognitive benchmarks for large language models (LLMs) often fail to predict real-world behavior, even when models achieve high benchmark scores. Prior work has attributed this evaluation-deployment gap to problems of measurement and validity. While these critiques are insightful, we argue that they overlook a more fundamental issue: many socio-cognitive evaluations proceed without an explicit theoretical specification of the target capability, leaving the assumptions linking task performance to competence implicit. Without this theoretical grounding, benchmarks that exercise only narrow subsets of a capability are routinely misinterpreted as evidence of broad competence: a gap that creates a systemic validity illusion by masking the failure to evaluate the capability's other essential dimensions. To address this gap, we make two contributions. First, we diagnose and formalize this theory gap as a foundational failure that undermines measurement and enables systematic overgeneralization of benchmark results. Second, we introduce the Theory Trace Card (TTC), a lightweight documentation artifact designed to accompany socio-cognitive evaluations, which explicitly outlines the theoretical basis of an evaluation, the components of the target capability it exercises, its operationalization, and its limitations. We argue that TTCs enhance the interpretability and reuse of socio-cognitive evaluations by making explicit the full validity chain, which links theory, task operationalization, scoring, and limitations, without modifying benchmarks or requiring agreement on a single theory.
Language is far more than a communication tool; it encodes a wealth of information about a person's identity, psychological state and social context, providing valuable insights for diverse fields including psychology, marketing and healthcare. Across three studies spanning seven datasets in different domains and over 880,000 texts, we show that the widespread adoption of large language models (LLMs) as writing assistants is linked to declines in linguistic diversity, interfering with the societal and psychological insights language provides. While core content is retained when LLMs polish and rewrite texts, LLMs also homogenize writing styles, reducing writing-complexity variance by a statistically significant 21-50% across datasets and models (P ≤ 0.05), and amplify patterns associated with dominant characteristics while suppressing others, emphasizing conformity over individuality. These trends hold across different LLMs, prompts and contexts, with potential implications for diagnostic processes, personalization efforts, hiring assessments and cultural preservation.
Social media bios are ubiquitous yet understudied identity signals, persistently visible to diverse audiences. Despite often non-political intent, such cues may be politicized in perception, with consequences for intergroup bias. Across four studies (N = 2,084), we test how commonplace bio content—occupations, hobbies, family roles, religious affiliations, pronouns—can unintentionally signal political ideology and shape prejudice. In Study 1, partisan-leaning bios were perceived as politically motivated, especially by outgroup members, who attributed greater political intent than signalers intended. Social media users whose bios contained outgroup signals were seen as less warm, competent, and trustworthy, and more threatening, toxic, and misinformed. Studies 2–4 extended the influence of bio content on behavioral contexts: influencing the perceived toxicity and valence of user comments (Study 2), discrimination in online marketplace interactions (Study 3), and simulated hiring scenarios (Study 4), each revealing discrimination toward users with outgroup-congruent bios. Findings show how subtle online self-presentations can be politicized, fueling prejudice and discrimination, with implications for signaling theory, meta-perception, and interventions to reduce polarization.
Large language models (LLMs) often produce human-like moral judgments, but it is unclear whether this reflects an internal conceptual structure or superficial “moral mimicry.” Using Moral Foundations Theory (MFT) as an analytic framework, we study how moral foundations are encoded, organized, and expressed within two instruction-tuned LLMs: Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct. We employ a multi-level approach combining (i) layer-wise analysis of MFT concept representations and their alignment with human moral perceptions, (ii) pretrained sparse autoencoders (SAEs) over the residual stream to identify sparse features that support moral concepts, and (iii) causal steering interventions using dense MFT vectors and sparse SAE features. We find that both models represent and distinguish moral foundations in a structured, layer-dependent way that aligns with human judgments. At a finer scale, SAE features show clear semantic links to specific foundations, suggesting partially disentangled mechanisms within shared representations. Finally, steering along either dense vectors or sparse features produces predictable shifts in foundation-relevant behavior, demonstrating a causal connection between internal representations and moral outputs. Together, our results provide mechanistic evidence that moral concepts in LLMs are distributed, layered, and partly disentangled, suggesting that pluralistic moral structure can emerge as a latent pattern from the statistical regularities of language alone.
While research has documented clear regional differences in environmental attitudes and behaviors, less is understood about the role of shared moral values in shaping these variations. This gap poses a critical challenge to designing effective climate action strategies. Many environmental initiatives rely on "moral framing" to promote proenvironmental behavior, often targeting specific geographical areas like cities or counties. However, these strategies may falter if they fail to account for the unique moral landscapes that shape climate beliefs and actions in different regions. To maximize the success of these interventions, it is crucial to understand how collective moral values influence environmental engagement across diverse communities. Across two studies, we offer insights at the collective level into the moral psychology of climate change by investigating how county-level moral values can predict (i) green attitudes and (ii) household carbon emissions within those counties after accounting for political behavior and region-specific factors. Using Bayesian geospatial modeling, we find that counties that endorse purity and fairness show higher environmental concerns and lower emissions across 3,102 US counties in 48 states. While political orientation strongly predicts environmental attitudes, moral values appear to be a more important factor in predicting carbon footprints. We discuss how county-level dynamics deviate from individual-level dynamics. Our community-level evidence can be leveraged to enhance green interventions on a regional scale by aligning them with the local populace's prevailing values and lived experiences, thus bolstering public support and increasing the likelihood of successful climate action initiatives.
Human conflict is often attributed to threats against material conditions and symbolic values, yet it remains unclear how they interact and which dominates. Progress is limited by weak causal control, ethical constraints, and scarce temporal data. We address these barriers using simulations of large language model (LLM)-driven agents in virtual societies, independently varying realistic and symbolic threat while tracking actions, language, and attitudes. Representational analyses show that the underlying LLM encodes realistic threat, symbolic threat, and hostility as distinct internal states, that our manipulations map onto them, and that steering these states causally shifts behavior. Our simulations provide a causal account of threat-driven conflict over time: realistic threat directly increases hostility, whereas symbolic threat effects are weaker, fully mediated by ingroup bias, and increase hostility only when realistic threat is absent. Non-hostile intergroup contact buffers escalation, and structural asymmetries concentrate hostility among majority groups.
Honor is universally comprehensible, varies regionally in frequency, chronicity, and intensity, and looks different at each time and place. We use culture-as-situated-cognition theory (CSCT), an integrating situated social cognition account of culture, to understand why. Human culture addresses recurrent problems; how frequently, chronically, and intensely each comes to mind depends on their ecological niche; the practices addressing them vary in time and place. We articulate the costly morality theory of honor (CMTH) within CSCT to distinguish honor from related constructs by theorizing two axes (morality-immorality and costly-cost-free) at each of CSCT's three levels. In our formulation, honor is costly morality, resolving the recurrent problem of regulating relationships through costly signals of trustworthiness (human-universal). Societies embedded in harsher ecological niches require more cost to find a signal to be honest and focus on particular relational aspects of morality (niche-linked). Honor specifies how to be a person in the world (time-and-place-specific).Public AbstractPeople have an everyday understanding of honor, what it is, and who has it, but what they mean can be hard to put into words, and what actions in service of honor look like vary across times and societies. We build on culture-as-situated-cognition theory, which accounts for honor's importance in human culture, its variable centrality across societies, and differing specific norms and practices connected to it within societies, to posit that honor entails moral action, a duty of care, that is costly to the actor. We apply our honor-as-costly-morality theory to distinguish honor from related ideas in the hope that our framework helps people better understand and communicate across time-and-place divides, even while disagreeing on to whom and in what way the duty of care extends.
Sometimes people miss the “good old days” of their country. This national nostalgia canfuel populist movements worldwide, rallying supporters with promises of restored glory.Socio-psychological work has characterized national nostalgia as a coping mechanism forcollective self-discontinuity, but it remains theoretically unclear which aspect of collectiveidentity (e.g., national, political, or moral) is most central in shaping national nostalgia.We propose the Moral Nostalgia Hypothesis (MNH), which posits that moral identity iscentral. MNH predicts trait sensitivity, meaning that people who prioritize moral valuesare dispositionally prone to national nostalgia, and threat sensitivity, meaning thatperceived moral decline evokes nostalgia more strongly than economic decline as asituational response. We test MNH across four studies in 20 societies (N = 4,924) usingsurveys and two centuries of English and Chinese texts. Values tied to hierarchy andtradition emerged as the strongest dispositional predictors of national nostalgia, even aftercontrolling for multiple individual- and societal-level variables, while perceived fairnessdecline was the most powerful situational trigger. These findings identify moral identitycontent, rather than material welfare or generic group cues, as the key mechanism linkingcollective self-discontinuity to national nostalgia, illuminating why nostalgic politicalrhetoric framed as moral restoration resonates so powerfully.
Ensuring the moral reasoning capabilities of Large Language Models (LLMs) is a growing concern as these systems are used in socially sensitive tasks. Nevertheless, current evaluation benchmarks present two major shortcomings: a lack of annotations that justify moral classifications, which limits transparency and interpretability; and a predominant focus on English, which constrains the assessment of moral reasoning across diverse cultural settings. In this paper, we introduce MFTCXplain, a multilingual benchmark dataset for evaluating the moral reasoning of LLMs via multi-hop hate speech explanations using the Moral Foundations Theory. MFTCXplain comprises 3,000 tweets across Portuguese, Italian, Persian, and English, annotated with binary hate speech labels, moral categories, and text span-level rationales. Our results show a misalignment between LLM outputs and human annotations in moral reasoning tasks. While LLMs perform well in hate speech detection (F1 up to 0.836), their ability to predict moral sentiments is notably weak (F1 < 0.35). Furthermore, rationale alignment remains limited mainly in underrepresented languages. Our findings show the limited capacity of current LLMs to internalize and reflect human moral reasoning.
Large Language Models (LLMs) exhibit impressive reasoning abilities, yet their reliance on structured step-by-step processing reveals a critical limitation. In contrast, human cognition fluidly adapts between intuitive, heuristic (System 1) and analytical, deliberative (System 2) reasoning depending on the context. This difference between human cognitive flexibility and LLMs' reliance on a single reasoning style raises a critical question: while human fast heuristic reasoning evolved for its efficiency and adaptability, is a uniform reasoning approach truly optimal for LLMs, or does its inflexibility make them brittle and unreliable when faced with tasks demanding more agile, intuitive responses? To answer these questions, we explicitly align LLMs to these reasoning styles by curating a dataset with valid System 1 and System 2 answers, and evaluate their performance across reasoning benchmarks. Our results reveal an accuracy-efficiency trade-off: System 2-aligned models excel in arithmetic and symbolic reasoning, while System 1-aligned models perform better in commonsense reasoning tasks. To analyze the reasoning spectrum, we interpolated between the two extremes by varying the proportion of alignment data, which resulted in a monotonic change in accuracy. A mechanistic analysis of model responses shows that System 1 models employ more definitive outputs, whereas System 2 models demonstrate greater uncertainty. Building on these findings, we further combine System 1- and System 2-aligned models based on the entropy of their generations, without additional training, and obtain a dynamic model that outperforms across nearly all benchmarks. This work challenges the assumption that step-by-step reasoning is always optimal and highlights the need for adapting reasoning strategies based on task demands.
How do ideological threats influence people from different political ideologies? Prior work primarily focuses on the conservative-threat dynamic, but less is known about the progressive-threat dynamic and how Progressives’ group attitudes are influenced by threats against their values. We investigate this gap in three experimental studies (N1 = 400, N2 = 600, N3 = 600) and a time series analysis of political polarization polls (N4 = 34,000 respondents) that illustrate the impact of ideological threats, both on social media and in real-life contexts, on political prejudice in the US. We find experimental and correlational evidence that exposure to threats against ideology influences the endorsement of outgroup prejudice, with Progressives being more reactive to relevant threats. Our results highlight the psychological underpinnings of polarization and dangers associated with threatening content on social media, providing novel evidence showing that political ideology influences how people perceive threats, which in turn, influences prejudice.
The emergence of large language models (LLMs) has sparked considerable interest in their potential application in psychological research, either as a human-like entity used as a model for the human psyche or as a general text-analysis tool. However, carelessly using LLMs in psychological studies, a trend we rhetorically refer to as ``GPTology,'' can have negative consequences, especially given the convenient access to models such as ChatGPT. We elucidate the promises, limitations, and ethical considerations of using LLMs in psychological research. First, LLM-based research should pay attention to the substantial psychological diversity around the globe, as well as demographic diversity within populations. Second, while LLMs are convenient tools, we caution against treating them as a one-size-fits-all method for psychological text analysis. Third, LLM-based psychological research needs to develop methods and standards to compensate for LLMs' opaque black-box nature to facilitate reproducibility, transparency, and robust inference from AI-generated data.While acknowledging the prospects offered by LLMs for easy task automation (e.g., text annotation) and to expand our understanding of human psychology (e.g., by contrasting human and machine psychology), we make a case for diversifying human samples and expanding psychology's methodological toolbox to achieve a truly inclusive and generalizable science, rather than homogenizing samples and methods through over-reliance on LLMs.
Sexual assaults are a social problem in Iran; however, psychological factors that predict perceptions of sexual assault remain largely unexamined. Here, we examine the relationship between moral concerns, culture-specific gender roles, and victim blaming in sexual assault scenarios in Iranian culture. Relying on Moral Foundations Theory and recent theoretical developments in moral psychology in the Iranian context, we examined the correlations between five moral foundations (Care, Fairness, Loyalty, Authority, and Purity), a culture-specific set of values called Qeirat (which includes guarding and [over]protectiveness of female kin, romantic partners, broader family, and country), and victim blaming. In a community sample of Iranians ( N = 411), we found Qeirat values to be highly correlated with victim blaming, and that this link was mediated by a number of culture-specific proscriptions about women’s roles and dress code (i.e., Haya). In a regression analysis with all moral foundations, Qeirat values, Haya, and religiosity as predictors of victim blaming, only Haya, religiosity, high Authority values, and low Care values were found to predict how strongly Iranian participants blamed victims of sexual assault scenarios.
Infectious diseases have been an impending threat to the survival of individuals and groups throughout our evolutionary history. As a result, humans have developed psychological pathogen-avoidance mechanisms and groups have developed societal norms that respond to the presence of disease-causing microorganisms in the environment. In this work, we demonstrate that morality plays a central role in the cultural and psychological architectures that help humans avoid pathogens. We present a collection of studies which together provide an integrated understanding of the socio-ecological and psychological impacts of pathogens on human morality. Specifically, in Studies 1 (2,834 U.S. counties) and 2 (67 nations), we show that regional variation in pathogen prevalence is consistently related to aggregate moral Purity. In Study 3, we use computational linguistic methods to show that pathogen-related words co-occur with Purity words across multiple languages. In Studies 4 (n = 513) and 5 (n = 334), we used surveys and social psychological experimentation to show that pathogen-avoidance attitudes are correlated with Purity. Finally, in Study 6, we found that historical prevalence of pathogens is linked to Care, Loyalty, and Purity. We argue that particular adaptive moral systems are developed and maintained in response to the threat of pathogen occurrence in the environment. We draw on multiple methods to establish connections between pathogens and moral codes in multiple languages, experimentally induced situations, individual differences, U.S. counties, 67 countries, and historical periods over the last century.
Despite the widespread availability of COVID-19 vaccines, the United States has a depressed rate of vaccination relative to similar countries. Understanding the psychology of vaccine refusal, particularly the possible sources of variation in vaccine resistance across U.S. subpopulations, can aid in designing effective intervention strategies to increase vaccination across different regions. Here, we demonstrate that county-level moral values (i.e., Care, Fairness, Loyalty, Authority, and Purity) are associated with COVID-19 vaccination rates across 3,106 counties in the contiguous United States. Specifically, in line with our hypothesis, we find that fewer people are vaccinated in counties whose residents prioritize moral concerns about bodily and spiritual purity. Further, we find that stronger endorsements of concerns about Fairness and Loyalty to the group predict higher vaccination rates. These associations are robust after adjusting for structural barriers to vaccination, the demographic makeup of the counties, and their residents' political voting behavior. Our findings have implications for health communication, intervention strategies based on targeted messaging, and our fundamental understanding of the moral psychology of vaccination hesitancy and behavior. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
The study of moral judgements often centres on moral dilemmas in which options consistent with deontological perspectives (that is, emphasizing rules, individual rights and duties) are in conflict with options consistent with utilitarian judgements (that is, following the greater good based on consequences). Greene et al. (2009) showed that psychological and situational factors (for example, the intent of the agent or the presence of physical contact between the agent and the victim) can play an important role in moral dilemma judgements (for example, the trolley problem). Our knowledge is limited concerning both the universality of these effects outside the United States and the impact of culture on the situational and psychological factors affecting moral judgements. Thus, we empirically tested the universality of the effects of intent and personal force on moral dilemma judgements by replicating the experiments of Greene et al. in 45 countries from all inhabited continents. We found that personal force and its interaction with intention exert influence on moral judgements in the US and Western cultural clusters, replicating and expanding the original findings. Moreover, the personal force effect was present in all cultural clusters, suggesting it is culturally universal. The evidence for the cultural universality of the interaction effect was inconclusive in the Eastern and Southern cultural clusters (depending on exclusion criteria). We found no strong association between collectivism/individualism and moral dilemma judgements. Including participants from 45 countries, Bago et al. find that the situational factors that affect moral reasoning are shared across countries, with diminished observed cultural variation.
Assessment of intrasexual competition has largely relied on Intrasexual Competition Scale (ICS; Buunk & Fisher Journal of Evolutionary Psychology, 7:37–48, 2009). Based on recent developments in mating psychology and the notion that humans use multiple tactics to compete with same-sex individuals, we propose a new theory-driven assessment strategy for intrasexual rivalry in men and women. Here, we develop and initially validate the 16-item Intrasexual Rivalry Scale (IRS). Eight items represented self-promoting tactics in four mating areas and eight items represented rival-derogatory tactics in the same mating areas. We pre-registered our study design and statistical strategy and recruited a community sample in a non-Western culture, Iran. Consistent with our theoretical expectation, exploratory factor analysis (N = 211) clearly suggested extraction of two distinct factors (self-promotion and rival-derogation). Results suggested that scores on the ICS are strongly correlated with rival-derogation, but only weakly associated with self-promotion. Findings are explained in the light of evolutionary psychological perspective and future directions with the newly developed scale are outlined.
Much research on moral judgment is centered on moral dilemmas in which deontological perspectives (i.e., emphasizing rules, individual rights and duties) are in conflict with utilitarian judgements (i.e., following the greater good defined through consequences). A central finding of this field Greene et al. showed that psychological and situational factors (e.g., the intent of the agent, or physical contact between the agent and the victim) play an important role in people’s use of deontological versus utilitarian considerations when making moral decisions. As their study was conducted with US samples, our knowledge is limited concerning the universality of this effect, in general, and the impact of culture on the situational and psychological factors of moral judgments, in particular. Here, we empirically test the universality of deontological and utilitarian judgments by replicating Greene et al.’s experiments on a large (N = X,XXX) and diverse (WEIRD and non-WEIRD) sample across the world to explore the influence of culture on moral judgment. The relevance of this exploration to a broad range of policy-making problems is discussed.
Religiosity as a significant cultural aspect can impact an array of reproductive behaviors. In particular, religiosity can influence intrasexual rivalry as a competitive strategy and content of mate retention behaviors among men and women. However, a few studies have examined the relationship between religiosity, intrasexual rivalry, and mate retention behaviors in non-Western cultures. In Study 1, we examine the province-level relationship between religiosity and reproductive outcomes (i.e., fertility, divorce, family values, and sex ratio) in Iran, a non-Western understudied culture, In Study 2, we use a multi-item measure of religiosity, a new multi-dimensional measure to assess intrasexual rivalry (Intrasexual Rivalry Scale; two components of rival-derogation and self-promotion), and Mate Retention Inventory-Short Form (MRI-SF) in a community sample (N = 211). Results suggested that province-level religiosity in Iran is associated with male-biased sex ratio, lower degrees of divorce, and higher levels of fertility. Study 2's findings showed that religiosity is inversely associated with self-promoting intrasexual traits. We demonstrated that self-promotion is related to benefit-provisioning and rival-derogation attitudes in a same-sex individual or is associated with cost-inflicting mate retention behaviors. We demonstrated that religiosity can predict important mating outcomes in both province- and individual-levels in Iran.