Can Large Language Models (LLMs) accurately estimate various societies' moral values? Here, we query the perceptions of LLMs regarding the moral norms of the "average" person from 48 nations and compare them to a large-scale (n = 90,802) survey of six moral values (Care, Equality, Proportionality, Loyalty, Authority, and Purity) from those populations. Our findings indicate that LLMs poorly capture the moral diversity around the globe, systematically overestimating some moral values (particularly Care) and underestimating others (especially Purity). Notably, examining various versions of Generative Pre-trained Transformer (GPT) shows that these LLMs may overestimate the overall moral concerns of some Western countries (e.g., the United States, Canada, and Australia) while underestimating those of non-Western countries (e.g., Nigeria, Morocco, and Indonesia). Our work demonstrates that LLMs are inaccurate generators of cross-cultural estimations in the moral domain; in other words, they stereotype the moral values of non-Western populations in predictable ways. Our results highlight the ethical and epistemic risks of relying on LLMs to estimate the endorsement of moral values around the globe.
ABSTRACT People often endorse multiple moral values. In many real‐world contexts, these values come into conflict, with each pushing people to act in a different way. Moral conflicts can occur not only within individuals but also between individuals and groups. Hence, we need metamoral principles—that is, principles governing how to adjudicate between competing moral values—to guide moral decision‐making at intrapersonal, interpersonal, and intergroup levels. We argue that metamorality (second‐order moral principles) is not monolithic; various metamoral principles can culturally evolve. There is sufficient evidence for at least four metamoral systems: utilitarianism, honor, divinity, and self‐interest. We explain how this descriptive pluralistic approach to metamorality can explain real‐world phenomena. Implications of this model and future directions in moral psychology are discussed.
Moral language often travels widely online, but does more moral content always correspond to higher engagement? We analysed 1,621,147 observations across 13 socio-political topics on Twitter (n = 530,104), Reddit (n = 1,048,653) and 8chan (n = 42,390). Using Distributed Dictionary Representations-word embeddings scored against an expert-validated moral dictionary-we measured moral loading (that is, a post's overall moral relevance) and moral density (that is, concentration of moral content across words). Negative-binomial models showed that moral loading was positively associated with engagement (range 1.12 [0.96, 1.28] to 9.07 [8.21, 9.93], all P < 0.001). Conditional on moral loading, however, moral density was 'negatively' associated with engagement (range -4.71 [-5.41, -4.02] to -0.40 [-0.52, -0.27], all P < 0.001). Engagement peaked at density 0.30 [0.30048, 0.30070], P < 0.001, with lower engagement below (2.28-fold) and above (2.78-fold), consistent with an engagement advantage for moral language that is bounded by an overmoralization penalty pattern.
Large language models offer new opportunities for behavioural science, but their rapid evolution poses challenges for research rigour. We introduce a consensus-based reporting checklist to improve transparency, reproducibility and ethical accountability of large-language-model-based research in the behavioural sciences.
Historical psychology is an emerging area of research aimed at understanding how historical processes influence the mind, including people’s beliefs, values, and attitudes. As an introduction to this special issue on historical psychology, we review the origins, current pressing questions, and the future of this new field. Our review of the field’s development underscores how historical psychology has been a longstanding topic in psychological science, even though it has only recently emerged as a formal area of inquiry. Our review of pressing questions situates each of the papers in the present special issue within broader lines of research concerning how economic development shapes values, how and why intergroup attitudes change over time, and how contact and conflict shape behavioral outcomes. In our section on the future, we call for more theory-driven research in historical psychology, more recognition of path-dependent cultural change, and historical analysis of a broader range of psychological variables beyond attitudes and beliefs. As historical psychology continues to mature, we look ahead to a richer and more rigorous understanding how the mind has changed over time.
In some societies, people find excessive wealth immoral, while others are structured so that having too much money is morally neutral or even praised. Here, we show that moral judgments of excessive wealth are distinguishable from moral judgments of economic inequality and examine how people's moral concerns and national inequality predict the immorality of excessive wealth around the globe. Using demographically stratified samples from 20 nations ( N = 4,351 ), we find that across all countries, people do not find excessive wealth very immoral, with notable variability such that more equal and wealthy societies (e.g. Belgium, Switzerland) consider having too much money more wrong. People's equality and purity concerns reliably predicted their condemnation of excessive wealth, whereas loyalty, authority, and proportionality concerns were negatively associated with condemnation of excessive wealth across societies after controlling for the moralization of inequality, religiosity, political ideology, and demographic variables. We conducted a follow-up study in the United States ( N = 315 ), showing that moral purity is more broadly linked to the moralization of excess beyond wealth, even after controlling for different ways of wealth acquisition and spending. Collectively, these cross-cultural results demonstrate that some moral intuitions shape our moral judgment of excessive wealth above and beyond economic inequality.
Groups can be diverse along many dimensions like gender, race, and national background. These forms of (demographic) diversity are celebrated by many and well-studied in social sciences. Much less is known, however, about moral diversity — the presence of people with different moral priorities in a group. Here, we argue that increased moral diversity may lead to greater tolerance for deviant behaviors and a lower propensity to punish norm violators within a group. In Studies 1-2, we examine two large-scale data sets (total N = 234,164), and in Studies 3-7, we conduct seven controlled laboratory studies (N = 1,384), showing that higher moral diversity results in: greater perceptions that norm violations are prevalent and tolerated, greater acceptance of such violations, and reduced willingness to punish rule breakers. Our findings complement prior work on moral echo chambers and nuance extant recommendations for enhancing moral diversity
In this work, we develop a pipeline for historical-psychological text analysis in classical Chinese. Humans have produced texts in various languages for thousands of years; however, most of the computational literature is focused on contemporary languages and corpora. The emerging field of historical psychology relies on computational techniques to extract aspects of psychology from historical corpora using new methods developed in natural language processing (NLP). The present pipeline, called Contextualized Construct Representations (CCR), combines expert knowledge in psychometrics (i.e., psychological surveys) with text representations generated via transformer-based language models to measure psychological constructs such as traditionalism, norm strength, and collectivism in classical Chinese corpora. Considering the scarcity of available data, we propose an indirect supervised contrastive learning approach and build the first Chinese historical psychology corpus (C-HI-PSY) to fine-tune pre-trained models. We evaluate the pipeline to demonstrate its superior performance compared with other approaches. The CCR method outperforms word-embedding-based approaches across all of our tasks and exceeds prompting with GPT-4 in most tasks. Finally, we benchmark the pipeline against objective, external data to further verify its validity.
Psychology is a young science. Its definition and scope have shifted over the discipline’s short history. Here, we call for psychology to become a historical and geographical science. We list four underlying reasons as to why such a transformation, a chronospatial revolution, has yet to take off: problems in data, scope, synergy, and theory. We discuss the need for psychology to adopt a more holistic lens—one that incorporates the rich mosaic of shared history, the dynamism of cultural shifts, and the variations ingrained by ecology and cross-regional differences. Such an integrated approach not only enriches our microscopic understanding of Homo sapiens but also draws a more telescopic map of human psychology that encapsulates the human journey.
The emergence of large language models (LLMs) has sparked considerable interest in their potential application in psychological research, either as a human-like entity used as a model for the human psyche or as a general text-analysis tool. However, carelessly using LLMs in psychological studies, a trend we rhetorically refer to as ``GPTology,'' can have negative consequences, especially given the convenient access to models such as ChatGPT. We elucidate the promises, limitations, and ethical considerations of using LLMs in psychological research. First, LLM-based research should pay attention to the substantial psychological diversity around the globe, as well as demographic diversity within populations. Second, while LLMs are convenient tools, we caution against treating them as a one-size-fits-all method for psychological text analysis. Third, LLM-based psychological research needs to develop methods and standards to compensate for LLMs' opaque black-box nature to facilitate reproducibility, transparency, and robust inference from AI-generated data.While acknowledging the prospects offered by LLMs for easy task automation (e.g., text annotation) and to expand our understanding of human psychology (e.g., by contrasting human and machine psychology), we make a case for diversifying human samples and expanding psychology's methodological toolbox to achieve a truly inclusive and generalizable science, rather than homogenizing samples and methods through over-reliance on LLMs.
Why do some people morally justify excessive wealth in a world where so many struggle? In some cultures, people find excessive wealth immoral, while others are structured so that having too much money is morally neutral or even praised. Here, we examine how people’s moral values and national inequality predict the moralization of excessive wealth around the globe. Using demographically stratified samples from 20 nations (N = 4,351), we find notable variability in the moralization of excessive wealth such that more equal societies (e.g., Belgium, Switzerland) consider having too much money more wrong.
Honor requires that individuals demonstrate their worth in the eyes of others. However, it is unclear how honor and its implications for behavior vary between societies. Here, we explore the tension between competing views about how to make sense of honor – as narrowly defined through self-reliance and self-defense or as broadly defined through strength of character. The former suggests that demonstrating the ability to defend one’s self, is a crucial component of honor, while the latter allows the centrality of self-reliance to vary depending on circumstances. To examine these implications, we conducted studies in the U.S., where self-reliance is central to honor , and in Iran, where individual agency must be balanced against the interests of kin. Americans (Studies 1, 2a; n = 978) who endorsed honor values tended to ignore governmental COVID-19 measures because they preferred relying on themselves. In contrast, honor-minded Iranians (Study 2b; n = 201) adhered to public-health guidelines and did not prefer self-reliance. Moreover, honor-minded Iranians endorsed family-reliance, but did not moralize self-reliance (Study 3; n = 107), while honor-minded Americans endorsed family-reliance and moralized self-reliance (Study 3; n = 120). Results suggest that local norms may shape how honor is expressed across cultures.
Social stereotypes negatively impact individuals' judgments about different groups and may have a critical role in understanding language directed toward marginalized groups. Here, we assess the role of social stereotypes in the automated detection of hate speech in the English language by examining the impact of social stereotypes on annotation behaviors, annotated datasets, and hate speech classifiers. Specifically, we first investigate the impact of novice annotators' stereotypes on their hate-speech-annotation behavior. Then, we examine the effect of normative stereotypes in language on the aggregated annotators' judgments in a large annotated corpus. Finally, we demonstrate how normative stereotypes embedded in language resources are associated with systematic prediction errors in a hate-speech classifier. The results demonstrate that hate-speech classifiers reflect social stereotypes against marginalized groups, which can perpetuate social inequalities when propagated at scale. This framework, combining social-psychological and computational-linguistic methods, provides insights into sources of bias in hate-speech moderation, informing ongoing debates regarding machine learning fairness.
Boyer presents a compelling account of ownership as the outcome of interaction between two evolved cognitive systems. We integrate this model into current discussions of moral pluralism, suggesting that ownership meets the criteria to be a moral foundation. We caution against ignoring cultural variation in ownership norms and against explaining complex, contested moral phenomena using a monist approach.
The emergence of large language models (LLMs) has sparked considerable interest in their potential application in psychological research, either as a human-like entity used as a model for the human psyche or as a general text-analysis tool. However, carelessly using LLMs in psychological studies, a trend we rhetorically refer to as “GPTology,” can have negative consequences, especially given the convenient access to models such as ChatGPT. We elucidate the promises, limitations, and ethical considerations of using LLMs in psychological research. First, LLM-based research should pay attention to the substantial psychological diversity around the globe, as well as demographic diversity within populations. Second, while LLMs are convenient tools, we caution against treating them as a one-size-fits-all method for psychological text analysis. Third, LLM-based psychological research needs to develop methods and standards to compensate for LLMs’ opaque black-box nature to facilitate reproducibility, transparency, and robust inference from AI-generated data. While acknowledging the prospects offered by LLMs for easy task automation (e.g., text annotation) and to expand our understanding of human psychology (e.g., by contrasting human and machine psychology), we make a case for diversifying human samples and expanding psychology’s methodological toolbox to achieve a truly inclusive and generalizable science, rather than homogenizing samples and methods through over-reliance on LLMs.
Humans use language toward hateful ends, inciting violence and genocide, intimidating and denigrating others based on their identity. Despite efforts to better address the language of hate in the public sphere, the psychological processes involved in hateful language remain unclear. In this work, we hypothesize that morality and hate are concomitant in language. In a series of studies, we find evidence in support of this hypothesis using language from a diverse array of contexts, including the use of hateful language in propaganda to inspire genocide (Study 1), hateful slurs as they occur in large text corpora across a multitude of languages (Study 2), and hate speech on social-media platforms (Study 3). In post hoc analyses focusing on particular moral concerns, we found that the type of moral content invoked through hate speech varied by context, with Purity language prominent in hateful propaganda and online hate speech and Loyalty language invoked in hateful slurs across languages. Our findings provide a new psychological lens for understanding hateful language and points to further research into the intersection of morality and hate, with practical implications for mitigating hateful rhetoric online.
A growing body of evidence suggests that many aspects of psychology have evolved culturally over historical time. A combination of approaches, including experimental data collected over the past 75 years, cross-cultural comparisons, and studies of immigrants, points to systematic changes in psychological domains as diverse as conformity, attention, emotion, morality, and olfaction. However, these approaches can go back in time only for a few decades and typically fail to provide continuous measures of cultural change, posing a challenge for testing deeper historical psychological processes. To tackle this challenge most directly, computational methods emerging from natural language processing can be adapted to extract psychological information from large-scale historical corpora. Here, we first review the benefits of psychology as a historical science and then present three useful classes of text-analytic techniques for historical psychological inquiry: dictionary-based methods, distributed-representational methods, and human-annotation-based methods. These represent an excellent suite of methodologies that can be used to examine the record of “dead minds.” Finally, we discuss the importance of going beyond English-centric text analysis in historical psychology to foster a more generalizable and inclusive science of human behavior. We propose that historical psychology should incorporate and further develop a variety of text-analytic approaches to reliably quantify the historical processes that gave rise to contemporary social, political, and psychological phenomena.
This account of puritanical morality is useful and innovative, but makes two errors. First, it mischaracterizes the purity foundation as being unrelated to cooperation. Second, it makes the leap from cooperation (broadly construed) to a monist account of moral cognition (as harm or fairness). We show how this leap is both conceptually incoherent and inconsistent with empirical evidence about self-control moralization.