Women remain underrepresented in the workplace, partly due to stereotypes associating competence traits with men rather than women. Efforts to change such stereotypes often yield mixed results. As language models become integrated into daily life, AI writing assistants offer an opportunity to shift gender images. In a preregistered experiment (N = 672), participants evaluated résumés for a female (“Jennifer”) and a male (“John”) candidate applying to a financial analyst role. They wrote evaluations using AI-generated suggestions in one of three conditions: suggestions for Jennifer integrated stereotypically male, female, or neutral traits. Suggestions for John remained neutral. Participants exposed to male-trait suggestions evaluated Jennifer as more competent, selected her as the leader, and offered higher salaries. However, we also observed signs of backlash: participants were less willing to work with competent Jennifer. We discuss implications for designing AI writing assistants to mitigate gender bias in hiring contexts.
Humans learn not only from personal experience but also by observing others. Integrating social information often allows groups to make decisions more efficiently and accurately than individuals, yet it can also generate persistent bias. These opposing outcomes are typically attributed to different mechanisms, with optimal outcomes linked to how people rationally update their beliefs, share information, and make decisions, whereas biased outcomes are attributed to frictions in these processes. Here, we show that the same underlying process -- rational social learning -- can produce both effects. Whether social learning improves or degrades group performance depends on the structure of the environment, particularly how clearly options differ. Using hiring decisions as a relevant context, we study networks of Bayesian-rational learners, generative AI agents, human online participants, and hiring professionals making decisions independently or collectively. Our design strips away common biases in social learning by embedding agents in fully connected networks and sharing information in their original form, while causally varying the decision environment and network structure. Across all four cases, integrating social information improves efficiency when one option is objectively optimal. However, when multiple options are equally optimal, the same learning process amplifies early random signals and prematurely reduces exploration, leading to bias. In these cases, biased outcomes do not reflect a failure of rationality, but rather a predictable consequence of rational inference. These results identify a unifying psychological mechanism underlying both collective intelligence and collective bias, with implications for designing fair decision-making systems in human and AI collectives.
Social stereotypes are traditionally understood as direct associations between groups and traits. Here, we show people also systematically link social groups to concepts across nonsocial domains, such as women to wine and harps, men to beer and drums. These seemingly nonsocial objects form a hidden architecture of stereotypes, which we bring into view. By asking participants, ``If X were a Y, what Y would it be?", we mapped 100 social groups across eight semantic domains such as beverages, musical instruments, and color, and found that these cross-domain mappings closely mirrored the perceived warmth and competence of social groups. By further randomly assigning participants to reason through counter-stereotypical cross-domain mappings using the prompt “If X were a Y, it would be y,” we shifted their subsequent justifications toward portraying the target groups as more competent and warmer than they otherwise would have. Analogical reasoning that emphasizes structural similarity was more effective at changing evaluations than associative reasoning based on contextual or surface-level similarity. Our findings suggest that stereotypical traits can be emergent properties of rich associative networks, and that manipulating which everyday objects people link to social groups-even seemingly irrelevant ones-can change how people make sense of social groups.
As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In this paper, we argue that the predominant approach of simply removing existing biases from models is not enough. Using a paradigm from the psychology literature, we demonstrate that LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist. These biases result in highly stratified task allocations, which are less fair than assignments by human participants and are exacerbated by newer and larger models. In social science, emergent biases like these have been shown to result from exploration-exploitation trade-offs, where the decision-maker explores too little, allowing early observations to strongly influence impressions about entire demographic groups. To alleviate this effect, we examine a series of interventions targeting model inputs, problem structure, and explicit steering. We find that explicitly incentivizing exploration most robustly reduces stratification, highlighting the need for better multifaceted objectives to mitigate bias. These results reveal that LLMs are not merely passive mirrors of human social biases, but can actively create new ones from experience, raising urgent questions about how these systems will shape societies over time.
Modern artificial intelligence systems, such as large language models, are increasingly powerful but also increasingly hard to understand. Recognizing this problem as analogous to the historical difficulties in understanding the human mind, we argue that methods developed in cognitive science can be useful for understanding large language models. We propose a framework for applying these methods based on the levels of analysis that David Marr proposed for studying information processing systems. By revisiting established cognitive science techniques relevant to each level and illustrating their potential to yield insights into the behaviour and internal organization of large language models, we aim to provide a toolkit for making sense of these new kinds of minds. This article is part of the theme issue 'World models in natural and artificial intelligence'.
Stereotype change is usually characterized using broad dimensions such as warmth and competence. We argue that these evaluative stereotypes rest on a richer web of group-related associations that we call distributional stereotypes. Using decade-specific word embeddings trained on 385 million words of historical American English, we trace the distributional stereotypes of 12 national groups across a century. Three findings emerge. First, distributional and evaluative stereotypes are correlated, but distributional stereotypes preserve historically specific content that evaluative stereotypes do not capture. Second, distributional stereotypes change primarily as group-related concepts enter and exit the cultural record, rather than because persistent concepts become more positive or negative. Third, distributional stereotypes are more strongly associated with groups’ occupational outcomes, whereas evaluative stereotypes are more strongly associated with their treatment in political discourse. Stereotypes are evolving networks of culturally available ideas that ebb and flow with historical events.
Although value-aligned language models (LMs) appear unbiased in explicit bias evaluations, they often exhibit stereotypes in implicit word association tasks, raising concerns about their fair usage. We investigate the mechanisms behind this discrepancy and find that alignment surprisingly amplifies implicit bias in model outputs. Specifically, we show that aligned LMs, unlike their unaligned counterparts, overlook racial concepts in early internal representations when the context is ambiguous. Not representing race likely fails to activate safety guardrails, leading to unintended biases. Inspired by this insight, we propose a new bias mitigation strategy that works by incentivizing the representation of racial concepts in the early model layers. In contrast to conventional mitigation methods of machine unlearning, our interventions find that steering the model to be more aware of racial concepts effectively mitigates implicit bias. Similar to race blindness in humans, ignoring racial nuances can inadvertently perpetuate subtle biases in LMs.
Large language models (LLMs) can pass explicit social bias tests but still harbor implicit biases, similar to humans who endorse egalitarian beliefs yet exhibit subtle biases. Measuring such implicit biases can be a challenge: As LLMs become increasingly proprietary, it may not be possible to access their embeddings and apply existing bias measures; furthermore, implicit biases are primarily a concern if they affect the actual decisions that these systems make. We address both challenges by introducing two measures: LLM Word Association Test, a prompt-based method for revealing implicit bias; and LLM Relative Decision Test, a strategy to detect subtle discrimination in contextual decisions. Both measures are based on psychological research: LLM Word Association Test adapts the Implicit Association Test, widely used to study the automatic associations between concepts held in human minds; and LLM Relative Decision Test operationalizes psychological results indicating that relative evaluations between two candidates, not absolute evaluations assessing each independently, are more diagnostic of implicit biases. Using these measures, we found pervasive stereotype biases mirroring those in society in 8 value-aligned models across 4 social categories (race, gender, religion, health) in 21 stereotypes (such as race and criminality, race and weapons, gender and science, age and negativity). These prompt-based measures draw from psychology's long history of research into measuring stereotypes based on purely observable behavior; they expose nuanced biases in proprietary value-aligned LLMs that appear unbiased according to standard benchmarks.
Social stereotypes are prevalent and consequential, yet sometimes inaccurate. How do people learn these inaccurate beliefs in the first place and why do these beliefs persist in the face of counter evidence? Building on past research on cognitive limitations and environmental sample biases, we propose an integrative perspective: Insufficient statistical learning (Insta-learn). Instalearn posits that humans are active learners of the environment. Starting from a small sample, people are able to extract statistical patterns within the sample accurately and quickly. However, people do not continue sampling sufficiently. If they decide not to collect more samples once they are (prematurely) satisfied, inaccurate stereotypes can emerge even when more data would show otherwise. We investigated this hypothesis across six online experiments (N = 1565), using novel pairs of computer-generated faces and social behaviors. Fixing the population level statistics of face-behavior associations to zero and varying the initial sample statistics, we found that participants quickly learned the initial sample statistics (from as few as three examples) and persisted in using such spurious associations in their final decisions. Granting the sampling power to participants — samples were endogenously generated by participants and not defined by the experimenters — we found insufficient sampling caused spurious associations to persist. Insta-learn provides a domain-general framework for a mechanistic explanation of the emergence and persistence of social stereotypes.
As machine learning applications proliferate, we need an understanding of their potential for harm. However, current fairness metrics are rarely grounded in human psychological experiences of harm. Drawing on the social psychology of stereotypes, we use a case study of gender stereotypes in image search to examine how people react to machine learning errors. First, we use survey studies to show that not all machine learning errors reflect stereotypes nor are equally harmful. Then, in experimental studies we randomly expose participants to stereotype-reinforcing, -violating, and -neutral machine learning errors. We find stereotype-reinforcing errors induce more experientially (i.e., subjectively) harmful experiences, while having minimal changes to cognitive beliefs, attitudes, or behaviors. This experiential harm impacts women more than men. However, certain stereotype-violating errors are more experientially harmful for men, potentially due to perceived threats to masculinity. We conclude that harm cannot be the sole guide in fairness mitigation, and propose a nuanced perspective depending on who is experiencing what harm and why.
Traditional explanations for stereotypes assume that they result from deficits in humans (ingroup-favoring motives, cognitive biases) or their environments (majority advantages, real group differences). An alternative explanation recently proposed that stereotypes can emerge when exploration is costly. Even optimal decision makers in an ideal environment can inadvertently form incorrect impressions from arbitrary encounters. However, all these existing theories essentially describe shortcuts that fail to explain the multidimensionality of stereotypes. Stereotypes of social groups have a canonical multidimensional structure, organized along dimensions of warmth and competence. We show that these dimensions and the associated stereotypes can result from feature-based exploration: When individuals make self-interested decisions based on past experiences in an environment where exploring new options carries an implicit cost and when these options share similar attributes, they are more likely to separate groups along multiple dimensions. We formalize this theory via the contextual multiarmed bandit problem, use the resulting model to generate testable predictions, and evaluate those predictions against human behavior. We evaluate this process in incentivized decisions involving as many as 20 real jobs and successfully recover the classic dimensions of warmth and competence. Further experiments show that intervening on the cost of exploration effectively mitigates bias, further demonstrating that exploration cost per se is the operating variable. Future diversity interventions may consider how to reduce exploration cost, in ways that parallel our manipulations.
Abstract Psychological theories continue to expand our understanding of stereotype content and processes. Stereotype content refers to what people think about social groups’ characteristics. Stereotype processes reflect how people integrate information to navigate social interactions. Advances in artificial intelligence introduce innovative analytical tools to revolutionize how psychologists understand stereotypes. Content research builds on theory-driven surveys (e.g. warmth and competence), to data-driven multidimensional scaling (e.g. belief), to large-scale linguistic analysis (e.g. emotion, appearance) to describe a myriad of dimensions people use for social evaluations. Process research starts from an information-processor metaphor (e.g. decoding, encoding), to the predictive brain (e.g. statistical learning, probabilistic modeling), and now to a feedback loop framework (e.g. reinforcement learning, algorithmic bias), paving the way to understand how and why people evaluate others. Understanding stereotypes is a collective enterprise, as evidenced by the scholarly debate that has helped move the field forward.
Large language models (LLMs) can pass explicit bias tests but still harbor implicit biases, similar to humans who endorse egalitarian beliefs yet exhibit subtle biases. Measuring such implicit biases can be a challenge: as LLMs become increasingly proprietary, it may not be possible to access their embeddings and apply existing bias measures; furthermore, implicit biases are primarily a concern if they affect the actual decisions that these systems make. We address both of these challenges by introducing two measures of bias inspired by psychology: LLM Implicit Association Test (IAT) Bias, which is a prompt-based method for revealing implicit bias; and LLM Decision Bias for detecting subtle discrimination in decision-making tasks. Using these measures, we found pervasive human-like stereotype biases in 6 LLMs across 4 social domains (race, gender, religion, health) and 21 categories (weapons, guilt, science, career among others). Our prompt-based measure of implicit bias correlates with embedding-based methods but better predicts downstream behaviors measured by LLM Decision Bias. This measure is based on asking the LLM to decide between individuals, motivated by psychological results indicating that relative not absolute evaluations are more related to implicit biases. Using prompt-based measures informed by psychology allows us to effectively expose nuanced biases and subtle discrimination in proprietary LLMs that do not show explicit bias on standard benchmarks.
Interaction and cooperation with humans are overarching aspirations of artificial intelligence (AI) research. Recent studies demonstrate that AI agents trained with deep reinforcement learning are capable of collaborating with humans. These studies primarily evaluate human compatibility through "objective" metrics such as task performance, obscuring potential variation in the levels of trust and subjective preference that different agents garner. To better understand the factors shaping subjective preferences in human-agent cooperation, we train deep reinforcement learning agents in Coins, a two-player social dilemma. We recruit participants for a human-agent cooperation study and measure their impressions of the agents they encounter. Participants' perceptions of warmth and competence predict their stated preferences for different agents, above and beyond objective performance metrics. Drawing inspiration from social science and biology research, we subsequently implement a new "partner choice" framework to elicit revealed preferences: after playing an episode with an agent, participants are asked whether they would like to play the next round with the same agent or to play alone. As with stated preferences, social perception better predicts participants' revealed preferences than does objective performance. Given these results, we recommend human-agent interaction researchers routinely incorporate the measurement of social perception and subjective preferences into their studies.
Artificial intelligence increasingly suffuses everyday life. However, people are frequently reluctant to interact with A.I. systems. This challenges both the deployment of beneficial A.I. technology and the development of deep learning systems that depend on humans for oversight, direction, and regulation. Nine behavioral studies (N = 3,300) demonstrate that social-cognitive processes guide human interactions across a diverse range of real-world A.I. systems. Across studies, perceived warmth and competence emerge prominently in participants’ impressions of A.I. systems. Judgments of warmth and competence systematically depend on human-A.I. interdependence and autonomy. In particular, participants perceive systems that optimize interests aligned with human interests as warmer and systems that operate independently from human direction as more competent. Finally, a prisoner’s dilemma game shows that warmth and competence judgments predict participants’ willingness to cooperate with a deep-learning system. These results underscore the generality of intent detection to interactions with a broad array of algorithmic actors. Researchers and policymakers should carefully consider the degree and alignment of interdependence between humans and new artificial intelligence systems.
Errors in clinical decision-making are disturbingly common. Here, we show that structured information–sharing networks among clinicians significantly reduce diagnostic errors, and improve treatment recommendations, as compared to groups of ...Errors in clinical decision-making are disturbingly common. Recent studies have found that 10 to 15% of all clinical decisions regarding diagnoses and treatment are inaccurate. Here, we experimentally study the ability of structured information–sharing ...
Mental representations of human social groups – social stereotypes – are widespread and consequential. Such mental representations are systematic and multidimensional (e.g., stereotypes of immigrant groups are organized by perceived warmth and competence). We show that adaptive exploration alone can create structured societal stereotypes that cascade from historical affordances – such as which group happened to be the first one adequate at a job – without requiring decision-makers to have malicious intentions or cognitive limitations, or social groups to differ in information accessibility or intrinsic quality. Rather than framing social perception as a static, one-shot event, we consider the consequences of sequential decisions in a setting where exploring new options carries an implicit cost (resulting in an “explore-exploit tradeoff”). In this setting, when groups have equal and high potential to succeed in diverse jobs, decision-makers nonetheless settle on a single social group to perform each job and form impressions of that group accordingly. Using stereotypes of immigrant groups based on warmth and competence as an example, we formalize this process as a contextual multi-armed bandit problem, show testable predictions from computational simulations, and demonstrate that human participants act consistently with these predictions in behavioral experiments. Our results show how rich, multidimensional stereotypes can emerge in an absolutely minimal setting.
The spontaneous stereotype content model (SSCM) describes a comprehensive taxonomy, with associated properties and predictive value, of social-group beliefs that perceivers report in open-ended responses. Four studies (N = 1,470) show the utility of spontaneous stereotypes, compared to traditional, prompted, scale-based stereotypes. Using natural language processing text analyses, Study 1 shows the most common spontaneous stereotype dimensions for salient social groups. Our results confirm existing stereotype models' dimensions, while uncovering a significant prevalence of dimensions that these models do not cover, such as Health, Appearance, and Deviance. The SSCM also characterizes the valence, direction, and accessibility of reported dimensions (e.g., Ability stereotypes are mostly positive, but Morality stereotypes are mostly negative; Sociability stereotypes are provided later than Ability stereotypes in a sequence of open-ended responses). Studies 2 and 3 check the robustness of these findings by: using a larger sample of social groups, varying time pressure, and diversifying analytical strategies. Study 3 also establishes the value of spontaneous stereotypes: compared to scales alone, open-ended measures improve predictions of attitudes toward social groups. Improvement in attitude prediction results partially from a more comprehensive taxonomy as well as a construct we refer to as stereotype representativeness: the prevalence of a stereotype dimension in perceivers' spontaneous beliefs about a social group. Finally, Study 4 examines how the taxonomy provides additional insight into stereotypes' influence on decision-making in socially relevant scenarios. Overall, spontaneous content broadens our understanding of stereotyping and intergroup relations. (PsycInfo Database Record (c) 2022 APA, all rights reserved).