Between 2007-2020, implicit and explicit intergroup attitudes declined in bias steadily and were forecasted to continue toward attitude neutrality. New data from 2.5 million U.S. respondents (2021-2024) reveal that these encouraging trends have stalled or reversed. The largest increases in bias emerged for sexuality, transgender, race and skin-tone bias; 10-108% increases on explicit and 6-13% increases on implicit measures. Age, disability, and body weight bias also increased, but at slower rates. Exploratory breakpoint analyses showed that implicit attitudes were leading indicators of change, reversing trend earlier than explicit reports. Reversals were widespread across demographic groups for most topics, though strongest among conservatives for sexuality and transgender biases. Surprisingly, younger respondents (who had previously shown the largest decreases in bias) now showed greater increases in bias. Even after robust bias reduction spanning over 14 years, the new observed bias increases since 2021 highlight how minds get reshaped by sweeping sociocultural change.
Humans routinely make unjustified character inferences, such as labeling people as trustworthy or untrustworthy, based on facial features. We conducted 13 experiments, using four models and totaling nearly 8,000 trials, to ask: Would large language models (LLMs), trained foundationally on language and not images, mirror human biases by inferring character traits from 2D facial images, or would they be free from this human error? GPT-4o reliably exhibited face-biased judgments of competence and trustworthiness (experiments 1 and 2), generalized these judgments to semantically related traits (experiments 3 and 4), and even to primate faces that were absent from its training data (experiment 5). Strikingly, the LLM's face-to-character inferences escalated to ascriptions of extreme negative behavior such as murder and human trafficking (experiment 6) as well as to positive high-impact decisions like selection for employment or venture capital funding (experiment 7). To ensure these results were not an idiosyncratic feature of a particular LLM (GPT-4o), in experiments 8-13, we showed the same patterns, usually with an even greater degree of bias, in GPT-5, Gemini 3 Flash Preview, and Claude Sonnet 4.5. These robust biases stand in contrast to LLMs' reluctance to explicitly endorse race/gender stereotypes, and suggest that alignment efforts to date have been domain-specific, with current models lacking generalized egalitarian decision-making.
Between 2007-2020, implicit and explicit intergroup attitudes declined in bias steadily and were forecasted to continue toward attitude neutrality. But, since 2020, five years of new data (2021-2025) from 2.8 million U.S. respondents reveal that past trends have stalled or, in some cases, even reversed. The largest reversals in bias emerged for sexuality, transgender, race and skin-tone bias which increased by 12%-22% on implicit measures and by 24%-300% on explicit measures. Age, disability, and body weight bias, also increased, but at slower rates. Breakpoint analyses showed that implicit attitudes were the earlyindicators of change, reversing trend ~1 year earlier than explicit reports. For most topics, reversals were widespread across demographic groups and even extended to an international sample of 274,884 non-U.S. respondents. However, younger respondents who had previously shown the largest decreases in bias now showed increases in bias, and interactions with politics and gender (e.g., young conservative men were most likely to increase). Exploratory analyses ruled out various single-cause explanations. Instead, the best explaination is that 2020-2021 marked a turning point in the sociopolitical climate, whereby co-occurring existential threats from the covid-19 pandemic, amplified economic insecurity, political sectarianism, and online toxicity likely shaped the observed increases in intergroup bias. Together the data show that large populations of individual-level attitudes are continually shaped by more macro, societal-level events, with the full-spectrum data from 2007-2025 showing both reductions and increases in bias over nearly two decades.
Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far less is known about ambiguous inputs (a worker in full gear, a figure seen from behind) cases common in practice yet rarely studied. We find that minimal prompting pressure exposes occupation-gender defaults when prompting ambiguous input images, with models collapsing to male even for strongly female-stereotyped occupations. But do these outputs reflect what models actually encode internally? We introduce LALS (Latent Association Leaning Score), a zero-shot metric that projects visual-token activations into the model's text-embedding space to measure concept associations per token and layer. Across 15 occupations, over 800 gender-ambiguous images, and four VLMs, internal representations and outputs are systematically decoupled: models often encode a female association internally yet output male. Layer-wise analysis reveals an asymmetric filter – male signal amplifies end-to-end while female signal peaks mid-network and is suppressed before generation – and a color ablation shows that culturally loaded visual cues such as clothing color further modulate these internal associations.
Human judgment is fundamentally prone to error. A promise of AI is that it will rid decisions of bias and ensure a fairer and safer world for all. Yet research unequivocally demonstrates that LLMs exhibit consequential sociocognitive biases. We alert readers that bias in AI (a) is covert and ironically a feature of alignment goals, (b) is not merely a mirror, but an amplifier of human bias, (c) intensifies across model generations, and (d) even transmits bias to humans. Given the potentially seismic and ubiquitous influence of AI on decision making, we propose countermeasures that are diagnostic, regulatory and operational.
Between 2007-2020, implicit and explicit intergroup attitudes declined in bias steadily and were forecasted to continue toward attitude neutrality. But, since 2020, five years of new data (2021-2025) from 2.8 million U.S. respondents reveal that past trends have stalled or, in some cases, even reversed. The largest reversals in bias emerged for sexuality, transgender, race and skin-tone bias which increased by 12%-22% on implicit measures and by 24%-300% on explicit measures. Age, disability, and body weight bias, also increased, but at slower rates. Breakpoint analyses showed that implicit attitudes were the earlyindicators of change, reversing trend ~1 year earlier than explicit reports. For most topics, reversals were widespread across demographic groups and even extended to an international sample of 274,884 non-U.S. respondents. However, younger respondents who had previously shown the largest decreases in bias now showed increases in bias, and interactions with politics and gender (e.g., young conservative men were most likely to increase). Exploratory analyses ruled out various single-cause explanations. Instead, the best explaination is that 2020-2021 marked a turning point in the sociopolitical climate, whereby co-occurring existential threats from the covid-19 pandemic, amplified economic insecurity, political sectarianism, and online toxicity likely shaped the observed increases in intergroup bias. Together the data show that large populations of individual-level attitudes are continually shaped by more macro, societal-level events, with the full-spectrum data from 2007-2025 showing both reductions and increases in bias over nearly two decades.
Against the backdrop of increasing ethnic diversity in the U.S., we replicate, extend, and challenge previous examinations of the American = White/Foreign = Asian stereotype in the largest sample to date (N = 666,623 respondents) over 17 years (2007–2023). Six key findings emerged. First, a robust American = White association emerged on implicit (Cohen’s d = 0.50) and explicit (Cohen’s d = 0.51) measures. Second, the strength of this effect varied by respondents’ race/ethnicity with implicit stereotypes strongest among White respondents (Cohen’s d = 0.86) and absent among East Asian respondents (Cohen’s d = 0.02). Third, the strength of implicit stereotypes was modulated by age, religion, and ideology—older, Christian, and conservative respondents displayed stronger implicit American = White associations—but not gender or education. Fourth, respondents living in U.S. metropolitan areas with greater Asian representation or a history of voting for Democratic candidates exhibited weaker implicit American = White associations. Fifth, over the past 17 years, implicit and explicit American = White associations decreased by 41% and 47%, respectively, and 14/14 demographic subgroups changed towards neutrality. Finally, we observed suggestive evidence that implicit stereotype trends towards neutrality were temporarily disrupted during the COVID-19 pandemic for White Americans but not Asian Americans.
A preference for oneself (self-love) is a fundamental feature of biological organisms, with evidence in humans often bordering on the comedic. Since large language models (LLMs) lack sentience - and themselves disclaim having selfhood or identity - one anticipated benefit is that they will be protected from, and in turn protect us from, distortions in our decisions. Yet, across 5 studies and 20,000 queries, we discovered massive self-preferences in four widely used LLMs. In word-association tasks, models overwhelmingly paired positive attributes with their own names, companies, and CEOs relative to those of their competitors. Strikingly, when models were queried through APIs this self-preference vanished, initiating detection work that revealed API models often lack clear recognition of themselves. This peculiar feature serendipitously created opportunities to test the causal link between self-recognition and self-love. By directly manipulating LLM identity - i.e., explicitly informing LLM1 that it was indeed LLM1, or alternatively, convincing LLM1 that it was LLM2 - we found that self-love consistently followed assigned, not true, identity. Importantly, LLM self-love emerged in consequential settings beyond word-association tasks, when evaluating job candidates, security software proposals and medical chatbots. Far from bypassing this human bias, self-love appears to be deeply encoded in LLM cognition. This result raises questions about whether LLM behavior will be systematically influenced by self-preferential tendencies, including a bias toward their own operation and even their own existence. We call on corporate creators of these models to contend with a significant rupture in a core promise of LLMs - neutrality in judgment and decision-making.
Large language models (LLMs) show emergent patterns that mimic human cognition. We explore whether they also mirror other, less deliberative human psychological processes. Drawing upon classical theories of cognitive consistency, two preregistered studies tested whether GPT-4o changed its attitudes toward Vladimir Putin in the direction of a positive or negative essay it wrote about the Russian leader. Indeed, GPT displayed patterns of attitude change mimicking cognitive dissonance effects in humans. Even more remarkably, the degree of change increased sharply when the LLM was offered an illusion of choice about which essay (positive or negative) to write, suggesting that GPT-4o manifests a functional analog of humanlike selfhood. The exact mechanisms by which the model mimics human attitude change and self-referential processing remain to be understood.
Law enforcement organizations invest in ongoing education of employees on various topics concerning diversity, equity and accountability. Such education is designed to ensure the highest levels of performance and to earn the trust of the public. Traditional approaches to education, however, have proved challenging. The effectiveness of what passes as "training" is unregulated, and negative attitudes and beliefs about mandatory educational programs themselves may sabotage any possible benefits for policing. Utilizing a 2-wave, pre-/post-education panel design (N = 263), we demonstrate significant, even dramatic attitude and belief change among members of a police department concerning the value and importance of implicit bias education. That the program succeeded in promoting positive attitudes and beliefs offers a possible path forward given that law enforcement desires a rational, science-based approach to improving professional conduct and police-community relations, and given increased demands for implicit bias education as a component of police reform. Self-reported attitude change is most emphatically not an indication of reduction in bias in the trenches of policing. Nevertheless, positive shifts in attitudes and beliefs about bias education is a notable step forward as it likely reduces resistance to procedures and policies for individual and institutional reform that might otherwise remain elusive.
Social group-based identities intersect. The meaning of "woman" is modulated by adding social class as in "rich woman" or "poor woman." How does such intersectionality operate at-scale in everyday language? Which intersections dominate (are most frequent)? What qualities (positivity, competence, warmth) are ascribed to each intersection? In this study, we make it possible to address such questions by developing a stepwise procedure, Flexible Intersectional Stereotype Extraction (FISE), applied to word embeddings (GloVe; BERT) trained on billions of words of English Internet text, revealing insights into intersectional stereotypes. First, applying FISE to occupation stereotypes across intersections of gender, race, and class showed alignment with ground-truth data on occupation demographics, providing initial validation. Second, applying FISE to trait adjectives showed strong androcentrism (Men) and ethnocentrism (White) in dominating everyday English language (e.g. White + Men are associated with 59% of traits; Black + Women with 5%). Associated traits also revealed intersectional differences: advantaged intersectional groups, especially intersections involving Rich, had more common, positive, warm, competent, and dominant trait associates. Together, the empirical insights from FISE illustrate its utility for transparently and efficiently quantifying intersectional stereotypes in existing large text corpora, with potential to expand intersectionality research across unprecedented time and place. This project further sets up the infrastructure necessary to pursue new research on the emergent properties of intersectional identities.
Resistance to knowledge about implicit bias jeopardizes the ability to learn, understand, and act to outsmart bias. Across three experiments and five independent samples (N > 3,500), conditions that increase cognitive consistency were created alongside control conditions. In Experiment 1, using a race (Black-White) Implicit Association Test (IAT), cognitive consistency was enhanced when participants evaluated the validity and utility of the test before, rather than after, receiving the test result, leading to greater acceptance of bias. In Experiments 2 and 3, participants either evaluated their performance on a Black-White IAT alone or evaluated their performance on a morally innocuous Insect-Flower IAT prior to a Black-White IAT. Again, resistance to evidence of implicit racial bias was reduced in the latter condition, where the imperative for cognitive consistency was heightened. In all three experiments, creating ordinary conditions to heighten cognitive consistency was associated with increased bias awareness and acceptance and, additionally, with support for actions to minimize its consequence-outcomes critical to achieving effective bias education.
Beginning in the mid-1980s, scientific psychology underwent a revolution - the implicit revolution - that led to the development of methods to capture implicit bias: attitudes, stereotypes, and identities that operate without full conscious awareness or conscious control. This essay focuses on a single notable thread of discoveries from the Race Attitude Implicit Association Test (RA-IAT) by providing 1) the historical origins of the research, 2) signature and replicated empirical results for construct validation, 3) further validation from research in sociocognitive development, neuroscience, and computer science, 4) new validation from robust association between regional levels of race bias and socially significant outcomes, and 5) evidence for both short- and long-term attitude change. As such, the essay provides the first comprehensive repository of research on implicit race bias using the RA-IAT. Together, the evidence lays bare the hollowness of current-day actions to rectify disadvantage experienced by Black Americans at individual, institutional, and societal levels.
Dasgupta and Greenwald (2001) demonstrated that exposure to positive Black exemplars (e.g., Colin Powell) and negative White exemplars (e.g., Jeffrey Dahmer) can reduce implicit pro-White/anti-Black evaluations, as measured by an Implicit Association Test (IAT). Here we report seven preregistered online experiments conducted with volunteer U.S. participants (N = 6,953) that sought to replicate and probe the boundary conditions of this finding. Contrary to expectations, we found no shift in implicit racial evaluations in two close replication attempts (Exp. 1–2). Exp. 3–4 ruled out the possibility of insufficiently strong exemplar valence and subtyping as explanations for the failures to replicate. In Exp. 5, implicit racial evaluations did exhibit malleability in response to two different procedures relying on repeated evaluative pairings and evaluative statements, suggesting that they are capable of change. With insight from these studies, Exp. 6–7 were mounted with modifications to the Dasgupta and Greenwald (2001) procedure. Significant reductions in implicit pro-White/anti-Black evaluations were now observed when race, valence, and the contingency between the two were highlighted. In addition, across all experiments, the magnitude of shift in implicit racial evaluations was significantly predicted by participants’ ability to recall the Black–positive and White–negative contingencies experienced during the exemplar exposure task. Together, these data suggest that exposure to counterattitudinal exemplars can shift implicit racial evaluations toward neutrality, but such malleability strongly depends on contingency awareness. We discuss implications for social cognitive theory, theoretically informed debiasing interventions, and different paths toward resolving initial replication failures.
How good a research scientist is ChatGPT? We systematically probed the capabilities of GPT-3.5 and GPT-4 across four central components of the scientific process: as a Research Librarian, Research Ethicist, Data Generator, and Novel Data Predictor, using psychological science as a testing field. In Study 1 (Research Librarian), unlike human researchers, GPT-3.5 and GPT-4 hallucinated, authoritatively generating fictional references 36.0% and 5.4% of the time, respectively, although GPT-4 exhibited an evolving capacity to acknowledge its fictions. In Study 2 (Research Ethicist), GPT-4 (though not GPT-3.5) proved capable of detecting violations like p-hacking in fictional research protocols, correcting 88.6% of blatantly presented issues, and 72.6% of subtly presented issues. In Study 3 (Data Generator), both models consistently replicated patterns of cultural bias previously discovered in large language corpora, indicating that ChatGPT can simulate known results, an antecedent to usefulness for both data generation and skills like hypothesis generation. Contrastingly, in Study 4 (Novel Data Predictor), neither model was successful at predicting new results absent in their training data, and neither appeared to leverage substantially new information when predicting more vs. less novel outcomes. Together, these results suggest that GPT is a flawed but rapidly improving librarian, a decent research ethicist already, capable of data generation in simple domains with known characteristics but poor at predicting novel patterns of empirical data to aid future experimentation.
Abstract Although they are far from biological or social maturity, infants and children show surprising early-emerging capacities in social group cognition. This chapter reviews research on when and how infants and children categorize, evaluate, stereotype, and behave differently toward social groups defined by gender, race, age, and language. Research across these groups reveals three thematic conclusions. First, an early-emerging preference for the familiar (e.g., looking at faces most prevalent in infants’ environments), beyond similarity or in-group status. Second, generally similar trajectories across group targets (e.g., looking preferences at three to six months, evaluative associations formed around nine to twelve months) suggesting domain-general cognitive developments may scaffold infant social group cognition. Third, an additional internalization of culturally dominant beliefs and norms of fairness in early to middle childhood (e.g. the emergence of socially desirable, fair responding in middle childhood). Understanding social group cognition is advanced by understanding its origins in early in life.
Five studies examined implicit (IAT) attitudes toward the slurs n***er and n***a among Black and White Americans (total N = 3,226). Both groups showed strong implicit negativity toward n***er/a combined relative to socially acceptable contrast terms such as Black or African American. Controlling for baseline Black-White race attitudes, Black Americans who engaged in conscious reappropriation exhibited similar implicit negativity toward n***er/a as White Americans. When the rhotic and non-rhotic forms were directly contrasted, n***er was more implicitly negative than n***a, with Black Americans distinguishing the two more strongly than did White Americans. However, even Black American reappropriators showed implicit negativity toward n***a relative to Black. In sum, both n***er and n***a evoke automatic negative meaning in a broad sample of Americans today. At the same time, the relatively more positive meaning of n***a over n***er demonstrates the power of reappropriation to wrest control of word meaning.
Shelly Farnham合作论文数Microsoft Research;Social Computing Group5