Popular advice in media and self-help literature encourages people to trust their intuition when forming impressions of strangers. We tested this widespread lay belief by examining whether intuitive judgments are preferred over deliberate ones and whether such preferences are justified by differences in accuracy. In Study 1 (n = 401), participants from the general population reported using intuition more frequently, preferring it, and expressing greater confidence in it compared to deliberate reasoning when judging strangers. In Study 2, we analyzed 21,739 judgments from 380 users of the quiz app “Who Knows,” who inferred personality-related characteristics from short video introductions. Judgment mode was experimentally manipulated within persons across multiple game rounds. While intuitive judgments were again preferred, accuracy did not differ between conditions. Thus, although intuition is favored when judging strangers, it does not yield higher accuracy but may still be adaptive by achieving comparable accuracy with less time.
In first social encounters, people regularly try to figure out whether others like them. When accurate, these meta-liking judgments enable individuals to more optimally engage in social interactions. Although previous research shows that people do know how much others like them, little is known about the processes leading to accurate meta-liking judgments. In the present, pre-registered research, we aimed to close this gap by investigating the role of a meta-perceiver’s appearance and behavior, others’ behavior towards that meta-perceiver, and the reciprocity heuristic in forming accurate meta-liking judgments. To this end, we used data from a round-robin study, in which N = 144 participants (Mage = 22.91, 67.6% female) took part in two group meetings, where they interacted with unacquainted others. After each participant had given a short self-introduction, participants indicated liking and meta-liking for each other participant. Additionally, trained raters judged participants’ appearance and behavior during the self-introduction. Results showed that people were able to correctly infer how much others generally liked them but hardly knew how much specific other liked them. Lens model analyses revealed that although one’s appearance and behavior were valid cues for being generally liked, meta-perceivers did not utilize most of these observable cues to inform their meta-liking judgments. However, the assumed reciprocity mediated both individual and dyadic meta-liking accuracy. In light of these findings, we discuss the role of self-observed behavior and stable personality characteristics as well as situational and motivational factors, presenting an integrative theoretical framework to explain generalized and dyadic meta-liking accuracy.
Most social groups punish freeriders (i.e., individuals who receive the same benefits from the group as others, despite contributing less to its success). In small groups, individual group members (rather than established authorities) typically implement this punishment spontaneously, and punishers may consequently be awarded social status by their peers. Here, we tested the preregistered hypothesis that this way of acquiring status works best when fellow group members are high, rather than low, in right-wing authoritarianism (RWA). The hypothesis was supported in a laboratory-based behavioral study (N = 667) in which small groups interacted in a financially incentivized repeated public goods game involving punishment (i.e., a social dilemma that puts immediate individual benefits at odds with long-term collective interests). Linking the process of status acquisition to peer RWA significantly advances the understanding of social dynamics in small groups.
Perhaps surprisingly, personality science has not yet adequately addressed a very basic aspect of individuality: The extent to which an individual differs from others, a phenomenon we refer to as unusualness. In this work, we provide a framework for defining, measuring, and uncovering the elements of unusualness. First, we define unusualness as the extent to which an individual’s personality deviates from that of others in the population, either through rare expression of individual variables (unusual values) or through atypical co-occurrence of multiple variables (unusual combinations), and distinguish it from related concepts in personality research. Second, we introduce and compare various methods to measure unusualness in a simulation study and demonstrate how reliability and validity of unusualness measures can be established. Third, we propose methods to uncover the elements of unusualness (i.e., unusual values, unusual combinations) and validate these methods through simulation. Fourth, we illustrate our framework using three empirical datasets (combined N = 1,834) and multiple assessment methods, showing that unusualness can be reliably measured and offering first insights into its generalizability, nomological network, incremental validity, and possible explanatory factors. Finally, we outline conceptual, methodological, and empirical directions for future research on unusualness.
Individual differences in psychological reactivities (i.e., the degree to which individuals react differently to social interactions) are central to psychological research. Previous theory-based research has identified substantial individual differences in reactivities but few robust predictors of these differences. This work aimed to address two questions: First, can individual differences in reactivities to social interactions be accurately predicted at all? Second, what are the most important person-level variables for this prediction? A data-driven machine learning approach was applied to three large-scale experience-sampling data sets (overall N = 5,047) to predict the extent to which individuals reacted with positive and negative affect to momentary social interaction characteristics (e.g., interaction depth). Individual differences in reactivities were extracted via multilevel modeling (i.e., random slopes) and then predicted with machine learning methods using a variety of person-level variables (i.e., socio-demographics, personality traits, political and societal attitudes). The robustness of predictions was examined by built-in cross-validation and across independent samples. Feature importance and interactions were analyzed with SHapley Additive exPlanations values. Our results suggest that, whereas complex prediction models outperformed a baseline model in predicting individual differences in reactivities in most analyses, the overall predictive performance was limited. This finding underlines the importance of replicating machine learning results across outcomes and independent samples. We revealed several predictive patterns that can stimulate future research, elaborate on limitations of current machine learning approaches for intensive within-person data, and discuss the results against the background of dynamic conceptualizations of personality.
Personalized medicine requires physicians to adapt clinical communication and decision-making to patients’ individual motives, emotions, and interpersonal behaviors. However, training these skills remains challenging, as established simulation-based formats—particularly actor-based simulations—are resource-intensive and difficult to scale. Consequently, there is a need for scalable and flexible training approaches that allow repeated practice across diverse patient personalities and clinical contexts. The PerTRAIN (Personalization Training in Medicine) project addresses this need by leveraging large language models (LLMs) to simulate virtual patients with dynamically adapting personality expression at scale. Grounded in Contemporary Integrative Interpersonal Theory (CIIT), patient behavior is modelled along the dimensions of agency and communion and updated in response to the medical trainee’s behavior, enabling systematic variation and dynamic adaptation of interpersonal behavior within medical scenarios. This allows trainees to learn how patient personality shapes communication, and to practice adaptive, patient-centered communication strategies and appropriate clinical decision-making. For the personalization trainings, we developed clinical cases in which personality expression substantially influences behavior in doctor–patient interactions, that represent common encounters in primary care, and that align with national guidelines for patient-centered care. An initial chat-based implementation enables structured interactions, dynamic personality adaptation, and iterative refinement. The system is designed as a complementary tool to existing simulation formats, offering a scalable, low-threshold environment for repeated practice and reflection. A first application in curricular teaching is scheduled for 2026. Future extensions of the framework are discussed, such as large-scale empirical validation, modelling long-term interpersonal trajectories, and the extension of interactions to multimodal formats.
Narcissism is a personality trait with far-reaching individual, social, and societal consequences. Thus, it is important to understand the sources of individual differences in this trait. Existing hypotheses on the development of narcissism have focused on familial and parental environments that act to make siblings in a family more alike. However, the relative importance of shared environments as opposed to other environmental and genetic sources is still unclear. Using a large extended twin family design, we found that parents’ and children’s narcissism scores were correlated, but this association was entirely genetically driven. Across age and measures, genetics and individual-specific environmental factors each explained 50% of the variance in narcissism and there was no evidence of environmental sources shared within families. This finding calls for a fundamental shift in the search for the origins of narcissism, including extra-familial environmental factors (e.g., educational and occupational pathways, experiences with peers, and romantic partners).
Self-knowledge plays a central role in contemporary psychological science across various domains, including interpersonal relationships, moral behaviour and health. Despite its importance, many fundamental questions remain. We conducted a pre-registered, expert-based consensus process to address four key gaps in research on self-knowledge: its conceptualization, measurement, outcomes and changeability. Seventeen experts from diverse subfields of psychology participated in a structured Delphi process guided by four facilitators and an external advisor. The panel developed a consensus definition of self-knowledge as the extent to which a person has accurate perceptions of their own relatively stable characteristics and momentary states. Experts further agreed that self-knowledge is largely domain-specific, context-dependent in its benefits, and malleable in principle but difficult to change in practice. Measurement was identified as a central challenge, and avenues for refinement in future work were proposed. Consensus was weaker regarding the existence of a domain-general factor of self-knowledge and shared underlying processes across domains. Overall, the findings clarify where experts converge, where debates persist and what should be prioritized in future research, providing a crucial foundation for advancing the study of self-knowledge across fields. Self-knowledge plays a central role in contemporary psychological science. However, unresolved conceptual and methodological issues have hindered theoretical integration and cumulative scientific progress. In this Consensus Statement, Thielmann et al. identify gaps in four key areas of self-knowledge research: its conceptualization, measurement, outcomes and changeability.
Concerns about the generalizability of machine learning models in mental health arise, partly due to sampling effects and data disparities between research cohorts and real-world populations. We aimed to investigate whether a machine learning model trained solely on easily accessible and low-cost clinical data can predict depressive symptom severity in unseen, independent datasets from various research and real-world clinical contexts. This observational multi-cohort study included 3021 participants (62.03% females, MAge = 36.27 years, range 15–81) from ten European research and clinical settings, all diagnosed with an affective disorder. We firstly compared research and real-world inpatients from the same treatment center using 76 clinical and sociodemographic variables. An elastic net algorithm with ten-fold cross-validation was then applied to develop a sparse machine learning model for predicting depression severity based on the top five features (global functioning, extraversion, neuroticism, emotional abuse in childhood, and somatization). Model generalizability was tested across nine external samples. The model reliably predicted depression severity across all samples (r = 0.60, SD = 0.089, p < 0.0001) and in each individual external sample, ranging in performance from r = 0.48 in a real-world general population sample to r = 0.73 in real-world inpatients. These results suggest that machine learning models trained on sparse clinical data have the potential to predict illness severity across diverse settings, offering insights that could inform the development of more generalizable tools for use in routine psychiatric data analysis.
In studies using the increasingly popular Experience Sampling Method (ESM), design decisions are often guided by theoretical or practical considerations. Yet limited empirical evidence exists on how these choices impact data quantity (e.g., response probabilities), data quality (e.g., response latency), and potential biases in study outcomes (e.g., characteristics of study variables). In a preregistered, four-week study (N = 395), we experimentally manipulated two key ESM protocol characteristics for sending ESM surveys: timing (fixed versus varying times) and contingency (directly versus indirectly after unlocking the smartphone). We evaluated the ESM protocols resulting from the combination of these two characteristics with regard to different criteria: As hypothesized for contingency, indirect protocols resulted in higher response probabilities (increased data quantity). But they also led to higher response latencies (reduced data quality). Contrary to our expectations, the combined effect of contingency and timing did not significantly influence response probability. We did also not observe other effects of timing or contingency on data quality. In exploratory follow-up analyses, we discovered that timing significantly affected response probability and smartphone usage behaviors, as measured by screen logs; however, these effects were likely attributable to time of day effects. Notably, self-reported states showed no differences based on the chosen ESM protocol, and similar trends were found when correlating primary outcomes with external criteria such as trait affect and well-being. Based on the study’s findings, we discuss the trade-offs that researchers should consider when choosing their ESM protocols to optimize data quantity, data quality, and biases in study outcomes.
Individual differences in social traits such as the affiliation motive are closely linked to the formation and maintenance of social relationships. Most previous research focused on long-term characteristics or momentary assessments of social relationships (e.g., social network size, relationship quality), whereas theoretical accounts have emphasized the temporal dynamics, that is, how social interactions unfold over time. The present studies examined how social interactions unfold within days as well as between days, taking personality traits and situational affordances into account. In two multimethod studies (Study 1: N = 307, age 18-80 years, 51% female; Study 2: N = 385, age 19-84 years, 48% female), we assessed participants' social interactions in daily life using ecological momentary assessments and mobile sensing over 2 and 14 days, respectively. Furthermore, participants answered questionnaires on affiliation, additional social traits, and situational affordances, for example, the voluntariness of social situations. Multilevel lead-lag analyses showed that affiliation predicted momentary social desires but not future social interactions, except when social interactions were assessed with unobtrusive mobile sensing. Situational affordances, such as the valence and voluntariness of social interactions, additionally predicted social desires and future contact. The results were largely specific to affiliation and not observed for extraversion. Future research on social interactions would benefit from (a) examining and specifying meaningful timescales of social relationship processes, (b) following the renewed interactionist call for considering person and situation factors, and (c) integrating the myriad of social trait concepts in theories and measurements. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
INTRODUCTION:Everyday experience as well as the research literature on trait attributions suggest that people use nonverbal cues when judging the personality of a person. However, little research has reported on people's explicitly held beliefs about these associations. METHODS:Two hundred forty-five participants recruited through Amazon's Mechanical Turk rated how strongly they thought 20 nonverbal cues are related to each of the Big Five traits. Their beliefs were then compared to a previous meta-analysis to see how explicit beliefs compare to implicit beliefs measured in lens models (cue utilizations) and to actual links between the Big Five and nonverbal cues (cue validities). RESULTS:Participants' explicit beliefs formed coherent constellations for each trait. The explicit beliefs corresponded generally well with implicit beliefs as well as with cue validities. CONCLUSION:The results support the validity of explicit beliefs about nonverbal cues and the Big Five, offering new opportunities for researchers interested in how beliefs affect interpersonal interactions.
Which behavioral and visual information do teachers rely on when judging relevant characteristics of their students and which cues should they rely on? Drawing on Brunswik’s Lens Model (Perception and the representative design of psychological experiments, University of California Press, 1956. https://doi.org/10.1525/9780520350519 ), we investigated the role of students’ expression of nonverbal behavioral cues (e.g., friendly facial expression) and physical appearance (e.g., wearing eyeglasses) and how this information is utilized during the judgment process by pre-service teachers and psychology students (N = 102). Perceivers provided ratings of students’ (N = 45) academic self-concept, intelligence and motivation in brief nonverbal video clips showing one student each in a physics classroom. Numerous behavioral and physical cues (in total 165) were extracted from the stimulus material by two independent raters. Perceivers achieved highest accuracy for students’ motivation, whereas intelligence was judged with the lowest accuracy. Lens model parameter analysis indicated that perceivers strongly relied on students’ sex, an attentive and self-assured facial expression, and whether or not a student was wearing eyeglasses in their judgments. Cues that were actually related to students’ characteristics, on the other hand, involved students’ sex, a masculine and distinctive appearance, and a tensed as well as friendly facial expression. An overall favorable judgment for boys points into the direction of a gender bias. Implications for our understanding of teacher judgment processes and outcomes are discussed.
In the context of growing international migration, it is crucial to understand factors that might alleviate or amplify threat perceptions by outgroups. Hereby, the role of subjective societal status (SSS), religiosity, Right-Wing Authoritarianism (RWA) and Social Dominance Orientation (SDO) are not fully understood. In a large online survey ( N = 1257), we investigated the joint and interactive effects of RWA, SDO, SSS, and religiosity on German residents’ threat perceptions by Middle Eastern immigrants. Higher RWA and SDO, and lower SSS, predicted both symbolic and realistic threat, even after controlling for income, education, age, and gender. Furthermore, higher SSS buffered the effect of RWA and SDO for realistic threat, while religiosity was not related to threat perceptions and did not moderate RWA or SDO threat associations. We discuss methodological limitations and implications of our findings for the understanding of societal conflict.
The Hypersensitive Narcissism Scale (HSNS) is a an economical, widely used self-report measure of vulnerable narcissism. Developed and mostly used as a unidimensional scale, previous structural examinations suggest two correlated dimensions, one emphasizing hypersensitive/neurotic aspects and the other highlighting egocentric/antagonistic aspects of vulnerable narcissism. The few extant factor analyses of the HSNS, however, differ profoundly in their methodological approach, the resulting item-to-factor assignment, and lack a thorough validation of the two putative subscales. To fill these gaps, we systematically examined and compared alternative measurement models for the HSNS and conducted comprehensive correlation analyses to map the proposed HSNS dimensions onto current models of general personality, narcissism, and psychopathology. In a first study, we constructed a German adaptation using data from three large samples (accumulated N = 3,655). In-depth examination of this German HSNS (Study 2, N = 1,359) confirmed the dimensions Oversensitivity and Egocentrism and suggested at least metric model invariance across gender and age. The two dimensions displayed distinct nomological nets and differed with respect to various personality traits, personality pathology markers, the Hierarchical Taxonomy of Psychopathology, and psychological (mal-)adjustment. HSNS-Oversensitivity corresponds with measures of neurotic narcissism and predicts internalizing pathology and intrapersonal dysfunctions, whereas Egocentrism overlaps with antagonistic narcissism, low agreeableness, and externalizing problems. Taken together, our research reconciles the HSNS with other multidimensional narcissism measures as well as current dimensional models of personality and psychopathology and attests to its utility to capture vulnerable narcissistic traits at a finer grained level. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Right-wing authoritarianism (RWA) refers to an adherence to conventional values and authorities with the power to penalize groups that are perceived to challenge the cohesion of ingroup norms. Correspondingly, RWA has repeatedly been linked to negative perceptions of minoritized groups, such as refugees or religious minorities. To investigate whether and how sociocultural factors add to and moderate how RWA influences perceptions that minoritized groups pose a threat (i.e. threat perceptions), we examined (a) the value of RWA, religiosity and perceived societal marginalization in predicting these threat perceptions across countries, (b) potential moderating effects of individual- and country-level religiosity and marginalization on the RWA-threat link and (c) the robustness of cross-sectional findings when daily threat perceptions were assessed longitudinally. We used cross-sectional survey data from Germany N = 1896; Study (1) and Europe N = 3227; Study (2) and global cross-sectional and longitudinal daily diary data N = 3154 individuals; N >52,447 assessments; N = 41 countries; Study (3). Our studies point to the significance of contextual conditions and the generalizability of cross-sectional findings to day-to-day assessments of threat perceptions.
Judging individual differences of interaction partners is a key mechanism of human social functioning. However, investigating the behavioral underpinnings of these judgments at a larger scale has traditionally been difficult. We present a machine learning-based approach for cue extraction and integration allowing for large-scale, fine-grained behavioral analyses of social judgments and showcase its application for the case of language behavior and judgments of individual differences in performance. We used a natural language processing approach to extract granular verbal and paraverbal language cues from audio streams and transcripts from a high-stakes assessment center (N = 556, C = 20,289 cues). We subsequently leveraged machine learning models to examine how well targets' language behavior predicted perceivers' performance judgments. We then analyzed the predictivity of different language domains and identified the cues that drove our predictions. We found that both verbal and paraverbal language behavior predicted the performance judgments, but a combination of the two domains led to limited improvement in predictive performance. Additional cue level analyses revealed that the utilized cues from both subdomains expressed similar information. We discuss contributions to the performance judgment literature as well as implications for future research on judgments of individual differences in general.
Machiavellianism (Mach) is a personality trait characterized by cold rationality, cynicism, duplicity, and the strategic and egotistical pursuit of goals. Despite recent advances in the measurement of Mach, most Mach scales show limited content validity because they have not systematically integrated recurring Mach themes in the areas of affect (A), behavior (B), cognition (C), and desire (D). To overcome this and other issues, we developed a new Mach scale, the M4. We created the M4 by using Ant Colony Optimization (ACO) to select 16 items from a pool of 92 newly generated and expert-rated Mach items. In two studies with German-speaking participants (N1 = 765; N2 = 1,288), the M4 total score showed high reliability and high convergent validity with established Mach scales and M4 informant reports. Furthermore, the nomological network of the M4 aligned in several respects with the theoretical conceptualization of Mach. However, some unexpected associations suggested the need to refine the conceptualization of Mach regarding its relationship with certain forms of impulsivity and neuroticism. Commonality analyses further indicated that the M4 predicted incentivized cheating behavior better than five other recently developed Mach measures. Hence, the M4 holds promise for advancing the assessment of Mach.
Interpersonal judgments play a central role in human social interactions, influencing decisions ranging from friendships to presidential elections. Despite extensive research on the accuracy of these judgments, an overreliance on broad personality traits and subjective judgments as criteria for accuracy has hindered progress in this area. Further, most individuals involved in past studies (either as judges or targets) came from ad-hoc student samples which hampers generalizability. This paper introduces Who Knows (https://whoknows.uni-muenster.de), an innovative smartphone application designed to address these limitations. Who Knows was developed with the aim to create a comprehensive and reliable database for examining first impressions. It utilizes a gamified approach where users judge personality-related characteristics of strangers based on short video introductions. The project incorporates multifaceted criteria to evaluate judgments, going beyond traditional self-other agreement. Additionally, the app draws on a large pool of highly specific and heterogenous items and allows users to judge a diverse array of targets on their smartphones. The app's design prioritizes user engagement through a responsive interface, feedback mechanisms, and gamification elements, enhancing their motivation to provide judgments. The Who Knows project is ongoing and promises to shed new light on interpersonal perception by offering a vast dataset with diverse items and a large number of participants (as of fall 2024, N = 9,671 users). Researchers are encouraged to access this resource for a wide range of empirical inquiries and to contribute to the project by submitting items or software features to be included in future versions of the app.