Two different approaches for calculating confidence intervals (CIs) for individual scores in psychological testing practice have been discussed in the literature within the framework of classical test theory, but unfortunately, both approaches can lead to biased estimates. The traditional approach (CI: observed score ± z · standard error of measurement) fails to take into account the phenomenon that, with imperfect measurement, true scores will be closer to the population average than observed scores will be (regression to the mean). On the other hand, the regression approach (CI: regression-based true score estimate ± z · standard error of the estimate) takes regression to the mean into account. However, this approach leads to confidence intervals that are on a different scale than the observed scores. The different scaling occurs because true scores have a smaller standard deviation than observed scores, and the extent of this shrinkage depends on the reliability of the test. It is thus incorrect, for example, to interpret true scores as falling on a standard scale (like T-scores) even when the observed scores were measured on such a scale. Here, I suggest a scale correction for the regression-based true score estimate to preserve the scaling, and thus to ensure both the accuracy and interpretability of the confidence intervals. Simulations indicate that this approach has the desired properties and outperforms the two existing approaches. The regression approach with scale correction is therefore recommended for calculating confidence intervals for individual scores in psychological testing practice.
Interest in emotional variability as an interindividual difference is growing. Yet, basic features of the construct, such as its stability and reliability, are not well understood. To address this gap, we examined two longitudinal data sets, each comprising two waves of daily assessments, one with a 3-month and the other with a 16-month retest interval. Overcoming a key methodological limitation of past approaches, we used Bayesian censored location scale models as an alternative modeling approach that accounts for biases introduced by bounded rating scales. The results showed that the variability estimates from the models had reliabilities around rel = .64, which can be sufficient for group-level predictions. Additionally, the latent stability of r = .60 provides evidence for stable individual differences in emotional variability, suggesting that it is more than just a transient state. We ran exploratory analyses to further examine the influence of external events and individual life transitions on emotional variability and found that emotional variability was responsive to environmental changes.
When faced with self-threat, people often engage in self-protective reactions. Yet not everyone does. The extent to which self-protection occurs, is thought to depend on people’s personality. We put this claim to test by considering multiple personality traits from different content domains. In an experiment, participants ( N = 1744) performed a bogus performance test and received negative versus positive feedback. As an indicator of self-protective reactions, we assessed participants’ tendency to question the validity of the test. We found that high self-esteem and high narcissism predicted stronger self-protective reactions, particularly when these traits were assessed in the content domain targeted by the negative feedback. Contrary to our hypotheses, the self-insight motive and mindfulness also predicted stronger (and not weaker) self-protective reactions. These findings provide better understanding of the role that personality plays in motivated reasoning.
In this registered report ( N = 423), we investigated in a competitive intergroup context to what extent the perception of targets scoring high in grandiose narcissism varies depending on whether they belong to one’s own group or to an opposing outgroup. In a laboratory study, members of newly formed groups had direct contact with another group and competed for scarce resources. Contrary to our hypothesis, perceivers did not ascribe targets scoring high in narcissistic admiration higher status when they belonged to their ingroup versus the outgroup. Also unexpectedly, they did not like targets scoring high in narcissistic rivalry better when they belonged to their ingroup. Instead, our findings indicate that narcissistic admiration was generally linked to more dominant-expressive behavior and that participants had a stronger inclination to interpret a specific behavior as aggressive when it was shown by a member of the outgroup, rather than a member of the ingroup.
People often attribute success to themselves and failure to others. Past research indicates that this tendency toward self-serving attributions is pronounced among individuals high in trait narcissism. The aim of this registered report was to re-visit the link between narcissism and self-serving attributions by studying attributions in a group context and by distinguishing between two major dimensions of grandiose narcissism, admiration, and rivalry. We conducted a group study, (N = 422 participants nested in 54 groups), in which participants of each group were randomly assigned to one of two teams which then engaged in an intergroup competition. In line with our hypotheses, admiration predicted the tendency to take personal credit for success. Contrary to our hypotheses, rivalry did not uniquely predict the tendency to blame others for failure. Instead, admiration uniquely predicted the tendency to attribute negative team outcomes to unfairness of the competing outgroup. Explorative analyses further revealed that both admiration and rivalry were associated with the tendency to attribute negative, rather than positive, team outcomes to chance. Taken together, the findings indicate that narcissism goes along with an increased propensity for self-serving attributions in competitive intergroup settings and that this tendency is mainly driven by the admiration dimension.
People care about different domains of life (e.g., their health, social life, work) to varying degrees. It thus seems plausible that how satisfied they are with those domains matters for their general life satisfaction to varying degrees. This idea has been investigated in the importance-weighting literature with at best mixed results, but variations of it can be found across different fields of psychology and include claims that values, personality, and age moderate the extent to which different life domains affect life satisfaction. In this study, we investigated the effects of satisfaction with 14 different life domains on general life satisfaction in a study of 439 individuals who provided up to 15 diary entries, resulting in a total of 6,071 observations. All domains had positive effects on average, with the largest effects for satisfaction with leisure time usage (b = 0.19, bstd = 0.25 relative to the within-person variability) and relationship satisfaction (b = 0.16, bstd = 0.17). Beyond these averages, there was robust interindividual variability; the standard deviation of the individual-level effects was of a similar magnitude as the average effect (and sometimes even larger). But when exploring correlations between these individual-level effects with third variables (e.g., self-reported importance of the respective domain, gender and age, Big Five personality traits), no convincing overall patterns arose. This may at least in part result from the high uncertainty with which individual-level effects were estimated, with reliabilities of ~.30, and the resulting low statistical power.
How does personality change when people get older? Numerous studies have investigated this question, overall supporting the idea of so-called personality maturation. However, heterogeneous findings have left open questions, such as whether maturation continues in old age and how large the effects are. We suggest that the heterogeneity is partly rooted in methodological issues. First, studies may have failed to recover age effects, as they did not stringently separate within-person changes from confounding between-person differences. Second, items supposedly belonging to the same trait may show different individual trajectories, thus rendering results sensitive to the specific set of items used. We analyzed panel data from Australia (N = 15,268; Study 1), Germany (N = 22,833; Study 2), and the Netherlands (N = 10,163; Study 3) to investigate age trends in the Big Five on the levels of both scores and items. We applied a fixed effects approach that incorporates only within-person changes over time. Developmental trends in the Big Five scores were generally moderate to large and broadly confirmed personality maturation at younger ages. At older ages, maturation consistently continued for Neuroticism, whereas we found mixed evidence for such changes in Conscientiousness and Agreeableness. Furthermore, in each study, individual items showed age trends that diverged from the rest of the corresponding trait; and these differential patterns could be partly replicated across the three studies. Our results highlight the importance of items in the study of personality development and provide an explanation for previously unaccounted for variability in age trends.
The self-insight motive (SIM; also known under the label self-assessment motive) describes the dispositional tendency to strive for accurate self-knowledge. The current research includes five multimethodological studies (total N = 3667) that comprehensively investigated the SIM’s nomological network, its antecedents, and cognitive-behavioral consequences, comprising longitudinal, round-robin, and population-representative data. Among the personality correlates of the SIM were curiosity, the intimacy and self-improvement motives, private self-consciousness, narcissistic admiration, and openness to experience. Further, the SIM was more pronounced among younger and highly educated people. A key environmental antecedent of the SIM was the instability of life circumstances, in the sense that the motive became stronger after life circumstances had changed. Concerning the cognitive-behavioral consequences, the results suggest that the SIM fosters feedback-seeking behavior. Nevertheless, the motive was not linked to more accurate self-perception across three studies. We discuss several reasons for this unexpected finding.
The personality trait neuroticism is tightly linked to mental health, and neurotic people experience stronger negative emotions in everyday life. But, do their negative emotions also show greater fluctuation? This commonsensical notion was recently questioned by [Kalokerinos et al. Proc Natl Acad Sci USA 112, 15838-15843 (2020)], who suggested that the associations found in previous studies were spurious. Less neurotic people often report very low levels of negative emotion, which is usually measured with bounded rating scales. Therefore, they often pick the lowest possible response option, which severely constrains the amount of emotional variability that can be observed in principle. Applying a multistep statistical procedure that is supposed to correct for this dependency, [Kalokerinos et al. Proc Natl Acad Sci USA 112, 15838-15843 (2020)] no longer found an association between neuroticism and emotional variability. However, like other common approaches for controlling for undesirable effects due to bounded scales, this method is opaque with respect to the assumed mechanism of data generation and might not result in a successful correction. We thus suggest an alternative approach that a) takes into account that emotional states outside of the scale bounds can occur and b) models associations between neuroticism and both the mean and variability of emotion in a single step with the help of Bayesian censored location-scale models. Simulations supported this model over alternative approaches. We analyzed 13 longitudinal datasets (2,518 individuals and 11,170 measurements in total) and found clear evidence that more neurotic people experience greater variability in negative emotion.
In psychology, causal inference—both the transport from lab estimates to the real world and estimation on the basis of observational data—is often pursued in a casual manner. Underlying assumptions remain unarticulated; potential pitfalls are compiled in post-hoc lists of flaws. The field should move on to coherent frameworks of causal inference and generalizability that have been developed elsewhere.
Several studies have suggested that the rank-order stability of personality increases until midlife and declines later in old age. However, this inverted U-shaped pattern has not consistently emerged in previous research; in particular, a recent investigation implementing several methodological advances failed to support it. To resolve the matter, we analyzed data from two representative panel studies and investigated how certain methodological decisions affect conclusions regarding the age trajectories of stability. The data came from Australia (N = 15,465; Study 1) and Germany (N = 21,777; Study 2), and each study included four waves of personality assessment. We investigated the life span development of the rank-order stability of the Big Five for 4-, 8-, and 12-year intervals. Whereas Study 1 provided strong evidence for an inverted U-shape with rank-order stability declining past age 50, Study 2 provided more mixed results that nonetheless generally supported the inverted U-shape. This developmental trend held for single personality traits as well as for the overall pattern across traits; and it held for all three retest intervals—both descriptively and in formal tests. Additionally, we found evidence that health-related changes accounted for the decline in rank-order stability in older age. This suggests that if analyses are implicitly conditioned on health (e.g., by excluding participants with missing data on later waves), the decline in stability in old age will be underestimated or even missed. Our results provide further evidence for the inverted U-shaped age pattern in personality stability development but also extend knowledge about the underlying processes.
Science is often perceived to be a self-correcting enterprise. In principle, the assessment of scientific claims is supposed to proceed in a cumulative fashion, with the reigning theories of the day progressively approximating truth more accurately over time. In practice, however, cumulative self-correction tends to proceed less efficiently than one might naively suppose. Far from evaluating new evidence dispassionately and infallibly, individual scientists often cling stubbornly to prior findings. Here we explore the dynamics of scientific self-correction at an individual rather than collective level. In 13 written statements, researchers from diverse branches of psychology share why and how they have lost confidence in one of their own published findings. We qualitatively characterize these disclosures and explore their implications. A cross-disciplinary survey suggests that such loss-of-confidence sentiments are surprisingly common among members of the broader scientific population yet rarely become part of the public record. We argue that removing barriers to self-correction at the individual level is imperative if the scientific community as a whole is to achieve the ideal of efficient self-correction.
This study compared the impacts of actual individual task competence, speaking time and physical expressiveness as indicators of verbal and nonverbal communication behavior, and likability on performance evaluations in a group task. 164 participants who were assigned to 41 groups first solved a problem individually and later solved it as a team. After the group interaction, participants’ performance was evaluated by both their team members and qualified external observers. We found that these performance evaluations were significantly affected not only by task competence but even more by speaking time and nonverbal physical expressiveness. Likability also explained additional variance in performance evaluations. The implications of these findings are discussed for both the people being evaluated and the people doing the evaluating.
Algorithmic automatic item generation can be used to obtain large quantities of cognitive items in the domains of knowledge and aptitude testing. However, conventional item models used by template-based automatic item generation techniques are not ideal for the creation of items for non-cognitive constructs. Progress in this area has been made recently by employing long short-term memory recurrent neural networks to produce word sequences that syntactically resemble items typically found in personality questionnaires. To date, such items have been produced unconditionally, without the possibility of selectively targeting personality domains. In this article, we offer a brief synopsis on past developments in natural language processing and explain why the automatic generation of construct-specific items has become attainable only due to recent technological progress. We propose that pre-trained causal transformer models can be fine-tuned to achieve this task using implicit parameterization in conjunction with conditional generation. We demonstrate this method in a tutorial-like fashion and finally compare aspects of validity in human- and machine-authored items using empirical data. Our study finds that approximately two-thirds of the automatically generated items show good psychometric properties (factor loadings above .40) and that one-third even have properties equivalent to established and highly curated human-authored items. Our work thus demonstrates the practical use of deep neural networks for non-cognitive automatic item generation.
Cote et al. (1) provided evidence that economic inequality moderates the effect of income on generosity. In their study, individuals with higher household income were less generous in a dictator game than poorer individuals only if they resided in a US state with comparatively large economic inequality. We questioned this finding because we did not find any evidence for the postulated moderation effect of economic inequality across three studies (ref. 2; for similar replication failures see ref. 3). However, our studies were conceptual rather than direct replications as we used different measures of generosity (charitable donations, behavior in a trust game, and volunteering) and also included non-US … [↵][1]1To whom correspondence may be addressed. Email: schmukle{at}uni-leipzig.de. [1]: #xref-corresp-1-1
The current research dealt with the stereotype that only children are more narcissistic than people with siblings. We first investigated the prevalence of this stereotype. In an online study (Study 1, N = 556), laypeople rated a typical only child and a typical person with siblings on narcissistic admiration and narcissistic rivalry, the two subdimensions of the Narcissistic Admiration and Rivalry Questionnaire. They ascribed both higher admiration and higher rivalry to the only child. We then tested the accuracy of this stereotype by analyzing data from a large and representative panel study (Study 2, N = 1,810). The scores of only children on the two narcissism dimensions did not exceed those of people with siblings, and this result held when major potentially confounding covariates were controlled for. Taken together, the results indicate that the stereotype that only children are narcissistic is prevalent but inaccurate.
Narcissists are assumed to lack the motivation and ability to share and understand the mental states of others. Prior empirical research, however, has yielded inconclusive findings and has differed with respect to the specific aspects of narcissism and socioemotional cognition that have been examined. Here, we propose a differentiated facet approach that can be applied across research traditions and that distinguishes between facets of narcissism (agentic vs. antagonistic) on the one hand, and facets of socioemotional cognition ability (SECA; self-perceived vs. actual) on the other. Using five nonclinical samples in two studies (total N = 602), we investigated the effect of facets of grandiose narcissism on aspects of socioemotional cognition across measures of affective and cognitive empathy, Theory of Mind, and emotional intelligence, while also controlling for general reasoning ability. Across both studies, agentic facets of narcissism were found to be positively related to perceived SECA, whereas antagonistic facets of narcissism were found to be negatively related to perceived SECA. However, both narcissism facets were negatively related to actual SECA. Exploratory condition-based regression analyses further showed that agentic narcissists had a higher directed discrepancy between perceived and actual SECA: They self-enhanced their socio-emotional capacities. Implications of these results for the multifaceted theoretical understanding of the narcissism-SECA link are discussed.
A landmark study published in PNAS [Côté S, House J, Willer R (2015) Proc Natl Acad Sci USA 112:15838-15843] showed that higher income individuals are less generous than poorer individuals only if they reside in a US state with comparatively large economic inequality. This finding might serve to reconcile inconsistent findings on the effect of social class on generosity by highlighting the moderating role of economic inequality. On the basis of the importance of replicating a major finding before readily accepting it as evidence, we analyzed the effect of the interaction between income and inequality on generosity in three large representative datasets. We analyzed the donating behavior of 27,714 US households (study 1), the generosity of 1,334 German individuals in an economic game (study 2), and volunteering to participate in charitable activities in 30,985 participants from 30 countries (study 3). We found no evidence for the postulated moderation effect in any study. This result is especially remarkable because (i) our samples were very large, leading to high power to detect effects that exist, and (ii) the cross-country analysis employed in study 3 led to much greater variability in economic inequality. These findings indicate that the moderation effect might be rather specific and cannot be easily generalized. Consequently, economic inequality might not be a plausible explanation for the heterogeneous results on the effect of social class on prosociality.