This review article seeks to take stock of personality psychology by sketching major accomplishments from recent past regarding (a) topics studied (cognitive abilities, structural models of basic tendencies, dynamics and processes, personality changes, biology, social relationships, culture, health/well-being, pathology/disorders), (b) methodologies (new technologies, large-scale surveys), and (d) meta-scientific aspects (Open Science, syntheses, causality, nomothetics-idiographics). Recent developments in the field and prospects are put into context regarding these accomplishments via reviews of the current literature. Our survey of the field suggests that there is great richness of approaches and methods, but also that more coordination, integration, and conceptual-terminological consistency is needed.
The city of Leipzig in Germany conducts large-scale school surveys of adolescents in secondary education schools. Following the regular surveys in 2010 and 2015, the 2020 survey had to be rescheduled to 2023 due to the COVID-19 pandemic. In this latest survey wave, the gender gap in general life satisfaction has significantly grown. While in 2010 and 2015 girls were somewhat less satisfied than boys (0.26 to 0.33 SD), in 2023 this gender gap had doubled (with girls 0.57 SD less satisfied). Why? Here, we probe various explanations, aiming to provide a template for researchers who are asking reverse causal questions (“What caused this?”). First, we find that the widening of the gender gap is much more pronounced among students with a migration background. This could plausibly be explained by a shift in the composition of the underlying population, with a strong increase of Syrian students, and a relative decrease of Vietnamese ones. Second, among students without a migration background, part of the increasing gender gap could potentially be attributed to survey mode: In 2023, for the first time, the survey was conducted on tablets—and unexpectedly, girls (but not boys) reported significantly lower satisfaction when surveyed on tablet rather than on paper. Third, beyond these two patterns, we still find significantly widening gender gaps in satisfaction with leisure time activities and relationships to friends. Thus, there may be a substantive increase in the gender gap in satisfaction in those two domains that is not readily attributable to changes in population and survey mode.
Different women experience hormonal contraceptives differently, reporting side effects on their sexuality that range from negative to positive. But research on such causal effects of hormonal contraceptives on psychological outcomes struggles both to identify average causal effects and capture the high heterogeneity in women’s treatment responses. In this study, we leveraged longitudinal data to improve our ability to separate the causal effects of hormonal contraceptives from other sources of association, including observed and unobserved confounding, reverse causality, and attrition. In this programmatic registered report (programmatic registered stage 1 protocol: https://osf.io/kj3h2 ; date of in-principle acceptance: 28/09/2023), we analyzed data from up to 5,041 women (23,130 observations), who participated in PAIRFAM, a German longitudinal panel dataset consisting of 14 waves, using Bayesian multilevel regressions. To deal with confounding and probe the robustness of findings, we implemented two analysis approaches: adjusted regression analysis and inverse probability of treatment weighting approach. We found evidence for positive average treatment effects of hormonal contraceptives on sexual frequency and sexual satisfaction, but no robust evidence for effects on desired sexual frequency. Furthermore, to move beyond average treatment effects, we analyzed heterogeneity in treatment responses. We found relatively high heterogeneity in individual treatment effects on sexual frequency and sexual satisfaction. Interindividual differences were not systematically related to individual treatment effects, and those treatment effects did not predict women’s decisions about which contraceptive method to use in the long run. Our results contribute to understanding the effects of hormonal contraceptives on sexuality in a naturalistic setting, where women adapt their choice of contraceptive method to their own experiences.
Collecting a sample that represents the population of interest well constitutes a challenge across the social and behavioural sciences. Psychology in particular frequently relies on convenience samples—most notably students and, increasingly, online participants—with a tendency to either (implicitly) assume representativeness without substantive justification, or to acknowledge a lack of it only in passing. In contrast, researchers rarely engage with the actual implications for their inferences, which undermines the generalisability of psychological findings. Critically, representativeness must be defined with respect to variables relevant to the target of inference, rather than superficial demographic diversity. Here we present the Total Survey Error (TSE) framework as a methodological tool that systematically addresses the multifaceted sources of error—particularly those related to representation—that emerge throughout the research cycle. Although TSE originated in survey research, its principles are broadly applicable to any psychological study seeking inference from sample to population. We offer practical strategies for identifying, preventing, and mitigating representation errors to improve the credibility and generalizability of psychological research.
Zusammenfassung: Dieser Übersichtsartikel stellt eine Bestandsaufnahme der Differentiellen und Persönlichkeitspsychologie dar, indem er zentrale Errungenschaften der letzten Jahre skizziert bezüglich (a) Inhalten (kognitive Leistungsunterschiede, Strukturmodelle von Basistendenzen, Dynamiken und Prozesse, Persönlichkeitsveränderungen, Biologie, soziale Beziehungen, Kultur, Gesundheit/Wohlbefinden, Pathologie/Störungen), (b) Methoden (neue Technologien, groß angelegte Umfragen) und (c) meta-wissenschaftlichen Aspekten (Open Science, Synthesen, Kausalität, Nomothetik–Idiographik). Anhand eines Fokus v. a. auf die aktuellste Literatur in diesen drei Bereichen erfolgt eine Einordnung von rezenten Entwicklungen im Fach sowie ein Ausblick. Der Überblick verdeutlicht, dass einerseits ein großer Reichtum an Ansätzen und Methoden vorhanden ist, das Fach andererseits von einer stärkeren Koordination, Integration und konzeptuell-terminologischen Konsistenz profitieren würde.
The city of Leipzig in Germany conducts large-scale school surveys of adolescents in secondary schools. While in 2010 and 2015 girls were somewhat less satisfied than boys, in 2023 this gender gap had doubled. Why? When asking such a reverse causal question, answers often focus on broad narratives, such as the psychological impact of social media. Here, we illustrate how to probe alternative explanations, such as demographic changes and methodological issues. First, we rule out that the observed pattern is a simple scaling artifact. Second, we find that the widening of the gender gap is much more pronounced among students with a migration background. This could plausibly be explained by a shift in the composition of the underlying population, with a strong increase in the proportion of Syrian students, and a relative decrease of Vietnamese students. Third, part of the increasing gender gap could potentially be attributed to survey mode: In 2023, for the first time, the survey was conducted on tablets—and unexpectedly, girls (but not boys) reported significantly lower satisfaction when surveyed on tablet rather than on paper. Lastly, beyond these patterns, we still find significantly widening gender gaps in satisfaction with leisure time activities and relationships to friends. Thus, there may be a substantive increase in the gender gap in satisfaction in those two domains that is not readily attributable to changes in demographics and survey mode.
Psychological researchers usually make sense of regression models by interpreting coefficient estimates directly. This works well enough for simple linear models, but is more challenging for more complex models with, for example, categorical variables, interactions, non-linearities, and hierarchical structures. Here, we introduce an alternative approach to making sense of statistical models. The central idea is to abstract away from the mechanics of estimation, and to treat models as “counterfactual prediction machines,” which are subsequently queried to estimate quantities and conduct tests that matter substantively. This workflow is model-agnostic; it can be applied in a consistent fashion to draw causal or descriptive inference from a wide range of models. We illustrate how to implement this workflow with the marginaleffects package, which supports over 100 different classes of models in R and Python, and present two worked examples. These examples show how the workflow can be applied across designs (e.g., observational study, randomized experiment) to answer different research questions (e.g., associations, causal effects, effect heterogeneity) while facing various challenges (e.g., controlling for confounders in a flexible manner, modelling ordinal outcomes, and interpreting non-linear models).
Psychological researchers are interested in how things change over time and routinely make claims about age effects (e.g., personality maturation), cohort effects (e.g., generational differences in narcissism), and sometimes, period effects (e.g., secular trends in mental health). The age-period-cohort identification problem means that these claims are not possible based on the data alone: Any possible temporal pattern can be explained by an infinite number of combinations of age, period, and cohort effects. This concern holds regardless of the study design (it also applies to longitudinal designs covering multiple cohorts) and the number of observations available (it also applies if researchers observe the whole population). Researchers usually rely on statistical models that impose constraints to pick one specific decomposition of effects. Unfortunately, these constraints often remain opaque, resulting in a lack of scrutiny of the underlying assumptions. How can researchers reason more transparently and systematically about age, period, and cohort? Here, I summarize advances in the understanding of the precise nature of the identification problem, provide an overview of ways to move forward, and highlight one approach that is particularly transparent about assumptions: bounding analysis, a framework developed by sociologists Ethan Fosse and Christopher Winship. To illustrate this approach, I analyze how age, period, and cohort affect attitudes toward working mothers in the German General Social Survey.
Measurement invariance is often touted as a necessary statistical prerequisite for group comparisons. Typically, when there is evidence against measurement invariance, the analysis ends. Here, we introduce readers to an alternative perspective on measurement invariance that shifts the focus from statistical procedures to causality. From that angle, violations of measurement invariance imply that there are potentially interesting differences in the measurement process between the groups, which could warrant explanations in their own right. We illustrate this with hypothetical examples of substantively meaningful violations of metric, scalar, and residual invariance. At the same time, standard procedures to test for measurement invariance rest on strong causal assumptions about the data-generating process that researchers may often be unwilling to endorse in other contexts. We point out two very different ways forward. First, for researchers who want to commit to latent factor models, violations of measurement invariance can be followed up with investigations into why those violations occur, turning them from a dead end into new research questions. Second, for researchers who feel more ambivalent about latent factor models, alternatives may be considered, and group differences on sum scores and item scores may be reported anyway as interesting descriptive findings-but they should be followed up with discussions of various explanations that take into account their plausibility.
Interest in emotional variability as an interindividual difference is growing. Yet, basic features of the construct, such as its stability and reliability, are not well understood. To address this gap, we examined two longitudinal data sets, each comprising two waves of daily assessments, one with a 3-month and the other with a 16-month retest interval. Overcoming a key methodological limitation of past approaches, we used Bayesian censored location scale models as an alternative modeling approach that accounts for biases introduced by bounded rating scales. The results showed that the variability estimates from the models had reliabilities around rel = .64, which can be sufficient for group-level predictions. Additionally, the latent stability of r = .60 provides evidence for stable individual differences in emotional variability, suggesting that it is more than just a transient state. We ran exploratory analyses to further examine the influence of external events and individual life transitions on emotional variability and found that emotional variability was responsive to environmental changes.
Psychological researchers are interested in how things change over time and routinely make claims about age effects (e.g., personality maturation), cohort effects (e.g., generational differences in narcissism), and sometimes period effects (e.g., secular trends in mental health). The age-period-cohort identification problem means that these claims are not possible based on the data alone: Any possible temporal pattern can be explained by an infinite number of combinations of age, period, and cohort effects. This concern holds regardless of the study design (it also applies to longitudinal designs covering multiple cohorts) and the number of observations available (it also applies if we observe the whole population). Researchers usually rely on statistical models that impose constraints to pick one specific combination of effects. Unfortunately, these constraints often remain opaque, resulting in a lack of scrutiny of the underlying assumptions. How can we reason more transparently and systematically about age, period, and cohort? Here, I summarize advances in our understanding of the precise nature of the identification problem, provide an overview of ways to move forward, and highlight one approach that is particularly transparent about assumptions: bounding analysis, a framework developed by sociologists Ethan Fosse and Christopher Winship. To illustrate this approach, I analyze how age, period, and cohort affect attitudes toward working mothers in the German General Social Survey.
How do life events affect life satisfaction? Previous studies focused on a single event or separate analyses of several events. However, life events are often grouped non-randomly over the lifespan, occur in close succession, and are causally linked, raising the question of how to best analyze them jointly. Here, we used representative German data (SOEP; N = 40,121 individuals; n = 41,402 event occurrences) to contrast three fixed-effects model specifications: First, individual event models in which other events were ignored, which are thus prone to undercontrol bias; second, combined event models which controlled for all events, including subsequent ones, which may induce overcontrol bias; and third, our favored combined models that only controlled for preceding events. In this preferred model, the events of new partner, cohabitation, marriage, and childbirth had positive effects on life satisfaction, while separation, unemployment, and death of partner or child had negative effects. Model specification made little difference for employment- and bereavement-related events. However, for events related to romantic relationships and childbearing, small but consistent differences arose between models. Thus, when estimating effects of new partners, separation, cohabitation, marriage, and childbirth, care should be taken to include appropriate controls (and omit inappropriate ones) to minimize bias.
Recent developments in the causal-inference literature have renewed psychologists’ interest in how to improve causal conclusions based on observational data. A lot of the recent writing has focused on concerns of causal identification (under which conditions is it, in principle, possible to recover causal effects?); in this primer, we turn to causal estimation (how do researchers actually turn the data into an effect estimate?) and modern approaches to it that are commonly used in epidemiology. First, we explain how causal estimands can be defined rigorously with the help of the potential-outcomes framework, and we highlight four crucial assumptions necessary for causal inference to succeed (exchangeability, positivity, consistency, and noninterference). Next, we present three types of approaches to causal estimation and compare their strengths and weaknesses: propensity-score methods (in which the independent variable is modeled as a function of controls), g-computation methods (in which the dependent variable is modeled as a function of both controls and the independent variable), and doubly robust estimators (which combine models for both independent and dependent variables). A companion R Notebook is available at github.com/ArthurChatton/CausalCookbook. We hope that this nontechnical introduction not only helps psychologists and other social scientists expand their causal toolbox but also facilitates communication across disciplinary boundaries when it comes to causal inference, a research goal common to all fields of research.