The causal impact of charisma in high-stakes economic settings is difficult to establish. We showcase a novel approach by exploiting a dataset using close-call U.S. gubernatorial elections. We exploit the randomness provided by such elections to estimate the causal effect of charisma on state-level economic variables. We account for fixed effects (constant terms) at the state and year level, and controlling for important observables, demonstrate an economically meaningful effect for charisma on state output. For example, ceteris paribus, a winner who is 10% more charismatic than the mean of governors, against an opponent at the average, increases state-level GDP by 0.78% each year of a four-year term. Potential mechanisms include state spending increases (i.e. the fiscal multiplier) and new business creation.
Are there benefits of mindfulness in the workplace? The answer is unclear. In this Point article, we argue that practice may be running ahead of robust scientific evidence. We highlight six systematic design and estimation challenges studies in the fields of management and applied psychology face: (1) Current survey studies are ill suited to inform policy because they do not employ designs that allow identifying causal effects. (2) The procedures used to explore how mindfulness may affect outcomes (i.e., mediating mechanisms) are inadequate for theory testing. (3) The majority of studies rely on self-report data instead of testing the effect of mindfulness on objective outcomes. (4) Intervention studies, typically considered the gold standard for causal evidence, often rely on self-selected employee samples, and (5) these studies also frequently use weak counterfactuals. (6) Finally, mindfulness reviews tend to be based on primary studies with clear methodological limitations, and methodological critiques are largely not reflected in research practice. We provide concrete recommendations for each of these issues and, where possible, highlight positive examples from the literature. We hope that these recommendations will inform future research and ultimately help organizations make better-grounded choices among competing interventions to foster employee health and well-being.
In this commentary on Liden, Wang, and Wang (2025), we identify three recurring problems in leadership style research that they—and much of the field—did not fully address. These omissions matter because they affect which conclusions about leadership styles can be drawn. First, many leadership style constructs conflate leader behaviors with followers’ evaluations, leading to causal indeterminacy—a conceptual problem for which methodological fixes do not suffice. Second, endogeneity frequently limits the causal interpretation of relationships between leadership styles and outcomes. Third, despite these limitations, the use of causal language is common, possibly misguiding practice. Sharing with Liden et al. the aim of advancing the field, we address these challenges and outline an agenda that seeks to revitalize leadership style research. We propose to move from single, conflated leadership style constructs to multi-construct theories that model relationships among behaviors, evaluations, and context, thereby improving causal explanation and practical relevance.
Obtaining valid ratings of verbal and nonverbal behavior to determine if they cause an outcome is difficult unless human coders are blinded. Blinding is achievable for verbal behaviors (e.g., analyzing transcripts). However, for nonverbal behaviors, coders typically observe the target, risking stereotype activation, which may cause bias in ratings and in determining whether it causes performance. In Study 1, we assess the extent of this potential bias by reviewing the predominant coding methods in a sample of current articles from top management and social/applied psychology journals. From the n = 16 coded articles, only two studies utilized valid nonverbal behavior ratings. We then evaluated the extent of bias in human-coded nonverbal behaviors in Study 2 using TED talks (n = 768) to code for charisma signaling. We relied on multiple data sources, including computer and human ratings (n = 5,385 raters). Human-coded—and not computer-coded—nonverbal charismatic signaling was influenced by target attractiveness. We also show that, when coded by computers, verbal charismatic signaling has a positive effect on outcome measures (e.g., social influence, as measured by TED views). Computer-coded nonverbal charismatic signaling, however, was unrelated to outcomes. Interestingly, we demonstrate bias in human-coded nonverbal charismatic behavior by showing that it correlates with outcomes when used in its uncorrected form; however, this result is spurious because, when using instrumental-variable regression to account for omitted variables (i.e., driven by stereotypes), the result becomes null. We discuss the value of computer coding as a better alternative to human coding.
Scholars have investigated the emergence of charismatic leaders in times of crisis. However, results from this research are usually descriptive, suffer from endogeneity bias, or rely on inappropriate causal modeling. Building on exogenous events, we explore the causal effect of crises on charismatic rhetoric and approval ratings of political leaders using regression discontinuity designs. In a reanalysis of Bligh et al. (2004), we find that the rhetoric of President George W. Bush changed after 9/11 to include more references to charismatic themes. We replicate these results using President Francois Hollande reactions to terrorist attacks in 2015 and 2016 (i.e., Charlie Hebdo, Paris, and Nice attacks). Across both studies, we find similar evidence for an upward shift in charismatic rhetoric and approval ratings at the time of crisis. Our findings contribute to the literature on charisma and crisis by showing that the emergence of charisma is not only a follower attributional process but that veritable behavior of leaders can change. Our manuscript also pedagogically re-introduces the regression discontinuity design, a quasi-experimental procedure largely unused in applied leadership and management research.
We argue and show empirically that constructs and measures of positive leadership styles, such as authentic, ethical, and servant leadership, are not veridical representations of leadership behaviors. Instead, these styles conflate behaviors with subjective evaluations of leaders. Labelling behaviors as, for example, “ethical” means evaluating leadership behaviors on positively valenced terms rather than describing these behaviors. Across four experiments, we show that positive leadership styles are outcomes that depend on non-behavioral, evaluative factors, such as information about a leader’s previous success or value alignment between leaders and followers. More importantly, the measures of these leadership styles create causal illusions by spuriously predicting objective outcomes, even when leader behaviors and other leader-specific factors are kept constant. Furthermore, these measures have predictive properties similar to those of a purely evaluative measure of leadership. In conclusion, our studies cast serious doubts on previous research claiming that positive leadership styles cause positive outcomes. Moreover, positive leadership style research is not only wrong but also practically futile because its constructs and measures are amalgams that do not isolate concrete and learnable behaviors. We call for a radical reorientation of leadership style research and sketch out options for more solid future research.
Latent moderated structural equation (LMS) is one of the most common techniques for estimating interaction effects involving latent variables (i.e., XWITH command in Mplus). However, empirical applications of LMS often overlook that this estimation technique assumes normally distributed variables and that violations of this assumption may lead to seriously biased parameter estimates. Against this backdrop, we study the robustness of LMS to different shapes and sources of nonnormality and examine whether various statistical tests can help researchers detect such distributional misspecifications. In four simulations, we show that LMS can be severely biased when the latent predictors or the structural disturbances are nonnormal. On the contrary, LMS is unaffected by nonnormality originating from measurement errors. As a result, testing for the multivariate normality of observed indicators of the latent predictors can lead to erroneous conclusions, flagging distributional misspecifications in perfectly unbiased LMS results and failing to reject seriously biased results. To solve this issue, we introduce a novel Hausman-type specification test to assess the distributional assumptions of LMS and demonstrate its performance. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
A key assumption in modern conceptualizations of charisma is that it is a costly signal. It thus should be easier for intelligent individuals to produce this signal: it requires one to be creative, communicate in symbolic ways, have the needed expertise, and be consistent in one's values and actions. At this time, it is unclear whether this assumption holds. Using data from an incentivized laboratory experiment (n = 1,998 general population) and two field settings (n = 134 public service leaders and n = 41 U.S. presidents), we show that individuals's charisma signaling scores strongly correlate with their scores on intelligence. A change of a standard deviation in intelligence was associated with changes in charisma signaling of 7.89% (Study 1), 11.01 % (Study 2), as well as 5.70 %, 6.80 %, and 12.23 % (Study 3), respectively. In addition, Studies 1 and 2 showed that scores on personality dimensions-whether the big five or the big six-do not correlate with charisma signaling. Our results lay the foundations for explaining a mechanism for why charisma signaling is a potent motivational tool and thus have important theoretical and policy implications.
To screen for careless responding, researchers have a choice between several direct measures (i.e., bogus items, requiring the respondent to choose a specific answer) and indirect measures (i.e., unobtrusive post hoc indices). Given the dearth of research in the area, we examined how well direct and indirect indices perform relative to each other. In five experimental studies, we investigated whether the detection rates of the measures are affected by contextual factors: severity of the careless response pattern, type of item keying, and type of item presentation. We fully controlled the information environment by experimentally inducing careless response sets under a variety of contextual conditions. In Studies 1 and 2, participants rated the personality of an actor that presented himself in a 5-min-long videotaped speech. In Studies 3, 4, and 5, participants had to rate their own personality across two measurements. With the exception of maximum longstring, intra-individual response variability, and individual contribution to model misfit, all examined indirect indices performed better than chance in most of the examined conditions. Moreover, indirect indices had detection rates as good as and, in many cases, better than the detection rates of direct measures. We therefore encourage researchers to use indirect indices, especially within-person consistency indices, instead of direct measures.
Using field and laboratory data, we show that leader charisma can affect COVID-related mitigating behaviors. We coded a panel of U.S. governor speeches for charisma signaling using a deep neural network algorithm. The model explains variation in stay-at-home behavior of citizens based on their smart phone data movements, showing a robust effect of charisma signaling: stay-at-home behavior increased irrespective of state-level citizen political ideology or governor party allegiance. Republican governors with a particularly high charisma signaling score impacted the outcome more relative to Democratic governors in comparable conditions. Our results also suggest that one standard deviation higher charisma signaling in governor speeches could potentially have saved 5,350 lives during the study period (02/28/2020–05/14/2020). Next, in an incentivized laboratory experiment we found that politically conservative individuals are particularly prone to believe that their co-citizens will follow governor appeals to distance or stay at home when exposed to a speech that is high in charisma; these beliefs in turn drive their preference to engage in those behaviors. These results suggest that political leaders should consider additional “soft-power” levers like charisma—which can be learned—to complement policy interventions for pandemics or other public heath crises, especially with certain populations who may need a “nudge.”
Does a greater representation of women in top management teams (TMTs) contribute to higher firm performance? Although several studies have investigated this question, they have failed to sufficiently account for endogeneity. We address the endogeneity problem by using an instrumental variable (IV) design to estimate the causal effect of women's representation in TMTs on firm performance. We use a shift-share or Bartik-type instrument, which is well-established in economics but has received little attention in management and leadership research. We analyze the effect of TMT gender diversity on four types of firm performance: profitability, market-based performance, liquidity, and growth. Our sample is based on firms in the S&P 1,500, which we observe over 24 years (1997–2020). Our findings indicate that TMT gender diversity positively affects the profitability, liquidity, and growth of firms but does not impact market-based performance. We also analyze whether the effect of TMT gender diversity was stronger during two economic crises, namely the 2008/2009 financial crisis and the COVID-19 pandemic, but our instrumental variable analysis provides no evidence for such an interaction effect. Our results are robust to multiple alternative specifications. This study contributes to research on strategic leadership, specifically regarding the effect of women leaders, as well as the crisis leadership literature.
Purpose - This study aims to provide a response to the commentary by Yuan on the paper "Marketing or Methodology" in this issue of EJM. Design/methodology/approach - Conceptual argument and statistical discussion. Findings - The authors find that some of Yuan's arguments are incorrect, or unclear. Further, rather than contradicting the authors' conclusions, the material provided by Yuan in his commentary actually provides additional reasons to avoid partial least squares (PLS) in marketing research. As such, Yuan's commentary is best understood as additional evidence speaking against the use of PLS in real-world research. Research limitations/implications - This rejoinder, coupled with Yuan's comment, continues to support the strong implication that researchers should avoid using PLS in marketing and related research. Practical implications - Marketing researchers should avoid using PLS in their work. Originality/value - This rejoinder supports the earlier conclusions of "Marketing or Methodology," with additional argumentation and evidence.
Purpose Over the past 20 years, partial least squares (PLS) has become a popular method in marketing research. At the same time, several methodological studies have demonstrated problems with the technique but have had little impact on its use in marketing research practice. This study aims to present some of these criticisms in a reader-friendly way for non-methodologists. Design/methodology/approach Key critiques of PLS are summarized and demonstrated using existing data sets in easily replicated ways. Recommendations are made for assessing whether PLS is a useful method for a given research problem. Findings PLS is fundamentally just a way of constructing scale scores for regression. PLS provides no clear benefits for marketing researchers and has disadvantages that are features of the original design and cannot be solved within the PLS framework itself. Unweighted sums of item scores provide a more robust way of creating scale scores. Research limitations/implications The findings strongly suggest that researchers abandon the use of PLS in typical marketing studies. Practical implications This paper provides concrete examples and techniques to practicing marketing and social science researchers regarding how to incorporate composites into their work, and how to make decisions regarding such. Originality/value This work presents a novel perspective on PLS critiques by showing how researchers can use their own data to assess whether PLS (or another composite method) can provide any advantage over simple sum scores. A composite equivalence index is introduced for this purpose.
For scientific discoveries to be valid-whether in theory or empirically-a phenomenon must be accurately described: The scientist must use appropriate counterfactuals and eliminate competing explanations. Empirical work must also use an appropriate design and method, and empirical claims made about the phe-nomenon must be correctly characterized. Moreover, valid empirical discoveries must be reliable in the sense that scientists who reexamine the data must be able to reproduce the finding or to replicate the effect from data gathered in a similar context. Only discoveries adhering to the above criteria can be scientifically informative, serve as building blocks for theory, or have policy implications. Unfortunately, as several recent surveys of the literature show, much of the published works in the management and applied psychology fields are uninforma-tive; contributing reasons include several intractable problems in the study design and analysis as well as the failure of the field to adopt open science practices. Against this backdrop, we identify common methodological mistakes made in applied work. We group these mistakes into three major categories: (a) study design and data collection (e.g., fit between hypotheses and methods, design, measurement, open science, literature reviews), (b) data analysis (e.g., data preprocessing, choice of estimators, analysis of data, issues concerning endogene-ity, and use of instrumental variables), and (c) diagnostics, inferences, and reporting. We also explain how to avoid these issues, so that published work makes for a useful contribution to the scientific record.
Leadership theories in sociology and psychology argue that effective leaders influence follower behavior not only through the design of incentives and institutions, but also through personal abilities to persuade and motivate. Although charismatic leadership has received considerable attention in the management literature, existing research has not yet established causal evidence for an effect of leader charisma on follower performance in incentivized and economically relevant situations. We report evidence from field and laboratory experiments that investigate whether a leader’s charisma—in the form of a stylistically different motivational speech—can induce individuals to undertake personally costly but socially beneficial actions. In the field experiment, we find that workers who are given a charismatic speech increase their output by about 17% relative to workers who listen to a standard speech. This effect is statistically significant and comparable in size to the positive effect of high-powered financial incentives. We then investigate the effect of charisma in a series of laboratory experiments in which subjects are exposed to motivational speeches before playing a repeated public goods game. Our results reveal that a higher number of charismatic elements in the speech can increase public good contributions by up to 19%. However, we also find that the effectiveness of charisma varies and appears to depend on the social context in which the speech is delivered. This paper was accepted by Yan Chen, behavioral economics and decision analysis.