
Abstract Elections are central to the study of politics. When studying parties’ vote shares across districts, scholars are encouraged to use compositional-outcome models in order to test their theories about what factors shape the dynamics of electoral support. Existing compositional modeling approaches incorrectly deal with scenarios in which not all parties compete in every electoral district. Because unobserved factors that affect a party’s decisions to contest a district are likely correlated with its performance in districts where the party fielded candidates, failing to account for partial contestation is likely to result in sample selection bias. Addressing sample selection in a compositional setting is challenging because the outcomes are in log-ratio form, and thus the errors often deviate from normality. To deal with these issues, we introduce a novel maximum likelihood approach which accounts for this type of sample selection and demonstrate through simulations that our method outperforms commonly used solutions, including the conventional Heckman correction. We illustrate the utility of our approach by analyzing the 2017 and 2019 U.K. parliamentary elections in English constituencies.
Regime types and transitions are central to a wide range of political phenomena. Reflecting this importance, prior research has produced a variety of regime measures. This diversity, however, comes with important challenges for applied research: selecting a measure among many options, having to define regime categories based on cut-offs, identifying regime transitions by specific magnitudes of change over a specific time window and dealing with measurement uncertainty and missing data. In this article, we introduce Unified Transitions and Stability (UNITAS), a new framework that offers a solution to these challenges. Combining information from commonly used regime indicators, this approach identifies regime types and transitions probabilistically, locates the most likely periods of regime transitions and incorporates measurement uncertainty. Through Monte Carlo simulations, we demonstrate the desirable properties and robustness of UNITAS under various scenarios. In an illustrative application, we show that stable semi-democracies are not inherently conflict-prone and that autocratization is consistently associated with higher civil war risk while democratization is not.
Recent advancements in language technology have opened new avenues in political science for automating and improving survey data analysis across diverse cultural contexts. This article examines the effectiveness of language models (LMs) in analyzing open-ended survey responses about democracy from ten countries, contrasting these modern tools with traditional survey methodologies. Utilizing a predefined coding scheme and a subset of pre-annotated survey data, it assesses the performance of fine-tuning pre-trained LMs in a multilingual setting to classify text spans. The findings suggest that LMs can capture democratic perceptions and handle data abstractions at levels comparable to human annotators. This study not only highlights the potential of LMs to transform political science research by augmenting traditional methods but also discusses the practical applications of pre-trained LMs in classifying complex survey responses, in collaboration with human annotators.
Text analysis typically focuses on content-such as sentiment or topic-but expression is also a form of effortful action. Building on this insight, I propose using simple features of open-ended tasks to study text as behavior. This approach treats expression, such as writing, as cognitively, emotionally and temporally "costly" for subjects but inexpensive for researchers. I show basic statistics like the number of characters can approximate effort and significantly improve estimation of quantities of interest, including candidate choice, the probability of turning out to vote and psychological states about which a subject may not be fully aware. Further, these methods can convert nonresponse into informative data; validate survey instruments; serve as mechanism checks; be hard for a subject to "game"; work across different languages and analogize well to real-world situations. In sum, text as behavior can help address a range of issues related to quantifying attitudes and actions.
When foundation models analyze political content, do they use demographic characteristics as shortcuts for ideological attribution? We conducted detailed experiments with GPT-4o-mini and validated key findings across GPT-4o and LLaVA, using identical, ideologically neutral campaign advertisements with systematically varied candidate demographics. All models consistently attributed more liberal ideologies to women than men. These effects exceeded real-world gender differences from a nationally representative survey. However, racial associations differed by model: strong in GPT-4o-mini (where Black candidates received substantially more liberal attributions), attenuated in GPT-4o, and insignificant in LLaVA. These demographic effects persisted across temperature settings, prompt variations, and even explicit debiasing instructions in GPT-4o-mini. Our findings reveal that visual demographic features can shape AI outputs in ways that vary across models, with implications for applications such as content classification.
Synthetic control methods are widely used for causal inference in case studies and panel data settings, often applied to model counterfactuals for proportional outcomes. However, conventional synthetic control methods are designed for univariate outcomes, leading researchers to model counterfactuals for each proportion separately. We make the case for jointly estimating synthetic controls across multiple compositional outcomes. Using the same weights for each proportion establishes a constant control comparison, improving comparability while adhering to compositional constraints on treatment effects. We illustrate the benefits of the method through a simulation and two applications to recent empirical studies. This implementation integrates naturally with a wide range of synthetic control approaches, providing interpretable estimates for compositional panel data common in political science.
Members of the majority party in Congress sometimes vote against bills that they prefer over the status quo. We estimate a model of congressional roll-call voting that allows for this kind of non-ideological protest voting. We find that protest voting has significant implications for roll-call-based estimates of ideology and other analyses that rely upon them. For example, a traditional item response theory model curiously identifies members of the Squad as relatively moderate Democrats, but our protest-voting-adjusted scores identify them as the most liberal members of Congress. We also find that previous studies may have underestimated responsiveness, the effects of ideology in elections, the utility of non-roll-call-based measures of ideology, and the increase in congressional polarization. Although the implications for most substantive applications are likely modest, our analyses suggest that future researchers can better measure legislative ideology by accounting for a small number of non-ideological votes.
One of the most robust empirical findings in political science is that in multiparty democracies cabinet ministries are distributed in rough proportion to parties' legislative seat shares, a pattern known as Gamson's Law. Yet existing research often overlooks the fact that portfolio and seat shares are compositional-mutually dependent parts of a whole. Standard methods treat them as unconstrained, risking bias, misleading uncertainty estimates, and flawed inference. Unfortunately, the most common strategy for handling compositions-the additive log-ratio (ALR) transformation combined with seemingly unrelated regression (SUR)-fails when the number and identity of compositional elements (like parties) vary across cases. We propose the isometric log-ratio (ILR) transformation-new to political science-as an alternative that both respects compositional geometry and adapts to differing compositional structures. Monte Carlo simulations show that ILR sharply outperforms standard approaches, reducing bias, improving coverage probability, and increasing statistical power. While we apply it to portfolio allocation, ILR provides a general solution for modeling compositional outcomes with other potential uses, including in electoral competition, where ALR+SUR has required strong assumptions or ad hoc adjustments. Using this improved methodology, we find that seat-portfolio proportionality is weaker overall than conventionally reported and varies substantially across governments.
We model attitude stability and constraint, using a dynamic discrete choice framework for multiple attitudes, to identify influential attitudes within attitude systems. Its value-added includes insights about different sources of (in)stability, the direction of causation between attitudes, and their relative degree of influence; capturing time-invariant individual traits with a multiple factor structure; and addressing the ordinal nature of attitudinal measures, together with heterogeneity in time intervals between interviews, across waves, and people. We examine five core political attitudes concerning how people view the political world and their role in it. Most of their variance reflects infrequently-changing individual characteristics and time-specific effects. Permanent heterogeneity plays a modest role. External efficacy is most influential concerning evaluations of the external political world, while internal efficacy is influential for views on one's role in politics. Another application examines the role of ideological and party identification on attitudes toward government spending and immigration. The attitudes form a weakly constrained attitude system. Party identification is the most influential, through spillovers to ideological identification. Party and ideological identifications are stable, time-invariant traits explaining most of their variance, with transitory shocks that hint at measurement error and/or expressive responding. Issue attitudes are unstable, driven mainly by transitory shocks.
Many inferential tasks involve fitting models to observed data and predicting outcomes at new covariate values, requiring interpolation or extrapolation. Conventional methods select a single best-fitting model, discarding fits that were similarly plausible in-sample but would yield sharply different predictions out-of-sample. Gaussian processes (GPs) offer a principled alternative. Rather than committing to one conditional expectation function, GPs deliver a posterior distribution over outcomes at any covariate value. This posterior effectively retains the range of models consistent with the data, widening uncertainty intervals where extrapolation magnifies divergence. In this way, the GP's uncertainty estimates reflect the implications of extrapolation on our predictions, helping to tame the "dangers of extreme counterfactuals" (King and Zeng, 2006). The approach requires (i) specifying a covariance function linking outcome similarity to covariate similarity and (ii) assuming Gaussian noise around the conditional expectation. We provide an accessible introduction to GPs with emphasis on this property, along with a simple, automated procedure for hyperparameter selection implemented in the R package gpss. We illustrate the value of GPs for capturing counterfactual uncertainty in three settings: (i) treatment effect estimation with poor overlap, (ii) interrupted time series requiring extrapolation beyond pre-intervention data, and (iii) regression discontinuity designs where estimates hinge on boundary behavior.
Political science is a field rich in multimodal information sources, from televised debates to parliamentary briefings. This paper bridges a gap between computer and political science in multimodal data analysis using audio. The adoption of multimodal analyses in political science (e.g., video/audio with text-as-data approaches) has been relatively slow due to unequal distribution of computational power and skills needed. We provide solutions to challenges encountered when analyzing audio, advancing the potential for multimodal data analysis in political science. Using a dataset of all televised U.S. presidential debates from 1960 to 2020, we focus on three features encountered when analyzing audio data: low-level descriptors (LLDs), such as pitch or energy; Mel-frequency cepstral coefficients (MFCCs); and audio embeddings/encodings, like Wav2Vec. We showcase four applications: (a) forced alignment of audio text using MFCCs, time-stamping transcripts, and speaker information; (b) speech characterization using LLDs; (c) custom-made classification models with audio embeddings and MFCCs; and (d) emotional recognition models using Wav2Vec for classification of discrete emotions and their valence-arousal dominance. We provide explanations to help understand how these features can be applied for different political research questions and advice on vigilance to naive interpretation, for both experienced researchers and those who want to start working with audio.
Multilevel modeling accounts for outcome dependence across lower-level units due to unobserved group effects, while spatial modeling accounts for outcome dependence across units in the same level of analysis due to diffusion. Outcome dependence can occur simultaneously due to both spatial diffusion in the lower-level units and spatial diffusion in the unobserved group effects. For example, counties are nested within states and diffusion processes might take place at both levels of analysis. Building on recent research from the spatial econometrics and multilevel modeling literature, we propose a class of spatial hierarchical models with binary outcomes. One method accounts for spatially independent, unobserved group effects and the other method accounts for spatially dependent unobserved group effects. We propose a Bayesian approach to estimate such effects while also accounting for lower-level diffusion in the outcome, and provide software to estimate these models. Our Monte Carlo results demonstrate that failing to correctly account for diffusion and/or the nested structure of data can lead to bias in both parameter estimates and substantive effects. We apply these models to analyze the causes of civil rights protests in the United States in the 1960s.
Many political surveys rely on post-stratification, raking, or related weighting adjustments to align respondents with the target population. But when respondents differ from nonrespondents on the outcome itself (nonignorable nonresponse), these adjustments can fail, introducing bias even into basic descriptives.We provide a practical method that corrects for nonignorable nonresponse by leveraging response-propensity proxies (e.g., interviewer-coded cooperativeness) observed among respondents to extrapolate toward nonrespondents, while directly integrating observable covariates and retaining the benefits of post-stratification with known population shares. The method generalizes the variable-response-propensity (VRP) framework of Peress (2010) from binary to ordinal outcomes, which are widely used to measure trust, satisfaction, and policy attitudes. The resulting estimator is computed by maximum likelihood and implemented in a compact R routine that handles both ordinal and binary outcomes. Using the 2024 American National Election Study (ANES), we show that accounting for nonignorable nonresponse produces substantively meaningful shifts for life satisfaction (estimated latent correlation ρ≈ 0.49), while yielding negligible changes for retrospective economic evaluations (ρ≈ 0), highlighting when nonignorable nonresponse substantively affects survey estimates.
In this note, we offer a cautionary tale on the dangers of drawing inferences from low-quality online survey datasets. We reanalyze and replicate a survey experiment studying the effect of acquiescence bias on estimates of conspiratorial beliefs and political misinformation. Correcting a minor data coding error yields a puzzling result: respondents with a postgraduate education appear to be the most prone to acquiescence bias. We conduct two preregistered replication studies to better understand this finding. In our first replication, conducted using the same survey platform as the original study, we find a nearly identical set of results. But in our second replication, conducted with a larger and higher-quality survey panel, this apparent effect disappears. We conclude that the observed relationship was an artifact of inattentive and fraudulent responses in the original survey panel, and that attention checks alone do not fully resolve the problem. This demonstrates how "survey trolls" and inattentive respondents on low-quality survey platforms can generate spurious and theoretically confusing results.
This article proposes a topic modeling method that scales linearly to billions of documents. We make three core contributions: i) we present a topic modeling method, tensor latent Dirichlet allocation, that has identifiable and recoverable parameter guarantees and sample complexity guarantees for large data; ii) we show that this method is computationally and memory efficient (achieving speeds over 3 $\times $ –4 $\times $ those of prior parallelized latent Dirichlet allocation methods), and that it scales linearly to text datasets with over a billion documents; and iii) we provide an open-source, GPU-based implementation of this method. This scaling enables previously prohibitive analyses, and we perform two real-world, large-scale new studies of interest to political scientists: we provide the first thorough analysis of the evolution of the #MeToo movement through the lens of over two years of Twitter conversation and a detailed study of social media conversations about election fraud in the 2020 presidential election. Thus, this method provides social scientists with the ability to study very large corpora at scale and to answer important theoretically-relevant questions about salient issues in near real-time.
Coalition research increasingly emphasizes party-level explanations of coalition outcomes. However, this work does not account for the complex multilevel structure between parties and governments: many parties participate in multiple governments and governments often comprise multiple parties. In this paper, I show that this crisscrossing structure creates dependencies among observations both across and within governments. If ignored, these dependencies produce downward-biased uncertainty estimates that cluster-robust standard errors fail to fully correct. To address this issue, I then introduce a model that extends the Multiple Membership Multilevel Model to represent the multilevel structure of coalition government data. The model accounts for party-level dependencies across governments through party-specific effects in each coalition they join, and for dependencies within governments by representing the total party effect on a government as a weighted sum of its members’ contributions. By allowing party weights to vary with covariates describing their interrelationships, the model enables researchers to examine the interdependent nature of coalition outcomes. I validate the model through simulation and an empirical application to coalition government survival, showing that ignoring party-level dependencies can produce misleading conclusions at all levels of analysis. The model is estimated via Bayesian MCMC and implemented in the accompanying R package ‘bml’.
Human choices are often both multi-dimensional and interactive. For example, a person deciding which of two immigrants is more worthy of admission to a country might weigh their education, and the weight placed on education may depend on other factors such as their age, country of origin, and employment history. We develop a response-adaptive experimental design that summarizes the range of effects of one attribute as a function of all other attributes. Our approach changes several aspects of the experimental design based on the ex ante choice to study the heterogeneous effects of one focal attribute (i.e., education). We update treatment assignment probabilities over the course of the experiment to search for the attribute vector at which the focal attribute has the most positive and most negative effects. By summarizing the full range of effects that exist, our approach complements existing approaches to conjoint experiments that typically aggregate over heterogeneity by marginalizing. We illustrate through two online experiments and provide customizable code infrastructure via a Docker container that other researchers can use to deploy adaptive randomization in online conjoint experiments.
The design-based paradigm may be adopted in causal inference and survey sampling when we assume Rubin’s stable unit treatment value assumption (SUTVA) or impose similar frameworks. While often taken for granted, such assumptions entail strong claims about the data-generating process. We develop an alternative design-based approach: we first invoke a generalized, non-parametric model that allows for unrestricted forms of interference, such as spillover. We define an associated set of inferential targets and discuss their interpretation under SUTVA and a weaker assumption that we call the “no unmodeled revealable variation assumption” (NURVA). We then reconstruct the standard paradigm, reconsidering SUTVA at the end rather than assuming it at the beginning. Despite its similarity to SUTVA, we demonstrate the practical limitations of NURVA alone for identifying substantively interesting quantities. In so doing, we provide clarity on the nature and importance of SUTVA for applied research.
Investigators are often interested in how a treatment affects an outcome for units responding to treatment in a certain way. We may wish to know the effect among units that, for example, meaningfully implemented an intervention, passed an attention check or demonstrated some important mechanistic response. Simply conditioning on the observed value of the post-treatment variable introduces problematic biases. Further, the identification assumptions required by several existing strategies are often indefensible. We propose the treatment reactive average causal effect (TRACE), which we define as the total effect of treatment in the group that, if treated, would realize a particular value of the relevant post-treatment variable. By reasoning about the effect among the "non-reactive" group, we can identify and estimate the range of plausible values for the TRACE. We demonstrate the use of this approach with three examples: (i) learning the effect of police-perceived race on police violence during traffic stops, a case where point identification may be possible; (ii) estimating effects of a community policing intervention in Liberia, in communities that meaningfully implemented it; and (iii) studying how in-person canvassing affects support for transgender rights, among participants for whom the intervention would result in more positive feelings toward transgender people.