
Causal inference approaches often emphasize binary treatments. But in many applications, the underlying constructs are continuous. In the potential outcomes framework, a continuous treatment can take on numerous values, each corresponding to a potential outcome that may be realized. In this setting, common estimands may be intractable because of a common issue in social research, particularly research on social inequality: the exposure is highly stratified by confounders. The authors show how to avoid drawing inferences about counterfactuals where data are unlikely to exist by carefully selecting the causal estimand. The authors adopt an additive shift estimand that adds a small, fixed amount to each unit’s income. This approach is preferable to population-average dose-response curves in settings in which some treatment values rarely occur in some subgroups. The authors also show how to estimate and summarize patterns of nonlinearity and effect heterogeneity with continuous treatments. As a motivating example, the authors consider the causal effect of parental income on college attendance, a setting in which the exposure is highly stratified by confounders (e.g., parental education). This approach applies to a wide range of possible treatment conditions in sociology.
In the thriving field of network studies, there has been an emerging practice of comparing optimal modularity across social networks to evaluate the variation of network module-related substantive concepts, such as the level of consensus, polarization, or community boundary rigidity. This practice offers valuable insights and is often thoughtfully motivated, but it faces several conceptual and empirical challenges that merit careful consideration. Conceptually, the selected modularity metric may misalign with the substantive concepts researchers aim to measure. Empirically, estimated optimal modularity is highly sensitive to algorithm choice and to network characteristics unrelated to those substantive concepts, which can bias comparison results. The authors illustrate these issues with toy examples and systematic simulations and offer suggestions for more rigorous comparison practices. To show the practical significance of these lessons, the authors replicate an empirical study that examines the temporal trend of modularity scores for job mobility networks to evaluate the evolution of mobility boundary rigidity in the U.S. labor market.
Protest surveys are one of the most common methods for studying participation in social movements. As the use of digital tools has become more common in academic research, they are being applied to facilitate data collection at protest events. The authors analyze the benefits and drawbacks of these digital tools for data collection at protest events. They present the results of a field experiment that tests two increasingly common forms of data collection: surveys collected through electronic tablets and through QR (quick response) codes. The field experiment was conducted at four anti-Trump rallies on the east coast of the United States in 2025. The authors find benefits and drawbacks to each technique. The findings focus on differences in response rates, delayed refusal, and the demographics of survey respondents. The authors conclude by discussing how to apply these findings to collect the most robust sample of protest participants.
Large language models (LLMs) increasingly serve as humanlike decision-making agents in social science and applied settings. These LLM agents are typically assigned humanlike characters and placed in real-life contexts. However, how these characters and contexts shape an LLM's behavior remains underexplored. In this study the author proposes and tests methods for probing, quantifying, and modifying an LLM's internal representations in a dictator game, a classic behavioral experiment on fairness and prosocial behavior. The author extracts "vectors of variable variations" (e.g., "male" to "female") from the LLM's internal state. Manipulating these vectors during the model's inference can substantially alter how those variables relate to the model's decision making. This approach offers a principled way to study and regulate how social concepts can be encoded and engineered within transformer-based models, with implications for alignment, debiasing, and designing artificial intelligence agents for social simulations in both academic and commercial applications, strengthening sociological theory and measurement.
A growing number of social scientists are using ecological momentary assessment (EMA) to observe how social forces operate in real time. However, the validity of EMA for measuring features of daily life-what people are doing, where, and with whom-remains uncertain. A key challenge is the lack of consensus across studies about how validity in EMA methods is defined and assessed. The authors address that gap by comparing EMA data (n = 1,174) with time diary data (n = 1,113) using two population-based samples. An advantage of large samples is the ability to evaluate the magnitude of bias rather than relying solely on p-values, as is common in small-sample studies. The authors find that both methods yield similar estimates of moments captured at home and in the workplace, supporting their validity in those contexts. However, EMA tends to overestimate moments spent alone compared with time diaries, likely because of moment selection bias. Moreover, large discrepancies in estimates for eating and drinking and household chores suggest that relying on primary activity reports can introduce significant bias for multitasked activities. Comparing these methods provides insight into their relative strengths and limitations, helping researchers assess the validity, potential biases, and interpretive implications of each across key domains of daily life.
Automated text analysis is becoming extremely popular and image analysis is gaining interest. However, multimodal analysis that combines both text and image information remains rare, even though many real-world data are intrinsically multimodal, such as social media posts. The authors compare three practical workflows for clustering text–image pairs: (1) label-level combination, which clusters text and image separately and combines the resulting labels; (2) vector-level combination, which clusters concatenated embeddings extracted from each modality; and (3) joint embedding, which clusters unified representations from multimodal embedding models such as Contrastive Language-Image Pre-training. The authors also introduce a set of reusable evaluation tools to help researchers compare, validate, and benchmark multimodal clustering workflows: adjusted mutual information to assess text–image alignment, the S_DbW index to evaluate number of clusters, and within-cluster consistency to validate interpretability. The authors validate the methods on a Chinese protest data set from social media with 336,921 text–image pairs and test robustness and scope conditions using a smaller U.S. news data set on gun violence with 1,297 news headlines. The authors find that when text and image provide distinct, nonoverlapping information, the second and third methods outperform the first. This study serves as a bridge between the text-as-data and image-as-data communities.
Racism, discriminatory practices, institutional bias, and systematic exclusion can take lasting physical form in the built environment. There has been growing attention to the long-term consequences of housing policies and practices on social and economic outcomes, including residential segregation, but comparatively less attention to other aspects of the built environment, such as road networks and spatial (dis)connectivity. In this article, the authors introduce a novel method that constructs counterfactual road networks by identifying missing road segments that would be expected to exist in a city's road network, given the surrounding infrastructure. The authors demonstrate the empirical application of the method by analyzing differences in racial composition and residential segregation for the observed and counterfactual road networks of five U.S. cities. The authors find that unexpected disconnectivity in a city's road network is associated with greater differences in the racial composition of nearby areas and higher levels of segregation at the local and city levels. The present findings suggest that road networks warrant more attention as a factor that may contribute to the persistence of segregation.
Inducing semantic relations in word vector spaces and analyzing how other words or entire documents discursively engage these relations is a popular form of cultural analysis. We propose a reliability metric that is easily interpretable and agnostic to the type of relation. The metric, which we call the anchor reliability coefficient (or relco), is found by creating a synthetic document-term matrix of simulated documents that sequentially shift more of their probability mass from relation-relevant anchor terms to randomly drawn words, and then regressing the documents' similarity to an induced relation by the inverse randomness rank of the documents. We validate the metric at the word-level with both expert- and crowd-sourced dictionaries and at the document-level with expert-annotated social media posts.
The opportunities for understanding how treatment effects vary across different segments of the population have led to a rise in the use of quantile regressions for identifying unconditional quantile treatment effects (QTEs). However, existing quantile regression models fall into two categories: those that are unsuitable for identifying unconditional QTEs, and those that often struggle with the complex data structures common in sociology and other social sciences. Therefore, existing methods to identify unconditional QTEs are incomplete: the propensity score framework of Firpo (2007) allows for only a binary treatment variable, and the generalized quantile regression model of Powell (2020) faces difficulties with large data sets and high-dimensional fixed effects. This paper introduces a two-step approach to estimating unconditional QTEs, which is easy to use and aligns with the needs of sociologists. First, the treatment variable is decomposed into a systematic and random part, and then, the random variation in the treatment status is used in a bivariate quantile regression model. Through a series of simulations and three empirical applications, we demonstrate that the RQR approach provides unbiased estimates of unconditional QTEs. Moreover, the RQR approach offers greater flexibility and enhances computational speed compared to existing models, and it can easily handle high-dimensional fixed effects. In sum, the RQR approach fills a pressing void in quantitative research methodology, offering a much-needed tool for studying treatment effect heterogeneity.
Mobility scholars are increasingly turning to computational methods to analyze mobility tables. Most of these approaches start with the detection of mobility clusters, namely, sets of occupations within which the flow of workers is dense and across which the flow is sparse. Yet clustering is not the only way worker flows can be structured. This article shows how a degree-corrected stochastic blockmodel can detect patterns of mobility that are more general than clustering and consistent with the homogeneity criterion laid out by Goodman as well as the internal homogeneity thesis proposed by Breiger. Because of the intractable marginal likelihood of the model, parameters are estimated using a variational expectation maximization algorithm. Simulation results suggest the estimation algorithm successfully recovers (conditionally) stochastically equivalent mobility classes. Analysis of two real-world examples shows the model is able to detect meaningful mobility patterns, even in situations in which commonly used community detection algorithms fail.
This article addresses two prominent theses in social stratification research, the great equalizer thesis and Mare’s school transition thesis. Both theses describe the role of an intermediate educational transition in the association between socioeconomic status and an outcome variable. However, the descriptive regularities of the two theses may be driven by differential selection into the intermediate transition, which prevents the two theses from having substantive interpretations. The authors propose a set of novel counterfactual slope estimands, which capture these theses under hypothetical interventions that would eliminate the differential selection. The authors thereby construct selection-free tests for these theses. The authors are the first to explicitly provide nonparametric causal estimands for the great equalizer thesis and the school transition thesis, which enable them to conduct more principled analysis. The authors are also the first to develop flexible, efficient, and robust estimators for the two theses on the basis of efficient influence functions. The authors apply this framework to a nationally representative data set in the United States and reevaluate the two theses. Findings from the selection-free tests suggest that the descriptive regularities are misleading for the substantive interpretation of the great equalizer thesis, but not for the school transition thesis. Additionally, the counterfactual slopes provide a new framework for evaluating the inequality impacts of policy interventions.
Research, advocacy, and archival projects related to incarceration often lack knowledge about the ongoing conditions of carceral facilities, and myriad challenges prevent stakeholders from successfully conducting outreach with incarcerated people. Using a case study of archival materials contributed to PrisonPandemic, wherein letters and phone calls were invited and accepted from people incarcerated in California during the coronavirus disease 2019 pandemic, the authors demonstrate a novel outreach method using web scraping and postal mailing. They analyze PrisonPandemic’s outreach, rejection rates, and response rates to 20 county jail systems from October to December 2021. The authors find that scraping and mailing help overcome typical challenges associated with conducting outreach with this population. Scraping and mailing can create a comprehensive sampling frame to achieve response rates comparable with that of traditional outreach methods to nonincarcerated populations. The authors discuss applications beyond pandemic periods and incarcerated populations, as well as the benefits, challenges, and ethical implications of using scraping and mailing.
Aggregated relational data (ARD), derived from questions of the form "How many people do you know who [belong to subpopulation X]?" are widely used to estimate the size and composition of social networks, often adopting the network scale-up method (NSUM). However, their measurement properties are insufficiently studied. The authors address this gap by assessing (1) the test-retest reliability of a large set of ARD questions and NSUM-estimated network sizes and (2) the convergent validity of these network size estimates. This mixed-methods study involved a heterogeneous quota sample of 50 citizens in Barcelona, Spain, in 2023. Respondents were interviewed twice over a 10- to 15-day period, answering a series of ARD questions on each occasion. Qualitative debriefing provided valuable insights into their response behaviors. Our findings indicate that NSUM accurately ranked respondents' network sizes but did not estimate their values consistently across measurements. Respondents gave lower answers in the second interview than in the first. In particular, the network sizes of people with large networks ("hubs") fluctuated significantly. NSUM-estimated network size moderately correlated with estimates from the summation method and Facebook friend counts. The authors discuss the implications and provide practical recommendations for ARD item selection and the use of NSUM instruments.
Amazon's Mechanical Turk (MTurk) and Prolific are popular online platforms for connecting academic researchers with respondents. A broad literature has sought to assess the extent to which these respondents are representative of the U.S. population in terms of their demographic background, yet no work has assessed the representativeness of their daily lives. The authors provide this analysis by collecting time diaries from 136 MTurk and 156 Prolific respondents, which they compare with diary responses from 468 contemporaneous responses to the American Time Use Survey (ATUS). Responses from MTurk and Prolific respondents include several notable differences relative to ATUS responses, including doing less housework and care work, spending less time traveling, spending more time at home, and spending more time alone. In general, MTurk respondents worked more than ATUS respondents, and Prolific respondents spent more time in leisure. These differences persist even after adjusting for demographic differences. The present findings highlight time use as a potential major source of differences across samples that go beyond demographic differences. Thus, scholars interested in these samples should consider how time use may moderate processes of interest.
Validation is at the heart of methodological discussions about topic modeling. The authors argue that validation based on human reading hinges on distinctive words and readers’ labeling of a topic, and it overlooks the probability of conflicting results from semantically similar models, such as regressions or other methods. This runs counter to the presumption that topic modeling can reveal features of documents that have some measurable association with social aspects outside the text. The authors develop a similar topic identifying procedure to verify that semantically similar solutions yield similar results in further analysis. The authors argue that future validations of topic modeling must consider such procedures.
The authors introduce BERTNN (Bidirectional Encoder Representations from Transformers Neural Network), a novel methodology designed to expand affective lexicons, a critical component in sociological research. BERTNN estimates the affective meanings and their distribution for new concepts, bypassing the need for extensive surveys by leveraging their contextual usage in language. The cornerstone of BERTNN is the use of nuanced word embeddings from Bidirectional Encoder Representations from Transformers. BERTNN uniquely encodes words within the framework of synthesized social event sentences, preserving their meaning across actor-behavior-object positions. The model is fine-tuned on the basis of the implied sentiment changes, providing a more refined estimation of affective meanings. BERTNN outperforms previous approaches, setting a new standard in deriving multidimensional affective meanings for novel concepts. It efficiently replicates sentiment ratings that traditionally require extensive survey hours, demonstrating the power of automated modeling in sociological research. The expanded affective lexicons that can be produced with BERTNN cater to shifting cultural meanings and diverse subgroups, demonstrating the potential of computational linguistics to enrich the measurement tools in sociological research. This article underscores the novelty and significance of BERTNN in the broader context of sociological methodology.
In addition to overall dispersion, the distributional shape of economic status has attracted growing attention in the inequality literature. Economic polarization is a specific form of distributional change, characterized by a shrinking middle of the distribution and a growing top and bottom, with potentially important and unique social consequences. Building on relative distribution methods and drawing from the literature on job polarization, the authors develop an approach for analyzing economic polarization at the individual level. The method has three useful features. First, it offers intuitive and flexible measurement of economic polarization both between and within categories. Second, it helps disentangle two potential sources of economic polarization: compositional change, which involves changes to the allocation of workers across categories, and relative economic status change, which involves changes to the allocation of economic rewards between individuals. Third, it enables researchers to uncover and examine potential heterogeneity in economic polarization, for example, across occupations, geographic units, demographic and educational groups, and firms. The authors demonstrate the utility of this approach through two empirical applications: (1) an analysis of trends in wage polarization between and within occupations and (2) an examination of geographic variation in income polarization.
Sociological approaches to digital and community-engaged research experienced significant innovation in recent years. This article examines developing and implementing a primarily virtual community-driven research (CDR) project with the National Survivors Union, the American national drug-users union, during the COVID-19 pandemic. Relationships between researchers and directly impacted people, such as people who use drugs, face many barriers. These issues were exacerbated during COVID-19 when in-person research decreased while drug-related harms increased. In response, this project modified the CDR model for drug-use research. The CDR model is particularly beneficial for studies with marginalized populations who may mistrust researchers. In CDR, impacted community members are fundamental project drivers. This project’s data are based on 29 months of weekly group meetings in National Survivors Union online spaces, group and individual text conversations, phone calls, and shared-document group work. The project co-developed methods for CDR with directly impacted people, including community-initiated research questions, low-threshold methods, collaborative writing strategies, coauthorship practices foregrounding directly impacted perspectives, and multiple dissemination forms. Modified CDR expands sociological methods for digital research, citizen science, and community-engaged research with vulnerable, criminalized groups. This approach may aid inclusive, innovative sociological scholarship and effective public health policy for reducing morbidity and mortality during multiple crises.
The qualitative interview has been a core technique in the sociological methods toolkit for generations. Interviews provide essential insights into how participants experience the world around them. New opportunities have emerged to adapt traditional in-depth interview techniques through the use of evolving technologies available to interview participants. This article describes the integration of ecological momentary assessment techniques to augment qualitative in-depth interviews focused on specific events, which we term event-centered interviewing. By incorporating photo data captured systematically through smartphone apps designed for ecological momentary assessment, event-centered interviews can extend the strengths of traditional qualitative interviews. We describe the processes and procedures for conducting event-centered interviews, and we highlight how the approach may create opportunities for qualitative analysis and minimize certain limitations of traditional in-depth interviews. We also highlight the positive participant responses to the approach from a pilot study. Although traditional in-depth interviews may remain at the core of qualitative sociological inquiry, event-centered interviewing may be especially useful for interviews about behavior and experiences that occur during specific events.
The proportion of explained variance is well defined in linear models, but Snijders and Bosker demonstrated that this concept is ill defined in linear multilevel models. Whenever a researcher adds a level 1 predictor to the model, the level 2 variance may increase because the level 2 variance also depends on the level 1 variance. This problem is more pronounced when there are few observations per cluster. The authors present a solution that allows researchers to decompose variance components from null models into parts explained and unexplained by level 1 predictors. The authors also offer an extension that incorporates level 2 predictors. This approach is based on multivariate multilevel modeling and provides a complete decomposition of the gross (or null model) variance components. The approach is also implemented in the user-written Stata program twolevelr2, and the online supplement contains worked code for implementation in R. The authors illustrate this method with an example analyzing sibling similarities in lifetime income.