
The paper argues for the value of conducting surveys in vulnerable neighborhoods and provides a detailed account of a cost-effective strategy for surveying a recognized hard-to-survey population. The approach is illustrated through insights from the Vulnerable Neighborhoods Survey, conducted in three Nordic countries. The strategy focuses on a small number of specific neighborhoods and implements a range of measures to lower participation barriers. A key component involves combining random and non-random sampling techniques to facilitate the recruitment of a broad segment of residents. According to comparisons with registry data, the strategy produces samples that resemble the population on multiple demographic factors.
Randomized paired comparisons (RPC) for social values have various advantages over a matrix format of multiple items; however, their use cannot exhaust all possible pairs if there are too many items to compare one-to-one. This article proposes (1) applying a dimension reduction method, structural topic modeling (STM), to RPC survey data by restructuring answers into ordered pairs to estimate latent answering patterns, (2) visualizing them into directed graphs, and (3) interpreting them as respondents’ preference structures among social values. For empirical validation, we randomly divided 920 respondents into RPC and matrix-format groups and asked about the seriousness of ten social problems. Our STM from the RPC group revealed five preference structures beyond a linear order among the 10 items, which are interpretable and incorporate statistical tests with respondents’ traits as covariates. We also discuss how to improve topic modeling with RPC and contribute to various research streams, such as cultural value networks and gamification, by pairwise wiki survey.
A series of papers uses administrative data on school students’ grades to assess whether teachers discriminate against certain demographic groups. Often, differences in teacher and test grades are regressed on student-level variables. However, it is unclear under what circumstances such an estimation strategy is valid. We conceptualize teacher bias as a direct causal effect of student-level attributes on teacher grades, fixing student ability. Standardized tests merely proxy for student ability; additionally, there may be confounders of ability and teacher grade. Accordingly, teacher bias is nonparametrically unidentified. However, we suggest substantive and parametric assumptions that ensure identification using difference-in-grades estimators. Estimators based on regression control for test grades are shown to be inconsistent even under these strong assumptions. We then develop a parametric sensitivity analysis that allows researchers to investigate the consequences of departures from critical assumptions. We illustrate our methodology using administrative data from Denmark.
Estimating the size of the undocumented migrant population remains a critical challenge for researchers and policymakers. This study assesses the viability of using social media platforms, specifically Facebook and Instagram, to recruit a survey sample of migrants and elicit their status. The research focuses on Mexican and Venezuelan immigrants in Texas, Florida, Illinois, and California. Three methods for eliciting legal status are tested: direct questions, indirect sequential questions (stepwise exclusion), and a list experiment. The study ( N = 2,027) finds that while social media recruitment is cost-effective and rapid, it faces challenges such as selection bias, misclassification, and platform-imposed restrictions. The list experiment suggests the presence of response bias in traditional surveys to sensitive legal status questions. Estimates of the share of undocumented migrants deviate considerably from available reference estimates. We argue that social media surveys are best applied in preparation for traditional surveys rather than in their place.
This research evaluates methodologies to mitigate misreporting in intimate partner violence (IPV) data collection in a middle-income country. We conducted surveys in Russia involving three list experiments, a self-administered tablet questionnaire, a self-administered online survey, and conventional face-to-face interviews. Results show that list experiments yield lower disclosure rates for the complex IPV definitions suggested by the UN. The tablet-based self-administered questionnaire, conducted with an interviewer present, also did not increase IPV reporting. Conversely, the self-administered online survey increased lifetime IPV disclosures by 51% (physical) and 26% (psychological) compared to face-to-face interviews. Women showed greater sensitivity to the online survey mode. This increase is linked to the absence of interviewer bias, enhanced safety by minimizing potential perpetrators' presence, and reduced cognitive burden. We argue that self-administered online surveys-using sampling bias mitigation-may thus be an optimal, low-cost method for surveying the general population in middle- and high-income countries.
Identities are fundamental to our understanding of social and political behavior, but are challenging to measure and are rarely observed in real-world settings. We introduce a method for measuring the identity-relevant aspects of brief self-descriptions regularly used online (e.g., on social media). Our approach combines the benefits of word embeddings for finding related identity terms with the ability of clustering algorithms to aggregate terms into discrete categories. To illustrate our approach, we apply it to daily observations of bios from millions of US Twitter/X users. We present three applications of our approach with substantive findings. First, we track users' social and political identities over time and find, among other things, that direct expressions of political affiliations are rare. Second, we map the identities that are most characteristic of each US state. Third, we show that users' political identities are highly predictable based on non-political identity markers. With the growing availability of user self-descriptions on social media platforms and elsewhere, our approach enables researchers to map and analyze expressions of identity at scale.
Large-N, inference-based approaches are gaining increasing prominence in video-based social science research across sociology, social psychology, political science, and other fields. However, existing methodological publications on video methods do not discuss sampling methodology and empirical video-based research often includes only cursory discussions of the issue. To address this gap, this article applies insights from sampling methodology to video-based social science research. We review how sampling has been addressed in video-based social science research, reflect on its specific challenges, and propose a decision-tree flowchart to help researchers identify appropriate sampling strategies and common pitfalls. We then illustrate how the flowchart can be used in three common video-based sampling scenarios. The article thereby contributes to establishing clear guidelines for sampling in video-based social research as a reference point and as a resource for current and future practitioners, as well as reviewers and readers of such studies.
Status is central to understanding collaborative behavior, yet it is often difficult to measure in cultural fields where perceived standings are only partially observable. This study develops a scalable supervised machine learning approach to infer directed deference in collaboration networks using a partially observed status hierarchy derived from a ritualized site of status conferral (a televised competition series). Drawing on a longitudinal "featuring" network of more than 3,000 South Korean hip-hop artists, we train a classifier to learn how differences in status-relevant characteristics map onto observed deference patterns and then use it to estimate preferential attachment across all collaboration dyads. The resulting measure aligns closely with external expert assessments of artists' relative standing. Applying this metric to streaming performance data, we show that collaboration improves listener engagement and that its effect varies nonlinearly with status distance: artists benefit both from partnering with higher-status collaborators and from featuring emerging talents.
This article introduces treatment effect on the association between outcomes (TEA), a new causal estimand that measures how a treatment influences the covariance between two post-treatment variables. TEA enables researchers to estimate how interventions affect associations that characterize social inequalities. I define TEA, provide identification results under standard causal inference assumptions, and outline estimation strategies including regression-imputation, weighting, and double machine learning estimators. I compare and contrast TEA with other common estimands in similar research settings, highlighting its unique use. I demonstrate the use of TEA through two applications: the effect of college completion on income gradient in health and the effect of college completion on issue alignment, using NLSY97 and GSS, respectively. By exploring how treatments modify associations between outcomes, TEA offers a valuable tool for sociological research on inequality, stratification, and public opinion, providing insights into the mechanisms sustaining social inequalities and informing policy interventions.
Studying the relationship between neighborhoods and individual-level outcomes such as crime, labor market success, or intergenerational mobility has a long history in the social sciences. As local processes like gentrification constantly change neighborhoods' composition and spatial expansion, time-constant one-size-fits-all neighborhood measures fail to capture important local dynamics. This article presents a flexible and data-driven approach for efficiently estimating overlapping and arbitrarily shaped neighborhoods with time-dynamic boundaries. Constructed in a two-stage clustering design, the first stage identifies homogeneous groups within a city, while the second stage clusters homogeneous groups by spatial proximity. In an analysis of 86 million person-year observations from 76 German cities, the paper shows that a larger spatial expansion of affluent neighborhoods negatively correlates with city crime cases, while higher neighborhood fragmentation and heterogeneity correlate positively with crime rates. The findings stress the importance of flexible neighborhood estimation techniques and the necessity to view neighborhoods as nonconstant entities.
Researchers routinely face suspicion during fieldwork. This article presents findings from interviews with 34 ethnographers who were suspected of being spies while conducting fieldwork in Turkey. I find that the way the ethnographers experienced and responded to this suspicion depended on whether they reported being questioned about whether they were spies versus accused of spying. Questioning was interpreted as sense-making, and researchers reported several common strategies for addressing the suspicions they faced. Accusations, in contrast, were associated with threats and motivated the researchers to mitigate risks to themselves and their interlocutors. Engaging with scholarship on social cognition, high-risk fieldwork, and reflexivity, I discuss how my findings offer practical insights for navigating suspicion and risk during fieldwork-even in seemingly low-risk environments-and I make the case that interrogating how researchers react to suspicion can help them clarify their positionality and aid reflexivity.
This study addresses the complexities of midpoint and non-substantive responses, such as "Don't Know," in Likert scale surveys on gender attitudes. While existing research often assumes these responses are neutral or random, this study challenges that notion by applying the item response tree model to disentangle respondents' attitudes from their response tendencies. Analysis of data from the Chinese General Social Survey shows that traditional gender attitudes are associated with a higher likelihood of such responses, indicating biases in conventional methods. After disentangling these elements, I reassemble them through latent profile analysis to examine the dominant configurations of gender attitudes and response tendencies. Five distinct profiles emerge: Passionate Egalitarians, Genuine Neutrals, Forthright Moderates, Evasive Traditionalists, and Unembellished Traditionalists. Passionate Egalitarians advocate strongly for gender equality, while Evasive Traditionalists use non-substantive responses to conceal traditional views. This study provides a refined approach to rating scale analysis, advancing both sociological methodology and gender studies.
Extant work has identified two discursive forms of racism: overt and covert. While both forms have received attention in scholarly work, research on covert racism has been limited. Its subtle and context-specific nature has made it difficult to systematically identify covert racism in text, especially in large corpora. In this article, we first propose a theoretically driven and generalizable process to identify and classify covert and overt racism in text. This process allows researchers to construct coding schemes and build labeled datasets. We use the resulting dataset to train XLM-RoBERTa, a cross-lingual large language model (LLM) for supervised classification with a cutting-edge contextual understanding of text. We show that XLM-R and XLM-R-Racismo, our pretrained model, outperform other state-of-the-art approaches in classifying racism in large corpora. We illustrate our approach using a corpus of tweets relating to the Ecuadorian ind & iacute;gena community between 2018 and 2021.
Analyzing social change requires detecting patterns of continuity and difference over time. While time-series clustering offers a valuable approach, existing techniques are often limited by assuming fixed cluster definitions and static assignments of entities to clusters. To address these limitations, we introduce a unified framework of temporal clustering methods that allows for both dynamic cluster definitions and the transition of entities between clusters, generalizing and extending previous work. We also provide new algorithms for this dynamic clustering that optimize global objectives, with optional constraints on the transitions of entities across clusters. This framework expands the methodological toolkit for analyzing social change, and we provide guidelines for its application. We illustrate our approach with three case studies: polarization of social and political attitudes across U.S. states; cross-national cultural change; and the evolution of neighborhood business patterns. We conclude with directions for further research.
Social science theories often postulate systems of causal relationships among variables, which are commonly represented using directed acyclic graphs (DAGs). As non-parametric causal models, DAGs require no assumptions about the functional form of the hypothesized relationships. Nevertheless, to simplify empirical evaluation, researchers typically invoke such assumptions anyway, even though they are often arbitrary and do not reflect any theoretical content or prior knowledge. Moreover, functional form assumptions can engender bias, whenever they fail to accurately capture the true complexity of the system. In this article, we introduce causal-graphical normalizing flows (cGNFs), a novel approach to causal inference that leverages deep neural networks to empirically evaluate theories represented as DAGs. Unlike conventional methods, cGNFs model the full joint distribution of the data using a DAG specified by the analyst, without relying on stringent assumptions about functional form. This enables flexible, non-parametric estimation of any causal estimand identified from the DAG, including total effects, direct and indirect effects, and path-specific effects. We illustrate the method with a reanalysis of Blau and Duncan’s ( 1967 ) model of status attainment and Zhou’s ( 2019 ) model of controlled mobility. The article concludes with a discussion of current limitations and directions for future development.
Can aggregated composite scores be used to compare countries or other groups despite measurement non-invariance? We propose a pragmatic approach, emphasizing that measurement invariance is valuable but not strictly necessary for all such comparisons. For descriptive analyses of group differences, composite scores may outperform factor-analytic approaches, because they are more intuitive and can capture multiple dimensions. Using data from the European Social Survey (39 countries, 11 measurement occasions, 546,954 respondents), we examined social and political trust. Composite scores aggregated to the country level were practically indistinguishable from countries' factor scores based on approximate measurement invariance testing. We conclude that composite scores can suffice for simple group comparisons, though their suitability depends on the data. They can, however, underestimate uncertainty, producing overly narrow confidence intervals. We further show that measurement invariance does not guarantee measurement equivalence. Finally, we highlight how researchers can leverage data even if measurement invariance fails.
In the social sciences, most process tracing evidence is gathered through individual or atomized sources. However, there are some cases in which individualized data collection methods are not enough to capture collective social processes. We propose using focus groups for process tracing (FGFPT) to gather and analyze qualitative evidence about causal processes and mechanisms by leveraging interaction and discussions. We present three key benefits of using FGFPT: instant fact-checking, obtaining mechanistic evidence through the interactive process, and enhancing participants' collective agency. Additionally, we propose general guidelines for designing and implementing focus groups with the aim of process tracing: specifying observable implications, forming the focus group, question design, and training the moderator. Focus groups can be the most adequate data collection method to support and enhance process tracing exercises for collective phenomena.
Multilevel modelling (MM) is widely utilized in the social sciences, with over 20% of articles in leading sociological journals employing this technique. Despite its prevalence, few studies address whether the variables used in MM are invariant across groups or allow to construct reliable indicators. This study investigates the effects of both measurement noninvariance and random measurement error on MM using Monte Carlo simulations. Our findings reveal significant biases in MM results when random measurement errors are overlooked. Attaining high reliability in the indicators - above 0.94 - can mitigate these biases. While measurement noninvariance introduces bias in MM, its impact is smaller compared to that of the bias caused by unaddressed measurement error. Multilevel structural equation modelling (SEM), which controls for random measurement errors, performs effectively in complete measurement invariance (MI) scenarios. However, the absence of MI can create significant challenges. While multilevel SEM is a powerful analytical tool, it is not immune to the effects of MI assumption violations.
Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy (MAIHDA) is a multilevel regression approach grounded in intersectionality theory. It examines inequalities across intersections of social identities (e.g., gender, ethnicity, class) and is argued to provide more accurate predictions of intersectional means than conventional methods that estimate group means directly or via regressions with all interactions. This study evaluates that claim using analytic expressions and an empirical illustration to compare simple and MAIHDA-predicted means against population values. Predictive accuracy is assessed via variance, correlation, bias, and mean squared error. Results show that MAIHDA estimates generally outperform simple means, particularly when decomposing intersectional means into additive and non-additive identity effects. The magnitude of the advantage depends on inequality patterns and group sample sizes. MAIHDA is especially valuable when inequalities are subtle or data for marginalized intersections are sparse-conditions common in practice. These findings highlight MAIHDA's practical relevance for quantitative intersectionality research.
Age-period-cohort analysis is often done in the context of two samples. This could be samples for women and men or for two countries. It is of interest to ask if some time effects could be common across samples. We clarify how the well-known age-period-cohort problem for one sample carries over to the two sample situation. This is done through a reparametrization in terms of parameters that are invariant to the identification issues. The new parametrization shows which hypotheses can be tested and their degrees of freedom. Testable hypotheses can be formulated for non-linear effects, but not for the linear parts of the individual time effects. This conclusion remains when imposing cross-sample restrictions. The analysis is extended to the mixed frequency situation where age and period are measured at different scales. As an empirical illustration a study of Swiss suicide rates is revisited.