Factorial surveys use a population of vignettes to elicit respondents' attitudes or beliefs about different hypothetical scenarios. However, the vignette population is frequently too large to be assessed by each respondent. Experimental designs such as randomized block confounded factorial (RBCF) designs, D-optimal designs, or random sampling designs can be used to construct small subsets of vignettes. In a simulation study, we compare the three vignette designs with respect to their biases in effect estimates and show how the biases arise from the designs' confounding structure, nonorthogonality, and unbalancedness. We particularly focus on the designs' sensitivity to context effects and misspecifications of the analytic model. We argue that RBCF designs and D-optimal designs are preferable to random sampling designs because they offer a stronger protection against undesirable confounding, context effects, and model misspecifications. We also discuss strategies for dealing with context and order effects since none of the basic vignette designs can satisfactorily handle them.
"Abstract: An Evaluation of Planned Missing Data Designs in Large Surveys." Multivariate Behavioral Research, 54(1), p. 146
This paper extends a recent study by Kaplan and Su (J Educ Behav Stat 41: 51–80, 2016) examining the problem of matrix sampling of context questionnaire scales with respect to the generation of plausible values of cognitive outcomes in large-scale assessments.
In survey research, vignette experiments typically employ short, systematically varied descriptions of situations or persons (called vignettes) to elicit the beliefs, attitudes, or behaviors of respondents with respect to the presented scenarios. Using a case study on the fair gender income gap in Austria, we discuss how different design elements can be used to increase a vignette experiment’s validity and reliability. With respect to the experimental design, the design elements considered include a confounded factorial design, a between-subjects factor, anchoring vignettes, and blocking by respondent strata and interviewers. The design elements for the sampling and survey design consist of stratification, covariate measurements, and the systematic assignment of vignette sets to respondents and interviewers. Moreover, the vignettes’ construct validity is empirically validated with respect to the real gender income gap in Austria. We demonstrate how a broad range of design elements can successfully increase a vignette study’s validity and reliability.
Randomized controlled trials (RCTs) and quasi-experimental designs like regression discontinuity (RD) designs, instrumental variable (IV) designs, and matching and propensity score (PS) designs are frequently used for inferring causal effects. It is well known that the features of these designs facilitate the identification of a causal estimand and, thus, warrant a causal interpretation of the estimated effect. In this article, we discuss and compare the identifying assumptions of quasi-experiments using causal graphs. The increasing complexity of the causal graphs as one switches from an RCT to RD, IV, or PS designs reveals that the assumptions become stronger as the researcher's control over treatment selection diminishes. We introduce limiting graphs for the RD design and conditional graphs for the latent subgroups of com-pliers, always takers, and never takers of the IV design, and argue that the PS is a collider that offsets confounding bias via collider bias.
This article presents findings on the consequences of matrix sampling of context questionnaires for the generation of plausible values in large-scale assessments. Three studies are conducted. Study 1 uses data from PISA 2012 to examine several different forms of missing data imputation within the chained equations framework: predictive mean matching, Bayesian linear regression, and proportional odds logistic regression. We find that predictive mean matching accurately reproduces the marginal distributions of the missing context questionnaire data due to matrix sampling. Study 2 uses data from PISA 2006 to examine the consequences of imputing context questionnaire data in terms of the generation of plausible values. We find that imputing context questionnaire data with predictive mean matching and using the imputed data to produce the plausible values yields very close approximation of the original marginal distributions but leads to underestimation of the correlation structure of the context questionnaire items. Study 3 examines imputation and plausible values generation within a partially balanced incomplete block design. We find that imputation within this design accurately reproduces the original marginal distributions and retains the correlation structure of the data. Implications for context questionnaire development are discussed.