Consider a study whose primary results are "not statistically significant". How often does it lead to the following published conclusion that "there is no effect of the treatment/exposure on the outcome"? We believe too often and that the requirement to report counternull values could help to avoid this! In statistical parlance, the null value of an estimand is a value that is distinguished in some way from other possible values, for example a value that indicates no difference between the general health status of those treated with a new drug versus a traditional drug. A counternull value is a nonnull value of that estimand that is supported by the same amount of evidence that supports the null value. Of course, such a definition depends critically on how "evidence" is defined. Here, we consider the context of a randomized experiment where evidence is summarized by the randomization-based p-value associated with a specified sharp null hypothesis. Consequently, a counternull value has the same p-value from the randomization test as does the null value; the counternull value is rarely unique, but rather comprises a set of values. We explore advantages to reporting a counternull set in addition to the p-value associated with a null value; a first advantage is pedagogical, in that reporting it avoids the mistake of implicitly accepting a not-rejected null hypothesis; a second advantage is that the effort to construct a counternull set can be scientifically helpful by encouraging thought about nonnull values of estimands. Two examples are used to illustrate these ideas.
The Rubin Causal Model (RCM), a framework for causal inference, has three distinctive features. First, it uses ‘potential outcomes’ to define causal effects at the unit level, first introduced by Neyman in the context of randomized experiments and randomization-based inference, but not used formally in non-randomized studies or with other modes of inference until Rubin (1974, 1975). Second is its formal use of a probabilistic assignment mechanism, which mathematically describes how treatments are given to units, with possible dependence on background variables and the potential outcomes themselves. Third is an optional probability distribution on all variables, including the potential outcomes, which thereby unifies frequentist and model-based forms of statistical inference for causal effects within one framework.
Rerandomization utilizes modern computing ability to improve covariate balance while adhering to the randomization principle originally advocated by RA Fisher. Affinely invariant rerandomization has the ``Equal Percent Variance Reducing'' (EPVR) property. When dealing with covariates of varying importance and/or mixed types, the conditionally EPVR property is often more desired. We discuss a general class of conditionally affinely invariant rerandomization methods and obtain their conditionally EPVR property. In addition, we set up a decision-theoretical framework to evaluate balance criteria for rerandomization. Popular rerandomization methods, such as the covariate balance table check, are found to be inadmissible. We suggest an admissible complete class of conditionally affinely invariant balance criteria, which can be applied to experimental designs involving tiers of covariates, stratification, and multiple treatment arms.
Catalytic prior distributions provide general, easy-to-use, and interpretable specifications of prior distributions for Bayesian analysis. They are particularly beneficial when the observed data are inadequate to stably estimate a complex target model. A catalytic prior distribution is constructed by augmenting the observed data with synthetic data that are sampled from the predictive distribution of a simpler model estimated from the observed data. We illustrate the usefulness of the catalytic prior approach using an example from labor economics. In the example, the resulting Bayesian inference reflects many important aspects of the observed data, and the estimation accuracy and predictive performance of the inference based on the catalytic prior are superior to, or comparable to, that of other commonly used prior distributions. We further explore the connection between the catalytic prior approach and a few popular regularization methods. We expect the catalytic prior approach to be useful in many applications.
Re-randomization has gained popularity as a tool for experiment-based causal inference due to its superior covariate balance and statistical efficiency compared to classic randomized experiments. However, the basic re-randomization method, known as ReM, and many of its extensions have been deemed sub-optimal as they fail to prioritize covariates that are more strongly associated with potential outcomes. To address this limitation and design more efficient re-randomization procedures, a more precise quantification of covariate heterogeneity and its impact on the causal effect estimator is in a great appeal. This work fills in this gap with a Bayesian criterion for re-randomization and a series of novel re-randomization procedures derived under such a criterion. Both theoretical analyses and numerical studies show that the proposed re-randomization procedures under the Bayesian criterion outperform existing ReM-based procedures significantly in effectively balancing covariates and precisely estimating the unknown causal effect.
Mahalanobis distance between treatment group and control group covariate means is often adopted as a balance criterion when implementing a rerandomization strategy. However, this criterion may not work well for high-dimensional cases because it balances all orthogonalized covariates equally. Here, we propose leveraging principal component analysis (PCA) to identify proper subspaces in which Mahalanobis distance should be calculated. Not only can PCA effectively reduce the dimensionality for high-dimensional cases while capturing most of the information in the covariates, but it also provides computational simplicity by focusing on the top orthogonal components. We show that our PCA rerandomization scheme has desirable theoretical properties on balancing covariates and thereby on improving the estimation of average treatment effects. We also show that this conclusion is supported by numerical studies using both simulated and real examples.
Existing methods that use propensity scores for heterogeneous treatment effect estimation on non-experimental data do not readily extend to the case of more than two treatment options. In this work, we develop a new propensity score-based method for heterogeneous treatment effect estimation when there are three or more treatment options, and prove that it generates unbiased estimates. We demonstrate our method on a real patient registry of patients in Singapore with diabetic dyslipidemia. On this dataset, our method generates heterogeneous treatment recommendations for patients among three options: Statins, fibrates, and non-pharmacological treatment to control patients' lipid ratios (total cholesterol divided by high-density lipoprotein level). In our numerical study, our proposed method generated more stable estimates compared to a benchmark method based on a multi-dimensional propensity score.
This study utilized the large-scale, multi-institutional CASSIE dataset to examine the impact of education abroad participation on academic outcomes for first-generation college students. Using robust multivariate matching methodology that effectively minimized self-selection bias, results showed the magnitude of benefit offered by studying abroad was greater for first-generation students than for continuing-generation students. Even after matching on a variety of background and prior achievement variables, first-generation students who studied abroad had higher 4- and 6-year graduation rates, had higher cumulative GPA scores, and took less time to graduate-relative to first-generation students who did not study abroad. These findings suggest that education abroad programming can be leveraged as a high-impact educational practice to promote college completion rates among first-generation students.
The propensity score is the conditional probability of assignment to a particular treatment given a vector of observed covariates. Both large and small sample theory show that adjustment for the scalar propensity score is sufficient to remove bias due to all observed covariates. Applications include: (i) matched sampling on the univariate propensity score, which is a generalization of discriminant matching, (ii) multivariate adjustment by subclassification on the propensity score where the same subclasses are used to estimate treatment effects for all outcome variables and in all subpopulations, and (iii) visual representation of multivariate covariance adjustment by a two-dimensional plot.
Meta-analysis can be a critical part of the research process, often serving as the primary analysis on which the practitioners, policymakers, and individuals base their decisions. However, current literature synthesis approaches to meta-analysis typically estimate a different quantity than what is implicitly intended; concretely, standard approaches estimate the average effect of a treatment for a population of imperfect studies, rather than the true scientific effect that would be measured in a population of hypothetical perfect studies. We advocate for an alternative method, called response-surface meta-analysis, which models the relationship between the quality of the study design as predictor variables and its reported estimated effect size as the outcome variable in order to estimate the effect size obtained by the hypothetical ideal study. The idea was first introduced by Rubin several decades ago, and here we provide a practical implementation. First, we reintroduce the idea of response-surface meta-analysis, highlighting its focus on a scientifically-motivated estimand while proposing a straightforward implementation. Then we compare the approach to traditional meta-analysis techniques used in practice. We then implement response-surface meta-analysis and contrast its results with existing literature-synthesis approaches on both simulated data and a real-world example published by the Cochrane Collaboration. We conclude by detailing the primary challenges in the implementation of response-surface meta-analysis and offer some suggestions to tackle these challenges.
Consider a situation with two treatments, the first of which is randomized but the second is not, and the multifactor version of this. Interest is in treatment effects, defined using standard factorial notation. We define estimators for the treatment effects and explore their properties when there is information about the nonrandomized treatment assignment and when there is no information on the assignment of the nonrandomized treatment. We show when and how hidden treatments can bias estimators and inflate their sampling variances.
The Gibbs sampler is an iterative simulation scheme for generating samples that converge to draws from a target distribution π(X) of random variable X. An iterative simulation scheme is used because direct simulation from the target distribution can not be easily implemented. To define the iterative Gibbs sampler, partition the components of X as X = (x1, . . . , xd) where xi is gi-dimensional, and thus X is a g = ∑d i=1 gi dimensional random vector. The Gibbs sampler is easy to implement when the set of d conditional distributions,
Although education abroad in the US offers participants demonstrable ben fits, direct and opportunity costs are cited as primary barriers to broader participation. Yet the degree to which low-income status deters studying abroad and whether additional need-based aid beyond Pell Grants encourages participation remain uncertain. Moreover, not all education abroad programs are equivalent in terms of costs. This study is the first to examine whether need-based aid recipients differentially choose programs of varying duration or programs offered by various provider types. The sample consisted of 221,981 students from 36 institutions of the Consortium for Analysis of Student Success through International Education (CASSIE). Within that sample, 60,477 received Pell grants. Of those recipients, 39% received additional need-based aid. Regression models controlling for student background and context indicated that Pell grant recipients were 3% less Rely to study abroad than peers receiving no such aid, and receipt of additional aid increased likelihood by 1% relative to Pell-only recipients. While aid was unrelated to study abroad duration, low-income students were less Rely to study with third party providers. The findings invite financial aid officers to determine thresholds of additional aid necessary to increase participation and to collaborate more systematically with counterparts in international education.
A common complication that can arise with analyses of high-dimensional data is the repeated use of hypothesis tests. A second complication, especially with small samples, is the reliance on asymptotic p-values. Our proposed approach for addressing both complications uses a scientifically motivated scalar summary statistic, and although not entirely novel, seems rarely used. The method is illustrated using a crossover study of seventeen participants examining the effect of exposure to ozone versus clean air on the DNA methylome, where the multivariate outcome involved 484,531 genomic locations. Our proposed test yields a single null randomization distribution, and thus a single Fisher-exact p-value that is statistically valid whatever the structure of the data. However, the relevance and power of the resultant test requires the careful a priori selection of a single test statistic. The common practice using asymptotic p-values or meaningless thresholds for "significance" is inapposite in general.
This study explores connections between design features of faculty-led short-term study abroad programs and resulting changes in students’ global perspectives. Over 2,000 students provided data for this study, completing the Global Perspective Inventory (GPI) before and after studying abroad. Results indicated that program features such as participation in an internship and opportunities for reflection are positively associated with global perspective development while abroad, whereas features such as number of students traveling together and coursework in English are negatively associated with such development. Given the increasing numbers of students who participate in faculty-led short-term abroad programs, research that provides evidence-based recommendations concerning program design is essential to enhancing global perspectives through study abroad.
Formal guidelines for statistical reporting of non-randomized studies are important for journals that publish results of such studies. Although it is gratifying to see some journals providing guidelines for statistical reporting, we feel that the current guidelines that we have seen are not entirely adequate when the study is used to draw causal conclusions. We therefore offer some comments on ways to improve these studies. In particular, we discuss and illustrate what we regard as the need for an essential initial stage of any such statistical analysis, the conceptual stage, which formally describes the embedding of a non-randomized study within a hypothetical randomized experiment.
Blocking is commonly used in randomized experiments to increase efficiency of estimation. A generalization of blocking removes allocations with imbalance in covariate distributions between treated and control units, and then randomizes within the remaining set of allocations with balance. This idea of rerandomization was formalized by Morgan and Rubin (Annals of Statistics, 2012, 40, 1263–1282), who suggested using Mahalanobis distance between treated and control covariate means as the criterion for removing unbalanced allocations. Kallus (Journal of the Royal Statistical Society, Series B: Statistical Methodology, 2018, 80, 85–112) proposed reducing the set of balanced allocations to the minimum. Here we discuss the implication of such an ‘optimal’ rerandomization design for inferences to the units in the sample and to the population from which the units in the sample were randomly drawn. We argue that, in general, it is a bad idea to seek the optimal design for an inference because that inference typically only reflects uncertainty from the random sampling of units, which is usually hypothetical, and not the randomization of units to treatment versus control.
We describe a new method to combine propensity‐score matching with regression adjustment in treatment‐control studies when outcomes are binary by multiply imputing potential outcomes under control for the matched treated subjects. This enables the estimation of clinically meaningful measures of effect such as the risk difference. We used Monte Carlo simulation to explore the effect of the number of imputed potential outcomes under control for the matched treated subjects on inferences about the risk difference. We found that imputing potential outcomes under control (either single imputation or multiple imputation) resulted in a substantial reduction in bias compared with what was achieved using conventional nearest neighbor matching alone. Increasing the number of imputed potential outcomes under control resulted in more efficient estimation, with more efficient estimation of the estimated risk difference when increasing the number of the imputed potential outcomes. The greatest relative increase in efficiency was achieved by imputing five potential outcomes; once 20 outcomes under control were imputed for each matched treated subject, further improvements in efficiency were negligible. We also examined the effect of the number of these imputed potential outcomes on: (i) estimated standard errors; (ii) mean squared error; (iii) coverage of estimated confidence intervals. We illustrate the application of the method by estimating the effect on the risk of death within 1 year of prescribing beta‐blockers to patients discharged from hospital with a diagnosis of heart failure.
The weaponization of digital communications and social media to conduct disinformation campaigns at immense scale, speed, and reach presents new challenges to identify and counter hostile influence operations (IOs). This paper presents an end-to-end framework to automate detection of disinformation narratives, networks, and influential actors. The framework integrates natural language processing, machine learning, graph analytics, and a network causal inference approach to quantify the impact of individual actors in spreading IO narratives. We demonstrate its capability on real-world hostile IO campaigns with Twitter datasets collected during the 2017 French presidential elections and known IO accounts disclosed by Twitter over a broad range of IO campaigns (May 2007 to February 2020), over 50,000 accounts, 17 countries, and different account types including both trolls and bots. Our system detects IO accounts with 96% precision, 79% recall, and 96% area-under-the precision-recall (P-R) curve; maps out salient network communities; and discovers high-impact accounts that escape the lens of traditional impact statistics based on activity counts and network centrality. Results are corroborated with independent sources of known IO accounts from US Congressional reports, investigative journalism, and IO datasets provided by Twitter.