A recurring debate in the philosophy of statistics concerns what, exactly, should count as a measure of evidence for or against a given hypothesis. P-values, likelihood ratios, and Bayes factors all have their defenders. In this paper we add two additional candidates to this list: the e-value and its sequential analogue, the e-process. E-values enjoy several desirable properties as measures of evidence: they combine naturally across studies, handle composite hypotheses, provide long-run error rates, and admit a useful interpretation as the wealth accrued by a bettor in a game against the null distribution. E-processes additionally handle optional stopping and optional continuation. This work examines the extent to which e-values and e-processes satisfy the evidential desiderata of different statistical traditions, concluding that they combine attractive features of p-values, likelihood ratios, and Bayes factors, and merit serious consideration as interpretable and intuitive measures of statistical evidence.
We consider the problem of nonasymptotic mean estimation for heavy-tailed random variables taking values in a smooth Banach space. In particular, we extend and improve a simple truncation-based mean estimator by Catoni and Giulini. We show that this estimator enjoys finite-sample, high probability deviation bounds that hold uniformly over the class of distributions with a pth central moment bounded above by a known constant v > 0 for p is an element of (1, 2]. Our main contributions exploit connections between truncation-based mean estimation and the concentration of martingales with bounded increments in smooth Banach spaces. We prove two types of time-uniform bounds on the distance between the estimator and unknown mean: line-crossing inequalities, which can be optimized for a fixed sample size n, and iterated logarithm inequalities, which match the tightness of line-crossing inequalities at all points in time up to a doubly logarithmic factor in n. Our results do not depend on the dimension of the Banach space, hold under martingale dependence, and all employed constants are known and small.
We derive inferential procedures for large sample sizes that remain valid under data-dependent significance levels (so-called "post-hoc valid inference"). Classical statistical tools require that the significance level – the "type-I error" – is selected prior to seeing or analyzing any data. This restriction leads to some drawbacks. For instance, if an analyst generates an inconclusive confidence interval, repeating the process with a larger significance level is not an option – the result is final. Recently, e-values have emerged as the solution to this problem, being both necessary and sufficient tools for performing various forms of post-hoc inference. All such results, however, have thus far been nonasymptotic. As a result, they inherit familiar limitations of nonasymptotic inferential procedures such as requiring strong moment assumptions and being conservative in general. This paper develops a theory of post-hoc inference in the asymptotic setting, yielding asymptotic post-hoc confidence sets and asymptotic post-hoc p-values that make weaker assumptions and are sharper than their nonasymptotic counterparts.
Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, where significance levels may be data-dependent. Existing asymptotic e-values, however, suffer from the “missing factor,” a scaling inefficiency resulting in overly conservative inference. Drawing on the framework of near-optimal concentration inequalities developed by Bentkus in the 2000s, we introduce Bentkus-type asymptotic e-values and prove that they successfully eliminate the missing factor. We also demonstrate both theoretically and empirically that Bentkus-type e-values consistently deliver sharper inference than existing alternatives, leading to tighter post-hoc confidence intervals and higher rejection rates in multiple testing procedures.
The validity of classical hypothesis testing requires the significance level alpha be fixed before any statistical analysis takes place. This is a stringent requirement. For instance, it prohibits updating alpha during (or after) an experiment due to changing concern about the cost of false positives, or to reflect unexpectedly strong evidence against the null. Perhaps most disturbingly, witnessing a p-value p << alpha vs p = alpha - & varepsilon; for tiny & varepsilon; > 0 has no (statistical) relevance for any downstream decision-making. Following recent work of Gr & uuml;nwald [1], we develop a theory of post-hoc hypothesis testing, enabling alpha to be chosen after seeing and analyzing the data. To study "good" post-hoc tests we introduce Gamma-admissibility, where Gamma is a set of adversaries which map the data to a significance level. We classify the set of Gamma-admissible rules for various sets Gamma, showing they must be based on e-values, and recover the Neyman-Pearson lemma when Gamma is the constant map.
We derive a new closed-form variance-adaptive confidence sequence (CS) for estimating the average conditional mean of a sequence of bounded random variables. Empirically, it yields the tightest closed-form CS we have found for tracking time-varying means, across sample sizes up to ≈ 10^6. When the observations happen to have the same conditional mean, our CS is asymptotically tighter than the recent closed-form CS of Waudby-Smith and Ramdas [38]. It also has other desirable properties: it is centered at the unweighted sample mean and has limiting width (multiplied by √(t/log t)) independent of the significance level. We extend our results to provide a CS with the same properties for random matrices with bounded eigenvalues.
We study the self-normalized concentration of vector-valued stochastic processes. We focus on bounds for sub-ψ processes, a tail condition that encompasses a wide variety of well-known distributions (including sub-exponential, sub-Gaussian, sub-gamma, and sub-Poisson distributions). Our results recover and generalize the influential bound of Abbasi-Yadkori et al. (2011) and fill a gap in the literature between determinant-based bounds and those based on condition numbers. As applications we prove a Bernstein inequality for random vectors satisfying a moment condition (which is more general than boundedness), and also provide the first dimension-free, self-normalized empirical Bernstein inequality. Our techniques are based on the variational (PAC-Bayes) approach to concentration.
We show that for any concave utility, the expected utility of an e-variable can only increase after conditioning on a sufficient statistic. The simplest form of the result has an extremely straightforward proof, which follows from a single application of Jensen's inequality. Similar statements hold for compound e-variables, asymptotic e-variables, and e-processes. These results echo the Rao-Blackwell theorem, which states that the expected squared error of an estimator can only decrease after conditioning on a sufficient statistic. We provide several applications of this insight, including a simplified derivation of the log-optimal e-variable for linear regression with known variance.
We study sequential mean estimation in $\mathbb{R}^d$. In particular, we derive time-uniform confidence spheres -- confidence sphere sequences (CSSs) -- which contain the mean of random vectors with high probability simultaneously across all sample sizes. Our results include a dimension-free CSS for log-concave random vectors, a dimension-free CSS for sub-Gaussian random vectors, and CSSs for sub-$\psi$ random vectors (which includes sub-gamma, sub-Poisson, and sub-exponential distributions). Many of our results are optimal. For sub-Gaussian distributions we also provide a CSS which tracks a time-varying mean, generalizing Robbins' mixture approach to the multivariate setting. Finally, we provide several CSSs for heavy-tailed random vectors (two moments only). Our bounds hold under a martingale assumption on the mean and do not require that the observations be iid. Our work is based on PAC-Bayesian theory and inspired by an approach of Catoni and Giulini.
We present a unified framework for deriving PAC-Bayesian generalization bounds. Unlike most previous literature on this topic, our bounds are anytime-valid (i.e., time-uniform), meaning that they hold at all stopping times, not only for a fixed sample size. Our approach combines four tools in the following order: (a) nonnegative supermartingales or reverse submartingales, (b) the method of mixtures, (c) the Donsker-Varadhan formula (or other convex duality principles), and (d) Ville's inequality. Our main result is a PAC-Bayes theorem which holds for a wide class of discrete stochastic processes. We show how this result implies time-uniform versions of well-known classical PAC-Bayes bounds, such as those of Seeger, McAllester, Maurer, and Catoni, in addition to many recent bounds. We also present several novel bounds. Our framework also enables us to relax traditional assumptions; in particular, we consider nonstationary loss functions and non-i.i.d. data. In sum, we unify the derivation of past bounds and ease the search for future bounds: one may simply check if our supermartingale or submartingale conditions are met and, if so, be guaranteed a (time-uniform) PAC-Bayes bound.
We consider estimating the shared mean of a sequence of heavy-tailed random variables taking values in a Banach space. In particular, we revisit and extend a simple truncation-based mean estimator by Catoni and Giulini. While existing truncation-based approaches require a bound on the raw (non-central) second moment of observations, our results hold under a bound on either the central or non-central $p$th moment for some $p > 1$. In particular, our results hold for distributions with infinite variance. The main contributions of the paper follow from exploiting connections between truncation-based mean estimation and the concentration of martingales in 2-smooth Banach spaces. We prove two types of time-uniform bounds on the distance between the estimator and unknown mean: line-crossing inequalities, which can be optimized for a fixed sample size $n$, and non-asymptotic law of the iterated logarithm type inequalities, which match the tightness of line-crossing inequalities at all points in time up to a doubly logarithmic factor in $n$. Our results do not depend on the dimension of the Banach space, hold under martingale dependence, and all constants in the inequalities are known and small.
We derive and study time-uniform confidence spheres - termed confidence sphere sequences (CSSs) - which contain the mean of random vectors with high probability simultaneously across all sample sizes. Inspired by the original work of Catoni and Giulini, we unify and extend their analysis to cover both the sequential setting and to handle a variety of distributional assumptions. More concretely, our results include an empirical-Bernstein CSS for bounded random vectors (resulting in a novel empirical-Bernstein confidence interval), a CSS for sub-$\psi$ random vectors, and a CSS for heavy-tailed random vectors based on a sequentially valid Catoni-Giulini estimator. Finally, we provide a version of our empirical-Bernstein CSS that is robust to contamination by Huber noise.
Entropy regularization is known to improve exploration in sequential decision-making problems. We show that this same mechanism can also lead to nearly unbiased and lower-variance estimates of the mean reward in the optimize-and-estimate structured bandit setting. Mean reward estimation (i.e., population estimation) tasks have recently been shown to be essential for public policy settings where legal constraints often require precise estimates of population metrics. We show that leveraging entropy and KL divergence can yield a better trade-off between reward and estimator variance than existing baselines, all while remaining nearly unbiased. These properties of entropy regularization illustrate an exciting potential for bridging the optimal exploration and estimation literatures.
We provide practical, efficient, and nonparametric methods for auditing the fairness of deployed classification and regression models. Whereas previous work relies on a fixed-sample size, our methods are sequential and allow for the continuous monitoring of incoming data, making them highly amenable to tracking the fairness of real-world systems. We also allow the data to be collected by a probabilistic policy as opposed to sampled uniformly from the population. This enables auditing to be conducted on data gathered for another purpose. Moreover, this policy may change over time and different policies may be used on different subpopulations. Finally, our methods can handle distribution shift resulting from either changes to the model or changes in the underlying population. Our approach is based on recent progress in anytime-valid inference and game-theoretic statistics---the ``testing by betting'' framework in particular. These connections ensure that our methods are interpretable, fast, and easy to implement. We demonstrate the efficacy of our approach on three benchmark fairness datasets.
We introduce a new setting, optimize-and-estimate structured bandits . Here, a policy must select a batch of arms, each characterized by its own context, that would allow it to both maximize reward and maintain an accurate (ideally unbiased) population estimate of the reward. This setting is inherent to many public and private sector applications and often requires handling delayed feedback, small data, and distribution shifts. We demonstrate its importance on real data from the United States Internal Revenue Service (IRS). The IRS performs yearly audits of the tax base. Two of its most important objectives are to identify suspected misreporting and to estimate the "tax gap" — the global difference between the amount paid and true amount owed. Based on a unique collaboration with the IRS, we cast these two processes as a unified optimize-and-estimate structured bandit. We analyze optimize-and-estimate approaches to the IRS problem and propose a novel mechanism for unbiased population estimation that achieves rewards comparable to baseline approaches. This approach has the potential to improve audit efficacy, while maintaining policy-relevant estimates of the tax gap. This has important social consequences given that the current tax gap is estimated at nearly half a trillion dollars. We suggest that this problem setting is fertile ground for further research and we highlight its interesting challenges. The results of this and related research are currently being incorporated into the continual improvement of the IRS audit selection methods.
Concentrated animal feeding operations (CAFOs) pose serious risks to air, water, and public health, but have proven to be challenging to regulate. The U.S. Government Accountability Office notes that a basic challenge is the lack of comprehensive location information on CAFOs. We use the U.S. Department of Agriculture's National Agricultural Imagery Program 1 m/pixel aerial imagery to detect poultry CAFOs across the continental USA. We train convolutional neural network models to identify individual poultry barns and apply the best-performing model to over 42 TB of imagery to create the first national open-source dataset of poultry CAFOs We validate the model predictions against held-out validation set on poultry CAFO facility locations from ten hand-labeled counties in California and demonstrate that this approach has significant potential to fill gaps in environmental monitoring.
We explore the promises and challenges of employing sequential decision-making algorithms -- such as bandits, reinforcement learning, and active learning -- in law and public policy. While such algorithms have well-characterized performance in the private sector (e.g., online advertising), the tendency to naively apply algorithms motivated by one domain, often online advertisements, can be called the "advertisement fallacy." Our main thesis is that law and public policy pose distinct methodological challenges that the machine learning community has not yet addressed. Machine learning will need to address these methodological problems to move "beyond ads." Public law, for instance, can pose multiple objectives, necessitate batched and delayed feedback, and require systems to learn rational, causal decision-making policies, each of which presents novel questions at the research frontier. We discuss a wide range of potential applications of sequential decision-making algorithms in regulation and governance, including public health, environmental protection, tax administration, occupational safety, and benefits adjudication. We use these examples to highlight research needed to render sequential decision making policy-compliant, adaptable, and effective in the public sector. We also note the potential risks of such deployments and describe how sequential decision systems can also facilitate the discovery of harms. We hope our work inspires more investigation of sequential decision making in law and public policy, which provide unique challenges for machine learning researchers with potential for significant social benefit.
This paper introduces a new, highly consequential setting for the use of computer vision for environmental sustainability. Concentrated Animal Feeding Operations (CAFOs) (aka intensive livestock farms or "factory farms") produce significant manure and pollution. Dumping manure in the winter months poses significant environmental risks and violates environmental law in many states. Yet the federal Environmental Protection Agency (EPA) and state agencies have relied primarily on self-reporting to monitor such instances of "land application." Our paper makes four contributions. First, we introduce the environmental, policy, and agricultural setting of CAFOs and land application. Second, we provide a new dataset of high-cadence (daily to weekly) 3m/pixel satellite imagery from 2018-20 for 330 CAFOs in Wisconsin with hand labeled instances of land application (n=57,697). Third, we develop an object detection model to predict land application and a system to perform inference in near real-time. We show that this system effectively appears to detect land application (PR AUC = 0.93) and we uncover several outlier facilities which appear to apply regularly and excessively. Last, we estimate the population prevalence of land application events in Winter 2021/22. We show that the prevalence of land application is much higher than what is self-reported by facilities. The system can be used by environmental regulators and interest groups, one of which piloted field visits based on this system this past winter. Overall, our application demonstrates the potential for AI-based computer vision systems to solve major problems in environmental compliance with near-daily imagery.
This cohort study compares 3 protocols for allocating health workers to deliver COVID-19 tests door-to-door within East San Jose, California, including index area selection, uncertainty sampling, and local knowledge of the health workers. Question What are effective mechanisms to identify and reach vulnerable populations and equalize access to COVID-19 testing resources in the presence of substantial demographic disparities? Findings In this cohort study of 756 participants, a door-to-door program with community-based health workers was associated with a substantial increase in the proportion of Latinx and elderly individuals undergoing testing, relative to neighborhood testing sites. The protocol associated with the greatest increase in testing at-risk individuals was uncertainty sampling, followed by local knowledge, and then targeting households in areas with a high number of index cases. Meaning These findings suggest that community-based testing programs that allocate resources using uncertainty sampling might effectively reduce COVID-19 testing disparities. Importance Overcoming social barriers to COVID-19 testing is an important issue, especially given the demographic disparities in case incidence rates and testing. Delivering culturally appropriate testing resources using data-driven approaches in partnership with community-based health workers is promising, but little data are available on the design and effect of such interventions. Objectives To assess and evaluate a door-to-door COVID-19 testing initiative that allocates visits by community health workers by selecting households in areas with a high number of index cases, by using uncertainty sampling for areas where the positivity rate may be highest, and by relying on local knowledge of the health workers. Design, Setting, and Participants This cohort study was performed from December 18, 2020, to February 18, 2021. Community health workers visited households in neighborhoods in East San Jose, California, based on index cases or uncertainty sampling while retaining discretion to use local knowledge to administer tests. The health workers, also known as promotores de salud (hereinafter referred to as promotores) spent a mean of 4 days a week conducting door-to-door COVID-19 testing during the 2-month study period. All residents of East San Jose were eligible for COVID-19 testing. The promotores were selected from the META cooperative (Mujeres Empresarias Tomando Accion [Entrepreneurial Women Taking Action]). Interventions The promotores observed self-collection of anterior nasal swab samples for SARS-CoV-2 reverse transcriptase-polymerase chain reaction tests. Main Outcomes and Measures A determination of whether door-to-door COVID-19 testing was associated with an increase in the overall number of tests conducted, the demographic distribution of the door-to-door tests vs local testing sites, and the difference in positivity rates among the 3 door-to-door allocation strategies. Results A total of 785 residents underwent door-to-door testing, and 756 were included in the analysis. Among the 756 individuals undergoing testing (61.1% female; 28.2% aged 45-64 years), door-to-door COVID-19 testing reached different populations than standard public health surveillance, with 87.6% (95% CI, 85.0%-89.8%) being Latinx individuals. The closest available testing site only reached 49.0% (95% CI, 48.3%-49.8%) Latinx individuals. Uncertainty sampling provided the most effective allocation, with a 10.8% (95% CI, 6.8%-16.0%) positivity rate, followed by 6.4% (95% CI, 4.1%-9.4%) for local knowledge, and 2.6% (95% CI, 0.7%-6.6%) for index area selection. The intervention was also associated with increased overall testing capacity by 60% to 90%, depending on the testing protocol. Conclusions and Relevance In this cohort study of 785 participants, uncertainty sampling, which has not been used conventionally in public health, showed promising results for allocating testing resources. Community-based door-to-door interventions and leveraging of community knowledge were associated with reduced demographic disparities in testing.
In many public health settings, there is a perceived tension between allocating resources to known vulnerable areas and learning about the overall prevalence of the problem. Inspired by a door-to-door Covid-19 testing program we helped design, we combine multi-armed bandit strategies and insights from sampling theory to demonstrate how to recover accurate prevalence estimates while continuing to allocate resources to at-risk areas. We use the outbreak of an infectious disease as our running example. The public health setting has several characteristics distinguishing it from typical bandit settings, such as distribution shift (the true disease prevalence is changing with time) and batched sampling (multiple decisions must be made simultaneously). Nevertheless, we demonstrate that several bandit algorithms are capable out-performing greedy resource allocation strategies, which often perform worse than random allocation as they fail to notice outbreaks in new areas.
Peter Grunwald合作论文数Leiden University3