E-values have recently emerged as a robust and flexible alternative to p-values for hypothesis testing, especially under optional continuation, i.e., when additional data from further experiments are collected. In this work we define optimal e-values for testing between maximum entropy models, in both the microcanonical (hard constraints) and canonical (soft constraints) settings. We show that, when testing between two hypotheses that are both microcanonical, the so-called growth-rate optimal e-variable admits an exact analytical expression, which also serves as a valid e-variable in the canonical case. For canonical tests, where exact solutions are typically unavailable, we introduce a microcanonical approximation and verify its excellent performance via both theoretical arguments and numerical simulations. We then consider constrained binary models, focusing on 2×k contingency tables-an essential framework in statistics and a natural representation for various models of complex systems. Our microcanonical optimal e-variable performs well in both settings, constituting a tool that remains effective even in the challenging case when the number k of groups grows with the sample size, as in models with growing features used for the analysis of real-world heterogeneous networks and time series.
The t-statistic is a widely-used scale-invariant statistic for testing the null hypothesis that the mean is zero. Martingale methods enable sequential testing with the t-statistic at every sample size, while controlling the probability of falsely rejecting the null. For one-sided sequential tests, which reject when the t-statistic is too positive, a natural question is whether they also control false rejection when the true mean is negative. We prove that this is the case using monotone likelihood ratios and sufficient statistics. We develop applications to the scale-invariant t-test, the location-invariant χ^2-test and sequential linear regression with nuisance covariates.
A recurring debate in the philosophy of statistics concerns what, exactly, should count as a measure of evidence for or against a given hypothesis. P-values, likelihood ratios, and Bayes factors all have their defenders. In this paper we add two additional candidates to this list: the e-value and its sequential analogue, the e-process. E-values enjoy several desirable properties as measures of evidence: they combine naturally across studies, handle composite hypotheses, provide long-run error rates, and admit a useful interpretation as the wealth accrued by a bettor in a game against the null distribution. E-processes additionally handle optional stopping and optional continuation. This work examines the extent to which e-values and e-processes satisfy the evidential desiderata of different statistical traditions, concluding that they combine attractive features of p-values, likelihood ratios, and Bayes factors, and merit serious consideration as interpretable and intuitive measures of statistical evidence.
We analyze common types of e-variables and e-processes for composite exponential family nulls: the optimal e-variable based on the reverse information projection (RIPr), the conditional (COND) e-variable, and the universal inference (UI) and sequentialized RIPr e-processes. We characterize the RIPr prior for simple and Bayes-mixture based alternatives, either precisely (for Gaussian nulls and alternatives) or in an approximate sense (general exponential families). We provide conditions under which the RIPr e-variable is (again exactly vs. approximately) equal to the COND e-variable. Based on these and other interrelations which we establish, we determine the e-power of the four e-statistics as a function of sample size, exactly for Gaussian and up to o(1) in general. For d-dimensional null and alternative, the e-power of UI tends to be smaller by a term of (d/2) log n + O(1) than that of the COND e-variable, which is the clear winner.
The validity of classical hypothesis testing requires the significance level alpha be fixed before any statistical analysis takes place. This is a stringent requirement. For instance, it prohibits updating alpha during (or after) an experiment due to changing concern about the cost of false positives, or to reflect unexpectedly strong evidence against the null. Perhaps most disturbingly, witnessing a p-value p << alpha vs p = alpha - & varepsilon; for tiny & varepsilon; > 0 has no (statistical) relevance for any downstream decision-making. Following recent work of Gr & uuml;nwald [1], we develop a theory of post-hoc hypothesis testing, enabling alpha to be chosen after seeing and analyzing the data. To study "good" post-hoc tests we introduce Gamma-admissibility, where Gamma is a set of adversaries which map the data to a significance level. We classify the set of Gamma-admissible rules for various sets Gamma, showing they must be based on e-values, and recover the Neyman-Pearson lemma when Gamma is the constant map.
We show that for any concave utility, the expected utility of an e-variable can only increase after conditioning on a sufficient statistic. The simplest form of the result has an extremely straightforward proof, which follows from a single application of Jensen's inequality. Similar statements hold for compound e-variables, asymptotic e-variables, and e-processes. These results echo the Rao-Blackwell theorem, which states that the expected squared error of an estimator can only decrease after conditioning on a sufficient statistic. We provide several applications of this insight, including a simplified derivation of the log-optimal e-variable for linear regression with known variance.
The (non)equivalence of canonical and microcanonical ensembles is a fundamental question in statistical physics, concerning whether the use of soft and hard constraints in the maximum-entropy construction leads to the same description of a system. Despite the fact that maximum-entropy models are also commonly used in statistical inference, pattern detection, and hypothesis testing, a complete understanding of the effects of ensemble nonequivalence on statistical modeling is still missing. Here, we study this problem from a rigorous model selection perspective by comparing canonical and microcanonical models via the minimum description length principle, which yields a trade-off between likelihood, measuring model accuracy, and complexity, measuring model flexibility and its potential to overfit data. We compute the normalized maximum likelihood (NML) of both formulations and find that (1) microcanonical models always achieve higher likelihood but are always more complex; (2) the optimal model choice depends on the empirical values of the constraints—the canonical model performs best when its fit to the observed data exceeds its uniform average fit across all realizations; (3) in the thermodynamic limit, the difference in description length per node vanishes when ensemble equivalence holds but persists otherwise, showing that nonequivalence implies extensive differences between large canonical and microcanonical models. Finally, we compare the NML approach to Bayesian methods, showing that (4) the choice of priors, practically irrelevant in equivalent models, becomes crucial when an extensive number of constraints are enforced, possibly leading to very different outcomes.
In this paper we review basic elements of Frequentist inference, specifically maximum likelihood (ML) and M-estimation, to point out a critical flaw of Bayesian methods for hydrologic model training and uncertainty quantification. Under model misspecification, the sensitivity matrix Âₙ and variability matrix B̂ₙ of the ML model parameter estimates θ̂ₙ provide conflicting information about the observed Fisher information Îₙ of the data ω₁,…,ωₙ for θ = (θ₁,…,θ_d)ᵀ. As a result, the estimated ML parameter covariance matrix does not simplify to the inverse of the observed Fisher information, Îₙ⁻¹, as suggested by naïve ML estimators and Bayesian methods, but instead corresponds to the so-called sandwich matrix Ĝₙ⁻¹ = n⁻¹ Âₙ⁻¹ B̂ₙ Âₙ⁻¹, where the observed Godambe information Ĝₙ is the fundamental currency of data informativeness under model misspecification. The sandwich matrix is a metaphor for a meat matrix B̂ₙ between two bread matrices Âₙ and yields asymptotically valid “robust standard errors” even when the likelihood function Lₙ(θ) is incorrectly specified. The implications of the sandwich variance estimator are demonstrated in three case studies involving modeling of soil water infiltration, watershed hydrologic fluxes, and rainfall–discharge transformation. First and foremost, our analytic and numerical results demonstrate that the sandwich variance estimator substantially increases hydrologic model parameter and predictive uncertainty. The sandwich estimator is invariant to likelihood stretching as practiced by the GLUE method as a remedy for over-conditioning and requires magnitude and/or curvature adjustments to the likelihood function to yield asymptotically valid sandwich parameter estimates and inference via Monte Carlo simulation.
The preference for simple explanations, known as the parsimony principle, has long guided the development of scientific theories, hypotheses, and models. Yet recent years have seen a number of successes in employing highly complex models for scientific inquiry (e.g., for 3D protein folding or climate forecasting). In this paper, we reexamine the parsimony principle in light of these scientific and technological advancements. We review recent developments, including the surprising benefits of modeling with more parameters than data, the increasing appreciation of the context-sensitivity of data and misspecification of scientific models, and the development of new modeling tools. By integrating these insights, we reassess the utility of parsimony as a proxy for desirable model traits, such as predictive accuracy, interpretability, effectiveness in guiding new research, and resource efficiency. We conclude that more complex models are sometimes essential for scientific progress, and discuss the ways in which parsimony and complexity can play complementary roles in scientific modeling practice.
Science is justly admired as a cumulative process (“standing on the shoulders of giants”), yet scientific knowledge is typically built on a patchwork of research contributions without much coordination. This lack of efficiency has specifically been addressed in clinical research by recommendations against avoidable research waste and for living systematic reviews and prospective meta-analysis. We propose to further those recommendations with ALL-IN meta-analysis: Anytime Live and Leading INterim meta-analysis. ALL-IN provides meta-analysis based on e-values and anytime-valid confidence intervals that can be updated at any time—reanalyzing after each new observation while retaining type-I error and coverage guarantees, live—no need to prespecify the looks, and leading—in the decisions on whether individual studies should be initiated, stopped or expanded, the meta-analysis can be the leading source of information without losing validity to accumulation bias. The analysis design requires no information about the trial sample sizes or the number of trials eventually included. So ALL-IN meta-analysis can be applied retrospectively as well as prospectively, to evaluate the evidence once or sequentially. Because the intention of the analysis does not change the validity of the results, the results of the analysis can change the intentions (‘optional stopping’ and ‘optional continuation’ based on the results so far). On the one hand: any analysis can be turned into a living one, or even become prospective and real-time by updating with new trial data and including interim data from trials that are still ongoing — without any changes in the cut-offs for testing or the method for interval estimation. On the other hand: no stopping rule needs to be enforced for the analysis to remain valid, so participating in a prospective meta-analysis does not require outside control over data collection. Hence ALL-IN meta-analysis breathes life into living systematic reviews, and offers better and simpler statistics, efficiency, collaboration and communication.
We provide a general condition under which e-variables in the form of a simple-vs.-simple likelihood ratio exist when the null hypothesis is a composite, multivariate exponential family. Such `simple' e-variables are easy to compute and expected-log-optimal with respect to any stopping time. Simple e-variables were previously only known to exist in quite specific settings, but we offer a unifying theorem on their existence for testing exponential families. We start with a simple alternative Q and a regular exponential family null. Together these induce a second exponential family Q containing Q, with the same sufficient statistic as the null. Our theorem shows that simple e-variables exist whenever the covariance matrices of Q and the null are in a certain relation. Examples in which this relation holds include some k-sample tests, Gaussian location- and scale tests, and tests for more general classes of natural exponential families.
We develop E-variables for testing whether two or more data streams come from the same source or not, and more generally, whether the difference between the sources is larger than some minimal effect size. These E-variables lead to exact, nonasymptotic tests that remain safe, i.e., keep their type-I error guarantees, under flexible sampling scenarios such as optional stopping and continuation. In special cases our E-variables also have an optimal ‘growth’ property under the alternative. While the construction is generic, we illustrate it through the special case of k×2 contingency tables, i.e. k Bernoulli streams, allowing for the incorporation of different restrictions on the composite alternative. Comparison to p-value analysis in simulations and a real-world 2 × 2 contingency table example show that E-variables, through their flexibility, often allow for early stopping of data collection — thereby retaining similar power as classical methods — while also retaining the option of extending or combining data afterwards.
A standard practice in statistical hypothesis testing is to mention the P-value alongside the accept/reject decision. We show the advantages of mentioning an e-value instead. With P-values, it is not clear how to use an extreme observation (e.g. [Formula: see text]) for getting better frequentist decisions. With e-values it is straightforward, since they provide Type-I risk control in a generalized Neyman-Pearson setting with the decision task (a general loss function) determined post hoc, after observation of the data-thereby providing a handle on "roving [Formula: see text]'s." When Type-II risks are taken into consideration, the only admissible decision rules in the post hoc setting turn out to be e-value-based. Similarly, if the loss incurred when specifying a faulty confidence interval is not fixed in advance, standard confidence intervals and distributions may fail, whereas e-confidence sets and e-posteriors still provide valid risk guarantees. Sufficiently powerful e-values have by now been developed for a range of classical testing problems. We discuss the main challenges for wider development and deployment.
We develop the theory of hypothesis testing based on the e-value, a notion of evidence that, unlike the p-value, allows for effortlessly combining results from several studies in the common scenario where the decision to perform a new study may depend on previous outcomes. Tests based on e-values are safe, i.e. they preserve Type-I error guarantees, under such optional continuation. We define growth-rate optimality (GRO) as an analogue of power in an optional continuation context, and we show how to construct GRO e-variables for general testing problems with composite null and alternative, emphasizing models with nuisance parameters. GRO e-values take the form of Bayes factors with special priors. We illustrate the theory using several classic examples including a one-sample safe t-test and the 2 x 2 contingency table. Sharing Fisherian, Neymanian and Jeffreys-Bayesian interpretations, e-values may provide a methodology acceptable to adherents of all three schools.
Information projections have found important applications in probability theory, statistics, and related areas. In the field of hypothesis testing in particular, the reverse information projection (RIPr) has recently been shown to lead to growth-rate optimal (GRO) e-statistics for testing simple alternatives against composite null hypotheses. However, the RIPr as well as the GRO criterion are undefined whenever the infimum information divergence between the null and alternative is infinite. We show that in such scenarios, under some assumptions, there still exists a measure in the null that is closest to the alternative in a specific sense. Whenever the information divergence is finite, this measure coincides with the usual RIPr. It therefore gives a natural extension of the RIPr to certain cases where the latter was previously not defined. This extended notion of the RIPr is shown to lead to optimal e-statistics in a sense that is a novel, but natural, extension of the GRO criterion. We also give conditions under which the (extension of the) RIPr is a strict sub-probability measure, as well as conditions under which an approximation of the RIPr leads to approximate e-statistics. For this case we provide tight relations between the corresponding approximation rates.
We consider growth-optimal e-variables with maximal e-power, both in an absolute and relative sense, for simple null hypotheses for a d-dimensional random vector, and multivariate composite alternatives represented as a set of d-dimensional means _1. These include, among others, the set of all distributions with mean in _1, and the exponential family generated by the null restricted to means in _1. We show how these optimal e-variables are related to Csiszár-Sanov-Chernoff bounds, first for the case that _1 is convex (these results are not new; we merely reformulate them) and then for the case that _1 `surrounds' the null hypothesis (these results are new).
We show how e-values simplify the design and the conduct of experiments. These e-values yield anytime-valid tests and confidence intervals that preserve type I error guarantee regardless of the sample size. This enables real-time monitoring of evidence as data are collected, permitting early termination of experiments without intolerably inflating the risk of making a false discovery. Early stopping not only preserves resources, but also mitigates risk for participants in clinical settings. Anytime-valid tests always allow for optional continuation, that is, the extension of an experiment regardless of the motivation. For instance, if more funds become available, or if the evidence looks promising and the funding agency, a reviewer, or an editor urges the experimenter to collect more data. Analogously, a researcher can be assured that a 95% anytime-valid confidence interval will, with at least 95% chance, cover the true effect size regardless of how, or even if, data collection is stopped. We use the free and open-source software library safestats implemented in R to illustrate the practical benefits of this novel inference framework.
We develop and compare e-variables for testing whether $k$ samples of data are drawn from the same distribution, the alternative being that they come from different elements of an exponential family. We consider the GRO (growth-rate optimal) e-variables for (1) a `small' null inside the same exponential family, and (2) a `large' nonparametric null, as well as (3) an e-variable arrived at by conditioning on the sum of the sufficient statistics. (2) and (3) are efficiently computable, and extend ideas from Turner et al. [2021] and Wald [1947] respectively from Bernoulli to general exponential families. We provide theoretical and simulation-based comparisons of these e-variables in terms of their logarithmic growth rate, and find that for small effects all four e-variables behave surprisingly similarly; for the Gaussian location and Poisson families, e-variables (1) and (3) coincide; for Bernoulli, (1) and (2) coincide; but in general, whether (2) or (3) grows faster under the alternative is family-dependent. We furthermore discuss algorithms for numerically approximating (1).
Teemu Roos合作论文数Helsinki Institute for Information Technology HIIT
Department of Computer Science
University of Helsinki12
Petri Myllymaki合作论文数Probabilistic Adaptive Systems Research Programme (PAS);Helsinki Institute for Information Technology (HIIT);Department of Computer Science;Intelligent Systems Specialisation Area;Complex Systems Computation Research Group (CoSCo);University of Helsinki11
Tomi Silander合作论文数Complex Systems Computation Group (CoSCo), P.O. Box 26, Department of Computer Science, FIN-00014 University of Helsinki, Finland5
Jorma Rissanen合作论文数Tampere University2