We use location model methodology to guide the least squares analysis of the Lasso problem of variable selection and inference. The nuisance parameter is taken to be an indicator for the selection of explanatory variables and the interest parameter is the response variable itself. Recent theory eliminates the nuisance parameter by marginalization on the data space and then uses the resulting distribution for inference concerning the interest parameter. We develop this approach and find: that primary inference is essentially one-dimensional rather than $n$-dimensional; that inference focuses on the response variable itself rather than the least squares estimate (as variables are removed); that first order probabilities are available; that computation is relatively easy; that a scalar marginal model is available; and that ineffective variables can be removed by distributional tilt or shift.
This article has two objectives. The first and narrower is to formalize the p-value function, which records all possible p-values, each corresponding to a value for whatever the scalar parameter of interest is for the problem at hand, and to show how this p-value function directly provides full inference information for any corresponding user or scientist. The p-value function provides familiar inference objects: significance levels, confidence intervals, critical values for fixed-level tests, and the power function at all values of the parameter of interest. It thus gives an immediate accurate and visual summary of inference information for the parameter of interest. We show that the p-value function of the key scalar interest parameter records the statistical position of the observed data relative to that parameter, and we the describe an accurate approximation to that p-value function which is readily constructed.
The need to combine likelihood information is common in analyses of complex models and in meta-analyses, where information is combined from several studies. We work to first order, and show that full accuracy when combining scalar or vector parameter information is available from an asymptotic analysis of the score variables. Then we use this approach to combine p-values for scalar parameters of interest.
The foundations of statistics have evolved over many centuries, perhaps millennia, with major paradigm shifts of the form described in Kuhn (1962). We briefly consider these important transitions and how they have led to major shifts in the foundations of statistical inference. Clearly there is no conventional mathematical or axiomatic basis. But there is a progressive clarification in the processes of statistical inference so that current theory can now coherently and definitively handle a wide range of inference problems.
At a recent conference on Bayes, fiducial and frequentist inference, David Cox presented eight illustrative examples, chosen to highlight potential difficulties for the theory of inference. We discuss these examples in light of the efforts of the conference, and related meetings, to study the similarities and differences between the approaches to inference. Emphasis is placed on the goal of finding a distribution for an unknown parameter.
1.1. Exponential and general models. For an exponential model, suppose φ and s are the canonical parameter and canonical variable, with {1, φ1(θ), . . . , φp(θ)} linearly independent. . If not an exponential model, a third order approximation is available using φ′(θ) = (∂/∂V )`(θ; y)|y0 where V = (v1, . . . , yp) are inference directions at y. If these directions V are not obvious, they can be extracted by differentiating a full quantile function at the observed data: V = (∂/∂θ)q(θ; y)|θ̂0,y0 ; see Fraser and Reid [3] 1.2. Density for canonical variable. The density for the canonical variable s can be presented in saddlepoint format to third order as
I introduce a p-value function that derives from the continuity inherent in a wide range of regular statistical models. This provides confidence bounds and confidence sets, tests, and estimates that all reflect model continuity. The development starts with the scalar-variable scalar-parameter exponential model and extends to the vector-parameter model with scalar interest parameter, then to general regular models, and then references for testing vector interest parameters are available. The procedure does not use sufficiency but applies directly to general models, although it reproduces sufficiency-based results when sufficiency is present. The emphasis is on the coherence of the full procedure, and technical details are not emphasized.
For a scalar or vector parameter of interest with a regular statistical model, we determine the definitive null density for testing a particular value of the interest parameter: continuity gives uniqueness without reference to sufficiency but the use of full available information is presumed. We start with an exponential family model, that may be either the original model or an approximation to it obtained by ancillary conditioning. If the parameter of interest is linear in the canonical parameter, then the null density is third order equivalent to the conditional density given the nuisance parameter score; and when the parameter of interest is also scalar then this conditional density is the familiar density used to construct unbiased tests. More generally but with scalar parameter of interest, linear or curved, this null density has distribution function that is third order equivalent to the familiar higher-order p-value Φ(r∗). Connections to the bootstrap are described: the continuity-based ancillary of the null density is the natural invariant of the bootstrap procedure. Also ancillarity provides a widely available general replacement for the sufficiency reduction. Illustrative examples are recorded and various further examples are available in Davison et al. (2014) and Sartori et al. (2015).
We consider the use of default priors in the Bayes methodology for seeking information concerning the true value of a parameter. By default prior, we mean the mathematical prior as initiated by Bayes [Philos. Trans. R. Soc. Lond. 53 (1763) 370–418] and pursued by Laplace [Theorie Analytique des Probabilites (1812) Courcier], Jeffreys [Theory of Probability (1961) Clarendon Press], Bernardo [J. Roy. Statist. Soc. Ser. B 41 (1979) 113–147] and many more, and then recently viewed as “potentially dangerous” [Science 340 (2013) 1177–1178] and “potentially useful” [Science 341 (2013) 1452]. We do not mean, however, the genuine prior [Science 340 (2013) 1177–1178] that has an empirical reference and would invoke standard frequency modelling. And we do not mean the subjective or opinion prior that an individual might have and would be viewed as specific to that individual. A mathematical prior has no referenced frequency information, but on occasion is known otherwise to lead to repetition properties called confidence. We investigate the presence of such supportive property, and ask can Bayes give reliability for other than the particular parameter weightings chosen for the conditional calculation. Thus, does the methodology have reproducibility? Or is it a leap of faith. For sample-space analysis, recent higher-order likelihood methods with regular models show that third-order accuracy is widely available using profile contours [In Past, Present and Future of Statistical Science (2014) 237– 252 CRC Press]. But for parameter-space analysis, accuracy is widely limited to first order. An exception arises with a scalar full parameter and the use of the scalar Jeffreys [J. Roy. Statist. Soc. Ser. B 25 (1963) 318–329]. But for vector full parameter even with a scalar interest parameter, difficulties have long been known [J. Roy. Statist. Soc. Ser. B 35 (1973) 189–233] and with parameter curvature, accuracy beyond first order can be unavailable [Statist. Sci. 26 (2011) 299–316].We show, however, that calculations on the parameter space can give full second-order information for a chosen scalar interest parameter; these calculations, however, require a Jeffreys prior that is used fully restricted to the one-dimensional profile for that interest parameter. Such a prior is effectively data-dependent and parameter-dependent and is focally restricted to the one-dimensional contour; these priors fall outside the usual Bayes approach and yet with substantial calculations can still give less than frequency analysis. We provide simple examples using discrete extensions of Jeffreys prior. These serve as counter-examples to general claims that Bayes can offer accuracy for statistical inference. To obtain this accuracy with Bayes, more effort is required compared to recent likelihood methods, which still remain more accurate. And with vector full parameters, accuracy beyond first order is routinely not available, as a change in parameter curvature causes Bayes and frequentist values to change in opposite direction, yet frequentist has full reproducibility. An alternative is to view default Bayes as an exploratory technique and then ask does it do as it overtly claims? Is it reproducible as understood in contemporary science? The posterior gives a distribution for an interest parameter and, thereby, a quantile for the interest parameter; an oracle could record whether it was left or right of the true value. If the average split in evaluative repetitions is in accord with the nominal level, then the approach is providing accuracy. And if not, then what is up, other than performance specific to the parameter frequencies in the prior. No one has answers although speculative claims abound.
We consider statistical inference for a vector-valued parameter of interest in a regular asymptotic model with a finite-dimensional nuisance parameter. We use highly accurate likelihood theory to derive a directional test, in which the p-value is obtained by one-dimensional numerical integration. This extends the results of Davison et al. (2014) for linear exponential families to nonlinear parameters of interest and to more general models. Examples and simulations provide comparisons with the likelihood ratio test and adjusted versions of the likelihood ratio test. The directional approach gives extremely accurate inference, even in high-dimensional settings where the likelihood ratio versions can fail catastrophically.
Fiducial inference provides probabilities for parameters on the basis of a statistical model and data; structural inference also provides such probabilities but requires a specialized transformation structure in the model. A comparison of the fiducial and structural methodologies with default Bayesian analysis is given and shows that both can be viewed as specializations of the Bayesian. Examples are given and comparisons made.
We consider default priors for Bayes analysis as initiated in Bayes (1763), then Laplace (1812), Jeffreys (1961), Bernardo (1979), and many more, and viewed recently as “potentially dangerous” (Efron, 2013) or potentially useful (Fraser, 2013). We use nominal data size n to develop an explicit new prior that has second-order O(n−1) reproducibility for any specified parameter of interest. As part of this we see that untargetted Bayes inference is routinely just first-order O(n−1/2), and thus comparable to maximum likelihood and signed-likelihood-root methods using Central Limit Theorem Normality. A related result is that such mathematical priors do not give valid probability statements by means of the conditional probability lemma, as necessarily they are just mathematical, without physical reference. But they can acquire properties based on repetition or reproducibility as proposed later by Fisher (1930) but implicitly in Laplace (1812). This of course means confidence which in turn was not otherwise developed at the time of Bayes or Laplace.
The public image of statistics is changing, and recently the changes have been mostly for the better, as we've all seen. But occasional court cases, a few conspicuous failures, and even appeals to personal feelings suggest that careful thought may be in order. Actually, statistics itself has more than one theory, and these approaches can give contradictory answers, with the discipline largely indifferent. Saying "we are just exploring!" or appealing to mysticism can't really be appropriate, no matter the spin. In this paper for the COPSS 50th Anniversary Volume, I would like to examine three current approaches to central theory. As we will see, if continuity that is present in the model is also required for the methods, then the conflicts and contradictions resolve.
Discussion of "On the Birnbaum Argument for the Strong Likelihood Principle" by Deborah G. Mayo [arXiv:1302.7021].
Ancillary Statistics, First Derivative† D. A. S. Fraser, D. A. S. FraserSearch for more papers by this authorN. Reid, N. ReidSearch for more papers by this author D. A. S. Fraser, D. A. S. FraserSearch for more papers by this authorN. Reid, N. ReidSearch for more papers by this author First published: 29 September 2014 https://doi.org/10.1002/9781118445112.stat01491 †This article was originally published online in 2006 in Encyclopedia of Statistical Sciences, © John Wiley & Sons, Inc. and republished in Wiley StatsRef: Statistics Reference Online, 2014. Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat No abstract is available for this article. Wiley StatsRef: Statistics Reference OnlineBrowse other articles of this reference work:BROWSE BY TOPICBROWSE A-Z RelatedInformation