It is shown that, under usual regularity conditions, the maximum likelihood estimator of a structural parameter is strongly consistent, when the (infinitely many) incidental parameters are independently distributed chance variables with a common unknown distribution function. The latter is also consistently estimated although it is not assumed to belong to a parametric class. Application is made to several problems, in particular to the problem of estimating a straight line with both variables subject to error, which thus after all has a maximum likelihood solution.
Publisher Summary The chapter discusses treatment of the concave case. A usual treatment of the concave case is that in life testing, where α 0 (F) is assumed to be zero. The chapter also discusses the obvious modifications in the definition of C n required under certain variations. It presents a theorem based on an assumption in which X 1 , …, X n are independent chance variables with the common concave distribution function F, F n are their empiric distribution function, and C n is the least concave majorant of F n .
Let X = {1, …, a} and Y = {1, …, a} be the input and output alphabets, respectively. In the (unsynchronized) channels studied in this paper, when an element of X is sent over the channel, the receiver receives either nothing or a sequence of k letters, each a member of Y, where k, determined by chance, can be 1, or 2, or … or L, a given integer. The channel is called unsynchronized because the sequence received (for each letter sent) is not separated from the previous sequence or the following sequence, so that the receiver does not know which letters received correspond to which letter transmitted. In Sections 1 and 2 we give the necessary definitions and auxiliary results. In Section 3 we extend the results of Dobrushin [2] by proving a strong converse to the coding theorem1 and making it possible to compute the capacity to within any desired accuracy. In Section 4 we study the same channel with feedback, prove a coding theorem and strong converse, and give an algorithm for computing the capacity. In Section 5 we study the unsynchronized channel where the transmission of each word is governed by an arbitrary element of a set of channel probability functions. Again we obtain the capacity of the channel, prove a coding theorem and strong converse, and give an algorithm for computing the capacity. In Section 6 we apply results of Shannon [4] and supplement Dobrushin's results on continuous transmission with a fidelity criterion.
are, respectively, asymptotically efficient estimators of # and o-, under certain regularity conditions on G. The coefficients A~ ") and B~ ") are constants given in Section 2 below. Here "asymptotically efficient" is meant in the classical sense of minimal variance of the limiting normal distribution. Now let p and q be constants such that 0 < p < q < 1. Throughout this paper, whenever we write np and nq we always mean the largest integer in np and nq,
Let X={1,..., a} be the “input alphabet” and Y={1,2} be the “output alphabet”. Let X t =X and Y t =Y for t=1,2,..., X n =\(\mathop \prod \limits_{t = 1}^n \)X t and Y n =\(\mathop \prod \limits_{t = 1}^n \)Y t . Let S be any set, C=={w(·¦·¦)s)¦s∈S} be a set of (a×2) stochastic matrices w(·∥·¦s), and St=S, t=1,..., n. For every s n =(s1,...,s n )∈\(\mathop \prod \limits_{t = 1}^n \)S t define P(·¦·¦sn)=\(\mathop \prod \limits_{t = 1}^n \)w(y t ¦x t ¦s t ) for every xn=x1, ⋯, xnεXn and every yn=(y1, ⋯, yn)εYn. Consider the channel C n ={P(·¦·¦)s n )¦s n ∈S n } with matrices (·¦·¦s), varying arbitrarily from letter to letter. The authors determine the capacity of this channel when a) neither sender nor receiver knows sn, b) the sender knows sn, but the receiver does not, and c) the receiver knows sn, but the sender does not.
X 1,⋯,X> n are independent, identically distributed random variables with common density function f(x¦θ 1 ,⋯,θ k ,θ k+1 ), assumed to satisfy certain standard regularity conditions. The k+1 parameters are unknown, and the problem is to test the hypothesis that θ k+1 =b against the alternative that θ k+1 =b+cn −1/2 . θ 1 ,⋯,θ k are nuisance parameters. For this problem, the following artificial problem is temporarily substituted. It is known that ¦θ 1 -a i ¦≦n −1/2 M(n) for i=1,⋯,k, where a 1 , ⋯,a k are known, and M(n) approaches infinity as n increases but n −1/2 M(n) approaches zero as n increases. A Bayes decision rule is constructed for this artificial problem, relative to the a priori distribution which assigns weight A to θ k+1 =b, and weight 1-A to θ k+1 =b+cn −1/2 , in each case the weight being spread uniformly over the possible values of θ 1 ,⋯,θ k in the artificial problem. An analysis of the structure of the Bayes rule shows that if estimates of θ 1 ,...,θ k are substituted for a 1 ..., a k respectively, the resulting rule is a solution to the original problem, and this rule has the same asymptotic properties as a solution to the artificial problem as the Bayes rule for the artificial problem, no matter what the values a 1 ..., a k are.
Without Abstract
Recent results [5] of Hoel and Levine (1964), which assert that designs on $\lbrack -1, 1\rbrack$ which are optimum for certain polynomial regression extrapolation problems are supported by the "Chebyshev points," are extended to cover other nonpolynomial regression problems involving Chebyshev systems. In addition, the large class of linear parametric functions which are optimally estimated by designs supported by these Chebyshev points is characterized.
For regression problems where observations may be taken at points in a set X which does not coincide with the set Y on which the regression function is of interest, we consider the problem of finding a design (allocation of observations) which minimizes the maximum over Y of the variance function (of estimated regression). Specific examples are calculated for one-dimensional polynomial regression when Y is much smaller than or much larger than X. A related problem of optimum estimation of two regression coefficients is studied. This paper contains proofs of results first announced at the 1962 Minneapolis Meeting of the Institute of Mathematical Statistics. No prior knowledge of design theory is needed to read this paper.
exists and all the rows of Q are the same. SIA matrices are defined differently in books on probability theory; see, for example, [1] or [2]. The latter definition is more intuitive, takes longer to state, is easier to verify, and explains why the probabilist is interested in SIA matrices. A theorem in probability theory or matrix theory then says that the customary definition is equivalent to the one we have given. The latter is brief and emphasizes the property which will interest us in this note. We define S(P) by
Journal Article Enterocolic Amœbic Fistulæ Get access B L Shaff, B L Shaff Johannesburg, South Africa Search for other works by this author on: Oxford Academic Google Scholar J Wolfowitz J Wolfowitz Johannesburg, South Africa Search for other works by this author on: Oxford Academic Google Scholar British Journal of Surgery, Volume 49, Issue 217, March 1962, Pages 535–538, https://doi.org/10.1002/bjs.18004921708 Published: 08 December 2005
The purpose of this paper is to prove Theorem 1 stated in Section 1 below and Theorem 2 of Section 6 and the results of Section 7. These theorems are the generalizations to vector chance variables of Theorems 4 and 5 and Section 6 of [1], and state that the sample distribution function (d.f.) is asymptotically minimax for the large class of weight functions of the type described below. The main difficulties are embodied in the proof of Theorem 1 (Sections 2 to 5), where the loss function is a function of the maximum difference between estimated and true d.f. The proof utilizes the results of [2] and is not a straight-forward extension of the result of [1], because the sample d.f. is no longer "distribution free" (even in the limit), and hence it is necessary to prove the uniformity of approach, to its limit, of the d.f. of the normalized maximum deviation between sample and population d.f.'s (for a certain class of d.f.'s). The latter fact enables us essentially to infer the existence of a uniformly (with the sample number) approximately least favorable (to the statistician) d.f., by means of which the proof of the theorem is achieved. Theorem 2 (Section 6) considers loss functions of integral type, and more general loss functions are treated in Section 7.