
It is shown that the variable bandwidth density estimator proposed by McKay (1993a and b) following earlier findings by Abramson (1982) approximates density functions in $C^4(\mathbb R^d)$ at the minimax rate in the supremum norm over bounded sets where the preliminary density estimates on which they are based are bounded away from zero. A somewhat more complicated estimator proposed by Jones McKay and Hu (1994) to approximate densities in $C^6(\mathbb R)$ is also shown to attain minimax rates in sup norm over the same kind of sets. These estimators are strict probability densities.
We discuss a number of estimates of the hazard under the assumption that the hazard is monotone on an interval [0,a]. The usual isotonic least squares estimators of the hazard are inconsistent at the boundary points 0 and a. We use penalization to obtain uniformly consistent estimators. Moreover, we determine the optimal penalization constants, extending related work in this direction by Woodroofe and Sun (1993) and Woodroofe and Sun (1999). Two methods of obtaining smooth monotone estimates based on a non-smooth monotone estimator are discussed. One is based on kernel smoothing, the other on penalization.
A problem of estimation of a large Hermitian nonnegatively definite matrix of trace 1 (a density matrix of a quantum system) motivated by quantum state tomography is studied.The estimator is based on a modified least squares method suitable in the case of models with random design with known design distributions.The bounds on Hilbert-Schmidt error of the estimator, including low rank oracle inequalities, have been proved.The proofs rely on Bernstein type inequalities for sums of independent random matrices.
If the log likelihood is approximately quadratic with constant Hessian, then the maximum likelihood estimator (MLE) is approximately normally distributed. No other assumptions are required. We do not need independent and identically distributed data. We do not need the law of large numbers (LLN) or the central limit theorem (CLT). We do not need sample size going to infinity or anything going to infinity. Presented here is a combination of Le Cam style theory involving local asymptotic normality (LAN) and local asymptotic mixed normality (LAMN) and Cramér style theory involving derivatives and Fisher information. The main tool is convergence in law of the log likelihood function and its derivatives considered as random elements of a Polish space of continuous functions with the metric of uniform convergence on compact sets. We obtain results for both one-step-Newton estimators and Newton-iterated-to-convergence estimators.
The semiparametric normal copula model is studied with a correlation matrix that depends on a covariate. The bivariate version of this regression-copula model has been proposed for statistical analysis of Quantitative Trait Loci (QTL) via twin data. Appropriate linear combinations of Van der Waerden’s normal scores rank correlation coefficients yield $\sqrt{n}$-consistent estimators of the coefficients in the correlation function, i.e. of the regression parameters. They are used to construct semiparametrically efficient estimators of the regression parameters.
We use results from modern empirical process theory to establish a uniform in bandwidth central limit theorem, laws of the iterated logarithm and Glivenko–Cantelli theorem for kernel distribution function estimators.
Developing an effective multi-stage treatment strategy over time is one of the essential goals of modern medical research. Developing statistical inference, including constructing confidence intervals for parameters, is of key interest in studies applying dynamic treatment regimens. Estimation and inference in this context are especially challenging due to non-regularity caused by the non-smoothness of the problem in the parameters. While various bootstrap methods have been proposed, there is a lack of theoretical validation for most bootstrap inference methods. Recently, Song et al. [Penalized Q-learning for dynamic treatment regimes (2011) Submitted] proposed the penalized Q-learning procedure, that enables valid inference without the need of bootstrapping. As a major drawback, penalized Q-learning can only handle discrete covariates. To overcome this issue, we propose an adaptive Q-learning procedure which is an adaptive version of penalized Q-learning. We show that the proposed method can not only handle continuous covariates, but it can also be more efficient than penalized Q-learning.
Bruno de Finetti was one of the most convinced advocates of finitely additive probabilities. The present work describes the intellectual pro- cess that led him to support that stance and provides a detailed account both of the first paper by de Finetti on the subject and of the ensuing correspondence with Maurice Fréchet. Moreover, the analysis is supplemented by a useful picture of de Finetti's interactions with the international scientific community at that time, when he elaborated his subjectivistic conception of probability.
: A scoring function is consistent for the α -quantile functional if, and only if, it is generalized piecewise linear (GPL) of order α , up to equivalence. Expressed differently, loss functions that yield quantiles as Bayes rules are GPL functions. We review and discuss this basic decision-theoretic result with focus on Thomson’s pioneering characterization.
Large-scale multiple testing problems require the simultaneous assessment of many p-values. This paper compares several methods to assess the evidence in multiple binomial counts of p-values: the maximum of the binomial counts after standardization (the `higher-criticism statistic'), the maximum of the binomial counts after a log-likelihood ratio transformation (the `Berk-Jones statistic'), and a newly introduced average of the binomial counts after a likelihood ratio transformation. Simulations show that the higher criticism statistic has a superior performance to the Berk-Jones statistic in the case of very sparse alternatives (sparsity coefficient $\beta \gtrapprox 0.75$), while the situation is reversed for $\beta \lessapprox 0.75$. The average likelihood ratio is found to combine the favorable performance of higher criticism in the very sparse case with that of the Berk-Jones statistic in the less sparse case and thus appears to dominate both statistics. Some asymptotic optimality theory is considered but found to set in too slowly to illuminate the above findings, at least for sample sizes up to one million. In contrast, asymptotic approximations to the critical values of the Berk-Jones statistic that have been developed by Wellner and Koltchinskii (2003) and Jager and Wellner (2007) are found to give surprisingly accurate approximations even for quite small sample sizes.
We consider the regression model with observation error in the design: y=Xθ* + e, Z=X+N. Here the random vector y in R^n and the random n*p matrix Z are observed, the n*p matrix X is unknown, N is an n*p random noise matrix, e in R^n is a random noise vector, and θ* is a vector of unknown parameters to be estimated. We consider the setting where the dimension p can be much larger than the sample size n and θ* is sparse. Because of the presence of the noise matrix N, the commonly used Lasso and Dantzig selector are unstable. An alternative procedure called the Matrix Uncertainty (MU) selector has been proposed in Rosenbaum and Tsybakov (2010) in order to account for the noise. The properties of the MU selector have been studied in Rosenbaum and Tsybakov (2010) for sparse θ* under the assumption that the noise matrix N is deterministic and its values are small. In this paper, we propose a modification of the MU selector when N is a random matrix with zero-mean entries having the variances that can be estimated. This is, for example, the case in the model where the entries of X are missing at random. We show both theoretically and numerically that, under these conditions, the new estimator called the Compensated MU selector achieves better accuracy of estimation than the original MU selector.
Let $P$ be a probability distribution on $q$-dimensional space. The so-called Diaconis-Freedman effect means that for a fixed dimension $d << q$, most $d$-dimensional projections of $P$ look like a scale mixture of spherically symmetric Gaussian distributions. The present paper provides necessary and sufficient conditions for this phenomenon in a suitable asymptotic framework with increasing dimension $q$. It turns out, that the conditions formulated by Diaconis and Freedman (1984) are not only sufficient but necessary as well. Moreover, letting $\hat{P}$ be the empirical distribution of $n$ independent random vectors with distribution $P$, we investigate the behavior of the empirical process $\sqrt{n}(\hat{P} - P)$ under random projections, conditional on $\hat{P}$.
This paper introduces and analyzes a stochastic search method for parameter estimation in linear regression models in the spirit of Beran and Millar (1987). The idea is to generate a random finite subset of a parameter space which will automatically contain points which are very close to an unknown true parameter. The motivation for this procedure comes from recent work of Duembgen, Samworth and Schuhmacher (2011) on regression models with log-concave error distributions.
We present a construction showing that a class of sets $\mathcal{C}$ that is Glivenko-Cantelli for an i.i.d. process need not be Glivenko-Cantelli for every stationary ergodic process with the same one dimensional marginal distribution. This result provides a counterpoint to recent work extending uniform strong laws to ergodic processes, and a recent characterization of universal Glivenko Cantelli classes.
For a bivariate random vector (X,Y), symmetry conditions are presented that yield stochastic orderings among |X|, |Y|, |max(X,Y)|, and | min(X, Y)|. Partial extensions of these results for multivariate random vectors (X1,...,Xn) are also given.
In this paper we formulate a corporate bond (CB) pricing model for deriving the term structure of default probabilities (TSDP) and the recovery rate (RR) for each pair of industry factor and credit rating grade, and these derived TSDP and RR are regarded as what investors imply in forming CB prices in the market at each time. A unique feature of this formulation is that the model allows each firm to run several business lines corresponding to some industry categories, which is typical in reality. In fact, treating all the cross-sectional CB prices simultaneously under a credit correlation structure at each time makes it possible to sort out the overlapping business lines of the firms which issued CBs and to extract the TSDPs for each pair of individual industry factor and rating grade together with the RRs. The result is applied to a valuation of CDS (credit default swap) and a loan portfolio management in banking business.
The general asymptotic distribution theory for the functional regression model in Ruymgaart et al. [Some asymptotic theory for functional regression and classification (2009) Texas Tech University] simplifies considerably if an extra assumption on the random regressor is made. In the special case where the regressor is a stochastic process on the unit interval, Johannes [Privileged communication (2008)] assumes the regressor to be stationary, in which case the eigenfunctions of their covariance operator turn out to be known, so that only the eigenvalues are to be estimated. In the present paper we will also assume the eigenvectors to be known, but within an abstract setting. The simplification mentioned above is due to the circumstance that the covariance operator of the regressor commutes with its estimator as it can be constructed under the current conditions. Moreover, it is now possible to test linear hypotheses for the regression parameter that correspond to linear subspaces spanned by a finite number of the known eigenvectors.
Nemirovski's inequality states that given independent and centered at expectation random vectors $X_{1},\ldots,X_{n}$ with values in $\ell^p(\mathbb{R}^d)$, there exists some constant $C(p,d)$ such that \[\mathbb{E}\Vert S_n\Vert _p^2\le C(p,d)\sum_{i=1}^{n}\mathbb{E}\Vert X_i\Vert _p^2.\] Furthermore $C(p,d)$ can be taken as $\kappa(p\wedge \log(d))$. Two cases were studied further in [ Am. Math. Mon. 117(2) (2010) 138–160]: general finite-dimensional Banach spaces and the special case $\ell^{\infty}(\mathbb{R}^{d})$. We show that in these two cases, it is possible to replace the quantity $\sum_{i=1}^n\mathbb{E}\Vert X_i\Vert _p^2$ by a smaller one without changing the order of magnitude of the constant when $d$ becomes large. In the spirit of [ Am. Math. Mon. 117(2) (2010) 138–160], our approach is probabilistic. The derivation of our version of Nemirovski's inequality indeed relies on concentration inequalities.