The problem of comparing the entire second order structure of two functional processes is considered and a L2-type statistic for testing equality of the corresponding spectral density operators is investigated. The test statistic evaluates, over all frequencies, the Hilbert–Schmidt distance between the two estimated spectral density operators. Under certain assumptions, the limiting distribution under the null hypothesis is derived. A novel frequency domain bootstrap method is introduced, which leads to a more accurate approximation of the distribution of the test statistic under the null than the large sample Gaussian approximation derived. Under quite general conditions, asymptotic validity of the bootstrap procedure is established for estimating the distribution of the test statistic under the null. Furthermore, consistency of the bootstrap-based test under the alternative is proved. Numerical simulations show that, even for small samples, the bootstrap-based test has a very good size and power behavior. An application to a bivariate real-life functional time series illustrates the methodology proposed.
We consider strictly stationary stochastic processes of Hilbert space‐valued random variables and focus on fully functional tests for the equality of the lag‐zero autocovariance operators of several independent functional time series. A moving block bootstrap (MBB)‐based testing procedure is proposed which generates pseudo random elements that satisfy the null hypothesis of interest. It is based on directly bootstrapping the time series of tensor products which overcomessome common difficulties associated with applications of the bootstrap to related testing problems. The suggested methodology can be potentially applied to a broad range of test statistics of the hypotheses of interest. As an example, we establish validity for approximating the distribution under the null of a test statistic based on the Hilbert–Schmidt distance of the corresponding sample lag‐zero autocovariance operators, and show consistency under the alternative. As a prerequisite, we prove a central limit theorem for the MBB procedure applied to the sample autocovariance operator which is of interest on its own. The finite sample size and power performance of the suggested MBB‐based testing procedure is illustrated through simulations and an application to a real‐life dataset is discussed.
We consider infinite-dimensional Hilbert space-valued random variables that are assumed to be temporal dependent in a broad sense. We prove a central limit theorem for the moving block bootstrap and for the tapered block bootstrap, and show that these block bootstrap procedures also provide consistent estimators of the long run covariance operator. Furthermore, we consider block bootstrap-based procedures for fully functional testing of the equality of mean functions between several independent functional time series. We establish validity of the block bootstrap methods in approximating the distribution of the statistic of interest under the null and show consistency of the block bootstrap-based tests under the alternative. The finite sample behaviour of the procedures is investigated by means of simulations. An application to a real-life dataset is also discussed.
We consider a Gaussian sequence model that contains ill-posed inverse problems as special cases. We assume that the associated operator is partially unknown in the sense that its singular functions are known and the corresponding singular values are unknown but observed with Gaussian noise. For the considered model, we study the minimax goodness-of-fit testing problem. Working with certain ellipsoids in the space of squared-summable sequences of real numbers, with a ball of positive radius removed, we obtain lower and upper bounds for the minimax separation radius in the non-asymptotic framework, i.e., for fixed values of the involved noise levels. Examples of mildly and severely ill-posed inverse problems with ellipsoids of ordinary-smooth and super-smooth sequences are examined in detail and minimax rates of goodness-of-fit testing are obtained for illustrative purposes.
We investigate the properties of a simple bootstrap method for testing the equality of mean functions or of covariance operators in functional data. Theoretical size and power results are derived for certain test statistics, whose limiting distributions depend on unknown infinite-dimensional parameters. Simulations demonstrate good size and power of the bootstrap-based functional tests.
We consider multichannel deconvolution in a periodic setting with long-memory errors under three different scenarios for the convolution operators, i.e., super-smooth, regular-smooth and box-car convolutions. We investigate global performances of linear and hard-thresholded non-linear wavelet estimators for functions over a wide range of Besov spaces and for a variety of loss functions defining the risk. In particular, we obtain upper bounds on convergence rates using the Lp-risk (1≤p<∞). Contrary to the case where the errors follow independent Brownian motions, it is demonstrated that multichannel deconvolution with errors that follow independent fractional Brownian motions with different Hurst parameters results in a much more involved situation. An extensive finite-sample numerical study is performed to supplement the theoretical findings.
We are concerned with minimax signal detection. In this setting, we discuss non-asymptotic and asymptotic approaches through a unified treatment. In particular, we consider a Gaussian sequence model that contains classical models as special cases, such as, direct, well-posed inverse and ill-posed inverse problems. Working with certain ellipsoids in the space of squared-summable sequences of real numbers, with a ball of positive radius removed, we compare the construction of lower and upper bounds for the minimax separation radius (non-asymptotic approach) and the minimax separation rate (asymptotic approach) that have been proposed in the literature. Some additional contributions, bringing into light links between non-asymptotic and asymptotic approaches to minimax signal, are also presented. An example of a mildly ill-posed inverse problem is used for illustrative purposes. In particular, it is shown that tools used to derive `asymptotic' results can be exploited to draw `non-asymptotic' conclusions, and vice-versa.
We consider the problem of estimating the unknown response function in the multichannel deconvolution model with long-range dependent Gaussian errors. We do not limit our consideration to a specific type of long-range dependence rather we assume that the errors should satisfy a general assumption in terms of the smallest and larger eigenvalues of their covariance matrices. We derive minimax lower bounds for the quadratic risk in the proposed multichannel deconvolution model when the response function is assumed to belong to a Besov ball and the blurring function is assumed to possess some smoothness properties, including both regular-smooth and super-smooth convolutions. Furthermore, we propose an adaptive wavelet estimator of the response function that is asymptotically optimal (in the minimax sense), or near-optimal within a logarithmic factor, in a wide range of Besov balls. It is shown that the optimal convergence rates depend on the balance between the smoothness parameter of the response function, the kernel parameters of the blurring function, the long memory parameters of the errors, and how the total number of observations is distributed among the total number of channels. Some examples of inverse problems in mathematical physics where one needs to recover initial or boundary conditions on the basis of observations from a noisy solution of a partial differential equation are used to illustrate the application of the theory we developed. The optimal convergence rates and the adaptive estimators we consider extend the ones studied by Pensky and Sapatinas (2009, 2010) for independent and identically distributed Gaussian errors to the case of long-range dependent Gaussian errors.
We investigate properties of a bootstrap-based methodology for testing hypotheses about equality of certain characteristics of the distributions between different populations in the context of functional data. The suggested testing methodology is simple and easy to implement. It resamples the original dataset in such a way that the null hypothesis of interest is satisfied and it can be potentially applied to a wide range of testing problems and test statistics of interest. Furthermore, it can be utilized to the case where more than two populations of functional data are considered. We illustrate the bootstrap procedure by considering the important problems of testing the equality of mean functions or the equality of covariance functions (resp. covariance operators) between two populations. Theoretical results that justify the validity of the suggested bootstrap-based procedure are established. Furthermore, simulation results demonstrate very good size and power performances in finite sample situations, including the case of testing problems and/or sample sizes where asymptotic considerations do not lead to satisfactory approximations. A real-life dataset analyzed in the literature is also examined.
We consider the estimation of a density function on the basis of a random sample from a weighted distribution. We propose linear and nonlinear wavelet density estimators, and provide their asymptotic formulae for mean integrated squared error. In particular, we derive an analogue of the asymptotic formula of the mean integrated square error in the context of kernel density estimators for weighted data, admitting an expansion with distinct squared bias and variance components. For nonlinear wavelet density estimators, unlike the analogous situation for kernel or linear wavelet density estimators, this asymptotic formula of the mean integrated square error is relatively unaffected by assumptions of continuity, and it is available for densities which are smooth only in a piecewise sense. We illustrate the behavior of the proposed linear and nonlinear wavelet density estimators in finite sample situations both in simulations and on a real-life dataset. Comparisons with a kernel density estimator are also given.
We consider the nonparametric regression estimation problem of recovering an unknown response function f on the basis of spatially inhomogeneous data when the design points follow a known compactly supported density g with a finite number of well separated zeros. In particular, we consider two different cases: when g has zeros of a polynomial order and when g has zeros of an exponential order. These two cases correspond to moderate and severe data losses, respectively. We obtain asymptotic minimax lower bounds for the global risk of an estimator of f and construct adaptive wavelet nonlinear thresholding estimators of f which attain those minimax convergence rates (up to a logarithmic factor in the case of a zero of a polynomial order), over a wide range of Besov balls. The spatially inhomogeneous ill-posed problem that we investigate is inherently more difficult than spatially homogeneous problems like, e.g., deconvolution. In particular, due to spatial irregularity, assessment of minimax global convergence rates is a much harder task than the derivation of minimax local convergence rates studied recently in the literature. Furthermore, the resulting estimators exhibit very different behavior and minimax global convergence rates in comparison with the solution of spatially homogeneous ill-posed problems. For example, unlike in deconvolution problem, the minimax global convergence rates are greatly influenced not only by the extent of data loss but also by the degree of spatial homogeneity of f. Specifically, even if 1/g is not integrable, one can recover f as well as in the case of an equispaced design (in terms of minimax global convergence rates) when it is homogeneous enough since the estimator is borrowing strength in the areas where f is adequately sampled.
A novel functional time-series methodology for short-term load forecasting is introduced. The prediction is performed by means of a weighted average of past daily load segments, the shape of which is similar to the expected shape of the load segment to be predicted. The past load segments are identified from the available history of the observed load segments by means of their closeness to a so-called reference load segment. The latter is selected in a manner that captures the expected qualitative and quantitative characteristics of the load segment to be predicted. As an illustration, the suggested functional time-series forecasting methodology is applied to historical daily load data in Cyprus. Its performance is compared with some recently proposed alternative methodologies for short-term load forecasting.
Statistics for high-dimensional data, by Peter Buhlmann and Sara van de Geer, Berlin, Springer-Verlag, 2011, xvii + 556 pp., £81.00 or US$99.00 (hardback), ISBN 978-3-642-20191-2 During the last fe...
Ill-posed inverse problems arise in various scientific fields. We consider the signal detection problem for mildly, severely and extremely ill-posed inverse problems with $l^q$-ellipsoids (bodies), $q\in(0,2]$, for Sobolev, analytic and generalized analytic classes of functions under the Gaussian white noise model. We study both rate and sharp asymptotics for the error probabilities in the minimax setup. By construction, the derived tests are, often, nonadaptive. Minimax rate-optimal adaptive tests of rather simple structure are also constructed.
We consider the nonparametric estimation problem of time-dependent multivariate functions observed in a presence of additive cylindrical Gaussian white noise of a small intensity. We derive minimax lower bounds for the $L^2$-risk in the proposed spatio-temporal model as the intensity goes to zero, when the underlying unknown response function is assumed to belong to a ball of appropriately constructed inhomogeneous time-dependent multivariate functions, motivated by practical applications. Furthermore, we propose both non-adaptive linear and adaptive non-linear wavelet estimators that are asymptotically optimal (in the minimax sense) in a wide range of the so-constructed balls of inhomogeneous time-dependent multivariate functions. The usefulness of the suggested adaptive nonlinear wavelet estimator is illustrated with the help of simulated and real-data examples.
We consider the problem of estimating the unknown response function in the multichannel deconvolution model with a boxcar-like kernel which is of particular interest in signal processing. It is known that, when the number of channels is finite, the precision of reconstruction of the response function increases as the number of channels $M$ grow (even when the total number of observations $n$ for all channels $M$ remains constant) and this requires that the parameter of the channels form a Badly Approximable $M$-tuple. Recent advances in data collection and recording techniques made it of urgent interest to study the case when the number of channels $M=M_n$ grow with the total number of observations $n$. However, in real-life situations, the number of channels $M = M_n$ usually refers to the number of physical devices and, consequently, may grow to infinity only at a slow rate as $n \rightarrow \infty$. When $M=M_n$ grows slowly as $n$ increases, we develop a procedure for the construction of a Badly Approximable $M$-tuple on a specified interval, of a non-asymptotic length, together with a lower bound associated with this $M$-tuple, which explicitly shows its dependence on $M$ as $M$ is growing. This result is further used for the evaluation of the $L^2$-risk of the suggested adaptive wavelet thresholding estimator of the unknown response function and, furthermore, for the choice of the optimal number of channels $M$ which minimizes the $L^2$-risk.
We consider the detection problem of a two-dimensional function from noisy observations of its integrals over lines. We study both rate and sharp asymptotics for the error probabilities in the minimax setup. By construction, the derived tests are non-adaptive. We also construct a minimax rate-optimal adaptive test of rather simple structure.
Using an approach based, amongst other things, on Proposition 1 of Kaluza (1928), Goldie (1967) and, using a different approach based especially on zeros of polynomials, Steutel (1967) have proved that each nondegenerate distribution function (d.f.) $F$ (on $\mathbb{R}$, the real line), satisfying $F(0-)=0$ and $F(x)=F(0)+(1-F(0))G(x), x > 0$, where $G$ is the d.f. corresponding to a mixture of exponential distributions, is infinitely divisible. Indeed, Proposition 1 of Kaluza (1928) implies that any nondegenerate discrete probability distribution $\{p_x:x=0,1,\ldots\}$ that is log-convex or, in particular, completely monotone, is compound geometric, and, hence, infinitely divisible. Steutel (1970), Shanbhag & Sreehari (1977) and Steutel & van Harn (2004, Chapter VI) have given certain extensions or variations of one or more of these results. Following a modified version of the C.R. Rao et al. (2009, Section 4) approach based on the Wiener-Hopf factorization, we establish some further results of significance to the literature on infinite divisibility.
We consider the problem of estimating the unknown response function and its derivatives in the standard nonparametric regression model. Recently, Abramovich et al. (2010) applied a Bayesian testimation procedure in a wavelet context and proved asymptotical minimaxity of the resulting adaptive level-wise maximum a posteriori wavelet testimator of the unknown response function and its derivatives in the Gaussian white noise model. Using the boundary-modified coiflets of Johnstone and Silverman (2004), we show that dicretization of the data does not affect the order of magnitude of the accuracy of a discrete version of the suggested level-wise maximum a posteriori wavelet testimator, obtaining thus its adaptivity and asymptotical minimaxity in the standard nonparametric regression model that is usually considered in practical applications. Simulated examples are used to illustrate the performance of the developed wavelet testimation procedure and compared with three recently proposed empirical Bayes wavelet estimators and a block thresholding wavelet estimator.