We propose in this work to derive a CLT in the functional linear regression model. The main difficulty is due to the fact that estimation of the functional parameter leads to a kind of ill-posed inverse problem. We consider estimators that belong to a large class of regularizing methods and we first show that, contrary to the multivariate case, it is not possible to state a CLT in the topology of the considered functional space. However, we show that we can get a CLT for the weak topology under mild hypotheses and in particular without assuming any strong assumptions on the decay of the eigenvalues of the covariance operator. Rates of convergence depend on the smoothness of the functional coefficient and on the point in which the prediction is made.
This article presents a selected bibliography on functional linear regression (FLR) and highlights the key contributions from both applied and theoretical points of view. It first defines FLR in the case of a scalar response and shows how its modelization can also be extended to the case of a functional response. It then considers two kinds of estimation procedures for this slope parameter: projection-based estimators in which regularization is performed through dimension reduction, such as functional principal component regression, and penalized least squares estimators that take into account a penalized least squares minimization problem. The article proceeds by discussing the main asymptotic properties separating results on mean square prediction error and results on L2 estimation error. It also describes some related models, including generalized functional linear models and FLR on quantiles, and concludes with a complementary bibliography and some open problems.
The paper considers functional linear regression, where scalar responses Y1,..., Yn are modeled in dependence of i.i.d. random functions X1,..., Xn. We study a generalization of the classical functional linear regression model. It is assumed that there exists an unknown number of "points of impact," that is, discrete observation times where the corresponding functional values possess significant influences on the response variable. In addition to estimating a functional slope parameter, the problem then is to determine the number and locations of points of impact as well as corresponding regression coefficients. Identifiability of the generalized model is considered in detail. It is shown that points of impact are identifiable if the underlying process generating X1,..., Xn possesses " specific local variation." Examples are well-known processes like the Brownian motion, fractional Brownian motion or the Ornstein-Uhlenbeck process. The paper then proposes an easily implementable method for estimating the number and locations of points of impact. It is shown that this number can be estimated consistently. Furthermore, rates of convergence for location estimates, regression coefficients and the slope parameter are derived. Finally, some simulation results as well as a real data application are presented.
In this paper we introduce a modified version of the BUS test, which we call NBUS (New Borovkov–Utev Statistic). This latter defines a family of goodness of fit tests that can be used to detect normality against alternative hypothesis of which all moments up to the fifth exist. The test statistic depends on empirical moments and real parameters that have to be chosen appropriately. The good abilities of the NBUS with respect to BUS and other powerful normality tests are illustrated by means of a Monte Carlo experiment for finite samples. Besides, we show how an adaptation of NBUS for testing departing from normality due only to kurtosis, leads to comparable performances with classical tests based on the fourth moment.
An important field of investigation in Geography is the modelization of the evolution of land cover in view of analyzing the dynamics of this evolution and then to build predictive maps. This is possible with the apparatus of measure : sattelite image... In this paper, we propose to use a polychotomous regression model to modelize and to predict land cover of a given area : we shox how to adapt this model in order to take into account the spatial correlation and the temporal evolution of the vegetation indexes. This study concerns an area in the Pyrenees mountains.
This paper focuses on combining association measures using corresponding receiver operating characteristic curves. The approach is motivated by a problem of automatic bigram collocation extraction from the field of computational linguistics. It is based on supervised machine learning techniques and the fact that different association measures discover different collocation types. Clusters of equivalent ROC curves are first determined by a testing procedure. The paper’s major contribution is an investigation of the possibility of combining representatives of the clusters of equivalent association measures into more complex models, thus improving performance of the collocation extraction.
The paper considers linear regression problems where the number of predictor variables is possibly larger than the sample size. The basic motivation of the study is to combine the points of view of model selection and functional regression by using a factor approach: it is assumed that the predictor vector can be decomposed into a sum of two uncorrelated random components reflecting common factors and specific variabilities of the explanatory variables. It is shown that the traditional assumption of a sparse vector of parameters is restrictive in this context. Common factors may possess a significant influence on the response variable which cannot be captured by the specific effects of a small number of individual variables. We therefore propose to include principal components as additional explanatory variables in an augmented regression model. We give finite sample inequalities for estimates of these components. It is then shown that model selection procedures can be used to estimate the parameters of the augmented model, and we derive theoretical properties of the estimators. Finite sample performance is illustrated by a simulation study.
A new test of normality based on Poincaré inequality is proposed and analyzed. It rests on the characterization of the normal distribution given by Borovkov and Utev, i.e., a r.v. is normal if and only if its Poincaré constant is equal to its variance.
A functional linear regression model linking observations of a functional response variable with measurements of an explanatory functional variable is considered. This model serves to analyse a real data set describing electricity consumption in Sardinia. The interest lies in predicting either oncoming weekends' or oncoming weekdays' consumption, provided actual weekdays' consumption is known. A B-spline estimator of the functional parameter is used. Selected computational issues are addressed as well.
The paper considers functional linear regression, where scalar responses $Y_1,...,Y_n$ are modeled in dependence of random functions $X_1,...,X_n$. We propose a smoothing splines estimator for the functional slope parameter based on a slight modification of the usual penalty. Theoretical analysis concentrates on the error in an out-of-sample prediction of the response for a new random function $X_{n+1}$. It is shown that rates of convergence of the prediction error depend on the smoothness of the slope function and on the structure of the predictors. We then prove that these rates are optimal in the sense that they are minimax over large classes of possible slope functions and distributions of the predictive curves. For the case of models with errors-in-variables the smoothing spline estimator is modified by using a denoising correction of the covariance matrix of discretized curves. The methodology is then applied to a real case study where the aim is to predict the maximum of the concentration of ozone by using the curve of this concentration measured the preceding day.
Functional linear regression model linking observations of a functional response variable with measurements of an explanatory functional variable is considered. The slope function is estimated with a tensor product splines. Some computational issues are addressed by means of a simulation study. This model serves to analyze a real data set concerning electricity consumption in Sardinia. The interest lies in predicting either incoming weekend or incoming weekdays consumption curves if actual weekdays consumption is known.
This article considers a generalization of the functional linear regression in which an additional real variable influences smoothly the functional coefficient. We thus define a varying-coefficient regression model for functional data. We propose two estimators based, respectively, on conditional functional principal regression and on local penalized regression splines and prove their pointwise consistency. We check, with the prediction one day ahead of ozone concentration in the city of Toulouse, the ability of such nonlinear functional approaches to produce competitive estimations.
The working group STAPH is pleased to organize the First International Workshop on Functional and Operational Statistics (IWFOS). After several years of fruitful collaboration and exchange with national and international experts in the field “Statistics in infinite dimensional spaces”, the need for such a workshop was becoming increasingly evident. The workshop will ofier participants an overview of the current state of knowledge in this area, whilst at the same time providing them with an opportunity to share their own experience.
We consider functional linear regression where a real variable Ydepends on a functional variable X. The functional coeficient of the model is estimated by means of smoothing splines. We derive the rates of convergence with respect to the semi-norm induced by the covariance operator of X, which comes to evaluate the error of prediction. These rates, which essentially depend on the smoothness of the function parameter and on the structure of the predictor, are shown to be optimal over a large class of functions parameters and distributions of the predictor.