Nonparametric regression models offer a way to understand and quantify relationships between variables without having to identify an appropriate family of possible regression functions. Although many estimation methods for these models have been proposed in the literature, most of them can be highly sensitive to the presence of a small proportion of atypical observations in the training set. A review of outlier robust estimation methods for nonparametric regression models is provided, paying particular attention to practical considerations. Since outliers can also influence negatively the regression estimator by affecting the selection of bandwidths or smoothing parameters, a discussion of robust alternatives for this task is also included. Using many of the “classical” nonparametric regression estimators (and their robust counterparts) can be very challenging in settings with a moderate or large number of explanatory variables, so recently proposed robust nonparametric regression methods that scale well with a growing number of covariates are also discussed.
In this article we propose a boosting algorithm for regression with functional explanatory variables and scalar responses. The algorithm uses decision trees constructed with multiple projections as the "base-learners", which we call "functional multi-index trees". We establish identifiability conditions for these trees and introduce two algorithms to compute them. We use numerical experiments to investigate the performance of our method and compare it with several linear and nonlinear regression estimators, including recently proposed nonparametric and semiparametric functional additive estimators. Simulation studies show that the proposed method is consistently among the top performers, whereas the performance of existing alternatives can vary substantially across different settings. In a real example, we apply our method to predict electricity demand using price curves and show that our estimator provides better predictions compared to its competitors, especially when one adjusts for seasonality.
Gradient boosting algorithms construct a regression predictor using a linear combination of ``base learners''. Boosting also offers an approach to obtaining robust non-parametric regression estimators that are scalable to applications with many explanatory variables. The robust boosting algorithm is based on a two-stage approach, similar to what is done for robust linear regression: it first minimizes a robust residual scale estimator, and then improves it by optimizing a bounded loss function. Unlike previous robust boosting proposals this approach does not require computing an ad-hoc residual scale estimator in each boosting iteration. Since the loss functions involved in this robust boosting algorithm are typically non-convex, a reliable initialization step is required, such as an L1 regression tree, which is also fast to compute. A robust variable importance measure can also be calculated via a permutation procedure. Thorough simulation studies and several data analyses show that, when no atypical observations are present, the robust boosting approach works as well as the standard gradient boosting with a squared loss. Furthermore, when the data contain outliers, the robust boosting estimator outperforms the alternatives in terms of prediction error and variable selection accuracy.
In this paper we review existing methods for robust functional principal component analysis (FPCA) and propose a new method for FPCA that can be applied to longitudinal data where only a few observations per trajectory are available. This method is robust against the presence of atypical observations, and can also be used to derive a new non-robust FPCA approach for sparsely observed functional data. We use local regression to estimate the values of the covariance function, taking advantage of the fact that for elliptically distributed random vectors the conditional location parameter of some of its components given others is a linear function of the conditioning set. This observation allows us to obtain robust FPCA estimators by using robust local regression methods. The finite sample performance of our proposal is explored through a simulation study that shows that, as expected, the robust method outperforms existing alternatives when the data are contaminated. Furthermore, we also see that for samples that do not contain outliers the non-robust variant of our proposal compares favourably to the existing alternative in the literature. A real data example is also presented.
In the last 25 years there has been an important increase in the amount of data collected from animal-mounted sensors (bio-probes) which are often used to study the animals' behaviour or environment. We focus here on an example of the latter, where the interest is in sea surface temperature (SST), and measurements are taken from sensors mounted on elephant seals in the southern Indian Ocean. We show that standard geostatistical models may not be reliable for this type of data, due to the possibility that the regions visited by the animals may depend on the SST. This phenomenon is know in the literature as preferential sampling, and, if ignored, it may affect the resulting spatial predictions and parameter estimates. Research on this topic has been mostly restricted to stationary sampling locations such as monitoring sites. The main contribution of this manuscript is to extend this methodology to observations obtained by devices that move through the region of interest, as is the case with the tagged seals. More specifically, we propose a flexible framework for inference on preferentially sampled fields where the process that generates the sampling locations is stochastic and moving over time through a two-dimensional space. Our simulation studies confirm that predictions obtained from the preferential sampling model are more reliable when this phenomenon is present, and they compare very well to the standard ones when there is no preferential sampling. Finally, we note that the conclusions of our analysis of the SST data can change considerably when we incorporate preferential sampling in the model.
In large-scale quantitative proteomic studies, scientists measure the abundance of thousands of proteins from the human proteome in search of novel biomarkers for a given disease. Penalized regression estimators can be used to identify potential biomarkers among a large set of molecular features measured. Yet, the performance and statistical properties of these estimators depend on the loss and penalty functions used to define them. Motivated by a real plasma proteomic biomarkers study, we propose a new class of penalized robust estimators based on the elastic net penalty, which can be tuned to keep groups of correlated variables together in the selected model and maintain robustness against possible outliers. We also propose an efficient algorithm to compute our robust penalized estimators and derive a data-driven method to select the penalty term. Our robust penalized estimators have very good robustness properties and are also consistent under certain regularity conditions. Numerical results show that our robust estimators compare favorably to other robust penalized estimators. Using our proposed methodology for the analysis of the proteomics data, we identify new potentially relevant biomarkers of cardiac allograft vasculopathy that are not found with nonrobust alternatives. The selected model is validated in a new set of 52 test samples and achieves an area under the receiver operating characteristic (AUC) of 0.85.
SummaryPreferential sampling in geostatistics occurs when the locations at which observations are made may depend on the spatial process that underlines the correlation structure of the measurements. We show that previously proposed Monte Carlo estimates for the likelihood function may not be approximating the desired function. Furthermore, we argue that, for preferential sampling of moderate complexity, alternative and widely available numerical methods to approximate the likelihood function produce better results than Monte Carlo methods. We illustrate our findings on the Galicia data set analysed previously in the literature.
Classical statistical techniques fail to cope well with deviations from a standard distribution. Robust statistical methods take into account these deviations while estimating the parameters of parametric models, thus increasing the accuracy of the inference. Research into robust methods is flourishing, with new methods being developed and different applications considered. Robust Statistics sets out to explain the use of robust methods and their theoretical justification. It provides an up-to-date overview of the theory and practical application of the robust statistical methods in regression, multivariate analysis, generalized linear models and time series. This unique book: Enables the reader to select and use the most appropriate robust method for their particular statistical model. Features computational algorithms for the core methods. Covers regression methods for data mining applications. Includes examples with real data and applications using the S-Plus robust statistics library. Describes the theoretical and operational aspects of robust methods separately, so the reader can choose to focus on one or the other. Supported by a supplementary website featuring time-limited S-Plus download, along with datasets and S-Plus code to allow the reader to reproduce the examples given in the book. Robust Statistics aims to stimulate the use of robust methods as a powerful tool to increase the reliability and accuracy of statistical modelling and data analysis. It is ideal for researchers, practitioners and graduate students of statistics, electrical, chemical and biochemical engineering, and computer vision. There is also much to benefit researchers from other sciences, such as biotechnology, who need to use robust statistical methods in their work.
SummaryUnder certain regularity conditions, maximum-likelihood-based inference enjoys several optimality properties, including high asymptotic efficiency. However, if the distribution of the data deviates slightly from the model proposed, the statistical properties of inference methods based on maximum likelihood can quickly deteriorate. We focus on the situation when the interest lies in one of the tails of the distribution, e.g. when we are estimating a high or low quantile. In this case, it may be natural, if slightly unorthodox, to consider models that fit well the corresponding tail of the sample, rather than its whole range. For example, if we are interested in estimating the fifth percentile, we can pretend that all observations above the 10th percentile have been censored and fit a parametric censored model to the lower tail of the sample. Such an approach, which we call ‘artificial censoring’, has been studied in the engineering literature. We study a data-dependent method to select the amount of artificial censoring and show that it compares favourably with the optimally chosen (‘oracle’) method, which is generally unavailable in practice. We also show that the artificial censoring approach can be applied to estimate tail dependence parameters in copula models, and that it performs well both in simulation and in real data studies.
Penalized regression estimators have been widely used in recent years to improve the prediction properties of linear models, particularly when the number of explanatory variables is large. It is well-known that different penalties result in regularized estimators with varying statistical properties. Motivated by the analysis of plasma proteomic biomarkers that tend to form groups of correlated predictors, we focus here on estimators with an Elastic Net penalty, in order to keep these groups of variables together as they enter or leave the model. Given the presence of potential outliers in our data, we propose a class of penalized S-estimators which have very good robustness properties. Furthermore, these penalized S-estimators can be used as initial values to compute more efficient penalized M-estimators. In this paper we derive an algorithm to compute our proposed estimators, and also a data-driven method to select the penalty term, which is a critical part of any application with real data. Our robust penalized estimators have very good robustness properties and are also consistent under relatively weak assumptions. Our numerical experiments show that our proposals compare favourably to other robust penalized estimators. When applied to our motivating example, the robust estimators identify new potentially relevant biomarkers that are not found with non-robust alternatives. Moreover, the robust estimators identify two patients with a suspected low obstruction in the artery examined. Further measurements by a more accurate technique validated our predictions. 1
Additive models provide an attractive setup to estimate regression functions in a nonparametric context. They provide a flexible and interpretable model, where each regression function depends only on a single explanatory variable and can be estimated at an optimal univariate rate. Most estimation procedures for these models are highly sensitive to the presence of even a small proportion of outliers in the data. In this paper, we show that a relatively simple robust version of the backfitting algorithm (consisting of using robust local polynomial smoothers) corresponds to the solution of a well-defined optimisation problem. This formulation allows us to find mild conditions to show Fisher consistency and to study the convergence of the algorithm. Our numerical experiments show that the resulting estimators have good robustness and efficiency properties. We illustrate the use of these estimators on a real data set where the robust fit reveals the presence of influential outliers.
Witten and Tibshirani (2010) proposed an algorithim to simultaneously find clusters and select clustering variables, called sparse K-means (SK-means). SK-means is particularly useful when the dataset has a large fraction of noise variables (that is, variables without useful information to separate the clusters). SK-means works very well on clean and complete data but cannot handle outliers nor missing data. To remedy these problems we introduce a new robust and sparse K-means clustering algorithm implemented in the R package RSKC. We demonstrate the use of our package on four datasets. We also conduct a Monte Carlo study to compare the performances of RSK-means and SK-means regarding the selection of important variables and identification of clusters. Our simulation study shows that RSK-means performs well on clean data and better than SK-means and other competitors on outlier-contaminated data.
ANOVA tests are the standard tests to compare nested linear models fitted by least squares. These tests are equivalent to likelihood ratio tests, so they have high power. However, least squares estimators are very vulnerable to outliers in the data, and thus the related ANOVA type tests are also extremely sensitive to outliers. Therefore, robust estimators can be considered to obtain a robust alternative to the ANOVA tests. Regression ¿ -estimators combine high robustness with high efficiency which makes them suitable for robust inference beyond parameter estimation. Robust likelihood ratio type test statistics based on the ¿ -estimates of the error scale in the linear model are a natural alternative to the classical ANOVA tests. The higher efficiency of the ¿ -scale estimates compared with other robust alternatives is expected to yield tests with good power. Their null distribution can be estimated using either an asymptotic approximation or the fast and robust bootstrap. The robustness and power of the resulting robust likelihood ratio type tests for nested linear models is studied.
Automatic modulation recognition (AMR) enables the detection of different data-transmission formats sharing the same frequency band. One such band is the European 868MHz band that is dedicated for short-range devices. Recently, an AMR method that applies a feature-based tree has been proposed for this band. In this paper, we present alternative feature-based classifiers that enable a more accurate AMR. In particular, we propose the use of classification tree and random forest classifiers, and we devise an extended set of features for the modulation classification problem at hand. Through simulation experiments we demonstrate a significant improvement in recognition success rate for typical transmission types in the 868MHz band.