
SummaryDistance-weighted discrimination (DWD) is a modern margin-based classifier with an interesting geometric motivation. It was proposed as a competitor to the support vector machine (SVM). Despite many recent references on DWD, DWD is far less popular than the SVM, mainly because of computational and theoretical reasons. We greatly advance the current DWD methodology and its learning theory. We propose a novel thrifty algorithm for solving standard DWD and generalized DWD, and our algorithm can be several hundred times faster than the existing state of the art algorithm based on second-order cone programming. In addition, we exploit the new algorithm to design an efficient scheme to tune generalized DWD. Furthermore, we formulate a natural kernel DWD approach in a reproducing kernel Hilbert space and then establish the Bayes risk consistency of the kernel DWD by using a universal kernel such as the Gaussian kernel. This result solves an open theoretical problem in the DWD literature. A comparison study on 16 benchmark data sets shows that data-driven generalized DWD consistently delivers higher classification accuracy with less computation time than the SVM.
SummaryWe propose a framework for constructing goodness-of-fit tests in both low and high dimensional linear models. We advocate applying regression methods to the scaled residuals following either an ordinary least squares or lasso fit to the data, and using some proxy for prediction error as the final test statistic. We call this family residual prediction tests. We show that simulation can be used to obtain the critical values for such tests in the low dimensional setting and demonstrate using both theoretical results and extensive numerical studies that some form of the parametric bootstrap can do the same when the high dimensional linear model is under consideration. We show that residual prediction tests can be used to test for significance of groups or individual variables as special cases, and here they compare favourably with state of the art methods, but we also argue that they can be designed to test for as diverse model misspecifications as heteroscedasticity and non-linearity.
Summary We define an isotropic Lévy-driven continuous auto-regressive moving average CARMA(p, q) random field on Rn as the integral of a radial CARMA kernel with respect to a Lévy sheet. Such fields constitute a parametric family characterized by an auto-regressive polynomial a and a moving average polynomial b having zeros in both the left and the right complex half-planes. They extend the well-balanced Ornstein–Uhlenbeck process of Schnurr and Woerner to a well-balanced CARMA process in one dimension (with a much richer class of autocovariance functions) and to an isotropic CARMA random field on Rn for n > 1. We derive second-order properties of these random fields and extend the results to a larger class of anisotropic CARMA random fields. If the driving Lévy sheet is compound Poisson it is trivial to simulate the corresponding random field on any bounded subset of Rn. A method for joint estimation of the CARMA kernel parameters and knot locations is proposed for compound-Poisson-driven fields and is illustrated by applications to simulated data and Tokyo land price data.
SummaryOver the last decade a variety of models to analyse incomplete multivariate and longitudinal data have been proposed, many of which allowing for the missingness to be not at random, in the sense that the unobserved measurements influence the process governing missingness, in addition to influences coming from observed measurements and/or covariates. The fundamental problems that are implied by such models, to which we refer as sensitivity to unverifiable modelling assumptions, has, in turn, sparked off various strands of research in what is now termed sensitivity analysis. The nature of sensitivity originates from the fact that a missingness not at random (MNAR) model is not fully verifiable from the data, rendering the empirical distinction between MNAR and missingness at random (MAR), where only covariates and observed outcomes influence missingness, difficult or even impossible, unless we are willing to accept the posited MNAR model in an unquestioning way. We show that the empirical distinction between MAR and MNAR is not possible, in the sense that each MNAR model fit to a set of observed data can be reproduced exactly by an MAR counterpart. Of course, such a pair of models will produce different predictions of the unobserved outcomes, given the observed outcomes. Theoretical considerations are supplemented with an illustration that is based on the Slovenian public opinion survey, which has been analysed before in the context of sensitivity analysis.
SummarySemiparametric regression models play a central role in formulating the effects of covariates on potentially censored failure times and in the joint modelling of incomplete repeated measures and failure times in longitudinal studies. The presence of infinite dimensional parameters poses considerable theoretical and computational challenges in the statistical analysis of such models. We present several classes of semiparametric regression models, which extend the existing models in important directions. We construct appropriate likelihood functions involving both finite dimensional and infinite dimensional parameters. The maximum likelihood estimators are consistent and asymptotically normal with efficient variances. We develop simple and stable numerical techniques to implement the corresponding inference procedures. Extensive simulation experiments demonstrate that the inferential and computational methods proposed perform well in practical settings. Applications to three medical studies yield important new insights. We conclude that there is no reason, theoretical or numerical, not to use maximum likelihood estimation for semiparametric regression models. We discuss several areas that need further research.
Increasingly, scientific studies yield functional data, in which the ideal units of observation are curves and the observed data consist of sets of curves that are sampled on a fine grid. We present new methodology that generalizes the linear mixed model to the functional mixed model framework, with model fitting done by using a Bayesian wavelet-based approach. This method is flexible, allowing functions of arbitrary form and the full range of fixed effects structures and between-curve covariance structures that are available in the mixed model framework. It yields nonparametric estimates of the fixed and random-effects functions as well as the various between-curve and within-curve covariance matrices. The functional fixed effects are adaptively regularized as a result of the non-linear shrinkage prior that is imposed on the fixed effects' wavelet coefficients, and the random-effect functions experience a form of adaptive regularization because of the separately estimated variance components for each wavelet coefficient. Because we have posterior samples for all model quantities, we can perform pointwise or joint Bayesian inference or prediction on the quantities of the model. The adaptiveness of the method makes it especially appropriate for modelling irregular functional data that are characterized by numerous local features like peaks.
We consider the prediction problem of a continuous-time stochastic process on an entire time-interval in terms of its recent past. The approach we adopt is based on functional kernel nonparametric regression estimation techniques where observations are segments of the observed process considered as curves. These curves are assumed to lie within a space of possibly inhomogeneous functions, and the discretized times series dataset consists of a relatively small, compared to the number of segments, number of measurements made at regular times. We thus consider only the case where an asymptotically non-increasing number of measurements is available for each portion of the times series. We estimate conditional expectations using appropriate wavelet decompositions of the segmented sample paths. A notion of similarity, based on wavelet decompositions, is used in order to calibrate the prediction. Asymptotic properties when the number of segments grows to infinity are investigated under mild conditions, and a nonparametric resampling procedure is used to generate, in a flexible way, valid asymptotic pointwise confidence intervals for the predicted trajectories. We illustrate the usefulness of the proposed functional wavelet-kernel methodology in finite sample situations by means of three real-life datasets that were collected from different arenas.