Robust regression methods have many potential applications in big data problems. In this paper, we consider two such applications using publicly available data. The first application looks at modeling taxi fares based on the trip distance of n = 49, 800 taxi rides in New York City on Tuesday January 15, 2013. The second application focuses on modeling the airfare from the miles flown of n = 78, 905 round trip itineraries for single passengers which consisted of 2 direct one-way flights within the contiguous domestic US market on Southwest Airlines in the fourth quarter of 2014. The robust estimates were obtained for both applications using PROC ROBUSTREG in SAS 9.4. In both cases, we find that the confidence intervals around the robust estimates of the parameters in the regression models are very narrow, typically $0.01 or lower. With these confidence intervals being so narrow, one is left with the impression that these robust estimates differ in some meaningful way across at least some of the robust methods. Finally, utilizing findings in Cox (Biometrika, 102:712–716, 2015) we argue that in such applications it is not surprising that the confidence intervals around the robust estimates are very narrow, thus producing the illusion of apparently very high precision.
We present a new program, gvselect, that helps users perform variable selection in regression. Best subsets variable selection is performed and provides the user with the best combinations of predictors for each level of model complexity. The leaps-and-bounds (Furnival and Wilson, 1974, Technometrics 16: 499–511) algorithm is applied using the log likelihoods of candidate models. This allows the user to perform variable selection on a wide variety of normal and non-normal regression models. Our method is described in Lawless and Singhal (1978, Biometrics 34: 318–327).
In business to business markets users are typically very experienced with the product or service. This study examines the effect of different levels of experience on satisfaction judgements, and on subsequent word of mouth intention in a business to business market. It finds that customer experience has an additive effect on disconfirmation of expectations. At any level of disconfirmation of expectations, more experienced customers were more satisfied. It also finds an interaction between experience and satisfaction in predicting word of mouth. At lower levels of satisfaction, more experienced customers were much less willing to recommend the service.
gvselect performs best subsets variable selection. The Furnival-Wilson (Technometrics, 1974) leaps-and-bounds algorithm is applied using the log likelihoods of candidate models, allowing variable selection to be performed on a wide family of normal and non-normal regression models. This method is described in Lawless and Singhal (Biometrics, 1978). The log likelihood, Akaike's information criterion, and the Bayesian information criterion are reported for the best regressions at each predictor quantity.
. Inverse response plots are a useful tool in determining a response transformation function for response linearization in regression. Under some mild conditions it is possible to seek such transformations by plotting ordinary least squares fits versus the responses. A common approach is then to use nonlinear least squares to estimate a transformation by modelling the fits on the transformed response where the transformation function depends on an unknown parameter to be estimated. We provide insight into this approach by considering sensitivity of the estimation via the influence function. For example, estimation is insensitive to the method chosen to estimate the fits in the initial step. Additionally, the inverse response plot does not provide direct information on how well the transformation parameter is being estimated and poor inverse response plots may still result in good estimates. We also introduce a simple robustified process that can vastly improve estimation.
To evaluate potential factors related to avian atherosclerosis, plasma cholesterol and triglyceride values were measured in 35 apparently healthy captive monk parakeets (Myiopsitta monachus). Birds were categorized as healthy or at risk based on body condition score and weight and were also evaluated based on their aviary environmental conditions. Plasma cholesterol mean was 8.008mmol/L (range: 4.655 to 20.33mmol/L) or 309.65mg/dl (range: 180 to 786mg/dl) for all birds sampled. Plasma triglyceride mean for all birds sampled was 4.364mmol/L (range: 0.960 to 44.62mmol/L) or 386.54mg/dl (range: 85 to 3952mg/dl). Thirty plasma samples were evaluated through density gradient ultracentrifugation lipid profiling techniques used to examine risk of cardiovascular disease in humans. The resultant lipid density profile graph was determined from the hydrated densities of the following lipids: triglyceride-rich lipoproteins, low-density lipoproteins (LDL) and subfractions, and high-density lipoproteins and subfractions. When analyzed using linear discriminant analysis, lipid profiles of triglyceride-rich lipoproteins, LDL1, LDL2, and high-density lipoprotein 2b subfractions were increased (P < 0.05) in at-risk monk parakeets when compared with healthy cohorts. Gender and diet had no apparent effect on plasma cholesterol or triglyceride concentrations. Cholesterol and triglyceride concentrations and lipoprotein density profiles from captive monk parakeets, a species known to be affected by atherosclerosis, may prove useful as markers for use in future investigation of atherosclerosis in birds. However, the consequence of increased plasma lipid concentrations and changes of lipoprotein profiles on avian health requires additional investigation.
The sliced mean variance–covariance inverse regression (SMVCIR) algorithm takes grouped multivariate data as input and transforms it to a new coordinate system where the group mean, variance, and covariance differences are more apparent. Other popular algorithms used for performing graphical group discrimination are sliced average variance estimation (SAVE, targetting the same differences but using a different arrangement for variances) and sliced inverse regression (SIR, which targets mean differences). We provide an improved SMVCIR algorithm and create a dimensionality test for the SMVCIR coordinate system. Simulations corroborating our theoretical results and comparing SMVCIR with the other methods are presented. We also provide examples demonstrating the use of SMVCIR and the other methods, in visualization and group discrimination by k-nearest neighbors. The advantages and differences of SMVCIR from SAVE and SIR are shown clearly in these examples and simulation.
This chapter contains sections titled: Introduction and Examples Location–Scale Parameter Families Estimators of Location Estimators of Dispersion Joint Estimation of Location and Dispersion Confidence Intervals for the Median Examples Problems Complements
In this paper we provide insight into the empirical properties of indirect cross-validation (ICV), a new method of bandwidth selection for kernel density estimators. First, we describe the method and report on the theoretical results used to develop a practical-purpose model for certain ICV parameters. Next, we provide a detailed description of a numerical study which shows that the ICV method usually outperforms least squares cross-validation (LSCV) in finite samples. One of the major advantages of ICV is its increased stability compared to LSCV. Two real data examples show the benefit of using both ICV and a local version of ICV.
This chapter contains sections titled: Matrix Results Vector Space Results Problems