A varying coefficients single-index regression model with responses missing at random is considered. Rank-based estimators of the index coefficient and the functional coefficients are studied, and their asymptotic properties (consistency and asymptotic normality) are established under mild conditions. To demonstrate the performance of the proposed approach, Monte Carlo simulation experiments are carried out and show that the proposed approach provides robust and more efficient estimators compared to its least-squares counterpart. This is demonstrated under different model error structures, including the standard normal, the t and the contaminated model error distributions. Finally, a real data example is given to illustrate our proposed method.
Social environments can profoundly affect the behavior and stress physiology of group-living animals. In many territorial species, territory owners advertise territorial boundaries to conspecifics by scent marking. Several studies have investigated the information that scent marks convey about donors' characteristics (e.g., dominance, age, sex, reproductive status), but less is known about whether scents affect the behavior and stress of recipients. We experimentally tested the hypothesis that scent marking may be a potent source of social stress in territorial species. We tested this hypothesis for Columbian ground squirrels (Urocitellus columbianus) during lactation, when territorial females defend individual nest-burrows against conspecifics. We exposed lactating females, on their territory, to the scent of other lactating females. Scents were either from unfamiliar females, kin relatives (a mother, daughter, or sister), or their own scent (control condition). We expected females to react strongly to novel scents from other females on their territory, displaying increased vigilance, and higher cortisol levels, indicative of behavioral and physiological stress. We further expected females to be more sensitive to unfamiliar female scents than to kin scents, given the matrilineal social structure of this species and known fitness benefits of co-breeding in female kin groups. Females were highly sensitive to intruder (both unfamiliar and kin) scents, but not to their own scent. Surprisingly, females reacted more strongly to the scent of close kin than to the scent of unfamiliar females. Vigilance behavior increased sharply in the presence of scents; this increase was more marked for kin than unfamiliar female scents, and was mirrored by a marked 131% increase in free plasma cortisol levels in the presence of kin (but not unfamiliar female) scents. Among kin scents, lactating females were more vigilant to the scent of sisters of equal age, but showed a marked 318% increase in plasma free cortisol levels in response to the scent of older and more dominant mothers. These results suggest that scent marks convey detailed information on the identity of intruders, directly affecting the stress axis of territory holders.
Generalised additive models (GAMs) provide flexible models for a wide array of data sources. In the past, improvements of GAM estimation have focused on the smoothers used in the local scoring algorithm used for estimation, but poor prediction for non-Gaussian data motivates the need for robust estimation of GAMs. In this paper, rank-based estimation, as a robust and efficient alternative to the likelihood-based estimation of GAMs, is proposed. It is shown that rank GAM estimators can be obtained through iteratively reweighted likelihood-based GAM estimation which we call the iterated regularised rank quasi-likelihood (IRRQL). Simulation experiments support the use of rank-based GAM estimation for heavy-tailed or contaminated sources of data.
Abstract The Arctic is undergoing rapid and accelerating change in response to global warming, altering biodiversity patterns, and ecosystem function across the region. For Arctic endemic species, our understanding of the consequences of such change remains limited. Spectacled eiders (Somateria fischeri), a large Arctic sea duck, use remote regions in the Bering Sea, Arctic Russia, and Alaska throughout the annual cycle making it difficult to conduct comprehensive surveys or demographic studies. Listed as Threatened under the U.S. Endangered Species Act, understanding the species response to climate change is critical for effective conservation policy and planning. Here, we developed an integrated population model to describe spectacled eider population dynamics using capture–mark–recapture, breeding population survey, nest survey, and environmental data collected between 1992 and 2014. Our intent was to estimate abundance, population growth, and demographic rates, and quantify how changes in the environment influenced population dynamics. Abundance of spectacled eiders breeding in western Alaska has increased since listing in 1993 and responded more strongly to annual variation in first‐year survival than adult survival or productivity. We found both adult survival and nest success were highest in years following intermediate sea ice conditions during the wintering period, and both demographic rates declined when sea ice conditions were above or below average. In recent years, sea ice extent has reached new record lows and has remained below average throughout the winter for multiple years in a row. Sea ice persistence is expected to further decline in the Bering Sea. Our results indicate spectacled eiders may be vulnerable to climate change and the increasingly variable sea ice conditions throughout their wintering range with potentially deleterious effects on population dynamics. Importantly, we identified that different demographic rates responded similarly to changes in sea ice conditions, emphasizing the need for integrated analyses to understand population dynamics.
There is considerable interest in the culture of whiteleg shrimp (Litopenaeus vannamei) in inland low-salinity water in Alabama and other states in the Sunbelt region of the US. However, the growing season is truncated as compared with tropical or subtropical areas where this species is typically cultured, and temperature is thought to be a major factor influencing shrimp production in the US. This study, conducted at Greene Prairie Aquafarm located in west-central Alabama, considered water temperature patterns on a shrimp farm in different ponds and different years; and sought possible effects of bottom water temperature in ponds on variation in shrimp survival, growth and production. Water temperature at 1.2 m depth in 22 ponds and air temperature were monitored at 1-hr intervals during the 2012, 2013, 2014 and 2015 growing seasons. Records of stocking rates, survival rates and production were provided by the farm owner. Correlation analysis and linear mixed model analysis of variance were used. Results showed that hourly water temperatures differed among ponds. The range of water temperature in each pond explained 41% of the variance in average final weight of shrimp harvested from each pond. In conclusion, the results suggest that variation in water temperature patterns has considerable influence on shrimp growth and survival in ponds.
In this paper, we consider a single-index regression model for which we propose a robust estimation procedure for the model parameters and an efficient variable selection of relevant predictors. The proposed method is known as the penalized generalized signed-rank procedure. Asymptotic properties of the proposed estimator are established under mild regularity conditions. Extensive Monte Carlo simulation experiments are carried out to study the finite sample performance of the proposed approach. The simulation results demonstrate that the proposed method dominates many of the existing ones in terms of robustness of estimation and efficiency of variable selection. Finally, a real data example is given to illustrate the method.
In this paper, a regression semi-parametric model is considered where responses are assumed to be missing at random. From the empirical likelihood function defined based on the rank-based estimating equation, robust confidence intervals/regions of the true regression coefficient are derived. Monte Carlo simulation experiments show that the proposed approach provides more accurate confidence intervals/regions compared to its normal approximation counterpart under different model error structure. The approach is also compared with the least squares approach, and its superiority is shown whenever the error distribution in the simulation study is heavy tailed or contaminated. Finally, a real data example is given to illustrate our proposed method.
Industry demands workers that can retrieve useful information from very complex, unstructured data. While undergraduate computer science degree courses in mathematics and statistics extensively focus on the traditional logical and problem-solving skills, they fall short of providing adequate training to students in big data concepts that integrate theory and computation. More universities are becoming aware of the need to have dedicated big data analytics programs. Most of the courses are offered at the graduate level though. After careful analysis, it is has become evident that there is a gap in the curriculum as it relates to training for big data concepts in foundational courses for students. To be more specific, there is relatively little focus on the infusion of big data concepts in mathematics and statistics courses and its impact on classroom practices. This is definitely the case in many undergraduate computer science curriculums. This paper serves to address this gap by providing an experience in infusing, teaching, and assessing big data modules in various computer science undergraduate mathematics and statistics courses that immerse students in real-world big data practices through active learning.
Annual variation in adult salmon migration timing makes the interpretation of in-season assessment data difficult, leading to much in-season uncertainty in run size. We developed and evaluated a run timing forecast model for the Kuskokwim River Chinook salmon stock, located in western Alaska, intended to aid in reducing this source of uncertainty. An objective and adaptive approach (using model-averaging and a sliding window algorithm to select predictive time periods, both calibrated annually) was adopted to deal with multidimensional selection of four climatic variables and was based entirely on predictive performance. Forecast cross-validation was used to evaluate the performance of three forecasting approaches: the null (i.e., intercept only) model, the single model with the lowest mean absolute error, and a model-averaged forecast across 16 nested linear models. As of 2016, the null model had the lowest mean absolute error (2.64days), although the model-averaged forecast performed as well or better than the null model in the majority of retrospective years. The model-averaged forecast had a consistent mean absolute error regardless of the type of year (i.e., average or extreme early/late) the forecast was made for, which was not true of the null model. The availability of the run timing forecast was not found to increase overall accuracy of in-season run assessments in relation to the null model, but was found to substantially increase the precision of these assessments, particularly early in the season.
A robust rank-based estimator for variable selection in linear models, with grouped predictors, is studied. The proposed estimation procedure extends the existing rank-based variable selection [Johnson, B. A., and Peng, L. (2008), 'Rank-based Variable Selection', Journal of Nonparametric Statistics, 20(3): 241-252] and the ww-scad [Wang, L., and Li, R. (2009), 'Weighted Wilcoxon-type Smoothly Clipped Absolute Deviation Method', Biometrics, 65(2): 564-571] to linear regression models with grouped variables. The resulting estimator is robust to contamination or deviations in both the response and the design space. The Oracle property and asymptotic normality of the estimator are established under some regularity conditions. Simulation studies reveal that the proposed method performs better than the existing rank-based methods [Johnson, B. A., and Peng, L. (2008), 'Rank-based Variable Selection', Journal of Nonparametric Statistics, 20(3): 241-252; Wang, L., and Li, R. (2009), 'Weighted Wilcoxon-type Smoothly Clipped Absolute Deviation Method', Biometrics, 65(2): 564-571] for grouped variables models. This estimation procedure also outperforms the adaptive hlasso [Zhou, N., and Zhu, J. (2010), 'Group Variable Selection Via a Hierarchical Lasso and its Oracle Property', Interface, 3(4): 557-574] in the presence of local contamination in the design space or for heavy-tailed error distribution.
The growing need for dealing with big data has made it necessary to find computationally efficient methods for identifying important factors to be considered in statistical modeling. In the linear model, the Lasso is an effective way of selecting variables using penalized regression. It has spawned substantial research in the area of variable selection for models that depend on a linear combination of predictors. However, work addressing the lack of optimality of variable selection when the model errors are not Gaussian and/or when the data contain gross outliers is scarce. We propose the weighted signed-rank Lasso as a robust and efficient alternative to least absolute deviations and least squares Lasso. The approach is appealing for use with big data since one can use data augmentation to perform the estimation as a single weighted L 1 optimization problem. Selection and estimation consistency are theoretically established and evaluated via simulation studies. The results confirm the optimality of the rank-based approach for data with heavy-tailed and contaminated errors or data containing high-leverage points.
Repeated measurement designs occur in many areas of statistical research. In 1986, Liang and Zeger offered an elegant analysis of these problems based on a set of generalized estimating equations (GEEs) for regression parameters, that specify only the relationship between the marginal mean of the response variable and covariates. Their solution is based on iterated reweighted least squares fitting. In this paper, we propose a rank-based fitting procedure that only involves substituting a norm based on a score function for the Euclidean norm used by Liang and Zeger. Our subsequent fitting, while also an iterated reweighted least squares solution to GEEs, is robust to outliers in response space and the weights can easily be adapted for robustness in factor space. As with the fitting of Liang and Zeger, our rank-based fitting utilizes a working covariance matrix. We prove that our estimators of the regression coefficients are asymptotically normal. The results of a simulation study show that the our proposed estimators are empirically efficient and valid. We illustrate our analysis on a real data set drawn from a hierarchical (three-way nested) design.
Abstract Energetic trade‐offs in resource allocation form the basis of life‐history theory, which predicts that reproductive allocation in a given season should negatively affect future reproduction or individual survival. We examined how allocation of resources differed between successful and unsuccessful breeding female Columbian ground squirrels to discern any effects of resource allocation on reproductive and somatic efforts. We compared the survival rates, subsequent reprodction, and mass gain of successful breeders (females that successfully weaned young) and unsuccessful breeders (females that failed to give birth or wean young) and investigated “carryover” effects to the next year. Starting capital was an important factor influencing whether successful reproduction was initiated or not, as females with the lowest spring emergence masses did not give birth to a litter in that year. Females that were successful and unsuccessful at breeding in one year, however, were equally likely to be successful breeders in the next year and at very similar litter sizes. Although successful and unsuccessful breeding females showed no difference in over winter survival, females that failed to wean a litter gained additional mass during the season when they failed. The next year, those females had increased energy “capital” in the spring, leading to larger litter sizes. Columbian ground squirrels appear to act as income breeders that also rely on stored capital to increase their propensity for future reproduction. Failed breeders in one year “prepare” for future reproduction by accumulating additional mass, which is “carried over” to the subsequent reproductive season.
Strict food safety measures adopted by the US have resulted in a number of detentions and refusals of foods imported from Latin American and Caribbean (LAC) countries. These food detentions have resulted in serious economic consequences to the food industries in these exporting countries. The objective of this study was to examine the trends in food detentions and refusals and to develop a model to explain the rising increases in food detentions in LAC countries. Food detentions from 14 LAC countries from 1992 to 2007 were examined. Poisson and Negative Binomial regression models were used to determine the factors that influenced food detentions. It is noted that food detentions have been increasing with exports from LAC countries. The increase has been more pronounced since 2002. The reasons for detention have changed over the period studied. Pesticides, detention without physical examination (DWPE) and filth were the most common reasons before 2002, whereas, labeling, pesticides, and DWPE were the most common reasons after 2002. The products most commonly detained are vegetables (42%), snack and processed foods (34%), and fish (9%). The model results show that foreign direct investment (FDI) influenced the number of detentions whereas the number of detentions decreased with time. The models produced some insightful results which show that detention patterns are likely to change in emphasis, increasing with FDI, but reducing with time.
We consider a nonlinear regression model when the index variable is multidimensional. Such models are useful in signal processing, texture modelling, and spatio-temporal data analysis. A generalised form of the signed-rank estimator of the nonlinear regression coefficients is proposed. This general form of the signed-rank (SR) estimator includes estimators and hybrid variants. Sufficient conditions for strong consistency and asymptotic normality of the estimator are given. It is shown that the rate of convergence to normality can be different from . The sufficient conditions are weak in the sense that they are satisfied by harmonic-type functions for which results in the current literature may not apply. A simulation study shows that certain generalised SR estimators (e.g. signed rank) perform better in recovering signals than others (e.g. least squares) when the error distribution is contaminated or is heavy-tailed.
The main purpose of this paper is to investigate macroeconomic variables that are predictive of banking crisis. We focus on selecting variables that have high predictive power to discriminate between two groups of countries: the sound and the distressed. We consider a sample of 50 emerging market and developing countries during 1990–2005 time period, and apply generalized estimating equations as well as univariate, bivariate and trivariate transvariation analysis to choose the variables that have high predictive and discriminative powers to separate the distressed countries from the sound countries. In order to compare the predictive performance of these selected variables, we calculate the leave one out predictive error rate for the top ranking variables with low transvariation probability using discrimination procedures. Not surprisingly the results show that countries with better macroeconomic and monetary environments and healthier banking institutions are less likely to suffer systemic banking crisis.
We obtain an atomic decomposition of weighted Lorentz spaces for a class of weights satisfying the Δ2 condition. Consequently, we study operators such as the multiplication and composition operators and also provide Hölder’s-type and duality-Riesz type inequalities on these weighted Lorentz spaces.
We provide an estimate of the score function for rank regression using compactly supported wavelets. This estimate is then used to find a rank-based asymptotically efficient estimator for the slope parameter of a linear model. We also provide a consistent estimator of the asymptotic variance of the rank estimator. For related mixed models, the asymptotic relative efficiency is also discussed
In this paper, the estimation of parameters of a generalized linear regression model is considered. The proposed estimator is defined iteratively starting from an initial obtained by minimizing the Wilcoxon dispersion function for independent errors. It is shown that the iterative estimator converges to the rank version of the maximum quasi-likelihood estimator as the number of iterations increases. The consistency and the asymptotic normality of the rank version of the maximum quasi-likelihood estimator are given. As in the linear model, the procedure results in estimators that are robust in the response space. This is proven theoretically via the influence function. A simulation study and a real world data example illustrate the robustness in the response space and efficiency of the estimator.