The instrument Cometary Secondary Ion Mass Analyzer (COSIMA) on board of the European Space Agency mission Rosetta to the comet 67P/Churyumov‐Gerasimenko is a secondary ion mass spectrometer with a time‐of‐flight mass analyzer. It collected near the comet several thousand particles, imaged them, and analyzed the elemental and chemical compositions of their surfaces. In this study, variables have been generated from the spectral data covering the mass ranges of potential C‐, H‐, N‐, and O‐containing ions. The variable importance in binary discriminations between spectra measured on cometary particles and those measured on the target background has been estimated by the univariate t test and the multivariate methods discriminant partial least squares, random forest, and a robust method based on the log ratios of all variable pairs. The results confirm the presence of organic substances in cometary matter—probably a complex macromolecular mixture.
Fully robust versions of the elastic net estimator are introduced for linear and logistic regression. The algorithms used to compute the estimators are based on the idea of repeatedly applying the non-robust classical estimators to data subsets only. It is shown how outlier-free subsets can be identified efficiently, and how appropriate tuning parameters for the elastic net penalties can be selected. A final reweighting step improves the efficiency of the estimators. Simulation studies compare with non-robust and other competing robust estimators and reveal the superiority of the newly proposed methods. This is also supported by a reasonable computation time and by good performance in real data examples.
A novel approach for supervised classification analysis for high dimensional and flat data (more variables than observations) is proposed. We use the information of class-membership of observations to determine groups of observations locally describing the group structure. By projecting the data on the subspace spanned by those groups, local projections are defined based on the projection concepts from Ortner et al. (2017a) and Ortner et al. (2017b). For each local projection a local discriminant analysis (LDA) model is computed using the information within the projection space as well as the distance to the projection space. The models provide information about the quality of separation for each class combination. Based on this information, weights are defined for aggregating the LDA-based posterior probabilities of each subspace to a new overall probability. The same weights are used for classifying new observations. In addition to the provided methodology, implemented in the R-package lop, a method of visualizing the connectivity of groups in high-dimensional spaces is proposed on the basis of the posterior probabilities. A thorough evaluation is performed using three different real-world datasets, underlining the strengths of local projection based classification and the provided visualization methodology.
Partial robust M regression (PRM), as well as its sparse counterpart sparse PRM, have been reported to be regression methods that foster a partial least squares‐alike interpretation while having good robustness and efficiency properties, as well as a low computational cost. In this paper, the partial robust M discriminant analysis classifier is introduced, which consists of dimension reduction through an algorithm closely related to PRM and a consecutive robust discriminant analysis in the latent variable space. The method is further generalized to sparse partial robust M discriminant analysis by introducing a sparsity penalty on the estimated direction vectors. Thereby, an intrinsic variable selection is achieved, which yields a better graphical interpretation of the results, as well as more precise coefficient estimates, in case the data contain uninformative variables. Both methods are robust against leverage points within each class, as well as against adherence outliers (points that have been assigned a wrong class label). A simulation study investigates the effect of outliers, wrong class labels, and uninformative variables on the proposed methods and its classical PLS counterparts and corroborates the robustness and sparsity claims. The utility of the methods is demonstrated on data from mass spectrometry analysis (time‐of‐flight secondary ion mass spectrometry) of meteorite samples. Copyright © 2016 John Wiley & Sons, Ltd.