This work discusses the article Independent component analysis by robust distance correlation appeared within this same issue of this journal. Particularly, it focuses into the novel bowl transform, a proposal embedded into an Independent Component Analysis strategy to obtain solutions which are resistant to the presence of contamination in the data. The proposal is discussed, and an alternative solution is suggested. This latter exploits data depth contours to embed a multivariate winsorization step within the same proposed strategy.
Circular variables that represent directions or periodic observations arise in many fields, such as biology and environmental sciences. An important issue when dealing with circular data is how to estimate their dispersion robustly, avoiding undue effects of anomalies. This work extends three robust dispersion measures from the line to the circle. Their robustness is studied via their influence functions and relative bias curves. From these dispersion measures, robust estimators of parameters of circular distributions can be derived. This yields robust estimators for the concentration parameter of the von Mises distribution and the dispersion parameter of the wrapped normal distribution. Their breakdown values and statistical efficiencies are obtained, and they are compared in a simulation study. Building on the best performing estimator, a robust circular anomaly detection procedure is developed, and employed to visualize outliers through a circular violin plot. Three real datasets are analyzed.
This special issue of Statistical Analysis and Data Mining contains a selection of the papers presented at the 13th Scientific Meeting of the Classification and Data Analysis Group (CLADAG), scheduled for September 9–11, 2021 in Florence, Italy. Due to the COVID-19 pandemic, the conference was held online. The CLADAG is a Section of the Italian Statistical Society (SIS), and a member of the International Federation of Classification Societies (IFCS). It was founded in 1997 to promote advanced methodological research in multivariate statistics, focusing on Data Analysis and Classification. The Section organizes a biennial international scientific meeting, offers classification and data analysis courses, publishes a newsletter, and collaborates on planning conferences and meetings with other IFCS societies. The previous 12 CLADAG meetings were held in various locations throughout Italy: Pescara (1997), Roma (1999), Palermo (2001), Bologna (2003), Parma (2005), Macerata (2007), Catania (2009), Pavia (2011), Modena and Reggio Emilia (2013), Cagliari (2015), Milano (2017), and Cassino (2019). Following a blind peer-review process, six papers presented at the conference and submitted to this special issue have been selected for publication. The articles cover a broad range of data analysis topics: gender gap analysis, income clustering, structural equation modeling, multivariate nonparametric methods, and classifier selection. Their content is briefly described below. In studying the gender gap, a relevant topic for promoting equality and social justice, Greselin et al. propose a new parametric approach utilizing the relative distribution method and Dagum parametric inference. Additionally, they assessed how to select covariates that impact gender gaps. The proposed approach is applied to measure and compare the gender gap in Poland and Italy, using data from the 2018 European Survey of Income and Living Conditions. On a related field, Condino proposes a procedure for clustering income data using a share density-based dynamic clustering algorithm. The paper compares subgroups’ income inequality using a dissimilarity measure based on information theory. This measure is then utilized for clustering, providing a prototype descriptor of income inequality for the clustered earners. The proposal is applied to data from the Survey on Households Income and Wealth by the Bank of Italy. The paper by Yu et al. introduces a refinement of the so-called Henseler–Ogasawara specification that integrates composites, linear combinations of variables, into structural equation models. This refined version addresses some concerns of the Henseler–Ogasawara specification, and it is less complex and less prone to misspecification mistakes. Additionally, the paper provides a strategy to compute standard errors. Statistical depth functions are a valuable tool for multivariate nonparametric data analysis, extending the concept of ranks, orderings, and quantiles to the multivariate setup. The paper by Laketa and Nagy investigates one of the fundamental open problems of contemporary depth research, the so-called characterization and reconstruction questions, focusing on the simplicial depth. Their results are illustrated via several insightful examples. On the same topic, Nagy revisits the classical definition of the simplicial depth and explores its theoretical properties. Particularly, properties of the simplicial median are investigated. The author provides the exact simplicial depth in several scenarios, outlining undesirable behaviors of this depth function. Carpita and Golia tackle the problem of choosing the rule to assign a unit to a category given the estimated probabilities. In particular, the paper compares the classical Bayesian Classifier, which minimizes the expected classification error rate, with the Max Difference Classifier and the Max Ratio Classifier, showing when these classifiers should be preferred. Findings are illustrated by means of a broad simulation study and an application on benchmark data sets. To conclude, we believe that this special issue accurately portrays the scientific features of the CLADAG community nowadays and supports the CLADAG mission of facilitating the exchange of ideas in Classification and Data Analysis. We warmly encourage all readers to attend the
Depth functions offer an array of tools that enable the introduction of quantile- and ranking-like approaches to multivariate and non-Euclidean datasets. We investigate the potential of using depths in the problem of nonparametric supervised classification of directional data, that is classification of data that naturally live on the unit sphere of a Euclidean space. In this paper, we address the problem mainly from a theoretical side, with the final goal of offering guidelines on which angular depth function should be adopted in classifying directional data. A set of desirable properties of an angular depth is put forward. With respect to these properties, we compare and contrast the most widely used angular depth functions. Simulated and real data are eventually exploited to showcase the main implications of the discussed theoretical results, with an emphasis on potentials and limits of the often disregarded angular halfspace depth.
Contaminated training sets can highly affect the performance of classification rules. For this reason, robust supervised classifiers have been introduced. Amongst the many, this work focuses on depth-based classifiers, a class of methods which have been proven to enjoy some robustness properties. However, no robustness studies are available for them within a directional data framework. Here, their performance under some directional contamination schemes is evaluated. A comparison with the directional Bayes rule is also provided. Different directional specific contamination scenarios are introduced and discussed: antipodality and orthogonality of the contaminated distribution mean, and the directional mean shift outlier model.
Directional data lies on the surface of the unit sphere. Exploiting new results on the computation and the properties of the angular halfspace depth, we introduce the spherical version of the bagdistance, applicable to directional data. A bagdistance-based classification method for directional data is considered. The proposed method will be compared with other directional classifiers by means of a simulation study.
Statistical Analysis and Data Mining: The ASA Data Science JournalVolume 14, Issue 4 p. 295-296 INTRODUCTION CLADAG 2019 Special Issue: Selected Papers on Classification and Data Analysis Francesca Greselin, Department of Statistics and Quantitative Methods, University of Milano-Bicocca, Milan, ItalySearch for more papers by this authorThomas Brendan Murphy, School of Mathematics and Statistics & Insight Research Centre, University College Dublin, Dublin, IrelandSearch for more papers by this authorGiovanni C. Porzio, Department of Economics and Law, University of Cassino and Southern Lazio, Cassino, ItalySearch for more papers by this authorDomenico Vistocco, Department of Political Science, University of Naples Federico II, Naples, ItalySearch for more papers by this author Francesca Greselin, Department of Statistics and Quantitative Methods, University of Milano-Bicocca, Milan, ItalySearch for more papers by this authorThomas Brendan Murphy, School of Mathematics and Statistics & Insight Research Centre, University College Dublin, Dublin, IrelandSearch for more papers by this authorGiovanni C. Porzio, Department of Economics and Law, University of Cassino and Southern Lazio, Cassino, ItalySearch for more papers by this authorDomenico Vistocco, Department of Political Science, University of Naples Federico II, Naples, ItalySearch for more papers by this author First published: 16 June 2021 https://doi.org/10.1002/sam.11533Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinked InRedditWechat No abstract is available for this article. Volume14, Issue4Special Issue: CLADAG 2019: Selected Papers on Classification and Data AnalysisAugust 2021Pages 295-296 RelatedInformation
A directional random variable is rotationally symmetric around a location parameter if its distribution only depends on the angle between the value the variable can take and the location parameter itself. This is clearly an oversimplified model. On the other hand, the performance of depth-based classifiers for directional data has been widely studied only under the case of rotational symmetry of the underlying class distributions. For this reason, this work aims at evaluating the efficacy of some of them under non-rotational symmetry. Particularly, the DD-classifiers exploiting the linear, quadratic, and KNN discriminant rules when associated with angular distance-based depths are examined. Their performances under Kent distributions are investigated by means of a simulation study. As a benchmark, the directional Bayes rule is considered. In passing, these classifiers are also reviewed, noting that the way they work within the directional data domain has been given a bit for granted within the literature.
The main goal of supervised learning is to construct a function from labeled training data which assigns arbitrary new data points to one of the labels. Classification tasks may be solved by using some measures of data point centrality with respect to the labeled groups considered. Such a measure of centrality is called data depth. In this paper, we investigate conditions under which depth-based classifiers for directional data are optimal. We show that such classifiers are equivalent to the Bayes (optimal) classifier when the considered distributions are rotationally symmetric, unimodal, differ only in location and have equal priors. The necessity of such assumptions is also discussed.
Directions, rotations, axes, clock, or calendar measurements can be represented as angles or equivalently as unit vectors. As points lying on the boundary of circles, spheres, or hyper-spheres, they are also referred as directional data, and they require dedicated methods to be analyzed. In the framework of supervised classification, this work introduces a directional data classifier based on a data depth function. Depth functions provide an inner–outer ordering of the data in a reference space according to some centrality measure, and have appeared as a powerful tool in many fields of multivariate statistics. The recently introduced distance-based depth functions for directional data are considered here. More specifically, this work introduces a cosine depth based distribution method which aims at assigning directional data to classes, given that a training set with class labels is already available. A simulation study evaluating the performance of the proposed method is provided.
Robust location estimators for directional data are known for about 30 years. Scientific literature has focused on studying the asymptotic properties of these estimators like consistency and influence function. Apart from the finite-sample breakdown point, the finite-sample performance of robust directional location estimators has attracted less attention. Hence, it is discussed how the finite-sample max-bias of directional location estimators can be evaluated. Additionally, two new robust estimators of the mean direction are introduced: the spherical Minimum Covariance Determinant estimator (sMCD) and the spherical Minimum Spanning Tree estimator (sMST). The sMCD seeks to identify the densest subset of a given size while the sMST seeks for a well-separated subset. Finally, the robust estimators are compared with respect to the max-bias and to the bias under shift outlier scenarios by means of an extensive simulation study. The results indicate that –in contrast to linear data– the maximum likelihood estimator shows high robustness in terms of the finite-sample max-bias. However, robust estimators are clearly superior to the maximum likelihood estimator in shift outlier contamination schemes.
This work discusses how to test antipodal symmetry of circular distributions through depth functions. Two notions of depths for circular data are adopted, and their performances are evaluated and compared through a simulation study.
A procedure is developed in order to deal with the classification problem of objects in circular statistics. It is fully nonparametric and based on depth functions for directional data. Using the so-called DD-plot, we apply the k-nearest neighbors method in order to discriminate between competing groups. Three different notions of data depth for directional data are considered: the angular simplicial, the angular Tuley and the arc distance. We investigate and compare their performances through the average accuracy rate by means of simulated and real data sets.
Directional data are constrained to lie on the unit sphere of~$\mathbb{R}^q$ for some~$q\geq 2$. To address the lack of a natural ordering for such data, depth functions have been defined on spheres. However, the depths available either lack flexibility or are so computationally expensive that they can only be used for very small dimensions~$q$. In this work, we improve on this by introducing a class of distance-based depths for directional data. Irrespective of the distance adopted, these depths can easily be computed in high dimensions too. We derive the main structural properties of the proposed depths and study how they depend on the distance used. We discuss the asymptotic and robustness properties of the corresponding deepest points. We show the practical relevance of the proposed depths in two applications, related to (i) spherical location estimation and (ii) supervised classification. For both problems, we show through simulation studies that distance-based depths have strong advantages over their competitors.
The box-and-whiskers plot is an extraordinary graphical tool that provides a quick visual summary of an observed distribution. In spite of its many extensions, a really suitable boxplot to display circular data is not yet available. Thanks to its simplicity and strong visual impact, such a tool would be especially useful in all fields where circular measures arise: biometrics, astronomy, environmetrics, Earth sciences, to cite just a few. For this reason, in line with Tukey's original idea, a Tukey-like circular boxplot is introduced. Several simulated and real datasets arising in biology are used to illustrate the proposed graphical tool.
This work evaluates the finite sample behavior of ML estimators in network autocorrelation models, a class of auto-regressive models studying the network effect on a variable of interest. Through an extensive simulation study, we examine the conditions under which these estimators are normally distributed in the case of finite samples. The ML estimators of the autocorrelation parameter have a negative bias and a strongly asymmetric sampling distribution, especially for high values of the network effect size and the network density. In contrast, the estimator of the intercept is positively biased but with an asymmetric sampling distribution. Estimators of the other regression parameters are unbiased, with heavy tails in presence of non-normal errors. This occurs not only in randomly generated networks but also in well-established network structures.
Among the measures of a distribution's location, the mode is probably the least often used, although it has some appealing properties. Estimators for the mode of univariate distributions are widely available. However, few contributions can be found for the multivariate case. A consistent direct multivariate mode estimation procedure, called minimum volume peeling, can be outlined as follows. The approach iteratively selects nested subsamples with a decreasing fraction of sample points, looking for the minimum volume subsample at each step. The mode is then estimated by calculating the mean of all points in the final set. The robustness of the method is investigated by analyzing its finite sample breakdown point and algorithms to determine minimum volume sets are discussed. Simulation results confirm that using minimum volume peeling leads to efficient mode estimates both in uncontaminated as well as contaminated situations.
The paper investigates the link between student relations and their performances at university. A social influence mechanism is hypothesized as individuals adjusting their own behaviors to those of others with whom they are connected. This contribution explores the effect of peers on a real network formed by a cohort of students enrolled at a graduate level in an Italian University. Specifically, by adopting a network effects model, the relation between interpersonal networks and university performance is evaluated assuming that student performance is related to the performance of the other students belonging to the same group. By controlling for individual covariates, the network results show informal contacts, based on mutual interests and goals, are related to performance, while formal groups formed temporarily by the instructor have no such effect.
The von Mises-Fisher distribution is probably the most widely used distribution to model data on the (hyper)sphere. Characterized by two parameters, the mean direction and concentration parameter, it is symmetrical about the mean direction. The Maximum Likelihood estimator for this parameter is the directional sample mean. However, this estimator is not robust. Although the estimation result cannot be corrupted arbitrarily (but only up to a certain amount) the need for some alternative (robust) estimators has been recognized. Within this work, it is first discussed how the robustness of a directional location estimator can be evaluated. Then, a new robust estimator of the mean direction is introduced.