Standard multidimensional scaling takes as input a dissimilarity matrix of general term δ _ij which is a numerical value. In this paper we input δ _ij=[δ _ij,δ _ij] where δ _ij and δ _ij are the lower bound and the upper bound of the “dissimilarity” between the stimulus/object S_i and the stimulus/object S_j respectively. As output instead of representing each stimulus/object on a factorial plane by a point, as in other multidimensional scaling methods, in the proposed method each stimulus/object is visualized by a rectangle, in order to represent dissimilarity variation. We generalize the classical scaling method looking for a method that produces results similar to those obtained by Tops Principal Components Analysis. Two examples are presented to illustrate the effectiveness of the proposed method.
Pyramidal clustering method generalizes hierarchies by allowing non-disjoint classes at a given level instead of a partition. Moreover, the clusters of the pyramid are intervals of a total order on the set being clustered. [Diday 1984], [Bertrand, Diday 1990] and [Mfoumoune 1998] proposed algorithms to build a pyramid starting with an arbitrary order of the individual. In this paper we present two new algorithms name CAPS and CAPSO. CAPSO builds a pyramid starting with an order given on the set of the individuals (or symbolic objects) while CAPS finds this order. These two algorithms allows moreover to cluster more complex data than the tabular model allows to process, by considering variation on the values taken by the variables, in this way, our method produces a symbolic pyramid. Each cluster thus formed is defined not only by the set of its elements (i.e. its extent) but also by a symbolic object, which describes its properties (i.e. its intent). These two algorithms were implemented in C++ and Java to the ISO-3D project.
Standard multidimensional scaling takes as input a dissimilarity matrix of general term $\delta _{ij}$ which is a numerical value. In this paper we input $\delta _{ij}=[\underline{\delta _{ij}},\overline{\delta _{ij}}]$ where $\underline{\delta _{ij}}$ and $\overline{\delta _{ij}}$ are the lower bound and the upper bound of the ``dissimilarity'' between the stimulus/object $S_i$ and the stimulus/object $S_j$ respectively. As output instead of representing each stimulus/object on a factorial plane by a point, as in other multidimensional scaling methods, in the proposed method each stimulus/object is visualized by a rectangle, in order to represent dissimilarity variation. We generalize the classical scaling method looking for a method that produces results similar to those obtained by Tops Principal Components Analysis. Two examples are presented to illustrate the effectiveness of the proposed method.
The paper draws attention to the use of Symbolic Data Analysis (SDA) in the field of Official Statistics. It is composed of three sections presenting three pilot techniques in the field of SDA. The three contributions range from a technique based on the notion of exactly unified summaries for the creation of symbolic objects, a model-based approach for interval data as an innovative parametric strategy in this context, and measures of similarity defined between a class and a collection of classes based on the frequency of the categories which characterize them. The paper shows the effectiveness of the proposed approaches as prototypes of numerous techniques developed within the SDA framework and opens to possible further developments.
Our aim is to introduce a new measure which expresses the “concordance” and “discordance” between a class and a collection of classes denoted P, of a given population. We first define two basic functions $$f_{c}$$ and $$g_{x}$$ , where $$f_{c} \left( x \right)$$ expresses the fit of the representation $$x$$ with $$c$$ and $$g_{x} \left( {c,P} \right)$$ expresses the proportion of classes $$c^{\prime}$$ of $$P$$ having a fit and a representation $$x^{\prime}$$ to the class $$c^{\prime}$$ close to that of $$c$$ . We show, for example, that by using dual scaling (Nishisato in Analysis of categorical data: dual scaling and its applications. University of Toronto Press, Toronto, 1980; Nishisato in elements of dual scaling: an introduction to practical data analysis. Lawrence Erlbaum Associates, Hillsdale, NJ, 1994; Nishisato in Multidimensional nonlinear descriptive analysis. Chapman & Hall/CRC, Boca Raton, FL, 2014), from a table describing each European country by socio-demographic variables, we can obtain from the higher to the lower concordance or discordance a ranking of all European countries. Then, we give the Axiomatic definitions of an s-concordance and s-discordance and examples of s-concordance and s-discordance families. We show that there exist useful links between concordances and copulas. A useful general formulation of the classical likelihood function where there underlying classes are given. In the case where $$P$$ is unknown at the beginning, we give a general formulation of mixture decomposition—by the dynamic clustering method (DCM)—taking into account the concordance or discordance and allowing to construct P and the probability density representation of each of its classes. We finally give a way to visualise (in 2D or 3D) clusters of classes of P based on concordance.
In this paper we present a model of the stock exchange domain using symbolic dataanalysis and we use the SODAS software to analyze this domain. After a short presentationof the software, we present the analysis in three steps: choice of the symbolic objects, theirdefinition and their analysis with SODAS. We give details for each of these steps and thereimportance is underlined. Two examples of results are described to show the analysis interestand pertinence. The conclusion describes perspectives after the improvement of SODAS forits application in the stock exchange domain.
We compare two extensions of principal component analysis to distributional variables and we introduce a generalization strategy when the knowledge experts want to perform histogram PCA on several distributional datasets. Actually, most of proposed approaches only consider the situation where users have one dataset and from a technical standpoint, the proposed solutions are based on either the first order moments or the quantiles. In that, we present two Histogram PCAs respectively based on the barycenters of distributions and on the average correlation matrix induced by the quantiles of distributions. We review the benefits and the flaws of using the barycenters versus the quantiles and we present a generalization framework when there are more than one histogram dataset. The methods we described are applicable in many domains like people analytics, risk and control management, internal audit, anti money laundering, healthcare analytics, sports analytics, etc.
Different mortality patterns across countries require different health and demographic policies. Positioning of the countries according to their characteristic mortality pattern can help allocate scarce resources appropriately. We use symbolic data analysis within SYR software to analyse 28 European Union countries' sex-, age-, and cause-specific mortality in 2015. There are two main advantages for using symbolic analysis: (i) it permits more transparent and informative data descriptions along with contextual relations, and (ii) advanced methods adapted for complex data representations can be employed to analyse such data, taking contextual relations into account. Clustering results based on symbolic data analysis show that groups of countries are strongly related to the geographical position of countries, with a clear east–west cut on the first-level partition and with an even more geographically consistent lower-level partition compared to the classical clustering result. Relations between the obtained clusters of countries and their external social and health indicators are well pronounced. We also identify the mortality rates as symbolic variables that discriminate the most between individual countries as well as between the resulting clusters. Knowledge of a country's mortality pattern and its position among comparable countries is valuable information for health and demographic policymakers and can be exploited to exchange good practices.
Covers everything readers need to know about clustering methodology for symbolic data—including new methods and headings—while providing a focus on multi-valued list data, interval data and histogram dataThis book presents all of the latest developments in the field of clustering methodology for symbolic data—paying special attention to the classification methodology for multi-valued list, interval-valued and histogram-valued data methodology, along with numerous worked examples. The book also offers an expansive discussion of data management techniques showing how to manage the large complex dataset into more manageable datasets ready for analyses.
Chapter 2 Likelihood in the Symbolic Context Richard Emilion, Richard EmilionSearch for more papers by this authorEdwin Diday, Edwin DidaySearch for more papers by this author Richard Emilion, Richard EmilionSearch for more papers by this authorEdwin Diday, Edwin DidaySearch for more papers by this author Book Editor(s):Edwin Diday, Edwin DidaySearch for more papers by this authorRong Guan, Rong GuanSearch for more papers by this authorGilbert Saporta, Gilbert SaportaSearch for more papers by this authorHuiwen Wang, Huiwen WangSearch for more papers by this author First published: 15 January 2020 https://doi.org/10.1002/9781119695110.ch2 AboutPDFPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShareShare a linkShare onFacebookTwitterLinked InRedditWechat Summary In this chapter, the authors propose a probabilistic framework for properly defining symbolic data as statistical units, modifying the framework proposed in a little bit. They consider the problem of defining distributions on symbols or aim to propose some likelihood functions for finite-dimensional symbols. The authors discuss symbols and likelihood on symbols with respect to a class variable is rigorously defined in a probabilistic setting. They present density models for one variable of probability vector symbols in the parametric and nonparametric case. The authors also discuss the interesting latent Dirichlet allocation model which is popular in text mining, text classification, and can be used in various domains such as collaborative filtering. Instead of proposing parametric models, a nonparametric or semiparametric approach is possible to estimate the density. Advances in Data Science: Symbolic, Complex and Network Data, Volume 4 RelatedInformation
Data science unifies statistics, data analysis and machine learning to achieve a better understanding of the masses of data which are produced today, and to improve prediction. Special kinds of data (symbolic, network, complex, compositional) are increasingly frequent in data science. These data require specific methodologies, but there is a lack of reference work in this field. Advances in Data Science fills this gap. It presents a collection of up-to-date contributions by eminent scholars following two international workshops held in Beijing and Paris. The 10 chapters are organized into four parts: Symbolic Data, Complex Data, Network Data and Clustering. They include fundamental contributions, as well as applications to several domains, including business and the social sciences.
Une nouvelle facon d analyser les donnees classiques, complexes et massives a partir des classes Applications avec Syr et R La numerisation croissante de notre societe alimente des bases de donnees de taille grandissante (Big Data). Ces donnees sont souvent complexes (heterogenes et multi-tables) et peuvent etre la source de creation de valeur considerable a condition qu elles soient exploitees avec des methodes d analyse adequates. Un « Data Scientist » a justement pour objectif d extraire des connaissances de ce type de donnees et c est l objectif de cet ouvrage. Les classes constituent un pivot central de la decouverTe de connaissances. En Analyse des Donnees Symboliques (ADS), les classes sont decrites par des variables dites symboliques prenant en compte leur variabilite interne sous forme de distributions, d intervalles, d histogrammes, de diagrammes de frequences, etc. Le livre debute par la construction de differents types de variables symboliques a partir de classes donnees. Des statistiques descriptives, une methode de discretisation automatique adaptee aux donnees massives (Big Data) suivies par des indices de proximite etendus aux donnees symboliques y sont presentes. Vient ensuite un ensemble de methodes presente dans le contexte de l ADS. Il s agit de la methode des nuees dynamiques (MND), de la decomposition de melange par partition (issue de la MND) ou par partition floue (EM), de l analyse en composantes principales, de l algorithme Apriori, des regles d association et des arbres de decision. Pour la prevision, le livre presente des methodes de regressions dont celles penalisees « ridge », « lasso » et « elastic », et des series temporelles. Pour la mise en application de ces premieres methodes, des exercices et des applications concretes realisees aupres d administrations, d industriels, de financiers et de scientifiques sont proposes. Leur mise en uvre s appuie aussi bien sur le logiciel innovant Syr que sur le logiciel statistique R. Cet ouvrage d introduction a l ADS s adresse aux etudiants, aux ingenieurs, aux universitaires, ainsi qu a tous ceux qui desirent comprendre cette nouvelle facon de penser en Science des Donnees.
Many nice machine learning methods are black box producing very efficient rules but hard to be understandable by the users. The aim of this paper is to help user by tools allowing a better comprehension of these rules. These tools are based on characteristic properties of the original variables in order to remain in the natural language of the user. They are based on three principles, first on local models fitting at best clusters to be found, second on a symbolic description of these clusters and their Symbolic Data Analysis, third on characteristic criterion increasing the explanatory power of the rules by an adaptive process filtering explanatory sub populations.