
This special issue is a following to the workshop entitled " Theory and Applications of High dimensional Complex and Symbolic Data Analysis in Economics and Management Science " which took place on October, 27 to October, 30 2011, at the School of Economics and Management, Beihang University, Beijing, China. 37 participants from 9 different countries (Belgium, Brasil, China, France, Italy, Japan, Portugal, elevenia, USA) attended this workshop. 24 communications have been presented. This issue corresponds to an extended version of the best contributions as well as original submissions in the same field; all have been submitted to a long and severe refereeing process. The major part of the 12 papers of this special issue are about symbolic data analysis as well in an exploratory context (PCA of interval data), as in a predictive one (regression with histogram data). The other contributions give an overview of applications in some very active fields of Data Mining: social networks, recommendation systems, grouped data.
We address the problem of clustering individuals described with several mixed variables divided in homogeneous blocks. We propose a hierarchical method with two levels to partition the individuals. The method is based on two successive steps using mixed topological maps combined with agglomerative hierarchical clustering. The proposed approach allows to take into account simultaneously qualitative and quantitative variables as well as the variable blocking. A real example on indoor air quality illustrates the proposed method.
With the expansion of internet to advertise, the number of potential channels is increasing every day. In the Human Resource domain, recruiters have to choose between hundreds of job search web sites when they post a job offer on the internet. In order to save costs, assessing job board expected performance has become necessary. In this paper, three recommender systems providing job board performance estimation for a given job posting are introduced. This work refers principally to the new item problem, which is still a challenging topic in the literature. The first system (PLS-R) is a content-based approach, while others are hybrid recommendation approaches. Estimation is made on item neighborhood according to a ?naive? similarity or a supervised similarity measure. These predictive algorithms are compared through experiments on a real dataset. In this application, supervised similarity-based system overcomes the lacks of other approaches and outperforms them.
. Histograms are commonly used for representing summaries of observed data and they can be considered non parametric estimates of probability distributions. Symbolic Data Analysis formalized the concept of histogram symbolic variable , as a variable which allows to describe statistical units by histograms instead of single values. In this paper we present a linear regression model for multivariate histogram variables. We use a Least Square estimation method where the sum of squared errors is defined according to the (cid:96) 2 Wasser-stein metric between the observed and the predicted histogram data. Consistently with the l 2 Wasserstein metric, we solve the Least Square computational problem by introducing a suitable inner product between two vectors of histogram data. Finally, measures of goodness of fit are discussed and an application on real data shows some interpretative advantages of the proposed method.
In the enterprise context, people need to exploit and mainly visualize different types of interactions between heterogeneous objects. Graph model seems to be the most appropriate way to represent those interactions. However, the graphs extracted have in general a huge size which makes it difficult to analyze and visualize. An aggregation step is needed to have more understandable graphs in order to allow users discovering underlying information and hidden relationships between entities. In this work, we propose new measures to evaluate the quality of summaries based on an existing algorithm named k-SNAP that produces a summarized graph according to user-selected node attributes and relationships.
Clustering is one of the most common operation in data analysis while constrained is not so common. We present here a clustering method in the framework of Symbolic Data Analysis (S.D.A) which allows to cluster Symbolic Data. Such data can be constrained relations between the variables, expressed by rules which express the domain knowledge. But such rules can induce a combinatorial increase of the computation time according to the number of rules. We present in this paper a way to cluster such data in a quadratic time. This method is based first on the decomposition of the data according to the rules, then we can apply to the data a clustering algorithm based on dissimilarities.
Adaptive Dynamic Clustering Algorithm for Interval-valued Data based on Squared-Wasserstein Distance