The influence of authors is mostly based on their capacity to form specific self-contained and/or active research communities or topics while also inspiring fruitful spin-off research derived from those communities or topics. Accurately estimating author influence and its indirect effect in inspiring external research creativity must thus be based on very precise community roles as well as topic-role analysis on citing material. In this paper we thus implement, compare, and combine two different kinds of knowledge mapping approaches used to track authors’ influence over time—a recent topic-based mapping approach and an original hybrid community-based mapping approach. The experimental context of our study is the analysis of the dynamics and of the influence of the research of Prof. Liu Zeyuan, the most important contributor in the field of Science of Science in China. This analysis focuses on papers citing Professor Liu Zeyuan's research and highlights the full extent of his creative thinking.
Dans une première partie de cet article, nous mettons en lumière le contexte historique de la Science de la Science en Chine et à l'échelle mondiale.Dans une deuxième partie, en utilisant la combinaison d'un clustering GNG (gaz de neurones), des mesures de maximisation des traits et des graphes de contraste, nous effectuons une analyse du contenu d'articles de revues académiques sélectionnées dans le domaine de la Science de la Science en Chine et construisons une carte globale de la recherche au cours des 40 dernières années.De plus, nous mettons en évidence l'évolution du domaine en exploitant les dates de publication et les informations auteurs afin de clarifier le contenu des sujets.Les résultats obtenus, validés par l'expertise, montrent clairement que la Science de la Science en Chine a progressivement mûri au cours des 40 dernières années, passant de la nature générale de la discipline aux disciplines connexes et à leurs interactions potentielles, de l'analyse qualitative à l'analyse quantitative et visuelle, et de la recherche générale sur la fonction sociale de la science aux études plus spécifiques sur sa fonction économique et stratégique.La méthode originale proposée permet d'obtenir sans supervision, sans paramètres et sans connaissances externes une vision à la fois très claire et très précise du développement d'un domaine scientifique.ABSTRACT.In a first part of this paper, we highlight the historical context of Science of Science both in China and at a world level.In a second part, based on the unsupervised combination of GNG (neural gas) clustering with feature maximization metrics and associated contrast graphs, we perform an analysis of the contents of selected academic journal papers in Science of Science in China and the construction of an overall map of the research topic structure during the last 40 years.Furthermore, we highlight the topic evolution by the exploitation of the publication dates and make additional use of the author's information for the sake of clarifying topics content.The obtained results, validated by domain experts, interestingly show that the Chinese Science of Science has gradually become mature in the last 40 years, turning from the general nature of the discipline to the relative disciplines and their potential interactions, from the qualitative analysis to the quantitative and visual analysis, and from the general research on social function of science to more specific economic function and strategic function studies.Consequently, the proposed novel method permits without supervision, without parameters and without help of any external knowledge to have very clear and very precise insights of the development of a scientific domain.
In the first part of this paper, we shall discuss the historical context of Science of Science both in China and at world level. In the second part, we use the unsupervised combination of GNG clustering with feature maximization metrics and associated contrast graphs to present an analysis of the contents of selected academic journal papers in Science of Science in China and the construction of an overall map of the research topics' structure during the last 40 years. Furthermore, we highlight how the topics have evolved through analysis of publication dates and also use author information to clarify the topics' content. The results obtained have been reviewed and approved by 3 leading experts in this field and interestingly show that Chinese Science of Science has gradually become mature in the last 40 years, evolving from the general nature of the discipline itself to related disciplines and their potential interactions, from qualitative analysis to quantitative and visual analysis, and from general research on the social function of science to its more specific economic function and strategic function studies. Consequently, the proposed novel method can be used without supervision, parameters and help from any external knowledge to obtain very clear and precise insights about the development of a scientific domain. The output of the topic extraction part of the method (clustering + feature maximization) is finally compared with the output of the well-known LDA approach by experts in the domain which serves to highlight the very clear superiority of the proposed approach.
[étude] Le projet ISTEX (initiative d’excellence en Information Scientifique et Technique) a pour objectif de permettre à la communauté scientifique française d’accéder à une bibliothèque numérique pluridisciplinaire en texte intégral regroupant l’essentiel des publications scientifiques mondiales. Nous développerons ici les actions R&D engagées pour enrichir les données brutes ainsi qu’un nouveau processus de diffusion d’ISTEX selon les standards du web sémantique (LOD).
Feature maximization (F-max) is an unbiased quality estimation metric of unsupervised classification (clustering) that favours clusters with a maximal feature F-measure value. In this article we show that an adaptation of this metric within the framework of supervised classification allows efficient feature selection and feature contrasting to be performed. We experiment the method on different types of textual data. In this context, we demonstrate that this technique significantly improves the performance of classification methods as compared with the use of state-of-the art feature selection techniques, notably in the case of the classification of unbalanced, highly multidimensional and noisy textual data gathered in similar classes.
In this paper we first propose a state of the art on the methods for the visualization and for the interpretation of textual data, and in particular of scientific data. We then shortly present our contributions to this field in the form of original methods for the automatic classification of documents and easy interpretation of their content through characteristic keywords and classes created by our algorithms. In a second step, we focus our analysis on the data evolving over time. We detail our diachronic approach, especially suitable for the detection and for visualization of topic changes. This allows us to conclude with Diachronic’Explorer, our upcoming visualization tool for visual exploration of evolutionary data.
We introduce Diachronic'Explorer, a toolbox to produce and visualize diachronic results, which is based on a new complete theoretic framework that we detail. This toolbox, which is dedicated to run diachronic algorithms from clustering results, allows also to explore the complex results at all the granularity levels through a web application.
This paper deals with a major challenge in clustering that is optimal model selection. It presents new efficient clustering quality indexes relying on feature maximization, which is an alternative measure to usual distributional measures relying on entropy, Chi-square metric or vector-based measures such as Euclidean distance or correlation distance. First Experiments compare the behavior of these new indexes with usual cluster quality indexes based on Euclidean distance on different kinds of test datasets for which ground truth is available. This comparison clearly highlights altogether the superior accuracy and stability of the new method on these datasets, its efficiency from low to high dimensional range and its tolerance to noise. Further experiments are then conducted on "real life" textual data extracted from a multisource bibliographic database for which ground truth is unknown. These experiments show that the accuracy and stability of these new indexes allow to deal efficiently with diachronic analysis, when other indexes do not fit the requirements for this task.
Feature maximization is an alternative measure to usual distributional measures relying on entropy or on Chi-square metric or vector-based measures such as Euclidean distance or correlation distance. One of the key advantages of this measure is that it is operational in an incremental mode both on clustering and on traditional classification. In the classification framework, it does not present the limitations of the aforementioned measures in the case of the processing of highly unbalanced, heterogeneous and highly multidimensional data. We shall present a new application of this measure in the clustering context for the creation of new cluster quality indexes which can be efficiently applied for a low-to-high dimensional range of data and which are tolerant to noise. We shall compare the behavior of these new indexes with usual cluster quality indexes based on Euclidean distance on different kinds of test datasets for which ground truth is available. This comparison clearly highlights the superior accuracy and stability of the new method.
Feature maximization is a cluster quality metric which favors clusters with maximum feature representation as regard to their associated data. In this paper we show that a simple adaptation of such metric can provide a highly efficient feature selection and feature contrasting model in the context of supervised classification. The method is experienced on different types of textual datasets. The paper illustrates that the proposed method provides a very significant performance increase, as compared to state of the art methods, in all the studied cases even when a single bag of words model is exploited for data description. Interestingly, the most significant performance gain is obtained in the case of the classification of highly unbalanced, highly multidimensional and noisy data, with a high degree of similarity between the classes.
This paper focuses on a subtask of the QUAERO research program, a major innovating research project related to the automatic processing of multimedia and multilingual content. The objective discussed in this paper is to propose a new method for the classification of scientific papers, developed in the context of an international patents classification plan related to the same field. The practical purpose of this work is to provide an assistance tool to experts in their task of evaluation of the originality and novelty of a patent, by offering to the latter the most relevant scientific citations. This issue raises new challenges in categorisation research as the patent classification plan is not directly adapted to the structure of scientific documents, classes have high citation or cited topic and that there is not always a balanced distribution of the available examples within the different learning classes.