This paper presents the methodology of the quantitative study of the dynamics of changes in the thematic structure of collections and vocabulary of terms that were aimed at selecting documents into thematic collections, creating thematic (classification and subject) profiles of these collections in different chronological periods and a comparative analysis of these profiles. The profiles are treated as descriptive sets, with TF-IDF used as the keyword weight. The methodology is tested in two collections, “Robotics and Robotic Systems” and “Intelligent Systems, Artificial Intelligence, and Machine Learning.”
Articles scattered in the VINITI abstract database dated 2020–2022 are studied using the example of three thematic fragments: chemistry and chemical technologies, mechanical engineering, and metallurgy and welding. The mathematical formulas of Bradford and Vickery were used to calculate the number of journals in the third scattering zone. It is argued that the Bradford–Vickery law of publication scattering allows us to predict the minimum number of journals that cover more than 90
Представлен анализ реализаций методов векторизации текстовой информации; описаны выбор модели классификации научных трудов и обучение лингвистической модели BERT на домене научных текстов. Приведены результаты экспериментов по обучению моделей классификации научных статей по первому и второму уровням ГРНТИ. This paper discusses modern approaches to natural language processing and application of machine learning models in the task of short scientific texts classification in Russian. The research is devoted to the analysis of methods for vectorization of textual information, selection of a model for scientific papers classification, training of linguistic model BERT on the domain of scientific texts. The paper presents the results of experiments to train scientific article classification models at the first and second levels of the Russian State Rubricator of Scientific and Technical Information (SRSTI).
The study investigates the dynamics of the Russian scientific journals input stream based on the results of article classification, according to the State Rubricator of Scientific and Technical Information in the VINITI RAS Database. The authors aim to trace the changes in the information provision of the codes of the subject field Chemistry, chemical technology, and the chemical industry. The research procedure presented includes the methods of calculating the indicators that allow to automate the assessment of journals thematic profiles—the similarity and scattering coefficients. The proposed methodology facilitates the identification of trends in the thematic profiles of serial publications, i.e., growth or degradation of their headings, thus contributing to the optimization of the information center documents input stream.
This paper discusses modern approaches to natural language processing and the application of machine learning models to the task of classifying short scientific texts in Russian. This study is devoted to the analysis of methods for vectorization of textual information, selection of a model for scientific paper classification, and training of linguistic model BERT on the domain of scientific texts. This paper presents the results of experiments to train scientific article classification models at the first and second levels of the Russian State Rubricator of Scientific and Technical Information (SRSTI).
Представлены результаты разработки и тестирования системы автоматической классификации научных текстов, позволяющей определять тематику текстов по трём классификационным схемам в пакетном и диалоговом режимах. Описаны структурно-функциональные компоненты, используемые методы оценки качества классификации, методика обучения, выбор оптимальной модели классификации, основные направления внедрения автоматического классификатора в технологию обработки электронного документального потока в ВИНИТИ РАН.
This paper presents the results of the development and testing of an automatic classification system for scientific texts that provides the functionality to determine the topic of texts by three classification schemes in batch and dialog modes. The structural and functional components, the methods used to assess the quality of classification, the teaching methodology, the selection of the optimal classification model, and the main areas for the introduction of an automatic classifier in the processing of electronic document flow at the VINITI RAS are described.
The standard impact factor allows one to compare scientific journals only within particular scientific subjects. To overcome this limitation, another indicator of citation, viz., the thematically weighted impact factor (TWIF), is proposed. This indicator allows one to compare journals of various subjects and takes the fact that a journal belongs to several subjects into account. Information on the thematic headings of a journal and the value of a standard impact factor is necessary for calculation of the indicator. The TWIF, which is calculated according to the citation index of Journal Citation Reports, is investigated in this article.
A new version of the VINITI RAS Electronic Catalog of scientific and technical literature is considered that allows selective and navigational searching for a set of interrelated objects in the sphere of scientific and technical information, such as scientific publications, events, persons, and organizations, whose descriptions are extracted from the input flow of literature during its primary processing. The position of the Electronic Catalog is identified among other documentary information retrieval systems. A conceptual model of the integrated object environment is presented along with the principles of its exposure to the relational database. The features of the retrieval language are described, which allow selective navigational access to data.