The article is devoted to the problem of constructing a semantic library of resources on mathematics and mathematical physics based on classical encyclopedias. The mechanism of integration of encyclopedias into the content of the semantic library and the unification of the mathematical encyclopedia edited by academician I.M. Vinogradov and the encyclopedia of mathematical physics edited by academician L.D. Faddeev are investigated. During the integration process, intersections of multiple articles in these encyclopedias are discovered, as well as mutual enrichment of descriptions of their terms. The library’s tools made it possible to form a knowledge graph into which both encyclopedias were integrated. Thanks to the ontological approach, the knowledge graph of the semantic library is saturated with new nodes and links, which in turn leads to the enrichment of the subject areas of the semantic library itself and the subject areas of integrated scientific publications. The library search is accompanied by navigation based on a knowledge graph, which allows you to rely on reliable information from classic encyclopedic sources. The work is addressed to specialists in the field of semantic modeling of scientific subject areas.
The paper considers an approach to a knowledge graph construction based on the ontological representation of scientific subject areas. The presentation is based on concepts related to information and data mining, such as knowledge, knowledge extraction, domain ontology, scientific domain, thesaurus, semantic digital library, user information need, ontological design method and, in fact, the knowledge graph. The digital semantic library LibMeta is presented as a repository of various structured data with the possibility of their integration with other data sources. Assumes the possibility of specifying personal content by describing a local subject area within LibMeta. The ontology of the content of the semantic library acts as a means of formalization. This paper addresses the experience of building semantic libraries based on thesauri and ontological design. Building ontologies based on the thesaurus of the subject area LibMeta allows us to say that the presence of internal semantic links ensures the consistency and reliability of search results, which is a necessary condition for extracting scientific knowledge. The digital library ontology defines the data structure of the library content. Each data element loaded into the library can be associated with an ontology vertex (top) that determines the position of the data element in the ontology. Based on the ontology links and the links defined at the design stage, you can build a data graph. On the example of the ontology of the LibMeta semantic library, the technology of forming the knowledge graph of modern applications in mathematics is discussed. The problems of filling a graph, embedding in a graph, extracting links and nodes of a graph are discussed.
The problem of presenting scientific results of academic institute in a digital environment is considered. A new look at the knowledge space of a scientific institute constitutes a natural stage in the development of WEB technologies. The data structure inherent in previous studies allows you to organize search and navigation through them using a knowledge graph, like a version of the semantic library LibMeta. The knowledge graph gives a more complete and high-quality idea of the knowledge space, often removing the cognitive load in the perception of complex structures and data connections.
В работе исследуется тематическое многообразие междисциплинарного журнала. Цель исследований составляет построение графа знаний журнала для тематического представления и систематизации электронного архива и новых публикаций журнала. Исходные данные представляют собой статьи журнала, посвященные различным информационным и математическим технологиям в науке и управлении, то есть междисциплинарным исследованиям. Предлагается систематизация текстов с помощью методов векторного анализа. В процессе тематического анализа контента журнала предлагается разбиение на рубрики, устанавливаются связи рубрик и статей с соответствующими описаниями специальностей ВАК. Для анализа тематики используется разведочный анализ исходных текстов, далее применяются методы интеллектуального анализа данных. Результаты разбиения предоставляются экспертам журнала, после чего вырабатывается решение о формировании тематической рубрики и включении в нее специальностей ВАК. Статьи журнала интегрируются в семантическую библиотеку LibMeta, в силу чего онтология библиотеки достраивается и формируется онтология журнала, и на этой основе строится граф знаний журнала. Предлагается процедура навигации по контенту журнала с помощью графа знаний в семантической библиотеке LibMeta, которая может стать основой для информационного сопровождения научных исследований и создания цифрового ассистента в междисциплинарной предметной области. Примеры приведены для конкретного контента журнала, но предложенная технология может быть распространена на другие журналы, так как большинство журналов, относящихся к нескольким специальностям ВАК, естественным образом захватывают несколько дисциплин.
The paper discusses an approach to constructing a personal knowledge graph based on an ontological description of a scientific subject area and its data, presented in the form of a knowledge graph. The presentation is based on concepts related to information and data mining, such as subject domain ontology, scientific subject area, thesaurus, semantic digital library, personal knowledge graph. The procedure for constructing a personal knowledge graph is presented using the example of the mathematics subject area, the ontology of which is based on extensive material from the mathematical encyclopedia edited by academician I.M. Vinogradov within the framework of the semantic library LibMeta. The proposed approach will make it possible to use the content of the Mathematics semantic library for scientific research, minimizing the process of searching for information in the local subject area, without losing more general results contained outside this area. This work refers to the experience of building semantic libraries based on thesauruses and ontological design. Construction of ontologies based on the thesaurus of the subject area is a necessary condition for constructing a personal knowledge graph of the scientific subject area.
The problem of constructing a knowledge graph for previously unprocessed scientific texts is studied. The task is to analyze the use of various machine processing methods for extracting text semantics using the example of an interdisciplinary subject area. The goal of the research is to create a semantic model of the subject area of an interdisciplinary scientific journal and use a knowledge graph to navigate through data. The semantic model is based on the ontology of the LibMeta library. Texts and metadata receive connections due to the processing procedure and become part of the ontology of the semantic library. The interdisciplinary subject area is integrated into the library content as a related area of the ‘‘mathematics’’ ontology of the LibMeta library. A comparison of texts is carried out according to different characteristics, thematic proximity, commonality of research methods and approaches, problem setting and their symbolic representation. The results of preliminary text processing and the results in the form of a knowledge graph integrated into the library content are presented. The formulation of the problem of constructing an ontology and a knowledge graph for an interdisciplinary subject area adjacent to mathematics is considered by the authors as part of the general problem of managing mathematical knowledge.
An ontological approach to the construction of a thesaurus of the applied subject area of a scientific academic journal based on the content of the LibMeta semantic digital library and subject area sources is considered. A thesaurus and a knowledge graph of the subject area of scientific publications of the journal of the Russian Academy of Sciences Mechanics of Composite Materials and Structures , devoted to the modeling of composite materials and structures, is being built. The approach is based on the use of the content of the semantic library LibMeta, classical primary sources on the theory of elasticity and metadata of the journal articles. As a result, a thesaurus and a knowledge graph of the subject area of the journal are built, and on their basis the user interface is modified and a variant of navigation through the links of articles is proposed. Examples of constructions are given.
The paper studies the problem of developing a semantic library by adding a new applied scientific area.The authors use the example of a journal on applied issues of composite materials in order to build an
The paper uses the ontological design approach to describe the semantics of some boundary value problems from the field of solid mechanics. Modern materials modeled within the framework of the theory of elasticity require new formulations of classical problems of mathematical physics. To describe new problems, it is necessary to establish connections between new terms and concepts with the classical definitions of the mathematical encyclopedia and other primary sources. Establishing links allows you to form a dictionary and thesaurus of the applied subject area of new boundary value problems and place the results in the semantic environment of the digital library. Examples of this approach are demonstrated using the capabilities of the LibMeta semantic library, which contains a digitized version of the mathematical encyclopedia and encyclopedia of mathematical physics, classifiers, and applied mathematical thesauri and dictionaries. The purpose of the research is to provide the user with additional services in the search for publications in the applied scientific field.
The article focus on problem of developing a semantic library by adding a new applied scientific field. The addition to the main content of the library is built using publications of a journal on applied issues of composite materials. The description of the original subject area is expanded. Universal Decimal Classification and Mathematics Subject Classification articles are detailed by corresponding to the local subject area. At the same time, the tasks of adding terms to the thesaurus, building a reference corpus of the applied subject area of mathematics, and creating a custom interface are solved. Formulas and equations of the local subject area are semantically linked to the main content of the library. The main advantage of using semantic libraries for this kind of tasks is to enrich the existing knowledge base of the library and identify relationships in the data. To study these problems, it is necessary to interact with subject matter experts and use modern tools and methods for natural language processing, machine learning, and approaches to knowledge representation. Representation of knowledge in the form of ontologies and thesauri connects the definitions of key concepts of the subject area not only within the same scientific school, but also allows experts and users to ''communicate'' in the same language. As a global thesaurus of the subject area, terminologically delineating its boundaries, the data of the Mathematical Encyclopedia, a Soviet encyclopedic edition integrated into the library, are used. Integration of data within the library allows expanding the description of subject areas related to the applications of mathematics in interdisciplinary research and technology. On the example of one of the applied sections of problems of mathematical physics, the procedure is shown for including arrays of publications of a specialized journal into the ontology of a semantic library. The ontology is based on available data sources, library content, and specific dictionaries and thesauri. The proposed approach will allow using the content of the semantic library ''Mathematics'' for scientific research, minimizing the process of searching for information in the local subject area, without losing more general results contained outside this area.
The work is devoted to the problem of constructing an ontology of an applied subject area as part of a semantic library. The initial data are presented by publications of a scientific thematic journal, an array of archival and current versions. The main stages of ontology formation are considered in accordance with the definition and requirements of Web Ontology Language. Examples of the structure of thesaurus articles are given.
The work is devoted to the problem of customizing the user interfaces of an information system that integrates data. An adaptive interface serves as one of the means of organizing the presentation of subject domain data. The issue of using the semantic relations of ontology to select data corresponding to the objectives of the study is investigated. A model of an adaptive interface is considered, which allows the most accurate reflection of the needs of a researcher within a particular subject domain. It is shown how the adaptive interface is formed by means of the semantic library model.
This article is devoted to the problem of including a subject area in a semantic library. The subject area is considered, which was not previously presented in part of the ontology of the digital library. A method based on the technology of establishing semantic links with the already accumulated content of the library is proposed. Main idea of the inclusion method is to use dictionaries, thesauri and links to the mathematical encyclopedia. They are used to attach a new scientific publication metadata set to an existing library metadata set. In the course of preliminary processing of texts of scientific publications, their semantic analysis is carried out and a local description of the subject area of this array is formed within the content of the library. Such subject areas as ordinary differential equations, partial differential equations, equations of continuum mechanics, equations of composite mechanics and their solutions expressed in terms of special functions of mathematical physics are considered. The description of the mathematical encyclopedia in its classical form is the central resource. It allows us to identify the appropriate classifiers through the links of the equations of applied problems and their solutions and use them also to form a description of the subject area. This subject area is new and has not been described in our digital library before. Proposed procedure is made possible only by data integration in the library. This takes into account the expansion of the description of subject areas related to the applications of mathematics in interdisciplinary research and technology. The implementation of this procedure is one of the new directions for filling the semantic libraries. It uses the approach of data integration and saturation with connections that are revealed only in the process of analyzing new data. If new references are found in the library during the process of adding a new description, then this method will also make it possible to identify and fill in the gaps in the description of related areas of mathematics, such as applications of solutions to classical problems of mathematical physics.
The problem of finding the most relevant documents as a result of an extended and refined query is considered. To solve it, a search model and a text preprocessing mechanism are proposed. It is proposed to use a search engine and a model based on an index using word2vec algorithms to generate an extended query with synonyms. To refine the search results, the idea of selecting similar documents in the digital semantic library is used. The paper investigates the construction of a vector representation of documents in relation to the data array of the digital semantic library LibMeta. Each piece of text is labeled. Both the whole document and its separate parts can be marked. Search through the library content, search for new terms and new semantic relationships between terms of the subject area becomes more meaningful and accurate. The task of enriching user queries with synonyms was solved. When building a search model in conjunction with word2vec algorithms, a "indexing first, then learning" approach is used, which allows obtaining more accurate search results. This work can be considered one of the first stages in the formation of a training data array for the subject area of problems of mathematical physics and the formation of a dictionary of synonyms for this subject area. The model was trained on the basis of the library's mathematical content. Examples of training, extended query and search quality assessment using training and synonyms are given.
The problem of finding the most relevant documents as a result of an extended and refined query is considered. For this, a search model and a text preprocessing mechanism are proposed, as well as the joint use of a search engine and a neural network model built on the basis of an index using word2vec algorithms to generate an extended query with synonyms and refine search results based on a selection of similar documents in a digital semantic library. The paper investigates the construction of a vector representation of documents based on paragraphs in relation to the data array of the digital semantic library LibMeta. Each piece of text is labeled. Both the whole document and its separate parts can be marked. The problem of enriching user queries with synonyms was solved, then when building a search model together with word2vec algorithms, an approach of "indexing first, then training" was used to cover more information and give more accurate search results.
The peculiarities of the task of authors identifying and determining author's contribution to publications in digital bibliographic codes are considered. The features of the problem of insufficient identification are manifested in the repetition of information, doubling, the presence of authors with completely coincidental names, self-quotation, autoplagiate and plagiarism itself. It is proposed to use publication information that has already been accumulated in the digital library in the form of related object area data and a variety of target thesaurus data, as the author and user of the library. This information contains links whereby keyword contexts, multiple co-authors, and term associations in dictionaries and thesauruses can be used to identify authorship. It is important that an array of scientific publications is considered, since they have an established traditional structure, which allows comparing fixed text elements (annotations, keywords, classifier codes, etc.). Thus, even if the names in the publications are fully matched, the question of authorship can be raised if the publications in the digital library correspond to different subject areas. Resolution of such contradictions is accomplished by evaluating a plurality of links of all elements of secondary publication information. The result of the comparison could be the addition of the author to a specific area, i.e. the extension of the addressee's thesaurus and the author's personal thesaurus, or the appearance of full namesakes in the library, but from different areas of knowledge. It has been shown that modern data analysis tools allow you to evaluate the author's contribution to publication, despite the fact that of course, only the scientific community can evaluate the real contribution to scientific research.
The problems of extracting the most complete information from the semantic library by accounting for related documents are considered. Expert knowledge encrypted in the subject area can be made available when the user obtains additional information from linked documents. A feature of the approach is the use of a shallow neural network algorithm to expand the search query in mathematical subject areas, where expert knowledge is available with a significant scientific background of users. The solution to this problem can be achieved by means of semantic analysis in the knowledge space using machine learning algorithms. The paper investigates the construction of a vector representation of documents based on paragraphs in relation to the data array of the digital semantic library LibMeta. Each piece of text is labelled. Both the whole document and its separate parts can be marked. Since the problem of enriching user queries with synonyms was solved, when building a search model in conjunction with word2vec algorithms, an approach of “indexing first, then training” was used to cover more information and give more accurate results.
The volume of scientific information appearing in the world today creates serious problems with its processing and use. In this regard, it becomes necessary to filter scientific information and extract from it knowledge that is novel and unique. The aggregate of such information in digital form, together with the tools to ensure its updating, preservation and provision to users, is defined as a common digital space of scientific knowledge. It consists of a set of thematic subspaces related to various areas of science, built on the same principles. Despite the fact that there are some examples of formalization of knowledge in different subject areas, there is no generalized approach to defining the digital space of scientific knowledge. This paper discusses the problems and tasks of the formation of a such space and the development of tools that provide the ability to study information in it.
The paper considers an information system designed to represent a subject area related to science and its features. Highlighted general concepts for formal description of such a subject area in the knowledge base of the semantic library. The peculiarity of these areas is that the data structure is subject to frequent changes. Therefore, the means of organizing knowledge, which is a semantic library, should be sufficiently universal and not require deep technical knowledge. The paper describes the functionality of the system and its use.
The paper addresses the issue of filling the gaps in the semanticlibrary based on the distributional semantics of the terms of itsthesaurus and the ontology relations. The goal of the study is tofully reflect the actual structure of relations betweenmathematical subject domains. This is done through identifyingcontext-sensitive semantic relations, and with the use of analgorithm that is based on the word2vec feedforwardneural networks. The understanding of the query is analyzed afterpreliminary processing of set of articles and metadatasaturation. The proposed procedure helps to improve the work withthe full-text index and, as a result, improves the quality ofsearch in the library. Using a full-text index of a digitalsemantic library as an example, we demonstrate the process offilling gaps by saturating the semantic relations of the ontologyof mathematical subject domains.