Описан подход к построению предметных онтологий пространства знаний на примере ряда научных дисциплин. В качестве входов в систему предлагается использовать тезаурус ключевых терминов, которые представлены в метаданных информационных ресурсов. Построена структура классификаций и тезаурусов, развитие которой до степени полноценной онтологии предполагает разработку тезауруса смысловых связей тематических рубрик, отражающего онтологические свойства объектов информационного поля.
Представлены результаты разработки и тестирования системы автоматической классификации научных текстов, позволяющей определять тематику текстов по трём классификационным схемам в пакетном и диалоговом режимах. Описаны структурно-функциональные компоненты, используемые методы оценки качества классификации, методика обучения, выбор оптимальной модели классификации, основные направления внедрения автоматического классификатора в технологию обработки электронного документального потока в ВИНИТИ РАН.
This paper presents the results of the development and testing of an automatic classification system for scientific texts that provides the functionality to determine the topic of texts by three classification schemes in batch and dialog modes. The structural and functional components, the methods used to assess the quality of classification, the teaching methodology, the selection of the optimal classification model, and the main areas for the introduction of an automatic classifier in the processing of electronic document flow at the VINITI RAS are described.
The article proposes an algorithm for decoding and representation in natural language of the Universal Decimal Classifycation (UDC) complex class numbers. The algorithm is based on the formal definition of correct class numbers using a generative grammar, which sets the list of structures starting with simple table codes of UDC classes. Then separate integers, auxiliary and independent class numbers are sequentially attached to the codes with special symbols of relations of classes, which compose the complex class number. The algorithm expresses the values of the analyzed complex indices by descriptions (names and notes) of the table classes included in the structure of the analyzed string. The class descriptions are accompanied with the logical connectors based on the functions of the auxiliary characters. They provide the idea on connection of concepts denoted in the class number. The algorithm action is described evidently for the analysis of combined index 539.4.019: [535-15+537.8.029.6]. The proposed algorithm is applicable both to visualize the meaning of complex class numbers, and to ensure the completeness and accuracy of the documents retrieval by the UDC classes.
This paper describes a method for constructing a network of classifiers that forms a multidimensional representation of the ontology of scientific and technical information, as well as illustrating a method for investigating the effectiveness of automated determination of semantic relationships between classification headings.
A new version of the VINITI RAS Electronic Catalog of scientific and technical literature is considered that allows selective and navigational searching for a set of interrelated objects in the sphere of scientific and technical information, such as scientific publications, events, persons, and organizations, whose descriptions are extracted from the input flow of literature during its primary processing. The position of the Electronic Catalog is identified among other documentary information retrieval systems. A conceptual model of the integrated object environment is presented along with the principles of its exposure to the relational database. The features of the retrieval language are described, which allow selective navigational access to data.
This paper considers a subsystem for the management of working hours during the operation of automated technology for processing information about the scientific activities of the VINITI RAS, the goals and tasks of this subsystem, and the methods of their solution, as well as forms of presentation of the results. It describes an automated system for processing the information about the scientific activities of the VINITI RAS as the major source of input data for a system for the management of the work hours of personnel, i.e., a subsystem for gathering statistics and salary estimation for the workers of the VINITI RAS who are employed in the processing of information about scientific activities.