Modern product retrieval systems are becoming increasingly complex due to the use of extra product representations, such as user behavior, language semantics and product images. However, adding new information and complicating machine learning models does not necessarily lead to an improvement in online and business search performance, since after retrieval the product list is ranked, which introduces its own bias. Nevertheless, the business performance of a product search will be worse from ranking an incomplete list of products than a complete one, and the relevance of search results will not improve from perfect sorting of products that do not match the search query. Therefore, the main quality indicators for the products retrieval phase remain Recall and Precision at the k threshold. This paper compares several architectures of product retrieval systems in product search for e-commerce. To do this, the concepts of threshold Recall and Precision for information retrieval are investigated and the dependence of these measures on the order of issuance is revealed. An automatic procedure has been developed for calculating R@k and P@k, which allows us to compare the effectiveness of information retrieval systems. The proposed automatic procedure has been tested on the WANDS public dataset for several key architectures. The obtained values R@1000 = 84% ± 9% and P@10 = 67% ± 17% are at the level of SOTA models.
Product search is uniquely different from search for documents, Internet resources or vacancies, therefore it requires the development of specialized search systems. The present work describes the H1 embdedding model, designed for an offline term indexing of product descriptions at e-commerce platforms. The model is compared to other state-of-the-art (SoTA) embedding models within a framework of hybrid product search system that incorporates the advantages of lexical methods for product retrieval and semantic embedding-based methods. We propose an approach to building semantically rich term vocabularies for search indexes. Compared to other production semantic models, H1 paired with the proposed approach stands out due to its ability to process multi-word product terms as one token. As an example, for search queries "new balance shoes", "gloria jeans kids wear" brand entity will be represented as one token - "new balance", "gloria jeans". This results in an increased precision of the system without affecting the recall. The hybrid search system with proposed model scores mAP@12 = 56.1 other SoTA analogues.
A method is developed for conducting comparative analysis on the content of full text patents collections. Named T4C, the approach is based on topic modelling and machine learning and extends comparative text mining. The idea of T4C was inspired by the possibility of precise topics extracting from a joint collection of texts and following analysing the parts of collection on the topics. The different aspects of meta information of the patents full texts collection are considered. The ownership of a patent in a particular country can be identified with an accuracy of 97.5% by using supervised machine learning. By studying how patents vary with time, those belonging to a specific period can be identified with an accuracy of 85% for a given country. Also developed is a visual representation of the thematic correlation between groups of patents. In terms of the text composition of patent descriptions, Chinese patents differ fundamentally from US patents. T4C method is valid for structured medium-sized collections of texts in English. The experimental results are used to manage the patenting process at GazpromNeft STC.
В работе рассматривается задача поиска в научной статье фрагментов с недостающими библиографическими ссылками с помощью автоматической бинарной классификации. Для обучения модели предложен метод контрастного семплирования, новшеством которого является рассмотрение контекста ссылки с учетом границ фрагмента, максимально влияющего на вероятность нахождения в нем библиографической ссылки. Обучающая выборка формировалась из автоматически размеченных семплов — фрагментов из трех предложений с метками классов «без ссылки» и «со ссылкой», удовлетворяющих требованию контрастности: семплы разных классов дистанцируются в исходном тексте. Пространство признаков строилось автоматически по статистике встречаемости термов и расширялось за счет конструирования дополнительных признаков — выделенных в тексте сущностей ФИО, чисел, цитат и аббревиатур.
Visualization of multidimensional data is the most important stage of data research. Often, decisions on the further stages of the study are made from the flat view of the data based on \"rough proportions\". High visibility and persuasiveness of representation on the plane of multidimensional vectors with the preservation of distances is used in models of distributive semantics (Word2Vec, GloVe, NaVec) successfully. On the other hand, the inaccuracy of the two-dimensional projection can lead to time being spent searching for non-existent multidimensional structures. The author set the task to evaluate the accuracy of dimensionality reduction methods with the following limitations: multi-dimensionality arises as a result of vector representation of text documents, dimensionality reduction is aimed at visualization on the plane. In numerous methods of dimension reduction, there is no separate class of approaches specifically for visualization. To measure the accuracy, an approach was chosen using marked-up data and quantifying the preservation of the markup while reducing the dimension. The author investigated 12 methods of reducing the dimension on two labeled data sets in Russian and English. Using the Silhouette Coefficient metric, the most accurate visualization method for text data was determined as UMAP with the Hellinger distance as the metric.
В статье рассматривается задача поиска схожих по смыслу текстовых документов в корпусе. Исследуется проблема невыявления алгоритмом TF-IDF части решений, возникающая при разработке прикладных интеллектуальных информационных систем: потеря пар, схожих согласно человеческой оценке, но получающих низкую оценку схожести от программы. Предложена модификация алгоритма с заменой общего словаря на словарь специализированных терминов. Добавление тезаурусов при построении векторной модели корпуса, основанной на ранжирующей функции, не было ранее исследовано; применение тезаурусов до сих пор изучалось лишь для улучшения тематической модели. Цель работы - повысить качество решения, минимизируя потерю значимой его части и не добавляя «ложно-схожие» пары документов, за счет применения при векторном разложении TF-IDF словаря терминов, выделенного из текста анализируемых документов. Эксперимент проведен поочередно на двух корпусах структурированных нормативно-технических документов, объединенных тематически: стандартов в отношении информационных технологий и в сфере железных дорог. Словарь терминов составлен при автоматическом анализе текста рассматриваемых документов методами выделения именованных сущностей, основанных на правилах. Продемонстрировано, что разложение ТF-IDF по словарю терминов дает больше релевантных результатов для исследуемой задачи, что подтвердило выдвинутую гипотезу. Предложенный метод в меньшей степени зависит от недостатков текстового слоя (таких как ошибки распознавания), чем расчет близости документов по полному словарю корпуса. Определены факторы, способные повлиять на качество решения: способ составления словаря терминов, выбор диапазона n-грамм для словаря, корректность формулировки терминов и обоснованность их включения в глоссарий документа. Полученные выводы могут использоваться при решении прикладных задач, связанных с поиском близких по смыслу документов, таких как семантический поиск с учетом предметной области, корпоративный поиск в многопользовательском режиме, обнаружение скрытого плагиата, выявление противоречий в коллекции документов, определение новизны в документах при построении базы знаний.
Рассмотрено влияние авторизации пользователей на релевантность поисковой выдачи для различных ранжирующих функций. Экспериментально показана степень искажения поисковой выдачи при подокументном доступе к коллекции. Создана модель, позволяющая количественно оценить порог для полного пересчета весов ранжирующей функции. Предложен эффективный алгоритм перерасчета весов с меньшей вычислительной сложностью, чем полный пересчет матрицы «документ – терм».
The average coherence is the key quality assessment parameter for topic models; it reflects most closely the human topic evaluation. Scientific literature describes a variety of methods to calculate topic coherence, each featuring their pros and cons from the point of view of both scientific validity and practical utility. The current paper studies the potential for practical application of different coherence calculation methods in real informational systems, which were tested on two text corpora, namely << Taiga >> (7695 documents). << GOST >> (1066 documents), using several algorithms of topic modeling (LDA, ARTM, PLSA). The problem of soft clustering of a text document collection was considered. The article describes the conducted experiment aimed to reduce the total number of the surveyed model topics to the set of highcoherent topics (HCT). Various ways of documents connection to the high-, mid-, and low-coherent topics have been examined; it has been demonstrated that the applied quality of the topic model is not diminished when the low-coherent topics are removed from the scoop, while depending on the specific features of the objective and the data being processed the coherence threshold value can be reduced. The topic coherence used to analyze the results of transition from the complete model to the set of HCT was calculated using the following formula, determined as the most valid one for the task: PPMI(w(di), w(dj)) = [log n(wd(i,) w(dj))n/n(w(di))/n(w(dj))](+,) where n(w(di), w(dj)) = Sigma(vertical bar D vertical bar)(d=1)Sigma(Nd)(i=1)Sigma(Nd)(j=1) left perpendicular 0 < vertical bar i - j vertical bar <= k right perpendicular is the CoocTF value, defining the co-occurrence of the terms w(i) and w(i) within the sliding window of a preset width in all the document; |D| is the document collection size , N-d is the volume of the document d. The main outcome of the research is the conclusion that topic model application for practical use in informational systems requires focusing not on the average coherence values, but on the topics featuring high values of coherence. To assess the applied quality of a topic model, a complex procedure was developed based on topic coherence calculating and involving some extra criteria. It demonstrated better effectiveness in soft clustering of a text collection, than evaluating a topic model by the metrics of its average coherence.
This article considers the problem of finding text documents similar in meaning in the corpus. We investigate a problem arising when developing applied intelligent information systems that is non-detection of a part of solutions by the TF-IDF algorithm: one can lose some document pairs that are similar according to human assessment, but receive a low similarity assessment from the program. A modification of the algorithm, with the replacement of the complete vocabulary with a vocabulary of specific terms is proposed. The addition of thesauri when building a corpus vector model based on a ranking function has not been previously investigated; the use of thesauri has so far been studied only to improve topic models. The purpose of this work is to improve the quality of the solution by minimizing the loss of its significant part and not adding “false similar” pairs of documents. The improvement is provided by the use of a vocabulary of specific terms extracted from the text of the analyzed documents when calculating the TF-IDF values for corpus vector representation. The experiment was carried out on two corpora of structured normative and technical documents united by a subject: state standards related to information technology and to the field of railways. The glossary of specific terms was compiled by automatic analysis of the text of the documents under consideration, and rule-based NER methods were used. It was demonstrated that the calculation of TF-IDF based on the terminology vocabulary gives more relevant results for the problem under study, which confirmed the hypothesis put forward. The proposed method is less dependent on the shortcomings of the text layer (such as recognition errors) than the calculation of the documents’ proximity using the complete vocabulary of the corpus. We determined the factors that can affect the quality of the decision: the way of compiling a terminology vocabulary, the choice of the range of n-grams for the vocabulary, the correctness of the wording of specific terms and the validity of their inclusion in the glossary of the document. The findings can be used to solve applied problems related to the search for documents that are close in meaning, such as semantic search, taking into account the subject area, corporate search in multi-user mode, detection of hidden plagiarism, identification of contradictions in a collection of documents, determination of novelty in documents when building a knowledge base.
The problem of detecting anomalous documents in text collections is considered. The existing methods for detecting anomalies are not universal and do not show a stable result on different data sets. The accuracy of the results depends on the choice of parameters at each step of the problem solving algorithm process, and for different collections different sets of parameters are optimal. Not all of the existing algorithms for detecting anomalies work effectively with text data, which vector representation is characterized by high dimensionality with strong sparsity.The problem of finding anomalies is considered in the following statement: it is necessary to checking a new document uploaded to an applied intelligent information system for congruence with a homogeneous collection of documents stored in it. In such systems that process legal documents the following limitations are imposed on the anomaly detection methods: high accuracy, computational efficiency, reproducibility of results and explicability of the solution. Methods satisfying these conditions are investigated.The paper examines the possibility of evaluating text documents on the scale of anomaly by deliberately introducing a foreign document into the collection. A strategy for detecting novelty of the document in relation to the collection is proposed, which assumes a reasonable selection of methods and parameters. It is shown how the accuracy of the solution is affected by the choice of vectorization options, tokenization principles, dimensionality reduction methods and parameters of novelty detection algorithms.The experiment was conducted on two homogeneous collections of documents containing technical norms: standards in the field of information technology and railways. The following approaches were used: calculation of the anomaly index as the Hellinger distance between the distributions of the remoteness of documents to the center of the collection and to the foreign document; optimization of the novelty detection algorithms depending on the methods of vectorization and dimensionality reduction. The vector space was constructed using the TF-IDF transformation and ARTM topic modeling. The following algorithms have been tested: Isolation Forest, Local Outlier Factor and One-Class SVM (based on Support Vector Machine).The experiment confirmed the effectiveness of the proposed optimization strategy for determining the appropriate method for detecting anomalies for a given text collection. When searching for an anomaly in the context of topic clustering of legal documents, the Isolating Forest method is proved to be effective. When vectorizing documents using TF-IDF, it is advisable to choose the optimal dictionary parameters and use the One-Class SVM method with the corresponding feature space transformation function.
Дополнительный материал к научной статье на тему оценки прикладного качества тематических моделей для задач кластеризации
This article reveals the possibilities of applying the results of spectral inversion on the example of a field in Eastern Siberia and shows the advantages of this approach over standard methods.
С развитием все более сложных методов автоматического анализа текста повышается важность задачи объяснения пользователю, почему прикладная интеллектуальная информационная система выделяет некоторые тексты как схожие по смыслу. В работе рассмотрены ограничения, которые такая постановка накладывает на используемые интеллектуальные алгоритмы. Проведенный авторами эксперимент показал, что абсолютное значение схожести документов не универсально по отношению к интеллектуальному алгоритму, поэтому оптимальную пороговую величину схожести необходимо устанавливать отдельно для каждой решаемой задачи. Полученные результаты могут быть использованы при оценке применимости различных методов установления смысловой схожести между документами в прикладных информационных системах, а также при выборе оптимальных параметров модели с учетом требований объяснимости решения. The problem of providing a comprehensive explanation to any user why the applied intelligent information system suggests meaning similarity in certain texts imposes significant requirements on the intelligent algorithms. The article covers the entire set of technologies involved in the solution of the text clustering problem and several conclusions are stated thereof. Matrix decomposition aimed at reducing the dimension of the vector representation of a corpus does not provide clear explanatiom of the algorithmic principles to a user. Ranking using the TF-IDF function and its modifications finds a few documents that are similar in meaning, however, this method is the easiest for users to comprehend, since algorithms of this type detect specific matching words in the compared texts. Topic modeling methods (LSI, LDA, ARTM) assign large similarity values to texts despite a few matching words, while a person can easily tell that the general subject of the texts is the same. Yet the explanation of how topic modeling works requires additional effort for interpretation of the detected ones. This interpretation gets easier as the model quality grows, while the quality can be optimized by its average coherence. The experiment demonstrated that the absolute value of documents similarity is not invariant for different intelligent algorithms, so the optimal threshold value of similarity must be set separately for each problem to be solved. The results of the work can be further used to assess which of the various methods developed to detect meaning similarity in texts can be effectively implemented in applied information systems and to determine the optimal model parameters based on the solution explicability requirements.
Один и тот же автор можем быть первым в списке авторов одной статьи и последним в другой. Порядковый номер автора изменяется со временем и отражает вклад автора в исследование. В данном исследовании изучен феномен изменения порядкового номера авторов от статьи к статье. Для этого предложена новая методика на основе анализа и сравнения временных рядов. Основным результатом исследования является методика выявления кластеров авторов характеризующих их научную карьеру и издательскую политику журнала.
Авторами предложена новая методика для парного сравнения коллекций научных статей с помощью тематической модели. Разработанная методика получила название Сравнительного Тематического Анализа (СТА). СТА позволяет получить не только количественную оценку похожести коллекций, но и структурные различия сравниваемых коллекций, как в количественном виде, так и с помощью средств визуализации, разработанных авторами. В данном исследовании проведено сравнение существующих подходов к тематическому моделирования применительно к рассматриваемой задаче сравнения коллекций научных статей. Рас- смотрены вероятностные и генеративные тематические модели. Проведен анализ требований к текстовым коллекциям для корректного применения СТА. Методика СТА показала высокую эффективность на выделении структурных различий близких по тематике коллекций. Автора- ми разработана интегральная метрика «Коэффициент контентной аутентичности», позволяющая сравнивать коллекций между собой. В результате цифрового эксперимента, наиболее информативной показала себя тематическая модель с аддитивной регуляризацией (АRТМ).
The authors developed an approach to comparative analysis of scientific journals collections based on the analysis of co-authors graph and the text model. The use of time series of co-authorship graphs metrics allowed the authors to analyze trends in the development of journal authors. The text model was built using machine learning techniques. The journals content was classified to determine the authenticity degree of various journals and different issues of a single journal via a text model. The authors developed a metric of Content Authenticity Ratio, which allows quantifying the authenticity of journal collections in comparison. Comparative thematic analysis of journals collections was carried out using the thematic model with additive regularization. Based on the created thematic model, the authors constructed thematic profiles of the journals archives in a single thematic basis. The approach developed by the authors was applied to archives of two journals on the Rheumatology for the period 2000–2018. As a benchmark for comparing the co-author’s metrics, public data sets from the SNAP research laboratory at Stanford University were used. As a result, the authors adapted the existing examples of the effective functioning of the authors collaborations in order to improve the work of journals editorial staff. Quantitative comparison of large volumes of texts and metadata of scientific articles was carried out. As a result of the experiment conducted using the developed methods, it was shown that the content authenticity of the selected journals is 89%, co-authorships in one of the journals have a pronounced centrality, which is a distinctive feature of the policy editor. The clarity and consistency of the results confirm the effectiveness of the approach proposed by the authors. The code developed in the course of the experiment in the Python programming language can be used for comparative analysis of other collections of journals in the Russian language.
Structural differences between scientific articles that arise from their translation from Russian into English are studied using the modal topic modeling technique. Each collected document is represented by two modes, that is, English and Russian. As a result of the topic modeling, the Φ and Θ bimodal matrices are obtained. Analysis of the Φ matrix showed that the topics were divided according to the degree of conformity between Russian and English terms when the words are considered in descending order of probability. For 90% of the topics, the English words fully match the Russian ones. Analysis of the Θ matrix showed that for 99% of the documents there is a subject with a value greater than 0.95. Thus, most of the documents are monotopical, which does not depend on the document language.
This paper presents the results of studies aimed at analyzing the effectiveness of a research center. The study focuses on the process of self-organization of project teams (groups of co-authors) for project implementation (writing a scientific article). The initiative to create a team comes from one of its members. The paper describes a formal model, based on a competence approach, which considers the types of tasks to be solved and the necessary skills of the staff. The paper also presents the results of simulation in the AnyLogic environment and problems for further research. The competency profile of each employee is a vector where each coordinate describes the level of mastery of the corresponding skill. The competency profile of the team is a vector obtained as the result of simple addition of the competency profiles of the participants. The proposed model assumes that each task requires a certain set of competencies and that the list of competencies and the level of experience are the criteria for deciding whether to join the team. The logic of decision making at various stages of team creation is modelled by functions. At each step of the modelling, the next employee is chosen randomly. To calibrate the team member's competency profile, internal data on employee qualifications of the Gazpromneft Research Center was used. The constructed model is the basis for further studies of the process by which project teams are created and function in a scientific environment and for developing a methodology to assess the effectiveness of the work of research teams. It helps to predict the need for personnel with different competencies, plan activities to improve the skills of employees and strengthen communication in the team.