
This study aims to improve the detection of small-scale microorganisms in complex microscopic environments using advanced deep-learning techniques. Specifically, we target the detection of diplococci in images captured through microscopy of live (dynamic) samples. We propose a method featuring a Localized and Global Context Attention mechanism within a multi-head framework and an innovative Feature Emphasis Adjustment to enhance detection accuracy. Our approach addresses the challenges posed by small microorganisms, which are often difficult to detect due to their small size, intricate backgrounds, and complex structures. By dynamically adjusting the attention scale across different heads, our method captures detailed features at multiple levels, significantly improving the detection and description of small microorganisms. Additionally, we introduce a softmax-based reweighting function that selectively emphasizes essential features for object recognition, reducing noise and irrelevant information. Our model surpasses the accuracy of existing state-of-the-art solutions. Furthermore, we provide a newly curated dataset specifically designed for microorganism detection, featuring a variety of annotated microscopic images for training and evaluation. These contributions advance the understanding of attention mechanisms in deep neural networks and offer practical improvements for applications requiring precise microorganism detection.
Speaker verification (SV) is the process of verifying whether speech from two audio signals originate from the same speaker or different speakers. Current state-of-the-art SV systems are based on deep neural networds, predominantly trained using the VoxCeleb dataset. This may lead to varying SV performance when using the models for inference on real-world data. To research these possible variations in performance, three establised SV models, namely the ECAPA-TDNN, ResNet and WavLM, are evaluated on the UCLA variability, CommonVoice, FRIDA and Wyred datasets. The ECAPA-TDNN and ResNet models are found to perform slightly worse when compared with the VoxCeleb evaluation results while the WavLM model performs significantly worse. The ResNet model shows the best performance on all four datasets. After evaluation, the ResNet model is improved by fine-tuning the model on the UCLA dataset and, further by creating a Deep Weight Space Ensemble (WSE) model between the pre-trained and fine-tuned models. Between the pre-trained, fine-tuned and WSE models, the WSE model has the best overall performance, attaining the best scores on the UCLA test set. Scores for the other three datasets show a lower decrease than the fine-tuned model. This indicates that fine-tuning with WSE can alleviate the loss in model performance on real-world data.
Studying the structure of social aspects of interactions in spoken speech is important for creating artificial intelligence systems capable of understanding everyday human talk and contributing to it in a natural manner. Much of the recent research in computational pragmatics has focused on modeling and testing specific linguistic phenomena like metaphor, irony, conversational maxims, etc., while studies in dialogue pragmatics tend to pay more attention to task-oriented conversations with the research of social aspects mostly left to the functioning of formulaic expressions. This paper presents the ongoing research on the annotation of social talk in Russian using the taxonomy from ISO standard 24617-2:2020 “Semantic annotation framework, Part 2: Dialogue acts”. Specifically, we report on the analysis of dialogue acts used in establishing social contact (namely, greetings and introductions) in dialogues from the Russian Multimedia Politeness Corpus, including multi-party interactions, and propose new communicative functions to cover widespread conversational intentions. Additionally, we provide preliminary conclusions on characteristic features of oral Russian communication within these forms of interactional exchange.
Nowadays, the problem of natural language understanding models evaluation is an essential part of theoretical and practical concern. In this paper, we build ( https://github.com/yaroslav-i-am/paramsum ) the annotated dataset for quality of in-domain aspect extraction assessment. We propose to design this task in domain of movie reviews. We used a crowdsourcing method to build a marked-up dataset. We use fine-tuned large language models to build the ultimate set of annotated reviews. We ended with MoRAE dataset of 986 reviews with 8413 extracted aspects.
Speech plays a crucial role in effective communication for teachers. Therefore, it is essential to choose a communication strategy that can help achieve goals quickly. In order to captivate students’ attention and implement a successful communication strategy, teachers often use certain multiword expressions. These expressions are typically manually analyzed in contemporary research on pedagogical discourse. The importance of this research lies in the growing need for a blend of linguistic research methods and artificial intelligence techniques in the domain of pedagogical discourse. Analyzing multiword expressions can help identify language and communication elements that have a significant informational impact, as well as how to properly utilize them. This study aims to identify multiword expressions in a corpus of teachers’ speech, specifically transcripts of recorded lessons from secondary school teachers in the Russian Federation. The corpus consists of lessons from both more and less effective teachers, with more effective teachers meeting specific criteria such as working in non-selective schools with diverse student populations and achieving above-average results in State Final Certification (the 9th Grade). By employing statistical metrics, contextualized vector models, and clustering algorithms, we are able to detect and describe the unique vocabulary used by teachers. The findings reveal that the speech of more effective teachers is distinguished by specific lexical markers related to interaction with students and structuring lessons. These results could be beneficial for speech technology specialists developing voice assistants for teachers, as well as linguists creating speech corpora in the Russian language.
Devanagari script, used in more than 120 South Asian languages including Hindi, Nepali, and Marathi, poses unique challenges due to complex character structures and ligatures. This work focuses on Offline Nepali Handwritten Word Recognition, involving the creation of a novel handwritten dataset of Nepali words and the implementation of deep learning networks, based on Convolutional Recurrent Neural Networks (CRNN) and attention-based encoder-decoder architectures. This work curates a diverse dataset, addressing limitations of existing Hindi words datasets, such as the repetition of words. The resulting dataset consists of 8558 samples taken from 28 individuals. The Deep Learning models were pre-trained on a generated Synthetic Nepali Fonts Dataset and then fine-tuned on our Nepali handwritten dataset and evaluated with accuracy and Character Error Rate (CER), achieving 81
The development of machine-readable lexical resources for low-resource languages, such as Kyrgyz, faces significant challenges due to limited NLP tools and poorly structured linguistic data. In this paper, we introduce an innovative method for extracting structured lexical information from Yudakhin’s Russian-Kyrgyz dictionary, a bilingual resource with inconsistent entry formatting. Our approach utilizes GPT-4o to bootstrap a dataset and explores both few-shot learning and fine-tuning techniques to convert dictionary entries into a structured JSON schema. We assess the impact of varying few-shot example sizes on model performance and compare the effectiveness of few-shot learning against fine-tuning across several models, including an open-source option. Our results demonstrate notable success, with the highest-performing model achieving 92.70
The paper is devoted to the investigation into the possibility of content dissemination in social networks, considering users’ influence potential. The process of information dissemination is studied using simulation methods, virtual social networks and real social networks (real data are extracted from the VK social network). During the study of virtual social networks SIR and SEIR diffusion models are used. However, to simulate the real social network only SEIR model is considered, because this model reflects better the actual behavior of social network users. The diffusion models in both cases are modified by altering the set of states. Within the investigation of VKontakte social network, the “Infected” state is divided into two: the first represents users who have liked a post; the second represents users who have shared the post on their social network page, specifically through reposting. Additionally, the developed simulation model allows to vary individual parameters for each agent: the level of influence potential and the probability that the agent will see the post with the disseminated content. An ontological approach is used to store data about real network. AnyLogic is chosen as the tool for conducting the simulation experiments.
Visual Question Answering is one of the essential parts of machine reasoning. Datasets are created to train a model to perform this task. However, there are only a few datasets for the Russian language. Moreover, existing sets may have strong biases, allowing models to score high without reasoning. In this paper, we adapt the idea of the English diagnostic dataset for compositional language and elementary visual reasoning, CLEVR, to Russian. We also evaluate multiple baselines and models on this dataset to see how well they perform. The results may be used to improve the performance of Russian multimodal LLMs.
Despite the promising results of large language models (LLMs) in labeling tasks, further exploration is needed to leverage them effectively for linguistic data annotation. One of the most challenging tasks in this regard is labeling discourse structures, which is highly subjective and often involves ambiguity in class description. In this paper, we address the challenge of using LLMs for hybrid annotation of the discourse structure in open-domain dialogues, relying on Eggins and Slade’s speech function theory. We conduct a comparative analysis between model-generated annotations and human annotations, exploring the potential of LLM-assisted annotation as a viable alternative to crowdsourcing.
In triadic data setting, implications can be extracted in different forms: those introduced by Biedermann (conditional attribute and attributional condition implications) and those introduced by Ganter and Obiedkov (attribute × condition, conditional attribute and attributional condition implications). We provide in this paper an optimal set of implications for triadic data, based on pseudo-features, a notion similar to pseudo-intent for dyadic data.
This research focuses on the cross-lingual text summarization between Russian and Chinese languages, addressing the growing need for effective information exchange amidst globalized relations between China and Russia. It highlights the linguistic challenges posed by these languages’ complex grammatical structures and unique alphabets. The study investigates existing methods and datasets, emphasizing the importance of improving cross-lingual summarization technology to overcome language barriers. The key findings demonstrate that Direct Preference Optimization, a standard reinforcement learning algorithm for LLM training, significantly improves summarization quality, particularly in many-to-one training scenarios, when compared to traditional pipeline methods. The study utilizes datasets such as WikiLingua and CrossSum, alongside manually collected data, to ensure comprehensive evaluation. The best results for the Russian-to-Chinese model showed a ROUGE-2 score of 11.71 and a LaSE (Language-agnostic Summary Evaluation) score of 31.18. Additional metrics include Language Confidence and Length Penalty. GPT-4 assessments further confirm the improvements in the generated summaries.
We consider the problem of handwritten text recognition (HTR) for collections of historical archive documents. The known HTR models can be split into two major categories – line-level models and page-level models. Line-level models for their training and inference require costly preprocessing procedures of careful extraction of image fragments corresponding to separate text lines. Page-level models do not require such preprocessing but usually show weaker recognition results. In this paper we propose a new HTR model YOLO-HTR that is aimed at closing a performance gap between line-level and page-level models. The new model simultaneously solves the tasks of detecting text lines and recognizing them, and thus does not require text line extraction for inference stage. The model architecture combines ideas from object detection YOLO model and HTR model Vertical Attention Network. In the paper we also propose a modification of CTC-loss that allows using text lines supervision with partially unknown text symbols – a common feature of expert supervision for challenging historical documents. Experiments were conducted on two collections of handwritten texts: the archive of diaries of the Russian navigator Fyodor Petrovich Litke and the archive of letters from prisoners of the Smolensk convict prison. The experiment results show that the proposed approach allows achieving recognition quality comparable to line-level models, with less labor costs.
Graphical abbreviation is a method of shortening words to save time and space on the page. This paper explores the distinction between graphical abbreviations and acronyms and provides a classification system for different types of abbreviations. We focus on the development and evaluation of a novel model designed for automatic abbreviation expansion in Russian—a task complicated by the language’s rich inflectional morphology. Our approach combines dictionary-based methods with a masked language model (BERT) to handle both unambiguous and ambiguous abbreviations while ensuring the correct expansion based on grammatical context. To further enhance model performance, we augmented a custom dataset with examples from news articles, providing diverse contexts for abbreviation use. We evaluate the model’s performance using perplexity, Word Error Rate (WER), and Lemma Error Rate (LER). Our results indicate that the model achieves high accuracy, with a WER of 0.0376 and an LER of 0.0299, making it highly effective for practical applications in natural language processing tasks such as text normalization, automatic speech recognition, and digital assistants. This research highlights the importance of context-aware models in abbreviation expansion, particularly for languages with complex morphological systems like Russian.
We consider the problem of finding minimum-weight spanning tree with a diameter, which is either at most or equal to a given bound d, in a complete edge-weighted undirected graph. We propose a new simple polynomial-time approximation algorithm for this problem and provide probabilistic analysis of the algorithm on random inputs in which the weights of edges are i.i.d. random variables with either uniform continuous distribution on [a_n,b_n] or uniform discrete distribution on segment [a_n,b_n] ∩ℕ , 0
We present a new corpus for the task of grammatical error correction for the Russian language. In contrast to previous works, our data consists of middle school essays written by native speakers. In total, the training data includes more than 4500 sentences and the test partition – more than 1000 sentences. The corpus contains a detailed annotation of grammatical errors in .M2 format, fine-grained error types are also available. The distribution of errors in our data differs from other corpora, containing more punctuation and less wordform errors. We study the performance of several models on our data and find that the finetuned YandexGPT model is the best. It shows F _0.5 -score about 73
Emotion recognition from multimodal sources is essential for advancing human-computer interaction. This paper introduces a comprehensive approach integrating robust methodologies to enhance multimodal emotion recognition. By harnessing both facial video and audio signals through Temporal Convolutional Networks (TCN) and Transformer models, alongside other advanced neural architectures, our ensemble effectively captures the complex dynamics of emotional expressions. Our approach has been rigorously evaluated on the Expression Classification challenge from the recent Affective Behavior Analysis in-the-Wild competition, demonstrating significant improvements compared to state-of-the-art methods. Specifically, our model achieved a 12
This paper proposes a new method for conformal inference in time series forecasting that works under any changes in the data generation process. It can be used with any black box forecasting models that predict the future values of a time series. Unlike existing methods, the proposed method adapts to changes in the distribution faster and in a more predictable way. We achieve this by modifying the adaptive conformal inference (ACI) algorithm of Gibbs and Candès (2021) by replacing the constant learning rate parameter with the one that adapts to changes in the data generation process. We tested our method on real datasets and showed the absence of sharp explosions in the width of intervals common for other adaptive conformal prediction approaches.
This paper presents a comprehensive analytical review of contemporary mathematical models of information influence and control in social networks, emphasizing the integration of agent-level factors such as trust, reputation, decision-making, and action execution. We introduce extensions to classical models, including the DeGroot model, by incorporating these critical components to more accurately reflect social interactions. Special attention is given to control strategies and the application of game-theoretic methods for analyzing interactions among controlling entities. We apply these models and methods to real-world social network data, including analyses of ideological preferences and public opinions on health measures during the COVID-19 pandemic. Our approach provides new perspectives for future research in social process modeling and offers insights for designing effective strategies for information influence and control.