
The article introduces an experiment conducted during classes on text tonality as part of a Master’s Degree course of Computational Linguistics. These classes develop competencies in vocabulary methods, ready-made software libraries, and neural network language models. The experiment involved 16 first-year Digital Linguistics students majoring in Smart Systems in the Humanities at St. Petersburg Polytechnic University in 2024–2025. The didactic design of assignments made it possible to gradually increase the complexity of tasks and combine the acquisition of theoretical knowledge with practical skill development. Module 1 focused on the use of English and Russian tonality dictionaries. The students learned to identify the fundamental limitations of the lexicographic approach, e.g., incomplete dictionaries, as well as difficulties associated with contextual meaning, negations, sarcasm, and cultural specificity. Module 2 involved training a DistilBERT model and working with various datasets. The students reflected on the role of data in prediction quality, as well as on the limitations of modern neural architectures in analyzing complex semantic phenomena. Module 3 involved a group project. The students compared lexicon-based, neural, and library-based (VADER, TextBlob, Flair) approaches for advantages, disadvantages, and applicability. The experiment fostered critical thinking and the ability to make informed decisions when electing tools for specific tasks and available resources. A comprehensive synergy of methods is efficient as it allows students to test both the strengths and the weaknesses of each approach. Integrating sentiment analysis into digital linguistics curricula is a key challenge for modern higher education.
Metadata extraction and processing are crucial for social media analysis as they help to systematize information about publications, authors, and audience engagement. The article introduces the web application OmniTrack, designed for automated collection and parameterization of metadata from Russian-language and English-language travel blogs. The application integrates web-scraping methods with a multi-layer architecture and a modular approach, which provides scalability, reproducibility, and extensibility even with changing external platform interfaces. The backend is implemented in Python (Flask); the frontend utilizes HTML, CSS, and JavaScript for an interactive user experience. Data-extraction algorithms are independent modules: undetected_chromedriver for TikTok’s dynamic rendering via browser emulation, yt-dlp for direct JSON-formatted metadata retrieval from YouTube, and Instaloader for high-level access to Instagram’s object model. Collected metadata are normalized to a unified schema in Excel format using the Openpyxl library, which facilitates subsequent statistical analysis. The application underwent usability testing: 42 participants processed 400 posts, evaluating installation simplicity, processing speed, and interface intuitiveness. The mean ease-of-use score was as high as 4.9 out of 5. Some critical issues were identified and resolved, including incompatibility of the pywebview backend module and incorrect handling of shortened TikTok links. The OmniTrack web application provides a robust framework for constructing a representative metadata corpus, supporting further linguistic research into the discursive, genre, and communicative features of Russian-language and English-language travel blogs.
As the digital environment transforms the Tajik language, it turns into a new hybrid functional style that diverges from literary standards. The article describes the graphical evolution of digital Tajic, as well as the Russian-Tajik code switching in online comments in social networks. Based on digital ethnography and discourse analysis, the author compares the actual online usage with the norms codified in the Tajik National Corpus. The Tajik net-speak is examined from the perspective of the law of least effort (Zipf’s law). It demonstrates an obvious typological parallel with the digital erratives of Russian online communication known as "Olbanian language" or "Scumbags’ slang". The graphical deviations from standard Tajik include vowel reduction, with readability maintained by the consonantal frame. These violations are systemic, driven primarily by technological factors. Russian borrowings are integrated into the agglutinative matrix of the Tajik language via the mechanisms of morphological adaptation and code-switching strategies. Virtual communication legitimizes dialectal forms, transforming them into a sociolect, which reflects the high vitality of the language in the digital age.
The media environment evolves from a purely broadcasting medium into an active factor that structures cognitive experience. For instance, the English-language media concept resilience adapts to the Russian and the Uzbek digital environments. In this respect, it can act as a cognitive tool for shaping and transforming interpretation models of virtual communication within a trilingual virtual continuum, which is a framework for analyzing semantic transformation. The interpretation model is a stable media-cognitive scheme for conceptualizing experience, actualized through a media concept. In the trilingual virtual space, these models gravitate toward national axiological dominants. The parametric model for grouping media concepts systematizes cognitive characteristics of the focal concept resilience and reveals its adaptation logic across different linguistic cultures. A cognitive-discourse analysis traces the mechanisms of media concept transformation within the gradual cognitive-discourse continuum of the trilingual virtual space. The Russian-language discourse reframes the concept from "strength of spirit" to an "expert resource". In the Uzbek segment, resilience corresponds to the sabr (patience) category, which shifts the focus from personal resilience to a collective ethical norm.
The article presents the results of a study on the relationship between communication in native and learned languages in offline and online communication systems. This study tests the hypothesis that the time of entry into and intensity of online communication influence instructed bilinguals’ self-assessments of their proficiency in basic forms of communication. These forms include passive skills (listening comprehension, reading) and active skills (speaking, writing). The study was conducted on the material of bilingualism variants involving the Russian language, implemented in variable language situations in which Russian is native for bilinguals and the majority (Russian- English bilingualism), the second, studied, minority (Uzbek-Russian and Tajik-Russian). The relevance of the study is determined by the need to study two global trends in the development of modern society – the increasing role of virtual communication and bilingual practices in the world in general, and in the Russian Federation, in particular. The main research methods are a questionnaire to collect subjective assessments by respondents of the target characteristics of social and linguistic experience, statistical analysis in processing the obtained materials, sociolinguistic analytics. As a result of the analysis, zones of use of the second language in virtual speech practices of the active and passive type were identified in different language situations in three regions. It was found that, overall, the correlation between levels of the second language proficiency and the time to engage in online communication in the second language ranges from average to zero, varying significantly across the three bilingual groups studied, operating in variable language situations. The data obtained generally agree with the results of related studies conducted on other language pairs of bilinguals and in other regions, while also revealing unique features due to the unique language situations within which the studied types of bilingualism are formed. The findings presented are limited to a sample of respondents, primarily students and recent graduates of humanities faculties with a language specialization. As a research prospect, we consider expanding the sample by involving respondents with other aspects of social and linguistic experience.