Transparent orthographies, such as Bulgarian and Italian, feature highly consistent grapheme-phoneme correspondences, enabling rapid acquisition of decoding skills. Despite belonging to different language families and using distinct scripts (i.e., Cyrillic vs. Latin), these languages provide an ideal framework to investigate whether orthographic transparency can outweigh script differences in shaping reading development. We conducted a cross-sectional study with primary school children from Grades 2 to 5 in Bulgaria and Italy. Reading performance was recorded using a novel finger-tracking technique, which allows the capture of temporal dynamics of reading in a portable, low-cost, and classroom-friendly format. Measures of reading time and text comprehension accuracy were compared across grades and languages. Developmental trajectories for both speed and comprehension accuracy showed remarkable similarity across Bulgarian and Italian, with both languages exhibiting steady improvement from grade 2 to grade 5. Our cross-linguistic results showed that reading development in primary school children follows both universal and language-specific trajectories. While broad developmental trajectories were similar, cross-linguistic differences emerged in the impact of morphological complexity, pointing to both universal and language-specific mechanisms. Our findings indicate that orthographic transparency may exert a stronger influence on early reading development than script type, even across languages from different families. The study also highlights the potential of finger-tracking for large-scale literacy research. Establishing comparable developmental benchmarks in transparent orthographies may inform cross-linguistic screening tools and early interventions.
The paper presents a model for the design and management of metadata that enables the efficient compilation of specialised datasets from large, heterogeneous data collections. The metadata are represented as a typed property graph that facilitates the FAIR principles in data compilation: Findable, Accessible, Interoperable, Reusable. The representation is general and independent of the modality and format of the data. The graph-based design of the metadata supports the incremental extension of categories and relationships without requiring the migration of existing data. Its feasibility is demonstrated through the use of a graph database, in which the metadata for 689,645 Bulgarian textual data units are combined with a web-based filtering interface. Metadata retrieval is implemented through Cypher queries executed as graph traversals, enabling the extraction of thematic and application-oriented data subsets based on combinations of selection criteria. The application validates the suitability of the graph-based metadata design for compiling specialised datasets for training and fine-tuning large language models and other NLP applications.
The article presents an infrastructure that integrates freely available tools and large language models into a personalized chatbot. The aim is to show how large language models can be used with affordable graphics processors and what possibilities exist to enrich user queries with additional context depending on the objective and area of application. Working with pre-trained large language models will be demonstrated using RAG (Retrieval-Augmented Generation) technology for Bulgarian to solve a specific task in a particular thematic domain. The chosen task is document summarization, and the domain is science and education. The presented solution can also be used to create other applications for the Bulgarian language.
This article presents the general principles for the inflectional description of verbal multiword expressions in Bulgarian. Two main criteria are applied: the maintenance of the relation to the inflectional description of simple words in Bulgarian and the presentation of the minimal grammatical information necessary for a standardised description of the forms of verbal multiword expressions. The paper also provides a brief outline of the formalism used to define local grammars, which are compiled into finite-state transducers that recognise and generate the forms of verbal multiword expressions, along with the associated grammatical information.
The article introduces the journal Computational Linguistics in Bulgaria, an annual open access peer-reviewed journal published by the Department of Computational Linguistics at the Institute for Bulgarian Language of the Bulgarian Academy of Sciences. The relationship between the terms computational linguistics, natural language processing and artificial intelligence is briefly commented on in order to clarify the concept behind the journal’s name. The focus is then placed on the Bulgarian language and the Bulgarian research community, emphasising the importance of international contributions for the development of scientific cooperation and progress. The scope of the journal Computational Linguistics in Bulgaria is presented: It publishes articles on all areas of theoretical computational linguistics as well as on existing language resources, datasets and technologies for natural language processing and artificial intelligence. The journal promotes new approaches and methods, especially those aimed at applying language technologies to small and still resource-poor languages such as Bulgarian.
Despite the remarkable progress made in the field of Machine Translation (MT), current systems still struggle when translating ambiguous words, especially when these express infrequent meanings. In order to investigate and analyze the impact of lexical ambiguity on automatic translations, several tasks and evaluation benchmarks have been proposed over the course of the last few years. However, work in this research direction suffers from critical shortcomings. Indeed, existing evaluation datasets are not entirely manually curated, which significantly compromises their reliability. Furthermore, current literature fails to provide detailed insights into the nature of the errors produced by models translating ambiguous words, lacking a thorough manual analysis across languages. With a view to overcoming these limitations, we propose Disambiguation Biases in MT (DiBiMT), an entirely manually curated evaluation benchmark for investigating disambiguation biases in eight language combinations and assessing the ability of both commercial and noncommercial systems to handle ambiguous words. We also examine and detail the errors produced by models in this scenario by carrying out a manual error analysis in all language pairs. Additionally, we perform an extensive array of experiments aimed at studying the behavior of models when dealing with ambiguous words. Finally, we show the ineffectiveness of standard MT evaluation settings for assessing the disambiguation capabilities of systems and highlight the need for additional efforts in this research direction and ad-hoc testbeds such as DiBiMT. Our benchmark is available at: https://nlp.uniroma1.it/dibimt/.
This article presents the combined results of three international surveys that were carried out in the context of the Horizon 2020 European Lexicographic Infrastructure project (ELEXIS). The aim of these surveys was to gain more insight into lexicographic practices and the needs of lexicographers in Europe. The surveys were delivered via online platforms. Based on the combined results, we sketch a map of lexicographic practices in Europe, both for born-digital and retrodigitized resources, analyze current needs in terms of tools, functionalities and training, and identify emerging trends that will affect lexicography in the short and long term.
In this paper, we present two experiments focussing on linguistic classification and annotation of examples, using zero-shot prompting. The aim is to show how large language models can confirm or reject the linguistic judgements of experts in order to increase the productivity of their work. In the first experiment, new lexical units evoking a particular FrameNet semantic frame are selected simultaneously with the annotation of examples with the core frame elements. The second experiment attempts to categorise verbs into the aspectual classes, assuming that only certain combinations of verbs belonging to different aspectual classes evoke a semantic frame. The linguistic theories underlying the two experiments, the development of the prompts and the results of the experiments are presented.
The article presents the challenges of implementing a System for data retrieval and visualisation from the Internet by crawling language resources from the Hugging Face repository and extracting the associated data. The data in the system is updated at regular intervals to track the dynamics of language resource creation for different time periods. The article presents: a) the analysis of the available data and its structure; b) the chosen method for crawling the pages and extracting the data. The shared experience of overcoming the specific challenges can serve to solve similar problems related to the extraction of data from the Internet, a task that often has to be solved in various projects (including school projects).
The 10th edition of the Forum on Bulgarian Grammar, titled The Spoken Language Phenomenon, was held in Sofia on 20 and 21 October 2023. The Forum was organised by the Institute for Bulgarian Language Prof. Lyubomir Andreychin at the Bulgarian Academy of Sciences and the Department of Bulgarian Language at the Konstantin Preslavsky University of Shumen. It was dedicated to one of the most distinguished Bulgarian linguists of our time – Corresponding Member Prof. Dr. Todor Boyadzhiev.
The paper reports on the first steps in developing a time-stamped multimodal dataset of reading data by Bulgarian children. Data are being collected, structured and analysed by means of ReadLet, an innovative infrastructure for multimodal language data collection that uses a tablet as a reader’s front-end. The overall goal of the project is to quantitatively analyse the reading skills of a sample of early Bulgarian readers collected over a two-year period, and compare them with the reading data of early readers of Italian, collected using the same protocol. We illustrate design issues of the experimental protocol, as well as the data acquisition process and the post-processing phase of data annotation/augmentation. To evaluate the potential and usefulness of the Bulgarian dataset for reading research, we present some preliminary statistical analyses of our recently collected data. They show robust convergence trends between Bulgarian and Italian early reading development stages.
The paper presents the bilateral project Assessing reading literacy and comprehension of early graders in Bulgaria and Italy, its aims and expected results, as well as the principles underlying the creation of texts assessing reading literacy in Italian and Bulgarian. The ultimate goal of the project is to improve the literacy skills of primary school children through education. In order to contribute to the achievement of this goal, a thorough investigation focused on the assessment of reading literacy and comprehension among early school children in Bulgaria and Italy. The article focuses on the principles for preparing materials in Italian and Bulgarian for assessing reading literacy. The linguistic features used to predict reading ability are divided into four main groups: raw text, lexical, morpho-syntactic and syntactic features.
WordNet contains a fair number of synsets with multiple hyperonyms. In parent–child relations, a child can have only one parent (ancestor). Consequently, multiple hyperonymy represents distinct semantic relations. In order to reclassify the multiple hyperonyms, we define a small set of new semantic relations (such as function, origin and form) that cover the various instances of multiple hyperonyms. The synsets with multiple hyperonyms that lead to the same root and belong to the same semantic class were grouped automatically, resulting in semantic patterns that serve as a point of departure for the classification. The proposed changes are based on semantic analysis and may involve the redefinition of one or several multiple hyperonymy relations to new ones, the removal of one or several multiple hyperonymy relations, and rarely the addition of a new hyperonymy relation. As a result, we incorporate the newly defined semantic relations that resolve the former multiple hyperonymy relations and propose an updated WordNet structure without multiple hyperonyms. The resulting WordNet structure without multiple hyperonyms may be used for a variety of purposes that require proper inheritance.
The article presents a system that dynamically displays the availability of language datasets and language models found in large repositories such as Hugging Face. The goal of developing such a system is to demonstrate that, outside of English, the datasets and language models required for advancements based on or utilizing language technologies and artificial intelligence have either moderate or fragmented support. At the same time, the description of the system architecture introduces readers to easy-to-use instruments such as Node-RED, MariaDB, and Grafana, which offer a wide range of application opportunities in solving various tasks like crawling and collecting data from the Internet, storing information in a database, and data visualization in a clear and functional manner. Each of these tasks, as well as their combinations, can be used to carry out student projects at the senior high school level of education.
The paper presents some general facts about Bulgarian, which is spoken by over 8 million people all over the world and is the official language of the Republic of Bulgaria. It is shown that language diversity within the country has been relatively constant and modest over the last 90 years (with a clear dominance of Bulgarian). The focus of the paper is the presentation of the legislation in Bulgaria that regulates language teaching and use in different spheres of life, the main activities of the national institute of language, and the role of language education. The main findings are as follows: There is no dedicated Bulgarian Language Act. Instead, over 100 legislative acts govern issues concerning the usage and study of the Bulgarian language. The Institute for Bulgarian Language at the Bulgarian Academy of Sciences supports the Bulgarian state's language policy through: researching the present state, history, and dialect variety of the Bulgarian language; monitoring changes in written and spoken language; publishing grammars and dictionaries; and developing language resources and technologies. The key factors that influence language education in Bulgaria include early childhood education and care, equity in education and educational outcomes, and the rise of new learning styles.
AbstractThis chapter reports on the current status of technology support for Bulgarian and highlights certain gaps. The analysis is based on the services and resources available in the European Language Grid in early 2022. While the LT field as a whole has significantly progressed in the last ten years, we conclude that there is still a yawning technological gap between English and Bulgarian, and even between German, French, Italian, Spanish and Bulgarian. It is exactly this distance that needs to be ideally eliminated, if not at least reduced, in order to move towards Digital Language Equality for Bulgarian.
This article presents the project Assessing the Reading Literacy and Comprehen-sion of Early Graders in Bulgaria and Italy, which is carried out as part of an international collaboration between two partner organisations – the Institute for Bulgarian Language Prof. Lyubomir Andreychin (BAS) with participants from the Department of Computational Linguistics and the Institute for Computational Linguistics A. Zampolli in Pisa, Italy. The main goal of the project is to research and assess the reading skills of primary school students using modern language technologies.