
The Oceanic languages of Melanesia are generally small, low-resource languages, of which very little primary data is available. For our study on tense, aspect, and modality (TAM), we have access to richly annotated corpora from seven endangered Oceanic languages. In this paper, we describe some of the methods we used to investigate the category of habitual aspect in these languages. We show that information can be recovered from the English translations and from metadata on genres. For a more in-depth study, we relied on clausebased tags labeling clause type, tense, aspect, mood, and polarity. The process of tagging aspect, in particular, revealed the theoretically and practically important fact that habituality is sometimes a property of larger spans of texts (Carlson and Spejewski, 1997), rather than just a property of clauses, and can combine with more specific clause-level aspect.
We present a procedure for generating a valence resources for Norwegian (Bokmål) from a deep grammar. The corpus is presented in the form of IGT (interlinear glossed text) augmented by valence information. Our deep parser is the HPSG-based computational grammar Norsource (Hellan and Bruland 2015), our online IGT repository is TypeCraft (Beermann andMihaylov 2014), while the sentences of the corpus are taken from the Leipzig Corpus Collection (Goldhahn et al. 2012). We create a common structure for the resources. Our aim is to make the grammatical information encoded in a deep parser more readily accessible for humans and for further processing.
The development of corpora inevitably involves the need for segmentation. For most of the corpora, the first segmentation to operate consist in determining silences vs Inter-Pausal Units - IPUs, i.e. sounding segments. This paper presents the "Search for IPUs" feature included in SPPAS - the automatic annotation and analysis of speech software tool distributed under the terms of public licenses. Particularly, this paper is focusing on its evaluation on Cheese! corpus, a corpus of reading then conversational speech between two participants. The paper reports the number of manual actions which was performed manually by the annotators in order to obtain the expected segmentation: add new IPUs, ignore irrelevant ones, split an IPU, merge two consecutive ones and move boundaries. The evaluation shows that the proposed fully automatic method is relevant.
We focus on a collaboration between community members and visiting linguists in Erakor, Vanuatu, aiming to build the capacity of community-based researchers to undertake and sustain documentation of Nafsan, the local indigenous language. We focus on the technical and procedural skills required to collect, manage, and work with audio and video data, and give an overview of the outcomes of a community-led documentation after initial training. We discuss the benefits and challenges of this type of project from the perspective of the community researchers and the external linguists. We show that community-led documentation such as this project in Erakor, in which data management and archiving are incorporated into the documentation process, has crucial benefits for both the community and the linguists. The two most salient benefits are: a) long-term documentation of linguistic and cultural practices calibrated towards community’s needs, and b) collection of larger quantities of data by community members, and often of better quality and scope than those collected by visiting linguists, which, besides being readily available for research, have a great potential for training and testing emerging language technologies for less-resourced languages, such as Automatic Speech Recognition (ASR).
A suite of related online and offline analysis and visualisation tools for training students of phonetics in the acoustics of prosody is described in detail. Prosody is informally understood as the rhythms and melodies of speech, whether relating to words, sentences, or longer stretches of discourse, including dialogue. The aim is to contribute towards bridging the epistemological gap between phonological analysis, based on the linguist’s intuition together with structural models, on the one hand, and, on the other hand, phonetic analysis based on measurements and physical models of the production, transmission (acoustic) and perception phases of the speech chain. The toolkit described in the present contribution applies to the acoustic domain, with analysis of the low frequency (LF) amplitude modulation (AM) and frequency modulation (FM) of speech, with spectral analyses of the demodulated amplitude and frequency envelopes, in each case as LF spectrum and LF spectrogram. Clustering functions permit comparison of utterances.
Spelling error correction is an important problem in natural language processing, as a prerequisite for good performance in downstream tasks as well as an important feature in user-facing applications. For texts in Polish language, there exist works on specific error correction solutions, often developed for dealing with specialized corpora, but not evaluations of many different approaches on big resources of errors. We begin to address this problem by testing some basic and promising methods on PlEWi, a corpus of annotated spelling extracted from Polish Wikipedia. We focus on isolated correction (without context) of non-word errors (ones producing forms that are out-of-vocabulary). The modules may be further combined with appropriate solutions for error detection and context awareness. Following our results, combining edit distance with cosine distance of semantic vectors may be suggested for interpretable systems, while an LSTM network, particularly enhanced by contextualized character embeddings such as ELMo, seems to offer the best raw performance.
In this article we present extended results obtained on the multidomain dataset of Polish text reviews collected within the Sentimenti project. We present preliminary results of classification models trained and tested on 7,000 texts annotated by over 20,000 individuals using valence, arousal, and eight basic emotions from Plutchik's model. Additionally, we present an extended evaluation using deep neural multilingual models and language-agnostic regressors on the translation of the original collection into 11 languages.
This paper describes an automatic segmentation and transcription module for Polish and its integration with the Annotation Pro software tool. The module is an extended desktop version of the CLARIN-PL online tool and has been named ANNPRO. Thanks to developing the module, it becomes possible to combine the functionality of Annotation Pro desktop program and the web-based automatic aligner. The results can be immediately used as the input for further acoustic-phonetic analyses with Annotation Pro native functions or annotation mining plugins. Annotation Pro enables using any number of external alignment modules, provided that certain basic format requirements are kept. We discuss these requirements and exemplify them with the ANNPRO module functionality and the integration steps. As an illustration, we present a brief report on experiences gained in the process of annotation of a multimodal corpus with the use of ANNPRO. Both Annotation Pro and the ANNPRO module are publicly available for download and can be freely used for research.
The present paper describes a set of tools created and used for fundamental frequency extraction and manipulation, prosody and speech perception analysis and speech synthesis. The tools were implemented as Praat and Python scripts and were created for different purposes in projects over the last few years. The paper presents the functionality and possible usage of the tools in phonetic research.
Nowadays, the Kazakh language belongs to the category of less-resourced languages, as there is a small number of resources developed and accessible to a wide range of users, such as text corpora, electronic dictionaries, morphological analyzers, thesauri, which allow to analyze text documents. The aim of this work is the design and development of pipeline of preprocessing tools for media-corpus of the Kazakh language. Media-corpus is hosted by al-Farabi Kazakh National University and serves linguists as an empirical basis for research in the contemporary written Kazakh language. The development of pipeline of preprocessing tools for media-corpus, the lexical and grammatical features of the Kazakh language were analyzed, on the basis of which the composition of the fundamental rules for changing the words (inflection) of the Kazakh language was determined. In the process of research, the tools for generation and lemmatization of the word forms of the Kazakh language were created. The proposed tools can be applied at the stage of morphological analysis in the systems of automatic analysis of the texts, in the creation of thesauruses and ontologies. For the case of the presence of homonymy, the template method was used, which allow to reduce the level of homonymy.
Corpus is one of the essential parts of language research, especially for the low resource language. To ensure the researching result to be most effective, the corpus that has been used also requires effectiveness and accuracy. The Thai language has some special characteristics that cause difficulty in building the corpus and affect the error of those corpora. Therefore, this paper proposes an effective and efficient approach to clean up the existing Named Entity corpus before using it in any language research. The THAI-NEST corpus is adopted to verify the consistency and integrity of the data and re-design with our proposed model. The revised corpus is verified by the BiLSTM-CNN-CRF model that combined the features among word, POS, and Thai character clusters (TCCs). Experimental results show the effectiveness of the verification, which increased the accuracy by up to 12%, and the model can effectively detect and handle errors of word segmentation and NE tag consistency.
The paper addresses an experiment in detecting metaphorical usage of adjectives and nouns in Polish data. First, we describe the data developed for the experiment. The corpus consists of 1833 excerpts containing adjective-noun phrases which can have both metaphorical and literal senses. Annotators assign literal or metaphorical senses to all adjectives and nouns in the data. Then, we describe two methods for literal/metaphorical sense classification. The first method uses Bi-LSTM neural network architecture and word embeddings of both tokenand character-level. We examine the influence of adversarial training and perform analysis by part-of-speech. The second method uses the BERT token-level classifier. On our relatively small data, the LSTM based approach gives significantly better results and achieves an F1 score equal to 0.81.
Detecting emotions from a text can be challenging, especially if we do not have any annotated corpus. We propose to use book dialogue lines and accompanying phrases to obtain utterances annotated with emotion vectors. We describe two different methods of achieving this goal. Then we use neural networks to train models that assign a vector representing emotions for each utterance. These solutions do not need any corpus of texts annotated explicitly with emotions because information about emotions for training data is extracted from dialogues’ reporting clauses. We compare the performance of both solutions with other emotion detection algorithms.
The research project described in the paper aimed at creating a semantically and grammatically annotated corpus of Polish synesthetic metaphors—Synamet. The texts in the corpus were excerpted from blogs devoted to perfume, wine, beer, cigars, Yerba Mate, tea, or coffee, as well as culinary blogs, music blogs, art blogs, massage, and wellness blogs. Most recent corpus-based studies on metaphors utilize the Conceptual Metaphor Theory by Lakoff and Johnson. Recently, however, a ‘domain’ has been replaced with the concept of frame. In this project, frames were built up from scratch and were adjusted to the texts. The paper outlines the analytical procedure employed during the corpus compilation, and the main results—statistics of source and target frames and their elements and frame-based models of synesthesia in the corpus.
Morphological synthesis of Georgian words requires to compose the word-forms by indication unchanged parts and morphological categories. Also, it is necessary by using a stem of the given word to get by the computer all grammatically right word-forms. In case of morphological analysis of Georgian words, it is essential to decompose the given word into morphemes and get the definition each of them. For solving these tasks we have developed some specific approaches and created software. Its tools are efficient for a language, which has free order of words and morphological structure is like Georgian. For example, a Georgian verb (in Georgian: “ ” - ts’era, in English: Writing) has several thousand verb-forms. It is very difficult to express morphological analysis’ rules by finite automaton and it will be inefficient as well. Splitting of some Georgian verb-forms into morphemes requires non-deterministic search algorithm, which needs many backtracks. To minimize backtracking, it is necessary to put constraints, which exist among morphemes and verify them as soon as possible to avoid false directions of search. Sometimes the constraints can be as a description type of specific cases of verbs. Thus, proposed software tools have many means to construct efficient parser, test and correct it. We realized morphological and syntactic analysis of Georgian texts by these tools. Besides this, for solving such problems of artificial intelligence, which requires composing of natural language’s word-form by using the information defining this word-form, it is convenient to use the software developed by us.
The article deals with the problem of assessing a visualization of the similarity of documents. A well-known approach for showing the similarity of text documents is a scatter plot generated by projecting text documents into a multidimensional feature space and then reducing the dimensionality to two. The problem stems from the fact that there is a large set of possible document vectorization methods, dimensionality reduction methods and their hyperparameters. Therefore, one can generate many possible charts. To enable a qualitative comparison of different scatter plots, the authors propose a set of metrics that assume that the documents are labeled. Proposed measures quantify how the similarity/dissimilarity of original text documents (described by labels) is maintained within a low-dimensional space. The authors verify the proposed metrics on three corpora, seven different vectorization methods, and three reduction algorithms (PCA, t-SNE, UMAP) with many values of their hyperparameters. The results suggest that t-SNE and fast Text trained on the KGR10 dataset is the best solution for visualizing the semantic similarity of text documents in Polish.
PolEval is a SemEval-inspired evaluation campaign for natural language processing tools for Polish. Submitted tools compete against one another within certain tasks selected by organizers, using available data and are evaluated according to pre-established procedures. It is organized since 2017 and each year the winning systems become the state-of-the-art in Polish language processing in the respective tasks. In 2019 we have organized six different tasks, creating an even greater opportunity for NLP researchers to evaluate their systems in an objective manner.
Ontology Repository Tool is a piece of software aimed to build wordnet-based ontologies which are an example of an information language designed to represent knowledge of various kinds, including general, domain, and application knowledge. This language extends the structure of the wordnet with new types of relations that make it possible to build synset hierarchies parallel to the wordnet structure. Ontology Repository Tool is equipped with functionalities that facilitate the development and management of polyhierarchical and polyrelational knowledge structures. Moreover, the software is intended for integration with information systems such as e-learning content repositories and information systems with multimodal data in which the wordnet-based ontology is deployed and expanded while indexing the documents stored there.
The rapid technological development has created new opportunities for language digitalization and the development of language technology applications. The core element of language technology is language resources, which is in a broad sense, can be considered as a scope of the databases that consists of the myriad of texts both in oral and written forms and used in the machine-learning algorithm. The creation of language resources requires two processes: the first one is language digitalization, meaning the transformation of the speech and texts into the machine-responsible form. The second process refers to text mining, which analyzes data by using a machine-learning algorithm. Adoption of the General Data Protection Regulation (GDPR) and Directive on copyright and related rights in the Digital Single Market (DSM Directive) has been building a renewed legal framework that addresses the demands of the digital economies and unseals challenges, opens prospects for further development. We examine the language resources from two perspectives. Firstly, the language resources are considered a database covered by the protection regulation (the person's rights who created the LR database). Within the second perspective, the legal analysis focuses on the materials used for the language resource creation (data subject's rights, copyright, related rights). The result of the research can be used for further legal investigations and policy design in the field of language technology development.
In this paper we present the use of NLP tools for lexical structure studies of the literary output of a writer. We present the usage made of several tools of our own design or developed externally. In this number were POLEX and Text SubCorpora Creator (TSCC1.3.) systems developed at AMU, as well as NLTK libraries, Corpusomat (pol. Korpusomat), and others. In particular these systems and tools were used to characterize the lexical component of the linguistic instrumentarium of Tadeusz Boy-Żeleński prose and journalistic author active from 1920 to 1941 and Julia Hartwig, an outstanding Polish poet and prose author active from 1954 to 2016.