
Treebanks can play a crucial role in developing natural language processing systems and to have a gold-standard treebank data it becomes necessary to adopt a uniform framework for the annotations. Universal Dependencies (UD) aims to develop cross-linguistically consistent annotations for the world’s languages. The current paper presents the essential pivots of a UD-based syntactically annotated treebank for Malayalam. Sentences extracted from the IndicCorp corpus were manually annotated for morphological features and dependency relations. Language-specific properties are discussed which shed light on many of the grammatical ar-eas in the Dravidian language syntax which needs to be examined in depth. This paper also discusses some pertaining issues in UD taking into consideration the Dravidian languages and provides insights for further improvements in the existing treebanks.
The Prague and Penn styles of discourse annotation are close to each other in basic theoretical views and also in taxonomies of semantic types of discourse relations. A transformation from one of the annotation styles to the other should seemingly be a straightforward process. And yet, slight differences in the taxonomies and significant differences in the technical ap-proaches present several interesting theoretical and practical challenges. The paper focuses on handling the most important issues in the transformation process from the Prague style to the Penn style of discourse annotation, in an effort to bring a valuable data resource – the Prague Discourse Treebank – closer to the international scientific community.
We present a method for supervised cross-lingual construction of word-formation networks (WFNs). WFNs are resources capturing derivational, compositional and other relations be-tween lexical units in a single language. Current state-of-the-art methods for automatically creating them typically rely on supervised or unsupervised pattern-matching of affixes in string representations of words, with few recent inroads into deep learning. All methods known to us work purely in a monolingual setting, limiting the use of higher-quality supervised models to high-resource languages. In this paper, we present two methods, one based on cross-lingual word alignments and translation and another based on cross-lingual word embeddings and neural networks. Both methods are capable of transfer of WFNs into languages for which no word-formational data are available. We evaluate our models on manually-annotated word-formation data from the Universal Derivations and UniMorph projects.
This paper presents the architecture of a derivational database of Modern Hebrew (and more generally of Semitic languages) called Hebrewnette . The methodology adopted is based on adjusting the structure and properties of a database developed for the description of the derivational relations in the lexicon of a Romance language ( Démonette ), and providing it with additional features to account for the specificities of the morphology of Semitic languages, with special reference to root-and-pattern non-concatenative morphology. We present the properties of Hebrewnette and the type of information it consists of, with emphasis on both structural and semantic relations between words. We show how this is implemented and examine two case studies, where we demonstrate how the annotations that are used allow us to verify theoretical hypotheses about non-concatenative morphology. The design of Démonette ’s annotation system allow its features, initially designed for French, to capture morphological and semantic relations between Hebrew words, regardless of the type of morphology (concatenative or non-concatenative).
In the paper, we report results of our experiments on identifying distributional semantic characteristics of different types of lexical items used in argumentation: connectors, meta-ar-gumentative words, key notions of a given discourse and the evaluative/connotative lexicon. These characteristics are contrasted within monolingual English corpora of different genres (Europarl-EN and Cord COVID-19) and in translation context for German-into-English direc-tion (Europarl-DE and Europarl-EN-from-DE). For the analysis, we propose a number of new methods that better characterize distributional semantic differences between the argumentatively relevant lexical items, such as measuring the knee in the mutual information-ranked list curve, testing categories for span variation and different selection procedures. In our experiments, meta-argumentative lexical items show the biggest differences in their distribution with other word types on several of such measures. The analysis based on word vector allows us to create a selection heuristic for candidate lists for different categories of argumentative lexicon.
Reflexivity represents one of the core research tasks in current linguistics. As the use of reflexives, encoding a variety of meanings, typically brings about changes in verb valency, the description of reflexivity is highly relevant – among others – also for valency oriented studies. In this paper, we address the reflexive in Czech categorized as a derivational morpheme (e.g., zlomit pf ‘to break something’ → zlomit se pf ‘to break; to crack’), with the focus on valency behavior of reflexive verbs as represented in the valency lexicon of Czech verbs VALLEX . In the data component of the lexicon, reflexive verbs, i.e., verbs with reflexive lexemes, are captured in separate lexicon entries, represented by respective verb lemma(s) containing the free reflexive morpheme se or si . In VALLEX , there are 922 lexical entries for reflexive verbs described in 1545 lexical units represented by 1525 verb lemmas (this number covers almost one quarter of lexical units and one third of verb lemmas in the lexicon). Reflexive verbs can be divided into two groups: into those without any non-reflexive counterpart ( reflexiva tantum , 208 lexical units represented by 177 verb lemmas) and into those for which a non-reflexive base verb can be identified ( derived reflexive verbs , 1337 lexical units represented by 1348 verb lemmas). Those derived reflexive verbs that are directly related to their non-reflexive base verbs are classified into seven types on the ground of their relation to the non-reflexive counterparts, captured in the data component of the lexicon by the value of the attribute reflexverb . Further, the relation of derived reflexive verbs and their respective non-reflexive counterparts is described by formal rules comprised in the grammar component of the lexicon (19 rules in total), which provide the information on changes in the mapping of
We present a study of discourse connectives in English-German and German-English translation and interpreting where we focus on the phenomena of explicitation and implicitation. Apart from distributional analysis of translation patterns in parallel data, we also look into surprisal, i.e. an information-theoretic measure of cognitive effort, which helps us to interpret the observed tendencies.
In this paper we present our method to build a derivational database of French deanthroponyms, which we call MONOPOLI for Mo ts construits sur No ms propres de personnalités Poli tiques (‘complex words based on politician proper names’). MONOPOLI contains 6,545 complex words amounting to a total of 55,030 tokens and includes almost only neologistic forms. The Web is the only conceivable resource for collecting them: it alone gives massive access to discourse genres that contain neologisms. To feed the database, a program automatically generates the set of all possible derived words. Generated forms are then used as queries on the Web. Attested forms are kept with their context. This method provides a potential solution to collect data that cannot be found elsewhere. Finally, this article describes some of the remarkable results obtained with the analysis of the deanthroponyms of MONOPOLI. We show that the original nature of our data is reflected both in the use of new extragrammatical patterns (e.g., X istan secretion pattern) and in the subversion of grammatical processes (e.g., foreign suffixation X ix ).
In this paper 1 we document both the structural and the diachronic extension of the derivational information provided in the LiLa Knowledge Base of interoperable linguistic resources for Latin. Structurally, to the flat information on families (i.e., groups of lemmas that share the same base) and affixes that is already available for the collection of lemmas of the LiLa Lemma Bank, we add hierarchical information on derivation processes provided by the Word Formation Latin (WFL) lexical resource, which in turn is characterised by a step-to-step mor-photactic approach, where lexemes that are directly derived from one another are connected through word formation rules of different kinds. This is done by modelling WFL data into an ontology that adheres to the principles of the Linked Data paradigm, and connecting these data to the LiLa Lemma Bank. From a diachronic point of view, while the previous version of WFL only took Classical Latin lemmas into account, in this paper we describe the work conducted to produce a new version of WFL that is enhanced with derivational information on Medieval Latin lemmas. We then show how the data of this new version of WFL were used to extract derivational information in the format required by the LiLa Lemma Bank.
We present a deep-learning tool called Word Formation Analyzer for Czech , which, given an input lexeme, automatically retrieves the lemma or lemmas from which the input lexeme was formed. We call this task parent retrieval. Furthermore, based on the number of words in the output sequence and its comparison to the input, the input word is classified into one of three categories: compound , derivative or unmotivated . We call this task word formation classification. In the task of parent retrieval, Word Formation Analyzer for Czech achieved an accuracy of 71%. In word formation classification, the tool achieved an accuracy of 87%.
This study focuses on cases of suffixal rivalry in denominal adjective formations in Russian, namely on two adjectival suffixes: -nand -sk-. We use statistical modelling (multivariate logistic regression) to shed light on properties of base nouns that contribute to the choice of one of the competing suffixes. In the first part, we provide model interpretation through traditional metrics (accuracy, confusion matrix and model coefficients with their respective pvalues). However, model accuracy may not be uniform if we compare different samples of the data set and may take a wide range of values. In the second part of this study, we complete our interpretation of model results by performing error analysis in order to get a better understanding of the underlying properties of base nouns that cause model failure. We explore Responsible AI Toolbox widgets for this purpose. Onemain result of this study is that the same semantic base noun properties are related to both high model performances and model errors.
Wepresent amethod for extending coverage of the Lexicon of Czech Discourse Connectives – CzeDLex – using annotation projection. We take advantage of two language resources: (i) the Penn Discourse Treebank 3.0 as a source of manually annotated discourse relations in English, and (ii) the Prague Czech–English Dependency Treebank 2.0 as a translation of the English texts to Czech and a link between tokens on the two language sides. Although CzeDLex was originally extracted from a large Czech corpus, the presented method resulted in an addition of a number of new connectives and new types of usages (discourse types) for already present entries in the lexicon. We classify and elaborate on reasons why the rest of automatically preselected candidateswere excluded from the process, and give examples of actual newadditions.
Reflexives, encoding a variety of meanings, pose a great challenge for both theoretical and lexicographic description. As they are associated with changes in morphosyntactic properties of verbs, their description is highly relevant for verb valency. In Czech, reflexives function as the reflexive personal pronoun and as verbal affixes. In this paper, we address those language phenomena that are encoded by the reflexive personal pronoun, i.e., reflexivity and reciprocity. We introduce the lexicographic representation of these two language phenomena in the VALLEX lexicon, a valency lexicon of Czech verbs, accounting for the role of the reflexives with respect to the valency structure of verbs. This representation makes use of the division of the lexicon into a data component and a grammar component. It takes into account that reflexivity and reciprocity are conditioned by the semantic properties of verbs on the one hand and that morphosyntactic changes brought about by these phenomena are systemic on the other. About one third of the lexical units contained in the data component of the lexicon are assigned the information on reflexivity and/or reciprocity in the form of pairs of the affected valency complementations (2,039 on reflexivity and 2,744 on reciprocity). A set of rules is formulated in the grammar component (3 rules for reflexivity and 18 rules for reciprocity). These rules derive the valency frames underlying syntactically reflexive and reciprocal constructions from the valency frames describing non-reflexive and non-reciprocal constructions. Finally, the proposed representation makes it possible to determine which lexical units of verbs create ambiguous constructions that can be interpreted either as reflexive or as reciprocal. © 2021 PBML. Distributed under CC BY-NC-ND. Corresponding author: kettnerova@ufal.mff.cuni.cz Cite as: Václava Kettnerová, Markéta Lopatková, Anna Vernerová. Reflexives in the VALLEX Lexicon: Syntactic Reflexivity and Reciprocity. The Prague Bulletin of Mathematical Linguistics No. 117, 2021, pp. 27–60. doi: 10.14712/00326585.016. PBML 117 OCTOBER 2021
We propose a new architecture for diacritics restoration based on contextualized embeddings, namely BERT, and we evaluate it on 12 languages with diacritics. Furthermore, we conduct a detailed error analysis on Czech, a morphologically rich language with a high level of diacritization. Notably, we manually annotate all mispredictions, showing that roughly 44% of them are actually not errors, but either plausible variants (19%), or the system corrections of erroneous data (25%). Finally, we categorize the real errors in detail. We release the code at https://github.com/ufal/bert-diacritics-restoration.
The foundation for the research of summarization in the Czech language was laid by the work of Straka et al. (2018). They published the SumeCzech, a large Czech news-based summarization dataset, and proposed several baseline approaches. However, it is clear from the achieved results that there is a large space for improvement. In our work, we focus on the impact of named entities on the summarization of Czech news articles. First, we annotate SumeCzech with named entities. We propose a new metric ROUGE_NE that measures the overlap of named entities between the true and generated summaries, and we show that it is still challenging for summarization systems to reach a high score in it. We propose an extractive summarization approach Named Entity Density that selects a sentence with the highest ratio between a number of entities and the length of the sentence as the summary of the article. The experiments show that the proposed approach reached results close to the solid baseline in the domain of news articles selecting the first sentence. Moreover, we demonstrate that the selected sentence reflects the style of reports concisely identifying to whom, when, where, and what happened. We propose that such a summary is beneficial in combination with the first sentence of an article in voice applications presenting news articles. We propose two abstractive summarization approaches based on Seq2Seq architecture. The first approach uses the tokens of the article. The second approach has access to the named entity annotations. The experiments show that both approaches exceed state-of-the-art results previously reported by Straka et al. (2018), with the latter achieving slightly better results on SumeCzech's out-of-domain testing set.