In this paper we will introduce two new language resources, two NE-annotated corpora for Estonian: Estonian Universal Dependencies Treebank (EDT, 440,000 tokens) and Estonian Universal Dependencies Web Treebank (EWT, 90,000 tokens). Together they make up the largest publicly available Estonian named entity gold annotation dataset. Eight NE categories are manually annotated in this dataset, and the fact that it is also annotated for lemma, POS, morphological features and dependency syntactic relations, makes it more valuable. We will also show that dividing the set of named entities into clear-cut categories is not always easy.
The aim of the paper is to study the effect of pre-annotated clause boundaries on dependency parsing of Estonian new media texts. Our hypothesis is that correct identification of clause boundaries helps to improve parsing because as the text is split into smaller syntactically meaningful units, it should be easier for the parser to determine the syntactic structure of a given unit. To test the hypothesis, we performed two experiments on a 14,000-word corpus of Estonian web texts whose morphological analysis had been manually validated. In the first experiment, the corpus with gold standard morphological tags was parsed with MaltParser both with and without the manually annotated clause boundaries. In the second experiment, only the segmentation of the text was preserved and the morphological analysis was done automatically before parsing. The experiments confirmed our hypothesis about the influence of correct clause boundaries by a small margin: in both experiments, the improvement of LAS was 0.6%.
This release contains errors in several files. Please use http://hdl.handle.net/11234/1-1983 instead.
This article is about annotating clauses with nonverbal predication in version 2 of Estonian UD treebank. Three possible annotation schemas are discussed, among which separating existential clauses from copular clauses would be theoretically most sound but would need too much manual labor and could possibly yield inconcistent annotation. Therefore, a solution has been adapted which separates existential clauses consisting only of subject and (copular) verb olema be from all other olema-clauses.
This article gives an overview of the state of art of tools and resources for syntactic analysis of Estonian. A morphosyntactic disambiguator, surface-syntactic analyzer and dependency parser are all based on the Constraint Grammar formalism. As for language resources, a 400,000-word manually annotated dependency treebank has been created, its annotation scheme is compatible with the output of the Constraint Grammar dependency parser. Part of the treebank has been converted to the Universal Dependencies annotation scheme. Our tools have also been tested by large-scale corpus annotation.
This paper presents the first version of Estonian Universal Dependencies Treebank which has been semi-automatically acquired from Estonian Dependency Treebank and comprises ca 400,000 words (ca 30,000 sentences) representing the genres of fiction, newspapers and scientific writing. Article analyses the differences between two annotation schemes and the conversion procedure to Universal Dependencies format. The conversion has been conducted by manually created Constraint Grammar transfer rules. As the rules enable to consider unbounded context, include lexical information and both flat and tree structure features at the same time, the method has proved to be reliable and flexible enough to handle most of transformations.The automatic conversion procedure achieved LAS 95.2%, UAS 96.3% and LA 98.4%. If punctuation marks were excluded from the calculations, we observed LAS 96.4%, UAS 97.7% and LA 98.2%.Still the refinement of the guidelines and methodology is needed in order to re-annotate some syntactic phenomena, e.g. inter-clausal relations. Although automatic rules usually make quite a good guess even in obscure conditions, some relations should be checked and annotated manually after the main conversion.
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Artikkel käsitleb ühendverbide tuvastamist eesti keele automaatse pindsüntaktilise analüüsi käigus. Ühendverbide äratundmine on vajalik lause täpsemaks süntaktiliseks analüüsiks, sest lause osaliste süntaktilised funktsioonid, semantilised rollid ja nende keelendamine sõltub sellest, milline on lause keskmeks olev predikaatverb, sh sellest, kas predikaatverb on lihtverb või ühendverb. Ühendverbide tuvastamineks rakendatakse kahte strateegiat: leksikonipõhist ja reeglipõhist. Viimane tähendab seda, et osa korrapäraseid produktiivselt kombineeruvaid ühendverbe pannakse kokku reeglite abil. Artiklis kirjeldatakse kahte eksperimenti: esialgse ja täiustatud ühendverbide tuvastamise käiku. Täiustatud süsteemi tulemus on päris hea, saavutades saagise 97,4% ja täpsuse 96,6%.
Artikkel käsitleb ühendverbide tuvastamist eesti keele automaatse pindsüntaktilise analüüsi käigus. Ühendverbide äratundmine on vajalik lause täpsemaks süntaktiliseks analüüsiks, sest lause osaliste süntaktilised funktsioonid, semantilised rollid ja nende keelendamine sõltub sellest, milline on lause keskmeks olev predikaatverb, sh sellest, kas predikaatverb on lihtverb või ühendverb. Ühendverbide tuvastamineks rakendatakse kahte strateegiat: leksikonipõhist ja reeglipõhist. Viimane tähendab seda, et osa korrapäraseid produktiivselt kombineeruvaid ühendverbe pannakse kokku reeglite abil. Artiklis kirjeldatakse kahte eksperimenti: esialgse ja täiustatud ühendverbide tuvastamise käiku. Täiustatud süsteemi tulemus on päris hea, saavutades saagise 97,4% ja täpsuse 96,6%. DOI: http://dx.doi.org/10.5128/ERYa10.14
This article investigates the role of particle verbs in the Estonian computational syntax in the framework of Constraint Grammar, a rule-based system that performs morphological disambiguation, determines grammatical relations and analyses dependency structure of a sentence. For recognizing the particle verbs, two-fold approach is used: non-compositional particle verbs are listed in a lexicon and compositional ones are composed by the rules. The system achieves both 97.4% precision and 97.4% recall for particle verb recognition. The plans for future include building a valency lexicon of particle verbs and utilizing that in syntactic analysis.
Proceedings of the NODALIDA 2009 workshop Constraint Grammar and robust parsing. Editors: Eckhard Bick, Kristin Hagen, Kaili Muurisep and Trond Trosterud. NEALT Proceedings Series, Vol. 8 (2009), 22-29. © 2009 The editors and contributors. Published by Northern European Association for Language Technology (NEALT) http://omilia.uio.no/nealt . Electronically published at Tartu University Library (Estonia) http://hdl.handle.net/10062/14180 .
This paper introduces our work for adapting a rule based parser of spoken Estonian to the morphologically unambiguous part of the corpus of dialects. A Constraint Grammar based parser was used for shallow syntactic analysis of Estonian dialects. The recall of the grammar was 96-97% and the precision 87-89%.
This paper introduces our strategy for adapting a rule based parser of written language to transcribed speech. Special attention has been paid to disfluencies (repairs, repetitions and false starts). A Constraint Grammar based parser was used for shallow syntactic analysis of spoken Estonian. The modification of grammar and additional methods improved the recall from 97.5% to 97.6% and precision from 91.6% to 91.8%. Also, the paper gives a detailed analysis of the types of errors made by the parser while analyzing the corpus of disfluencies.
This paper introduces our work for adapting a rule based parser of spoken Estonian to the morphologically unambiguous part of the cor- pus of dialects. A Constraint Grammar based parser was used for shallow syntactic analysis of Estonian dialects. The recall of the grammar was 96-97% and the precision 87-89%.
This paper discusses some issues of developing a parser for spoken Estonian which is based on an already existing parser for written language, and employs the Constraint Grammar framework. When we used a corpus of face-to-face everyday conversations as the training and testing material, the parser gained the recall 97.6