The paper presents an online editor for lexical-semantic databases with relational structure similar to the structure of WordNet - Hydra for Web. It supports functionalities for editing of relational data (including query, creation, change, and linking of relational objects), simultaneous access of multiple user profiles, parallel data visualization and editing of the data on top of single- and parallel mode visualization of the language data.
The paper presents the CRUD functions of the Hydra for Web system for work on lexical-semantic databases with relational structure similar to the structure of WordNet. It supports functionalities for editing of relational data, simultaneous access of multiple users, parallel data visualisation. Hydra for Web has been used for the development of the Bulgarian wordnet.
This paper presents a machine learning method for automatic identification and classification of morphosemantic relations (MSRs) between verb and noun synset pairs in the Bulgarian WordNet (BulNet). The core training data comprise 6,641 morphosemantically related verb–noun literal pairs from BulNet. The core dataset were preprocessed quality-wise by applying validation and reorganisation procedures. Further, the data were supplemented with negative examples of literal pairs not linked by an MSR. The designed supervised machine learning method uses the RandomTree algorithm and is implemented in Java with the Weka package. A set of experiments were performed to test various approaches to the task. Future work on improving the classifier includes adding more training data, employing more features, and fine-tuning. Apart from the language specific information about derivational processes, the proposed method is language independent.
This paper presents a web interface for wordnets named Hydra for Web which is built on top of Hydra – an open source tool for wordnet development – by means of modern web technologies. It is a Single Page Application with simple but powerful and convenient GUI. It has two modes for visualisation of the language correspondences of searched (and found) wordnet synsets – single and parallel modes. Hydra for web is available at: http://dcl.bas.bg/bulnet/.
This paper presents work in progress on a machine learning method for classification of morphosemantic relations between verb and noun synsets. The training data comprises 5,584 verb–noun synset pairs from the Bulgarian WordNet, where the morphosemantic relations were automatically transferred from the Princeton Word-Net morphosemantic database. The machine learning is based on 4 features (verb and noun endings and their respective semantic primes). We apply a supervised machine learning method based on a decision tree algorithm implemented in Python and NLTK. The overall performance of the method reached F 1 -score of 0 . 936 . Our future work focuses on automatic iden-tification of morphosemantically related synsets and on improving the classification.
In the context of developing wordnets and using them in various applications, we have been enriching the Romanian and Bulgarian resources with morphosemantic relations that can aid broadening the wordnet content and improving the possible NLP applications. In this paper, we build on our previous results, adding to our presentation data from English. As a consequence, we offer a comparative study-on noun-verb derivation in the three languages (from three different branches of the Indo-European language family). Our results can serve as training material for the automatic identification and assignment of derivational and morphosemantic relations needed in various applications.
This paper presents Hydra for Web - a web interface for wordnets (and lexical-semantic databases with similar relational structure). Hydra for web is built on top of Hydra - an open source tool for wordnet development - and is a single page application with a simple GUI. It has two modes - single and parallel - for visualisation of the language correspondences of searched words. The main focus of the paper is on the user interface, its functionalities and some issues regarding the ongoing process of adapting the resources that are currently visualised - the Princeton WordNet, the Romanian wordnet, and the Bulgarian wordnet.
The paper presents Anaphora -an OS and language independent tool for clause annotation and alignment, developed at the Department of Computational Linguistics, Institute for Bulgarian Language, Bulgarian Academy of Sciences.The tool supports automated sentence splitting and alignment and modes for manual monolingual annotation and multilingual alignment of sentences and clauses.Anaphora has been successfully applied for the annotation and the alignment of the Bulgarian-English Sentence-and Clause-Aligned Corpus (Koeva et al. 2012a) and a number of other languages including French and Spanish.
The paper motivates a strategy for identification and annotation of derivational relations in the Bulgarian wordnet that aims at coping with the complex morphology of the language in an elegant way. Our method involves transfer of the Princeton WordNet (morpho)semantic relations into the Bulgarian wordnet, at the level of the synset, and further detection of derivational relations between literals in Bulgarian. Derivational relations have been annotated to reflect the complexity of Bulgarian morphology. Introduced literal relations improve the consistency and employability of the
This paper presents an overview of the software for wordnet processing Hydra. The system has fully-fledged GUI and API, both working with powerful modal query language. Hydra has been used for the development of the Bulgarian WordNet for the last 7 years and recently was improved, became open source and is distributed as part of the Meta-Share platform.
The paper presents the methodology and the outcome of the compilation and the processing of the Bulgarian X-language Parallel Corpus (Bul-X-Cor) which was integrated as part of the Bulgarian National Corpus (BulNC). We focus on building representative parallel corpora which include a diversity of domains and genres, reflect the relations between Bulgarian and other languages and are consistent in terms of compilation methodology, text representation, metadata description and annotation conventions. The approaches implemented in the construction of Bul-X-Cor include using readily available text collections on the web, manual compilation (by means of Internet browsing) and preferably automatic compilation (by means of web crawling - general and focused). Certain levels of annotation applied to Bul-X-Cor are taken as obligatory (sentence segmentation and sentence alignment), while others depend on the availability of tools for a particular language (morpho-syntactic tagging, lemmatisation, syntactic parsing, named entity recognition, word sense disambiguation, etc.) or for a particular task (word and clause alignment). To achieve uniformity of the annotation we have either annotated raw data from scratch or transformed the already existing annotation to follow the conventions accepted for BulNC. Finally, actual uses of the corpora are presented and conclusions are drawn with respect to future work.
The paper presents a new resource light flexible method for clause alignment which combines the Gale-Church algorithm with internally collected textual information. The method does not resort to any pre-developed linguistic resources which makes it very appropriate for resource light clause alignment. We experiment with a combination of the method with the original Gale-Church algorithm (1993) applied for clause alignment. The performance of this flexible method, as it will be referred to hereafter, is measured over a specially designed test corpus. The clause alignment is explored as means to provide improved training data for the purposes of Statistical Machine Translation (SMT). A series of experiments with Moses demonstrate ways to modify the parallel resource and effects on translation quality: (1) baseline training with a Bulgarian-English parallel corpus aligned at sentence level; (2) training based on parallel clause pairs; (3) training with clause reordering, where clauses in each source language (SL) sentence are reordered according to order of the clauses in the target language (TL) sentence. Evaluation is based on BLEU score and shows small improvement when using the clause aligned corpus.
espanolEl articulo describe la metodologia de la compilacion y la anotacion del Corpus Semanticamente Anotado Bulgaro - un corpus anotado de modo manual y consta de mas de 100 mil palabras donde en cada unidad linguistica se le ha atribuido un significado conforme al Wordnet Bulgaro. El articulo presenta, asimismo, el programa de anotacion Chooser. Han sido descritas las convenciones linguisticas y las soluciones practicas adoptadas en el proceso de la anotacion. Al final, el articulo describe una de las aplicaciones esenciales de BulSemCor como un corpus de entrenamiento orientado a desarrollar el sistema de desambiguacion de la lengua bulgara. EnglishThis paper describes the methodology of compilation and annotation of the Bulgarian Sense-Annotated Corpus - a manually annotated corpus of over 100,000 words in which each lexical unit (LU) is assigned a sense according to the Bulgarian wordnet. The paper gives a brief outline of the corpus representation, the functionalities of the annotation tool Chooser, and sketches the linguistic conventions and practical considerations adopted in the process of corpus annotation. Finally, the paper describes one of the major applications of the Bulgarian Sense-Annotated Corpus as a training corpus for a word-sense disambiguation system for Bulgarian.
This paper presents a multipurpose system for wordnet (WN) development, named Hydra. Hydra is an application for data editing and validation, as well as for data retrieval and synchronization between wordnets for different languages. The use of modal language for wordnet, the representation of wordnet as a relational database and the concurrent access are among its main advantages (Rizov, 2006).
The paper presents a tool assisting manual annotation of linguistic data developed at the Department of Computational linguistics, IBL-BAS. Chooser is a general-purpose modular application for corpus annotation based on the principles of commonality and reusability of the created resources, language and theory independence, extendibility and user-friendliness. These features have been achieved through a powerful abstract architecture within the Model-View-Controller paradigm that is easily tailored to task-specific requirements and readily extendable to new applications. The tool is to a considerable extent independent of data format and representation and produces outputs that are largely consistent with existing standards. The annotated data are therefore reusable in tasks requiring different levels of annotation and are accessible to external applications. The tool incorporates edit functions, pass and arrangement strategies that facilitate annotators' work. The relevant module produces tree-structured and graph-based representations in respective annotation modes. Another valuable feature of the application is concurrent access by multiple users and centralised storage of lexical resources underlying annotation schemata, as well as of annotations, including frequency of selection, updates in the lexical database, etc. Chooser has been successfully applied to a number of tasks - POS tagging, WS annotation, syntactic annotation.