This release contains errors in several files. Please use http://hdl.handle.net/11234/1-1983 instead.
Slovak Dependency Treebank (Slovenský zavislostný korpus) was created as part of the Slovak National Corpus at the Ľ. Stur Institute of the Slovak Academy of Sciences. The annotation follows the guidelines of the Prague Dependency Treebank (Czech), slightly modified in the spirit of Slovak grammatical tradition. Morphological tags, lemmas and dependency relations have been assigned manually to every word. The present dataset is a subset of the original treebank. We automatically selected the sentences where the two human annotators 100% agreed on the analysis. This increases the quality and trustworthiness of the data but it also results in selecting short sentences most of the time. An extended version may be published in the future when manually merged and checked annotation is available. The selected sentences have been converted to the CoNLL-X file format (original token IDs are preserved in the FEATS column). This PDT-style annotation will serve as the source for the first Slovak dataset in the Universal Dependencies (to be published separately).
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
The paper presents a comparison of time expressions translations from Slovak to Bulgarian language using existing Bulgarian-Slovak/Slovak-Bulgarian electronic text corpora. It uses techniques of concordances and collocations generation from sentence aligned parallel bilingual text corpora to compare semantic content and related translations. The approach is applicable in language teaching and is useful for teaching less resourced languages. (C) 2015 The Authors. Published by Elsevier Ltd.