The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, especially coreference and discourse relations. We present its second consolidated version (PDT-C 2.0), which concludes almost 30-years long project of sustained development of the resource to a uniformly and coherently annotated, genre-diversified, almost 4 million token language resource of Czech language, with accompanying fully compatible lexicons. In addition to continuous linguistic research, the richly linguistically annotated corpus is also widely used in international comparisons of the development of traditional and novel NLP tools as well as in conversions into other formalisms. The corpus and the trained parsers are available under the CC BY-NC-SA licence.
As previous research on annotator disagreement in discourse phenomena has shown, understanding text coherence varies considerably from one individual to another. To explore this phenomenon, we created two corpora with multiple annotations of Czech texts, accompanied by annotators' explanations of their choices. The first corpus consists of 1,024 contexts annotated in parallel by three annotators. It captures differences in the identification of coreference across various text types and grammatical-semantic categories, including pronouns, full noun phrases, and anaphoric adverbials. The second corpus comprises 512 contexts, annotated in parallel by five annotators, and focuses on identifying discourse relations in attributive and non-attributive constructions. Both corpora achieve a comparable inter-annotator agreement of approximately 60-65
NomVallex 2.5 is a valency lexicon of Czech nouns and adjectives that besides valency of particular noun and adjectival lexical units captures several valency-related syntactic phenomena, namely (i) active and passive syntax, (ii) systemic and non-systemic valency behavior, (iii) reflexivity and reciprocity, and (iv) negation. These phenomena, represented in the lexicon mostly by deverbal and deadjectival derivatives, are described here with respect to the derivational category of denominal adjectives. Though being regarded as less typical representatives of valency bearers in the non-verbal domain, denominal adjectives turned out to be involved in all of the studied syntactic phenomena. Their syntactic behavior thus can be viewed as similar to other derivational categories, especially to deverbal adjectives.
We introduce a first step to modelling valency frames of selected types of nominals. We work on the assumption that nominals inherit – at least to some extent – valency from their base verbs. We illustrate this task in a case study focused on modelling valency frames of Czech deverbal adjectives -teln ý ‘able’. First, the valency frames of the adjectives -teln ý contained in NomVallex are compared with the valency frames of their base verbs in VALLEX. Based on this comparison, two formal rules describing valency changes in the valency frames of adjectives -teln ý are formulated. Second, for each lexical unit of a verb that satisfies the conditions imposed by some of the rules, the derived adjective -teln ý is extracted from DeriNet, if such an adjective is available. Third, the valency frame of the adjective is derived from the valency frame of the verb based on the respective rule. Lastly, the accuracy of both rules is verified in the corpus data. The experiment has shown that the valency of these adjectives can be modeled on the rule basis. However, if this task is to be accurate, it requires advanced linguistic information, namely the information on semantic class membership of verbs and on compound adjectives.
Discourse relations represent a relatively ambiguous area of language that can be a challenge for computational and corpus linguistics. In our study, we examine the reliability of the Prague Dependency Treebank – Consolidated 2.0 by employing multiple annotations that can reveal possible different readings. It turns out that the complexity of sentence structure has a very significant influence on the distinction of the left discourse argument; in contrast, neither the mode of the text (written vs. spoken) nor the presence of attributive constructions appears to influence the variability in the interpretation of discourse structure. Furthermore, we identify the technical organization of the annotation process itself as a potential source of bias in discourse annotation.
MasKIT is a command-line tool, an on-line web application and a REST API service for anonymization and pseudonymization of Czech legal texts. Taking a plain text as input (e.g. a letter sent by a legal authority to a citizen), it runs external services for dependency parsing and named entity recognition and then via a rule-based approach identifies and replaces sensitive information in the text.
Administrative and legal communication is often difficult for laypersons to understand, creating barriers to justice and undermining trust in public institutions. The PONK tool helps authors identify and revise unclear text using a multimodular approach that combines linguistic rules and lexical surprisal based on large language models. This integration ensures precise and adaptable detection of problematic passages. Human evaluation confirmed the tool's effectiveness in improving legal text comprehensibility. PONK's user-friendly interface highlights unclear segments at the word level, offering a practical solution for clearer and more accessible legal communication.
We present a data story, the central concept of our international onesemester data analytics course for students of social sciences and humanities. Namely, we demonstrate the four stages of the data lifecycle – gathering, analyzing, annotating, licensing & sharing – using the multilingual correspondence collection of French Slavist André Mazon and tools for data visualization (Tableau Public), text transcription (Transkribus, Pero), building and searching text corpora (TEITOK, Corpus Query Language) and natural language processing (UDPipe).
Determining a relative position of a discourse connective and the two arguments (text segments) it connects is an important part of a full discourse parsing task. This paper investigates discourse connectives whose position in a text deviates from the usual setting – namely connectives that occur in neither of the two arguments – and as such present a challenge for discourse parsers. We find syntactic patterns for this phenomenon and describe it linguistically on the basis of Czech discourse-annotated corpus material, with the aim to facilitate an automatic detection of such connectives and a correct localization of their arguments.
The article focuses on two different lexicons providing complementary information: MorfFlex covering general Czech morphology, and VALLEX giving information on the syntax and semantics of Czech verbs. We discuss different designs of these lexicons, concentrating primarily on variants and homographs in the Czech vocabulary. Within the project, we have verified the theoretical approaches and harmonized the treatment of variants in both lexicons, adopting the clear morphologically based criteria from MorfFlex for distinguishing variants in VALLEX. The two updated lexicons, MorfFlex and VALLEX, with interlinked records represent the project’s main outcome.
Abstract NomVallex is a manually annotated valency lexicon of Czech nouns and adjectives that enables a comparison of valency properties of derivationally related lexical units. We present new developments in how the lexicon facilitates research into changes in valency across part-of-speech categories and derivational types. In particular, it provides links from derived lexical units to their base lexical units and also allows to search and display a base lexical unit together with all lexical units directly derived from it. Using an automatic procedure, any difference in valency between two derivationally related lexical units is specified. As a case study, focusing on nouns and adjectives directly or indirectly motivated by verbs, the facilities provided by the lexicon are used to show differences in what ways the particular deverbal derivatives representing various derivational types express the valency complementation standing in the base verbal construction in the subject position.
The Prague and Penn styles of discourse annotation are close to each other in basic theoretical views and also in taxonomies of semantic types of discourse relations. A transformation from one of the annotation styles to the other should seemingly be a straightforward process. And yet, slight differences in the taxonomies and significant differences in the technical ap-proaches present several interesting theoretical and practical challenges. The paper focuses on handling the most important issues in the transformation process from the Prague style to the Penn style of discourse annotation, in an effort to bring a valuable data resource – the Prague Discourse Treebank – closer to the international scientific community.
Josef Psutka合作论文数Coordinator - Center of Computational Linguistics2