This paper presents a comparative study of two approaches to automated assessment of Czech spoken language exams for non-native speakers: one using large language models applied to transcripts, and the other based on pre-trained speech encoder models. To our knowledge, this is the first study to explore automatic speaking assessment (ASA) for the Czech language. We evaluate both methods on a dataset of authentic high-stakes oral exams, annotated with binary pass/fail labels and total exam scores. Our best-performing models reach a QWK score of 0.65. Our experiments demonstrate the feasibility of applying ASA techniques in Czech and illustrate challenges related to data scarcity, transcription quality, and performance variability between input types.
The Prague and Penn styles of discourse annotation are close to each other in basic theoretical views and also in taxonomies of semantic types of discourse relations. A transformation from one of the annotation styles to the other should seemingly be a straightforward process. And yet, slight differences in the taxonomies and significant differences in the technical ap-proaches present several interesting theoretical and practical challenges. The paper focuses on handling the most important issues in the transformation process from the Prague style to the Penn style of discourse annotation, in an effort to bring a valuable data resource – the Prague Discourse Treebank – closer to the international scientific community.
En):In the paper, we explore cohesive devices in Czech texts written by studentsnon-native speakers.Specifically, we focus on conjunctions (significantly contributing to text coherence) from the point of view of text genres; we analyse four basic genres, namely narration, information, argumentation and description, and we follow the relative occurrence of text conjunctions across them.Methodologically, we use data obtained from the CzeSL-SGT corpus (AKCES 5;Šebesta et al., 2014) containing texts written by non-native speakers of Czech.The texts are annotated with text genres, students' age, text evaluation (A-C levels according to CEFR, the Common European Framework of Reference for Languages) and other additional information.For our analysis, we selected texts of narration, information, argumentation and description written by students at the age of 16+.We examined the occurrence of conjunctions at the individual CEFR levels (A, B and C).The absolute numbers of occurrences were converted to relative frequencies (the i.p.m. value -Instances per Million Positions).The aim of the analysis is to find out whether the particular genres differ in frequency and diversity of text conjunctions and whether some difference also occurs across the CEFR levels -A: basic user, B: independent user, C: proficient user.Based on the corpus data, we conclude that conjunction devices contributing to text coherence occur most frequently in the genre of argumentation, followed by narration, less in description and information.The individual genres are explored in detail also with respect to the CEFR levels of A-C.A and B levels exhibit the same tendencies, i.e. conjunction devices occurred most frequently in genres with the following order: argumentation (A: 90,680.64i.p.m.,
In the paper, we study discourse connectives and other discourse phenomena in Czech based on corpus data. We focus on the delimitation of connectives primarily from the functional point of view and we present their further classification into primary and secondary connectives based especially on their degree of grammaticalization. Special attention is paid to the variability of secondary connectives occurring in several lexical realizations and grammatical variants. We also present the frequency of the individual connectives in written texts and we analyze their syntactic behavior. We then discuss the relation of connectives to reference (a special group consists of the anaphoric and cataphoric connectives), connectives occurring next to each other in the text, and the possibilities of modification of connectives. Our analysis is based especially on the language data of the Prague Discourse Treebank 2.0. The aim of the paper is to present a detailed study of discourse connectives in written Czech and to present connectives as an open class covering both grammaticalized expressions and looser lexical phrases.
A richly annotated and genre-diversified language resource, The Prague Dependency Treebank – Consolidated 1.0 (PDT-C 1.0, or PDT-C in short in the sequel) is a consolidated release of the existing PDT-corpora of Czech data, uniformly annotated using the standard PDT scheme. PDT-corpora included in PDT-C: Prague Dependency Treebank (the original PDT contents, written newspaper and journal texts from three genres); Czech part of Prague Czech-English Dependency Treebank (translated financial texts, from English), Prague Dependency Treebank of Spoken Czech (spoken data, including audio and transcripts and multiple speech reconstruction annotation); PDT-Faust (user-generated texts). The difference from the separately published original treebanks can be briefly described as follows: it is published in one package, to allow easier data handling for all the datasets; the data is enhanced with a manual linguistic annotation at the morphological layer and new version of morphological dictionary is enclosed; a common valency lexicon for all four original parts is enclosed. Documentation provides two browsing and editing desktop tools (TrEd and MEd) and the corpus is also available online for searching using PML-TQ.
We introduce the first version of GeCzLex, an online electronic resource for translation equivalents of Czech and German discourse connectives. The lexicon is one of the outcomes of the research on anaphoricity and long-distance relations in discourse, it contains at present anaphoric connectives (ACs) for Czech and German connectives, and further their possible translations documented in bilingual parallel corpora (not necessarily anaphoric). As a basis, we use two existing monolingual lexicons of connectives: the Lexicon of Czech Discourse Connectives (CzeDLex) and the Lexicon of Discourse Markers (DiMLex) for German, interlink their relevant entries via semantic annotation of the connectives (according to the PDTB 3 sense taxonomy) and statistical information of translation possibilities from the Czech and German parallel data of the InterCorp project. The lexicon is, as far as we know, the first bilingual inventory of connectives with linkage on the level of individual entries, and a first attempt to systematically describe devices engaged in long-distance, non-local discourse coherence. The lexicon is freely available under the Creative Commons License.
In the paper, we present research of most common errors in coherence occurring in learners' essays. We carry out corpus-based research of essays written by learners (non-native speakers) of Czech focussing especially on errors concerning discourse connectives: on their use in a relevant context, on most common semantic confusions, on the use of connectives in a right position in a sentence etc. We examine whether there are some shared issues or tendencies concerning coherence errors that could be generalized and formulated as recommendations to students, e.g. which aspects they should pay attention to when writing a well-structured essay. In the next step, we introduce how learners of Czech may improve their writing skills using an online tool EVALD (Evaluator of Discourse). EVALD provides an automated essay scoring and gives also feedback to the users - it points out the strong and weak aspects of the text in various language layers including text coherence and discourse connectives. EVALD is a suitable tool for e-learning which can be used not only by students (language learners) but also by teachers as an assistant tool in their classes.
We present a feature-rich system for automatic evaluation of surface text coherence in Czech essays written by native and non-native speakers. The EVALD system, in addition to basic features covering spelling, vocabulary, morphology and syntax, stands on two main pillars representing the features closely related to the phenomenon of surface coherence: discourse relations and coreference. Newly we add a third pillar, features targeting topic–focus articulation (sentence information structure). Therefore, we propose and implement a procedure for disclosing topic–focus articulation by marking contextual boundness in the text automatically. The experiments show that EVALD enriched with topic–focus articulation features succeeds in outperforming the original system. Further experiments show that the system for essays written by non-native speakers exhibits different signs in terms of importance of individual feature sets and the size of the training data than the system for native speakers.
The paper contributes to the research on automatic evaluation of surface coherence in student essays. We look into possibilities of using large unlabeled data to improve quality of such evaluation. Particularly, we propose two approaches to benefit from the large data: (i) n-gram language model, and (ii) density estimates of features used by the evaluation system. In our experiments, we integrate these approaches that exploit data from the Czech National Corpus into the evaluator of surface coherence for Czech, the EVALD system, and test its performance on two datasets: essays written by native speakers (L1) as well as foreign learners of Czech (L2). The system implementing these approaches together with other new features significantly outperforms the original EVALD system, especially on L1 with a large margin.
Abstract In the paper, we present EVALD applications (Evaluator of Discourse) for automated essay scoring. EVALD is the first tool of this type for Czech. It evaluates texts written by both native and non-native speakers of Czech. We describe first the history and the present in the automatic essay scoring, which is illustrated by examples of systems for other languages, mainly for English. Then we focus on the methodology of creating the EVALD applications and describe datasets used for testing as well as supervised training that EVALD builds on. Furthermore, we analyze in detail a sample of newly acquired language data – texts written by non-native speakers reaching the threshold level of the Czech language acquisition required e.g. for the permanent residence in the Czech Republic – and we focus on linguistic differences between the available text levels. We present the feature set used by EVALD and – based on the analysis – we extend it with new spelling features. Finally, we evaluate the overall performance of various variants of EVALD and provide the analysis of collected results.
As the quality of machine translation rises and neural machine translation (NMT) is moving from sentence to document level translations, it is becoming increasingly difficult to evaluate the output of translation systems. We provide a test suite for WMT19 aimed at assessing discourse phenomena of MT systems participating in the News Translation Task. We have manually checked the outputs and identified types of translation errors that are relevant to document-level translation.
In this paper, we introduce an experimental probe comparing how texts written by non-native speakers of Czech are evaluated by a software application (computer program) EVALD and by teachers of Czech as a foreign language. The hypothesis for the probe was that teachers, even if they are given structured instruction for evaluation and go through the standardization process, are not able to reach satisfactory results and to agree on the same evaluation, which depreciates the whole text evaluation process. This is a problem especially for objective assessment during certificated exams such as the Exam for Permanent Residence. A group of 44 teachers of Czech as a foreign language who underwent special training and the computer program evaluated 2 texts from the point of view of relevant features of the A1–C1 levels established by the Common European Framework of Reference for Languages. The task included evaluation of the overall level of the texts and evaluation of specific aspects of the texts: punctuation, morphology, lexis, syntax and coherence. We compare the evaluation of the text among the teachers and the teachers with the computer program. In the general evaluation, only 41% of persons agreed on the same level for text A and 50% for text B (we describe and interpret the agreement in the evaluation of orthography, morphology, syntax, lexis and coherence below). The lowest rate of interevaluator agreement was in orthography – 38% for text A and 40% for text B. In morphology, 59% of persons agreed on the same level for text A and 61% for text B. We further compared human and automatic agreement on the evaluation. 41% of teachers agreed with the program on the evaluation of text A and 50% of text B. Again, we also compared the results on the particular language and text levels. Our results clearly show that human evaluation is rather inconsistent and it would be advisable to use automatic evaluation in cases where consistency and high agreement is desired.
The paper introduces a lexical database covering anaphoric discourse connectives. The database contains 80 most frequently used connective units consisting of (in their current use or historically) a preposition and an anaphoric element, cf. the connective thereafter containing a preposition after and an anaphoric part there. The database has two parts. The first part includes German anaphoric connectives and their most common equivalents in Czech and English. The second one consists of Czech anaphoric connectives and their counterparts in German and English. The involvement of German, Czech and English enables the use of the database in foreign language teaching with a focus on deepening the productive language skills of students. In addition, the database can be used by translators as well as in the machine translation area.
EVALD 4.0 for Beginners is a software that serves for automatic evaluation of Czech texts written by non-native speakers of Czech – language beginners.
In the present paper, we introduce a contrastive study on text coherence in Czech and German. Specifically, we focus on the discourse-anaphoric devices (called anaphoric connectives) contributing to text coherence and we analyze their role in the overall communicative competence of both native and non-native speakers of Czech and German. In our analysis, we firstly examine anaphoric connectives in Czech and we present their frequencies in three corpora - SYN6 and PDiT 2.0 (for texts by native speakers), and MERLIN (for texts by non-native speakers). Then we take a closer look at anaphoric connectives in texts by non-native speakers (learners) of Czech. Specifically, we examine whether the students who actively use such expressions in their writings also have better communicative competence as a whole (i.e. they reach better overall grade). Subsequently, we focus on the ways anaphoric connectives in Czech are translated into German (using the corpus InterCorp) and we provide an analysis of these German counterparts. In the next step, we carry out the same type of analysis for anaphoric connectives in German (with the PCC and the DWDS corpora for native speakers and the MERLIN corpus for non-native speakers). We analyze the results for Czech and German separately, and finally, we carry out a comparative study of these two languages with a focus on the use of our findings in the teaching process and possibly in (automated) translation.
In the present paper, we examine discourse connectives from the perspective of reference (i.e. a presence of an anaphoric element). We introduce a division of connectives into: i) connectives without an inherent (internal) reference (e.g. and, but, or, if, however, so), and ii) connectives with an inherent (internal) reference that is either optional (e.g. as a result vs. as a result of this), or obligatory – cf. already grammaticalized connectives (e.g. thereafter, therefore or thereby) vs. not yet grammaticalized connectives (e.g. because of this or for this reason). We apply this general division on Czech and German connectives and conduct a contrastive study on the parallel data of the corpus InterCorp 10. Specifically, we focus on the group of Czech connectives in the form of prepositional phrases with an obligatory inherent reference that do not have any fully grammaticalized form in Czech (like kromě toho, lit. “except this”, ‘moreover’) and we search for their most frequent semantic counterparts in German. The results of our research demonstrate that the German counterparts of the selected connectives in Czech are mostly (in 72%) grammaticalized connectives containing a referential morpheme (e.g. außerdem, deswegen, stattdessen, dagegen, demgegenüber, daneben, infolgedessen).