Dieser Beitrag beschaftigt sich mit zwei eng miteinander verbundenen Fragen: Wie konnen die syntaktischen Strukturen in Chat-Texten beschrieben werden? Welche syntaktischen Eigenschaften haben deutsche Chat-Texte? Chats (und hier insbesondere sogenannte ‚Plauderchats‘) weichen in vielerlei Hinsicht von einer schriftlichen ‚Standardsprache‘ ab. Uns interessiert in diesem Beitrag vor allem die Syntax von Auserungen aus Plauderchats. In Abschnitt 1 werden wir zunachst kurz auf einige Grundannahmen von syntaktischen Beschreibungen eingehen und erlautern, warum diese fur die Analyse von Chatdaten nicht immer geeignet sind. Dann, in Abschnitt 2, werden wir die Chatdaten aus dem NoSta-D-Korpus und ihre Vorverarbeitung vorstellen, bevor wir in Abschnitt 3 auf einige syntaktische Eigenschaften der Daten genauer eingehen. Empirikom
Learner corpora consist of texts produced by non-native speakers. In addition to these texts, some learner corpora also contain error annotations, which can reveal common errors made by language learners, and provide training material for automatic error correction. We present a novel type of error-annotated learner corpus containing sequences of revised essay drafts written by non-native speakers of English. Sentences in these drafts are annotated with comments by language tutors, and are aligned to sentences in subsequent drafts. We describe the compilation process of our corpus, present its encoding in TEI XML, and report agreement levels on the error annotations. Further, we demonstrate the potential of the corpus to facilitate research on textual revision in L2 writing, by conducting a case study on verb tenses using ANNIS, a corpus search and visualization platform.
We describe the annotation of a new dataset for German Named Entity Recognition (NER). The need for this dataset is motivated by licensing issues and consistency issues of existing datasets. We describe our approach to creating annotation guidelines based on linguistic and semantic considerations, and how we iteratively refined and tested them in the early stages of annotation in order to arrive at the largest publicly available dataset for German NER, consisting of over 31,000 manually annotated sentences (over 591,000 tokens) from German Wikipedia and German online news. We provide a number of statistics on the dataset, which indicate its high quality, and discuss legal aspects of distributing the data as a compilation of citations. The data is released under the permissive CC-BY license, and will be fully available for download in September 2014 after it has been used for the GermEval 2014 shared task on NER. We further provide the full annotation guidelines and links to the annotation tool used for the creation of this resource.
Fur viele aktuelle Fragestellungen der Zweitund Fremdspracherwerbsforschung („L2Erwerbsforschung“) sind Lernerkorpora unverzichtbar geworden. Sie stellen Texte von L2Lernern1 zur Verfugung, oftmals erganzt durch vergleichbare Texte von Muttersprachlern der Zielsprache. Beschrankten sich Analysen der Lernerkorpusforschung in den ersten Jahren hauptsachlich auf einzelne Wortformen (vgl. Granger, 1998), hat sich das Forschungsinteresse bestandig hin zu komplexeren grammatischen Kategorien entwickelt. Dazu zahlen u.A. die Untersuchung tiefer syntaktischer Analysen (Dickinson und Ragheb, 2009; Hirschmann et al., 2013, u.a.) oder die Strategien der Markierung von Koharenzrelationen (z.B. Breckle und Zinsmeister, 2012). Derartige Analysen bauen dabei nur selten auf der Textoberflache selbst auf, sondern setzen i.d.R. die Annotation von Wortarten fur jedes Texttoken voraus und ggfs. weitere, darauf aufbauende Annotationsebenen. Annotationen dienen generell immer der Suche nach Klassen in den Daten, die anhand der Oberflachenformen allein nicht leicht zuganglich waren (im Kontext von Lernerkorpora vgl. Diaz-Negrillo et al., 2010). Ist man z.B. an einer Analyse von Possessivpronomen interessiert, wurde man bei einer Korpussuche, die nur Zugriff auf die Wortformen selbst hat, bei der ambigen Form meinen neben Beispielen fur das Possessivpronomen (1) auch alle Belege fur die gleichlautende Verbform (2) finden. Das Suchergebnis ware also sehr ‘unsauber’, da die Wortform selbst keinen Aufschluss uber ihre Interpretation gibt. Eine Annotation mit Wortarten wurde die beiden Lesarten disambiguieren und damit die Ruckgabe der Suchanfrage praziser machen. Die Ruckgabe wurde weniger ungewunschte Lesarten enthalten, die man andernfalls bei der Ergebnissichtung manuell ausschliesen musste. Kurz gesagt, eine Suchanfrage auf Wortarten-annotierten Daten ist fur den Nutzer effizienter als eine Suche auf reinen Wortformen.
Error annotation is a key feature of modern learner corpora. Error identification is always based on some kind of reconstructed learner utterance (target hypothesis). Since a single target hypothesis can only cover a certain amount of linguistic information while ignoring other aspects, the need for multiple target hypotheses becomes apparent. Using the German learner corpus Falko as an example, we therefore argue for a flexible multi-layer stand-off corpus architecture where competing target hypotheses can be coded in parallel. Surface differences between the learner text and the target hypotheses can then be exploited for automatic error annotation.
This paper shows how the automatic syntactic analysis of a corpus of advanced learners of German as a foreign language helps in understanding the acquisition of modification. In former corpus research modification has been studied only by comparing the distributions of single words (or groups of words) in learner and native speaker data. We argue that in order to study modification as a syntactic category it is necessary to work with syntactically analyzed corpora. In this vein, we sketch out our approach to parsing learner language and conduct two contrastive interlanguage studies on modification in the syntactically annotated corpus, showing that not only lexical modifiers can be underused (as shown in many other studies), but that modification as a whole category (including multi-word modifiers such as prepositional phrases, and clausal modifiers such as relative clauses) is underused in our learner corpus data.
Until recently, most research in computational linguistics has been done on newspaper texts. Nowadays, the focus has been extended to other types of language data. This means that many linguistic descriptions and automatic tools need to be adapted or extended to non-newspaper language. The non-standard varieties corpus of German (NoSta-D) will provide a first gold standard for evaluation and training data of dependency analysis, named entity recognition and coreference resolution for out-of-domain text types.
Parsing learner data poses a great challenge for standard tools, since non-canonical and unusual structures may lead to wrong interpretations on the part of the taggers and parsers. It is well known that providing a statistical parser with perfect part-of-speech (POS) tags is of great benefit for parsing accuracy, and that parsing results can decrease considerably when the parser has to predict its own POS tags. Therefore one might expect that even small improvements in POS accuracy have a positive effect on parsing performance. In this paper we test this assumption and assess the impact of POS tag accuracy on constituency parsing for German learner language. We compare different strategies to manual correction of the learner text and specific POS tags, and we measure the time requirements for each strategy. We show that tagging a canonical equivalent of the non-canonical learner text substantially improves POS tag accuracy. Correcting selected POS tags can only lead to parsing results comparable to a setting where all POS tags are corrected, while reducing annotation time substantially. However, the manual corrections of the POS tags do not result in a statistically significant improvement for parsing, giving evidence for the high quality of the automatically predicted parts-of-speech for the corrected learner data.
Das Wissenschaftliche Netzwerk „Kobalt-DaF“ : Korpusbasierte Analyse von Lernertexten fur Deutsch als Fremdsprache
This talk is concerned with using syntactic annotation of learner language and the corresponding target hypothesis to find structural acquisition difficulties in German as a foreign language. Using learner data for the study of acquisition patterns is based on the idea that learners do not produce random output but rather possess a consistent internal grammar (interlanguage; cf. [1] and many others). Analysing learner data is thus an indirect way of assessing the interlanguage of language learners. There are two main ways of looking at learner data, error analysis and contrastive interlanguage analysis [2, 3]. A careful analysis of errors makes it possible to understand learners’ hypotheses about a given grammatical phenomenon. Contrastive interlanguage analysis is not concentrated on errors but compares categories (of any kind) of learner language with the same categories in native speaker language. Learners’ underuse of a category (i.e. a significantly lower frequency in learner language than in native speaker language) can be seen as evidence for the perceived difficulty of that category (either because learners fail to acquire it, or because they deliberately avoid it). While some learner corpora are annotated (manually or automatically) with part-of-speech or lemma information [4], or even error types, there are as yet only very few attempts to annotate them syntactically (some exceptions are [5] or [6]. Parsing learner data is very difficult because of the learner errors but would be very helpful for the analysis of errors and overuse/underuse of syntactic structures and categories. In our paper we therefore discuss how the comparison of parsed learner data and the corresponding target hypotheses helps in understanding syntactic properties of learner language. We use the Falko corpus which contains essays of advanced learners of German as a foreign language and control essays by German native speakers [7]; the corpus is freely available1. Since it is very difficult to decide what an error is and often there can be different hypotheses about the ‘correct’ structure the learner utterance