Linguistic research frequently requires the categorization of language phenomena in corpus data (annotation). Since those may occur plentifully, a partial or full automation of the annotation process appears attractive. The filtering and recombination of existing annotation layers seems to further provide an elegant solution to the deduction of higher-level annotations. In this contribution, we show at the example of German split particle verbs that this approach results in a number of linguistic, technological, and epistemological challenges related to the precise definition of the various models employed and their interfaces. We argue that the manual annotation of corpus data is not merely a preprocessing task, but is itself an epistemological process central to the development of linguistic theory. We discuss why machine-based language processing can neither mimic nor replace this process; why it can generally not reach a level of precision that would be suitable for linguistic research without further integration, adaptation, and manual correction; and how its blind application systematically skews results in crucial areas of research. We close with the suggestion of several best practice approaches which help to prevent and resolve incompatibilities and delays arising from common problems of corpus-based language modeling.
Abstract The article presents the RUEG corpora. We begin by describing the basic principles and research questions, that influenced and shaped the corpora and introduce the method of data generation. We proceed by describing the fundamental components of the building process and the overall outcome. Said outcome provides interfaces for new researchers, who wish to conduct their own research using the RUEG corpora. Options for such research are discussed by showing that it can succeed within the range of re-using exisiting annotations and metadata up to building entirely new, but comparable corpora.
AbstractThis paper argues for incorporating corpus data into the teaching of historical linguistics. While deeply annotated historical corpora are becoming available and corpus data is already widely used to answer various research questions, corpora are as yet rarely used in teaching. We believe they are ideally suited to make the variation in historical data transparent and help students to explore contexts and parameters. In our first study, we show how the KaJuK corpus and its more elaborated version, the GiesKaNe corpus, can be exploited to study adverbial sentences. Using the RIDGES corpus, the second study deals with phrasal and lexical development. Both studies focus on explaining the method and its extension to other corpora and research questions.
Falko ist ein frei zugängliches Lernerkorpus des schriftsprachlichen Deutschen als Fremdsprache und umfasst nach jahrelanger Erschließung neuer Textressourcen und der Anreicherung mit diversen Annotationsebenen eine Reihe einzelner Korpora, die teilweise sehr komplex strukturiert sind. Im vorliegenden Beitrag stellen wir die komplexeste Datenressource aus der Reihe dieser Korpora vor – das Falko-Essay-Korpus, welches aktuell in einer neuen Version (3.0) erscheint und interessierten Forscherinnen und Forschern frei zur Verfügung steht. Falko is a freely available learner corpus with written learner texts of German as a foreign language. After years of data acquisition and data processing on multiple layers of annotation, Falko consists of several different corpora with partly avery complex architecture. In this article, we want to introduce the most complex data resource among these corpora – the Falko essay corpus. It is currently released in a new version (3.0) and can be used openly and freely by researchers that are interested in this resource.
Abstract Scientific herbal texts in vernacular German first emerged in the 15th century, and their diachronic analysis presents an opportunity to trace linguistic changes across centuries. This study deals with the linguistic strategies that are used to refer to the concept of menstruation, which we study in the RIDGES corpus, representing herbal texts from the late fifteenth century to the early twentieth century. This exploratory study focuses on terminology, i.e. expressions referring to concepts relevant to scientific communication, and investigates the morphological, syntactic and semiotic characteristics of references to menstruation. Our study reveals two tendencies: a decrease of formal variation as well as a decrease of semiotic variation. Both tendencies point towards the increase of terminologisation in the early and late modern period (1483–1914).
Anke Lüdelinga, Artemis Alexiadoub, Aria Adlic, Karin Donhausera, Malte Dreyera, Markus Egga, Anna Helene Feulnera, Natalia Gagarinad, Wolfgang Hocka, Stefanie Jannedyd, Frank Kammerzella, Pia Knoeferlea, Thomas Krausea, Manfred Krifkab, Silvia Kutschera, Beate Lütkea, Thomas McFaddend, Roland Meyera, Christine Mooshammera, Stefan Müllera, Katja Maquatea, Muriel Nordea, Uli Sauerlandd, Stephanie Soltd, Luka Szucsicha, Elisabeth Verhoevena, Richard Waltereita, Anne Wolfsgrubera & Lars Erik Zeigea aHumboldt-Universität zu Berlin bHumboldt-Universität zu Berlin / LeibnizZAS cUniversität zu Köln dLeibniz-ZAS
Die Sprache von Lerner/-innen einer Fremdsprache unterscheidet sich auf allen linguistischen Ebenen von der Sprache von Muttersprachler/-innen. Seit einigen Jahrzehnten werden Lernerkorpora gebaut, um Lernersprache quantitativ und qualitativ zu analysieren. Hier argumentieren wir anhand von drei Fallbeispielen (zu Modifikation, Koselektion und rhetorischen Strukturen) für eine linguistisch informierte, tiefe Phänomenmodellierung und Annotation sowie für eine auf das jeweilige Phänomen passende formale und quantitative Modellierung. Dabei diskutieren wir die Abwägung von tiefer, mehrschichtiger Analyse einerseits und notwendigen Datenmengen für bestimmte quantitative Verfahren andererseits und zeigen, dass mittelgroße Korpora (wie die meisten Lernerkorpora) interessante Erkenntnisse ermöglichen, die große, flacher annotierte Korpora so nicht erlauben würden.
This paper presents RST-Tace, a tool for automatic comparison and evaluation of RST trees. RST-Tace serves as an implementation of Iruskieta’s comparison method, which allows trees to be compared and evaluated without the influence of decisions at lower levels in a tree in terms of four factors: constituent, attachment point, nuclearity as well as relation. RST-Tace can be used regardless of the language or the size of rhetorical trees. This tool aims to measure the agreement between two annotators. The result is reflected by F-measure and inter-annotator agreement. Both the comparison table and the result of the evaluation can be obtained automatically.
zur Konferenz Digital Humanities im deutschsprachigen Raum 2018 Suche und Visualisierung von Annotationen historischer Korpora mit ANNIS. Kritik der korpuslinguistischen Analysemethoden in einem erweiterten Nutzungskontext
This paper describes a fundamental re-design and extension of the existing general multi-layer corpus search tool ANNIS, which simplifies its re-use in other tools. This embeddable corpus search library is called graphANNIS and uses annotation graphs as its internal data model. It has a modular design, where each graph component can be implemented by a so-called graph storage and allows efficient reachability queries on each graph component. We show that using different implementations for different types of graphs is much more efficient than relying on a single strategy. Our approach unites the interoperable data model of a directed graph with adaptable and efficient implementations. We argue that graphANNIS can be a valuable building block for applications that need to embed some kind of search functionality on linguistically annotated corpora. Examples are annotation editors that need a search component to support agile corpus creation. The adaptability of graphANNIS, and its ability to support new kinds of annotation structures efficiently, could make such a re-use easier to achieve.
zur Konferenz Digital Humanities im deutschsprachigen Raum 2017 Zwei grundlegende Fragen der digitalen Nachhaltigkeit: Wie können wir die heterogenen Forschungsfragen und die Community bei der Verfügbarmachung von Forschungsdaten miteinbeziehen?
The present study analyzes morphological productivity for complex verbs in second language acquisition by analyzing a corpus of German as a Foreign Language (GFL). It shows that advanced learners of GFL use prefix and particle verbs relatively frequently and productively but less so than native speakers do and discusses these findings in the light of different linguistic models and acquisition theories. It argues that corpus data must be evaluated against good models and that it is necessary to make the categorization decisions available as annotations.
Dieser Beitrag beschaftigt sich mit zwei eng miteinander verbundenen Fragen: Wie konnen die syntaktischen Strukturen in Chat-Texten beschrieben werden? Welche syntaktischen Eigenschaften haben deutsche Chat-Texte? Chats (und hier insbesondere sogenannte ‚Plauderchats‘) weichen in vielerlei Hinsicht von einer schriftlichen ‚Standardsprache‘ ab. Uns interessiert in diesem Beitrag vor allem die Syntax von Auserungen aus Plauderchats. In Abschnitt 1 werden wir zunachst kurz auf einige Grundannahmen von syntaktischen Beschreibungen eingehen und erlautern, warum diese fur die Analyse von Chatdaten nicht immer geeignet sind. Dann, in Abschnitt 2, werden wir die Chatdaten aus dem NoSta-D-Korpus und ihre Vorverarbeitung vorstellen, bevor wir in Abschnitt 3 auf einige syntaktische Eigenschaften der Daten genauer eingehen. Empirikom
AbstractIn this article, we explore the disfluencies of advanced learners and native speakers of German in spontaneous speech. We focus on the frequency, form, and place of silent and filled pauses as well as self-repairs. Frequency significantly differs for silent pauses only. As to form, the distribution for both filled pauses and repair types significantly differs between the groups, while the proportion of within-repair hesitations (‘interregna’) is similar. For the neighbouring tokens of filled pauses, learners adhere to the pattern of their native language English, which is significantly different from the pattern we find for native German. Our results indicate that for some aspects of disfluencies, it seems that learners can adapt to a native-like pattern, while others are imported from the L1. Still others are significantly different from both the target and the native pattern. We present different possible explanations for all these cases.
This paper introduces a multi-layer corpus architecture with multiple tokenizations using the open source historical, diachronic corpus of German called Register in Diachronic German Science. The corpus contains herbal texts printed between the fifteenth and nineteenth centuries and is concerned with the development of a German scientific register, independent of Latin. We will discuss difficulties of transcribing, normalizing and annotating historical texts and will thereby argue for the advantages of multiple layers and multiple tokenizations. A virtually infinite number of annotations can be added to the corpus, without the need for deciding between or discarding interpretations. Thus, this flexible architecture enables multiple normalizations and types of annotation and is open to a wide range of research questions in the humanities. We provide case studies concerning the exploitation of our different normalizations as well as structural, register-specific and linguistic annotations. The corpus architecture allows for its reuse as a resource for corpus-based research approaches.
We present graphANNIS, a fast implementation of the established query language AQL for dealing with deeply annotated linguistic corpora.AQL builds on a graphbased abstraction for modeling and exchanging linguistic data, yet all its current implementations use relational databases as storage layer.In contrast, graphANNIS directly implements the ANNIS graph data model in main memory.We show that the vast majority of the AQL functionality can be mapped to the basic operation of finding paths in a graph and present efficient implementations and index structures for this and all other required operations.We compare the performance of graphANNIS with that of the standard SQL-based implementation of AQL, using a workload of more than 3000 real-life queries on a set of 17 open corpora each with a size up to 3 Million tokens, whose annotations range from simple and linear part-of-speech tagging to deeply nested discourse structures.For the entire workload, graphANNIS is more than 40 times faster, and slower in less than 3% of the queries.graphANNIS as well as the workload and corpora used for evaluation are freely available at GitHub and the Zenodo Open Access archive.
This paper describes the construction of deeply annotated spoken dialogue corpora. To ensure a maximum of flexibility — in the degree of normalization, the types and formats of annotations, the possibilities for modifying and extending the corpus, or the use for research questions not originally anticipated — we propose a flexible multi-layer standoff architecture. We also take a closer look at the interoperability of tools and formats compatible with such an architecture. Free access to the corpus data through corpus queries, visualizations, and downloads — including documentation, metadata, and the original recordings — enables transparency, verifiability, and reproducibility of every step of interpretation throughout corpus construction and of any research findings obtained from this data.
Der Workshop setzt sich das Ziel, die Möglichkeiten, Aufgaben und Herausforderungen bei der Wiederverwendung von historischen Korpora zu identifizieren und zu diskutieren. Insbesondere sollen dabei deren Architektur, Dokumentation, Veröffentlichung und Speicherung betrachtet werden. So wollen wir versuchen, Methoden und Strategien für das interdisziplinäre Forschungsparadigma der Digital Humanities zu entwickeln und diese in den Fragestellungen der Konferenz der DHd 2016 zu verorten. Der Fokus wird auf die spezifischen und fächerübergreifenden Anforderungen historischer Texte in Bezug auf deren Aufbereitung und Speicherung in Repositorien zum Zweck der Wiederverwendung gelegt. Damit richtet sich der Workshop gleichermaßen an Korpuserstellende, Entwickelnde und an Betreiber_innen von Repositorien und deren Nutzer_innen.
Wir stellen eine neue OCR-Methode vor, mit der es erstmals möglich ist, gedruckte Texte von der Inkunabelzeit (1450-1500) bis heute mit hoher Genauigkeit (größer 95% Zeichenerkennungsrate) in elektronischen Text zu verwandeln. Bisherige Versuche der Konversion von Inkunabeln lieferten keinen brauchbaren Text (Rydberg-Cox, 2009). Die Methode beruht auf rekurrenten neuronalen Netzen mit langem Kurzzeitgedächtnis (Hochreiter und Schmidhuber, 1997), deren Anwendbarkeit auf OCR erstmals von Breuel u. a. (2013) beschrieben wurde. Durch den Vergleich von gedruckten Textzeilen mit einer diplomatischen Transkription (”ground truth“), die das Netzwerk als Input erhält, werden die internen Parameter in einem automatischen Verfahren so eingestellt, dass nach einer hinreichenden Anzahl von Lernschritten eine Erkennung neuer Textzeilen mit hoher Genauigkeit möglich wird. Das Verfahren benötigt daher Trainingsdaten in Form einer diplomatischen Transkription, wie sie im diachronen RIDGES-Korpus zur Verfügung stehen. Ein trainiertes Modell kann dann auf Drucke mit gleicher oder ähnlicher Typographie angewendet werden. Die Resultate sind sprachunabhängig und setzen kein Lexikon voraus. Damit eröffnet sich die Möglichkeit, das gedruckte Erbe auch bei sehr frühen Drucken maschinell in digitalen Text zu transformieren. Mögliche Anwendungen ergeben sich für die Suche (mit Trefferanzeige im Bild anstelle im OCR-Text) sowie, gegebenenfalls nach automatischer und manueller Nachkorrektur, für den automatisch unterstützten Aufbau von Korpora und Lexika.