At the heart of the »textklang« (Sound of Text) project is the development of a mixed-method approach to investigate the interrelation between written lyric poetry and its sonic realisation. It is an interdisciplinary collaboration between the German Literary Archive in Marbach and the University of Stuttgart that includes literary studies, digital humanities, computational linguistics, laboratory phonology and speech technology. The project’s corpus is centred on the poetry of Romanticism and is based on the holdings of the German Literature Archive. In our contribution, we illustrate a multi-perspective approach to one collection within the corpus entitled The Boy’s Magic Horn (Des Knaben Wunderhorn), edited in three volumes by A. v. Arnim and C. Brentano in 1806 and 1808 and including more than 700 poems. The Boy’s Magic Horn is considered one of the most influential poetry collections in German literature because of its vivid reception in both ‘low’ folkloristic cultures and ‘high’ culture, especially in musical settings (G. Mahler, J. Brahms, F. Silcher). We share some initial outcomes of the ongoing research process considering quantitative, textual, prosodic, and sonic aspects. As part of the project’s methodological and experimental toolbox, we present a speech synthesis model that has been trained on this sub-corpus and which results in a better realisation of poetic speech compared to synthesis models exclusively trained on prose data. Finally, we discuss the challenges this data pose to automatic processing tools.
The research project »text sound«: mixed-methods-analysis of lyric poetry in text and tonal sound (funded by the Federal Ministry for Education and Research, BMBF) aims to undertake a systematic and diachronic investigation of the relationship between literary texts, especially lyric poetry from the Romantic period, and their phonetic realisation in recitations or musical performances. Ideas of orality, sound and voice, which are particularly associated with poetry, are investigated empirically and also theorised in the line with modern approaches to the analysis of lyric poetry. Of particular importance is the experimental approach of speech synthesis, i.e. using computers to artificially produce a human sounding voice; this approach makes it possible to explore an ideal-typical realisation of the text and to test the aesthetic peculiarity of human realisations.
We present the steps taken towards an exploration platform for a multi-modal corpus of German lyric poetry from the Romantic era developed in the project "textklang". This interdisciplinary project develops a mixed-methods approach for systematic investigations of the relationship between written text (here lyric poetry) and its potential and actual sonic realisation (in recitations and musical performances). The multi-modal "textklang" platform will be designed to technically and analytically combine three modalities: the poetic text, the audio signal of a recorded recitation and, at a later stage, music scores of a musical setting of a poem. The methodological workflow will enable scholars to develop hypotheses about the relationship between textual form and sonic/prosodic realisation based on theoretical considerations, text interpretation and evidence from recorded recitations. The full workflow will support hypothesis testing either through systematic corpus analysis alone or with additional contrastive perception experiments. For the experimental track, researchers will be enabled to manipulate prosodic parameters in (re-)synthesised variants of the original recordings. The focus of this paper is on the design of the base corpus and on tools for systematic exploration - placing special emphasis on our response to challenges stemming from multi-modality and the methodologically diverse interdisciplinary setup.
Process metadata constitute a relevant part of the documentation of research processes and the creation and use of research data. As an addition to publications on the research question, they capture details needed for reusability and comparability and are thus important for a sustainable handling of research data. In the DH project SDC4Lit we want to capture process metadata when researchers work with literary material, conducting manual and automatic processing steps, which need to be treated with equal emphasis. We present a content-related mapping between two process metadata schemas from the area of Digital Humanities (GRAIN, RePlay-DH) and one from Computational Engineering (EngMeta) and find that there are no basic obstacles preventing the use of any of them for our purposes. Actually a basic difference rather exists between GRAIN on the one hand and RePlay-DH and EngMeta on the other, regarding the treatment of tools and actors in a workflow step.
Ein Grosteil der Forschung in den Digitalen Geisteswissenschaften und anderen textbasierten Disziplinen sieht sich mit dem Problem konfrontiert, dass die Weitergabe von Forschungsdaten durch das geltende Urheberrecht erschwert wird. Dies fuhrt dazu, dass Forschungsfragen so formuliert werden, dass diese nicht am eigentlichen wissenschaftlichen Interesse und Bedarf ausgerichtet sind, sondern auf der Verwendbarkeit der zugrundeliegenden Textressourcen. Insbesondere Annotationsprojekte arbeiten nahezu ausschlieslich auf gemeinfreiem Material, zeitgenossische Literatur bleibt weitestgehend unerforscht. Die Problematik wurde seitens des Gesetzgebers aufgegriffen: § 60d des Urheberrechtsgesetzes erlaubt seit 2018 die Nutzung urheberrechtlich geschutzter Werke zum Zwecke der Forschung mittels Text und Data Mining. Anderungen sind zudem zum Juni 2021 zu erwarten, wenn die sog. DSM-RL der Europaischen Union ins deutsche UrhG implementiert wird. Geschutzte Korpora durfen derzeit Bibliotheken zur Archivierung uberfuhrt werden. Unzureichend abgedeckt bleibt allerdings weiterhin die Frage, wie nach Projektende mit geschutzten Daten verfahren werden darf bzw. wie eine wunschenswerte Nachnutzbarkeit umgesetzt werden kann, ohne dabei Urheber zu benachteiligen. Im Projekt XSample wird auf Grundlage des § 60c UrhG gepruft, ob Auszuge geringen Umfangs von § 60d geschutzten Werken, Dritten zuganglich gemacht werden konnen. Dazu wird anhand von wissenschaftlichen Use Cases, die juristischen Rahmenbedingungen gepruft und ein Workflow-Konzept erarbeitet, wie bibliothekarische Einrichtungen die Nachnutzbarkeit von geschutzten Textkorpora unterstutzen konnen. Mittels eines Webinterfaces soll auf einfachem Wege ermoglicht werden, basierend auf inhaltlichen Kriterien, auszugsweise Zugang zu geschutzten Korpora zu erhalten, um beispielsweise damit zusammenhangende Veroffentlichungen zu validieren oder deren Eignung fur Anschlussforschung im Kontext einer neuen Forschungsfrage zu testen.
Die Digitalisierung bietet ein neues, effizientes Hilfsmittel, den wissenschaftlichen Fortschritt zu unterstutzen. In nahezu allen Bereichen lassen sich mit Hilfe moderner Informationssysteme Forschungsdaten digital archivieren und bei Bedarf leichter wiederverwenden. In der heutigen Zeit, in der das kollektive Wissen ein enormes Ausmas angenommen hat, ist die Systematisierung von Forschungsdaten zwingender denn je erforderlich. Die E-Science-Tage 2019, aus denen dieser Tagungsband hervorgegangen ist, haben neue Wege der Verarbeitung von Forschungsdaten aufgezeigt und durch den regen Austausch von Erfahrungen und Innovationen die digitale Wissenschaft weiter vorangetrieben.
Corpus query systems exist to address the multifarious information needs of any person interested in the content of annotated corpora. In this role they play an important part in making those resources usable for a wider audience. Over the past decades, several such query systems and languages have emerged, varying greatly in their expressiveness and technical details. This paper offers a broad overview of the history of corpora and corpus query tools. It focusses strongly on the query side and hints at exciting directions for future development.
In mobile and stationary applications, axial piston machines are often used as pumps or motors. The paper on hand deals with the pressurization behavior of swash plate axial piston units. The pressurization is influenced amongst others by the pressure on high and low pressure line pHP and pLP and the rotational speed ω, geometric properties, e.g. the dead vol-ume Vdead, the valve plate design and the fluid characteristics (density ρ, viscosity ν, bulk modulus E). In the following a theoretical and an experimental analysis of the pressurization in the piston chamber and the resulting compression work is presented.
At the Institute for Fluid Power Drives and Systems of RWTH Aachen University (ifas) the tribological behaviour of the piston–bushing contact in swash plate pumps has been investigated experimentally on a single piston test rig for several years. In the present publication simulation modelling is discussed with the help of the multi-body simulation tool FIRST, developed by the IST GmbH. Three different simulation models with different modeling depths are compared with respect to the effort of modelling, the required boundary conditions and the calculation duration. Finally, differences in the calculation results arec ritically compared on the basis of the loss characte-r istics of axial friction force and leakage. Furthermore, approaches for improving the simulation models are mentioned.
We describe the process metadata of GRAIN, a complex language data corpus, as a show case for application of metadata in the Digital Humanities. While the creation of language resources usually involves some automatic processing ranging from format conversion to labeling of structural features, data selection, inspection and interpretation are important manual steps, which tend to be neglected in the description of scientific workflows. GRAIN makes use of a format which (i) maps all workflow steps to flexible triples of \(\{input,operator,output\}\) and (ii) treats manual and automatic steps equally. Moreover, the process metadata has been semi-automatically generated and allows for a straightforward visualization describing the creation of the resource.
Large high-resolution displays (LHRDs) are entering into our daily life. Today, we already see them in installations where they display tailored applications, e.g. in exhibitions. However, while heavily studied under lab conditions, real-world applications for personal use, which utilize the extended screen space are rarely available. Thus, today's studies of LHRD are particularly designed to embrace the large screen space. In contrast, in this paper, we investigate a real-world application designed for researchers working on large text corpora to support them in deep text understanding. We conducted a study with 14 experts from the humanities and computational linguistics which solved a text analysis task using a standard desktop version on a 24 inch screen and an LHRD version on three 50 inch screens. Surprisingly, the smaller display condition outperformed the LHRD in terms of task completion time and error rate. While participants appreciated the overview provided by the large screen, qualitative feedback also revealed that the need for head movement and the scrolling mechanism decreased the usability of the LHRD condition.
Today, we see an ever growing number of tools supporting text annotation. Each of these tools is optimized for specific use-cases such as named entity recognition. However, we see large growing knowledge bases such as Wikipedia or the Google Knowledge Graph. In this paper, we introduce NLATool, a web application developed using a human-centered design process. The application combines supporting text annotation and enriching the text with additional information from a number of sources directly within the application. The tool assists users to efficiently recognize named entities, annotate text, and automatically provide users additional information while solving deep text understanding tasks.
We present GRAIN (German RAdio INterviews) as part of the SFB732 Silver Standard Collection. GRAIN contains German radio interviews and is annotated on multiple linguistic layers. The data has been processed with state-of-the-art tools for text and speech and therefore represents a resource for text-based linguistic research as well as speech science. While there is a gold standard part with manual annotations, the (much larger) silver standard part (which is growing as the radio station releases more interviews) relies completely on automatic annotations. We explicitly release different versions of annotations for the same layers (e.g. morpho-syntax) with the aim to combine and compare multiple layers in order to derive confidence estimations for the annotations. Therefore, parts of the data where the output of several tools match can be considered clear-cut cases, while mismatches hint at areas of interest which are potentially challenging or where rare phenomena can be found.
Present-day empirical research in computational or theoretical linguistics has at its disposal an enormous wealth in the form of richly annotated and diverse corpus resources. Especially the points of contact between modalities are areas of exciting new research. However, progress in those areas in particular suffers from poor coverage in terms of visualization or query systems. Many limitations for such tools stem from the non-uniform representations of very diverse resources and the lack of standards that address this problem from the perspective of processing or querying. In this paper we present our framework for modeling arbitrary multi-modal corpus resources in a unified form for processing tools. It serves as a middleware system and combines the expressiveness of general graph-based models with a rich metadata schema to preserve linguistic specificity. By separating data structures and their linguistic interpretations, it assists tools on top of it so that they can in turn allow their users to more efficiently exploit corpus resources.
Um Forschungsdaten auffindbar zu machen, müssen diese mit ausreichend Metadaten beschrieben werden. Damit die durch die Metadaten beschriebenen Forschungsdaten für andere Wissenschaftlerinnen und Wissenschaftler reproduzierbar sind, ist es notwendig, den Kontext ihrer Entstehung mit abzubilden. Gerade die Dokumentation dieses Entstehungsprozesses wird aber oft durch mangelnde Zeit im Forschungsalltag vernachlässigt. Auch fehlt es hier noch an niederschwelliger Unterstützung im Arbeitsprozess. Einige Methoden sind gerade dabei sich zu etablieren oder befinden sich in der Entwicklung. Im Folgenden werden Softwareanwendungen, die die Dokumentation erleichtern sollen, vorgestellt und mit der aktuell im Projekt RePlay-DH entwickelten Lösung verglichen. Der Ansatz der Virtuellen Forschungsumgebung setzt auf die Zusammenarbeit über eine gemeinsame Plattform. Das Elektronische Laborbuch unterstützt die Dokumentation im Labor. Das Workflow-Management definiert, im Gegensatz zum Workflow-Tracking, einen Workflow vor der Ausführung der einzelnen Arbeitsschritte. Dabei steht die prozessbegleitende Dokumentation im Mittelpunkt. Der Lösungsansatz, der im Projekt RePlay-DH verfolgt wird, besteht in der unterstützenden Dokumentation des Forschungsprozesses mit Metadaten durch ein vereinfachtes Workflow-Tracking. Die Integration in bestehende Arbeitsabläufe von Wissenschaftlerinnen und Wissenschaftlern und die einfache Bedienbarkeit stehen dabei im Vordergrund.
Jonas Kuhn合作论文数Institute for Natural Language Processing, University of Stuttgart16
Dieter Merkl合作论文数institut f??r softwaretechnik und interaktive systeme;electronic commerce group2