One of the great desiderata in the field of Classics unfulfilled for centuries has been to trace the reception of the tremendously influential thinker and writer Plato throughout antiquity.The task is not only challenging because of the vast amount of literature (even if we focus on Greek) but has also been impossible to solve with digital approaches so far, especially because many voices of this reception neither explicitly refer to Plato nor cite his words verbatim.A precondition to further explore these instances was, therefore, to be able to detect them.In an iterative process of discussion, implementation, and evaluation, the project developed two web-based solutions for paraphrase search on the Ancient Greek corpus fit for different research questions (based on close and distant reading) together with different supportive resources and tools.The shaping of the theoretical framework, hand in hand with the work on the algorithms, did not only yield innovative digital methods for searching Ancient Greek texts but also sparked a reconsideration of the literary concepts of 'paraphrase ' and 'intertextuality'.
The reception of Plato’s work in ancient literature is characterized by a tremendous diversity and has not nearly been researched sufficiently. Our research project aims to further reveal his impact on Ancient Greek literature, particularly by finding paraphrases in an interdisciplinary approach combining expertise from humanities, computational linguistics, and computer science. The research project is ambitious in terms of methodology: it explores which algorithms for paraphrase identification are suited for smaller volumes of text beyond Big Data, and how to adjust their parameters to yield valuable results for literary phenomena which are less frequent and partly more elaborate than those in everyday language. A crucial precondition to reach this goal is a thorough analysis of different types of paraphrases in Ancient Greek literature that grants us an understanding of what we are searching for. On the other hand, we need a stable basis to evaluate our approaches to paraphrase identification by validating our results and to ensure comparability of different algorithms and their combinations. For both demands we could not resort to parallel corpora and similar resources that already exist for many modern languages. In our presentation we want to introduce our gold standard of over 200 intertextual references to Plato as well as the annotation tool we developed to characterize and categorize these references. We consider the annotated gold standard as a starting point for a new view on the concept of paraphrasing adjusted to the actualities of ancient literature as well as a touchstone for the development and the systematic evaluation of our algorithms for finding yet undiscovered paraphrases.
zur Konferenz Digital Humanities im deutschsprachigen Raum 2018 Entwicklungsstand im Projekt 'Digital Plato'
zur Konferenz Digital Humanities im deutschsprachigen Raum 2017 Paraphrasenerkennung im Projekt Digital Plato
To find receptions of Plato‘s work within the ancient Greek literature, automatic methods would be a useful assistance. Unfortunately, such methods are often knowledge-based and thus restricted to extensively annotated texts, which are not available to a sufficient extent for ancient Greek. In this article, we describe an approach that is based on the distributional hypotheses instead, to overcome the problem of missing annotations. This approach uses word2vec and the related Word Mover‘s Distance to determine phrases with similar meaning. Despite its experimental state, the method produces some meaningful results as shown in three examples.
Vor allem durch die häufige Nutzung in den sozialen Medien sind Tag Clouds heutzutage weit verbreitete Visualisierungen um den Inhalt textbasierter Daten zu veranschaulichen. Im Forschungsbereich der Informatik wurden zahlreiche Verfahren zur Berechnung von Tag Clouds entwickelt, unter anderem Wordle (Viégas et al. 2009). Abbildung 1 zeigt eine durch Wordle generierte Tag Cloud für die häufigsten Schlagworte aus den fünf Shakespeare Werken As You Like It, Macbeth, Othello, Richard III und Romeo and Juliet. Wie für Tag Clouds üblich, wird die Häufigkeit eines Wortes mit Schriftgröße kodiert. Leider tragen die weiteren visuellen Eigenschaften – Farbe, Position und Orientierung – der Darstellung keine Informationen. In unserer Posterpräsentation möchten wir TagPies vorstellen, ein im Rahmen des Digital Humanities Projektes eXChange entwickeltes Tag Cloud Layout. Während das Design von TagPies bereits vorgestellt wurde (Jänicke et al. 2015), sollen hier repräsentative Anwendungsszenarien illustriert werden. Ein TagPie kann als Hybrid aus Tag Cloud und Pie Chart gesehen werden, bei dem mehrere einzelne Tag Clouds miteinander zu einer Einheit verschmelzen. Im Gegensatz zu herkömmlichen Tag Clouds verwenden TagPies neben Schriftgröße auch Farbe und Position als Informationsträger und ermöglichen damit den Vergleich von Schlagworten verschiedener Kategorien. Abbildung 2 zeigt einen TagPie für die häufigsten Schlagworte aus den oben genannten Shakespeare Werken. Wie in den folgenden Beispielen bestimmt die Größe einer Textklasse im Vergleich zu den anderen die Größe des zugehörigen Kuchenstückes im TagPie. Weitere Beispiele in den Abbildungen 3 und 4 veranschaulichen die sprachunabhängige, vielseitige Anwendbarkeit von TagPies (Jähnicke 2015) – welche als Web-basierte Open Source JavaScript Bibliothek implementiert und damit projektunabhängig einsetzbar sind – für verschiedene geisteswissenschaftliche Fragestellungen.