In the humanities, text comparison is essential for scholarly editing the traditions of a text with its different witnesses under study. For the computer-aided creation of a critical apparatus, there have been established approaches and tools for many years. However, they mainly focus on subsentences or sentences. An efficient and easy-to-use automatic comparison of entire chapters or books still represents a research desideratum. The Locate, Explore, Retrace and Apprehend complex text variants (LERA) working environment presented here solves this issue. It is based on a two-stage collation approach: an efficient, fully automatic alignment of text segments, which can be paragraphs, subparagraphs or sentences, with interactive post-processing options, followed by the detailed comparison at segment level. Because aligning text segments, such as paragraphs, for more than two text witnesses is algorithmically challenging, we discuss the heuristics we developed in more detail. LERA combines the entire process of document management, tokenization/segmentation, normalization, alignment, and visualization with interactive control options and exploratory tools. It has already been and is being successfully applied in several Digital Humanities projects of different languages, e.g. for Arabic, French, Hebrew as well as German and English texts.
One of the great desiderata in the field of Classics unfulfilled for centuries has been to trace the reception of the tremendously influential thinker and writer Plato throughout antiquity.The task is not only challenging because of the vast amount of literature (even if we focus on Greek) but has also been impossible to solve with digital approaches so far, especially because many voices of this reception neither explicitly refer to Plato nor cite his words verbatim.A precondition to further explore these instances was, therefore, to be able to detect them.In an iterative process of discussion, implementation, and evaluation, the project developed two web-based solutions for paraphrase search on the Ancient Greek corpus fit for different research questions (based on close and distant reading) together with different supportive resources and tools.The shaping of the theoretical framework, hand in hand with the work on the algorithms, did not only yield innovative digital methods for searching Ancient Greek texts but also sparked a reconsideration of the literary concepts of 'paraphrase ' and 'intertextuality'.
This article presents an approach to the collation problem of text witnesses which is correct for all distance functions and is applicable to the alignment of any types of tokens, be it letters, words, sentences, or paragraphs. In contrast to the existing approaches, which all are heuristics in nature, we specify a formal model in the form of an integer linear program (ILP) that formally describes the optimization problem. For a special case that usually occurs in practice, we then transfer the ILP into an efficient algorithm that optimally solves the alignment problem. In the second part of the article, we apply our approach to the alignment of paragraphs of text witnesses—an alignment problem that has received little attention in the literature so far—and compare the results of our approach, which we have called TSaligner, with those of CollateX.
Abstract In this paper, A shorter version of the paper appeared in German in the final report of the Digital Plato project which was funded by the Volkswagen Foundation from 2016 to 2019. [35], [28]. we present a method for paraphrase extraction in Ancient Greek that can be applied to huge text corpora in interactive humanities applications. Since lexical databases and POS tagging are either unavailable or do not achieve sufficient accuracy for ancient languages, our approach is based on pure word embeddings and the word mover’s distance (WMD) [20]. We show how to adapt the WMD approach to paraphrase searching such that the expensive WMD computation has to be computed for a small fraction of the text segments contained in the corpus, only. Formally, the time complexity will be reduced from O ( N · K 3 · log K )\mathcal{O}(N\cdot {K^{3}}\cdot \log K) to O ( N + K 3 · log K )\mathcal{O}(N+{K^{3}}\cdot \log K), compared to the brute-force approach which computes the WMD between each text segment of the corpus and the search query. N is the length of the corpus and K the size of its vocabulary. The method, which searches not only for paraphrases of the same length as the search query but also for paraphrases of varying lengths, was evaluated on the Thesaurus Linguae Graecae® (TLG®) [25]. The TLG consists of about 75 · 10 6 75\cdot {10^{6}} Greek words. We searched the whole TLG for paraphrases for given passages of Plato. The experimental results show that our method and the brute-force approach, with only very few exceptions, propose the same text passages in the TLG as possible paraphrases. The computation times of our method are in a range that allows its application in interactive systems and let the humanities scholars work productively and smoothly.
zur Konferenz Digital Humanities im deutschsprachigen Raum 2020 Keter Shem #ov Prozessualisierung eines Editionsprojekts mit 100 Textzeugen
Word-clouds are a useful tool for providing overviews over texts, visualising relevant words. Multiple word-clouds can also be used to visualise changes over time in a text. This requires that the words in the individual word-clouds have stable positions, as otherwise it is very difficult so see what changed between two consecutive word-clouds. Existing approaches have used coordinated positioning algorithms, which do not allow for their use in an online, dynamic context. In this paper we present a fast word-cloud algorithm that uses word orthogonality to determine which words can share the same space in the word-clouds combined with a simple, but fast spiral-based layout algorithm. The evaluation shows that the algorithm achieves its goal of creating series of word-clouds fast enough to enable use in an online, dynamic context.
The reception of Plato’s work in ancient literature is characterized by a tremendous diversity and has not nearly been researched sufficiently. Our research project aims to further reveal his impact on Ancient Greek literature, particularly by finding paraphrases in an interdisciplinary approach combining expertise from humanities, computational linguistics, and computer science. The research project is ambitious in terms of methodology: it explores which algorithms for paraphrase identification are suited for smaller volumes of text beyond Big Data, and how to adjust their parameters to yield valuable results for literary phenomena which are less frequent and partly more elaborate than those in everyday language. A crucial precondition to reach this goal is a thorough analysis of different types of paraphrases in Ancient Greek literature that grants us an understanding of what we are searching for. On the other hand, we need a stable basis to evaluate our approaches to paraphrase identification by validating our results and to ensure comparability of different algorithms and their combinations. For both demands we could not resort to parallel corpora and similar resources that already exist for many modern languages. In our presentation we want to introduce our gold standard of over 200 intertextual references to Plato as well as the annotation tool we developed to characterize and categorize these references. We consider the annotated gold standard as a starting point for a new view on the concept of paraphrasing adjusted to the actualities of ancient literature as well as a touchstone for the development and the systematic evaluation of our algorithms for finding yet undiscovered paraphrases.
zur Konferenz Digital Humanities im deutschsprachigen Raum 2018 Entwicklungsstand im Projekt 'Digital Plato'
zur Konferenz Digital Humanities im deutschsprachigen Raum 2017 Paraphrasenerkennung im Projekt Digital Plato
To find receptions of Plato‘s work within the ancient Greek literature, automatic methods would be a useful assistance. Unfortunately, such methods are often knowledge-based and thus restricted to extensively annotated texts, which are not available to a sufficient extent for ancient Greek. In this article, we describe an approach that is based on the distributional hypotheses instead, to overcome the problem of missing annotations. This approach uses word2vec and the related Word Mover‘s Distance to determine phrases with similar meaning. Despite its experimental state, the method produces some meaningful results as shown in three examples.
Der Vergleich von Textfassungen ist eine zentrale Aufgabe der Editionsphilologie und Grundlage für die Erstellung von Variantenapparaten in kritischen Editionen. Zudem lässt sich über den Textvergleich die inhaltliche Evolution von Texten nachvollziehen. Je umfänglicher oder komplexer die zu edierenden Werke sind, desto schwieriger ist es für die Editoren und die Nutzer von Editionen den Überblick über die Textfassungen zu behalten. Mit dem hier vorgestellten, in die Arbeitsumgebung LERA (Bremer et al. 2015) integrierten Ansatz, wird diesem Problem mit einer Kombination von graphischen Werkzeugen begegnet, deren Zusammenwirken den Überblick über große Textmengen und ihre inhaltliche Auswertung erleichtert und zudem die Möglichkeit zum ergbnisoffenen Erkunden der Textvarianten bietet. LERA ist eine interaktive, webbasierte Arbeitsumgebung zur Untersuchung mehrerer Fassungen eines Textes. Sie wurde im Rahmen des vom Bundesministeriums für Bildung und Forschung geförderten Projekts: Semiautomatische Differenzanalyse von komplexen Textvarianten ( SaDA ) entwickelt (Medek et al. 2015). Als Fallbeispiel für die Entwicklung von LERA wurde ein Buch aus der 'Histoire philosophique et politique des établissements et du commerce des Européens dans les deux Indes' von Thomas-Guillaume Raynal, einem der gesamteuropäisch einflussreichsten Erfolgsund Skandalbücher des 18. Jahrhunderts, gewählt. Das Werk ist eine kritische Auseinandersetzung mit der europäischen Kolonialexpansion in den ,beiden‘ Indien und liegt in vier stark überarbeiteten Fassungen aus den Jahren 1770, 1774, 1780 und 1820 vor (Schütz / Pöckelmann 2014). Für die Kollationierung verschiedener Textfassung gibt es bereits eine Reihe digitaler Werkzeuge. Drei der bekanntesten, CollateX , Juxta und TUSTEP , liefern für viele Fragestellungen gute Ergebnisse, unterscheiden sich in ihren Anwendungsszenarien aber erkennbar von LERA. Während LERA einen zweistufigen Ansatz verfolgt, bei dem vor dem detaillierten Textvergleich zunächst eine Alignierung größerer Textpassagen berechnet wird, was denEinsatz komplexer Signaturwerte für den Vergleich notwendig macht, liegt der Fokus von CollateX auf der Alignierung (normalisierter) Token, für den auf einfache Zeichenketten als Signaturwerte zurückgegriffen wird. Juxta zeigt mit seinen hilfreichen Visualisierungen die Unterschiede zu einer ausgewählten Leithandschrift auf, während LERA für den direkten Vergleich mehrerer Textfassungen konzipiert wurde und stets die Textänderungen zwischen allen Fassungen darstellt. Der wohl vielseitigste digitale Werkzeugkasten aus dem Bereich der Geisteswissenschaften, TUSTEP, bietet ebenfalls Möglichkeiten zum Vergleich verschiedener Textfassungen. Allerdings erfordert der Umgang mit dem textbasierten Interface eine längere Einarbeitungszeit, die sich zumindest für einige Anwender durch die neu entwickelte XML-basierte Variante TXSTEP verringert. LERA wurde hingegen von Beginn an durch interaktive graphische Elemente als intuitiv bedienbare Arbeitsumgebung entwickelt. Veränderte Vergleichsparameter sollen dabei durch schnelle Neuberechnung eine direkte Präsentation des Ergebnisses ermöglichen und so zum Experimentieren einladen.
AbstractThis paper presents a digital working environment for exploratory analysis of the changes to the text of Raynal’s Histoire philosophique et politique des établissements et du commerce des Europeéns dans les deux Indes, a French text from 1770 about the negative influence of European civilization during the colonization of the East and West Indies which was significantly revised in 1774, 1780 and post-mortem in 1820. After a brief summary of the printing history, the features of the environment are presented. Starting with a general map of the corpus by means of a navigation bar and word clouds, the online platform provides users with the opportunity to analyse the text genesis in an exploratory manner. The general map leads the users to text passages of interest with respect to the text genesis. Synoptic representations together with colour markings adaptable to the respective philological problem make the revision history easily comprehensible and the text differences easily identifiable. The approach is, in principle, transferable to scholarly editions of other texts and texts in other languages.
Eine der zentralen Aufgaben bei editionsphilologischen Vorhaben ist die Feststellung der Abhängigkeiten von zu unterschiedlichen Zeitpunkten entstandenen Textzeugnissen. Für die Edition handschriftlicher Textzeugnisse beispielsweise bezieht sich dies auf die editionsphilologisch zentrale Frage des Zusammenhangs unterschiedlicher Handschriften und ihrer Filiation untereinander und bei Texten, die starke Überarbeitungsprozesse erfahren haben, auf die Frage der Textgenese. Mit leistungsstarken Werkzeugen zum Textvergleich kann sich der Editor viel monotone Arbeit ersparen und auf die wirklich interessanten Inhalte konzentrieren. Die Vergleichswerkzeuge müssen in der Lage sein, über die verschiedenen Textvarianten hinweg ähnliche Textpassagen zu identifizieren, diese zu alignieren und die Abweichungen, soweit sie für die angedachte Edition von Interesse sind, präzise aufzulisten. Sie müssen mit großen Textmengen umgehen können und trotz dieser Datenmengen kurze Reaktionszeiten haben, um ein flüssiges Arbeiten zu erlauben. Der Beitrag diskutiert das Problem des Textvergleiches an zwei Fallbeispielen, der ‚Wundarznei‛ des Heinrich von Pfalzpaint, ein frühneuhochdeutscher Text aus dem 15. Jahrhundert mit Überlieferungen aus dem 15. und 16. Jahrhundert [8], und der ‚Histoire philosophique et politique des établissements et du commerce des Européens dans les deux Indes‛ von Guillaume Thomas François Raynal, ein französischer Text aus dem 18. Jahrhundert, der mehrere starke Überarbeitungen erfuhr.
P. Molitor合作论文数 Institut f?r Informatik;Martin-Luther-Universit?t Halle-Wittenberg8