In almost all current approaches, the collation of large texts is applied to a fixed given segmentation of the two texts witnesses to be compared and consists of two consecutive steps. First, the segments of the two texts are aligned, and then the aligned segments are compared in detail. For larger manuscripts or books consisting of many pages, the segments are usually the paragraphs of the texts. When comparing two texts, where the second text is a revised version of the first, poor local alignments can arise. This occurs in places where paragraphs have been split into two smaller paragraphs to insert a new paragraph in between, or where several consecutive sentences have been moved from one paragraph to the previous or next paragraph. Most paragraph collation tools cannot handle these scenarios properly because they align each paragraph with at most one paragraph of the other text. In this paper, we discuss this problem in detail and present a heuristic for resegmenting the two texts to be compared in order to achieve a better collation.
A main feature of VLSI design systems is the placement and routing aspect. A routing problem is given by a routing region, a set of multiterminal nets (the demands) and the number of available layers. In this chapter the routing region will always be a planar graph, called grid. Cross-overs, junctions, knock-knees, bends, and vias may only be placed on vertices of the routing region. Routing itself typically consists of two steps. The first step determines the placement of the routing segments, which is the wire layout. In the second step which is called wiring or layer assignment, each wire segment of the wire layout has to be assigned to one of the k available layers so that the segments are electrically connected in the right way.
In the humanities, text comparison is essential for scholarly editing the traditions of a text with its different witnesses under study. For the computer-aided creation of a critical apparatus, there have been established approaches and tools for many years. However, they mainly focus on subsentences or sentences. An efficient and easy-to-use automatic comparison of entire chapters or books still represents a research desideratum. The Locate, Explore, Retrace and Apprehend complex text variants (LERA) working environment presented here solves this issue. It is based on a two-stage collation approach: an efficient, fully automatic alignment of text segments, which can be paragraphs, subparagraphs or sentences, with interactive post-processing options, followed by the detailed comparison at segment level. Because aligning text segments, such as paragraphs, for more than two text witnesses is algorithmically challenging, we discuss the heuristics we developed in more detail. LERA combines the entire process of document management, tokenization/segmentation, normalization, alignment, and visualization with interactive control options and exploratory tools. It has already been and is being successfully applied in several Digital Humanities projects of different languages, e.g. for Arabic, French, Hebrew as well as German and English texts.
This article presents an approach to the collation problem of text witnesses which is correct for all distance functions and is applicable to the alignment of any types of tokens, be it letters, words, sentences, or paragraphs. In contrast to the existing approaches, which all are heuristics in nature, we specify a formal model in the form of an integer linear program (ILP) that formally describes the optimization problem. For a special case that usually occurs in practice, we then transfer the ILP into an efficient algorithm that optimally solves the alignment problem. In the second part of the article, we apply our approach to the alignment of paragraphs of text witnesses—an alignment problem that has received little attention in the literature so far—and compare the results of our approach, which we have called TSaligner, with those of CollateX.
Günter Hotz hat im Laufe der vielen Jahre, in denen er an der Universität des Saarlandes als akademischer Lehrer tätig war, insgesamt 54 „Kinder“ zur Promotion, manche von ihnen dann auch zur Habilitation geführt. Nachfolgend sind sie mit ihrem Promotions- und, wo zutreffend, Habilitationsthema aufgelistet, bevor in den nachfolgenden Kapiteln von jedem einzelnen Doktorkind (akademischer) Lebenslauf folgen, bei einigen auch mit einem mehr oder weniger umfangreichen Beitrag weiter ergänzt.
Abstract In this paper, A shorter version of the paper appeared in German in the final report of the Digital Plato project which was funded by the Volkswagen Foundation from 2016 to 2019. [35], [28]. we present a method for paraphrase extraction in Ancient Greek that can be applied to huge text corpora in interactive humanities applications. Since lexical databases and POS tagging are either unavailable or do not achieve sufficient accuracy for ancient languages, our approach is based on pure word embeddings and the word mover’s distance (WMD) [20]. We show how to adapt the WMD approach to paraphrase searching such that the expensive WMD computation has to be computed for a small fraction of the text segments contained in the corpus, only. Formally, the time complexity will be reduced from O ( N · K 3 · log K )\mathcal{O}(N\cdot {K^{3}}\cdot \log K) to O ( N + K 3 · log K )\mathcal{O}(N+{K^{3}}\cdot \log K), compared to the brute-force approach which computes the WMD between each text segment of the corpus and the search query. N is the length of the corpus and K the size of its vocabulary. The method, which searches not only for paraphrases of the same length as the search query but also for paraphrases of varying lengths, was evaluated on the Thesaurus Linguae Graecae® (TLG®) [25]. The TLG consists of about 75 · 10 6 75\cdot {10^{6}} Greek words. We searched the whole TLG for paraphrases for given passages of Plato. The experimental results show that our method and the brute-force approach, with only very few exceptions, propose the same text passages in the TLG as possible paraphrases. The computation times of our method are in a range that allows its application in interactive systems and let the humanities scholars work productively and smoothly.
It is undisputed that with the application of digital methods humanities issues can be addressed which could not be addressed so far.This is the third special issue on digital methods in the humanities that is published in it -Information Technology.The first one was published in 2009 by Thomas Burch, Claudine Moulin and Andrea Rapp, at that time all working at the Trier Center for Digital Humanities [1], the second one in 2016 by Manfred Thaller from University of Cologne [2].Both issues presented humanities questions where the use of digital methods is useful or even necessary.As Manfred Thaller wrote in his editorial, many of them are challenges definitely worthy of a computer scientist.However, the situation has hardly changed in the last decade.Digital Humanities continue to be an issue only in the humanities and are largely ignored by computer scientists.In Germany in particular, there are only a few working groups in computer science dealing with the counterpart of Digital Humanities, the so called eHumanities which is concerned with the development of new and non-trivial information technology approaches to support humanities scholars in addressing their issues.With this special issue we would like to draw our computer science colleagues' attention to the exciting topics and the associated challenges at the interface between humanities and computer science, once again.Even though the special issue only deals with digital methods for textbased studies, there are exciting issues for computer scientists in almost all humanities fields.Take art history or archaeology, for example, where computer vision or 3D-modelling and -reconstruction play central roles.Or take political and social sciences, where big data analysis and deep learning methods have become indispensable digital methods.To answer questions in the humanities, a sound knowledge of data structures and efficient algo-
zur Konferenz Digital Humanities im deutschsprachigen Raum 2020 Keter Shem #ov Prozessualisierung eines Editionsprojekts mit 100 Textzeugen
The paper presents a BDD-based post-synthesis technique to detect the redundant gates in a reversible circuit. Given a reversible circuit C, we are looking for a maximal (or most costly) subset of gates in C that can be removed from C without altering the functionality of the circuit. The runtime of the new algorithm is linear in the size of the involved binary decision diagrams (BDD). In order to lower the runtimes, the presented approach is extended to handle the restricted problem of only looking for up to k gates that can be removed from C for some constant k. This restriction should ensure that the sizes of the involved BDDs remain practicable for adequate constants k.
Heinz Zemanek†, Wien it–Information Technology is a strictly peer-reviewed scientific journal. It is the oldest German journal in the field of information technology. Today, the major aim of it–Information Technology is highlighting issues on ongoing newsworthy areas in information technology and informatics and their application. It aims at presenting the topics with a holistic view. It addresses scientists, graduate students, and experts in industrial research and development. it–Information Technology is the organ of the faculties Technical Informatics and Computer Science in the Life Sciences of the German Informatics Society (GI–Gesellschaft für Informatik eV) and faculty Technical Informatics of the Information Technology Society (ITG) within the VDE (Verband der Elektrotechnik, Elektronik und Informationstechnik eV). it–Information Technology was founded as Elektronische Rechenanlagen in 1959. To live up to …
H. Strack, S. Wefel, P. Molitor, M. Räckers, J. Becker, J. Dittmann, R. Altschaffel, J. Marx Gómez, N. Brehm, A. Dieckmann 1 Hochschule Harz, Fachbereich Automatisierung und Informatik, hstrack@hs-harz.de 2 Martin-Luther-Universität Halle-Wittenberg, Institut für Informatik, sandro.wefel@informatik.unihalle.de 3 Martin-Luther-Universität Halle-Wittenberg, Institut für Informatik, paul.molitor@informatik.unihalle.de 4 Westfälische Wilhelms-Universität Münster, Institut für Wirtschaftsinformatik, michael.raeckers@ercis.uni-muenster.de 5 Westfälische Wilhelms-Universität Münster, Institut für Wirtschaftsinformatik, becker@ercis.unimuenster.de 6 Otto-von-Guericke-Universität Magdeburg, Institut für Technische und Betriebliche Informationssysteme (ITI), jana.dittmann@iti.cs.uni-magdeburg.de 7 Otto-von-Guericke-Universität Magdeburg, Institut für Technische und Betriebliche Informationssysteme (ITI), robert.altschaffel@iti.cs.uni-magdeburg.de 8 Carl von Ossietzky Universität Oldenburg, jorge.marx.gomez@uni-oldenburg.de 9 Ernst-Abbe-Hochschule Jena, Fachbereich Wirtschaftsingenieurwesen, nico.brehm@fh-jena.de 10 Ministerium für Wirtschaft Wissenschaft und Digitalisierung des Landes Sachsen-Anhalt, Magdeburg, andreas.dieckmann@mw.sachsen-anhalt.de
One fundamental problem of bioinformatics is the computational recognition of DNA and RNA binding sites. Given a set of short DNA or RNA sequences of equal length such as transcription factor binding sites or RNA splice sites, the task is to learn a pattern from this set that allows the recognition of similar sites in another set of DNA or RNA sequences. Permuted Markov (PM) models and permuted variable length Markov (PVLM) models are two powerful models for this task, but the problem of finding an optimal PM model or PVLM model is NP-hard. While the problem of finding an optimal PM model or PVLM model of order one is equivalent to the traveling salesman problem (TSP), the problem of finding an optimal PM model or PVLM model of order two is equivalent to the quadratic TSP (QTSP). Several exact algorithms exist for solving the QTSP, but it is unclear if these algorithms are capable of solving QTSP instances resulting from RNA splice sites of at least 150 base pairs in a reasonable time frame. Here, we investigate the performance of three exact algorithms for solving the QTSP for ten datasets of splice acceptor sites and splice donor sites of five different species and find that one of these algorithms is capable of solving QTSP instances of up to 200 base pairs with a running time of less than two days.
AbstractThis paper presents a digital working environment for exploratory analysis of the changes to the text of Raynal’s Histoire philosophique et politique des établissements et du commerce des Europeéns dans les deux Indes, a French text from 1770 about the negative influence of European civilization during the colonization of the East and West Indies which was significantly revised in 1774, 1780 and post-mortem in 1820. After a brief summary of the printing history, the features of the environment are presented. Starting with a general map of the corpus by means of a navigation bar and word clouds, the online platform provides users with the opportunity to analyse the text genesis in an exploratory manner. The general map leads the users to text passages of interest with respect to the text genesis. Synoptic representations together with colour markings adaptable to the respective philological problem make the revision history easily comprehensible and the text differences easily identifiable. The approach is, in principle, transferable to scholarly editions of other texts and texts in other languages.
This article examines the impacts of capturing correspondence metadata through an exhaustive discussion of how details such as the sender and recipient of a letter, their respective addresses, and the date of writing can be entered in an intuitive and accurate fashion. The focus is on the development of fundamental input mechanisms, which can be reused for specific metadata items. Our discussion results in the implementation of a proof-of-concept application, which shows that only a single text box is needed for each metadata item, without violating any of the self-imposed goals and requirements.
The editors of it – Information Technology express their gratitude to all colleagueswho served as peer reviewers for 2013 and 2014, namely – Jörg Ackermann – Wolfgang Aigner – Nadine Amende – Anushka Anand – Gennady Andrienko – Michael Aupetit – Ann Barcomb – Sven Behnke – Karsten Berns – Christian Bischof – Roland Bless – Hamish Carr – Remco Chang – Stefan Dessloch – Christian Schlögl – Stefan Conrad – Carlos Denner – Stephan Eggersglüß – Niklas Elmqvist – Oliver Ferschke – Goerschwin Fey – Bernd Finkbeiner – Paul Groth – Tobias Gummer – Anja Hartmann – Lars Hetmank – Andreas Heuer – Gerhard Heyer – Ralf Hofestädt – Gottfried Hofmann – Kim Holmberg – Dieter Hutter – Nicolas Jullien – Wolfgang Kellerer – Andreas Kempf
Eine der zentralen Aufgaben bei editionsphilologischen Vorhaben ist die Feststellung der Abhängigkeiten von zu unterschiedlichen Zeitpunkten entstandenen Textzeugnissen. Für die Edition handschriftlicher Textzeugnisse beispielsweise bezieht sich dies auf die editionsphilologisch zentrale Frage des Zusammenhangs unterschiedlicher Handschriften und ihrer Filiation untereinander und bei Texten, die starke Überarbeitungsprozesse erfahren haben, auf die Frage der Textgenese. Mit leistungsstarken Werkzeugen zum Textvergleich kann sich der Editor viel monotone Arbeit ersparen und auf die wirklich interessanten Inhalte konzentrieren. Die Vergleichswerkzeuge müssen in der Lage sein, über die verschiedenen Textvarianten hinweg ähnliche Textpassagen zu identifizieren, diese zu alignieren und die Abweichungen, soweit sie für die angedachte Edition von Interesse sind, präzise aufzulisten. Sie müssen mit großen Textmengen umgehen können und trotz dieser Datenmengen kurze Reaktionszeiten haben, um ein flüssiges Arbeiten zu erlauben. Der Beitrag diskutiert das Problem des Textvergleiches an zwei Fallbeispielen, der ‚Wundarznei‛ des Heinrich von Pfalzpaint, ein frühneuhochdeutscher Text aus dem 15. Jahrhundert mit Überlieferungen aus dem 15. und 16. Jahrhundert [8], und der ‚Histoire philosophique et politique des établissements et du commerce des Européens dans les deux Indes‛ von Guillaume Thomas François Raynal, ein französischer Text aus dem 18. Jahrhundert, der mehrere starke Überarbeitungen erfuhr.
C. Scholl合作论文数Albert-Ludwigs-University Freiburg;Institute of Computer Science11
G. Hotz合作论文数Universit?t des Saarlandes9
Dietmar Saupe合作论文数Department of Computer and Information Science, University of Konstanz2