We have built a suite of tools in Python to proficiently analyze text reuse and intertextuality for a specific kind of set of medieval Arabic texts (commentaries) available in print. We take these printed editions, scan them, pre-process the images, give it to an OCR engine, clean the results, and store it in a data structure that mimics the explicit intertextual relation the texts have, and continue to perform data analysis on it. Digital approaches to medieval Arabic texts have either been at the micro-level in what has become known as a ‘digital edition’, i.e. the digital representation of one text, densely annotated, most commonly in TEI-XML, or it has been done at the macro-level in what is called a ‘digital corpus’, consisting of thousands of loosely encoded and sparsely annotated plain text files, accompanied by an entire infrastructure and high-performing software to perform broadly scoped queries. The micro-level generally is at the level of tens of thousands of words while the macro-level can be at the level of over a billion words. The micro-level is explicitly designed to be human readable first, while the macro-level is built to be machine readable first. At the micro-level, every little detail needs to be correct and in order, while at the macro-level a fairly large margin of error is still negligible as a mere rounding error. Amidst these levels we have been seeking a meso-level of digital analysis: neither edition nor corpus, but rather a group of texts at the level of hundreds of thousands to millions of words, with a small but perceptible margin of error, and a light but noticeable level of annotations, principally geared towards machine readability, but with ample opportunity for visual inspection and manual correction. In this paper we explain the rationale for our approach, the technical achievements it has led us to, and the results we so far obtained.
# This is a workflow that transforms scanned pages into readable text. The pages come from several printed Arabic books from the past few centuries. The workflow takes care of cleaning, OCR and postprocessing. A user can copy and paste image fragments of specks and symbols that must be removed before doing OCR. The workflow detects column layout and line boundaries. Individual lines will be passed to the OCR engine, which is Kraken using a model trained on many printed Arabic books. See [model](https://among.github.io/fusus/about/model.html). The result is stored in tab-separated files, with the transcription computed by the OCR step, plus position and confidence info resulting from that same step. The workflow can generate proofing pages that support manually checking the OCR results. # Next steps Once we have scanned a significant amount of pages, we'll construct a dataset in [Text-Fabric]() format out of it, with features that preserve positions of the words on the page and their confidence. From there we can implement steps to correct OCR mistakes and to perform intertextuality research between the ground work (the Fusus by Ibn Arabi) and its commentary books. # Authors This is work done by Cornelis van Lit and Dirk Roorda. There is more documentation about sources, the research project, and how to use this software in the [docs](https://among.github.io/fusus/).
In this article we are concerned with the question of what open access means in relation to research data. To that end we will (1) briefly discuss the definition of open access and the most important expected benefits; (2) highlight the publication - data distinction and the challenges regarding data; (3) discuss research data and open access in terms of its focus, licensing, and current practices at DANS in publishing data sets open access.
The BHSA (Biblia Hebraica Stuttgartensia Amstelodamensis) is the BHS text plus the linguistic annotations of the Eep Talstra Centre for Bible and Computer. The BHSA is available as a data set in Text-Fabric format. Text-Fabric is a minimalistic model to represent text: it provides addresses for all textual objects, so that it is easy to add arbitrary information at all textual levels, precisely and firmly anchored. A Text-Fabric resource resembles an IKEA ware house. The parts are nicely separated and stacked, so that they can be retrieved easily, to be combined into meaningful output later on. A consequence is that different teams with divergent purposes still can add to the same body of work, with a minimum of interference or duplication of work. Text-Fabric has helped with various types of data construction work, of which the most visible is the website SHEBANQ. We focus on two recent data combination jobs, (A) treebanks from the BHSA data and (B) a detailed comparison of the morphology in the BHSA and in the Open Scriptures effort. As the OSM is not yet finished, the comparison is repeatable.
One of the main products of the Humanities at Scale project was a profound (re)-conceptualisation of the DARIAH in-kinds and a web-based service [2] to collect, review and display in-kind contributions, in short DARIAH contributions. [1] This paper looks at the results of the implementation phase of this new service, and particularly focuses on baseline statistics and visual analytics of the content submitted (about 300 contributions for 2017 and 2018). In general, for European Research Infrastructures, so-called in-kind contributions are a way for the members to account for their national eorts under the umbrella of the ERIC. They may represent contributions available for all ERIC members (e.g., central services executed by an institution in a member country) and/or contributions which embody, complement, or enhance the mission and strategic actions of an ERIC on the national level. DARIAH's reference model [1] on the basis of which contributions are dened, introduces two main categories: `services' and `activities'. For them a detailed metadata scheme has been devised. Submitted contributions are further subject to a detailed self-assessment and reviewing process, one part of which is dedicated to determine those contributions which are put up for the financial accountability of a member's contributions. The web-based service replaces earlier forms of template-based and data-based submission of in-kinds, and enables immediate comparison of the submissions - also by a couple of visual interfaces (map, tables). In this paper, we aim to demonstrate the main benets of the tool and zoom into three aspects. The DARIAH contribution tool relies on the tedious and comprehensive work of the National Coordinators which are in charge of the submission process. To make their often invisible work more visible is one motivation behind this paper. The submitted content as such forms an interesting empirical base for reflection on what is seen as a DARIAH contribution by the DARIAH community. This, in turn, can inform the DARIAH strategy and help to monitor the success of its actions. This is the second motivation. Submissions can still vary in form greatly, despite the formal model. They remain human generated content. Analysing the variety helps to tune the tool to user requirements, to fix bugs, and to curate the data it collects. We conclude our paper with reflections on the future use of the tool and its connection to other DARIAH strategic actions, as designed in the Strategic Action Plan II.
The text of the Hebrew Bible is a subject of ongoing study in disciplines ranging from theology to linguistics to history to computing science. In order to study the text digitally, one has to represent it in bits and bytes, together with related materials. The author has compiled a dataset, called bhsa (Biblia Hebraica Stuttgartensia (Amstelodamensis)), consisting of the textual source of the Hebrew Bible according to the Biblia Hebraica Stuttgartensia (bhs), and annotations by the Eep Talstra Centre for Bible and Computer. This dataset powers the website shebanq and others, and is being used in education and research. The author has developed a Python package, Text-Fabric, to process ancient texts together with annotations. He shows how Text-Fabric can be used to process the bhsa. This includes creating new research data alongside it, and sharing it. Text-Fabric also supports versioning: as versions of the bhsa change over time, and people invest a lot in applications based on the data, measures are needed to prevent the loss of earlier results.
Within the DARIAH research infrastructure the need to collect, disseminate and monitor the contributions offered to the infrastructure has always been present. Work package 5 of the DARIAH Humanities at Scale (HaS) project has explored and described the activities and structure required to sustain these needs, resulting in the DARIAH contributions concept and procedure including a supporting online tool.ContributionsThe contributions to the DARIAH infrastructure consist either of services, being repeatable actions, or activities, which have a more unique/ oneBtime character. They are collected for several reasons. First of all, dissemination to the Arts & Humanities (AH contributions offered to the infrastructure need to be visible to the users within the community, so they can be found and used. Secondly the compatibility of the contributions to the infrastructure; how compatible are the contributions to the infrastructure, how easily can they be (re)used and/or combined with other services and last but not least monitoring of the variety and maturity of contributions as it is essential for the strategic planning and future development of the infrastructure.SpecificationsThe Contrib tool implements a registry of contributions with supports a workflow of (self)-assessing and reviewing contributions.The specifications of the contribution data and the workflow steps can be found in this report.
Annotation and research in the humanities are tightly coupled. Annotations can be seen as expressions of research activity which can be turned into input data for subsequent research. The digital paradigm has profoundly altered the ways in which we humans can handle the information content of our sources and it also affects the practice of annotation. We explore new ways of annotation that were not feasible before the digital times, and we list a few requirements for annotation to act as a reliable source of research information. Rather than conducting an academic discussion on the ontology of annotations, we highlight practical use cases for new kinds of annotations. We illustrate those in a concrete system for linguistic annotations to the Hebrew Bible, SHEBANQ.
Paper presented at ‘Technology, Software, Standards for the Digital Scholarly Edition’ DiXiT Convention, The Hague, September 14-18, 2015.
Digital Humanities (DH), as a growing field, generates lots of innovative new tools and methods to enhance work in humanities disciplines. To ensure a sustainable digital European infrastructure for long-term availability of digital research data and tools the DARIAH infrastructure (Digital Research Infrastructure for the Arts and Humanities) aims at the development of a strong cooperation and information exchange between research communities and institutions in the arts and humanities. The ambition to extend this community and create a European-wide network of expertise was one of the reasons to initiate the new Horizon 2020 project Humanities at Scale (HaS). The project coordinated by DARIAH-EU aims to evolve the DARIAH community, offer training and information material, set up a number of training workshops and develop core services within a sustainable framework. Within this framework a research infrastructure to connect DH research tools, services and data will be established and the challenge is to build technical systems to support this. The poster will highlight and answer the following questions from a technical perspective: • How can a decentralized framework (architecture of participation) of tools and services be developed and maintained? • Which kind of technologies are useful to establish a sustainable framework for DH-tools and -services? • What is needed to integrate externally developed tools into such a framework?
In this article we develop an algorithm to detect parallel texts in the Masoretic Text of the Hebrew Bible. The results are presented online and chapters in the Hebrew Bible containing parallel passages can be inspected synoptically. Differences between parallel passages are highlighted. In a similar way the MT of Isaiah is presented synoptically with 1QIsaa. We also investigate how one can investigate the degree of similarity between parallel passages with the help of a case study of 2 Kings 19-25 and its parallels in Isaiah, Jeremiah and 2 Chronicles.
In this article we reinvestigate the variation in Masoretic Hebrew of two linguistic features related to the object clause that have diachronic relevance according to various scholars. These features are the order of subject personal pronoun and nominal predicate in certain object clauses, and the variation between ×›×™ and ×שר introducing the object clause. We do this by using tools provided by the SHEBANQ project. By examining as many cases as possible throughout the Masoretic Text instead of only a few exemplary cases, we analyse the extent of variation in these syntactic constructions and their relevance for linguistic dating of biblical texts.