
The biodiversity literature is one of the longest-standing examples of recording heritage in the world. Today there are many efforts to standardise and integrate the literature to ensure access to the information, both for heritage and research purposes. Ontologies are increasingly being turned to as knowledge representation tools in these efforts. However, the validity of using ontological frameworks to represent biological taxonomies has been questioned. Biological taxonomies use the scientific nomenclature to assign names to described species. While the nomenclature is a useful classification tool, it can also be a source of confusion because of its synonymous, homonymous and fluid nature. Despite this, no empirical evaluation of scientific nomenclature use in the literature has ever been performed. Corpus-based analysis is already used in automatic ontology extraction, and this study explores the possibility of applying recently developed lexicography techniques to the problem to provide an evaluation of the empirical data in the literature, and serve as a comparison with existing ontologies. This paper focuses on the work flow, parameters and preliminary findings of the research investigating how to extract structures from the literature to perform these comparisons. It uses the manipulation of corpus analysis techniques, visualisation and filtering methods to do so and evaluates potential classification and disambiguation qualities of the resulting graphs for future work. Preliminary results look at the effects of frequency and salience when filtering the graphs, which indicate that these filter parameters could be used for different purposes in revealing relationships between organism mentions.
While morphological segmentation has always been a hot topic in Arabic, due to the morphological complexity of the language and the orthography, most effort has focused on Modern Standard Arabic. In this paper, we focus on pre-MSA texts. We use the Gradient Boosting algorithm to train a morphological segmenter with a corpus derived from Al-Manar, a late 19th/early 20th century magazine that focused on the Arabic and Islamic heritage. Since most of the cultural heritage Arabic available suffers from substandard orthography, we have trained a machine learner to standardize the text. Our segmentation accuracy reaches 98.47%, and the orthography standardization an F-macro of 0.98 and an F-micro of 0.99. We also produce stemming as a by-product of segmentation.
Part-of-speech tagging, morphological tagging, and lemmatization of historical texts pose special challenges due to the high spelling variability and the lack of large, high-quality training corpora. Researchers therefore often first map the words to their modern spelling and then annotate with tools trained on modern corpora. We show in this paper that high quality part-of-speech tagging and lemmatization of historical texts is possible while operating directly on the historical spelling. We use a part-of-speech tagger based on bidirectional long short-term memory networks (LSTMs) [11] with character-based word representations and lemmatize using an encoder-decoder system with attention. We achieve state-of-the-art results for modern German morphological tagging on the Tiger corpus and also on two historical corpora which have been used in previous work.
This paper explores the way in which the old catalogues from the Romanian folklore archives can be improved with updated information about two key aspects of the folklore collections: the informant/performer and the village/location where the recordings where made. Using the headers from three different archives we will show the solution found for enriching the equivalent to dc:creator and dc:coverage. We will particularly focus on the case study of a series of recordings from Western Transylvania and how the metadata was cleaned, spliced and improved.
This paper presents the development and implementation of a robust OCR tool and a related comprehensive workflow for the recognition of Greek printed polytonic scripts. This project is initiated and developed by an interdisciplinary team with expertise in the areas of document image processing, character segmentation and recognition, machine learning, corpus creation and digital humanities. Our paper aims to describe the design and development of the workflow around this project, including data gathering and structuring, OCR tool development, user interface development, experiments on the training procedure of the tool, evaluation, post-correction and quality control of the results.
In this paper, we argue that exploitation of historical corpus data requires text metadata which metadata accompanying digital objects from digital libraries, archives or other electronic text collections, do not provide. Most text collections describe in their metadata the object (book, newspaper) containing the text. To do research on the style of an author, or study the language of a certain time period, or a phenomenon through time, correct metadata is needed for each word in the text, which leads to a very intricate metadata scheme for some text collections. We focus on the Nederlab corpus. Nederlab is a research environment that gives access to a large diachronic corpus of Dutch texts from the 6th - 21st century, of more than 10 billion words. The corpus has been compiled using existing digitised text material from researchers, research organisations, archives and libraries. The aim of Nederlab is to provide tools and data to enable researchers to trace long-term changes in Dutch language, culture and society. This type of research sets high-level requirements on the metadata accompanying the texts. Since the Nederlab corpus consists of different collections, each with their own metadata, the task of adding the appropriate metadata was not straightforward, all the more so because of the difference in perspective content providers and corpus builders have. We will describe the desired metadata scheme and how we tried to realize this for a corpus of the size of Nederlab.
Historical ciphers, a special type of manuscripts, contain encrypted information, important for the interpretation of our history. The first step towards decipherment is to transcribe the images, either manually or by automatic image processing techniques. Despite the improvements in handwritten text recognition (HTR) thanks to deep learning methodologies, the need of labelled data to train is an important limitation. Given that ciphers often use symbol sets across various alphabets and unique symbols without any transcription scheme available, these supervised HTR techniques are not suitable to transcribe ciphers. In this paper we propose an un-supervised method for transcribing encrypted manuscripts based on clustering and label propagation, which has been successfully applied to community detection in networks. We analyze the performance on ciphers with various symbol sets, and discuss the advantages and drawbacks compared to supervised HTR methods.
The TRIADO project (2016-2019) is a cooperation between Netwerk Oorlogsbronnen (coordinator), NIOD Institute for War, Holocaust and Genocide Studies, Huygens ING/KNAW Humanities Cluster and the National Archives of the Netherlands (Nationaal Archief). TRIADO explores technological strategies to transform analogue text-based archival collections into digital data that can be used for research. The first part of the project is about trying out new techniques to open up collections, the second part is a 'reality check' to explore the research potential of the data created. Increasingly, archives, libraries and museums (ALMs) digitize their analogue historical collections. Yet, in 2017 it was estimated that only approximately one tenth of all heritage collections in Europe have been digitized so far. There is still a large gap between the specific needs of the digital humanities-community and the digital 'raw materials' supplied by the ALMs. Text-based historical collections are potentially interesting to a wide range of different scientific disciplines, but so far - in case of the Netherlands - only a few digitized archives are equipped to be used for digital research. The main aim of TRIADO is to bridge this gap by performing a 'laboratory to reality'-check with the most frequently consulted WWII archive in the Netherlands: the Central Archive of Special Jurisdiction (CABR). The CABR held by the Nationaal Archief (National Archives of the Netherlands) consists of the legal case files of some 300,000 persons accused of collaborating with the German occupier. The CABR contains approximately 4 kilometers of analogue documents (shelf space), ranging from minutes and verdicts to membership cards, forms and summons. Most documents are typed or hybrid (typed/handwritten). The experimental pilot project TRIADO focuses on two complementary research questions: 1. Which digital methods are best suited (in terms of quality, efficiency, etc) to make large corpora of unstructured, imperfect data, based on analogue collections, usable as a research facility? 2. Is it possible to answer specific, mainly quantitative statistical research questions on the basis of the digital data created under 1? A sample of 13.8 meters from the CABR was digitized to test technologies and perform experiments. Also, a workflow for mass digitization was devised and a demonstrator was built to showcase the results of the experiments. In this paper we discuss the main findings of the research done in part 1. This paper reports on processes for mass digitization, OCR quality and improvement, auto-classification of document types, named entity recognition, date extraction and matching of existing name lists to OCR'd data.
The paper describes the method and results of validation of 14 library catalogues. The format of the catalog record is Machine Readable Catalog (MARC21) which is the most popular metadata standards for describing books. The research investigates the structural features of the record and as a result finds and classifies different commonly found issues. The most frequent issue types are usage of undocumented schema elements, then improper values in places where a value should be taken from a dictionary, or should match to other strict requirements.
PoCoTo is known as a web-based interactive tool for the postcorrection of OCR-results on historical texts. In this paper we first introduce A-PoCoTo, a fully automated extension of PoCoTo designed for the use in large-scale digitization projects. Among other features, A-PoCoTo takes into account the recognition results of several OCR-engines on the given input text, and sentence context is used for refining rankings and decisions. Preliminary evaluation results are given. In view of the very high level of accuracy needed for many scholarly applications it is questionable if a fully automated process is always able to fully meet the standards expected by researchers in Digital Humanities. We describe the architecture of A-I-PoCoTo, a postcorrection system (under development) combining automated postcorrection as a first step and interactive postcorrection as an optional second step. In A-I-PoCoTo decisions and correction steps of the automated component are stored in a special protocol. Views offered by the graphical user interface help to efficiently confirm, reject, or improve these decisions as a first step of the manual postcorrection.
This paper describes first large scale article detection and extraction efforts on the Finnish Digi1 newspaper material of the National Library of Finland (NLF) using data of one newspaper, Uusi Suometar 1869-1898. The historical digital newspaper archive environment of the NLF is based on commercial docWorks2 software. The software is capable of article detection and extraction, but our material does not seem to behave well in the system in this respect. Therefore, we have been in search of an alternative article segmentation system and have now focused our efforts on the PIVAJ machine learning based platform developed at the LITIS laboratory of University of Rouen Normandy [11--13, 16, 17]. As training and evaluation data for PIVAJ we chose one newspaper, Uusi Suometar. We established a data set that contains 56 issues of the newspaper from years 1869-1898 with 4 pages each, i.e. 224 pages in total. Given the selected set of 56 issues, our first data annotation and experiment phase consisted of annotating a subset of 28 issues (112 pages) and conducting preliminary experiments. After the preliminary annotation and experimentation resulting in a consistent practice, we fixed the annotation of the first 28 issues accordingly. Subsequently, we annotated the remaining 28 issues. We then divided the annotated set into training and evaluation sets of 168 and 56 pages. We trained PIVAJ successfully and evaluated the results using the layout evaluation software developed by PRImA research laboratory of University of Salford [6]. The results of our experiments show that PIVAJ achieves success rates of 67.9, 76.1, and 92.2 for the whole data set of 56 pages with three different evaluation scenarios introduced in [6]. On the whole, the results seem reasonable considering the varying layouts of the different issues of Uusi Suometar along the time scale of the data.
The intensified circulation of people, commodities and ideas is one of the characteristics of a globalizing world. To understand the causes and consequences of these circulations, we have to know which commodities circulated when and where, on what scale and who made them circulate. In our paper we want to present the first results of a CLARIAH Research Pilot1 on diamonds in Borneo, using the large historical newspaper collection of the KB (Royal Library of the Netherlands) in Delpher2. So far the diamond industry in Borneo has been a true blind spot in our knowledge on the global diamond commodity chain. We have little information on where diamonds were found, who the miners and traders were and if there was really an 'age-old' diamond polishing industry as literature suggests. We believe that the newspapers can provide more information on this topic. To answer these questions, we developed a workflow that enables us to query the KB newspaper collection in an efficient and elaborate way that can also be used for research on other commodities.
The article aims to be an introduction to the dependency treebanks currently available for Ancient Greek and Latin, i.e., the Ancient Greek and Latin Dependency Treebank (AGLDT), the Index Thomisticus Treebank (IT-TB), the PROIEL Treebank, and the SEMATIA Treebank. Their pipelines for creation of morphosyntactic annotations are presented so as to highlight major commonalities and differences. All treebanks share the same basic underlying formalism, whereby syntactic words are connected to each other to form labeled directed acyclic graphs, and their annotation schemes, although different, are comparable to a very large extent.
Various research projects were concerned with the development and adaptation of methods for OCR specifically for historical printed documents (cf. METAe [20], IMPACT [1], eMOP [9]). However, these initiatives have ended before the wide adoption of deep neural networks and, despite the various project's achievements, there remains a lack of OCR software that is a) comprehensive with regard to the challenges presented by the wide variety of historical documents and b) available as ready-to-use Free Software. The OCR-D project aims to rectify that. In this paper we introduce the background of OCR-D, the main challenges and shortcomings in the availability of open tools and resources for OCR of historical printed documents and discuss the various software modules and related components (repositories, workflows) that are being made available through OCR-D. Finally we provide an outlook to a number of remaining challenges that are not addressed by OCR-D and point out several examples for the positive community aspects arisen through the creation and sharing of open resources for historical German OCR.
Historic itinerary research investigates the traveling paths of historic entities, to determine their influence and reach. A potential source of such information are the Regesta Imperii (RI), a large-scale resource for European medieval history research. However, two important intermediate problems must be addressed: 1. place names may be stated as unknown or are left empty; 2., place name queries return large candidate sets of points scattered all across Europe and the correct point must be selected. For 1., we perform a place name completion step to predict place names for regests referencing charters of unknown origin. To address 2., we formulate a graph framework which allows efficient reconstruction of the emperors' itineraries by means of shortest path finding algorithms. Our experiments show that our method predicts coordinates of places with significant correlation to human gold coordinates and significantly outperforms a baseline which selects points randomly from the candidate sets. We further show that the method can be leveraged to detect errors in human coordinate labels of place names.
In this paper, we show that the OCR engine Ocropy can be trained for fonts used in rather complex and varied Coptic typeset. For each of the three fonts presented in this paper, we used a number of texts from scholarly editions with different philological and editorial standards and texts from two different dialects of Coptic (Bohairic and Sahidic). Despite the complexity of the training data, we observed accuracy rates of 97.5%, for one font even up to 99%.
Current OCR has limited capability for Arabic because of script models lacking scientific basis. We propose a new OCR strategy for Arabic, based on 1. Islamic script grammar including extended shaping and 2. treating Arabic script as a multi-layered writing system. We analyse Arabic script as an allographic rendering of graphemic abstractions. Grapheme is a term adapted from phonology; it is analogous to the term phoneme. In phonology, the smallest functional unit of sound is the phoneme. This is not heard, but perceived. What one hears are contextually conditioned allophones. In Arabic orthography, the smallest functional unit of spelling is the grapheme. This is not seen, but perceived. What one sees are contextually conditioned allographs. In our analysis, the letter block is the minimum unit of Arabic script formation and therefore of script grammar. A letter block is a single allograph or of a group of fused allographs surrounded by graphic space. The analogy with phonology can be pushed further: the archiphoneme is a bundle of shared features between two or more phonemes, minus their distinctive features. The archigrapheme is the bundle of shared features between two or more graphemes, minus their distinctive features. An archigraphemic letter block consists of one or more reduced allographs between spaces. The letter block follows the base line. There can be ligatures between letter blocks. In our strategy the archigraphemic letter block also forms the minimum unit of OCR. We have (1) implemented an algorithm that reduces any Unicode text in Arabic script to archigraphemes and we used it to create a list in Unicode format of all attested unique archigraphemic letter blocks on the internet. (2) With this list, and applying extended Islamic script grammar, we can synthesize realistic images of all possible archigraphemic fusions in a given style. These two developments make it possible to create an OCR system for recognizing synthetic Arabic under controlled conditions for both basic and extended shaping in a given style. These two steps result in competence, after which the OCR system should be trained to apply tolerance for the variation of performance in real documents. To interpret the identified letter blocks linguistically, a technique for the parsing of archigraphemes must be developed. For example, the single sequence of the three archigraphemic letter blocks EBD A LLH can be interpreted as several different surface texts such as abda-n li llaahi, abdu l-laahi and inda l-laahi. To facilitate the linguistic phase of the process, the same list of unique archigraphemic letter blocks is designed to identify the language of the text under scrutiny. In this phase we can present • Islamic script synthesis • Unicode conversion from plene orthography to archigraphemic transliteration • the archigraphemic search algorithm • the list of unique archigraphemic letter blocks • samples of authentic shape generation These are the first steps towards static OCR technology. The next step is to create or find matching AI software to teach OCR to recognize any unmapped letter blocks in order to make the OCR dynamic.
The search through large corpora of unstructured text on the Web Domain is not an easy task and such services are not offered for the common user. One such corpus is the published works of east Christian fathers by Jacques Paul Migne, known as Patrologia Graeca (PG). In this paper, an application of a Databaseless model is presented for extracting information from the unstructured patristic works of PG on the Web Domain. The user queries terms that may exist in PG and retrieves all the paragraphs-fragments that contain these terms. The time for retrieving the information is faster than implementing this querying system in a common Relational Database Management System (RDMS). Our proposed system is portable, secure and can be easily maintained by institutions or organizations. The system auto-transforms the PG corpus into a Representational State Transfer Access Point Interface (REST API) for retrieving and processing the information, using the JavaScript Object Notation (JSON) format. The User Interface (UI) is completely distinguished from the backbone of the system while the system, on the other hand, can be easily extended for more complicated queries as well as to be applied to other corpora. Two kinds of user interfaces are described: The first one is completely static, useful for the average user using the Web browser. The second one illustrates the use of simple Shell Scripting, for searching and extracting statistical, syntactical and semantic information in real time. In both cases we try to strip down the complexity in order to accomplish the corpus transformation and searching, in a simple, secure and manageable way. Difficulties and key problems are also discussed.
Numerical data of considerable significance is present in historical documents in tabular form. Due to the challenges involved in the extraction of this data from the scanned documents it is not available to researchers in a useful representation that unlocks the underlying statistical information. This paper sets out to create a better understanding of the problem of extracting and representing statistical information from numerical tables, in order to enable the creation of appropriate technical solutions and also for collection holders to appropriately plan their digitisation projects to better serve their readers. To that effect, after an initial overview of current practices in digitisation and representation of historical numerical data, the authors' findings are presented from a scoping exercise of the Wellcome Library's high-profile collection of the Medical Officer of Health reports. In addition to users' perspectives and a detailed examination of the nature and structure of the data in the reports, a study of the extraction and integration of the data is also described.
This paper presents experiments on Optical character recognition (OCR) of historical newspapers and journals published in Finland. The corpus has two main languages: Finnish and Swedish and is written in both Blackletter and Antiqua fonts. Here we experiment with how much training data is enough to train high accuracy models, and try to train a joint model for both languages and all fonts. So far we have not been successful in getting one best model for all, but it is promising that with the mixed model we get the best results on the Finnish test set with 95 % CAR, which clearly surpasses previous results on this data set.