This paper introduces new methodologies for reliably identifying writers of Arabic historical manuscripts. We propose an approach that transforms key point-based features, such as SIFT, into a global form that captures high-level characteristics of writing styles. We suggest a modification for a common local feature, the contour direction feature, and show the contribution of combining local and global features for writer identification. Our work also presents a novel algorithm that determines the number of writers involved in writing a given manuscript. The experimental study confirms the significant improvement in this algorithm on writer identification once applied to historical manuscripts. Comprehensive experiments using different features and classification schemes demonstrate the vitality of the suggested methodologies for reliable writer identification. The presented techniques were evaluated on both historical and modern documents where the suggested features yielded very promising results with respect to state-of-the-art features.
Für die Digital Humanities im Bereich Mediävistik und Frühneuzeitforschung stellt die Digitalisierung von Handschriften ein zentrales Feld dar. Da jede Handschrift eigene Charakteristika aufweist, führt die automatische Erstellung eines maschinenlesbaren Textes durch Optical Character Recognition (OCR) anhand von Digitalisaten in den allermeisten Fällen zu fehlerhaften Ergebnissen. Andererseits können Charakteristika dieser Schrift wie Buchstabengröße und -abstand, Dichte des Schriftbildes, Neigung u.a. genutzt werden, um die Identifikation der schreibenden Hand bzw. Hände zu ermöglichen. In dem Beitrag wird gezeigt, wie die Analyse von Handschriftenabbildungen zur Identifikation der schreibenden Hand bzw. Hände genutzt werden kann. Ein Algorithmus soll sonstige paläographische oder kodikologische Befunde unterstützen und Argumente zur Verioder Falsifikation von unsicheren Zuschreibungen liefern. For Digital Humanities in medieval studies and early modern studies, the digitization of manuscripts is a central field. Since each manuscript displays its own unique characteristics, the automatic generation of a machine-readable text using Optical Character Recognition (OCR) as applied to digital images leads, in most cases, to error-prone results. However, characteristics of handwriting such as the size of letters and spacing, slope, and so on can be used to identify the scribe or scribes. This paper demonstrates how the analysis of manuscript images can be used to identify the scribe or scribes. An algorithym will support additional paleographic and codicological findings and provide evidence for the verification or falsification of uncertain attributions.
Determining the individuality of handwriting in ancient manuscripts is an important aspect of the manuscript analysis process. Automatic identification of writers in historical manuscripts can support historians to gain insights into manuscripts with missing metadata such as writer name, period, and origin. In this paper writer classification and retrieval approaches for multi-page documents in the context of historical manuscripts are presented. The main contribution is a learning-based rejection strategy which utilizes writer retrieval and support vector machines for rejecting a decision if no corresponding writer can be found for a query manuscript. Experiments using different feature extraction methods demonstrate the abilities of our proposed methods. A dedicated data set based on a publicly available database of historical Arabic manuscripts was used and the experiments show promising results.
Identification of writers of handwritten historical documents is an important and challenging task. In this paper we present several feature extraction and classification approaches for the identification of writers in historical Arabic manuscripts. The approaches are able to successfully identify writers of multipage documents. The feature extraction methods rely on different principles, such as contour-, textural- and key point-based and the classification schemes are based on averaging and voting. For all experiments a dedicated data set based on a publicly available database is used. The experiments show promising results and the best performance was achieved using a novel feature extraction based on key point descriptors.
In this paper, we present a new and freely available dataset comprising 80 pages of an historical handwritten Arabic document in conjunction with a detailed ground truth for the development and evaluation of segmentation-free word spotting approaches. Besides information on the underlying manuscript and technical details, we introduce a comprehensive list of tags that each word is labeled with. These tags can be used for research on specific issues such as dealing with text in different colors. For comparison of different word spotters, a fixed set of 25 keywords with different properties is included. Furthermore, some specifics of spotting on Arabic manuscripts are discussed. We exemplarily present a state-of-the-art word spotting algorithm in its original and a new extended implementation and evaluate both approaches on the new dataset. For comparison, they are also tested on the widely used George Washington dataset. It is shown that the extended word spotter outperforms the original version in terms of mean average precision on both datasets.
Recently, many big libraries all over the world have been scanning their collections to make them publicly available and to preserve historical documents. We present a modular software system which can be used as a tool for semi-automatical processing of historical handwritten Arabic documents. The development of this system is part of the HADARA project which aims for historical document analysis of Arabic manuscripts and consists of a project team including engineers and computer scientists but also users such as linguists and historians. The HADARA system is designed to support script and content analysis, identification, and classification of historical Arabic documents. The system has been created following an iterative development approach, and the current version assists the user in an interactive and partially already in an automatic manner. In this paper a system overview is given and the first modules are presented which support the annotation of a scanned manuscript in a semi-automatic manner They comprise page layout analysis, text line segmentation, and transcription. Word spotting is the first application implemented in the HADARA system and its concept is outlined in this paper
The problem of highly imbalanced datasets with only sparse data of the minority class in the context of two-class classification is investigated. A novel synthetic data oversampling technique is proposed which utilizes estimations of the probability density distribution in the feature space. First, a Gaussian mixture model (GMM) from the data of the well-sampled majority class is generated and with its help a new GMM is approximated by Bayesian adaptation using the sparse minority class data. Random synthetic data is generated from the adapted GMM and an additional assignment rule assigns this data to either the minority class or else discards it. The obtained synthetic data is employed in combination with the available original data to train a support vector machine classifier. The examined application in this paper is optical on-line process monitoring of laser brazing with only rare sporadic occurring defects. Experiments with different amounts of minority class data samples and comparisons to other methods show that this approach performs very well for highly imbalanced datasets.
This paper investigates on the training of classifiers with highly imbalanced datasets for industrial quality control. The application is on-line process monitoring of laser brazing processes and only a limited amount of data of an imperfection class is available for training. Bayesian adaptation is used to derive a model of the imperfection class from a well sampled model of the class representing a high grade joint surface. For this application, we are able to show that with the sparse training data a performance comparable to a training with a balanced dataset is achievable and even a moderate increase of training data quickly yields a performance gain.
This paper presents a joint imperfection detection method for on-line process monitoring of laser brazing processes. The method uses images from two spectral ranges and fuses their features with the scores of the separately calculated log-likelihood ratios to decide whether the currently generated part of a joint contains an imperfection or not. To avoid unconfident classifications, a decision reject is implemented. Furthermore, the classifier can be adapted to different quality requirements.
Laser brazing of zinc coated steel is a widely established manufacturing process in the automotive sector, where high quality requirements must be fulfilled. The strength, impermeablitiy and surface appearance of the joint are particularly important for judging its quality. The development of an on-line quality control system is highly desired by the industry. This paper presents recent works on the development of such a system, which consists of two cameras operating in different spectral ranges. For the evaluation of the system, seam imperfections are created artificially during experiments. Finally image processing algorithms for monitoring process parameters based the captured images are presented.
Recently, many big libraries all over the world have been scanning their collections to make them publicly available and to preserve historical documents. We present a modular software system which can be used as a tool for semi-automatical processing of historical handwritten Arabic documents. The development of this system is part of the HADARA project which aims for historical document analysis of Arabic manuscripts and consists of a project team including engineers and computer scientists but also users such as linguists and historians. The HADARA system is designed to support script and content analysis, identification, and classification of historical Arabic documents. The system has been created following an iterative development approach, and the current version assists the user in an interactive and partially already in an automatic manner. In this paper, a system overview is given and the first modules are presented which support the annotation of a scanned manuscript in a semi-automatic manner. They comprise page layout analysis, text line segmentation, and transcription. Word spotting is the first application implemented in the HADARA system and its concept is outlined in this paper. Introduction Nowadays, there is a trend to digitize printed or handwritten historical documents in many big libraries all over the world. Scanned images are published on library websites after manually adding metadata information. This information provides the ability to search for a specific document in a large database. Unfortunately, searching through the content of a document is not possible as long as the content itself is not digitally available in a textual form. However, a manual transcription requires large efforts in terms of time and costs. To overcome these limitations computer scientists and researchers in the field of document analysis and recognition develop algoritms for an automatic transcription of scanned historical documents or at least to support specific parts of this task. Pattern recognition methods, such as automatic text recognition or word spotting, are employed for this purpose. On the one hand, large scale digitization projects are underway at most big libraries and even private companies (e. g., Million Book Project [1] or Google Book Search [2]) to preserve paper documents in digital format. On the other hand, many researchers all over the world work on projects for historical document processing and recognition, as can be seen at conferences like ICDAR [3] , ICFHR [4], or DAS [5], workshops like HIP [6], and research projects like IMPACT [7]. The cooperation of exFigure 1. Example pages of a scanned Arabic handwritten book with side notes1 perts from digitization projects and researchers from the field of document analysis and recognition is gaining importance. For example, the very time consuming and expensive task of the generation of training data for recognition tasks can benefit from cooperation of librarians and computer scientists. Typical pages of an Arabic handwritten book written in the 18th century (the text dates back to the 12th century) are shown in Figure 1. In the center of each page the main body text is written, additionally comments and remarks are written on the page borders. These are added usually neither from the same writer nor from the same period as the main body text. This example shows one of the major problems for interpretation of historical Arabic manuscripts. An overview about the challenges of historical document processing is given in [8]. The HADARA project team consists of scientists from signal processing, computer science, science of history, and linguistics. The diversity inside this group ensures that the different needs of the involved areas of expertise are respected during the whole development process. The core of the HADARA system consists of an easy-to-use historical document processing tool chain 1Scan provided by the Damascene family library Refaiya at the University Library in Leipzig, Germany, Website: http://www.refaiya.unileipzig.de/ Scanned Documents Metadata Generation Database Word Spotting Origin Identification Browsing Writing Classification Book Search Data Retrieval Palaeography Codicology Text Analysis Segmentation Annotation Transcription Denoising Binarization Image Preprocessing Processing