Automatic pattern detection has become increasingly important for scholars in the humanities as the number of manuscripts that have been digitised has grown. Most of the state-of-the-art methods used for pattern detection depend on the availability of a large number of training samples, which are typically not available in the humanities as they involve tedious manual annotation by researchers (e.g. marking the location and size of words, drawings, seals and so on). This makes the applicability of such methods very limited within the field of manuscript research. We propose a learning-free approach based on a state-of-the-art Naïve Bayes Nearest-Neighbour classifier for the task of pattern detection in manuscript images. The method has already been successfully applied to an actual research question from South Asian studies about palm-leaf manuscripts. Furthermore, state-of-the-art results have been achieved on two extremely challenging datasets, namely the AMADI_LontarSet dataset of handwriting on palm leaves for word-spotting and the DocExplore dataset of medieval manuscripts for pattern detection. A performance analysis is provided as well in order to facilitate later comparisons by other researchers. Finally, an easy-to-use implementation of the proposed method is developed as a software tool and made freely available.
Presenting actual research questions from academia through publishing datasets is a practice of great importance in order to generate relevant solutions. Therefore, we propose a dataset of handwriting on papyri for the task of writer identification. This dataset is derived directly from research questions in the field of Papyrology, and the samples are selected by experts from the respective field of research. This dataset consists of 50 handwriting samples in Greek on papyri approximately from the 6th century A.D., which belong to 10 different scribes. It is prepared and made freely available for non-commercial research along with their confirmed groundtruth information related to the task of writer identification. This paper presents not only the details of the dataset but also its relation to research questions and how the results of computational analysis can support scholars from manuscript research. Some preprocessing and experimentation results are provided as well in order to highlight the difficulties posed by the image degradation of this dataset.
In this paper we propose a method for eliminating SIFT keypoints in document images. The proposed method is applied as a first step towards word spotting. One key issue when using SIFT keypoints in document images is that a large number of keypoints can be found in non-textual regions. It would be ideal if we could eliminate as much as irrelevant keypoints as possible in order to speed-up processing. This is accomplished by altering the original matching process of SIFT descriptors using an iterative process that enables the detection of keypoints that belong to multiple correct instances throughout the document image, which is an issue that the original SIFT algorithm cannot tackle in a satisfactory way. The proposed method manages a reduction over 99% of the extracted keypoints with satisfactory performance.
Several methods have been proposed for the task of writer identification for historical manuscripts. Most of these methods have been evaluated on private historical datasets only, while few have been evaluated on the recently published historical dataset provided by the Historical-WI competition at ICDAR-2017. Nevertheless, there is no thorough analysis in the literature available w.r.t. the degradation typically found in the digitized manuscripts. Furthermore, the currently proposed methods are beyond the reach of the scholars from the humanities; either because of the impracticality of the method itself, or because of the lack of an easy-to-use implementation. In this paper, we analyse a state-of-the-art method against common degradation types in historical manuscripts using images from the virtual manuscript library of Switzerland. Furthermore, we show that, by optimising a key parameter, we can enhance the performance of the method and significantly outperform the winner method of the Historical-WI competition. Finally, we demonstrate the practicality of our implementation yielding intuitively comprehensible results for direct use of scholars from the humanities.
This paper introduces new methodologies for reliably identifying writers of Arabic historical manuscripts. We propose an approach that transforms key point-based features, such as SIFT, into a global form that captures high-level characteristics of writing styles. We suggest a modification for a common local feature, the contour direction feature, and show the contribution of combining local and global features for writer identification. Our work also presents a novel algorithm that determines the number of writers involved in writing a given manuscript. The experimental study confirms the significant improvement in this algorithm on writer identification once applied to historical manuscripts. Comprehensive experiments using different features and classification schemes demonstrate the vitality of the suggested methodologies for reliable writer identification. The presented techniques were evaluated on both historical and modern documents where the suggested features yielded very promising results with respect to state-of-the-art features.
Background: Cognitive Muscular TherapyTM (CMT) is an integrated behavioural intervention developed for knee osteoarthritis.CMT teaches patients to reconceptualise the condition, integrates muscle biofeedback and aims to reduce muscle overactivity, both in response to pain and during daily activities.This nested qualitative study explored patient and physiotherapist perspectives and experiences of CMT.Methods: Five physiotherapists were trained to follow a well-defined protocol and then delivered CMT to at least two patients with knee osteoarthritis.Each patient received seven individual clinical sessions and was provided with access to online learning materials incorporating animated videos.Semi-structured interviews took place after delivery/completion of the intervention and data were analysed at the patient and physiotherapist level.Results: Five physiotherapists and five patients were interviewed.All described a process of changing beliefs throughout their engagement with CMT.A framework with three phases was developed to organise the data according to how osteoarthritis was conceptualised and how this changed throughout their interactions with CMT.Firstly, was an identification of pain beliefs to be challenged and recognition of how current beliefs can misalign with daily experiences.Secondly was a process of challenging and changing beliefs, validated through new experiences.Finally, there was an embedding of changed beliefs into self-management to continue with activities. Conclusion:This study identified a range of psychological changes which occur during exposure to CMT.These changes enabled patients to reconceptualise their condition, develop a new understanding of their body, understand psychological processes, and make sense of their knee pain.
Writer identification and verification can be viewed as a classification problem, where each writer represents a class. We propose a classifier for offline, text-independent, and segmentation-free writer identification based on the Local Naïve Bayes Nearest-Neighbour (Local NBNN) classification. Our proposed method takes into consideration the particularity of handwriting patterns by adding a constraint to prevent the matching of irrelevant keypoints. Furthermore, a normalisation factor is proposed to cope with the prevalent problem of unbalanced data. The method has been evaluated on several public datasets of different writing systems and state-of-the-art results are shown to be improved.
Welcome to the 2015 International Conference on Document Analysis and Recognition in Nancy, France. It is truly an honor to host our premier conference. This conference was originally planned, and fully organized to take place in Tunis, Tunisia. Sadly, for safety reasons following the terrorist attacks in Tunisia, the conference was relocated to Nancy, France before less than two months. It is saddening that terrorism and obscurantism has its way and deprives the document analysis community from discovering the beautiful country and culture of Tunisia. You will always be welcome in Tunisia!
The discrimination between languages is one of the first steps in the problem of automatic documents text recognition. In many documents, such as bank checks and application forms, printed and handwritten texts are mixed. In this paper, an automatic identification system of Arabic and French words in both handwritten and printed script based on Gaussian Mixture Models (GMMs) was presented. A fixed-length sliding window was used for the feature extraction. Experiments using some parts of the freely available AHTID/MW, APTI and RIMES databases show a remarkable performance of the proposed approach. RESUME. La discrimination entre les langues est l'une des premieres etapes dans le probleme de reconnaissance automatique des documents de textes. Dans de nombreux documents, tels que les cheques bancaires et les formulaires, les textes imprimes et manuscrits sont melanges. Dans cet article, nous proposons un systeme d'identification automatique des mots arabes et francais dans les deux formes: manuscrite et imprimee. Ce systeme est base sur les modeles de melanges gaussiens (GMMs). Pour l'extraction des caracteristiques, nous utilisons une fenetre glissante de longueur fixe. Des experimentations utilisant quelques parties des bases gratuitement disponibles AHTID/MW, APTI et RIMES montrent une performance remarquable de l'approche proposee.
Binarization is often used for pixel-wise document text extraction as preprocessing step for scanned historical documents. These documents are scanned in color and high resolution today. The reduction of color to grayscale images and the subsequent binarization implies a loss of information and often results in unsatisfying processing results. In this paper, a color segmentation instead of a binarization approach is used to segment text from background in historical manuscripts. A color segmentation approach based on Markov random fields with a reduced set of required parameters is presented to segment text written in different colors from noisy page background. First tests with historical Arabic manuscripts show promising results. In case of words written in light red color, our approach shows better results than a state-of-the-art binarization approach.
In this paper, we describe a novel method for handwriting style identification. A handwriting style can be common to one or several writer. It can represent also a handwriting style used in a period of the history or for specific document. Our method is based on Gaussian Mixture Models (GMMs) using different kind of features computed using a combined fixed-length horizontal and vertical sliding window moving over a document page. For each writing style a GMM is built and trained using page images. At the recognition phase, the system returns log-likelihood scores. The GMM model with the highest score is selected. Experiments using page images from historical German document collection demonstrate good performance results. The identification rate of the GMM-based system developed with six historical handwriting style is 100%.
In this paper, we will present a mathematical analysis of the transition proportion for the normal threshold (NorT) based on the transition method. The transition proportion is a parameter of NorT which plays an important role in the theoretical development of NorT. We will study the mathematical forms of the quadratic equation from which NorT is computed. Through this analysis, we will describe how the transition proportion affects NorT. Then, we will prove that NorT is robust to inaccurate estimations of the transition proportion. Furthermore, our analysis extends to thresholding methods that rely on Bayes rule, and it also gives the mathematical bases for potential applications of the transition proportion as a feature to estimate stroke width and detect regions of interest. In the majority of our experiments, we used a database composed of small images that were extracted from DIBCO 2009 and H-DIBCO 2010 benchmarks. However, we also report evaluations using the original (H-)DIBCO׳s benchmarks.
In this article, our goal is to describe mathematically and experimentally the gray-intensity distributions of the fore- and background of handwritten historical documents. We propose a local pixel model to explain the observed asymmetrical gray-intensity histograms of the fore- and background. Our pixel model states that, locally, the gray-intensity histogram is the mixture of gray-intensity distributions of three pixel classes. Following our model, we empirically describe the smoothness of the background for different types of images. We show that our model has potential application in binarization. Assuming that the parameters of the gray-intensity distributions are correctly estimated, we show that thresholding methods based on mixtures of lognormal distributions outperform thresholding methods based on mixtures of normal distributions. Our model is supported with experimental tests that are conducted with extracted images from DIBCO 2009 and H-DIBCO 2010 benchmarks. We also report results for all four DIBCO benchmarks.
Determining the individuality of handwriting in ancient manuscripts is an important aspect of the manuscript analysis process. Automatic identification of writers in historical manuscripts can support historians to gain insights into manuscripts with missing metadata such as writer name, period, and origin. In this paper writer classification and retrieval approaches for multi-page documents in the context of historical manuscripts are presented. The main contribution is a learning-based rejection strategy which utilizes writer retrieval and support vector machines for rejecting a decision if no corresponding writer can be found for a query manuscript. Experiments using different feature extraction methods demonstrate the abilities of our proposed methods. A dedicated data set based on a publicly available database of historical Arabic manuscripts was used and the experiments show promising results.
Since printed/handwritten Arabic text recognition is a very challenging research field and the recognition methodologies are different, it is important to separate these two types of texts before the recognition phase. In this paper, we introduce a simple and effective method to identify printed and handwritten Arabic words using local features. A Gaussian Mixture Models (GMMs) based approach is used to model the printed and handwritten classes. Experimental results using some parts of the freely available IFN/ENIT, AHTID/MW and APTI databases show that our method is robust and provides very good identification performance.
In this paper, we present a new and freely available dataset comprising 80 pages of an historical handwritten Arabic document in conjunction with a detailed ground truth for the development and evaluation of segmentation-free word spotting approaches. Besides information on the underlying manuscript and technical details, we introduce a comprehensive list of tags that each word is labeled with. These tags can be used for research on specific issues such as dealing with text in different colors. For comparison of different word spotters, a fixed set of 25 keywords with different properties is included. Furthermore, some specifics of spotting on Arabic manuscripts are discussed. We exemplarily present a state-of-the-art word spotting algorithm in its original and a new extended implementation and evaluate both approaches on the new dataset. For comparison, they are also tested on the widely used George Washington dataset. It is shown that the extended word spotter outperforms the original version in terms of mean average precision on both datasets.
Optical Font Recognition (OFR) has been proven to increase Optical Character Recognition (OCR) accuracy, but it can also help in harvesting semantic information from documents. It therefore becomes a part of many Document Image Analysis (DIA) pipelines. Our work is based on the hypothesis that Local Binary Patterns (LBP), as a generic texture classification method, can address several distinct DIA problems at the same time such as OFR, script detection, writer identification, etc. In this paper we strip down the Redundant Oriented LBP (RO-LBP) method, previously used in writer identification, and apply it for OFR with the goal of introducing a generic method that classifies text as oriented texture. We focus on Arabic OFR and try to perform a thorough comparison of our method and the leading Gaussian Mixture Model method that is developed specifically for the task. Depending on the nature of proposed OFR method, each method's performance is usually evaluated on different data and with different evaluation protocols. The proposed experimental procedure addresses this problem and allows us to compare OFR methods that are fundamentally different by adapting them to a common measurement protocol. In performed experiments LBP method achieves perfect results on large text blocks generated from the APTI database, while preserving its very broad generic attributes as proven by secondary experiments.
This paper describes the first edition of the Arabic writer identification competition using AHTID/MW and KHATT databases held in the context of the 14th International Conference on Frontiers in Handwriting Recognition (ICFHR2014). This competition has used the new freely available Arabic Handwritten Text Images Database written by Multiple Writers (AHTID/MW) and the Arabic handwritten text database called KHATT presented in ICFHR2012. We propose three tasks in this Arabic writer identification competition: the first and second are based respectively on word and text line level using the AHTID/MW database and the third one is paragraph based using the KHATT database. We received one system for the second task, three systems for the third task and none for the first task. All systems are tested in a blind manner using a set of images kept internal. A short description of the participating groups, their systems, the experimental setup, and the observed results are presented.