
Document structure analysis, such as zone segmentation and table recognition, is a complex problem in document processing and is an active area of research. The recent success of deep learning in solving various computer vision and machine learning problems has not been reflected in document structure analysis since conventional neural networks are not well suited to the input structure of the problem. In this paper, we propose an architecture based on graph networks as a better alternative to standard neural networks for table recognition. We argue that graph networks are a more natural choice for these problems, and explore two gradient-based graph neural networks. Our proposed architecture combines the benefits of convolutional neural networks for visual feature extraction and graph networks for dealing with the problem structure. We empirically demonstrate that our method outperforms the baseline by a significant margin. In addition, we identify the lack of large scale datasets as a major hindrance for deep learning research for structure analysis and present a new large scale synthetic dataset for the problem of table recognition. Finally, we open-source our implementation of dataset generation and the training framework of our graph networks to promote reproducible research in this direction.
This work summarizes the results of the first Competition on Harvesting Raw Tables from Infographics (ICDAR 2019 CHART-Infographics). The complex process of automatic chart recognition is divided into multiple tasks for the purpose of this competition, including Chart Image Classification (Task 1), Text Detection and Recognition (Task 2), Text Role Classification (Task 3), Axis Analysis (Task 4), Legend Analysis (Task 5), Plot Element Detection and Classification (Task 6.a), Data Extraction (Task 6.b), and End-to-End Data Extraction (Task 7). We provided a large synthetic training set and evaluated submitted systems using newly proposed metrics on both synthetic charts and manually-annotated real charts taken from scientific literature. A total of 8 groups registered for the competition out of which 5 submitted results for tasks 1-5. The results show that some tasks can be performed highly accurately on synthetic data, but all systems did not perform as well on real world charts. The data, annotation tools, and evaluation scripts have been publicly released for academic use.
The long-standing challenges for offline handwritten Chinese character recognition (HCCR) are twofold: Chinese characters can be very diverse and complicated while similarly looking, and cursive handwriting (due to increased writing speed and infrequent pen lifting) makes strokes and even characters connected together in a flowing manner. In this paper, we propose the template and instance loss functions for the relevant machine learning tasks in offline handwritten Chinese character recognition. First, the character template is designed to deal with the intrinsic similarities among Chinese characters. Second, the instance loss can reduce category variance according to classification difficulty, giving a large penalty to the outlier instance of handwritten Chinese character. Trained with the new loss functions using our deep network architecture HCCR14Layer model consisting of simple layers, our extensive experiments show that it yields state-of-the-art performance and beyond for offline HCCR.
With the emergence of the touchpad devices and drawing tablets, a new era of sketching started afresh. However, the recognition of sketches is still a tough task due to the variability of the drawing styles. Moreover, in some application scenarios there is few labelled data available for training, which imposes a limitation for deep learning architectures. In addition, in many cases there is a need to generate models able to adapt to new classes. In order to cope with these limitations, we propose a method based on few-shot learning and graph neural networks for classifying sketches aiming for an efficient neural model. We test our approach with several databases of sketches, showing promising results.
Handwriting production is a complex mechanism of fine motor control, associated with mainly two degrees of freedom in the horizontal and vertical directions. The relation between the horizontal and vertical velocities depends on the trajectory shape and its length. In this work, we explore the generation of handwriting velocities using two sinusoidal oscillations. The proposed method follows the motor equivalence theory and considers that the patterns are stored in the form of a sequence of corner shapes and its relative location in the letter. These points are referred to as the modulation points, where the parameters of the sinusoidal oscillations are modulated to generate required velocity profiles. Depending on the location and shape of the corners, the amplitude, phase, and frequency relations between the two underlying oscillations are changed. Accordingly, this paper presents an efficient method to synthesize the velocity profiles and hence the handwriting. Further, the shape variability in the synthesized data can also be introduced by modifying the position of the modulation points and its corner shapes. The quality of the synthesized handwriting is evaluated using both subjective and quantitative evaluation methods.
Historical Chinese character recognition faces problems including low image quality and lack of labeled training samples. We propose a generative adversarial network (GAN) based transfer learning method to ease these problems. The proposed TH-GAN architecture includes a discriminator and a generator. The network structure of the discriminator is based on a convolutional neural network (CNN). Inspired by Wasserstein GAN, the loss function of the discriminator aims to measure the probabilistic distribution distance of the generated images and the target images. The network structure of the generator is a CNN based encoder-decoder. The loss function of the generator aims to minimize the distribution distance between the real samples and the generated samples. In order to preserve the complex glyph structure of a historical Chinese character, a weighted mean squared error (MSE) criterion by incorporating both the edge and the skeleton information in the ground truth image is proposed as the weighted pixel loss in the generator. These loss functions are used for joint training of the discriminator and the generator. Experiments are conducted on two tasks to evaluate the performance of the proposed TH-GAN. The first task is carried out on style transfer mapping for multi-font printed traditional Chinese character samples. The second task is carried out on transfer learning for historical Chinese character samples by adding samples generated by TH-GAN. Experimental results show that the proposed TH-GAN is effective.
Alignment tasks generally seek to establish a spatial correspondence between two versions of a text, for example between a set of manuscript images and their transcript. This paper examines a different form of alignment problem, namely pixel-scale alignment between two renditions of a handwritten word or phrase. Using loopy inkball graph models, the proposed technique finds spatial correspondences between two text images such that similar parts map to each other. The method has applications to word spotting and signature verification, and can provide analytical tools for the study of handwriting variation.
This paper presents an attention-based convolutional sequence to sequence (ACseq2seq) model for recognizing an input image of multiple text lines from Japanese historical documents without explicit segmentation of lines. The recognition system has three main parts: a feature extractor using Convolutional Neural Network (CNN) to extract a feature sequence from an input image; an encoder employing bidirectional Long Short-Term Memory (BLSTM) to encode the feature sequence; and a decoder using a unidirectional LSTM with the attention mechanism to generate the final target text based on the attended pertinent features. We also introduce a residual LSTM network between the attention vector and softmax layer in the decoder. The system can be trained end-to-end by a standard cross-entropy loss function. In the experiment, we evaluate the performance of the ACseq2seq model on the anomalously deformed Kana datasets in the PRMU contest. The results of the experiments show that our proposed model achieves higher recognition accuracy than the state-of-the-art recognition methods on the anomalously deformed Kana datasets.
With the rapid emergence of new technologies, a voluminous number of images including document images is generated every day. Considering the volume of data and complexity of processes, manual analysis, annotation, recognition, classification, and retrieval, of such document images is impossible. To automatically deal with such processes, many document image analysis applications exist in the literature and many of them are currently in place in different organisation and institutes. The performance of those applications are directly affected by the quality of document images. Therefore, a document image quality assessment (DIQA) method is of primary need to allow users capture, compress and forward good quality (readable) document images to various information systems, such as online business and insurance, for further processing. To assess the quality of document images, this paper proposes a new full-reference DIQA method using first followed by second order Hast derivations. A similarity map is then created using second order Hast derivation maps obtained by employing Hast filters on both reference and distorted images. An average pooling is then employed to obtain a quality score for the distorted document image. To evaluate the proposed method, two different datasets were used. Both datasets are composed of images with the mean human opinion scores (MHOS) considered as ground truth. The results obtained from the proposed DIQA method are superior to the results reported in the literature.
This paper presents DECO (Dresden Enron COrpus), a dataset of spreadsheet files, annotated on the basis of layout and contents. It comprises of 1,165 files, extracted from the Enron corpus. Three different annotators (judges) assigned layout roles (e.g., Header, Data, and Notes) to non-empty cells and marked the borders of tables. Files that do not contain tables were flagged using categories such as Template, Form, and Report. Subsequently, a thorough analysis is performed to uncover the characteristics of the overall dataset and specific annotations. The results are discussed in this paper, providing several takeaways for future works. Furthermore, this work describes in detail the annotation methodology, going through the individual steps. The dataset, methodology, and tools are made publicly available, so that they can be adopted for further studies. DECO is available at: https://wwwdb.inf.tu-dresden.de/research-projects/deexcelarator/,
The problem of answering questions about an image is popularly known as visual question answering (or VQA in short). It is a well-established problem in computer vision. However, none of the VQA methods currently utilize the text often present in the image. These "texts in images" provide additional useful cues and facilitate better understanding of the visual content. In this paper, we introduce a novel task of visual question answering by reading text in images, i.e., by optical character recognition or OCR. We refer to this problem as OCR-VQA. To facilitate a systematic way of studying this new problem, we introduce a large-scale dataset, namely OCRVQA-200K. This dataset comprises of 207,572 images of book covers and contains more than 1 million question-answer pairs about these images. We judiciously combine well-established techniques from OCR and VQA domains to present a novel baseline for OCR-VQA-200K. The experimental results and rigorous analysis demonstrate various challenges present in this dataset leaving ample scope for the future research. We are optimistic that this new task along with compiled dataset will open-up many exciting research avenues both for the document image analysis and the VQA communities.
This paper presents an objective comparative evaluation of page segmentation and region classification methods for docu-ments with complex layouts. It describes the competition (mo-dus operandi, dataset and evaluation methodology) held in the context of ICDAR2019, presenting the results of the evaluation of twelve methods - nine submitted, three state-of-the-art sys-tems (commercial and open-source). Three scenarios are re-ported in this paper, one evaluating the ability of methods to accurately segment regions and two evaluating both segmenta-tion and region classification. Text recognition was a bonus challenge and was not taken up by all participants. The results indicate that an innovative approach has a clear advantage but there is still a considerable need to develop robust methods that deal with layout challenges, especially with the non-textual content.
The aim of this paper is to propose a new strategy adapted to the semantic segmentation of document images in order to extract baselines. Inspired by the work of Grüning [7], we used a convolutional model with residual layers enriched by an attention mechanism, called ARU-Net, a post-processing for the agglomeration of predictions and a data augmentation to enrich the database. Then, to consolidate the ARU-Net and help explicitly model dependencies between feature maps, we added a module of "Squeeze and Excitation" as proposed by Hu et al. [9]. Finally, to exploit the amount of unrated data available, we used a semi-supervised learning, based on ARUNet, through the use of adversary networks. This approach has shown some interesting predictive qualities, compared to Grüning's work, with easier processing and less task-specific error correction. The resulting performance improvement is a success.
In this paper, we present an evaluative study of pixel-labeling methods using the HBA 1.0 dataset for historical book analysis. This study is held in the context of the 2nd historical book analysis (HBA2019) competition and in conjunction with the 15th IAPR international conference on document analysis and recognition (ICDAR2019). The HBA2019 competition provides a large experimental corpus and a thorough evaluation protocol to ensure an objective performance benchmarking of pixel-labeling document image methods. Two nested challenges are evaluated in the HBA2019 competition: Challenge 1 and Challenge 2. Challenge 1 evaluates how image analysis methods could discriminate the textual content from the graphical ones at pixel level. Challenge 2 assesses the capabilities of pixel-labeling methods to separate the textual content according to different text fonts (e.g. lowercase, uppercase, italic, etc.) at pixel level. During the competition, we received 52 and 38 different teams' registrations for Challenge 1 and Challenge 2, respectively and finally 5 of them submitted their results in each challenge. Qualitative and numerical results of the participating methods in both challenges are reported and discussed in this paper in order to provide a baseline for future evaluation studies in historical document image analysis. The evaluation shows that the method submitted by the NLPR-CASIA team achieves the highest performance in both challenges.
Document image binarization, especially old handwritten documents, is a very important yet challenging task. There are various bottlenecks for binarizing historical documents due to different types of degradation present imultaneously such as back impression, ink bleed through, faded colours, and wear and tear of the writing media. We consider these degradation as various types of noise in the document image. Here we have proposed a 2D morphological network which consists of basic morphological operation like dilation and erosion to perform our targeted task. The network also includes linear combination of output from dilation and erosion operations. The aforementioned 2D morphological network is applied for image binarization, where the structuring elements (SEs) and the weights of the linear combination layer are learned through back-propagation. The proposed network has been evaluated on DIBCO 2017 and H-DIBCO 2018 and ISI-Letter dataset. Our results show more convincing as compared to the results of other state-of-the-art methods. Though the network is developed for old handwritten documents, it may be tuned to work for image processing task. The source code can be found here https://github.com/ranjanZ/ICDAR_Binarization.
Recognizing the layout of unstructured digital documents is an important step when parsing the documents into structured machine-readable format for downstream applications. Deep neural networks that are developed for computer vision have been proven to be an effective method to analyze layout of document images. However, document layout datasets that are currently publicly available are several magnitudes smaller than established computing vision datasets. Models have to be trained by transfer learning from a base model that is pre-trained on a traditional computer vision dataset. In this paper, we develop the PubLayNet dataset for document layout analysis by automatically matching the XML representations and the content of over 1 million PDF articles that are publicly available on PubMed Central. The size of the dataset is comparable to established computer vision datasets, containing over 360 thousand document images, where typical document layout elements are annotated. The experiments demonstrate that deep neural networks trained on PubLayNet accurately recognize the layout of scientific articles. The pre-trained models are also a more effective base mode for transfer learning on a different document domain. We release the dataset (https://github.com/ibm-aur-nlp/PubLayNet) to support development and evaluation of more advanced models for document layout analysis.
The following topics are dealt with: learning (artificial intelligence); document image processing; convolutional neural nets; text analysis; handwritten character recognition; feature extraction; neural nets; optical character recognition; image segmentation; image recognition.
State-of-the-art methods for document image classification rely on visual features extracted by deep convolutional neural networks (CNNs). These methods do not utilize rich semantic information present in the text of the document, which can be extracted using Optical Character Recognition (OCR). We first study the performance of state-of-the-art text classification approaches when applied to noisy text obtained from OCR. We then show that fusing this textual information with visual CNN methods produces state-of-the-art results on the RVL-CDIP classification dataset.
This study is different from most of the recent text detection work which focuses on creating a robust text detector system. In this work we studied how script languages affect a text detector's performance by using a multi-language synthetic dataset-namely, the Synthetic Octa-Language (SOL) dataset. The effect of script languages continues to be largely unexplored. Previously, this kind of experiment was infeasible because too many factors influence the performance of a text detector. We really cannot tell what role the factor X plays, neither positive nor negative. To overcome these difficulties, we used controlled synthesized data, which allows us to explicitly control factors such as base image, script language, text content, text color, font face, and font size. With the SOL dataset, we were able to investigate the effect that script languages have on on deep neural-network (DNN)-based methods under different scenarios. Moreover, this dataset can be used in other script-language-related text detection research as well.
The text recognition system for natural images or video frames containing multilingual text needs a method to first identify the written script and then recognize the word in the identified script. However, the occurrence of some scripts is rare as compared to others. Due to the availability of a few samples of the rare script, the supervised learning of the deep neural networks is difficult. To overcome this problem, we have proposed a zero-shot learning based method for script identification. We have also proposed architecture for script identification which fuses the global feature vector and the semantic embedding vector. The semantic embedding of the script is obtained by using the spatial dependency of the stroke's sequence via the recurrent neural network. The proposed architecture shows superior results as compared to the baseline approaches.