This paper defines a computational protocol to evaluate the performance of recognizers ofsolid line entities, circle entities, arc entities, dashed line entities, dashed circle entities, dashedarc entities, and text entities in engineering drawings. The protocol handles the one-to-manyand many-to-one matching problems so that detected or groundtruth entities are not multiplycounted.Keyword: Line-drawing recognition, benchmark, performance evaluation, documentimage database.1...
We study the 1/f noise currents and dark currents in LWIR HgCdTe photodiodes with different passivation. The diodes are fabricated by ion implanting boron on MBE HgCdTe with x=0.2173. One kind of photodiodes was passivated by ZnS and the other kind was passivated by CdTe/ZnS. Both dark currents and 1/f noise currents were measured at several reverse bias voltages. The measured dark currents of the photodiodes are analyzed using current model fitting methods. The different dark current components, such as diffusion current, generation-recombination current, trap assisted tunneling current and band-to-band tunneling current, at various biases voltages can be separated from the measured dark currents. The measurement results demonstrate that the dominant mechanism that produces 1/f noise in HgCdTe photodiodes with either passivation is tunneling. When the reverse bias voltages are less than 200mv, the main mechanism that produces 1/f noise is trap assisted tunneling. In this case, the 1/f noise currents of the photodiodes passivated by ZnS are smaller than those passivated by CdTe/ZnS. When the reverse biases are larger than 200mv, the band-to-band tunneling currents of the photodiodes passivated by ZnS are much larger than the photodiodes passivated by CdTe/ZnS. And the 1/f noise currents of the ZnS passivated photodiodes are larger than the different passivated one. In order to investigate the effect of surface passivation on the stability of two kinds of diodes, R-V characteristics and 1/f noise of the diodes were measured after vacuum baking for 10 hour at 80°C, the photodiodes passivated by CdTe/ZnS show higher performance compared with the diodes passivated by ZnS after baking.
This paper defines a computational protocol for evaluating the performance of raster to vector conversion systems. The graphical entities handled by this protocol are continuous and dashed lines, arcs, and circles, and text regions. The protocol allows matches of the type one-to-one, one-to-many, and many-to-one between the ground truth and the recognition results.
The Hg1-xCdxTe photovoltaic detectors with x=0.217 passivated by single ZnS layer ad dual (CdTe+ZnS) layers were fabricated in the same wafer. The fabricated devices were characterized by measurements of the diode low-frequency noise. The diode passivated by dual ( CdTe + ZnS) layers show higher performance compared to diode passivated by the single ZnS layer at high reverse bias, and. the modeling of diode dark current mechanisms indicate that the performance of the diode passivated by single ZnS is strongly affected by tunneling current related to the surface defects, which is responsible for the low frequency noise characteristics. By the analysis of X-ray reciprocal space map, it was found that the Q(y) direction broadening of HgCdTe epitaxial layer passivated by ZnS was wider than the CdTe + ZnS, which confirmed the existence of defects in the surface of HgCdTe epitaxial layer passivated by ZnS.
In this paper, we give a formal definition of a document image structure representation, and formulate document image structure extraction as a partitioning problem: finding an optimal solution partitioning the set of glyphs of an input document image into a hierarchical tree structure where entities within the hierarchy at each level have similar physical properties and compatible semantic labels. We present a unified methodology that is applicable to construction of document structures at different hierarchical levels. An iterative, relaxation-like method is used to find a partitioning solution that maximizes the probability of the extracted structure. All the probabilities used in the partitioning process are estimated from an extensive training set of various kinds of measurements among the entities within the hierarchy. The offline probabilities estimated in the training then drive all decisions in the online document structure extraction. We have implemented a text line extraction algorithm using this framework.
This paper presents a performance metric for the document structure extraction algorithms by finding the correspondences between detected entities and ground truth. We describe a method for determining an algorithm's optimal tuning parameters. We evaluate a group of document layout analysis algorithms on 1600 images from the UW-III Document Image Database, and the quantitative performance measures in terms of the rates of correct, miss, false, merging, splitting, and spurious detections are reported.
Presents a special symbol recognition system that incorporates the result of OCR to recognize the special symbols those nor handled by the current commercial OCR systems. Given a document image and the OCR output, we first refine the character coordinates produced by the OCR. Then, the special symbols are distinguished from the normal characters. Finally, we compute the features from the special symbol sub-images and a supervised classifier is used to assign the sub-images to one of the predefined special symbol categories. The system was tested on 5516 images from the National Library of Medicine. The evaluation results are reported in the paper.
This paper describes a text-line identification and segmentation technique that is probability based, where all probabilities are estimated from an extensive training set of various kind of measurements of distances between the terminal and non-terminal entities with which the algorithm works. The off-line probabilities estimated in the training then drive all decisions in the on-line segmentation algorithm. On the UW-III database of some 1600 scanned document image pages, having some 105020 text lines, the algorithm identifies and segments 104773 correctly, an accuracy of 99.76%.
The goal of document image structure analysis is to find an optimal solution partitioning the set of glyphs on a given document image into a hierarchical tree structure where entities within the hierarchy are associated with their physical properties and semantic labels. In this dissertation, we present a unified document image structure extraction algorithm that is probability based, where the probabilities are estimated from an extensive training set of various kinds of measurements of distances between the terminal and non-terminal entities with which the algorithm works. The off-line probabilities estimated in the training then drive all decisions in the on-line segmentation module. An iterative, relaxation-like method is used to find the partitioning solution that maximizes the joint probability. This approach can be uniformly apply to the construction of the document hierarchy at any level. We have implemented a text line segmentation algorithm and a text block extraction algorithm using this framework. Another example is the development of a system that detects and recognizes special symbols (Greek letters, mathematical symbols, etc.) on technical document pages, that are not handled by the current Optical Character Recognition (OCR) systems. A large quantity of ground-truth data, varying in quality, is required in order to give an accurate measurement of the performance of an algorithm under different conditions. We have constructed the University of Washington English Document Image Database-III, which contains 1600 scanned scientific/technical document image pages that come with manually edited ground-truth of entity bounding boxes and properties. Based on the ground-truth data, we can evaluate the performance of document analysis algorithms and build statistical models to characterize various types of document image structures. In this dissertation, we present a set of quantitative performance metrics for each kind of information a document image analysis technique infers. The text line and text block extraction algorithms were trained and evaluated on the UW-III database using a cross-validation method. The text line extraction algorithm identifies and segments 99.76% of text lines correctly, while the preliminary result of the text block extraction shows 91% accuracy.
In this paper, we discuss a performance evaluator for line-drawing recognition systems on images that contain binary digital logic schematic diagrams, a restricted subclass of engineering line drawings. The evaluator accepts inputs of IGES (the Initial Graphics Exchange Specification) files containing IGES primitives of straight lines, circles, partial arcs of circles, and IGES label block objects. Our evaluator takes two IGES files. One of these files is the recognition algorithm's output and the other IGES file is the corresponding groundtruth. The first step of processing involves parsing each IGES file and extracting IGES entities and the parameter information according to the IGES file format specification. The evaluator performs the evaluation for each pair of entities within these two files based on their types and the matching protocols and matching criteria defined in this paper. The results of our evaluator is a table of numbers which when weighted by application specific weights can be summed to produce an overall score relevant to the application.
A performance evaluation protocol for the layout analysis is discussed in this paper. In the University of Washington English Document Image Database-III, there are 1600 English document images that come with manually edited ground truth of entity bounding boxes. These bounding boxes enclose text and non-text zones, text-lines, and words. We describe a performance metric for the comparison of the detected entities and the ground truth in terms of their bounding boxes. The Document Attribute Format Specification is used as the standard data representation. The protocol is intended to serve as a model for using the UW-III database to evaluate the document analysis algorithms. A set of layout analysis algorithms which detect different entities have been tested based on the data set and the performance metric. The evaluation results are presented in this paper.
A document image analysis toolbox including a collection of data structures and algorithms to support a variety of applications, is described in this paper. An experimental environment is built to allow developers to develop, test and optimize their algorithms and systems. Appropriate and quantitative performance metrics for each kind of information a document analysis technique infers have been developed. The performance of each algorithm has been evaluated based on these metrics and the UW-III document image database which contains a total of 1600 English document images randomly selected from scientific and technical journals.
This paper describes the Document Image Understanding Toolbox currently under development at the University of Washington's Intelligent Systems Laboratory. The Toolbox provides a common data structure and a variety of document image analysis and understanding algorithms from which Toolbox users can construct document image processing systems. An algorithms for font attribute recognition based on the image analysis techniques available in the toolbox ISL DIU Toolbox is also presented.
The paper presents an efficient technique for document page layout structure extraction and classification by analyzing the spatial configuration of the bounding boxes of different entities on the given image. The algorithm segments an image into a list of homogeneous zones. The classification algorithm labels each zone as test, table, line-drawing, halftone, ruling, or noise. The text lines and words are extracted within text zones and neighboring text lines are merged to form text blocks. The tabular structure is further decomposed into row and column items. Finally, the document layout hierarchy is produced from these extracted entities
In this paper, we describe a feature based supervised zone classifier using only the knowledge of the widths and the heights of the connected-components within a given zone. The distribution of the widths and the heights of the connected-components is encoded into a n multiplied by m dimensional vector in the decision making. Thus, the computational complexity is in the order of the number of connected-components within the given zone. A binary decision tree is used to assign a zone class on the basis of its feature vector. The training and testing data sets for the algorithm are drawn from the scientific document pages in the UW-I database. The classifier is able to classify each given scientific and technical document zone into one of the eight labels: text of font size 8-12, text of font size 13-18, text of font size 19-36, display math, table, halftone, line drawing, and ruling, in real time. The classifier is able to discriminate text from non-text with an accuracy greater than 97%.
We consider the problem of zone classification in document image processing. Document blocks are labelled as text or nontext using texture features derived from a feature based interaction map (FBIM), a recently introduced general tool for texture analysis. The zone classification procedure proposed is tested on the comprehensive document image database UW-I created at the University of Washington in Seattle. Different classification procedures are considered. The performance ranges from 96% to 98% using 6 FBIM texture features only.
This paper discusses a method for binary morphological filter design to restore document images degraded by subtractive or additive noise, given a constraint on the size of filters. With a filter size restriction (for example 3 by 3), each pixel in output image depends only on its (3 by 3) neighborhood of input image. Therefore, we can construct a look-up table between input and output. Each output image pixel is determined by this table. So the filter design becomes the search for the optimal look-up table. By considering the degradation condition of the input image, we provide a methodology for knowledge based look-up table design, to achieve computational tractability. The methodology can be applied iteratively so that the final output image is the input image after being transformed through successive 3 by 3 operations. An experimental protocol is developed for restoring degraded document images, and improving the corresponding recognition accuracy rates of an OCR algorithm. We present results for a set of real images which are manually ground-truthed. The performance of each filter is evaluated by the OCR accuracy.