We describe algorithms for identifying the language of text in document images which are complex, unoriented, and degraded. We distinguish among seven lan-page layouts may be complex, containing text blocks in unknown roughly Manhat-tan arrangements. The pages may be unoriented, that is, upright or rotated by 90, 180, or 270 degrees. The images may be degraded by digitization at coarse and unequal spatial sampling rates as in FAXes. We begin by segmenting the page into text lines in a manner oblivious to page skew and both page and text-line orientation. Then we distinguish between Asian and Latin scripts at any orientation. Chinese versus Japanese is decided at any orientation, and then their orientation is detected. On Latin scripts, we detect rst orientation and then language. A variety of decision procedures are used, some hand-crafted (e.g. using spatial features and optical density distributions) and others trainable (e.g. using word unigram relative entropy models). Tests on 1088 standard (low) resolution FAX images show that our method accurately identiies scripts (98.16%), and language and page orientations (94.76%).
We describe an image analysis system for handling complex and noisy images of forms and bank documents, such as business checks, personal checks, or bank deposits. Some of these document types have no standardized layout, requiring a careful analysis of the whole image, to find out where the relevant information, for example the courtesy amount, is located. Each element in the image is first classified as being part of machine printed text, handwritten text, or as being a graphical element, such as a line. To obtain a reliable identification of these different elements under noisy conditions, a set of templates is scanned over the image, extracting such elements as strokes, line end stops and corners. From this representation a quick and robust analysis of the image's content is possible to identify the different parts. Once a set of candidate subimages has been found, they are sent to a field recognition system. We describe an example of one such system, which locates and reads courtesy amounts on US checks.
We developed a machine vision system around an analog neural net chip and used it in several applications. Some of them were: locating the address blocks on mail pieces, finding the identification numbers on rail cars, and discriminating between handwritten and machine-printed characters. The chip, operating as a coprocessor of a workstation, provides a speed-up of a factor of 1000, compared with the workstation. The computation speed achieved lies between one and ten billion multiply-accumulates/s. The neural net chip is based on building blocks,neurons, that can be arranged in various network architectures. The dataflow is optimized for implementing large, structured neural nets, and is also suited for any task in which signals are to be convolved with many kernels. Some of the networks are trained on the neural net chip with a weight-perturbation learning algorithm that was adapted to work with the coarse quantization of the weights and the states in the chip.
We describe a method, “Shortest Path Segmentation” (SPS), which combines dynamic programming and a neural net recognizer for segmenting and recognizing character strings. We describe the application of this method to two problems: recognition of handwritten ZIP Codes, and recognition of handwritten words. For the ZIP Codes, we also used the method to automatically segment the images during training: the dynamic programming stage both performs the segmentation and provides inputs and desired outputs to the neural network. Results are reported for a test set of 2642 unsegmented handwritten 212 dpi binary ZIP Code (5- and 9-digit) images. For handwritten word recognition, we combined SPS with a “Space Displacement Neural Network” approach, in which a single-character-recognition network is extended over the entire word image, and in which SPS techniques are then used to rank order a given lexicon. We report results on a test set of 3000 300 ppi gray scale word images, extracted from images of live mail pieces, for lexicons of size 10, 100, and 1000. Representing the problem as a graph as proposed in this paper has advantages beyond the efficient finding of the final optimal segmentation, or the automatic segmentation of images during training. We can also easily extend the technique to generate K “runner up” answers (for example, by finding the K shortest paths). This paper will also describe applications of some of these ideas.
We developed a neural net architecture for segmenting complex images, i.e., to localize two-dimensional geometrical shapes in a scene, without prior knowledge of the objects' positions and sizes. A scale variation is built into the network to deal with varying sizes. This algorithm has been applied to video images of railroad cars, to find their identification numbers. Over 95% of the characters were located correctly in a data base of 300 images, despite a large variation in lighting conditions and often a poor quality of the characters. A part of the network is executed on a processor board containing an analog neural net chip (Graf et al. 1991), while the rest is implemented as a software model on a workstation or a digital signal processor.
The authors describe a board system that integrates an analog neural net chip with a digital signal processor and fast memory. This system is in use as a coprocessor of a workstation where it accelerates computationally-intensive tasks for machine vision. A software environment has been developed to support image processing and testing of the system. The system was used to develop an application where the neural net determines the position and size of characters in complex images. For this task an increase in speed of a factor over 1000 over a workstation was achieved.< >
We present a semiclassical relativistic model for the orbital spectra of mesons, based on the assumption of a universal, flavor-independent linear confining interaction. Flavor dependence of the spectra arises from the quark masses.
We show that quantization of a nonrelativistic nonlinear wave equation is equivalent to the set of N-particle Schrödinger equations for all positive N. We compare the qualitative features of the quantized and unquantized field theory in a particular case, the cubic Schrödinger equation in one spatial dimension.We comment on the features of more general quantum field theories of interest in physics, and their possible relations to the properties of solutions of the corresponding classical field equations.During the past three years there has been a growing interest among physicists in the quantization of soluble classical field theories (equivalent generally to systems of coupled nonlinear partial differential equations) [1], [4], [5].It is the purpose of this review to sketch the quantization procedure as applied to relatively simple classical field theories, and to demonstrate that the resulting quantum field theories can be interpreted as describing an interesting physical system.Although the quantum field theory associated with a given classical field theory does not generally describe the same system as the classical theory, nonetheless solutions to the equations of motion of the classical theory give approximate information about physical observables in the quantized system.
We construct an infinite class of exact, time dependent solutions of the classical Yang-Mills equations for the group O(4) ≃ SU(2) × SU(2), depending upon a continuous parameter kϵ [0, 1]. In Euclidean space, our solutions interpolate continuously between the previously known solutions. For k = 0 they correspond to the solution of De Alfaro, Fubini and Furlan; for k = 1 they correspond to the pseudoparticle of Belavin, Polyakov, Schwartz and Tyupkin. In Minkowski space, for 0 ⩽ k < kc ≃ 0.173, they are everywhere regular and possess finite energy and action.