Most facsimile machines are manually operated, and accept only documents printed or handwritten on paper. Few of them are interfaced to computers. Currently there are a number of low–cost facsimile interface adapters designed for PCs which not only provide a link between a PC and a facsimile machine, but also perform most of the traditional facsimile functions. These adapters are able to transmit and receive messages as any ordinary facsimile machine does. They also have the added advantage of being able to send ASCII files from a PC directly and to display the received images on a VDU. They are thus more economic and easier to operate. Their only disadvantage is their inability to accept graphic inputs. This can be overcome by installing a hand scanner to act as a graphic input device. In this paper, we shall discuss the performance of two facsimile interface systems, namely the CWS–186F and the MFAX75S. The possibility of an efficient OCR system based on noise–free facsimile transmission is also discussed.
We present a front-end feature processor for distributed speech recognition for an integer-based DSP, and we employ block floating point and range reduction for the computation of elementary functions. We show that by reducing the numerical accuracy of the block floating point and the elementary functions, we are able to reduce the operational requirements to 12.6 wMOPs (weight million operations per second), 2.4 kWords of RAM, 3.7 kWords of ROM. When used on a small vocabulary of 800 words, 6.4 perplexity, and a large vocabulary of 20,200 words, 102.5 perplexity, our optimized DSP front-end produces recognition accuracy comparable to an equivalent implementation on a floating point processor, without requiring to retrain the recognition system with features produced by our DSP front-end.
While n-gram modeling is simple and dominant in speech recognition, it can only capture the short-distance context dependency within an n-word window where currently the largest practical n for natural language is three. However, many of the context dependencies in natural language occur beyond a three-word window. This paper proposes a new language modeling approach to capture the preferred relationships between words over a short or long distance through the concept of MI-Trigger pairs. Different MI-Trigger-based models are constructed in either a distance-dependent or a distance-independent way within a window from 1 to 10 words. This new MI-Trigger-based modeling is also compared and merged with word bigram modeling. It is found that the MI-Trigger-based modeling has better performance than word bigram modeling. It is also found that n-gram and MI-Trigger models have good complementarity and their proper merging can further increase the recognition rate when tested on Mandarin speech recognition. One advantage of MI-Trigger-based modeling is that the number of parameters needed for MI-Trigger modeling is much less than that of word bigram modeling. Another advantage is that the number of trigger pairs in an MI-Trigger model can be kept to a reasonable size without losing too much of its modeling power.
This paper proposes that the process of language understanding can be modeled as a collective phenomenon that emerges from a myriad of microscopic and diverse activities. The process is analogous to the crystallization process in chemistry. The essential features of this model are: asynchronous parallelism temperature-controlled randomness; and statistically emergent active symbols. A computer program that tests this model on the task of capturing the effect of context on the perception of ambiguous word boundaries in Chinese sentences is presented. The program adopts a holistic approach in which word identification forms an integral component of sentence analysis. Various types of knowledge, from statistics to linguistics, are seamlessly integrated for the tasks of word boundary disambiguation as well as sentential analysis. Our experimental results showed that the model is able to address the word boundary ambiguity problems effectively.
We report our development of a simple but fast and efficient inductive unsupervised semantic tagger for Chinese words. A POS hand-tagged corpus of 348,000 words is used. The corpus is being tagged in two steps. First, possible semantic tags are selected from a semantic dictionary(Tong Yi Ci Ci Lin), the POS and the conditional probability of semantic from POS, i.e., P(S|P). The final semantic tag is then assigned by considering the semantic tags before and after the current word and the semantic-word conditional probability P(S|W) derived from the first step. Semantic bigram probabilities P(S|S) are used in the second step. Final manual checking shows that this simple but efficient algorithm has a hit rate of 91%. The tagger tags 142 words per second, using a 120 MHz Pentium running FOXPRO. It runs about 2.3 times faster than a Viterbi tagger.
The infrared absorption spectrum of the ν2 band of deuterated nitric acid (DNO3) has been measured on a Bomem DA3.002 Fourier transform spectrometer in the wavenumber region 1640-1740 cm−1 with a resolution of 0.004 cm−1; 2036 transitions have been assigned in this hybrid type A and B band centered at 1688.3629 ± 0.0002 cm−1. The assigned transitions have been fitted to give nine rovibrational constants for the v2 = 1 state with a standard deviation of 0.00064 cm−1. The ratio of the transition moments, |μb/μa|, has been found to be 1.98 ± 0.10 for the band.
The high-resolution Fourier transform infrared spectrum of the ν6 and ν7 bands of deuterated nitric acid (DNO3) has been measured in the region between 510 and 667 cm−1 . The rovibrational transitions of both bands have been assigned and analyzed. Slightly improved ground state constants have been determined. The ν7 band is centered at 541.5847 ± 0.0002 cm−1. It shows no sign of perturbations and has been fit through J = 64 with an rms deviation of 0.0002 cm−1. The ν6 band is centered at 642.1383 ± 0.0002 cm−1 and is perturbed by 2ν9 through a Fermi resonance term and possibly a Coriolis term. The resonance is particularly noticeable above J = 40, where a crossing of rotational levels of the two vibrational states leads to a ΔK ± 2 interaction. A hot band centered at 632.7081 ± 0.0005 cm−1 has also been analyzed and identified as due to (ν6 + ν9) − ν9.
The ability to see through noise and distortion to a pattern is vital to the task of character recognition. Artificial neural networks exhibit such a capability as they are able to generalize automatically once they are trained. An application of an artificial neural network model, the Adaptive Resonance Theory (ART), to Chinese character classification is described. The ART classifier is used to classify 3755 Chinese characters. Our experimental results indicate that the classifier is able to achieve a high classification rate.
The Fourier transform infrared (FTIR) spectrum of the ν5 and 2ν9 bands of nitric acid (HNO3) has been measured with a resolution of 0.004 cm−1 in the frequency range of 856–910 cm−1. By fitting a total of 744 unperturbed transitions of ν5 with a standard deviation (SD) of 0.00051 cm−1, using a Watson Hamiltonian, a set of nine improved rovibrational constants for the upper state was obtained. Some transitions of 2ν9 were perturbed by a Femi resonance with the nearby ν5 state. A total of 280 unperturbed transitions of 2ν9 have been analyzed to provide nine rovibrational constants for the υ9 = 2 state with a SD of 0.00058 cm−1. The ν5 and 2ν9 bands are mainly A-type with band centres at 879.108200 ± 0.000045 and 896.44881 ± 0.00011 cm−1, respectively. A comparison of calculated line intensities of ν5 with measured values gives the transition moment as 0.155 ± 0.031 D and the integrated band intensity as 282 ± 56 cm−2 atm−1 (or 68 ± 14 km mol−1). The transition moment of the 2ν9 band is therefore estimated to be 0 17 ± 0.03 D which gives an integrated band intensity of 356 ± 71 cm−2 atm−1 (or 86 ± 17 km mol−1).
The infrared spectrum of the ν8 band of deuterated nitric acid (DNO3) has been measured with a resolution of 0.002 cm−1 in the frequency range of 730 to 794 cm−1. A total of 1888 assigned transitions have been analyzed to provide rovibrational constants for the upper state with a standard deviation of 0.00022 cm−1. In the course of the analysis, the constants for the ground state were improved. Because of the absence of perturbations, the constants can be used to accurately calculate the infrared line positions for the band. The band is C-type with a band center at 762.8738 ± 0.0001 cm−1. As was found for ν8 of HNO3, the P branch is much weaker than the R branch. Effective Herman-Wallis correction factors are given for the band.
We have measured the FT spectrum of natural OCS from 4800 to 8000 cm−1with a near Doppler resolution and a line-position accuracy between 2 and 8 × 10−4cm−1. For the normal isotopic species16O12C32S, 37 vibrational transitions have been analyzed for both frequencies and intensities. We also report six bands of16O12C34S, five bands of16O13C32S, two bands of16O12C33S, and two bands of18O12C32S. Important effective Herman–Wallis terms are explained by the anharmonic resonances between closely spaced states. As those results complete the study of the Fourier transform spectra of natural carbonyl sulfide from 1800 to 8000 cm−1, a new global rovibrational analysis of16O12C32S has been performed. We have determined a set of 148 molecular parameters, and a statistical agreement is obtained with all the available experimental data.
The high-resolution Fourier transform infrared spectrum of deuterated nitric acid (DNO3) has been measured in the ν9 region between 320 and 384 cm−1 with a resolution of 0.0035 cm−1. As expected for an out-of-plane vibrational fundamental, and accompanying hot bands, only C-type transitions were observed. Using a Watson Hamiltonian, 1772 infrared transitions have been assigned and fitted to give 9 rovibrational constants for the v9 = 1 state. With the knowledge of these constants, an analysis was made of 598 assigned transitions for the hot band 2ν9-ν9 in order to provide rovibrational constants for the v9 = 2 state. Both bands show no evidence of perturbation, although the v9 = 2 state is expected to be crossed by the v6 = 1 state above J = 40. The band centers are at 343.8496 ± 0.0001 cm−1 for ν9 and 333.7331 ± 0.0001 cm−1 for 2ν9-ν9.
The perception of a letter in the context of a word is easier than in the context of a random letter sequence. It appears that our knowledge about words can influence our perception process. McClelland and Rumelhart (1981) propose an interactive activation model to account for the interaction between our knowledge about words and our visual input. They use their model to explain how these interactions facilitate perception. In their account, word context effect is a constant independent of the identity of the words. In this paper, we propose the use of informatin theory to quantify word context effect. In this way, the strength of word context effect will depend on the identity of the words. We apply the method to quantify word context effect in Chinese words. This knowledge is encoded in an artificial neural network using the interactive activation and competition model. The network is used to recognize Chinese characters and we are able to achieve a high recognition rate.
This paper describes a new stroke and feature point extraction method for the recognition of printed Chinese characters of multiple fonts and various sizes. Based on our experimental study using 8 sets of Chinese character fonts, each comprising 3755 Chinese characters, this method is shown to be more stable compared to traditional methods.
The Geometric Arithmetic Parallel Processor (GAPP) is a systolic planar array processor with 72 (6 × 12) processing elements arranged in a mesh network. It runs at a clock speed of 10 MHz. The greatest advantage of such a design is that as many processor chips as are needed can be implemented without the loss of processing speed as interprocessor communication has been designed into the hardware architecture. It is therefore particularly suitable for processing of large image files with SIMD type of program execution. In this paper we report our study of the GAPP's performance in comparison with a number of other array processor systems, including: CLIP4, ICL DAP and Goodyear MPP. We also compare GAPP with some DSPs such as TMS32020, INMOS Transputer. We have found that GAPP's performance is in the same order of magnitude as these processors.