Colour blindness is a genetic mutation that alters the colour vision of the subjects by decreasing the sensitivity to certain colour wavelengths, depending on the defect. There are many forms of colour blindness ranging from monochromacy (black-white) to the most common form, the "red-green" variation where reds or greens are weakened, the vibrant shades are easily seen and the dull shades are difficult to perceive. A filter was designed based on the Ishihara colour tests in order to correct the colour blind deficiencies. This was successful for seeing the hidden objects within the test plates but did not translate well for real world images. The filter was modified, removing the dullest/lightest shades and shifting all the shades to the darker vibrant shades. The original image was shown to colour blind and normal vision subjects with results varying among all the subjects. After the modified filter was applied to a natural image, the colour blind and normal vision subjects were all able to correctly identify the test colours.
Excessive background noise is one of the most common complaints from hearing aid users. Background noise classification systems can be used in hearing aids to adjust the response based on the noise environment. This paper examines and compares two promising classification techniques, non-windowed artificial neural networks (ANN) and hidden Markov models (HMM), with an artificial neural network using windowed input. Results obtained show that an ANN with a windowed input gives an accuracy of up to 97.9%, which is more accurate than both the non-windowed ANN and the HMM. Overall, a windowed ANN is able to give excellent accuracy and reliability and is considered to be a good model for background noise classification in hearing aids
Neural networks can be of benefit in many image compression schemes. However, any system is constrained by the performance of the paradigm on which it is based. For example, although neural networks have been shown to improve differential pulse code modulation (DPCM) image compression, the overall performance of the system is still limited by the performance of DPCM. In this work a multiresolution neural network(MRNN) filter bank has been created for use within a state-of-the-art subband-coding framework. A polyphase implementation and training algorithm is presented. A filter bank that can synthesize the signal accurately from only the reference coefficients will be well suited for low-bitrate coding where the detail coefficients are coarsely quantized. Thus, the low-pass channel of the MRNN filter bank is trained to recreate the signal accurately The high-pass channel is trained for perfect reconstruction so that the MRNN filter bank will also be effective at high bitrates. This paper presents an analysis of the MRNN filter bank and its potential as a transform for coding. The MRNN filter bank has been used in place of a linear filter bank in the set partitioning in hierarchical trees (SPIHT) coder. The new filter bank shows advantages over the linear filter bank for coding at low bitrates, although its performance suffers at high bitrates. However, the results are encouraging and suggest that further work in this area warranted.
Storyboarding is a standard method for visual summarisation of shots in film and video preproduction. Reverse storyboarding is the generation of similar visualisations from existing footage. The key attributes of preproduction storyboards are identified, then computational techniques that extract corresponding features from video, render them appropriately, and composite them into a single storyboard image are developed. The result succinctly represents background composition, foreground object appearance and motion, and camera motion. For a variety of shots, it is shown that the visual representation conveys all the essential elements of shot composition.
We have developed a new method for speech decomposition and modeling. The purpose of this approach is to obtain better performance for modeling speech signals corrupted by non-stationary noise. The signal is first divided into frames and then each frame is decomposed into chirp-like partials which are linearly modulated in both amplitude and frequency. The proposed approach, based on the complex ambiguity function (CAF), can successfully estimate the parameters of each partial without assuming the harmonic structure of the signal. The effectiveness of this approach shows its potential for use in a speech enhancement system.
We analyze a number of corner detection algorithms and identify the advantages and disadvantages of each algorithm to evaluate their suitability for hardware implementation. We implemented three popular corner detectors, Plessy, Wang-Brady, and SUSAN, in software and compared them on the basis of their stability, accuracy, speed and computational requirements. The Plessy algorithm was found to have good stability and accuracy, but suffered from a large computational cost. The SUSAN method required the least computational resources and would therefore be suitable for implementation on a simple FPGA platform. However, it did not perform well on real world images. The Wang-Brady method was found to have better stability than SUSAN but worse than the Plessy algorithm while having a lower computational cost than Plessy and a higher cost than that for SUSAN. Despite the higher computational requirements, we conclude that the Plessy algorithm, because of its significantly better performance, is the most appropriate algorithm for hardware implementation.
The performance evaluation of a new approach to efficient disparity map computation for stereo vision is presented. Using Bayesian particle filtering to focus computational expenditure on image regions of primary importance to a navigational task allows the robust computation of a depth map sufficiently dense for most navigational tasks, while minimizing the computational load, and thus the resources, required.We show that a particleguided approach allows the efficient construction of a sparse disparity map which intelligently samples navigationally relevant scene information allowing real-time stereo computation on commodity hardware suitable for navigational tasks.
We propose a texture segmentation method based on frequency characteristics in a hybrid neural network approach using both unsupervised and supervised neural network classifiers. Our goal is to segment out tendons accurately and repeatedly from clinical ultrasound (US) images of horse tendons. The proposed method first extracts frequency-based texture features through the discrete cosine transform (DCT). A self-organizing-map (SOM) neural network is used for unsupervised classification. Following unsupervised training, a supervised neural network, learning vector quantization (LVQ), is used to improve further the performance and accuracy of segmentation. In terms of efficiency, only rotationally invariant features are adopted. The experimental results show that improvements can also be achieved by a feature selection scheme. The experimental images were all captured at a veterinary hospital. The results favourably compare to gold standards created by a radiologist.
Storyboarding is a standard method for visual summarization of shots in lm and video preproduction. Reverse storyboarding is the generation of similar visualizations from existing footage. We identify the key attributes of preproduction storyboards then develop computational techniques that extract corresponding features from video, render them appropriately, then composite them into a single storyboard image. The result succinctly represents background composition, foreground object appearance and motion, and camera motion. For tracking shots, we show that the visual representation conveys all the essential elements of shot composition. Visual summaries play an important role in the production and analysis of media. Practitioners, researchers and archivists all demand that the information presented is accurate and described in a consistent form using common metaphors derived from industry nomenclature. The goal is to enable quick access to details of specic shots or sequences without having to view the footage itself. In the media production industry, visual summarization is typically achieved through storyboards. Storyboards are drawn during preproduction then used throughout production and postproduction in tasks like set design, location lighting and image compositing. They provide for all participants a common reference to the ivisioni of the piece. Shorthand descriptions of all important visual components of each shot provide clear and accurate depictions of motion sequences in static form. These include specic methods of describing camera or subject movement through the use of various drawing techniques. While the term storyboard has been applied in the context of automated media analysis to a sequence of consecutive still images extracted from a lm or video programme, the representations traditionally used in the production industry are much richer. Our usage in this paper corresponds to lm production storyboarding: i.e., we seek to describe the temporal evolution of a shot through a single picture using rich visual cues. Storyboards incorporate the following types of information: 1. Composition of the shot, including start, end and notable intermediate camera positions
We consider optical flow as a means for determining the fundamental matrix for video assuming a small motion between frames. The mapping from one frame to a subsequent frame can be characterized by an eight parameter projective transformation or homography. We use optical flow to find this correspondence. Once found, this mapping can be used to establish the epipolar geometry, represented by the fundamental matrix. This is a basic tool in the analysis of scenes taken with two uncalibrated cameras or subsequent video frames in our case. We use SVD to restrict the rank to 2. The use of optical flow in this particular approach has not previously been well researched. However, we feel that it has a particular advantage when dealing with video because redundant information between frames can be more easily exploited. The results of this method were compared to those of other techniques for calculating the fundamental matrix and are found to be quite comparable.
This paper describes a novel matching pursuit based grouping approach for separating a speech signal from a mixture with non-Gaussian interference. At first, the mixture signal is decomposed into atoms by matching pursuit with a Gabor dictionary. Then a psychoacoustic based grouping algorithm is developed to cluster the atoms into groups to identify the atoms of a speech signal. These atoms are then used to reconstruct the desired speech signal. Simulations were performed on speech corrupted by factory noise and music. Preliminary results show that the proposed approach can remove almost all non-speech signal while the recovered speech signal possesses acceptable intelligibility.
The partitioned iterated function systems (PIFS) fractal image compression technique provides very competitive rate-distortion curves and fast decoding. However, it suffers from complicated encoding computation. Three novel neural network techniques, mixture of nonlinear principal components (MNLPC), mixture of independent components (MIC) and high-dimensional mixture of principal components (H-MPC) are developed to reduce the encoding complexity of the PIFS fractal coding. Applying these new techniques, the potential best range-domain matching search is confined to a relatively small size domain block pool. Using the new techniques, the encoding time is shortened dramatically, and the compression performance is improved as well.
Anew color image segmentation algorithm is presented in this paper. This algorithm is invariant to highlights and shading. This is accomplished in two steps. First, the average pixel intensity is removed form each RGB coordinate. This transformation mitigates the effects of highlights. Next, the Mixture of Principal Components algorithm is used to perform the segmentation. The MPC is implicitly invariant to shading due to the inner vector product or vector angle being used as similarity measure. Since the new coordinate system contains negative numbers, it is necessary to modify the MPC algorithm since in its original form it does not distinguish between positive and negative color space coordinates. Results on artificial and real images illustrate the effectiveness of the method. Finally, the use of the total within-cluster variance is investigated as possible criterion for selecting the number of clusters for the new algorithm.
A comprehensive color edge comparison across a variety color spaces is presented. Color adaptations of several well known edge detectors are compared in RGB, XYZ, CIELAB, CIELUV, rgb, l/sub 1/l/sub 2/l/sub 3/ and h/sub 1/h/sub 2/h/sub 3/-a new color space. The edge detectors studied are the Sobel operator, the modified Roberts operator, the vector gradient operator, and the 3/spl times/3 difference vector operator. Pratt's (1991) figure of merit is used for quantitative evaluation. The results indicate that the Sobel operator has the best performance across the edge detectors being compared and that results appear best for the h/sub 1/h/sub 2/h/sub 3/ color space.
In the past few years, researchers have been increasingly interested in color image segmentation. We analyze two different global image segmentation algorithms each using its own distance metric: k-means and a mixture of principal components (MPC) neural network. The k-means uses Euclidean distance for color comparisons while the MPC neural network uses vector angles. Two variants of the algorithms are examined. The first uses the RGB pixel itself for clustering while the second uses a 3/spl times/3 neighborhood. Preliminary results on a staged scene image are shown and discussed.
This paper introduces a new edge detection approach for color images. The method is based on the calculation of the vector angle between two adjacent pixels. Unlike Euclidean distance in RGB space, the vector angle distinguishes differences in chromaticity, independent of luminance or intensity. It is particularly well suited to applications where differences in illumination are irrelevant. Both metrics were implemented as modified Roberts edge operators to determine their effectiveness on an artificial image. The Euclidean method found edges across both luminance and chromatic boundaries whereas the vector angle method detected only chromatic differences.
A number of novel adaptive image compression methods have been developed using a new approach to data representation, a mixture of principal components (MPC). MPC, together with principal component analysis and vector quantization, form a spectrum of representations. The MPC network partitions the space into a number of regions or subspaces. Within each subspace the data are represented by the M principal components of the subspace. While Hebbian learning has been effectively used to extract principal components for the MPC, its stability is still a concern in practice. As a result, computationally more expensive methods such as batch eigendecomposition have produced more consistent results. This paper compares the performance of a number of Hebbian- based training schemes for the MPC network. These include training the entire network, network growing techniques, and a new tree-structured method. In the new tree-structured approach, each level in the tree, M, corresponds to an M- dimensional representation. A node and all its M - 1 parents represents a single M-dimensional subspace or class. The evaluation shows that the use of tree-structured approach improves training and results in reduced squared error.
This paper examines the relationship between iterated function systems (IFS), a fractal approach to image compression, and the mixture of principal components (MPC), a neural network approach to image compression. Both can be fundamentally expressed as a local linear transformation. In IFS, the basis vector comes from the image itself and evolves during the iterations while in MPC, the trained network contains the basis vectors. A new method of image compression is presented which uses an MPC network as a library for reducing the search for the large domain blocks in IFS. The resulting hybrid approach has better rate-distortion characteristics relative to standard IFS when tested on a standard image.
Shawki M. Areibi合作论文数School of Engineering, University of Guelph1